Short video generation cross-shot consistency guarantee and multi-modal material management method
Patent Information
- Application Number
- CN202611327581.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
在实际生成过程中,创意内容、分镜描述或局部画面经过调整后,相关镜头通常需要重新进行素材匹配和内容生成处理,由于镜头之间缺少有效关联机制,容易出现部分镜头完成更新而其他关联镜头未同步调整的问题,进而影响视频整体的视觉连续性
1、通过将创意简报与分镜序列中的角色、场景等信息转化为一致性锚点,并结合初始锚定图的动态质量评估、候选参考图优先级计算以及强制注入机制,使不同镜头在生成过程中能够持续获得与当前镜头语义相匹配的参考信息;同时结合三道校验与缺失补全机制,对生成计划中的参考信息进行逐级检查,从而降低连续镜头之间因参考信息缺失或选择偏差造成的视觉偏移,提高多镜头视觉内容的连续性。
Smart Images

Figure CN122845901A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of short video content generation technology, specifically to a method for ensuring cross-camera consistency and managing multimodal materials in short video generation. Background Technology
[0002] With the development of short video generation technology, the use of automated methods to generate continuous shot content has been gradually applied in video creation scenarios. In the actual generation process, after the creative content, storyboard description, or partial frame is adjusted, the relevant shots usually need to be reprocessed with material matching and content generation. Due to the lack of an effective correlation mechanism between shots, it is easy for some shots to be updated while other related shots are not adjusted synchronously, thus affecting the overall visual continuity of the video.
[0003] On the other hand, with the increase in the number of short video projects and the number of times videos are generated, a large amount of similar video content and repetitive visual materials continue to accumulate. This not only increases the pressure on material storage but also makes it difficult to have a unified judgment standard for reusing existing materials between different projects. At the same time, the changes in video materials during multiple generation, replacement, and optimization processes are difficult to record completely. When the generation results are abnormal or historical content needs to be restored, it is difficult to accurately locate the material status at the corresponding stage, which increases the maintenance difficulty of the short video generation system and the complexity of material management.
[0004] Therefore, there is an urgent need for a unified management method for shot association, material reuse, and version changes in the short video generation process, in order to improve cross-shot consistency and multimodal material management capabilities during video generation. Summary of the Invention
[0005] This invention aims to provide a method for ensuring consistency across shots and managing multimodal materials in short video generation. By establishing stable relationships between shots and managing the generation, updating, reuse, and lifecycle of materials in a unified manner, it improves the consistency control and material management capabilities in the continuous shot generation process.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for ensuring cross-shot consistency and managing multimodal materials in short video generation includes: Video content consistency anchors are extracted from short video creation scripts, shot planning information and shot sequences. Initial visual reference frames are bound to the consistency anchors, and video image feature parameters corresponding to the initial visual reference frames are extracted as cross-shot comparison benchmarks. Based on the video image feature parameters, a matching score is calculated between the candidate visual reference material and the current video shot, and the materials are sorted according to the matching score. The initial visual reference frame is then added as the highest priority material to the video reference material set corresponding to the current video shot. A three-level verification is performed on the video reference material set, and video visual materials corresponding to each video shot are generated based on the verified video reference material set; when the video reference material corresponding to the initial visual reference frame is replaced, the associated video shot in the corresponding storyboard is marked as a shot to be regenerated, and the corresponding video visual material is regenerated. The generated video footage is similar to existing video footage. When the deduplication threshold is reached, a reference relationship is established. A video footage association graph is constructed based on the footage generation record, and source tracing is performed.
[0007] Preferably, the step of extracting video content consistency anchor points and binding them to initial visual reference frames includes: The text description information in the short video creation script and the shot description information in the storyboard sequence are subjected to video content semantic analysis to extract video character objects, video scene objects, and video visual representation objects. The storyboard sequence is traversed, and the video visual material with the highest image quality evaluation parameter in the video shot in which the video character object or video scene object first appears is bound as the corresponding initial visual reference frame. The similarity of video image features between the video visual material generated in subsequent video shots and the corresponding initial visual reference frame is calculated. When the similarity of video image features is lower than a preset consistency threshold, the storyboard sequence is traversed backward to find alternative video shots in which the video character object or video scene object reappears and has a higher image quality evaluation parameter. The corresponding video visual material is then updated as a new initial visual reference frame until the similarity of video image features meets the preset consistency threshold.
[0008] Preferably, the calculation of the matching score for the candidate visual reference material includes: A shot type recognition module is used to predict the semantic type of a video shot based on the text description information and shot motion information corresponding to the current video shot. The semantic types of the video shot include close-up shots of people, medium shots of people, long shots of the environment, and dynamic tracking shots. The weight parameters corresponding to character object features, scene environment features, visual style features, motion trajectory features, and video composition features are adaptively adjusted according to the semantic type of the video shot. Specifically, when the semantic type of the video shot is a close-up shot of a person, the weight parameters corresponding to the character object features are increased; when the semantic type of the video shot is a dynamic tracking shot, the weight parameters corresponding to the video composition features and motion trajectory features are increased. The consistency of candidate visual reference materials in terms of character object consistency, scene environment consistency, visual style consistency, and video composition are calculated respectively.Figure 1 The matching degree in the consistency dimension is calculated, and the credibility coefficient of the material is determined according to the source type of the candidate visual reference material. Based on the matching degree of each dimension, the credibility coefficient of the material, and the adaptively adjusted weight parameters, a weighted calculation is performed to obtain the video matching score of the candidate visual reference material.
[0009] Preferably, adding the initial visual reference frame to the video reference material set includes: A forced association process, independent of video matching scores, is adopted to traverse the video character objects and video scene objects involved in the current video shot. When a corresponding initial visual reference frame exists in the video content consistency anchor point, the initial visual reference frame is added to the current video reference material set as the highest priority video consistency constraint material. The remaining candidate visual reference materials are sorted in descending order according to the video matching scores, and candidate visual reference materials with video matching scores greater than a preset reference threshold and whose video reference material set has not reached a preset maximum number of references are added to the video reference material set.
[0010] Preferably, the three-level verification and missing data completion of the video reference material set includes: The first layer of verification is the integrity check of the generated parameters. It checks whether the video shot generation configuration contains video reference frame information. If it is missing, it matches the corresponding initial visual reference frame based on the video image features and fills it in. The second layer of verification is the cross-shot consistency check. It checks whether the video reference material contains the corresponding consistency anchor point. If it is missing, it fills in the corresponding initial visual reference frame. The third layer of verification is the material conflict check. When the initial visual reference frame conflicts with other reference materials, it selects the optimal combination of reference materials based on the material compatibility score. At the same time, it counts the verification triggering situation and adjusts the video reference material selection strategy for the corresponding shot based on the abnormal results.
[0011] Preferably, the similarity calculation between the generated video visual material and existing video material includes: The generated video visual material and existing video material are sampled along the timeline to obtain multiple video keyframes. Image texture features and discrete cosine transform low-frequency coefficients are extracted from the video keyframes to generate a video frame feature sequence. Video segments of different lengths are aligned using a time window, and the distance between the video frame feature sequences is calculated to obtain a video-level visual content similarity score. When the video-level visual content similarity score is not less than a preset deduplication threshold, a video material reference record is established; otherwise, the corresponding video visual material is marked as material to be verified.
[0012] Preferably, the step of constructing a relationship graph of video materials and performing source tracing includes: Establish a versioned video material relationship graph and record the modification, replacement, and regeneration information of video materials; when the initial visual reference frame is replaced and triggers the regeneration of video shots, create a corresponding new version node in the video material relationship graph and retain the original version material relationship path; when a material tracking request is received, traverse the video material relationship graph in a forward and / or reverse direction according to the specified time node, and output the corresponding video material generation, replacement, and regeneration link.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By transforming information such as characters and scenes in the creative brief and storyboard sequence into consistent anchor points, and combining dynamic quality assessment of the initial anchor image, priority calculation of candidate reference images, and forced injection mechanism, different shots can continuously obtain reference information that matches the semantics of the current shot during the generation process. At the same time, by combining three-level verification and missing information completion mechanism, the reference information in the generation plan is checked step by step, thereby reducing the visual shift caused by missing or selection bias between consecutive shots and improving the continuity of visual content in multiple shots.
[0014] 2. By organizing the multi-frame features of video footage into a temporal hash sequence and combining time window alignment and edit distance for video-level similarity judgment, the identification of similar content is no longer limited to a single frame. Furthermore, by combining reference counting, cross-project reuse, and recycling condition judgment, duplicate footage can be managed through references, reducing the duplicate storage of identical or highly similar content and enabling footage to have a clearer reuse relationship between different projects.
[0015] 3. By coordinating the anchor map update trigger mechanism, the affected shot regeneration mechanism, the versioned material lineage directed acyclic graph, and the cross-project reuse mechanism, the replacement of materials can synchronously affect the generation process of the corresponding shots, and the old and new versions and their generation relationships can be continuously preserved. At the same time, the temporal similarity judgment results can provide a basis for material reference and reuse, so that "generation consistency control, version evolution record and material lifecycle management" form an interconnected processing link, improving the coordination of material status maintenance and historical traceability in complex short video projects. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of the short video generation method for ensuring consistency across shots and managing multimodal materials according to the present invention. Figure 2 This is a schematic diagram illustrating the priority calculation and three-stage verification process of the present invention. Figure 3 This is a schematic diagram illustrating the process of deduplication of temporal features and version lineage tracing of video materials in this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figures 1 to 3 This invention provides a method for ensuring cross-shot consistency and managing multimodal materials in short video generation, referring to... Figure 1 Step-by-step flowchart Figure 2 The reference diagram shows the flowchart of priority calculation and three-stage verification, and Figure 3 The flowchart of the present invention for deduplication of video material based on temporal features and version lineage tracing is shown below; the technical solution is as follows: Video content consistency anchors are extracted from short video creation scripts, shot planning information and shot sequences. Initial visual reference frames are bound to the consistency anchors, and video image feature parameters corresponding to the initial visual reference frames are extracted as cross-shot comparison benchmarks. Based on the video image feature parameters, a matching score is calculated between the candidate visual reference material and the current video shot, and the materials are sorted according to the matching score. The initial visual reference frame is then added as the highest priority material to the video reference material set corresponding to the current video shot. A three-level verification is performed on the video reference material set, and video visual materials corresponding to each video shot are generated based on the verified video reference material set; when the video reference material corresponding to the initial visual reference frame is replaced, the associated video shot in the corresponding storyboard is marked as a shot to be regenerated, and the corresponding video visual material is regenerated. The generated video footage is similar to existing video footage. When the deduplication threshold is reached, a reference relationship is established. A video footage association graph is constructed based on the footage generation record, and source tracing is performed.
[0019] Example 1: First, the step of extracting video content consistency anchor points and binding initial visual reference frames includes: performing semantic parsing of the text description information in the short video creation script and the shot description information in the storyboard sequence to extract video character objects, video scene objects, and video visual representation objects; traversing the storyboard sequence and binding the video visual material with the highest image quality evaluation parameter in the video shot where the video character object or video scene object first appears as the corresponding initial visual reference frame; calculating the video image feature similarity between the video visual material generated by subsequent video shots and the corresponding initial visual reference frame; when the video image feature similarity is lower than a preset consistency threshold, traversing backward along the storyboard sequence to find alternative video shots where the video character object or video scene object reappears and has a higher image quality evaluation parameter, and updating the corresponding video visual material as a new initial visual reference frame, until the video image feature similarity meets the preset consistency threshold.
[0020] Specifically, the extraction of video content consistency anchor points is completed by a semantic parsing module, where the input data of the module is the text description information of the short video creation script and the shot description information corresponding to each shot in the shot sequence, and the input format is structured text fields, including script title, character lines, scene description and shot action description. The interior of the semantic parsing module consists of a named entity recognition submodule and an object classification submodule. The named entity recognition submodule adopts a sequence labeling structure to perform label prediction character by character on the input text. The label categories include three types: character name, scene location and visual performance description. The recognition result is output with entity text fragments and their start and end positions in the original text; wherein, for entities with clear boundaries, namely character names and scene locations, the named entity recognition submodule directly adopts continuous character fragments corresponding to the character-by-character label prediction results as entity boundaries; for entities of the visual performance description category, the named entity recognition submodule uses punctuation marks such as commas and periods, and structural auxiliary words such as "de", "di" and "zhe" as the basis for segmentation, pre-segments the input text into a plurality of candidate phrase fragments, then performs label prediction on the whole of each candidate phrase fragment, and takes the candidate phrase fragments predicted to have the visual performance description label and whose prediction confidence is not lower than a preset confidence threshold (set to 0.8 in this embodiment, and can be adjusted according to actual conditions) as the boundary of this type of entity, so as to avoid boundary intersection or nesting conflict between visual performance description entities and character name and scene location entities caused by the character-by-character labeling method; the object classification submodule receives the recognition result, and merges the entities into a video character object list, a video scene object list and a video visual performance object list according to entity categories. The merging is based on the character overlap between entity text fragments and the matching result of a synonym mapping table. For example, "protagonist" appearing in the script and "the character" appearing in the subsequent shot description are merged into the same video character object after synonym mapping. The above three types of object lists are output in the form of structured records, and each record includes the object name, the object category and the number of the first shot where the object appears in the shot sequence, for reading in the subsequent binding step.
[0021] Furthermore, after obtaining the shot number of the first appearance of the video character and video scene objects, the system traverses the storyboard sequence and reads the set of candidate video visual materials generated during the generation phase for that first appearance shot. Each candidate video visual material carries an image quality evaluation parameter, which is obtained by weighting the image sharpness index, subject integrity index, and image noise index. These three sub-indicators are calculated synchronously by the image quality evaluation module when generating the visual material and are written into the material library along with the material. The system sorts the set of candidate video visual materials corresponding to the shot according to the image quality evaluation parameter, selects the one with the highest parameter value as the initial visual reference frame for the corresponding video character or video scene object, and writes the storage identifier, shot number, and entity name of the material into the consistency anchor record table to complete the binding process. For example, if the first appearance shot number of a video character in the storyboard sequence is shot 2, there are 3 candidate video visual materials corresponding to this shot, with image quality evaluation parameters of, for example, 82 points, 91 points, and 76 points. The system selects the material with a parameter of 91 points and binds it as the initial visual reference frame for the video character.
[0022] Specifically, after binding, the system performs dynamic quality comparison on each subsequent shot of the video character or scene object in the storyboard sequence. The input to the comparison module is the video visual material generated from subsequent shots and the bound initial visual reference frame. Both are first converted into video image feature parameters by the feature extraction submodule. These feature parameters cover three dimensions: color distribution, contour shape, and texture structure. The output is a fixed-length numerical sequence. The comparison module calculates the video image feature similarity between the subsequent video visual material and the initial visual reference frame based on the above numerical sequence, and compares this similarity value with a preset consistency threshold. The preset consistency threshold is set differently for video character objects and video scene objects. For example, the threshold for character anchor points is set to 0.75, and the threshold for scene anchor points is set to 0.7. A test set of video frames containing a large number of "consistency" and "offset (drift)" annotations is pre-constructed to calculate the true similarity distribution of characters and scenes in each feature dimension. By plotting evaluation curves (such as ROC curves), the optimal cutoff point is found that minimizes the combined costs of "false positive rate (judging different as the same)" and "false negative rate (judging the same as different)". When the calculated video image feature similarity is lower than the corresponding preset consistency threshold, it is determined that there is an appearance offset between the video visual material generated by the subsequent shot and the initial visual reference frame, triggering the subsequent reference frame update process.
[0023] Furthermore, after triggering the appearance offset determination, the system continues to traverse the storyboard sequence from the current shot backward, reading the candidate video visual materials and their image quality evaluation parameters associated with each shot in which the video character or scene object reappears. The system compares the image quality evaluation parameters of each candidate material with those of the current initial visual reference frame, selecting the candidate video shots with higher image quality evaluation parameters. The system updates the video visual material corresponding to the candidate video shot as a new initial visual reference frame, writes it to the consistency anchor record table and overwrites the original binding record, while retaining the storage identifier of the original initial visual reference frame as a historical version. After the update, the system re-executes the aforementioned video image feature similarity calculation on the updated initial visual reference frame. If the similarity is still lower than the preset consistency threshold, the system continues to traverse the storyboard sequence backward to find the next candidate video shot until the found video visual material meets the preset consistency threshold with the comparison benchmark before the update, or until the end of the storyboard sequence is reached. For example, after the offset determination is triggered in the 5th shot, the system continues to search backward to the 9th shot. The image quality evaluation parameter of the candidate material corresponding to this shot is, for example, 88 points, which is higher than the current evaluation value of the original initial visual reference frame after multiple reuses of the 91-point version. The system updates the material corresponding to the 9th shot as the new initial visual reference frame.
[0024] By automatically extracting consistency anchor points through semantic parsing and dynamically updating the initial visual reference frame in combination with image quality evaluation parameters, the reference frame maintains a superior comparison benchmark quality throughout the sequence of shots. This helps reduce the degree of appearance deviation of the same character or scene in cross-shot presentation and improves the continuity of visual content across multiple shots.
[0025] Further, the calculation of the matching score for candidate visual reference materials includes: using a shot type recognition module to predict the semantic type of the video shot based on the text description information and shot motion information corresponding to the current video shot; the semantic type of the video shot includes close-up shots of people, medium shots of people, long shots of the environment, and dynamic tracking shots; adaptively adjusting the weight parameters corresponding to character object features, scene environment features, visual style features, motion trajectory features, and video composition features according to the semantic type of the video shot; wherein, when the semantic type of the video shot is a close-up shot of people, the weight parameters corresponding to the character object features are increased; when the semantic type of the video shot is a dynamic tracking shot, the weight parameters corresponding to the video composition features and motion trajectory features are increased; calculating the consistency of candidate visual reference materials in terms of character object consistency, scene environment consistency, visual style consistency, and video composition, respectively. Figure 1The matching degree in the consistency dimension is calculated, and the credibility coefficient of the material is determined according to the source type of the candidate visual reference material. Based on the matching degree of each dimension, the credibility coefficient of the material, and the adaptively adjusted weight parameters, a weighted calculation is performed to obtain the video matching score of the candidate visual reference material.
[0026] Specifically, the input data for the shot type recognition module consists of text description information corresponding to the current video shot and shot motion information. The text description information is the scene content description recorded in the storyboard, and the shot motion information is the camera motion mode field marked in the storyboard entries, such as values like fixed, zoom in, and follow. The shot type recognition module internally consists of a text encoding layer and a motion feature concatenation layer. The text encoding layer uses a pre-trained BERT encoder to segment and vectorize the input text description information, taking the hidden state vector corresponding to the "CLS" position output by the encoder as the sentence-level semantic vector, with a vector dimension of 768. The motion feature concatenation layer pre-constructs a motion mode category embedding table containing values such as fixed, zoom in, and follow. After mapping the shot motion information into a 128-dimensional motion category vector, it concatenates it with the sentence-level semantic vector in the feature dimension to obtain an 896-dimensional concatenated vector, which is then input into a classification layer consisting of two fully connected layers. The first class layer uses the ReLU activation function, and the second class layer uses the Softmax activation function to output the corresponding probability values of the four categories. The motion feature concatenation layer maps the lens motion information into a fixed-dimensional motion category vector and concatenates it with the sentence-level semantic vector. The concatenated vector is input into the classification layer for category scoring. The classification layer outputs the corresponding probability values of the four categories: close-up shot of a person, medium shot of a person, long shot of an environment, and dynamic tracking shot. The category with the highest probability is taken as the semantic type determination result of the current video shot, and this determination result, along with the confidence probability, is written into the processing record of the current shot for subsequent weight adjustment.
[0027] Furthermore, the weight adjustment stage receives the video shot semantic type determination result output by the shot type recognition module and reads a set of initial weight parameters from a preset weight configuration table. This set of weight parameters covers four dimensions: character object features, scene environment features, visual style features, and video composition features. The weight adjustment stage has built-in adjustment rules: when the semantic type of the video shot is a close-up of a person, the weight parameters corresponding to the character object features are increased from their initial values, while the weight parameters corresponding to the scene environment features are decreased proportionally; when the semantic type of the video shot is a dynamic tracking shot, the weight parameters corresponding to the video composition features and motion trajectory features are increased, while the weight parameters corresponding to the character object features are decreased proportionally; for medium-shot shots of people and long-shot shots of the environment, the weight adjustment stage maintains a balanced set of weight parameter values without individually increasing them. The adjusted weight parameters are output as a numerical group and passed to the matching degree calculation stage. For example, if a shot is determined to be a close-up of a person, the weight parameter of the character object features is increased from the initial 0.3 to 0.5, and the weight parameter of the scene environment features is correspondingly decreased from 0.3 to 0.2.
[0028] Specifically, the input to the matching degree calculation stage is a set of candidate visual reference materials and consistent anchor point materials corresponding to the current video shot. Both are converted into four sets of vectors—character object features, scene environment features, visual style features, and video composition features—by the feature extraction submodule, and then compared dimension by dimension to obtain the matching degree of the candidate visual reference materials in terms of character object consistency, scene environment consistency, visual style consistency, and video composition consistency. Figure 1 The matching scores are calculated across four dimensions: consistency, reliability, and credibility. Simultaneously, the credibility coefficient of the candidate visual reference material is determined based on its source type, which includes three categories: initial visual reference frame source, historical material generated within the same storyboard, and material reused across projects. The initial visual reference frame source has the highest credibility coefficient, while the material reused across projects has a relatively lower value. The weighted calculation stage receives the matching scores across the four dimensions, the credibility coefficient of the material, and the weight parameters output from the weight adjustment stage. It multiplies the matching scores by their corresponding weight parameters for each dimension, sums the results, and then multiplies by the credibility coefficient of the material to obtain the video matching score for the candidate visual reference material. This score is output numerically and written into the candidate material ranking list.
[0029] By using lens semantic type recognition to drive the adaptive adjustment of the weight parameters of each feature dimension, and combining the material credibility coefficient for weighted scoring, the matching score of candidate visual reference materials can be aligned with the presentation emphasis of different lens types, which helps to improve the matching degree between the reference material selection results and the actual needs of the current lens.
[0030] Furthermore, adding the initial visual reference frame to the video reference material set includes: employing a forced association process independent of the video matching score to traverse the video character objects and video scene objects involved in the current video shot; when a corresponding initial visual reference frame exists in the video content consistency anchor point, the initial visual reference frame is added to the current video reference material set as the highest priority video consistency constraint material; the remaining candidate visual reference materials are sorted in descending order according to the video matching score, and candidate visual reference materials whose video matching score is greater than a preset reference threshold and whose video reference material set has not reached a preset maximum number of references are added to the video reference material set.
[0031] Specifically, the forced association process, as a parallel processing branch independent of the video matching scoring and sorting process, takes as input the list of video character objects and the list of video scene objects parsed from the shot description information corresponding to the current video shot, and outputs the corresponding set of initial visual reference frame identifiers. The forced association process reads each video character object and video scene object involved in the current shot one by one, and queries the consistency anchor record table using the object name as the search key. If an object is found to have a bound initial visual reference frame, the storage identifier of that initial visual reference frame is directly written to the beginning of the current video reference material set. This writing action does not go through the video matching scoring calculation process and is not affected by the video matching score value. When the current shot involves multiple objects with bound initial visual reference frames, the forced association process writes the corresponding initial visual reference frames sequentially to the beginning of the video reference material set, with the writing order determined by the order in which the objects appear in the shot description information.
[0032] Furthermore, after the forced association process completes the initial visual reference frame writing, the system reads the candidate visual reference material sorting list output by the aforementioned weighted calculation step. This sorting list has been sorted in descending order of video matching score from high to low. The system traverses and filters this sorting list, determining whether the video matching score of each candidate visual reference material is greater than a preset reference threshold, and simultaneously determining whether the number of existing materials in the current video reference material set has reached a preset maximum reference quantity. Only when the video matching score of a candidate visual reference material is greater than the preset reference threshold and the video reference material set has not reached the preset maximum reference quantity is the candidate visual reference material added to the video reference material set. If either of the two conditions is not met, the candidate material is skipped and the system continues processing the next item in the sorting list. After the filtering is completed, the video reference material set is output in the form of an ordered list. The first part is the initial visual reference frame that is forcibly written, and the second part is the candidate visual reference material added according to the score. For example, in this embodiment, the preset reference threshold value is 0.6 and the preset maximum number of references is 5. The specific values can be adjusted according to the actual situation. After forcibly writing 2 initial visual reference frames, the system continues to filter candidate materials with a score greater than 0.6 from the sorted list to add them until the total number of sets reaches 5 or the sorted list is traversed.
[0033] By setting up a mandatory association process independent of video matching scores, the initial visual reference frames are always included in the video reference material set without being constrained by the score ranking results. Combined with thresholds and quantity limits, the remaining candidate materials are supplemented in an orderly manner, which helps to reduce the situation where key consistency constraint materials are missed due to score ranking deviations.
[0034] Furthermore, the three-level verification and missing element completion for the video reference material set includes: the first level of verification is the generation parameter integrity detection, which checks whether the video shot generation configuration contains video reference frame information. If it is missing, the corresponding initial visual reference frame is matched and completed based on the video image features; the second level of verification is the cross-shot consistency detection, which checks whether the video reference material contains the corresponding consistency anchor point. If it is missing, the corresponding initial visual reference frame is added; the third level of verification is the material conflict detection, which selects the optimal reference material combination based on the material compatibility score when the initial visual reference frame conflicts with other reference materials; at the same time, the verification triggering situation is counted, and the video reference material selection strategy for the corresponding shot is adjusted according to the abnormal results.
[0035] Specifically, the input to the first-level verification is the video shot generation configuration. This configuration is a structured set of parameters formed during the visual generation process before material generation. It includes a video reference frame information field, which records the list of reference material identifiers actually used in this generation. The first-level verification module reads this configuration and checks whether the video reference frame information field is empty or missing. If a missing field is detected, the verification module uses the video image feature parameters corresponding to the current video shot as the query condition, matches the initial visual reference frame with the highest feature similarity in the consistency anchor record table, and writes the matched initial visual reference frame identifier back to the video reference frame information field, thus completing the configuration. The completed configuration is then passed to the next level of verification. If a complete field is detected, the configuration is passed directly to the next level without modification.
[0036] Furthermore, the second-layer verification receives the generation configuration output from the first-layer verification, reads the reference material identifier list recorded in the video reference frame information field, and compares this identifier list with the initial visual reference frame identifiers registered in the consistency anchor point record table corresponding to the current video shot item by item to check whether the reference material identifier list contains all the consistency anchor points involved in the current shot. If the comparison result shows that the initial visual reference frame corresponding to a certain video character object or video scene object is not included in the reference material identifier list, the second-layer verification module adds the missing initial visual reference frame identifier to the reference material identifier list and updates the video reference frame information field in the generation configuration; if the comparison result shows that all consistency anchor points are included, the generation configuration is not modified, and both outputs are passed to the third-layer verification stage, with the output form being the updated generation configuration.
[0037] Specifically, the third-layer verification receives the generation configuration output from the second-layer verification and performs content conflict detection on each pair of visual reference materials corresponding to the reference material identifier list. The conflict detection is based on the degree of difference between the materials in terms of character object features, scene environment features, and visual style features. When the initial visual reference frame and other reference materials in the reference material identifier list have a difference in the above dimensions that exceeds the preset conflict judgment condition, it is determined that there is a content conflict between the two. Upon detecting a conflict, the third-layer verification module enumerates and evaluates candidate material combinations in the reference material identifier list based on the material compatibility score. The compatibility score is derived from the mean of the pairwise differences between each material in the candidate combination. Specifically, the differences in the character object features, scene environment features, and visual style features are calculated using the L2 norm to determine the Euclidean distance between feature vectors, and then weighted and summed according to preset weights to obtain a comprehensive mean difference. Subsequently, a negative exponential mapping formula is used, where the compatibility score equals the natural exponent of the negative mean difference, to convert the comprehensive mean difference into a compatibility score between 0 and 1. This ensures that a lower mean difference results in a higher compatibility score and a smooth convergence trend, avoiding instability in the score results due to sudden changes in the mean. The module selects the candidate material combination with the highest compatibility score to replace the original reference material identifier list and outputs the final generated configuration after three layers of verification and completion for the visual generation stage to read and execute.
[0038] The preset conflict determination condition refers to the maximum permissible difference boundary used to measure whether there are mutually exclusive features among multiple visual reference materials. Specifically, it is calculated by separately calculating the distances (such as Euclidean distance) between the initial visual reference frame and other reference images in the character object, scene environment, and visual style feature vector spaces, and then summing them according to weights to obtain the comprehensive difference tolerance upper limit. When the comprehensive feature difference between materials exceeds this limit, it is defined as a content conflict, indicating that if they are used as input conditions at the same time, they will cause the visual generation model to receive contradictory guiding cues (such as severe stylistic discontinuity or contradictory subject space features), thus requiring the system to reselect a more harmonious combination of reference images based on compatibility scores.
[0039] Furthermore, a differentiated parameter tuning and feedback control closed-loop mechanism deeply bound to specific error codes is established to address various types of abnormal results that may occur during the three-level verification process. This mechanism relies on independent hierarchical anomaly capture and targeted feedback adjustment laws to replace the coarse strategy of globally uniformly modifying thresholds, thereby more accurately managing the candidate material library. When performing three-level verifications for the integrity of generated parameters, cross-shot consistency, and material conflicts, the system independently records the anomaly codes triggered by each level of interception, the semantic types of the associated shots, and the cumulative frequency. When the trigger frequency of a specific anomaly code exceeds the preset safety baseline within a certain statistical period, the feedback control module will automatically match and activate differentiated parameter adjustment rules. For example, if the consistency verification defense frequently reports errors, it indicates that the current candidate set lacks sufficient constraint samples. The system will dynamically expand the maximum reference quantity limit and reduce the credibility weight of historical generated materials to introduce new anchor points. Conversely, if the material conflict verification interception is high-frequency, it indicates that there is a serious mutual exclusion of feature dimensions among the selected reference frames. At this time, the system will specifically increase the preset reference threshold to strictly tighten the candidate range and enhance the truncation penalty coefficient in the subsequent compatibility score calculation. To prevent extreme oscillations in parameters during multiple feedback iterations, the system also configures absolute upper and lower bound protection conditions for various parameter tuning operations. All dynamic parameter update trajectories are written to the control log in real time, allowing the system to continuously optimize the initial feature allocation weights for similar shots. This closed-loop mechanism breaks the limitations of a single parameter tuning direction, enabling cross-shot generation configurations to perform refined self-correction based on multi-dimensional actual verification feedback. This not only significantly reduces the failure rate caused by missing materials or underlying conflicts in continuous generation tasks but also significantly enhances the system's robustness when handling complex scripts.
[0040] Furthermore, the system records the trigger count, the semantic type of the triggered shot, and the video shot number of each of the three layers of verification, forming a verification trigger log. This log is then aggregated according to a preset statistical period to obtain the verification trigger frequency corresponding to each video shot semantic type. When the aggregated verification trigger frequency for a certain video shot semantic type exceeds a preset threshold, the system determines that there is a systematic deviation in the reference material selection process for that type of shot. Based on this, the system adjusts the candidate visual reference material selection parameters corresponding to that shot semantic type. Adjustments include increasing the preset reference threshold for that type of shot or reducing the preset maximum number of references. The adjusted parameters are written back to the weight configuration table for use in the subsequent generation of video reference material sets for the same type of shot. For example, if the third-layer verification trigger frequency for a dynamic tracking shot reaches, for instance, 35% within a certain statistical period, exceeding the preset threshold of 25%, the system increases the preset reference threshold for that type of shot from 0.6 to 0.65.
[0041] By implementing a progressive three-layer check for integrity, consistency, and conflict, combined with trigger frequency statistics to adjust the filtering strategy, the generated configuration is checked and corrected from multiple angles before entering the visual generation stage. This helps reduce abnormal generation results caused by missing or conflicting reference information.
[0042] Furthermore, the similarity calculation between the generated video visual material and the existing video material includes: performing time-axis sampling on both the generated video visual material and the existing video material to obtain multiple video keyframes; extracting image texture features and discrete cosine transform low-frequency coefficients from the video keyframes to generate a video frame feature sequence; aligning video segments of different lengths through a time window and calculating the distance between the video frame feature sequences to obtain a video-level visual content similarity score; when the video-level visual content similarity score reaches a preset deduplication threshold, establishing a video material reference record; otherwise, marking the corresponding video visual material as material to be verified.
[0043] Specifically, the input to the timeline sampling stage is the complete duration information of the generated video visual material or existing video material, as well as the material data itself. The sampling module uniformly selects points within the total duration of the material according to a preset number. The video frame corresponding to each selected point is extracted as a video keyframe. For example, if a piece of material is 4 seconds long and the preset sampling number is 8, the sampling module uniformly selects 8 time points within that 4-second duration and extracts the corresponding frames. The extracted video keyframes are output as an image sequence and numbered according to their corresponding time point order. The numbering information and image data are passed to the subsequent feature extraction stage to ensure that the time order of the feature sequence is consistent with the playback order of the original material.
[0044] Furthermore, the feature extraction stage receives the video keyframe image sequence output from the time-axis sampling stage. For each video keyframe, image texture feature extraction and discrete cosine transform (DCT) processing are performed. Image texture feature extraction obtains the grayscale distribution information of the frame's local area. DCT transforms the frame from the pixel domain to the frequency domain, retaining the low-frequency coefficients. The preset deduplication threshold is determined by statistically analyzing the edit distance of a sample set of video materials marked as duplicates and non-duplicates in historical projects. Specifically, curves showing the change in duplicate recognition accuracy and false positive rate under different threshold values are plotted, and the threshold point with the optimal sum of the two is selected as the preset deduplication threshold. Different threshold values can be set according to the content type of the material (e.g., people, scenes). The feature extraction stage converts the low-frequency coefficients corresponding to each video keyframe into single-frame perceptual hash values according to preset quantization rules. The quantization rule compares each low-frequency coefficient with the mean of all coefficients in the frame, marking positions where the coefficient is greater than the mean as 1 and positions where it is less than or equal to the mean as 0, thus obtaining a fixed-length binary hash value. The feature extraction stage ultimately concatenates the single-frame perceptual hash values of each video keyframe in chronological order of sampling to form the corresponding video frame feature sequence and outputs it. At the same time, image texture features are retained as auxiliary information for use in subsequent distance calculation stages.
[0045] Specifically, the input to the time window alignment stage is two video frame feature sequences to be compared. These sequences originate from generated video visual material and existing video material, respectively, and their original durations may differ, resulting in different feature sequence lengths. The time window alignment stage slides along the longer feature sequence, truncating a subsequence of equal length to the shorter one, according to a preset window step. Each truncated subsequence is then aligned with the shorter feature sequence to ensure they are compared at the same relative time progression, avoiding misalignment due to differences in the total duration of the materials. After alignment, the distance calculation stage compares the differences in the perceptual hash values of each aligned feature sequence, and calculates the minimum number of insertion, deletion, and replacement operations required to convert one sequence into the other. This yields the edit distance between the two video frame feature sequences; a smaller edit distance indicates a closer similarity in temporal changes between the two segments.
[0046] Furthermore, the scoring conversion stage receives the edit distance value output from the distance calculation stage and normalizes and maps this edit distance to a video-level visual content similarity score with a fixed value range, based on the length of the feature sequences involved in the comparison. The smaller the edit distance, the higher the similarity score obtained. The scoring conversion stage compares the obtained video-level visual content similarity score with a preset deduplication threshold. When the similarity score reaches the preset deduplication threshold, it is determined that the generated video visual material and the corresponding existing video material are duplicate content. The system establishes a video material reference record accordingly and accumulates the reference count of the existing video material, while no longer storing the generated video visual material separately. When the similarity score does not reach the preset deduplication threshold, the system marks the generated video visual material as material to be verified and retains it for independent storage for further verification and processing by humans or subsequent processes. For example, if the similarity score mapped to the edit distance of a certain comparison is, for example, 0.88, and the preset deduplication threshold is 0.8, the system determines that the two are duplicates and establishes a reference record.
[0047] By organizing multi-frame perceptual hashes into a temporal feature sequence and combining time window alignment and edit distance for video-level similarity determination, the similarity recognition covers the temporal changes of the footage rather than a single frame. This helps to improve the applicability of duplicate video footage recognition and reduce the storage footprint of duplicate content.
[0048] Specifically, the input to the source tracing stage is a material tracing request, which carries a specified historical time node and the identifier of the target material or the target video shot number to be tracked. After receiving the request, the tracing module starts from the lineage node corresponding to the target material and performs a traversal along the directed edges of the video material relationship graph. The traversal direction is specified by the request parameters. Reverse traversal is used to find all source nodes of the target material before the specified historical time node, and forward traversal is used to find all subsequent nodes derived from the target material after the specified historical time node. During the traversal, the tracing module reads the incremental change logs of the nodes it passes through one by one, filters out log records whose occurrence time is earlier than or equal to the specified historical time node, and reconstructs the lineage topology state corresponding to the specified historical time node based on these records. The reconstruction result is represented as a list of nodes and their connections, showing the version status of each material at that time node.
[0049] Furthermore, after completing the lineage topology state reconstruction, the tracking module organizes the generation events, replacement events, and regeneration events involved in the reconstruction results into a chain format according to the chronological order of node occurrences and outputs it. Each link in the chain is labeled with the corresponding material storage identifier, event type, and occurrence time, allowing the requester to view the complete evolution process of the target material from its initial generation to a specified historical time node. For example, if a material tracking request specifies a historical time node as, say, the completion time of the 3rd project iteration, the chain output by the tracking module after reverse traversal shows that the material went through three stages: initial generation, one style adaptation replacement, and one regeneration triggered by an anchor map update. Each stage is labeled with its corresponding occurrence time for verification and reference.
[0050] By maintaining incremental change logs for lineage nodes and combining version node creation with a mechanism that preserves the link between old and new versions, the historical evolution of materials can be reconstructed in a forward or reverse manner along the relationship graph. This helps to improve the operability of verifying the historical status and locating problems after material replacement and regeneration.
[0051] The process of performing cross-project reuse on existing materials that have reached the reuse threshold includes: calculating the cosine similarity between the new project and candidate existing materials in the dimensions of character, scene, and style, and performing a weighted calculation to obtain a cross-project reuse score; when the cross-project reuse score reaches the preset reuse threshold, triggering a style transfer network to perform style adaptation on the existing materials, retaining the character and scene features of the existing materials, and mapping the visual style to the style space corresponding to the new project; injecting the adapted visual materials into the new project, and accumulating the reference count corresponding to the existing materials.
[0052] Specifically, the input to the cross-project reuse score calculation stage is the shot description information corresponding to the new project, as well as the character features, scene features, and style features associated with existing materials in the material library. The input data format is a structured feature vector, derived from the feature parameters extracted and stored during the generation stage of the corresponding materials, without the need for repeated extraction. The calculation stage calculates the cosine similarity between the new project and the candidate existing materials in the character, scene, and style dimensions. The similarity values of the three dimensions respectively reflect the degree of fit between the character image and scene environment required by the new project and the candidate existing materials. Subsequently, the calculation stage performs a weighted sum of the three similarity dimensions according to preset weight coefficients to obtain the cross-project reuse score corresponding to the candidate existing materials. This score is output in numerical form and written into the candidate existing materials reuse sorting list for subsequent reuse triggering stages to read. Specifically, in the calculation stage, after calculating the cosine similarity between the new project and the candidate existing materials in the dimensions of character, scene, and style, the similarity of the three dimensions is first normalized to the maximum value, uniformly set to the range of 0 to 1. Then, the normalized similarity of the three dimensions is weighted and summed according to preset weight coefficients (e.g., the weights of the character dimension, scene dimension, and style dimension are 0.4, 0.3, and 0.3 respectively, since the character is the main factor, it should be slightly higher than the other two, and the sum of the weights is 1). This yields the cross-project reuse score corresponding to the candidate existing materials. The preset weight coefficients are determined based on the correlation statistical analysis between the similarity of each dimension in historical projects and the results of manual reuse judgment. The higher the correlation, the larger the weight coefficient value of the dimension.
[0053] Furthermore, the reuse triggering step compares the cross-project reuse score output from the cross-project reuse scoring calculation step with a preset reuse threshold. Specifically, it extracts the similarity scores between candidate materials and the new project in three dimensions: role, scene, and style, and plots a performance curve against the results of the "actual reuse availability" assessed manually. By finding an optimal cutoff point—a balance that maximizes the cross-project reuse rate of materials while suppressing the false alarm risk of style transfer failure (such as identity feature destruction or scene incongruity) within a safe range—this threshold is determined and serves as a strict criterion for whether the system triggers the style transfer network to perform material reuse. When the cross-project reuse score corresponding to a candidate existing material reaches the preset reuse threshold, the style transfer network is triggered to perform style adaptation processing on that candidate existing material. Candidate existing materials that do not reach the preset reuse threshold do not enter the style adaptation stage and remain in the original material library for subsequent projects to query. The input to the style transfer network is the image data of the candidate existing material and the target style identifier corresponding to the new project; the output is the visual material image data after style adaptation. The style transfer network consists of a feature encoding layer, a style injection layer, and an image reconstruction layer. The feature encoding layer uses a dual-branch convolutional structure to encode the input candidate existing material. One branch extracts character feature components under the constraint of a character segmentation mask, while the other branch extracts scene feature components from the inverted region of the mask. An orthogonality loss function is used to constrain the correlation between the two sets of feature components in the feature space to be below a preset threshold (0.5 in this embodiment, but can be adjusted according to actual conditions) to ensure effective decoupling of character and scene feature components during the encoding stage. The style injection layer receives the style vector corresponding to the target style identifier and uses an adaptive... The style vector is fused with the scene feature components using instance normalization. The character feature components are then directly transmitted to the image reconstruction layer via a residual direct connection path, bypassing the style injection layer. This ensures that the character feature components are not affected by the style injection process. The image reconstruction layer generates the final adapted image based on the fused scene feature components and the directly transmitted character feature components. During the training phase, identity consistency loss is used to constrain the difference between the reconstructed image and the original image in the corresponding regions of the character feature components to not exceed the deviation range. This ensures that the invariance of the character feature components and scene feature components during the generation process has verifiable mathematical constraints, rather than relying solely on the convergence of the network itself.
[0054] Specifically, during the training phase, the style transfer network uses paired original style images and target style images as training samples. During training, the network's internal parameters are adjusted by comparing the content fidelity and style fit between the image reconstruction layer output and the target style image. The training optimizer uses Adam, and the training samples are derived from the inventory data of materials labeled with style categories from historical projects. After training, the style transfer network parameters are fixed and deployed on the visual generation server, and are uniformly invoked by the generation task scheduling module during runtime. After style adaptation processing, the adapted visual material image data is output by the style transfer network to the material injection stage. The material injection stage writes the adapted visual material into the material set corresponding to the new project and establishes an association record between the adapted visual material and the original candidate existing material in the material library. Simultaneously, the material injection stage increments the original reference count field of the candidate existing material, and the incremented reference count is written back to the corresponding record in the material library to reflect the number of times the existing material has been reused in cross-project scenarios.
[0055] By using multi-dimensional weighted scoring to select reusable existing materials and combining them with a style transfer network to achieve style adaptation while preserving character and scene characteristics, existing materials can be adapted to the style requirements of new projects at a lower cost. This helps reduce the duplication of similar content and improves the reuse efficiency of materials across different projects.
[0056] Example 2: In scenarios where multiple characters appear simultaneously in the same video shot, such as a two-person dialogue shot or a multi-person collaborative shot, the aforementioned forced association process will simultaneously write the initial visual reference frames corresponding to the multiple video characters involved in the shot to the first position of the video reference material set. At this time, if each initial visual reference frame comes from a single close-up shot in a different storyboard, the reference frames are often inconsistent with each other in terms of spatial relationship dimensions such as character proportions, viewing angles, and lighting directions. Relying solely on the compatibility score of the aforementioned material conflict detection step is insufficient to fully resolve such spatial relationship conflicts, which can easily lead to the visual generation agent generating images according to multiple contradictory spatial clues, resulting in problems such as disproportionate proportions or incorrect orientations of multiple characters in the same shot.
[0057] Furthermore, for the aforementioned scenario of forced injection of multiple roles, the forced association process, while writing multiple initial visual reference frames into the video reference material set, adds spatial relationship annotation information to each forcibly injected initial visual reference frame. This annotation information includes the relative position range of the subject in the frame, the facial orientation angle range, and the main light source direction category. The annotation information is extracted from the reference frame image by the spatial relationship parsing submodule when the initial visual reference frame is first bound and stored in the consistency anchor point record table along with the reference frame, eliminating the need for repeated parsing during each forced injection. The forced association process outputs all initial visual reference frames involved in the current shot, along with their respective spatial relationship annotation information, to the subsequent spatial conflict resolution stage. The output format is a list of correspondences between reference frame identifiers and spatial relationship annotations, for this stage to read and process.
[0058] Specifically, the spatial conflict resolution stage receives the reference frame identifiers and spatial relationship annotation correspondence list output by the forced association process. First, it calls the aforementioned shot type recognition module to determine the composition layout type of the current shot. The composition layout types include three categories: parallel dialogue layout, foreground / background layered layout, and surrounding interactive layout. The determination is based on the character relative position descriptions and shot motion information recorded in the current shot's text description. According to the determined composition layout type, the resolution stage reads the standard relative position range, standard orientation angle range, and unified light source direction category of each subject corresponding to that layout type from a preset layout template table. It then compares the read standard spatial parameters with the original spatial relationship annotation information of each initial visual reference frame. For initial visual reference frames whose relative position range, orientation angle range, or light source direction category deviates from the standard spatial parameters, a corresponding spatial adjustment description is generated. The adjustment description records in text form the relative position, orientation, and light source direction that the reference frame should be adjusted to in this generation. For example, if a two-person dialogue scene is determined to be a parallel dialogue layout, and the original orientation angle range of one character's initial visual reference frame deviates from the standard orientation angle range of the parallel dialogue layout by more than the deviation range, the spatial adjustment description generated in the resolution process records that the reference frame should be adjusted to face the side of the center line of the screen in this generation.
[0059] The layout template table is constructed by combining classic film and television composition rules with prior statistical analysis of massive amounts of high-quality video samples. It extracts features from a large number of standard multi-person interactive scenes (such as parallel dialogues, foreground and background layering), and statistically analyzes the relative position boundaries (coordinate range), facial orientation distribution (yaw and pitch angle range), and reasonable lighting direction categories of the main characters in each typical composition scene. Subsequently, this spatial relationship data, which conforms to human visual aesthetics and physical logic, is subjected to cluster analysis and parameter boundary delineation, and is structured and solidified into a template library. This provides the system with standardized and quantifiable spatial constraint benchmarks when generating multiple characters in the same frame.
[0060] Furthermore, the spatial conflict resolution process writes the generated spatial adjustment descriptions into the spatial constraint field of the current shot's corresponding generation plan. This field, together with the reference map identifier field after the aforementioned three-layer verification and completion, constitutes the final generation plan, and both are transmitted to the visual generation agent. While calling the corresponding initial visual reference frame based on the reference image identifier field, the visual generation agent reads the relative position, orientation, and light source direction adjustment requirements recorded in the spatial constraint field. It converts these adjustment requirements into a region mask and structural control map of the corresponding subject in the image. The structured condition injection module then inputs the region mask and structural control map, along with the initial visual reference frame, into the visual generation agent as bypass control conditions for the latent space feature trajectory during the generation process. During generation, the structured condition injection module only constrains and adjusts the position, orientation, and lighting-related latent space features within the region mask's coverage area. The latent space feature components of the character's facial features and clothing features outside the region mask maintain their values from the initial visual reference frame and are not adjusted. This ensures that the final position, orientation, and lighting direction of each character object in the image are consistent with the unified spatial parameters after resolution, while also reducing the degree to which facial features and clothing features undergo unexpected changes during the spatial adjustment process.
[0061] By adding spatial relationship annotations during the forced injection process and combining them with composition layout judgment to generate spatial adjustment descriptions, shots with multiple characters appearing simultaneously can obtain a unified spatial layout basis while retaining the appearance characteristics of their respective initial visual reference frames. This helps to alleviate the problem of abnormal screen proportions or orientations caused by inconsistent spatial relationships of reference frames in scenes with multiple subjects in the same shot.
[0062] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for ensuring cross-shot consistency and managing multimodal materials in short video generation, characterized in that, include: Video content consistency anchors are extracted from short video creation scripts, shot planning information and shot sequences. Initial visual reference frames are bound to the consistency anchors, and video image feature parameters corresponding to the initial visual reference frames are extracted as cross-shot comparison benchmarks. Based on the video image feature parameters, a matching score is calculated between the candidate visual reference material and the current video shot, and the materials are sorted according to the matching score. The initial visual reference frame is then added as the highest priority material to the video reference material set corresponding to the current video shot. Perform three-level verification on the video reference material set, and generate video visual materials corresponding to each video shot based on the verified video reference material set; When the video reference material corresponding to the initial visual reference frame is replaced, the associated video shot in the corresponding storyboard is marked as a shot to be regenerated, and the corresponding video visual material is regenerated. The generated video footage is similar to existing video footage. When the deduplication threshold is reached, a reference relationship is established. A video footage association graph is constructed based on the footage generation record, and source tracing is performed.
2. The method for ensuring cross-shot consistency and managing multimodal materials in short video generation according to claim 1, characterized in that, Extracting video content consistency anchor points and binding them to initial visual reference frames includes: The text description information in the short video creation script and the shot description information in the storyboard sequence are subjected to video content semantic analysis to extract video character objects, video scene objects, and video visual representation objects. The storyboard sequence is traversed, and the video visual material with the highest image quality evaluation parameter in the video shot in which the video character object or video scene object first appears is bound as the corresponding initial visual reference frame. The similarity of video image features between the video visual material generated in subsequent video shots and the corresponding initial visual reference frame is calculated. When the similarity of video image features is lower than a preset consistency threshold, the storyboard sequence is traversed backward to find alternative video shots in which the video character object or video scene object reappears and has a higher image quality evaluation parameter. The corresponding video visual material is then updated as a new initial visual reference frame until the similarity of video image features meets the preset consistency threshold.
3. The method for ensuring cross-shot consistency and managing multimodal materials in short video generation according to claim 1, characterized in that, The calculation of the matching score between the candidate visual reference material and the current video shot includes: A shot type recognition module is used to predict the semantic type of a video shot based on the text description information and shot motion information corresponding to the current video shot. The semantic type of the video shot includes close-up shots of people, medium shots of people, long shots of the environment, and dynamic tracking shots. The weight parameters corresponding to the character object features, scene environment features, visual style features, motion trajectory features, and video composition features are adaptively adjusted according to the semantic type of the video shot. Specifically, when the semantic type of the video shot is a close-up shot of people, the weight parameters corresponding to the character object features are increased; when the semantic type of the video shot is a dynamic tracking shot, the weight parameters corresponding to the video composition features and motion trajectory features are increased. The matching degree of candidate visual reference materials in the dimensions of character object consistency, scene environment consistency, visual style consistency, and video composition consistency is calculated respectively, and the material credibility coefficient is determined according to the source type of the candidate visual reference materials. The video matching score of the candidate visual reference materials is obtained by weighted calculation based on the matching degree of each dimension, the material credibility coefficient, and the adaptively adjusted weight parameters.
4. The method for ensuring cross-shot consistency and managing multimodal materials in short video generation according to claim 1, characterized in that, Adding the initial visual reference frame to the video reference material set includes: A forced association process, independent of video matching scores, is adopted to traverse the video character objects and video scene objects involved in the current video shot. When a corresponding initial visual reference frame exists in the video content consistency anchor point, the initial visual reference frame is added to the current video reference material set as the highest priority video consistency constraint material. The remaining candidate visual reference materials are sorted in descending order according to the video matching scores, and candidate visual reference materials with video matching scores greater than a preset reference threshold and whose video reference material set has not reached a preset maximum number of references are added to the video reference material set.
5. The method for ensuring cross-shot consistency and managing multimodal materials in short video generation according to claim 1, characterized in that, Performing three levels of validation and missing data completion on the video reference material set includes: The first layer of verification is the integrity check of the generated parameters. It checks whether the video shot generation configuration contains video reference frame information. If it is missing, it matches the corresponding initial visual reference frame based on the video image features and fills it in. The second layer of verification is the cross-shot consistency check. It checks whether the video reference material contains the corresponding consistency anchor point. If it is missing, it fills in the corresponding initial visual reference frame. The third layer of verification is the material conflict check. When the initial visual reference frame conflicts with other reference materials, it selects the optimal combination of reference materials based on the material compatibility score. At the same time, it counts the verification triggering situation and adjusts the video reference material selection strategy for the corresponding shot based on the abnormal results.
6. The method for ensuring cross-shot consistency and managing multimodal materials in short video generation according to claim 1, characterized in that, The similarity calculation between the generated video visual material and existing video material includes: The generated video visual material and existing video material are sampled along the timeline to obtain multiple video keyframes. Image texture features and discrete cosine transform low-frequency coefficients are extracted from the video keyframes to generate a video frame feature sequence. Video segments of different lengths are aligned using a time window, and the distance between the video frame feature sequences is calculated to obtain a video-level visual content similarity score. When the video-level visual content similarity score is not less than a preset deduplication threshold, a video material reference record is established; otherwise, the corresponding video visual material is marked as material to be verified.
7. The method for ensuring cross-shot consistency and managing multimodal materials in short video generation according to claim 1, characterized in that, The process of constructing a video material generation relationship graph and performing source tracing includes: establishing a versioned video material relationship graph and recording information on video material modifications, replacements, and regenerations; when the initial visual reference frame is replaced and triggers the regeneration of video footage, creating a corresponding new version node in the video material relationship graph and retaining the original version material relationship path; when a material tracing request is received, performing forward and / or reverse traversal along the video material relationship graph according to a specified time node, and outputting the corresponding video material generation, replacement, and regeneration links.