An intent-driven material retrieval revision and video generation method
Patent Information
- Application Number
- CN202611316633.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-25
AI Technical Summary
当前,现有技术中与视频素材检索、剪辑及成片相关的方案存在诸多局限,难以满足意图驱动下的全流程素材处理需求,具体主要集中在以下几个方向:
[0056]本发明具有以下优点:本发明以用户创作意图为导向,实现文本、语音、图片多模态指令与视频素材的精准语义对齐,自动检测素材库中残缺视频片段并进行针对性补全修正,同时结合专业镜头语法约束,将检索、修正后的素材整合生成语义准确、片段完整、镜头流畅的可直接发布成品,有效提升视频素材利用率和视频剪辑效率。
Smart Images

Figure CN122817508A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to an intent-driven method for material retrieval, correction, and video generation. Background Technology
[0002] In recent years, with the rapid development of content industries such as short videos, self-media, and corporate promotional videos, the scale of video material libraries has exploded, and users' demand for accurate retrieval, efficient modification, and rapid production of video materials has become increasingly urgent. Currently, existing solutions related to video material retrieval, editing, and production have many limitations and cannot meet the end-to-end material processing needs driven by intent. Specifically, these limitations mainly focus on the following areas:
[0003] (1) Existing technologies mostly use text keywords to match video titles and subtitles, or use single image features to match video frames to achieve preliminary material retrieval. They cannot achieve accurate semantic alignment between text, voice, and image multimodal commands and video materials, and it is difficult to deeply analyze the core creative intent behind the user's multimodal commands. The problem of "commands not matching material content" often occurs, causing the search results to deviate from the user's intent and failing to provide an accurate basis for subsequent correction and final film generation.
[0004] (2) Existing AI editing solutions mostly adopt the "complete segment splicing" mode, which can only mechanically splice complete video segments in the material library according to simple instructions. It cannot handle incomplete or fragmented video materials, nor can it make targeted corrections to materials that deviate from the user's intentions. Furthermore, it cannot integrate the corrected materials according to professional shot syntax based on the user's core intentions to generate a finished film that meets the requirements. The finished film is stiff and lacks logic.
[0005] (3) Existing video completion technology does not combine the scene, motion trajectory, and lighting style of the original material, nor does it associate with the user's core creative intent. The corrected segment is not naturally connected with the original material, has obvious splicing marks, and does not conform to the logic of professional shot editing. It cannot achieve precise linkage of "retrieval-correction", which leads to insufficient coherence and consistency in the subsequent production of finished film, and cannot meet the demand for high-quality finished film driven by intent.
[0006] (4) The videos generated by the existing technology do not follow the rules of professional editing and do not optimize the shot connection and final logic based on the user's core intent, resulting in poor smoothness and weak logic of the finished video, which cannot be directly published and used, thus violating the core goal of efficient and high-quality final video generation under the "intent-driven" approach. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intent-driven material retrieval, correction, and video generation method.
[0008] The objective of this invention is achieved through the following technical solution: an intent-driven material retrieval, correction, and video generation method, comprising the following steps:
[0009] S1: Extract multi-dimensional features from all video materials in the material library to construct a structured material feature library;
[0010] S2: Receive multimodal instructions from the user, parse the user's creative intent behind the multimodal instructions, and transform them into a unified semantic feature vector;
[0011] S3: Align multimodal semantic features with video material features based on the semantic feature vector corresponding to the user intent;
[0012] S4: Perform semantic deviation verification on the retrieved video segments, identify semantic deviations that do not conform to the user's intent, extract relevant contextual features of the segments, and generate correction instructions;
[0013] S5: Based on the correction instructions, and combined with the contextual features of the retrieved fragments and the structural features of the material library, perform attribute correction on the fragments with semantic deviations and complete the missing elements;
[0014] S6: Guided by the user's creative intent, it splices and merges the completed and corrected segments or the retrieved complete segments to generate the finished video.
[0015] Preferably, in step S1, the extracted features include visual features, audio features, and semantic tag features.
[0016] The extraction of visual features specifically involves: using the ResNet18 model to extract texture, color, and object features from each frame of the video; simultaneously, using optical flow to extract inter-frame motion features; and fusing the visual and motion features of each frame to obtain the video visual feature vector. ;
[0017] The audio feature extraction process specifically involves: using MFCC to extract frequency and rhythm features from the video's audio track; extracting the rhythm peaks using a peak detection algorithm; and outputting an audio feature vector. ;
[0018] Extracting semantic label features specifically involves using the CLIP lightweight model to label video frames with scene, object, and action tags, and then converting these tags into semantic feature vectors. ;
[0019] Will , and Weighted fusion is performed to obtain the comprehensive feature vector of each video segment. All video clips' comprehensive feature vectors and original material information are stored in the material feature library, and an index is created.
[0020] Preferably, step S2 further includes the following step:
[0021] S21: Receive text commands input by the user, encode the text using a lightweight BERT model, remove redundant information, extract core semantics, and output a text semantic feature vector. =[ ], where n is the feature dimension;
[0022] S22: Receive user-input voice commands, extract voice features using MFCC, convert the voice to text using a speech-to-text model, and then output a semantic feature vector with the same dimension as the text command using the text encoding method from step S21. ;
[0023] S23: Receive the sketch drawn by the user, use a CNN model to extract the outline and object features of the sketch, and map them into a semantic feature vector of the same dimension as the text instruction. Ultimately Transformed into a unified semantic feature vector .
[0024] Preferably, in step S3, the SemAlignNet algorithm is used to accurately align multimodal semantic features with video material features, specifically including a feature mapping module, a cross-modal attention alignment module, and a similarity calculation module;
[0025] The feature mapping module uses learnable fully connected layers to... and Mapping to the same feature space, the mapping formula is:
[0026] ;
[0027] ;
[0028] in, and All are learnable weight matrices. and For bias terms;
[0029] The cross-modal attention alignment module is used to highlight the correspondence between the core semantics of user commands and the key features in video footage, and to calculate the attention weight matrix. ,
[0030] ;
[0031] in, This is the attention weight matrix. The function is used to normalize the weights. To transpose video features;
[0032] The similarity calculation module combines attention weight calculation. and cosine similarity,
[0033] ;
[0034] in, To weight the special effects matching score, To normalize the feature length, · It is a norm.
[0035] Preferably, step S4 further includes the following step:
[0036] S41: Based on unified semantic feature vector This maps it to a structured set of ideal intent semantics:
[0037] ;
[0038] in, As a collection of core concepts, For a set of attribute constraints, It is a spatial relation vector. The feature dimension of the spatial relation vector;
[0039] S42: Based on extracted video segment features By combining visual recognition algorithms, the actual semantic set of the video clips can be parsed out:
[0040] ;
[0041] in, It represents the total number of semantic tuples parsed from the video clip;
[0042] S43: Quantitative Analysis and Deviation;
[0043] S44: Construct the material topology graph and the ideal intent topology graph, and calculate the material topology graph. Topology diagram with ideal intention The overall deviation value;
[0044] S45: Output correction instruction set .
[0045] Preferably, step S43 further includes the following step:
[0046] S43.1: Yes Each core concept in exist In performing semantic embedding matching, if exist If there is no match, then ( , , Add to missing semantic set ,like There is a match, but the attributes and Chinese correspondence If the similarity is below the threshold θ, then ( , , ) and incorrect match ( , , ) Forming an error semantic pair set ;
[0047] S43.2: Quantify the overall deviation using the distance formula.
[0048] ;
[0049] in, Weighting the difficulty of completing missing content. As the weight of spatial location deviation, Weights for differences in image attributes. This represents the total semantic deviation value. For concept similarity terms, For missing semantic concepts to be completed, To fill in the ideal semantic concept as expected. For spatial topological distance, For the spatial location topological features of the erroneous video / material, Topological features of the target spatial location for standard ideal materials. For attribute difference items, The distribution of style, lighting, and color attributes of the incorrect material. The attribute distribution of standard ideal materials.
[0050] Preferably, in step S44, the material topology map Topology diagram with ideal intention The formula for calculating the overall deviation value is:
[0051] ;
[0052] in, This represents the overall topological deviation value. , , These are learnable weight coefficients, corresponding to the weight proportions of node bias, edge relationship bias, and attribute bias, respectively. , For node topology deviation, This represents the topological deviation of the edge relationship. This refers to the topological deviation of the attribute.
[0053] Preferably, in step S5, for The specific processing of missing completion instructions is as follows: based on the user's semantic intent and the ideal topological graph structure, combined with the scene style, scale ratio, and perspective relationship of the original video, missing elements matching the scene logic are generated within the coordinate position and motion range specified by the instruction.
[0054] against The specific processing of error correction instructions is as follows: based on the element ID, modification dimension and target parameters marked in the instruction, the material, color, style and position attributes of objects with deviations in the video are adjusted frame by frame;
[0055] against The specific processing of the spatiotemporal constraint instructions is as follows: the generated model is forced to follow the optical flow direction, inter-frame transition rules and object spatial relationships of the original video, and the motion consistency and spatiotemporal topology compliance of the newly added content and the correction area are constrained.
[0056] This invention has the following advantages: Guided by user creative intent, this invention achieves precise semantic alignment between multimodal commands (text, voice, and images) and video materials. It automatically detects incomplete video segments in the material library and performs targeted completion and correction. At the same time, combined with professional shot grammar constraints, it integrates the retrieved and corrected materials to generate semantically accurate, complete, and smooth finished products that can be directly published, effectively improving the utilization rate of video materials and the efficiency of video editing. Attached Figure Description
[0057] Figure 1 This is a schematic diagram of the process of intent-driven material retrieval, correction, and video generation. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0059] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0060] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.
[0061] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0062] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0063] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0064] In this embodiment, as Figure 1 As shown, an intent-driven material retrieval, correction, and video generation method includes the following steps:
[0065] S1: Extract multi-dimensional features from all video materials in the material library to construct a structured material feature library; specifically, in step S1, the extracted features include visual features, audio features, and semantic tag features.
[0066] The extraction of visual features specifically involves: using the ResNet18 model to extract texture, color, and object features from each frame of the video; simultaneously, using optical flow to extract inter-frame motion features; and fusing the visual and motion features of each frame to obtain the video visual feature vector. ;
[0067] The audio feature extraction process specifically involves: using MFCC to extract frequency and rhythm features from the video's audio track; extracting the rhythm peaks using a peak detection algorithm; and outputting an audio feature vector. ;
[0068] Extracting semantic label features specifically involves using the CLIP lightweight model to label video frames with scene, object, and action tags, and then converting these tags into semantic feature vectors. ;
[0069] Will , and Weighted fusion is performed with weights of 0.5, 0.3, and 0.2 to obtain the comprehensive feature vector for each video segment. All video clips' comprehensive feature vectors and original material information are stored in the material feature library, and an index is created.
[0070] S2: Receive multimodal instructions from the user, parse the user's creative intent behind the multimodal instructions, and transform them into a unified semantic feature vector;
[0071] S3: Align multimodal semantic features with video material features based on the semantic feature vector corresponding to the user intent;
[0072] S4: Perform semantic deviation verification on the retrieved video segments, identify semantic deviations that do not conform to the user's intent, extract relevant contextual features of the segments, and generate correction instructions;
[0073] S5: Based on the correction instructions, and combined with the contextual features of the retrieved fragments and the structural features of the material library, perform attribute correction on the fragments with semantic deviations and complete the missing elements;
[0074] S6: Guided by the user's creative intent, the corrected and supplemented segments or the retrieved complete segments are spliced and merged to generate the finished video. Specifically, after the processing of the correction and generation modules mentioned above, all video segments to be synthesized have met the basic requirements of semantic accuracy, completeness of image, smooth timing, and logical compliance. However, individual segments still have problems such as disordered shot arrangement, abrupt transitions, chaotic rhythm, and non-compliance with film and television creation standards, making it impossible to directly form a complete and high-quality finished video. To solve the shortcomings of traditional intelligent editing algorithms that only perform simple splicing of segments, ignore professional shot grammar, and produce a stiff viewing experience, step S6 is guided by the user's core creative intent and introduces a professional film and television shot grammar constraint mechanism. It standardizes the splicing, merging, and rhythm optimization of the corrected and supplemented high-quality video segments and the original complete retrieved segments, standardizes the shot narrative logic and image connection rules, and finally outputs a finished video that conforms to the creative intent, fits film and television aesthetics, and can be directly published, realizing the intelligent generation of finished videos that are intent-driven and grammatically compliant. Specifically, this module abandons the traditional disordered splicing synthesis method and constructs a three-dimensional final product synthesis principle that is intent-driven, grammatically constrained, and temporally compliant. The entire synthesis process relies on multiple prior rules to constrain the synthesis process: First, user creative intent constraints, based on the narrative logic, creative theme, and content priority of the initial user instructions, determine the core arrangement order of video clips to ensure that the narrative main line of the final product is clear and the semantics fit the creative needs; Second, professional shot grammar constraints, following the basic rules of film and television editing, including professional standards such as shot scale progression, movement shot connection, dynamic and static shot combination, and narrative rhythm control, to avoid problems such as disordered splicing and shot conflicts; Third, spatiotemporal feature consistency constraints, continuing the STTCG spatiotemporal topology logic and image feature consistency rules mentioned above, to ensure that the style, lighting, motion trajectory, and audio atmosphere of the overall video after synthesis are unified and coherent, without local inconsistencies or logical conflicts.
[0075] For multiple video clips to be synthesized, this application uses professional shot grammar to intelligently sort and hierarchically arrange them, constructing a standardized narrative structure. First, based on the user's semantic intent, the creative narrative is broken down into three main video segments: introductory setup, main narrative, and concluding summary, matching corresponding video clips. Second, combining shot grammar rules, the shots are arranged according to a progressive logic of long shot setup—medium shot narrative—close-up focus—extreme close-up enhancement, while also adapting to the principle of alternating static and dynamic shots to avoid monotonous visuals from continuous static shots and visual fatigue from continuous dynamic shots. Furthermore, based on the results of pre-production optical flow detection and inter-frame difference detection, the motion states of the shots are classified and matched to ensure natural transitions in the direction and rhythm of movement between adjacent shots, preventing abrupt changes in shot movement and narrative jumps, ensuring that the overall shot arrangement of the final film conforms to public visual perception and the rules of film and television creation.
[0076] To address issues such as disjointed shots, abrupt transitions, and rhythmic breaks caused by splicing multiple segments, this application employs refined optimization at shot transition points to achieve seamless transitions and smooth flow. Based on inter-frame difference features and optical flow abrupt change detection results, it intelligently identifies shot transition boundaries and applies differentiated transition strategies to different shot groups: for adjacent shots with coherent content and consistent movement trends, an inter-frame gradual blending method is used to weaken the sense of shot boundaries and ensure smooth temporal transitions; for shots with changes in shot size or scene jumps, lightweight and compliant transition effects are matched to fit the overall video style and do not damage the original texture of the image. At the same time, it strictly adheres to shot duration syntax specifications and adaptively adjusts the display duration of individual shots according to the overall creative rhythm of the video, avoiding shots that are too long and dragging or too short and rushed, and accurately controlling the narrative rhythm of the final film.
[0077] After splicing multiple segments, a global secondary calibration and optimization is performed on the final product to eliminate local differences caused by multi-segment synthesis. At the visual level, the overall video's hue, brightness, saturation, and texture style are unified, smoothing out subtle image quality differences between segments. At the temporal level, the consistency of motion between frames is globally calibrated to ensure consistent object motion logic and optical flow patterns throughout the film, eliminating local jitter, jumps, and logical conflicts. At the audio-visual matching level, based on the global audio rhythm, shot transition nodes are fine-tuned to ensure a high degree of compatibility between shot transitions, visual rhythm, and audio beats and emotional atmosphere, achieving audio-visual synchronization and rhythmic unity. Simultaneously, relying on the spatiotemporal topology constraint graph, the spatial relationships of objects and the compliance of scene logic throughout the film are verified a second time, ensuring the final product is free of logical loopholes and visual inconsistencies.
[0078] After shot grammar constraints, transition optimization, and global consistency calibration, the final product video is generated with precise and complete semantics, standardized shot arrangement, smooth and natural transitions, and a well-paced rhythm. The finished product perfectly matches the user's core creative intent, solving the problems of logical inconsistencies, abrupt shots, and rough visuals in traditional intelligent editing, while preserving the authentic texture of the original material and the precise effect of the corrected segments, balancing semantic accuracy, visual aesthetics, and film and television professionalism. As the final output stage, this application completes a closed-loop process from intent parsing, deviation detection, precise correction to grammar synthesis, and can directly output finished videos that meet publishing standards, significantly lowering the threshold for user video creation and improving the quality and practicality of intelligent video editing. In other words, this invention is user-inspired, achieving precise semantic alignment between multimodal commands (text, voice, and images) and video materials, automatically detecting and correcting incomplete video segments in the material library, and integrating the retrieved and corrected materials with professional shot grammar constraints to generate semantically accurate, complete, and smoothly editable finished products that can be directly published, effectively improving the utilization rate of video materials and the efficiency of video editing.
[0079] Furthermore, step S2 also includes the following steps:
[0080] S21: Receive text commands input by the user, encode the text using a lightweight BERT model, remove redundant information, extract core semantics, and output a text semantic feature vector. =[ ], where n is the feature dimension;
[0081] S22: Receive user-input voice commands, extract voice features using MFCC, convert the voice to text using a speech-to-text model, and then output a semantic feature vector with the same dimension as the text command using the text encoding method from step S21. This ensures semantic consistency between voice commands and text commands;
[0082] S23: Receive the sketch drawn by the user, use a CNN model to extract the outline and object features of the sketch, and map them into a semantic feature vector of the same dimension as the text instruction. Ultimately Transformed into a unified semantic feature vector This step enables unified encoding of multimodal instructions. Specifically, guided by the user's core creative intent, this step transforms different types of user instructions into unified semantic feature vectors, ensuring consistency in subsequent semantic alignment and providing a precise intent benchmark for intent-driven material retrieval and correction.
[0083] Furthermore, in step S3, the SemAlignNet algorithm is used to accurately align multimodal semantic features with video material features, specifically including a feature mapping module, a cross-modal attention alignment module, and a similarity calculation module.
[0084] The feature mapping module uses learnable fully connected layers to... and Mapping to the same feature space, the mapping formula is:
[0085] ;
[0086] ;
[0087] in, and All are learnable weight matrices. and For bias terms;
[0088] The cross-modal attention alignment module is used to highlight the correspondence between the core semantics of user commands and the key features in video footage, and to calculate the attention weight matrix. ,
[0089] ;
[0090] in, This is the attention weight matrix, which maps text features to the same space as video features, allowing for direct comparison of their similarity. The function is used to normalize the weights, ensuring that the sum of the weights is 1. Transposing the video features is simply to allow the two matrices to be legally multiplied; it does not change the underlying meaning.
[0091] The similarity calculation module combines attention weight calculation. and cosine similarity,
[0092] ;
[0093] in, To weight the special effects matching score, To normalize the feature length, · The norm is used to eliminate computational biases caused by varying feature lengths. Specifically, the algorithm input for this step is the multimodal unified semantic feature vector. The combined feature vector of all video clips in the material library The output is sorted according to the similarity Sim and outputs the Top-K video segments that best match the user's instructions. It also outputs the start / end time and similarity score of each segment for use in subsequent steps.
[0094] In this embodiment, step S4 further includes the following step:
[0095] S41: Based on unified semantic feature vector This maps it to a structured set of ideal intent semantics:
[0096] ;
[0097] in, As a collection of core concepts, For a set of attribute constraints, It is a spatial relation vector. The feature dimension of the spatial relation vector;
[0098] S42: Based on extracted video segment features By combining visual recognition algorithms, the actual semantic set of the video clips can be parsed out:
[0099] ;
[0100] in, It represents the total number of semantic tuples parsed from the video clip;
[0101] S43: Quantitative Analysis and The deviation; furthermore, step S43 also includes the following steps:
[0102] S43.1: Yes Each core concept in exist In performing semantic embedding matching, if exist If there is no match, then ( , , Add to missing semantic set ,like There is a match, but the attributes and Chinese correspondence If the similarity is below the threshold θ, then ( , , ) and incorrect match ( , , ) Forming an error semantic pair set ;
[0103] S43.2: Quantify the overall deviation using the distance formula.
[0104] ;
[0105] in, Weighting the difficulty of completing missing content. As the weight of spatial location deviation, Weights for differences in image attributes. This represents the total semantic deviation value. For concept similarity terms, For missing semantic concepts to be completed, The ideal semantic concept for filling in missing parts is used to specifically measure the difficulty of completing the missing content; the lower the similarity, the greater the difficulty of completion. For spatial topological distance, For the spatial location topological features of the erroneous video / material, The topological features of the target spatial location for standard ideal materials are used to quantify the spatial offset error in the target position and structural layout of the image. For attribute difference items, The distribution of style, lighting, and color attributes of the incorrect material. To determine the attribute distribution of a standard ideal source material, the difference between two attribute distributions can be calculated using methods such as KL divergence. The aim is to quantify the degree of inconsistency in visual attributes such as style, lighting, and color tone. Specifically, this formula uses weighting coefficients... , , This method employs weighted fusion of conceptual completion difficulty, spatial topological position deviation, and visual attribute distribution differences. Embedding similarity is used to measure the semantic gap between the missing concept and the ideally completed concept, topological distance is used to characterize the degree of spatial positional offset of the material, and distribution differences are used to measure inconsistencies in attributes such as image style and lighting. Unlike traditional methods that only calculate global similarity, this method deconstructs the sources of deviation from a difference set perspective, separating two types of deviation dimensions: missing and incorrect, providing quantitative support for accurate matching and intelligent correction of video material. The total semantic deviation value is used to quantify the overall deviation in the process of video / material matching and completion. The larger the value, the more serious the difference and deviation between the current material and the ideal standard; the smaller the value, the closer the matching and completion effect is to expectations.
[0106] S44: Construct the material topology graph and the ideal intent topology graph, and calculate the material topology graph. Topology diagram with ideal intention The overall deviation value; further, in step S44, the material topology map Topology diagram with ideal intention The formula for calculating the overall deviation value is:
[0107] ;
[0108] in, This represents the overall topological deviation value. The larger the value, the greater the difference in spatiotemporal topological structure between the source video and the ideal scene, and the more severe the logical deviation. , , These are learnable weight coefficients, corresponding to the weight proportions of node bias, edge relationship bias, and attribute bias, respectively. , To address node topology deviation, the differences in entity nodes between two graphs are quantified, including three types of errors: missing nodes, redundant nodes, and mismatched node types. Statistics are then compiled. and The missing rate and error rate of matching core entity nodes To address topological deviations in edge relationships, this study quantifies differences in spatiotemporal relationships, encompassing matching deviations in spatial static positional relationships and temporal dynamic motion relationships. It characterizes the degree of disorder in object hierarchy, position, and motion logic. To address attribute topology deviation, this method quantifies the differences in attribute parameters of graph nodes and edges, including the degree of inconsistency in key attributes such as object style, motion direction, and scene materials. The overall approach replaces the traditional single-dimensional graph editing distance calculation method with a weighted fusion of three types of deviations, achieving full coverage quantification of topology deviation across three dimensions: entities, relationships, and attributes. Specifically, constructing the material topology graph involves automatically analyzing the content of the original video material to be corrected, constructing a spatiotemporal topology graph that closely matches the actual video state, and fully recording the object composition and motion logic of the current video. Node definition (entity elements): Extracts all core semantic entities from the video frames, including scene objects, moving targets, and background elements such as tanks, ground, sky, and vegetation; each independent object corresponds to a graph node. Edge definition: Divided into two categories: spatial static relationships and temporal dynamic relationships, comprehensively covering the video logic. Spatial static edges describe the position, subordination, and coverage relationships between objects, such as a tank above the ground and the sky covering the entire scene. Temporal dynamic edges, combined with optical flow detection results, describe the motion trajectory and motion logic of objects, such as a tank moving forward and a bird moving from right to left. The purpose of constructing the material topology map is to completely replicate the real scene structure and motion state of the original video, serving as a benchmark sample for deviation detection.
[0109] The construction of the ideal intent topology graph involves pre-constructing a standard ideal scene topology structure that meets the user's creative needs, based on the semantic parsing results of prior user commands and without relying on the original video footage. The nodes and edges of this graph are fully aligned with the semantic requirements of the user commands: the object elements, object positional relationships, scene style, object movement directions, and scene hierarchy relationships required by the user are all solidified into a standardized topology graph structure. The purpose of constructing the ideal intent topology graph is to define a correct, compliant, and user-intended standard video structure, serving as a truth template for deviation judgment. To quantify the structured logical differences between the source video topology graph and the user's ideal intent topology graph, and to address the problem that traditional similarity algorithms cannot quantify spatiotemporal topological deviations, this paper uses the graph edit distance concept to accurately calculate the source video topology graph. Topology diagram with ideal intention The overall deviation value enables the quantification, calculation, and traceability of topological differences.
[0110] S45: Output correction instruction set Specifically, the missing correction instruction set. This feature addresses the issue of missing scene elements detected during topology map comparison, generating instructions to add new elements to supplement the core content missing from the video and meeting user requirements. Standard format: ADD: Element type, position coordinates, dynamic parameters, style parameters, quantity parameters; Instruction parsing: Includes all attributes of the new object, precisely controlling the position, shape, movement, and style of the supplemented content, preventing blind or incorrect supplementation. Example: ADD: Flock of birds, (x:100, y:200), dynamic trajectory: from right to left, style: realistic, quantity: 5.
[0111] Error correction instruction set This function generates attribute modification and content replacement instructions for existing elements in a video that do not conform to user intent, have topological logic errors, or have mismatched attributes, correcting existing visual defects. Standard format: MOD: Element ID, Attribute Modification Item, Original Value, Target Value; Instruction Explanation: Precisely locates the erroneous visual element, clarifies the attribute dimensions that need modification, specifies the correction target, and achieves refined replacement. Example 1 (Scene Attribute Correction): MOD: Background Ground, Material: Desert, Target Value: Grassland; Example 2 (Target Attribute Correction): MOD: Tank, Color: Desert Camouflage, Target Value: Jungle Camouflage.
[0112] Spatiotemporal topology constraint instructions Unlike traditional pixel-level correction instructions, this instruction is a global logical constraint instruction, generated based on STTCG topological relationships, forcing subsequent video generation, completion, and editing algorithms to comply with physical spatiotemporal rules. Standard format: CONSTRAINT: constraint type, constraint dimension, target rule; core function: to solve problems such as image jitter, motion distortion, object relationship conflicts, and inter-frame non-smoothness, ensuring the temporal continuity and spatial compliance of the corrected video. Example: CONSTRAINT: motion consistency, optical flow direction: right. In summary, step S4 aims to overcome the limitations of existing technologies that rely solely on pixel-level or simple feature similarity comparisons. This application introduces a multimodal semantic difference recognition and topological constraint graph matching algorithm to achieve deep semantic deviation quantification and precise correction instruction generation for retrieved segments. The core logic of the algorithm is: treating the user intent as an "ideal semantic feature set" and the retrieved video segment as an "actual semantic feature set," calculating the semantic difference between the two, not only identifying missing semantic elements but also elements with inconsistent semantic attributes, and combining this with the spatiotemporal topological constraint graph to generate high-precision correction instructions.
[0113] In this embodiment, in step S5, for The specific processing of missing completion instructions is as follows: based on the user's semantic intent and the ideal topological graph structure, combined with the scene style, scale ratio, and perspective relationship of the original video, missing elements matching the scene logic are generated within the coordinate position and motion range specified by the instruction.
[0114] against The specific processing of error correction instructions is as follows: based on the element ID, modification dimension and target parameters marked in the instruction, the material, color, style and position attributes of objects with deviations in the video are adjusted frame by frame;
[0115] against The specific processing of spatiotemporal constraint instructions involves: forcing the generation model to follow the optical flow direction, inter-frame transition rules, and object spatial relationships of the original video, and constraining the consistency of motion and spatiotemporal topology compliance between the newly added content and the correction area. Specifically, this stage, based on generative AI algorithms, relies on correction instructions to drive intelligent video correction and content generation, completing precise optimization of deviation segments and completion of missing content, while ensuring a high degree of integration between the correction area and the original video segment, outputting high-quality video material that conforms to the user's semantic intent and physical spatiotemporal logic. The correction and generation work in this module is not unconstrained random generation, but rather driven by multi-dimensional prior features. The core input includes three types of effective information: first, the correction instruction set generated earlier. The application employs a multi-dimensional information fusion approach. First, it clearly defines the specific rules for element addition, attribute modification, and spatiotemporal constraints. Second, it utilizes the temporal characteristics of the retrieved original video clips, including inter-frame motion patterns and transition logic. Third, it incorporates pre-constructed structured feature information from a material library, covering prior parameters such as scene style, object attributes, light and shadow distribution, and audio features. This provides real-world prior constraints for accurate correction. By avoiding the generation of content that deviates from the original video context, the application prevents issues such as visual inconsistencies and logical errors. For the three types of correction instructions, this application adopts a layered and differentiated correction strategy to respectively complete error attribute correction, missing element completion, and spatiotemporal logic constraint optimization, achieving precise and structured video optimization.
[0116] against The missing element completion command intelligently generates and supplements missing elements in the video. Based on the user's semantic intent and ideal topological graph structure, combined with the scene style, scale, and perspective of the original video, it generates missing elements that match the scene logic within the coordinate positions and movement range specified in the command. The generation process strictly adheres to the overall context of the video, ensuring that the shape, proportion, and dynamic trajectory of the added elements are highly adapted to the original image, thus solving the problems of incomplete video content and missing key information.
[0117] against Error correction commands execute precise replacement and correction operations for video element attributes. Based on the element ID, modification dimension, and target parameters marked in the command, they finely adjust the material, color, style, position, and other attributes of objects in the video that have deviations, frame by frame. The original motion trajectory and spatial topology of the objects are preserved throughout the process. Only erroneous attributes are corrected, without altering the original normal content of the video, thus preserving the authenticity and integrity of the original material to the greatest extent.
[0118] against Spatiotemporal constraint instructions impose global logical constraints on the overall correction and generation process, forcing the generation model to follow the optical flow direction, inter-frame transition rules, and spatial relationships of objects in the original video. This constrains the consistency of motion and spatiotemporal topology compliance between the newly added content and the correction area, effectively avoiding defects such as object suspension, sudden motion changes, scene logic conflicts, and screen flickering after correction.
[0119] To address the issues of disconnect between the corrected areas and the original image, inconsistent styles, and abrupt transitions in traditional generation algorithms, this application introduces a full-dimensional feature consistency constraint during the generation process, ensuring overall video uniformity across visual, temporal, and auditory dimensions. At the visual level, it strictly matches the realistic style, lighting intensity, color tone, and texture details of the original video, guaranteeing that the visual quality of the added and corrected areas is indistinguishable from the original image. At the temporal level, it adheres to the original inter-frame optical flow patterns and motion trajectories, ensuring a smooth and natural rhythm for dynamic elements without abrupt jumps. At the auditory level, it preserves the original video's audio features throughout, achieving a high degree of adaptation between the image correction and the audio rhythm and atmosphere, avoiding audio-visual disconnect.
[0120] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An intent-driven method for material retrieval, correction, and video generation, characterized in that: Includes the following steps: S1: Extract multi-dimensional features from all video materials in the material library to construct a structured material feature library; S2: Receive multimodal instructions from the user, parse the user's creative intent behind the multimodal instructions, and transform them into a unified semantic feature vector; S3: Align multimodal semantic features with video material features based on the semantic feature vector corresponding to the user intent; S4: Perform semantic deviation verification on the retrieved video segments, identify semantic deviations that do not conform to the user's intent, extract relevant contextual features of the segments, and generate correction instructions; S5: Based on the correction instructions, and combined with the contextual features of the retrieved fragments and the structural features of the material library, perform attribute correction on the fragments with semantic deviations and complete the missing elements; S6: Guided by the user's creative intent, it splices and merges the completed and corrected segments or the retrieved complete segments to generate the finished video.
2. The intent-driven material retrieval, correction, and video generation method according to claim 1, characterized in that: In step S1, the extracted features include visual features, audio features, and semantic tag features. The extraction of visual features specifically involves: using the ResNet18 model to extract texture, color, and object features from each frame of the video; simultaneously, using optical flow to extract inter-frame motion features; and fusing the visual and motion features of each frame to obtain the video visual feature vector. ; The audio feature extraction process specifically involves: using MFCC to extract frequency and rhythm features from the video's audio track; extracting the rhythm peaks using a peak detection algorithm; and outputting an audio feature vector. ; Extracting semantic label features specifically involves using the CLIP lightweight model to label video frames with scene, object, and action tags, and then converting these tags into semantic feature vectors. ; Will , and Weighted fusion is performed to obtain the comprehensive feature vector of each video segment. All video clips' comprehensive feature vectors and original material information are stored in the material feature library, and an index is created.
3. The intent-driven material retrieval, correction, and video generation method according to claim 2, characterized in that: Step S2 further includes the following steps: S21: Receive text commands input by the user, encode the text using a lightweight BERT model, remove redundant information, extract core semantics, and output a text semantic feature vector. =[ ], where n is the feature dimension; S22: Receive user-input voice commands, extract voice features using MFCC, convert the voice to text using a speech-to-text model, and then output a semantic feature vector with the same dimension as the text command using the text encoding method from step S21. ; S23: Receive the sketch drawn by the user, use a CNN model to extract the outline and object features of the sketch, and map them into a semantic feature vector of the same dimension as the text instruction. Ultimately Transformed into a unified semantic feature vector .
4. The intent-driven material retrieval, correction, and video generation method according to claim 3, characterized in that: In step S3, the SemAlignNet algorithm is used to accurately align multimodal semantic features with video material features, specifically including a feature mapping module, a cross-modal attention alignment module, and a similarity calculation module. The feature mapping module uses a learnable fully connected layer to... and Mapping to the same feature space, the mapping formula is: ; ; in, and All are learnable weight matrices. and For bias terms; The cross-modal attention alignment module is used to highlight the correspondence between the core semantics in user commands and the key features in video footage, and to calculate the attention weight matrix. , ; in, Here is the attention weight matrix. The function is used to normalize the weights. To transpose video features; The similarity calculation module combines attention weight calculation. and cosine similarity, ; in, To weight the special effects matching score, To normalize the feature length, · It is a norm.
5. The intent-driven material retrieval, correction, and video generation method according to claim 4, characterized in that: Step S4 also includes the following steps: S41: Based on unified semantic feature vector This maps it to a structured set of ideal intent semantics: ; in, As a collection of core concepts, For a set of attribute constraints, It is a spatial relation vector. The feature dimension of the spatial relation vector; S42: Based on extracted video segment features By combining visual recognition algorithms, the actual semantic set of the video clips can be parsed out: ; in, It represents the total number of semantic tuples parsed from the video clip; S43: Quantitative Analysis and Deviation; S44: Construct the material topology graph and the ideal intent topology graph, and calculate the material topology graph. Topology diagram with ideal intention The overall deviation value; S45: Output correction instruction set .
6. The intent-driven material retrieval, correction, and video generation method according to claim 5, characterized in that: Step S43 further includes the following steps: S43.1: Yes Each core concept in exist In performing semantic embedding matching, if exist If there is no match, then ( , , Add to missing semantic set ,like There is a match, but the attributes and Chinese correspondence If the similarity is below the threshold θ, then ( , , ) and incorrect match ( , , ) Forming an error semantic pair set ; S43.2: Quantify the overall deviation using the distance formula. ; in, Weighting the difficulty of completing missing content. As the weight of spatial location deviation, Weights for differences in image attributes. This represents the total semantic deviation value. For concept similarity terms, For missing semantic concepts to be completed, To fill in the ideal semantic concept as expected. For spatial topological distance, For the spatial location topological features of the erroneous video / material, Topological features of the target spatial location for standard ideal materials. For attribute difference items, The distribution of style, lighting, and color attributes of the incorrect material. The attribute distribution of standard ideal materials.
7. The intent-driven material retrieval, correction, and video generation method according to claim 6, characterized in that: In step S44, the material topology map Topology diagram with ideal intention The formula for calculating the overall deviation value is: ; in, This represents the overall topological deviation value. , , These are learnable weight coefficients, corresponding to the weight proportions of node bias, edge relationship bias, and attribute bias, respectively. , For node topology deviation, This represents the topological deviation of the edge relationship. This refers to the topological deviation of the attribute.
8. The intent-driven material retrieval, correction, and video generation method according to claim 7, characterized in that: In step S5, for The specific processing of missing completion instructions is as follows: based on the user's semantic intent and the ideal topological graph structure, combined with the scene style, scale ratio, and perspective relationship of the original video, missing elements matching the scene logic are generated within the coordinate position and motion range specified by the instruction. against The specific processing of error correction instructions is as follows: based on the element ID, modification dimension and target parameters marked in the instruction, the material, color, style and position attributes of objects with deviations in the video are adjusted frame by frame; against The specific processing of the spatiotemporal constraint instructions is as follows: the generated model is forced to follow the optical flow motion direction, inter-frame transition rules and object spatial relationships of the original video, and the motion consistency and spatiotemporal topology compliance of the newly added content and the correction area are constrained.