Time anchor key frame video task scheduling and homing method and system
Patent Information
- Application Number
- CN202610709412.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请旨在解决传统多图或多宫格关键帧生成视频无法精确表达关键帧所在时间点位,以及异步视频结果返回后难以校验关键帧状态和目标视频槽位状态的问题
[0007]本申请将关键帧从普通参考图提升为具有时间锚点的时间线控制对象;能够在视频生成请求之外保留结构化时序元数据;在模型仅支持单图或多图输入时仍可保持关键帧与时间线点位之间的映射关系;通过目标视频槽位版本、关键帧槽位版本、图像摘要、时间锚点和输入载体元数据构成的状态快照或状态指纹,降低关键帧变更后错误覆盖目标视频槽位的可能性;通过保存视频与关键帧时间锚点的映射关系,便于后续时间线编辑和结果比对。
Smart Images

Figure CN122551243A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of generative video processing, keyframe control, timeline editing, asynchronous video task orchestration, and video result consistency and reallocation, and in particular to a method and system for binding keyframe slots to timeline time points in story engineering and using them for video generation task orchestration and result reallocation. Background Technology
[0002] Existing video generation tools already support input of the first frame, last frame, multiple reference images, or keyframes. Some systems can also stitch multiple images together into a grid image or input the model as an array of multiple images. However, simple multi-image input or multi-grid input usually only expresses the visual order between images, making it difficult to accurately define the absolute time point, relative time point, or time proportion in the target video slot for each keyframe on the story engineering timeline. For creative systems that need precise alignment with storyboards, timeline intervals, video slots, and subsequent editing, relying solely on multi-image input cannot guarantee a stable correspondence between keyframes and the timeline.
[0003] Furthermore, keyframe images may be replaced, deleted, or moved during the generation task. If the video result is written directly after the task is completed, it may cause video to overwrite the target slot due to inconsistencies with the current keyframe state. Therefore, a technical solution is needed to uniformly bind keyframe slots, time anchors, input metadata, asynchronous tasks, and result placement. Summary of the Invention Technical problems to be solved
[0004] This application aims to solve the problems that traditional multi-image or multi-grid keyframe video generation cannot accurately express the time position of the keyframe, and that it is difficult to verify the keyframe status and target video slot status after asynchronous video results are returned. Technical solution
[0005] The system identifies target video slots and multiple associated keyframe slots during story engineering. Based on timeline identifiers, storyboard time ranges, or target video slot time ranges, the system assigns time anchors to each keyframe slot. A time anchor can be an absolute time point relative to the global start of the timeline, a relative time point relative to the start of a storyboard, or a proportional value relative to the duration of the target video slot. The system serializes the keyframe slots according to the time anchors and generates keyframe input metadata.
[0006] Keyframe images can be provided to the video generation service in the form of a grid reference map, a multi-image array, or an image sequence. Regardless of the input medium, the system maintains the mapping relationship between keyframe slot identifiers, time anchors, validity flags, empty slot flags, and target video slot identifiers. The system creates asynchronous video generation tasks and binds the tasks to target video slots, keyframe state snapshots at the time of request, keyframe state fingerprints, and keyframe input metadata. The keyframe state snapshot at the time of request records at least the target video slot identifier, target video slot state version, keyframe slot list, each keyframe slot version, keyframe image summary, time anchor, input medium type, and grid layout summary. After the task is completed, the system verifies the current keyframe slot state and the target video slot state. If the verification passes, the result is written to the target video slot; if the verification fails, it is saved as a candidate video result. Beneficial effects
[0007] This application elevates keyframes from ordinary reference graphs to timeline control objects with time anchors; it can retain structured temporal metadata beyond video generation requests; it can maintain the mapping relationship between keyframes and timeline points even when the model only supports single or multiple graph inputs; it reduces the possibility of incorrectly overwriting target video slots after keyframe changes by using a state snapshot or state fingerprint composed of target video slot version, keyframe slot version, image summary, time anchor, and input carrier metadata; and it facilitates subsequent timeline editing and result comparison by saving the mapping relationship between video and keyframe time anchors. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the video task orchestration system for time anchor keyframes in this application. Figure 2 This is a schematic diagram of the keyframe arrangement and video result placement process in this application. Figure 3 This diagram illustrates the relationship between keyframe slots, time anchors, input carriers, and target video slots. Detailed Implementation Example 1: Keyframe Time Anchor Point Allocation
[0009] The system reads a target video slot for a storyboard scene, with a duration of five seconds. The user or system sets three keyframe slots, corresponding to the 0%, 50%, and 100% time positions of the video slot, respectively. The system generates time anchors for the three keyframes and records the keyframe slot identifier, keyframe image identifier, relative time scale, and absolute time point. Example 2: Separation of Grid Input and Metadata
[0010] When the video generation service only supports a single reference image input, the system can stitch multiple keyframe images into a grid reference image. Simultaneously, the system generates grid metadata, recording the keyframe slot identifier and time anchor point for each grid. Even if the model only receives grid images, the system retains the precise mapping between keyframes and timeline points. The same input metadata can be used for multi-image arrays or image sequences. Example 3: Empty Space Keyframe
[0011] When the target video slot requires four temporal positions but the user only provides two keyframes, the system can generate two empty placeholder keyframe units. These empty placeholder keyframe units are marked in the metadata as either not participating in visual constraints or only participating in temporal placeholders, thus ensuring the stability of the input structure and temporal mapping. Example 4: Result Relocation and Conflict Handling
[0012] During the video generation task, if the user replaces an intermediate keyframe or changes the time anchor point of that keyframe, the system will detect an inconsistency between the keyframe state snapshot at the time of the request and the current state upon task return. In this case, the system will not overwrite the target video slot, but will save the generated video as a candidate video result and record the changed keyframe slot and the reason for the anomaly. Example 5: Engineering Data Structure Example
[0013] In one implementation, keyframe control data includes keyframe identifiers, tags, image paths, image summaries, cue text, relative time, relative time scale, location tags, keyframe slot versions, and keyframe roles. The system constrains the relative time range based on the storyboard duration and writes roles such as the first frame, keyframe reference, and last frame into the keyframe input metadata.
[0014] In one implementation, the grid layout metadata includes the number of slots, columns, rows, filled slots, overflow slots, grid index, grid key, labels, image paths, tooltip text, relative time, relative time scale, position labels, and whether a slot is empty. Empty slots are used to maintain the temporal position of the input grid and are not considered valid visual constraints. The target slot for the video generation request is the storyboard video slot. Workflow parameters include the number of results, duration, resolution, frame rate, whether keyframe mode is used, input image path, output path, keyframe control set, and grid layout metadata. The system can generate a status fingerprint based on the target video slot identifier, target video slot status version, each keyframe slot version, keyframe image summary, time anchors, and grid layout summary. Example 6: Implementation of Exception and Conflict Handling
[0015] In one implementation, the system identifies duplicate video generation callbacks using a request identifier or idempotent key. If the generated video has already been written to the target video slot or saved as a candidate video result, subsequent duplicate callbacks will not change the current target video slot. If the task times out, the system releases the processing state of the target video slot and retains a snapshot of the keyframe state at the time of the request; when a late result is returned, the current keyframe slot state and the current target video slot state are still verified. If the task fails or is canceled, the system records the reason for failure or the cancellation status, and subsequent callbacks must not overwrite the target video slot.
[0016] In one implementation, the retry video generation task uses a new request identifier or idempotent key and regenerates the keyframe input metadata, grid layout metadata, and a snapshot of the keyframe state at the time of request. If any keyframe slot is replaced, deleted, moved, the time anchor point changes, the keyframe image summary changes, the input carrier type changes, the grid layout summary changes, or the target video slot state version changes during task execution, the system does not overwrite the target video slot but saves the generated video as a candidate video result. Thus, repeated callbacks, timeouts, failures, cancellations, retries, target object modifications, version inconsistencies, candidate saving, and idempotent processing can all be handled within the keyframe video task orchestration and repositioning chain. Example 7: Material Data Boundaries
[0017] In this application, user-inputted or system-generated text, images, audio, video, and other materials are treated as data objects with identifiers, versions, and states. This application focuses on the technical management methods between keyframe slots, time anchors, input metadata, grid or multi-image mapping, request-time state snapshots, and target video slots. Material source determination or content review can be completed by other processes and are not a necessary component of the technical solution of this application. Electronic devices, storage media and program products implementation
[0018] In one implementation, the aforementioned keyframe time anchor point allocation, input metadata generation, asynchronous video task orchestration, and video result placement are implemented by an electronic device. The electronic device includes at least one processor, a memory, a communication interface, and optional display or input / output interfaces. The memory stores engineering data, state records, referenced entities, task records, and a computer program for executing the aforementioned methods; when the processor reads and executes the computer program, it completes processes such as data object creation, state snapshot or state fingerprint generation, asynchronous task creation, state verification, result writing, or candidate result saving.
[0019] In one embodiment, a computer-readable storage medium stores a computer program. When executed by a processor, the computer program causes the electronic device to perform the aforementioned time-anchor keyframe video task orchestration and repositioning method. The computer-readable storage medium may be a non-volatile memory, a disk, an optical disk, a solid-state memory, a read-only memory, a random access memory, or a data carrier capable of storing program instructions.
[0020] In one implementation, the computer program product includes a computer program, a program code package, an installation package, or an online update package. When the computer program product is downloaded, installed, or executed, it causes the electronic device to perform the aforementioned time-anchor keyframe video task orchestration and repositioning method. The program can run on a client, server, browser, desktop, mobile, edge device, or a combination thereof, and the relevant modules can be implemented by software, hardware, firmware, or a combination thereof.
Claims
1. A method for arranging and repositioning video tasks using time-anchored keyframes, characterized in that, include: In the story engineering process, a target video slot and multiple keyframe slots associated with the target video slot are identified. The target video slot corresponds to a storyboard video slot or an episode video slot. Based on the timeline identifier, storyboard time range, or target video slot time range, time anchors are assigned to the multiple keyframe slots. The time anchors include at least one of absolute time points, relative time points, or relative time ratios. The multiple keyframe slots are serialized according to the time anchors to generate keyframe input metadata. The keyframe input metadata includes a keyframe slot identifier, a keyframe image identifier, a time anchor, a relative time ratio, a keyframe character, a location label, a prompt text, a validity flag, and a target video slot identifier. A video generation request is constructed based on the keyframe input metadata and keyframe images. The keyframe images can be provided to the video generation service in at least one of the following forms: a grid reference map, a multi-image array, or an image sequence. When a grid reference map is used, grid layout metadata is generated synchronously. This metadata includes the number of slots, columns, rows, filled slots, overflow slots, grid index, grid key, position label, whether a slot is empty, and the mapping relationship between these and the keyframe input metadata. An asynchronous video generation task is created, bound to the target video slot, the keyframe state snapshot at the time of request, the keyframe state fingerprint, and the keyframe input metadata. The keyframe state snapshot at the time of request includes at least the target video slot identifier, the target video slot state version, the keyframe slot list, each keyframe slot version, time anchor, and input carrier type. The keyframe state fingerprint is generated at least based on the target video slot identifier, the keyframe slot list, each keyframe slot version, time anchor, and input carrier type. After the asynchronous video generation task is completed, the current keyframe slot state and the current target video slot state are queried. When the current keyframe slot state and the current target video slot state match the keyframe state snapshot and keyframe state fingerprint at the time of request, a video will be generated and written to the target video slot; when either state does not match, the generated video will be saved as a candidate video result and the reason for the exception will be recorded.
2. The method according to claim 1, characterized in that, The time anchor point includes at least two of the following: a millisecond value relative to the start time of the scene, a proportional value relative to the duration of the target video slot, an absolute time value relative to the global start point of the timeline, and a segmentation identifier between adjacent keyframes.
3. The method according to claim 1, characterized in that, The keyframe role includes at least one of the first frame, keyframe reference, last frame, and ordinary reference frame. The system determines the keyframe role based on the time anchor order and keyframe slot position.
4. The method of claim 1, wherein, When the number of keyframes associated with the target video slot is less than the preset number of inputs, the system generates empty placeholder keyframe units and marks the empty placeholder keyframe units in the keyframe input metadata as not participating in visual constraints or only participating in temporal placeholders.
5. The method of claim 1, wherein, The requested keyframe state snapshot includes the target video slot identifier, the target video slot state version, the keyframe slot list, the version of each keyframe slot, the keyframe image summary, the time anchor point, the input carrier type, the grid layout summary, and the keyframe state fingerprint. The keyframe state fingerprint is generated based on at least the target video slot identifier, the keyframe slot list, the version of each keyframe slot, the time anchor point, and the input carrier type.
6. The method of claim 1, wherein, When the system generates a video and writes it into the target video slot, it simultaneously saves the mapping relationship between the generated video and the keyframe time anchor points so that the keyframe constraint points can be displayed in the timeline workbench or subsequent interval editing can be performed.
7. The method of claim 1, wherein, If any keyframe slot is replaced, deleted, or moved before the task is completed, or if its time anchor point changes, the keyframe image summary changes, the input carrier type changes, the grid layout summary changes, or the target video slot status version changes, the system will not overwrite the target video slot, but will save the generated video as a candidate video result.
8. A time anchor keyframe video task orchestration and homing system, characterized by, It includes a target video slot determination module, a keyframe time anchor point allocation module, a keyframe input metadata generation module, a video generation request construction module, an asynchronous video task orchestration module, and a video result placement module; wherein each module is used to perform the method described in any one of claims 1 to 7.
9. An electronic device, comprising: It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.