A dynamic metadata processing method, device, apparatus and storage medium

CN122824940APending Publication Date: 2026-09-25MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611091166.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]随着各类视频创作平台的普及,用户针对 HDR Vivid 视频进行编辑处理后合成视频时,往往会进行添加滤镜、调色、叠加贴纸、插入转场等操作,这些操作通常会导致原视频中的像素的亮度分布和色彩特征值被非线性改写,原始附带的 HDR Vivid 动态元数据在语义上已经失效,进一步导致改写后画面内容与原来的动态元数据无法实现精准匹配

Benefits of technology

[0014]本申请中,在处理目标视频的动态元数据时,获取目标视频中各视频帧对应的帧缓存条目,以得到所述目标视频对应的帧缓存条目集合,并在获取到针对所述目标视频中目标视频帧的视频编辑操作后,基于所述帧缓存条目集合中所述目标视频帧对应的目标帧缓存条目,判断所述视频编辑操作是否已在所述目标视频帧中的目标附着物中生效;所述目标视频为HDR Vivid视频;所述目标帧缓存条目携带有所述目标附着物对应的附着物句柄和所述目标帧缓存条目对应的目标关联存储区;所述附着物句柄指向所述目标视频帧对应的目标附着物;所述目标附着物为所述目标视频帧经特效渲染管线处理后的最终画面;若所述视频编辑操作已在所述目标视频帧中的所述目标附着物中生效,则基于所述目标帧缓存条目从所述帧缓存条目集合中确定目标参照条目,以基于所述目标参照条目确定所述目标视频帧对应的参照视频帧和相应的参照附着物句柄;基于所述目标视频帧对应的所述目标附着物和所述参照附着物句柄指向的参照附着物,对所述目标视频帧对应的原有动态元数据进行重算生成以获取目标动态元数据,并将所述目标动态元数据挂载至所述目标关联存储区;若获取到所述目标视频对应的视频导出请求,则对各所述帧缓存条目对应的关联存储区中保存的动态元数据进行编码封装得到所述目标视频对应的目标导出文件;所述目标导出文件为HDR Vivid视频文件。可见,本申请在获取到针对目标视频中目标视频帧的视频编辑操作后,先基于帧缓存条目集合中该目标视频帧对应的目标帧缓存条目,判断该视频编辑操作是否已经在目标视频帧的目标附着物中生效;由于该目标附着物即为目标视频帧经特效渲染管线处理后的最终画面,因此一旦确认编辑操作已生效,便意味着该帧的亮度分布与色彩极值已经偏离原始帧,其原有动态元数据在语义上已经失效,必须重算而不可沿用。此后,本申请基于目标帧缓存条目从帧缓存条目集合中确定目标参照条目,进而确定参照视频帧及其参照附着物句柄,并以目标附着物与参照附着物为依据对原有动态元数据进行重算生成,最后将所得的目标动态元数据挂载至目标关联存储区。由此带来如下有益效果:其一,由于动态元数据的分析面即为特效已全部生效的最终画面,因此所生成的动态元数据与播放端实际所要映射的画面内容严格对应,预览效果与最终播放效果在创作意图层面得以保持一致;其二,由于参照条目取自编辑引擎为时间线预览天然维护的既有帧缓存条目集合,而非另行开辟前瞻缓存或反复触发解码,因此在不引入额外数据加载延迟、不大幅增加内存开销的前提下,仍可利用前后帧信息完成场景切换检测与时域平滑;其三,由于目标动态元数据被挂载在目标帧缓存条目自身的目标关联存储区中,与帧标识一一对应且与条目同分配同淘汰,无需依赖以时间戳索引的全局元数据表进行跨表查询,因而从结构上消除了多线程流水线中因帧序漂移而导致的元数据与帧错位风险;其四,由于动态元数据的全部计算成本已被摊入编辑阶段,而编辑阶段的各个环节可以并行执行从而尽量屏蔽计算耗时,导出阶段仅需从各关联存储区直通复用预存载荷、无需重新扫描像素内容,因此可大幅提升视频导出速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824940A_ABST
    Figure CN122824940A_ABST
Patent Text Reader

Abstract

The application discloses a dynamic metadata processing method and device, equipment and storage medium, and relates to the field of video editing, and comprises the following steps: acquiring a frame cache entry set of a target video, and after acquiring a video editing operation for a target video frame in the target video, judging whether the video editing operation has taken effect in a target attachment in the target video frame; if the video editing operation has taken effect, determining a target reference entry from the frame cache entry set, determining a reference video frame and a corresponding reference attachment handle based on the target reference entry; recalculating original dynamic metadata based on the target attachment and a reference attachment pointed by the reference attachment handle to generate target dynamic metadata and mounting the target dynamic metadata to a target associated storage area; and if a video export request is acquired, encoding and packaging the dynamic metadata saved in the associated storage area to obtain a target export file. The application generates new dynamic metadata which is accurately matched with rewritten picture content in real time in the video editing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video editing, and in particular to a dynamic metadata processing method, apparatus, device, and storage medium. Background Technology

[0002] HDR Vivid is a Chinese-developed high dynamic range video standard led by the World Ultra HD Video Industry Alliance. This standard specifies that each frame or scene carries dynamic metadata describing the brightness range and distribution characteristics of the current image. Unlike static HDR (such as the fixed static metadata of HDR10), HDR Vivid's dynamic metadata is content-related, calculated for the brightness and color distribution of each frame or scene.

[0003] With the widespread use of various video creation platforms, when users edit and synthesize HDR Vivid videos, they often add filters, adjust colors, overlay stickers, and insert transitions. These operations usually cause the brightness distribution and color feature values ​​of the pixels in the original video to be non-linearly rewritten, and the original HDR Vivid dynamic metadata becomes semantically invalid, further resulting in the rewritten image content not being able to accurately match the original dynamic metadata. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a dynamic metadata processing method, apparatus, device, and storage medium, capable of generating new dynamic metadata in real time during video editing that precisely matches the rewritten screen content. The specific solution is as follows: In a first aspect, this application discloses a dynamic metadata processing method, including: The process involves obtaining frame cache entries corresponding to each video frame in the target video to obtain a set of frame cache entries corresponding to the target video. After obtaining a video editing operation for a target video frame in the target video, based on the target frame cache entries corresponding to the target video frame in the set of frame cache entries, it is determined whether the video editing operation has taken effect on the target attachment in the target video frame. The target video is an HDR Vivid video. Each target frame cache entry carries an attachment handle corresponding to the target attachment and a target associated storage area corresponding to the target frame cache entry. The attachment handle points to the target attachment corresponding to the target video frame. The target attachment is the final image of the target video frame after processing by the special effects rendering pipeline. If the video editing operation has been effective in the target attachment in the target video frame, then a target reference entry is determined from the set of frame cache entries based on the target frame cache entry, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry; Based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain target dynamic metadata, and the target dynamic metadata is mounted to the target associated storage area. If a video export request corresponding to the target video is obtained, the dynamic metadata stored in the associated storage area corresponding to each frame buffer entry is encoded and encapsulated to obtain the target export file corresponding to the target video; the target export file is an HDR Vivid video file.

[0005] Optionally, the step of recalculating and generating target dynamic metadata based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, and then mounting the target dynamic metadata to the target associated storage area, includes: Based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the target brightness distribution information is determined, and the target brightness distribution information is mapped to the target metadata semantic structure to obtain the target dynamic metadata corresponding to the target video frame, so as to write the target dynamic metadata into the target associated storage area; Accordingly, the dynamic metadata processing method further includes: If a timeline trimming operation is detected for the target video, and the target frame cache entry corresponding to the target video frame is evicted from the frame cache entry set after the timeline trimming operation is performed, then the target associated storage area and the target dynamic metadata are released simultaneously.

[0006] Optionally, determining the target reference entry from the set of frame buffer entries based on the target frame buffer entry includes: If the set of frame buffer entries contains a predecessor entry corresponding to the target frame buffer entry, then the predecessor entry is determined as the target reference entry; the predecessor entry is the frame buffer entry corresponding to the predecessor video frame of the target video frame on the time axis.

[0007] Optionally, the dynamic metadata processing method further includes: If there is a successor entry corresponding to the target frame buffer entry in the frame buffer entry set, and the successor attachment corresponding to the successor entry has been processed by the special effects rendering pipeline, then the successor entry is determined as the target reference entry. Accordingly, the step of recalculating and generating the original dynamic metadata corresponding to the target video frame based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle to obtain the target dynamic metadata includes: Based on the target brightness distribution information of the target attachment, the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry, and the subsequent brightness distribution information of the subsequent attachment corresponding to the subsequent entry, scene switching is determined to obtain the corresponding target determination result. Based on the target determination result, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain the target dynamic metadata. Wherein, the successor entry is the frame buffer entry corresponding to the successor video frame on the timeline of the target video frame, the predecessor attachment is the final image of the predecessor video frame after being processed by the special effects rendering pipeline, and the successor attachment is the final image of the successor video frame after being processed by the special effects rendering pipeline.

[0008] Optionally, the step of recalculating and generating target dynamic metadata based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle includes: If there is no successor entry corresponding to the target frame cache entry in the frame cache entry set, or if the successor attachment corresponding to the successor entry of the target frame cache entry has not yet been processed by the special effects rendering pipeline, then a scene switching determination is performed based on the target brightness distribution information of the target attachment and the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry to obtain the corresponding target determination result, and the original dynamic metadata corresponding to the target video frame is recalculated and generated based on the target determination result to obtain the target dynamic metadata.

[0009] Optionally, the dynamic metadata processing method further includes: If a target preview request for the target video frame is obtained, and the preview window corresponding to the target preview request meets the preset HDR rendering conditions, then the tone mapping intent corresponding to the target video frame is obtained based on the target attachment and the target dynamic metadata, so as to generate a first preview screen based on the tone mapping intent using the target compositing link, and present the first preview screen corresponding to the target video frame in the preview window.

[0010] Optionally, the dynamic metadata processing method further includes: If a target preview request for the target video frame is obtained, and the preview window corresponding to the target preview request does not meet the preset HDR rendering conditions, then based on the adaptation information described by the target dynamic metadata, tone mapping pre-folding is performed on the target attachment to obtain a second preview image, and the second preview image is presented in the preview window.

[0011] Secondly, this application discloses a dynamic metadata processing apparatus, comprising: The operation effect judgment module is used to obtain the frame cache entries corresponding to each video frame in the target video to obtain the frame cache entry set corresponding to the target video. After obtaining the video editing operation for the target video frame in the target video, it determines whether the video editing operation has taken effect in the target attachment in the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set. The target video is an HDRVivid video. The target frame cache entry carries the attachment handle corresponding to the target attachment and the target associated storage area corresponding to the target frame cache entry. The attachment handle points to the target attachment corresponding to the target video frame. The target attachment is the final image of the target video frame after processing by the special effects rendering pipeline. The reference information acquisition module is used to determine a target reference entry from the frame cache entry set based on the target frame cache entry if the video editing operation has been effective in the target attachment in the target video frame, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry; The metadata recalculation module is used to recalculate and generate the original dynamic metadata corresponding to the target video frame based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, to obtain the target dynamic metadata, and to mount the target dynamic metadata to the target associated storage area. The video export module is used to encode and encapsulate the dynamic metadata stored in the associated storage area corresponding to each frame cache entry to obtain the target export file corresponding to the target video if a video export request corresponding to the target video is obtained; the target export file is an HDR Vivid video file.

[0012] Thirdly, this application discloses an electronic device, including: Memory is used to store computer programs; A processor is used to execute the computer program to implement the aforementioned dynamic metadata processing method.

[0013] Fourthly, this application discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned dynamic metadata processing method.

[0014] In this application, when processing the dynamic metadata of the target video, frame cache entries corresponding to each video frame in the target video are obtained to obtain a set of frame cache entries corresponding to the target video. After obtaining a video editing operation for a target video frame in the target video, based on the target frame cache entries corresponding to the target video frame in the set of frame cache entries, it is determined whether the video editing operation has taken effect on the target attachment in the target video frame; the target video is HDR. Vivid video; the target frame cache entry carries the attachment handle corresponding to the target attachment and the target associated storage area corresponding to the target frame cache entry; the attachment handle points to the target attachment corresponding to the target video frame; the target attachment is the final image of the target video frame after processing by the special effects rendering pipeline; if the video editing operation has been effective in the target attachment in the target video frame, then a target reference entry is determined from the set of frame cache entries based on the target frame cache entry, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry; based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain target dynamic metadata, and the target dynamic metadata is mounted to the target associated storage area; if a video export request corresponding to the target video is obtained, then the dynamic metadata stored in the associated storage area corresponding to each frame cache entry is encoded and encapsulated to obtain the target export file corresponding to the target video; the target export file is an HDR Vivid video file. As can be seen, after obtaining the video editing operation for the target video frame in the target video, this application first determines whether the video editing operation has taken effect in the target attachment of the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set. Since the target attachment is the final image of the target video frame after processing by the special effects rendering pipeline, once it is confirmed that the editing operation has taken effect, it means that the brightness distribution and color extreme values ​​of the frame have deviated from the original frame, and its original dynamic metadata has become semantically invalid and must be recalculated and cannot be reused. Subsequently, this application determines the target reference entry from the frame cache entry set based on the target frame cache entry, and then determines the reference video frame and its reference attachment handle. Based on the target attachment and the reference attachment, the original dynamic metadata is recalculated and generated. Finally, the obtained target dynamic metadata is mounted to the target associated storage area.This brings the following beneficial effects: First, since the analysis surface of dynamic metadata is the final screen with all effects in effect, the generated dynamic metadata strictly corresponds to the actual screen content to be mapped by the playback end, ensuring consistency between the preview effect and the final playback effect at the creative intent level; Second, since the reference entries are taken from the existing frame cache entry set naturally maintained by the editing engine for timeline preview, rather than creating a separate lookahead cache or repeatedly triggering decoding, scene transition detection and temporal smoothing can still be completed using the information from preceding and following frames without introducing additional data loading delays or significantly increasing memory overhead; Third, because The target dynamic metadata is mounted in the target associated storage area of ​​the target frame cache entry itself, corresponding one-to-one with the frame identifier and allocated and evicted together with the entry. It does not need to rely on the global metadata table indexed by timestamp for cross-table queries, thus structurally eliminating the risk of metadata and frame misalignment caused by frame order drift in multi-threaded pipelines. Fourth, since the entire computational cost of dynamic metadata has been amortized into the editing stage, and the various stages of the editing stage can be executed in parallel to minimize computational time, the export stage only needs to directly reuse the pre-stored payload from each associated storage area without rescanning pixel content, thus greatly improving the video export speed. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a flowchart of a dynamic metadata processing method disclosed in this application; Figure 2 This is a schematic diagram of the overall data flow in a specific dynamic metadata processing process disclosed in this application; Figure 3 This is a schematic diagram of a specific dual-path generation architecture for preview screens disclosed in this application; Figure 4 This is a schematic diagram of a specific video file export and reuse process disclosed in this application; Figure 5 This is a schematic diagram of the structure of a dynamic metadata processing device disclosed in this application; Figure 6 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] HDR Vivid is a Chinese-developed high dynamic range video standard led by the World Ultra HD Video Industry Alliance. This standard specifies that each frame or scene carries dynamic metadata describing the brightness range and distribution characteristics of the current image. Unlike static HDR (such as the fixed static metadata of HDR10), HDR Vivid's dynamic metadata is content-related, calculated for the brightness and color distribution of each frame or scene. With the widespread use of various video creation platforms, when users edit and synthesize HDR Vivid videos, they often add filters, adjust colors, overlay stickers, and insert transitions. These operations typically cause the brightness distribution and color feature values ​​of pixels in the original video to be non-linearly rewritten, rendering the original HDR Vivid dynamic metadata semantically invalid. This further leads to a mismatch between the rewritten image content and the original dynamic metadata. To address these technical problems, this application discloses a dynamic metadata processing method that can generate new dynamic metadata that precisely matches the rewritten image content in real time during video editing.

[0019] See Figure 1 As shown, this embodiment of the invention discloses a dynamic metadata processing method, including: Step S11: Obtain the frame cache entries corresponding to each video frame in the target video to obtain the frame cache entry set corresponding to the target video. After obtaining the video editing operation for the target video frame in the target video, determine whether the video editing operation has taken effect in the target attachment in the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set. The target video is an HDR Vivid video. The target frame cache entry carries the attachment handle corresponding to the target attachment and the target associated storage area corresponding to the target frame cache entry. The attachment handle points to the target attachment corresponding to the target video frame. The target attachment is the final image of the target video frame after processing by the special effects rendering pipeline.

[0020] In this embodiment, the set of frame cache entries is not a newly created independent analysis cache for this application, but rather a set of recently used frame cache entries maintained in memory by the short video editing SDK (Software Development Kit) engine for purposes such as timeline preview. That is, this embodiment reuses the existing frame cache structure of the editing engine, adding an associated storage area to the existing entry structure to store the generated HDR Vivid dynamic metadata payload. This associated storage area lives and dies along with the entries. Compared to building a separate global metadata table indexed by timestamps, this approach neither increases memory usage nor triggers additional decoding, thus not compromising the real-time performance required by the editing scenario. The frame buffer entries maintained by the editing SDK engine carry at least the following three items: First, a frame identifier ID / timestamp, used to uniquely identify the position of the video frame corresponding to the entry on the timeline; second, an attachment handle pointing to the final image attachment of the GPU output, which is the render-to-texture handle of the FBO (Frame Buffer Object) color attachment, and the target attachment it points to is equivalent to "the final image with all effects applied"; third, an optional associated storage area, which exists with the allocation of the entry and is destroyed when the entry is evicted, and is not maintained externally in a global metadata table indexed by timestamps. It should be noted that the "attachment" mentioned here refers to the image data residing on the off-screen rendering target. The image of each entry maintained by the editing engine for timeline interaction has been submitted by the GPU rendering pipeline and resides on the off-screen rendering target. This step obtains the access point of the attachment and uses it as the input surface of the dynamic metadata recalculation module. This embodiment neither creates a new full-size decoded pixel copy nor starts a separate analysis surface from the original decoded frame domain, thereby avoiding the problem of data format conflict between the compressed rendering surface and the analysis surface in the editor frame buffer.

[0021] Regarding the method for determining whether video editing operations have taken effect in the target attachment, this embodiment provides the following specific implementation: When a user operation (e.g., changes to filter parameters, color grading parameters, sticker parameters, or transition parameters) causes a certain interval on the timeline to be marked as "needs redrawing," the effects rendering pipeline will resubmit the rendering of the entries within that interval. After rendering is complete, the target attachment of the corresponding entry is updated, and the entry will be marked with a semantic tag of "attachment overwritten / dirty." This tag can be implemented through a rendering completion callback or a dirty flag. Based on this, the system can determine that the old HDR Vivid dynamic metadata that may have existed on the entry no longer corresponds to this final image, and therefore should enter the recalculation process. Conversely, if the entry is hit but its attachment has not been overwritten, the original payload on it can be directly used to avoid meaningless duplicate statistics, thereby further saving computing power. In other words, this step structurally confirms the state of "old metadata semantic invalidation"—at this point, the brightness distribution and color extreme values ​​of the frame have deviated from the original frame (the effect rewriting involves non-linear compositing operations), and the original accompanying HDR Vivid dynamic metadata no longer corresponds to this image, so it needs to be recalculated instead of being reused.

[0022] Step S12: If the video editing operation has been effective in the target attachment in the target video frame, then a target reference entry is determined from the set of frame cache entries based on the target frame cache entry, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry.

[0023] In this embodiment, as Figure 2As shown, using the current target frame buffer entry as an anchor, adjacent rendered entries are taken from the frame buffer entry set along the time axis to form the analysis reference range. It is important to emphasize that these reference entries are associated with their respective final image attachments, not their respective original decoded frame planes; therefore, the analysis scope is consistent throughout the entire reference range. In one specific implementation, if a preceding entry corresponding to the target frame buffer entry exists in the frame buffer entry set, the preceding entry is determined as the target reference entry; the preceding entry is the frame buffer entry corresponding to the preceding video frame on the time axis of the target video frame. That is, the analysis reference range should at least include the current frame and the preceding frame. In another specific implementation, if a successor entry corresponding to the target frame cache entry exists in the frame cache entry set, and the successor attachment corresponding to the successor entry has been processed by the effects rendering pipeline, then the successor entry is determined as the target reference entry. Here, the successor entry is the frame cache entry corresponding to the successor video frame on the timeline of the target video frame, and the successor attachment is the final image of the successor video frame after processing by the effects rendering pipeline. This situation typically occurs when the user performs pre-scrolling on the timeline, or when the corresponding entry has already been cached and hit. In this case, the analysis reference range can be expanded from "current frame + predecessor frame" to "predecessor frame + current frame + successor frame," thereby enabling more accurate scene switching determination using information from future frames. It is understood that this embodiment, through this method of "using existing rendered entries," achieves capabilities equivalent to prospective analysis without adding a new look-ahead buffer or repeatedly triggering decoding, thus resolving the contradiction between prospective analysis and real-time performance.

[0024] In another specific implementation, if there is no successor entry corresponding to the target frame cache entry in the frame cache entry set, or if there is a successor entry but its corresponding subsequent attachment has not yet been processed by the effects rendering pipeline, the analysis reference range degenerates to include only the current frame and the preceding frame. Furthermore, if the preceding entry also does not exist (e.g., the target video frame is the first frame of the timeline, or its preceding content has been cropped), the analysis reference range further degenerates to a single-frame range containing only the current frame. In this case, the temporal recursive history can be initialized as agreed, and after the first valid statistic appears, it enters the normal continuous segment logic. This design ensures that the "switch reset / continuous smoothing" mechanism can be activated at any point on the timeline without failing due to the lack of a reference frame. Furthermore, those skilled in the art will understand that the specific size of the analysis reference range can also be dynamically adjusted according to the device capability strategy, expanding to a larger window when the device has sufficient computing power and shrinking to a smaller window when the device's computing power is limited.

[0025] Step S13: Based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, recalculate and generate the original dynamic metadata corresponding to the target video frame to obtain target dynamic metadata, and mount the target dynamic metadata to the target associated storage area.

[0026] In this embodiment, based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain target dynamic metadata, and the target dynamic metadata is mounted to the target associated storage area. Specifically, this may include: determining the target brightness distribution information based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, and mapping the target brightness distribution information to the target metadata semantic structure to obtain the target dynamic metadata corresponding to the target video frame, so as to write the target dynamic metadata into the target associated storage area.

[0027] Regarding the acquisition of target brightness distribution information, the dynamic metadata real-time generation engine can directly obtain brightness-related information from the aforementioned attachments. For example, it can statistically analyze the distribution of the maxRGB values ​​of the final image through GPU computational shaders, downsampling mipmap chains, or histogram statistics, thereby avoiding the bandwidth overhead and synchronization wait caused by reading back the entire frame's pixels to the CPU. It should be noted that this embodiment does not improve the specific calculation method of the aforementioned brightness statistics themselves. What this embodiment ensures is the correct attribution between "what load is generated" and "from which frame is it generated," that is, the load must be generated from the final image where all special effects have been applied, and must belong to the frame from which it was generated.

[0028] In this embodiment, regarding scene switching determination, the dynamic metadata real-time generation engine simultaneously utilizes brightness-related information from adjacent frames to differentiate between continuous shot segments and scene switching boundaries. Specifically, based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain target dynamic metadata. This may include: determining scene switching based on the target brightness distribution information of the target attachment, the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry, and the subsequent brightness distribution information of the subsequent attachment corresponding to the subsequent entry to obtain the corresponding target determination result, and recalculating and generating the original dynamic metadata corresponding to the target video frame based on the target determination result to obtain target dynamic metadata; wherein, the subsequent entry is the frame buffer entry corresponding to the subsequent video frame on the timeline of the target video frame, the preceding attachment is the final image of the preceding video frame after processing by the special effects rendering pipeline, and the subsequent attachment is the final image of the subsequent video frame after processing by the special effects rendering pipeline. If there is no successor entry corresponding to the target frame cache entry in the frame cache entry set, or if the successor attachment corresponding to the successor entry of the target frame cache entry has not yet been processed by the special effects rendering pipeline, then scene switching determination is performed based on the target brightness distribution information of the target attachment and the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry to obtain the corresponding target determination result, and the original dynamic metadata corresponding to the target video frame is recalculated and generated based on the target determination result to obtain the target dynamic metadata.

[0029] In other words, when the analysis reference range includes both preceding and succeeding entries, scene switching can be determined based on the target brightness distribution information of the target attachment, the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry, and the succeeding brightness distribution information of the succeeding attachment corresponding to the succeeding entry to obtain the corresponding target determination result. When the analysis reference range only includes the preceding entry, scene switching can be determined based on the target brightness distribution information of the target attachment and the preceding brightness distribution information to obtain the corresponding target determination result. Specifically, the difference metric between the brightness distribution information of adjacent frames (e.g., maximum component difference, average component difference, or histogram distance) can be calculated and compared with a preset threshold. If the difference metric exceeds the preset threshold, the current position is determined to be at the scene switching boundary. At this time, the temporal recursive history should be cleared and dynamic metadata should be output independently from the current frame to avoid incorrect smoothing across shots. If the difference metric does not exceed the preset threshold, the current position is determined to be within a continuous shot segment. At this time, temporal recursive smoothing is performed on the continuous segment to avoid brightness breathing and flickering caused by frame-by-frame statistical jitter within the same shot. When the analysis reference range includes both predecessor and successor frames, the judgment result can be further refined. For example, if the difference between the predecessor and the current frame is significant, while the difference between the current frame and the successor frame is subtle, the current frame can be determined as the first frame of a new scene; conversely, if the difference between the predecessor and the current frame is subtle, while the difference between the current frame and the successor frame is significant, the current frame can be determined as the last frame of an old scene, thus making the positioning of the switching boundary more accurate. Subsequently, based on the target judgment result, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain the target dynamic metadata.

[0030] In this embodiment, the aforementioned luminance distribution information is mapped to the dynamic metadata semantic structure (i.e., the target metadata semantic structure) specified in T / UWA 005.1. For example, system_start_code=0x01 and num_windows=1 are set, and the luminance distribution related field set (including but not limited to minimum_maxrgb, average_maxrgb, variance_maxrgb, maximum_maxrgb, etc.) and the corresponding tone mapping control identifier and parameter area of ​​the processing window are filled in, thereby obtaining a complete payload that can be directly written into the bitstream.

[0031] It is understood that in this embodiment, the newly generated payload is written into the target associated storage area of ​​the current target frame cache entry, and the payload version / verification flag is strictly aligned with the frame identifier ID, thereby achieving a one-to-one correspondence between the payload and the frame identifier. This associated storage area is allocated, evicted, and released together with the entry, without relying on any global timestamp metadata table for jump queries. The direct effect of this is that even if the timeline is not continuous, the metadata of frame A will never be loaded into the picture of frame B; even if the preview side or the encoding side reads the entry asynchronously, what it can read is only "its own copy", and there is no misaligned path in the system of "taking the timestamp to another table to retrieve someone else's payload". This is also the key to the structural elimination of the risk of metadata and frame misalignment in multi-threaded pipelines in this application. Accordingly, the dynamic metadata processing method of this embodiment also includes: if a timeline trimming operation for the target video is detected, and the target frame cache entry corresponding to the target video frame is evicted from the frame cache entry set after the timeline trimming operation is performed, then the target associated storage area and the target dynamic metadata are released synchronously. In other words, if subsequent timeline pruning causes the entry to be slide-out, the payload will disappear and will not become isolated data hanging in the global table, nor will it be mistakenly retrieved by subsequent frames.

[0032] Furthermore, such as Figure 3 As shown, this embodiment discloses a specific implementation of preview secondary rendering based on the aforementioned steps, including: if a target preview request for a target video frame is obtained, and the preview window corresponding to the target preview request meets the preset HDR rendering conditions, then the tone mapping intent corresponding to the target video frame is obtained based on the target attachment and target dynamic metadata, so as to generate a first preview image based on the tone mapping intent using the target compositing chain, and the first preview image corresponding to the target video frame is presented in the preview window. If a target preview request for a target video frame is obtained, and the preview window corresponding to the target preview request does not meet the preset HDR rendering conditions, then the tone mapping pre-folding is performed on the target attachment based on the adaptation information described by the target dynamic metadata to obtain a second preview image, and the second preview image is presented in the preview window. The preset HDR rendering conditions refer to the display subsystem confirming that it can submit an HDR intent surface, such as the HDR swapchain being available and the color space contract being satisfied. When this condition is met, the system maps the final image attachment of the same entry to the mapping intent expressed by the payload of that entry. Figure 1The image is then submitted to the target compositing pipeline, where the system renders it based on the known HDR intent. This is the "HDR surface penetration" path. It's important to note that neither of these preview paths overwrites the HDR encoding value range of the final image. The preview stage only changes "how it's presented," not "what's stored." Because of this, the creative intent reflected in the preview window closely matches the adaptation results made by the playback end using the same set of dynamic metadata. This achieves consistency between the preview and the final playback at the creative intent level, solving the problem of what you see is not what you get. Furthermore, since the HDR encoding value range of the final image is not contaminated by the preview stage, the subsequent export still uses the final image and its associated payload.

[0033] Step S14: If a video export request corresponding to the target video is obtained, the dynamic metadata stored in the associated storage area corresponding to each frame buffer entry is encoded and encapsulated to obtain the target export file corresponding to the target video; the target export file is an HDR Vivid video file.

[0034] In this embodiment, as Figure 4 As shown, when the user clicks "Publish," the encoder processes each frame according to the ordered frame reference sequence already determined in the timeline: for each frame, it retrieves its corresponding frame buffer entry, reads the pre-stored dynamic metadata payload from the same associated storage area of ​​that entry, associates the payload with the corresponding encoded frame, and writes the supplemental enhancement information (SEI) of the bitstream (e.g., user data registration SEI according to ITU-T T.35 encapsulation, terminal provider code 0x26, and terminal provider orientation code 0x0004) into the bitstream. Then, it completes the HDRVivid identifier encapsulation according to the container layer requirements (e.g., writing the corresponding color primary colors, transmission characteristics, matrix coefficients, and other color description information, as well as the HDR Vivid format identifier of the container layer), and finally outputs a compliant HDR Vivid video file. Since the dynamic metadata is calculated and bound during the editing period, the export stage does not rescan pixel content to calculate dynamic metadata, thus significantly reducing export time.

[0035] As can be seen, after obtaining the video editing operation for the target video frame in the target video, this application first determines whether the video editing operation has taken effect in the target attachment of the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set. Since the target attachment is the final image of the target video frame after processing by the special effects rendering pipeline, once it is confirmed that the editing operation has taken effect, it means that the brightness distribution and color extreme values ​​of the frame have deviated from the original frame, and its original dynamic metadata has become semantically invalid and must be recalculated and cannot be reused. Subsequently, this application determines the target reference entry from the frame cache entry set based on the target frame cache entry, and then determines the reference video frame and its reference attachment handle. Based on the target attachment and the reference attachment, the original dynamic metadata is recalculated and generated. Finally, the obtained target dynamic metadata is mounted to the target associated storage area. This brings the following beneficial effects: First, since the analysis surface of dynamic metadata is the final screen with all effects in effect, the generated dynamic metadata strictly corresponds to the actual screen content to be mapped by the playback end, ensuring consistency between the preview effect and the final playback effect at the creative intent level; Second, since the reference entries are taken from the existing frame cache entry set naturally maintained by the editing engine for timeline preview, rather than creating a separate lookahead cache or repeatedly triggering decoding, scene transition detection and temporal smoothing can still be completed using the information from preceding and following frames without introducing additional data loading delays or significantly increasing memory overhead; Third, because The target dynamic metadata is mounted in the target associated storage area of ​​the target frame cache entry itself, corresponding one-to-one with the frame identifier and allocated and evicted together with the entry. It does not need to rely on the global metadata table indexed by timestamp for cross-table queries, thus structurally eliminating the risk of metadata and frame misalignment caused by frame order drift in multi-threaded pipelines. Fourth, since the entire computational cost of dynamic metadata has been amortized into the editing stage, and the various stages of the editing stage can be executed in parallel to minimize computational time, the export stage only needs to directly reuse the pre-stored payload from each associated storage area without rescanning pixel content, thus greatly improving the video export speed.

[0036] See Figure 5 As shown, this application discloses a dynamic metadata processing apparatus, comprising: The operation effect judgment module 11 is used to obtain the frame cache entries corresponding to each video frame in the target video to obtain the frame cache entry set corresponding to the target video, and after obtaining the video editing operation for the target video frame in the target video, it determines whether the video editing operation has taken effect in the target attachment in the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set; the target video is an HDRVivid video; the target frame cache entry carries the attachment handle corresponding to the target attachment and the target associated storage area corresponding to the target frame cache entry; the attachment handle points to the target attachment corresponding to the target video frame; the target attachment is the final image of the target video frame after processing by the special effects rendering pipeline; The reference information acquisition module 12 is used to determine a target reference entry from the frame cache entry set based on the target frame cache entry if the video editing operation has been effective in the target attachment in the target video frame, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry; The metadata recalculation module 13 is used to recalculate and generate the original dynamic metadata corresponding to the target video frame based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle to obtain target dynamic metadata, and to mount the target dynamic metadata to the target associated storage area. The video export module 14 is used to encode and encapsulate the dynamic metadata stored in the associated storage area corresponding to each frame buffer entry to obtain the target export file corresponding to the target video if a video export request corresponding to the target video is obtained; the target export file is an HDR Vivid video file.

[0037] As can be seen, after obtaining the video editing operation for the target video frame in the target video, this application first determines whether the video editing operation has taken effect in the target attachment of the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set. Since the target attachment is the final image of the target video frame after processing by the special effects rendering pipeline, once it is confirmed that the editing operation has taken effect, it means that the brightness distribution and color extreme values ​​of the frame have deviated from the original frame, and its original dynamic metadata has become semantically invalid and must be recalculated and cannot be reused. Subsequently, this application determines the target reference entry from the frame cache entry set based on the target frame cache entry, and then determines the reference video frame and its reference attachment handle. Based on the target attachment and the reference attachment, the original dynamic metadata is recalculated and generated. Finally, the obtained target dynamic metadata is mounted to the target associated storage area. This brings the following beneficial effects: First, since the analysis surface of dynamic metadata is the final screen with all effects in effect, the generated dynamic metadata strictly corresponds to the actual screen content to be mapped by the playback end, ensuring consistency between the preview effect and the final playback effect at the creative intent level; Second, since the reference entries are taken from the existing frame cache entry set naturally maintained by the editing engine for timeline preview, rather than creating a separate lookahead cache or repeatedly triggering decoding, scene transition detection and temporal smoothing can still be completed using the information from preceding and following frames without introducing additional data loading delays or significantly increasing memory overhead; Third, because The target dynamic metadata is mounted in the target associated storage area of ​​the target frame cache entry itself, corresponding one-to-one with the frame identifier and allocated and evicted together with the entry. It does not need to rely on the global metadata table indexed by timestamp for cross-table queries, thus structurally eliminating the risk of metadata and frame misalignment caused by frame order drift in multi-threaded pipelines. Fourth, since the entire computational cost of dynamic metadata has been amortized into the editing stage, and the various stages of the editing stage can be executed in parallel to minimize computational time, the export stage only needs to directly reuse the pre-stored payload from each associated storage area without rescanning pixel content, thus greatly improving the video export speed.

[0038] In one specific implementation, the metadata recalculation module 13 may include: Metadata storage unit is used to determine target brightness distribution information based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, and to map the target brightness distribution information to the target metadata semantic structure to obtain the target dynamic metadata corresponding to the target video frame, so as to write the target dynamic metadata into the target associated storage area; Accordingly, the device may further include: The storage area release module is used to simultaneously release the target associated storage area and the target dynamic metadata if a timeline trimming operation for the target video is detected, and the target frame cache entry corresponding to the target video frame is evicted from the frame cache entry set after the timeline trimming operation is performed.

[0039] In one specific implementation, the reference information acquisition module 12 may include: The first reference entry determination unit is used to determine the predecessor entry as the target reference entry if there is a predecessor entry corresponding to the target frame buffer entry in the frame buffer entry set; the predecessor entry is the frame buffer entry corresponding to the predecessor video frame of the target video frame on the time axis.

[0040] In one specific embodiment, the device may further include: The second reference entry determination module is used to determine the successor entry as the target reference entry if there is a successor entry corresponding to the target frame buffer entry in the frame buffer entry set, and the successor attachment corresponding to the successor entry has been processed by the special effects rendering pipeline. Accordingly, the metadata recalculation module 13 may specifically include: The first metadata recalculation unit is used to perform scene switching determination based on the target brightness distribution information of the target attachment, the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry, and the following brightness distribution information of the following attachment corresponding to the following entry to obtain the corresponding target determination result, and to recalculate and generate the original dynamic metadata corresponding to the target video frame based on the target determination result to obtain the target dynamic metadata. Wherein, the successor entry is the frame buffer entry corresponding to the successor video frame on the timeline of the target video frame, the predecessor attachment is the final image of the predecessor video frame after being processed by the special effects rendering pipeline, and the successor attachment is the final image of the successor video frame after being processed by the special effects rendering pipeline.

[0041] In one specific implementation, the metadata recalculation module 13 may include: The second metadata recalculation unit is used to determine the scene switching based on the target brightness distribution information of the target attachment and the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry if there is no successor entry corresponding to the target frame cache entry in the frame cache entry set, or if the successor attachment corresponding to the successor entry of the target frame cache entry has not been processed by the special effects rendering pipeline. The unit then performs scene switching determination based on the target brightness distribution information of the target attachment and the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry to obtain the corresponding target determination result. Based on the target determination result, the unit then recalculates and generates the original dynamic metadata corresponding to the target video frame to obtain the target dynamic metadata.

[0042] In one specific embodiment, the device may further include: The first preview screen generation module is used to obtain the tone mapping intent corresponding to the target video frame based on the target attachment and the target dynamic metadata if a target preview request for the target video frame is obtained and the preview window corresponding to the target preview request meets the preset HDR rendering conditions, so as to generate a first preview screen based on the tone mapping intent using the target compositing link, and present the first preview screen corresponding to the target video frame in the preview window.

[0043] In one specific embodiment, the device may further include: The second preview image generation module is used to obtain a second preview image by performing tone mapping pre-folding on the target attachment based on the adaptation information described by the target dynamic metadata if a target preview request for the target video frame is obtained and the preview window corresponding to the target preview request does not meet the preset HDR rendering conditions, and then presenting the second preview image in the preview window.

[0044] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0045] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the dynamic metadata processing method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0046] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0047] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0048] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the dynamic metadata processing method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0049] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned dynamic metadata processing method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0050] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0051] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0052] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0053] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0054] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A dynamic metadata processing method, characterized in that, include: The process involves obtaining frame cache entries corresponding to each video frame in the target video to obtain a set of frame cache entries corresponding to the target video. After obtaining a video editing operation for a target video frame in the target video, based on the target frame cache entries corresponding to the target video frame in the set of frame cache entries, it is determined whether the video editing operation has taken effect on the target attachment in the target video frame. The target video is an HDR Vivid video. Each target frame cache entry carries an attachment handle corresponding to the target attachment and a target associated storage area corresponding to the target frame cache entry. The attachment handle points to the target attachment corresponding to the target video frame. The target attachment is the final image of the target video frame after processing by the special effects rendering pipeline. If the video editing operation has been effective in the target attachment in the target video frame, then a target reference entry is determined from the set of frame cache entries based on the target frame cache entry, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry; Based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain target dynamic metadata, and the target dynamic metadata is mounted to the target associated storage area. If a video export request corresponding to the target video is obtained, the dynamic metadata stored in the associated storage area corresponding to each frame buffer entry is encoded and encapsulated to obtain the target export file corresponding to the target video. The target exported file is an HDR Vivid video file.

2. The dynamic metadata processing method according to claim 1, characterized in that, The step of recalculating and generating target dynamic metadata based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, and then mounting the target dynamic metadata to the target associated storage area, includes: Based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, the target brightness distribution information is determined, and the target brightness distribution information is mapped to the target metadata semantic structure to obtain the target dynamic metadata corresponding to the target video frame, so as to write the target dynamic metadata into the target associated storage area; Accordingly, the dynamic metadata processing method further includes: If a timeline trimming operation is detected for the target video, and the target frame cache entry corresponding to the target video frame is evicted from the frame cache entry set after the timeline trimming operation is performed, then the target associated storage area and the target dynamic metadata are released simultaneously.

3. The dynamic metadata processing method according to claim 1, characterized in that, The step of determining the target reference entry from the set of frame buffer entries based on the target frame buffer entry includes: If the set of frame buffer entries contains a predecessor entry corresponding to the target frame buffer entry, then the predecessor entry is determined as the target reference entry; the predecessor entry is the frame buffer entry corresponding to the predecessor video frame of the target video frame on the time axis.

4. The dynamic metadata processing method according to claim 3, characterized in that, Also includes: If there is a successor entry corresponding to the target frame buffer entry in the frame buffer entry set, and the successor attachment corresponding to the successor entry has been processed by the special effects rendering pipeline, then the successor entry is determined as the target reference entry. Accordingly, the step of recalculating and generating the original dynamic metadata corresponding to the target video frame based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle to obtain the target dynamic metadata includes: Based on the target brightness distribution information of the target attachment, the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry, and the subsequent brightness distribution information of the subsequent attachment corresponding to the subsequent entry, scene switching is determined to obtain the corresponding target determination result. Based on the target determination result, the original dynamic metadata corresponding to the target video frame is recalculated and generated to obtain the target dynamic metadata. Wherein, the successor entry is the frame buffer entry corresponding to the successor video frame on the timeline of the target video frame, the predecessor attachment is the final image of the predecessor video frame after being processed by the special effects rendering pipeline, and the successor attachment is the final image of the successor video frame after being processed by the special effects rendering pipeline.

5. The dynamic metadata processing method according to claim 3, characterized in that, The step of recalculating and generating target dynamic metadata based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle includes: If there is no successor entry corresponding to the target frame cache entry in the frame cache entry set, or if the successor attachment corresponding to the successor entry of the target frame cache entry has not yet been processed by the special effects rendering pipeline, then a scene switching determination is performed based on the target brightness distribution information of the target attachment and the preceding brightness distribution information of the preceding attachment corresponding to the preceding entry to obtain the corresponding target determination result, and the original dynamic metadata corresponding to the target video frame is recalculated and generated based on the target determination result to obtain the target dynamic metadata.

6. The dynamic metadata processing method according to claim 1, characterized in that, Also includes: If a target preview request for the target video frame is obtained, and the preview window corresponding to the target preview request meets the preset HDR rendering conditions, then the tone mapping intent corresponding to the target video frame is obtained based on the target attachment and the target dynamic metadata, so as to generate a first preview screen based on the tone mapping intent using the target compositing link, and present the first preview screen corresponding to the target video frame in the preview window.

7. The dynamic metadata processing method according to claim 1, characterized in that, Also includes: If a target preview request for the target video frame is obtained, and the preview window corresponding to the target preview request does not meet the preset HDR rendering conditions, then based on the adaptation information described by the target dynamic metadata, tone mapping pre-folding is performed on the target attachment to obtain a second preview image, and the second preview image is presented in the preview window.

8. A dynamic metadata processing device, characterized in that, include: The operation effect judgment module is used to obtain the frame cache entries corresponding to each video frame in the target video to obtain the frame cache entry set corresponding to the target video. After obtaining the video editing operation for the target video frame in the target video, it determines whether the video editing operation has taken effect in the target attachment in the target video frame based on the target frame cache entry corresponding to the target video frame in the frame cache entry set. The target video is an HDR Vivid video. The target frame cache entry carries the attachment handle corresponding to the target attachment and the target associated storage area corresponding to the target frame cache entry. The attachment handle points to the target attachment corresponding to the target video frame. The target attachment is the final image of the target video frame after processing by the special effects rendering pipeline. The reference information acquisition module is used to determine a target reference entry from the frame cache entry set based on the target frame cache entry if the video editing operation has been effective in the target attachment in the target video frame, so as to determine the reference video frame and the corresponding reference attachment handle corresponding to the target video frame based on the target reference entry; The metadata recalculation module is used to recalculate and generate the original dynamic metadata corresponding to the target video frame based on the target attachment corresponding to the target video frame and the reference attachment pointed to by the reference attachment handle, to obtain the target dynamic metadata, and to mount the target dynamic metadata to the target associated storage area. The video export module is used to encode and encapsulate the dynamic metadata stored in the associated storage area corresponding to each frame cache entry to obtain the target export file corresponding to the target video if a video export request corresponding to the target video is obtained. The target exported file is an HDR Vivid video file.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the dynamic metadata processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer programs, wherein the computer programs, when executed by a processor, implement the dynamic metadata processing method as described in any one of claims 1 to 7.