A video marking method, a video playing method, a video marking system and a device

By identifying and calibrating the target coordinates in drone aerial videos, a video subtitle file of the flight path timeline is generated, solving the problem of insufficient matching accuracy between marker points and video frames, and realizing accurate association between marker points and flight path trajectories and efficient data analysis.

CN122138015APending Publication Date: 2026-06-02XIAN XUANJI ZHIHANG TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN XUANJI ZHIHANG TECHNOLOGY CO LTD
Filing Date
2026-02-12
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing aerial video marking and trajectory display systems, the matching accuracy between marker points and video frames is insufficient, causing marker points to shift with the video playback position, making it difficult to accurately anchor target areas. Furthermore, marker points cannot be effectively correlated with flight path trajectories, affecting the efficiency of scene tracing and data analysis.

Method used

By acquiring the unlabeled stream and photo data of drone aerial video, the coordinates in the target video frame are identified, and the coordinate system transformation relationship is determined by using the feature point sets of historical video frames and the target video frame. The marker points are then calibrated, and a video subtitle file of the flight path time axis is generated, achieving accurate matching and spatiotemporal correlation between the marker points and the video frames.

Benefits of technology

It improves the matching accuracy between video markers and video frames, enhances scene tracing capabilities and data analysis efficiency, and realizes the synchronous display and integrated visualization of markers and flight paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138015A_ABST
    Figure CN122138015A_ABST
Patent Text Reader

Abstract

This invention provides a video tagging method, a video playback method, a video tagging system, and a device, relating to the field of data processing technology. The video tagging method includes: acquiring a video stream to be tagged obtained from drone aerial photography and drone image data; for each target video frame in the video stream to be tagged, identifying the coordinates of the target and acquiring the coordinates of manually tagged targets; determining the transformation relationship between the image coordinate systems of historical video frames and target video frames based on their respective feature point sets; transforming the coordinates of the target in the historical video frames to the target video frames based on the transformation relationship; and calibrating the coordinates of the target in the target video frames based on the transformed coordinates; inserting the calibrated coordinates and image nodes of the target into the timeline according to the timestamps and image data of each target video frame to obtain the flight path timeline, and generating a video subtitle file, thereby improving the matching accuracy between the tagging points and video frames and enhancing scene tracing capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a video tagging method, a video playback method, a video tagging system, and a device. Background Technology

[0002] With the popularization of drone aerial photography technology, aerial videos have been widely used in fields such as surveying and exploration, environmental monitoring, power line inspection, and emergency rescue. In practical applications, key targets in aerial videos are usually automatically marked using artificial intelligence (AI) recognition technology, supplemented and corrected manually, and combined with the aerial flight path trajectory for scene tracing and data analysis.

[0003] There are two main types of methods for processing aerial videos in related technologies: one focuses on marker generation, mainly using AI recognition or manual annotation to mark key targets in the aerial video. This method can only display the markers separately in the video frame and cannot be associated with flight paths and time nodes. The other focuses on flight path visualization. Although this method can visualize the flight path, it cannot be associated with the aerial video markers. In an existing aerial video marking and trajectory display system, markers are generated by AI recognition and statically superimposed on the video frame. The creation time of the markers is presented in the form of a text list. Furthermore, the system generates flight paths using GPS data corresponding to the aerial video and displays them on a separate map interface, thus achieving aerial video marking and trajectory display.

[0004] However, in the aforementioned aerial video marking and trajectory display system, the generated markers are statically superimposed on the video frame. This can easily cause the markers to shift position as the video plays, making it difficult to accurately anchor the target area and resulting in insufficient spatial matching accuracy between the markers and the target area in the video frame. Furthermore, the static superposition of markers on the video frame while the flight path is displayed independently on the map interface makes it difficult to correlate the two when using markers and flight paths for scene tracing, thus weakening the tracing capability. Additionally, when using markers and flight paths for data analysis, the scattered data presentation requires repeated switching of interfaces, leading to low analysis efficiency. Summary of the Invention

[0005] The purpose of this invention is to provide a video tagging method, a video playback method, a video tagging system, and a device to improve the matching accuracy between video tagging points and video frames, and to enhance scene tracing capabilities and data analysis efficiency. The specific technical solution is as follows:

[0006] In a first aspect, embodiments of the present invention provide a video tagging method, the method comprising:

[0007] Acquire the video stream to be labeled obtained from drone aerial photography and the photo data taken by the drone;

[0008] For each target video frame in the video stream to be labeled, the coordinates of the target in the target video frame are identified as the first coordinates, and the manually labeled coordinates of the target in the target video frame are obtained as the second coordinates.

[0009] For each target video frame in the video stream to be labeled, based on the feature point set of historical video frames and the feature point set of the target video frame, a transformation relationship between the image coordinate system of the historical video frames and the image coordinate system of the target video frame is determined; based on the transformation relationship, the first and second coordinates of the target in the historical video frames are transformed to the target video frame to obtain the transformed first and second coordinates; and the first and second coordinates of the target in the target video frame are calibrated based on the transformed first and second coordinates respectively.

[0010] According to the timestamp of each target video frame, the calibrated first coordinate and calibrated second coordinate of the target in each target video frame are inserted into the time axis, and the photo-taking node of the UAV is inserted into the time axis according to the photo-taking data to obtain the flight path time axis;

[0011] Generate a video subtitle file to represent the timeline of the flight path.

[0012] Secondly, embodiments of the present invention provide a video playback method, the method comprising:

[0013] The video to be played and the corresponding video subtitle file are played synchronously, wherein the video subtitle file is pre-marked using any of the video marking methods described above.

[0014] Thirdly, embodiments of the present invention provide a video tagging system, the video tagging system comprising:

[0015] The video stream acquisition module is used to acquire the video stream to be labeled obtained by drone aerial photography and the photo data of the drone;

[0016] The marker generation module is used to identify the coordinates of the target in each target video frame in the video stream to be marked as the first coordinate, and to obtain the manually marked coordinates of the target in the target video frame as the second coordinate;

[0017] The labeling and calibration module is used to determine, for each target video frame in the video stream to be labeled, the transformation relationship between the image coordinate system of the historical video frames and the image coordinate system of the target video frames based on the feature point set of the historical video frames and the feature point set of the target video frames; based on the transformation relationship, the first and second coordinates of the target in the historical video frames are transformed to the target video frames to obtain the transformed first and second coordinates; and the first and second coordinates of the target in the target video frames are calibrated based on the transformed first and second coordinates.

[0018] The flight path trajectory timeline generation module is used to insert the calibrated first coordinates and calibrated second coordinates of the target within each target video frame into the timeline according to the timestamp of each target video frame, and to insert the UAV's photo-taking nodes into the timeline according to the photo-taking data, so as to obtain the flight path trajectory timeline.

[0019] The subtitle file generation module is used to generate video subtitle files that represent the timeline of the flight path.

[0020] Fourthly, embodiments of the present invention provide a video playback system, the video playback system comprising:

[0021] The video playback module is used to synchronously play the video to be played and the corresponding video subtitle file of the video to be played, wherein the video subtitle file is pre-marked using any of the video marking methods described above.

[0022] Fifthly, embodiments of the present invention provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0023] Memory, used to store computer programs;

[0024] A processor, when executing a program stored in memory, implements any of the methods described above.

[0025] Sixthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the methods described above.

[0026] In a seventh aspect, embodiments of the present invention also provide a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0027] Beneficial effects of the embodiments of the present invention:

[0028] This invention provides a video tagging method, video playback method, video tagging system, and device. The method involves acquiring a video stream to be tagged obtained from drone aerial photography and drone image data. For each target video frame in the video stream to be tagged, the coordinates of the target within the target video frame are identified as first coordinates, and manually marked coordinates of the target within the target video frame are acquired as second coordinates. For each target video frame in the video stream to be tagged, based on the feature point sets of historical video frames and the feature point sets of the target video frame, a transformation relationship between the image coordinate systems of the historical video frames and the target video frame is determined. Based on this transformation relationship, the first and second coordinates of the target within the historical video frames are transformed to those within the target video frame, resulting in transformed first and second coordinates. The transformed first and second coordinates of the target within the target video frame are then calibrated based on these transformed first and second coordinates. According to the timestamps of each target video frame, the calibrated first and second coordinates of the target within each target video frame are inserted into the timeline. The drone's image capture nodes are also inserted into the timeline according to the image data, resulting in a flight path timeline. Finally, a video subtitle file representing the flight path timeline is generated.

[0029] This invention utilizes feature point sets from historical video frames and target video frames to determine the transformation relationship between the image coordinate systems of the historical and target video frames. Based on this transformation, the first and second coordinates of the target within the historical video frames are transformed to those within the target video frames. The first and second coordinates of the target within the target video frames are then calibrated based on these transformed coordinates, achieving frame-level target localization of marker points. This reduces the spatial matching error between marker points and video frames, thereby improving the matching accuracy between marker points and video frames. By inserting the calibrated first and second coordinates of the targets within each target video frame, along with the drone's image capture node, into the timeline, a linkage relationship is established between the flight path timeline and the playback progress of the video stream to be tagged. This achieves spatiotemporal correlation of the video stream, improving scene tracing capabilities and data analysis efficiency. A video subtitle file representing the flight path timeline is generated, enhancing the user's visualization experience.

[0030] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0032] Figure 1 This is a flowchart illustrating a video tagging method in an embodiment of the present invention;

[0033] Figure 2 This is a schematic diagram of a method for obtaining feature point sets of video frames in an embodiment of the present invention;

[0034] Figure 3 This is a schematic diagram illustrating the generation of a video subtitle file in an embodiment of the present invention;

[0035] Figure 4 This is a schematic diagram illustrating a visual representation in an embodiment of the present invention;

[0036] Figure 5 This is a schematic diagram of a video tagging system according to an embodiment of the present invention;

[0037] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.

[0039] This invention provides a video tagging method, a video playback method, a video tagging system, and a device for processing drone aerial videos. For example... Figure 1 As shown, an embodiment of the present invention provides a video tagging method, comprising:

[0040] S101, acquire the video stream to be labeled and the drone's photo data obtained from aerial photography;

[0041] S102, for each target video frame in the video stream to be labeled, identify the coordinates of the target in the target video frame as the first coordinate, and obtain the coordinates of the target in the manually labeled target video frame as the second coordinate;

[0042] S103, for each target video frame in the video stream to be labeled, based on the feature point set of the historical video frames and the feature point set of the target video frame, determine the transformation relationship between the image coordinate system of the historical video frames and the image coordinate system of the target video frame; based on the transformation relationship, transform the first and second coordinates of the target in the historical video frames to the target video frames respectively, to obtain the transformed first coordinates and the transformed second coordinates; and calibrate the first and second coordinates of the target in the target video frames based on the transformed first coordinates and the transformed second coordinates respectively.

[0043] S104. According to the timestamp of each target video frame, insert the calibrated first coordinate and calibrated second coordinate of the target in each target video frame into the time axis, and insert the UAV's photo-taking node into the time axis according to the photo-taking data to obtain the flight path time axis.

[0044] S105, Generate video subtitle files to represent the timeline of the flight path.

[0045] The video tagging method provided in this invention utilizes feature point sets from historical video frames and target video frames to determine the transformation relationship between the image coordinate systems of the historical and target video frames. Based on this transformation, the first and second coordinates of the target within the historical video frames are transformed to those within the target video frames. The first and second coordinates of the target within the target video frame are then calibrated based on these transformed coordinates, achieving frame-level target localization of the tagged points. This reduces the spatial matching error between the tagged points and the video frames, thereby improving the matching accuracy between the tagged points and the video frames. By inserting the calibrated first and second coordinates of the targets within each target video frame, along with the drone's image capture node, into the timeline, a linkage relationship is established between the flight path timeline and the playback progress of the video stream to be tagged. This achieves spatiotemporal correlation of the video stream, improving scene tracing capabilities and data analysis efficiency. A video subtitle file representing the flight path timeline is generated, enhancing the user's visualization experience.

[0046] The video tagging method provided by the embodiments of the present invention will be described in detail below:

[0047] The video tagging method provided in this invention is applied to a video tagging system, which can be a software or hardware system deployed in ground electronic equipment, such as a server. The video tagging system includes a video stream acquisition module, a tag generation module, a tag calibration module, a flight path timeline generation module, a subtitle file generation module, and a visualization module. These modules work collaboratively to achieve frame-level precise positioning and overlay of AI-generated intelligent tags and manually tagged tags. They correlate UAV aerial trajectory points, image capture nodes, and tag coordinates along a timeline dimension, generating corresponding video subtitle files. Finally, all data is integrated to achieve a unified visual presentation, achieving a five-dimensional integrated effect of "spatial positioning—temporal tracing—trajectory association—source differentiation—video synchronization." The specific implementation process is detailed below.

[0048] In practical applications, when video tagging is required, the video tagging system is started, and each module in the system is initialized and loaded to implement the aforementioned steps S101-S105. The specific implementation process of steps S101-S105 is as follows:

[0049] like Figure 1 In the illustrated embodiment, in step S101, the video stream acquisition module receives, in real time, the video stream to be labeled, the drone's trajectory data, and the photographed data obtained from drone aerial photography via the drone communication interface. The drone communication interface can be, for example, a 4G, 5G, or WiFi interface. The video stream to be labeled can be a target video stream obtained from real-time drone aerial photography or a segment of that target video stream, where a segment contains multiple video frames. The following description uses a segment of the target video stream as an example. The drone's trajectory data can be GPS trajectory data or BeiDou Navigation Satellite System (BDS) trajectory data, etc. This embodiment uses GPS trajectory data as an example.

[0050] Subsequently, the video stream acquisition module performs data format conversion on each video frame in the acquired video stream to be labeled. For example, it converts the resolution of all video frames to the same resolution of 1920×1080, and performs coordinate transformation on the trajectory data. For example, it converts the coordinates of the trajectory data to coordinates in the World Geodetic System-1984 Coordinate System (WGS-84) to achieve preprocessing of the acquired data. The preprocessed video stream to be labeled is then sent to the label generation module, the label calibration module, and the visualization module. The captured data and the processed trajectory data are also sent to the flight path trajectory time axis generation module for subsequent processing.

[0051] In step S102, the marker generation module uses a target detection model to identify targets in each target video frame of the video stream to be marked, obtaining the AI ​​recognition result output by the target detection model. This AI recognition result includes the coordinates of the detection box corresponding to the target, the marker type, and the confidence level. Then, the coordinates of the detection box corresponding to the target obtained through recognition are used as the first coordinates. These coordinates are the coordinates of the target within the target video frame, corresponding to four coordinate points—the coordinates of the four vertices of the detection box corresponding to the target. The target video frame can be any video frame in the video stream to be marked. However, to reduce the amount of video marker data processing and computation while considering the processing capabilities of electronic devices, video frames in the video stream to be marked at intervals of a set duration or a set number of frames can be selected as target video frames for target recognition. The set duration can be, for example, 10ms, 20ms, or 50ms, and the set number of frames can be, for example, 5 frames, 10 frames, or 20 frames. Object detection models can be large-scale deep learning object detection algorithms such as YOLO (You Only Look Once) or Faster Region-based Convolutional Neural Network (Faster R-CNN) models. Targets can be people, buildings, vehicles, etc.

[0052] After target identification, AI-generated marker data can be generated, carrying an AI identifier, a first creation timestamp, and the aforementioned first coordinates. Each AI-generated marker data point corresponds to one target identification; the AI ​​identifier indicates that the marker point originates from an AI marker; and the first creation timestamp is the timestamp corresponding to the identified target video frame. In this embodiment of the invention, the timestamp is accurate to milliseconds.

[0053] In step S102, the marker generation module provides a visual marker interface, allowing operators to manually draw marker points or regions and input marker description information on the video frames corresponding to the video stream to be marked. The marker generation module can obtain the manual marker results for the target video frames in the video stream to be marked. These results may include the coordinates of the marker points or regions, marker description information, and the coordinates of the marker points or regions—that is, the coordinates of the target within the manually marked video frame—as the second coordinates. If it's the coordinates of a marker point, the second coordinate is a single coordinate point; if it's the coordinates of a marker region, the second coordinate corresponds to four coordinate points—that is, the coordinates of the four vertices of the marker region. The target video frame for manual marking may be the same as or different from the target video frame for AI marking. Similarly, manually annotated marker point data carrying a manual identifier, a second creation timestamp, the aforementioned second coordinates, and marker description information can be generated. One manually annotated marker point data corresponds to one manual annotation. The manual identifier indicates that the marker point originated from manual annotation, and the second creation timestamp is the timestamp corresponding to the manually annotated target video frame.

[0054] Once the AI ​​recognition results and manual labeling results are obtained, the AI-recognized label point data and the manually labeled label point data are sent to the labeling calibration module, the flight path timeline generation module, and the subtitle file generation module, etc., for subsequent processing.

[0055] In step S103, the marking calibration module receives the video stream to be marked from the video stream acquisition module and the AI-recognized marker point data and manually labeled marker point data from the marking generation module. By using the feature point sets of historical video frames and the feature point sets of the target video frame, it determines the transformation relationship between the image coordinate systems of the historical video frames and the target video frame. Using this transformation relationship, it transforms the first and second coordinates of the marked target within the historical video frame to the target video frame, calculating the precise pixel coordinates of the AI-recognized marker points and manually labeled marker points of the target within the historical video frame in the target video frame. Based on this, it adjusts the first and second coordinates of the target within the target video frame, ensuring that the marker points accurately cover the target area. The specific calibration process is described in detail below. Then, the marking calibration module sends the calibrated AI-recognized marker point data and the calibrated manually labeled marker point data to the flight path timeline generation module and the visualization module.

[0056] In step S104, the flight path trajectory timeline generation module receives trajectory data and photo data sent by the video stream acquisition module, as well as calibrated AI-recognized marker point data and calibrated manually labeled marker point data sent by the marker calibration module. Following the chronological order of the video frame timestamps, it inserts the calibrated first and second coordinates of the targets within each target video frame into the timeline. Simultaneously, following the chronological order of the photo data, it also inserts the drone's photo nodes into the timeline, thus obtaining the flight path trajectory timeline. This achieves multi-dimensional data association between the drone's photo nodes and the marker points of the video stream to be labeled, thereby establishing a linkage between the flight path trajectory timeline and the video playback progress corresponding to the video stream to be labeled, enabling the timeline to move synchronously with video playback. Then, the flight path trajectory timeline generation module sends the data corresponding to the flight path trajectory timeline, such as photo nodes, marker point data, and the association information between each data point and the timeline, to the subtitle file generation module and the visualization module.

[0057] In one example, different colors or icons can be set for the calibrated first coordinate, calibrated second coordinate, and photo capture node of each target in the video frame inserted into the coordinate axis, so that each marker point can be displayed differently on the time axis.

[0058] In step S105, the subtitle file generation module directly generates a video subtitle file representing the timeline of the flight path based on the data corresponding to the timeline. The specific generation process is described in detail below. The generated subtitle file is then sent to the visualization module. The visualization module integrates all the received data and displays it visually.

[0059] The following section introduces the calibration of the marker point coordinates (i.e., the acquired first and second coordinates) by the marker calibration module:

[0060] exist Figure 1 Based on the embodiment shown, the marker calibration module receives AI-identified marker data and manually labeled marker data. If the coordinates of the target contained in the AI-identified marker data and manually labeled marker data, i.e., the first coordinate and the second coordinate, are not coordinates in the image coordinate system, such as coordinates in the world geodetic coordinate system, then the first coordinate and the second coordinate are first converted to coordinates in the image coordinate system.

[0061] For example, the first and second coordinates are in the same WGS-84 coordinate system as the trajectory data. The first and second coordinates are converted to coordinates in the image coordinate system using the following expression. Taking the first coordinate as an example:

[0062] u=f×(Lng-Lng0) / Alt×cos(Pitch)+u0;

[0063] v=f×(Lat-Lat0) / Alt×cos(Roll)+v0;

[0064] Where (Lng, Lat, Alt) are the first coordinates before transformation, (u, v) are the coordinates of the first coordinates transformed to the image coordinate system, (Lng0, Lat0, Alt0) are the position coordinates of the UAV corresponding to the video frame where the first coordinate is located, (u0, v0) are the principal point coordinates of the UAV aerial camera corresponding to the video frame where the first coordinate is located, f is the focal length of the UAV aerial camera, Pitch is the pitch angle of the UAV aerial camera, and Roll is the roll angle of the UAV aerial camera.

[0065] Then, as Figure 2 As shown, the marker calibration module in this embodiment of the invention uses... Figure 2 The method shown illustrates the acquisition of feature point sets from historical video frames. This process includes:

[0066] S201, Extract corner points located within the target area in historical video frames.

[0067] The Shi-Tomasi corner detection algorithm is used, with a corner response threshold λ ≥ 0.3, to extract corners located within the target area in historical video frames, resulting in a corner feature point set F = {f1, f2, ..., fn}, where n is the number of corners. Setting the corner response threshold λ ≥ 0.3 filters out strong corners within the target area in historical video frames, such as the corners of power towers and the intersections of hazard area outlines, to select stable corners and avoid interference from weak feature points. Setting n ≥ 20 ensures the stability of the extracted features.

[0068] S202, Extract the edge points of the target in historical video frames.

[0069] Using the Canny edge detection operator and setting thresholds T1=50 and T2=150, the contours of targets in historical video frames are extracted to obtain the edge feature point set E={e1, e2, ..., em}, where m is the number of edge points.

[0070] S203 uses optical flow to track corner and edge points of historical video frames and the next video frame, and adds corner and edge points that meet the displacement threshold to the feature point set of historical video frames.

[0071] A feature point tracking algorithm based on optical flow (Lucas–Kanade, KLT) is used to track the corner and edge points of historical video frames and the corner and edge points of the next video frame. The displacement vectors (Δu, Δv) corresponding to each corner and edge point are calculated. Then, corner and edge points with displacement amounts ≤ a displacement threshold (e.g., 3 pixels) are selected and added to the feature point set of the historical video frames. The next video frame is the video frame immediately following the historical video frame in the temporal domain.

[0072] In one example, texture histograms within the neighborhood of each corner and edge point in historical video frames can be extracted as the texture features of each corner and edge point. For instance, Local Binary Patterns (LBP) can be used to extract the texture histogram of the target region in historical video frames, thus extracting the texture histograms within the neighborhood of each corner and edge point in historical video frames as the texture features of each corner and edge point. Then, the texture features of the corner and edge points are used as auxiliary conditions for feature point set selection. Based on the similarity between the texture features of each corner and edge point in historical video frames and the texture features of each corner and edge point in the next video frame, corner and edge points whose similarity meets the similarity threshold are added to the feature point set of historical video frames. The similarity threshold can be set according to requirements, for example, it can be set to any value in [0.85, 1).

[0073] In this embodiment of the invention, a multi-feature fusion screening method of "corner point + edge + optical flow tracking + texture assistance" is adopted to obtain the feature point set of historical video frames. This avoids the influence of lighting, occlusion, etc. on a single feature, making the obtained features more stable, so as to more accurately calibrate the marker point features.

[0074] Through the above Figure 2 The feature point set of historical video frames is obtained in the manner shown, and the feature point set of the target video frame is obtained in the same way. Based on this, Figure 1 In the illustrated embodiment, the marker calibration module in step S103 determines the transformation relationship H between the image coordinate system of the historical video frame and the image coordinate system of the target video frame through the expression [u1'; v1'; 1] = H × [u1; v1; 1], where (u1, v1) are the coordinates of the feature points in the feature point set of the historical video frame in the image coordinate system, (u1', v1') are the coordinates of the feature points in the feature point set of the target video frame in the image coordinate system, and H is a 3×3 homography matrix.

[0075] Then, using the expression (x1, y1) = H × (x0, y0), the first and second coordinates of the target within the historical video frames are transformed to the target video frame, resulting in the transformed first and second coordinates. Here, (x1, y1) represents the transformed coordinates, and (x0, y0) represents the target's coordinates within the historical video frames. The historical video frame is the k-th video frame whose timestamp precedes the timestamp of the target video frame; the value of k is set according to the actual situation.

[0076] In this embodiment of the invention, the Random Sample Consensus (RANSAC) algorithm is used to solve for H, with 1000 iterations and an interior point threshold of 2 pixels, to ensure the robustness of the homography matrix H solution and to remove outliers to ensure the accuracy of the transformation relationship.

[0077] After obtaining the first and second transformed coordinates, accordingly, Figure 1 In the illustrated embodiment S103, the marker calibration module calibrates the first and second coordinates of the target within the target video frame based on the transformed first and second coordinates, respectively, including:

[0078] Based on the attitude compensation factor, time decay factor, and the first and second coordinates of the target within the target video frame, the transformed first and second coordinates are calibrated to obtain the first calibrated coordinates and the second calibrated coordinates. Distortion correction is then performed on the first and second calibrated coordinates to obtain the first corrected coordinates and the second corrected coordinates, which are used as the calibrated first and second coordinates of the target within the target video frame.

[0079] In this embodiment of the invention, considering the randomness of UAV jitter, an attitude compensation factor and a time decay factor are introduced. The attitude compensation factor characterizes the attitude change of the UAV from the acquisition of historical video frames to the adoption of the target video frame, while the time decay factor characterizes the duration of this period.

[0080] In one example, the attitude compensation factor K a Through expression K a =1-(|ΔR|+|ΔP|+|ΔY|) / 180 is determined, where |ΔR|, |ΔP|, and |ΔY| represent the absolute changes in roll angle, pitch angle, and yaw angle of the UAV from the acquisition of historical video frames to the adoption of the target video frame, respectively. The roll angle, pitch angle, and yaw angle are measured by the UAV's inertial measurement unit (IMU). K a The value ranges from 0.9 to 1.0 to balance the impact of attitude changes on positioning; the smaller the attitude change, the weaker the attitude compensation.

[0081] Time decay factor K t =e ∧ (-k / 100), where k is the frame interval between historical video frames and the target video frame, and k=0 when K t When =1, k=200, K t ≈0.135. The time decay factor increases with the number of frame intervals k to avoid overcalibrating the target video frame using distant frames.

[0082] Based on attitude compensation factor K a Time decay factor K t The first and second coordinates of the target within the target video frame are then used to calibrate the transformed first and second coordinates using the following expressions, yielding the first and second calibrated coordinates:

[0083] x d+1 =x d +K a ×K t ×(x_feature-x d );

[0084] y d+1 =y d +K a ×K t ×(y_feature-y d );

[0085] Among them, (x) d+1 y d+1 ) represents the calibration coordinates obtained in the d-th iteration (which can be either the first or second calibration coordinates obtained in the d-th iteration), (x d y d Let (x_feature, y_feature) be the calibration coordinates obtained in the (d-1)th iteration. Let (x_feature, y_feature) be the mean coordinates of the feature point set of the target video frame, initially (x1, y1) being the transformed coordinates (either the first or second transformed coordinates). Set the iteration termination condition to |x_feature|_feature|_min. d+1- x d |≤0.1 pixels and|y d+1 -y d |≤0.1 pixels, to achieve sub-pixel level calibration.

[0086] In this embodiment of the invention, by using an attitude compensation factor, a time decay factor, and the first and second coordinates of the target within the target video frame, the first and second coordinates of the target in historical video frames are transformed to coordinates within the target video frame to achieve sub-pixel-level calibration. This allows the marker points of the target within the target video frame to be calibrated using the marker points of the target in historical video frames, thereby achieving frame-level target localization of the marker points, reducing the spatial matching error between the marker points and the video frames, and thus improving the matching accuracy between the marker points of the video markers and the video frames.

[0087] Considering the impact of lens distortion on video markers, given the first and second calibration coordinates obtained through calibration, distortion correction is performed on the first and second calibration coordinates using the following expressions:

[0088] x_corrected=x j ×(1+k1r 2 +k2r 4 )+2p1x j y j +p2(r) 2 +2x j 2 );

[0089] y_corrected=y j ×(1+k1r 2 +k2r 4 )+p1(r 2 +2y j 2 )+2p2x j y j ;

[0090] Where (x_corrected, y_corrected) are the coordinate correction values, (x j y j ) represents the calibration coordinates obtained from the calibration (which can be either the first or second calibration coordinates), r 2 =x j 2 + y j 2 k1 and k2 are radial distortion coefficients, and p1 and p2 are tangential distortion coefficients. k1, k2, p1 and p2 are obtained in advance by calibrating the drone aerial camera.

[0091] The calculated coordinate correction values ​​are used to correct the distortion of the first and second calibration coordinates to compensate for the impact of lens distortion on video markers, thereby improving the matching accuracy between the marker points and video frames.

[0092] In this embodiment of the invention, while correcting distortion in the calibration coordinates, the presence of motion blur in the video frame can be determined using the Laplacian variance. If the Laplacian variance is less than 100, motion blur is determined to exist in the video frame. In this case, motion blur compensation is performed on the first and second corrected coordinates (i.e., the calibrated first and second coordinates) obtained through distortion correction using inter-frame interpolation. The resulting first and second compensated coordinates are then used as the final calibrated first and second coordinates of the target within the target video frame. Specifically, the inter-frame interpolation can be performed linearly on the calibrated first and second coordinates of the target within two adjacent video frames in the temporal domain of the target video frame, and the interpolated coordinates are used as the first and second compensated coordinates.

[0093] In this embodiment of the invention, when motion blur exists in a video frame, motion blur compensation is performed on the calibrated first coordinates and calibrated second coordinates of the target within the target video frame to avoid positioning deviation caused by motion blur.

[0094] Using the embodiments of the present invention, the marker calibration module calibrates the first and second coordinates of the target within the target video frame in the above manner, with a response time ≤8ms / frame, meeting the real-time processing requirements (video frame resolution is 1080P, transmission frame rate is 30 frames / second), and a positioning error ≤±1 pixel, which is more than 80% higher than the ±5 to 10 pixel accuracy of the prior art.

[0095] exist Figure 1 In the illustrated embodiment, based on the route trajectory timeline obtained by the S104 route trajectory timeline generation module, it may further include:

[0096] Based on the timestamp of each target video frame, the first and second coordinates of the target within each target video frame are aligned with the flight segments of the UAV in the time domain to obtain the correspondence between each first and second coordinate and the flight segments.

[0097] Based on the trajectory data, a route map is generated. Based on the correspondence between each first coordinate and each second coordinate and the flight segment, as well as the photographed data, each first coordinate, each second coordinate, and the photographed node are marked on the route map.

[0098] Align the first and second coordinates of the target within each video frame with the drone's flight segment in the time domain to determine the flight segment position corresponding to each marker point, thereby achieving spatiotemporal coordinate association of drone aerial videos.

[0099] Then, based on the timestamps and location coordinates (such as latitude, longitude, and altitude) contained in the drone's trajectory data, a flight path map in the form of a line graph is generated. Based on the correspondence between each first and second coordinate and the flight segment, the timestamps of each target video frame, and the timestamps of the captured images contained in the image data, each first and second coordinate, and the captured image node are marked on the flight path map. The flight path map is then linked to the obtained flight path timeline, causing the timeline slider to move synchronously with the playback of the video stream to be marked, highlighting the flight segment and marker point corresponding to the current time in real time.

[0100] The calibrated first and second coordinates of the targets within each target video frame, along with the drone's photo capture node, are inserted into the timeline. The first coordinates, second coordinates, and photo capture node are then marked on the flight path map. This establishes a linkage between the flight path trajectory timeline and the playback progress of the video stream to be marked. Associating the flight path map with the obtained flight path trajectory timeline improves scene tracing capabilities and achieves spatiotemporal correlation of the video stream to be marked, thereby improving data analysis efficiency.

[0101] The following is an introduction to the video subtitle file generation module:

[0102] Figure 1 In the embodiment shown, the video subtitle file is a SubRip Subtitle (SRT) file. The subtitle file generation module receives AI-recognized marker point data and manually annotated marker point data sent by the marker generation module, data corresponding to the flight trajectory timeline sent by the flight trajectory timeline generation module, and video frame rate data sent by the video stream acquisition module, and then generates the file strictly following the SRT file standard of "serial number + timeline + content + blank line".

[0103] First, the marker data, which includes AI-recognized marker data and manually labeled marker data, as well as the image node data, are structured to define the core fields of each SRT entry:

[0104] 1) Serial Number:

[0105] In chronological order, incrementing sequence numbers are generated for AI-identified marker data, manually labeled marker data, and photographed node data. Specifically, the sequence numbers are automatically incremented based on the first creation timestamp of the AI-identified marker data, the second creation timestamp of the manually labeled marker data, and the photographed timestamp of the photographed data.

[0106] 2) Time axis interval:

[0107] Based on the first creation timestamp, the second creation timestamp, the photo capture timestamp, and preset rules, the timeline intervals corresponding to AI-recognized marker point data, manually labeled marker point data, and photo capture node data are determined respectively. The preset rules are determined based on the video frame rate.

[0108] In one example, for the timeline interval corresponding to AI-identified marker data / manually labeled marker data, the start time = first creation timestamp / second creation timestamp, and the end time = start time + A seconds, where A can be configured from 0.5 seconds to 3 seconds to ensure that the marker data can be continuously displayed during video playback. For the timeline interval corresponding to the photo capture node data, the start time = photo capture timestamp, and the end time = start time + 0.8 seconds. Setting this fixed short duration of 0.8 seconds highlights the instantaneous photo capture event.

[0109] 3) Content:

[0110] The format is “Type-Description-Related Information”. For example, the content of AI-identified marker data is: “AI Identifier | Type: Power Tower | Confidence: 0.85 | Flight Segment Identifier: S001 | GPS: 30.1234°N, 120.5678°E, Altitude 150m”; the content of manually marked marker data is: “Manual Identifier | Description: Hazardous Area | Coordinates: (520,380)-(610,450) | Flight Segment Identifier: S002 | GPS: 30.1245°N, 120.5689°E, Altitude 148m”; the content of photo node data is: “Photo Event | Flight Segment Identifier: S003 | GPS: 30.1256°N, 120.5690°E, Altitude 152m”.

[0111] The AI-identified marker data, manually labeled marker data, and image capture node data are structured using the aforementioned fields of sequence number, timeline range, and content. Then, abnormal timestamps and invalid GPS data are removed from these datasets to ensure compliance with the generated SRT file format. Abnormal timestamps include those exceeding the time range of the video stream to be labeled, and invalid GPS data includes data with GPS errors.

[0112] Secondly, time-axis mapping is performed on the structured data.

[0113] The timeline intervals corresponding to AI-identified marker data, manually labeled marker data, and image capture node data are converted into SRT format time strings. For example, the timeline intervals obtained from structured processing are converted into SRT format time strings: HH:mm:ss,SSS. This SRT format time string represents hours:minutes:seconds, millimeters, such as "10:05:32,123" which means 10 hours 5 minutes 32 seconds 123 millimeters. Setting the SRT format time string HH:mm:ss,SSS ensures that the timeline intervals do not overlap and that the order is consistent with the video playback progress.

[0114] Then, based on the sequence number, time axis interval (converted to SRT format time string), and content of the AI-identified marker point data, manually labeled marker point data, and photographed node data, SRT entries corresponding to each type of data are generated according to the SRT standard format. Each generated SRT entry follows the structure of "Sequence Number -> Time Axis Interval -> Content -> Blank Line".

[0115] Following the example above, the SRT entry corresponding to the generated AI-identified marker data is: 100:01:20,345 --> 00:01:21,345 AI Identifier | Type: Power Tower | Confidence: 0.85 | Segment Identifier: S001 | GPS: 30.1234°N, 120.5678°E, Altitude 150m. Here, "100:01:20,345 --> 00:01:21,345" indicates that the sequence number is 1, and the time interval is 00:01:20,345 --> 00:01:21,345. The SRT entry corresponding to the generated photo node data is: 200:01:25,678-->00:01:26,478 Photo Event | Segment Identifier: S003 | GPS: 30.1256°N, 120.5690°E, Altitude 152m.

[0116] When generating SRT entries, each SRT entry is named and stored according to a preset format so that it can be played synchronously with the video stream to be tagged. For example, the generated SRT entries are named "Aerial Mission Identifier_Time Track_SRT_Generation Timestamp.srt", with one SRT entry named as "DRONE_20240520_101530_TIMELINE_20240520143022.srt". Of course, user-defined naming methods are also supported for each SRT entry. The SRT file composed of the generated SRT entries can be stored in a user-specified directory / storage location for synchronous playback during video playback or for exporting the SRT file.

[0117] For long-duration aerial photography missions, such as drones operating continuously for more than 2 hours, the number of video frames in the acquired video stream to be labeled exceeds 200,000 frames. If the SRT file is generated by generating it from the entire stream, it suffers from poor real-time performance and is prone to data loss due to interruptions in SRT file generation caused by mission interruptions or restarts. To address this issue, this invention employs a strategy of "segment fragmentation + incremental appending + conflict checking + final merging" to generate SRT files in real time, achieving efficient and reliable SRT file generation while supporting breakpoint resume and real-time synchronization.

[0118] In this embodiment of the invention, the video stream acquisition module automatically divides the flight stream into segments based on the spatial continuity and time interval of the UAV trajectory data. Each flight segment corresponds to a slice (corresponding to the video stream to be labeled mentioned above). The number of flight segments can be configured during segment division, which also means configuring the slice threshold. For example, segmentation can be based on a time threshold: the video stream within a preset time period is divided into a slice. This preset time period can be any value within 5 to 30 minutes, such as 10 minutes. In this case, the slice threshold is the preset time period. Slicing can also be based on a spatial threshold: slices are triggered when the GPS distance difference between two adjacent video frames is greater than or equal to a preset distance threshold, or the altitude difference is greater than or equal to a preset altitude threshold. The preset distance threshold is, for example, 500 meters, 600 meters, etc., and the preset altitude threshold is, for example, 100 meters, 150 meters, etc. In this case, the slice threshold is the preset distance threshold or the preset altitude threshold.

[0119] For each segment obtained from the division, a unique segment identifier is assigned to each segment, such as S001, S002, etc.

[0120] When generating SRT files using an incremental append method, the increment is based on segments. Each segment's incremental unit can contain metadata and data content. The metadata includes: segment identifier, segment start timestamp (T_start), segment end timestamp (T_end), number of markers, and segment checksum. The number of markers includes the total number of AI-recognized markers, manually labeled markers, and image node data. The segment checksum is an MD5 checksum. The data content includes: marker data within the segment (including AI-recognized markers, manually labeled markers, and image node data), the trajectory segment corresponding to the segment (i.e., flight segment segments), and the video frame index range corresponding to the segment.

[0121] In one example, such as Figure 3 As shown, the process of generating an SRT file through incremental appending includes:

[0122] S301. Whenever an SRT entry is generated, the checksum of the temporary SRT file is calculated. If the checksum matches the cached checksum, the SRT entry is appended to the temporary SRT file. If the checksum does not match the cached checksum, the temporary SRT file is reverted to the state where the checksum matched the corresponding cached checksum, and the SRT entry is regenerated.

[0123] In this embodiment of the invention, for a target video stream obtained by drone aerial photography, video frames in the segments are processed in real time according to the segment order in the target video stream. For the current segment, whenever an SRT entry is generated, the SRT entry is appended and stored in the temporary SRT file of the current segment. The SRT entry is the SRT entry corresponding to AI-identified marker point data, manually labeled marker point data, or image node data. Here, the current segment is the aforementioned video stream to be labeled, and the temporary SRT file of the current segment is the cached SRT file corresponding to the video stream to be labeled. The storage rule of the temporary SRT file of the current segment is: it is stored according to the naming of "segment identifier_temporary file_generation timestamp.tmp.srt", for example, named S001_TMP_20240520101530.tmp.srt, and then the temporary SRT file of the current segment is stored in the cache directory.

[0124] To ensure data consistency during append storage, a data consistency check is performed. Whenever an SRT entry is generated, the MD5 checksum of the current fragment's temporary SRT file is first calculated and compared with the cached MD5 checksum of the current fragment. If they match, it means the temporary SRT file of the current fragment has not changed since the last SRT entry was stored (e.g., it has not been tampered with or data has not been added or deleted), and the SRT entry is appended to the current fragment's temporary SRT file. If they do not match, it means the temporary SRT file of the current fragment has changed since the last SRT entry was stored. In this case, the temporary SRT file of the current fragment is reverted to the state where the checksum matched the corresponding cached checksum, and the SRT entry is regenerated. This prevents the temporary SRT file of the current fragment from being tampered with or having its SRT entry generation interrupted, which could lead to file corruption.

[0125] S302: When the number of SRT entries in the temporary SRT file meets the preset threshold or the video stream to be marked is marked, the SRT entries in the temporary SRT file are persistently stored.

[0126] In one example, the preset threshold could be 100 or 200, etc. If the temporary SRT file of the current segment meets the above conditions, a local persistent storage is triggered so that the stored SRT file can be played synchronously in real time or exported.

[0127] S303, when the time axis interval of the SRT entry added to the temporary SRT file overlaps with the SRT entry already stored in the temporary SRT file, the SRT entry already stored in the temporary SRT file with overlapping time axis intervals is updated to the SRT entry added to the temporary SRT file.

[0128] Updating SRT entries that already exist in the temporary SRT file and have overlapping timeline intervals to newly added SRT entries ensures that the timeline intervals of the newly added SRT entries do not overlap with the timeline intervals of the SRT entries already stored in the temporary SRT file, thus resolving the time conflict problem.

[0129] S304, determine whether the GPS trajectory data corresponding to the SRT entry appended to the temporary SRT file matches the GPS trajectory segment of the video stream to be marked. If they do not match, mark the SRT entry appended to the temporary SRT file as an abnormal entry.

[0130] In one example, determining whether the GPS trajectory data corresponding to the SRT entry appended to the temporary SRT file matches the GPS trajectory segment of the video stream to be labeled can be done by checking whether the longitude and latitude deviations between the GPS trajectory data corresponding to the SRT entry appended to the temporary SRT file and the GPS trajectory segment of the video stream to be labeled are both less than preset values. If so, a match is determined; otherwise, a mismatch is determined. The preset value is, for example, 0.0001°. In cases of mismatch, the SRT entry appended to the temporary SRT file is marked as an abnormal entry and temporarily stored in a separate abnormal file for subsequent manual review.

[0131] In this embodiment of the invention, GPS association verification for generating SRT entries is achieved by determining whether the GPS trajectory data corresponding to the SRT entry appended to the temporary SRT file matches the GPS trajectory segment of the video stream to be tagged, thereby improving the reliability of generating SRT entries.

[0132] During the generation of SRT entries, the generation status of each shard SRT entry is recorded and stored in real time. Generation status includes not started, generating, completed, and abnormal. In one example, snapshots can be stored containing information such as the number of generated SRT entries for the current shard, the timestamp of the last generated SRT entry, and the cache location. These snapshots are automatically backed up to the local database at set intervals such as 20 seconds, 30 seconds, or 50 seconds. Then, based on the generation status of each shard SRT entry, interrupted SRT entry generation is resumed through steps S305-S307, and the final merging of generated SRT entries is achieved through steps S308-S310.

[0133] S305 identifies the generation status of SRT entries for each segment when an aerial photography mission or video marking mission is restarted or resumed after an interruption. The generation status includes not started, generation in progress, completed, and abnormal.

[0134] When an aerial photography mission or video tagging mission restarts or resumes after an interruption, the stored snapshots are read, and the generation status of SRT entries for each segment is identified through the content of the snapshots.

[0135] S306. For a fragment in the generation state, verify the checksum of the temporary SRT file of the fragment. If the checksum of the temporary SRT file of the fragment matches the corresponding cache checksum, then append the SRT entries starting from the latest stored SRT entry. Otherwise, delete the temporary SRT file of the fragment and regenerate the SRT entries of the fragment.

[0136] For segments in the "generating" state, calculate the MD5 checksum of the segment's temporary SRT file. Compare the calculated MD5 checksum with the cached MD5 checksum of the segment's temporary SRT file. If the MD5 checksum of the segment's temporary SRT file matches its corresponding cached MD5 checksum, it means that the segment's temporary SRT file has not changed during the restart or interruption recovery of the aerial photography or video marking task. In this case, SRT entries are appended starting from the latest stored SRT entry, that is, from the last cached SRT entry. If the MD5 checksums do not match, it means that the segment's temporary SRT file has changed during the restart or interruption recovery of the aerial photography or video marking task. In this case, delete the segment's temporary SRT file and regenerate the segment's SRT entries, that is, regenerate SRT entries starting from the segment's start timestamp to avoid SRT entry generation errors or duplicate generation.

[0137] S307: For a fragment whose generation status is abnormal, determine whether the abnormality has been recovered. If the abnormality has been recovered, regenerate the SRT entry for the fragment. If the abnormality has not been recovered, mark the fragment as pending processing.

[0138] For segments with an abnormal status, determine the cause of the abnormality, such as GPS data loss or video stream interruption. If it is determined that the abnormality has been resolved, such as GPS data returning to normal or video stream reception being normal, then regenerate the SRT entry for that segment. If it is determined that the abnormality has not been resolved, then mark that segment as pending processing, skip that segment, generate the SRT entry for the next segment, and complete the processing after the subsequent aerial photography task or video marking task is completed.

[0139] S308, upon completion of aerial photography or video marking tasks, sorts all completed SRT files in ascending order according to the identifier of each segment, and rearranges all SRT entries in the rearranged SRT files, verifying the consistency between the timeline interval of each SRT entry and the playback progress of each segment.

[0140] The process involves sorting all completed SRT files in ascending order according to their segment identifiers, including files transferred from temporary SRT files. All SRT entries in the rearranged SRT files are then renumbered, for example, starting from 1 and incrementing sequentially, to cover the local numbers of all SRT entries in each segment's SRT files. The consistency between the timeline intervals of each SRT entry and the playback progress of each segment is verified. This ensures that the overall order of all SRT entries in the rearranged SRT files matches the playback progress of the target video stream, with no overlapping or gaps in the timeline intervals. For any existing gaps, a corresponding SRT entry is generated for that gap, but this generated SRT entry is empty.

[0141] S309, In ​​the case of an abnormal entry, mark the abnormal entry as abnormal.

[0142] If there is an abnormal entry among all SRT entries in the rearranged SRT file, its position is determined according to the time axis interval of the abnormal entry, and the abnormal SRT entry is marked as abnormal, for example, by adding "Status: Abnormal" to the content field of the abnormal SRT entry.

[0143] S310, according to the SRT standard format, merges all SRT entries, generates a complete SRT file corresponding to the complete video stream of each segment, and deletes all temporary SRT files.

[0144] According to the SRT standard format, such as UTF-8 encoding, all SRT entries are merged to generate a complete SRT file corresponding to each segment of the complete video stream (i.e., the target video stream). At the same time, all temporary SRT files and segment files are deleted to free up storage resources.

[0145] This invention employs a strategy of "segment segmentation + incremental appending + conflict checking + final merging" to generate SRT files in real time. The video stream is segmented, and the segmentation threshold is configurable, adapting to aerial photography tasks of varying durations. For example, short tasks (≤30 minutes) can be segmented without segmentation, while long tasks are automatically segmented. For a 10-minute video stream segment, approximately 18,000 SRT entries are generated, with a single segment SRT file size ≤5MB (megabytes), reducing real-time storage usage by over 80% compared to generating a full SRT file. The incremental appending method for SRT entry generation has an incremental appending latency of ≤1 second, supporting real-time synchronization and export for long tasks, ensuring strong real-time performance. The interruption recovery and conflict checking mechanisms in SRT entry generation and appending storage prevent data loss or corruption during SRT entry generation. Marking abnormal SRT entries does not affect the validity of the entire SRT file, thus achieving efficient and reliable SRT file generation. It also supports breakpoint resumption and real-time synchronization.

[0146] In this embodiment of the invention, the visualization module can integrate the acquired video stream, calibrated first coordinates, calibrated second coordinates, flight path timeline, and generated SRT file to achieve integrated display and interaction of multi-dimensional data. Based on this, in the process of generating the SRT file, after persistently storing the SRT entries in the temporary SRT file, the method may further include: responding to a visualization operation command, synchronously displaying the SRT entries in the temporary SRT file and the video stream to be marked; responding to a file export command, exporting the SRT entries in the temporary SRT file; responding to an editing command for the target SRT entry, updating the edited SRT entry to the temporary SRT file, and after completing the editing of the SRT entries in the temporary SRT file, persistently storing the updated temporary SRT file.

[0147] During the SRT file generation process described above, after the SRT entries in the temporary SRT file are persistently stored, the SRT entries in the persistently stored SRT file are automatically synchronized to the visualization module. The visualization module provides a visual user interface, allowing users to interact with the visualization module in real time through the interface.

[0148] For example, a user can issue a visualization operation command through the visual operation interface to view or display the SRT entries for a certain segment. In response to the user's visualization operation command, the visualization module will synchronously display the SRT entries in the corresponding SRT file and the video stream to be marked, so as to support the user to view the effect of the synchronization between the SRT entries and the video in each segment in real time.

[0149] Users issue export commands for SRT files of a specific segment through the visual operation interface. In response, the visualization module exports the SRT entries from the stored SRT file. The naming format of the exported SRT entries can be, for example, "Aerial Mission Identifier_Segment Identifier_Incremental Export_Timestamp.srt" to meet the needs of rapid on-site analysis.

[0150] Users can edit a specific SRT entry within a segment through a visual interface, such as modifying the timeline range or content of the SRT entry. In response, the visualization module updates the edited SRT entry to a temporary SRT file based on the editing command for the target SRT entry. After completing the editing of the SRT entry in the temporary SRT file, the updated temporary SRT file is persistently stored, and the MD5 checksum of the updated temporary SRT file is updated synchronously. The updated temporary SRT file is also synchronized to the incremental generation process of subsequent SRT entries to ensure that the editing results of secondary editing of SRT entries are not lost.

[0151] In this embodiment of the invention, the visualization module provides a video playback unit that supports synchronous playback of video and SRT files. Specifically, the visualization module receives video stream data, calibrated first coordinates, calibrated second coordinates, flight path timeline, and generated SRT files. The video playback unit provides control functions such as real-time playback, pause, fast forward, rewind, and frame skipping, enabling the SRT file information to be synchronized with the video screen and timeline during playback.

[0152] By assigning different colors or icons to the calibrated first and second coordinates of targets within each video frame inserted into the coordinate axis, as well as to the captured nodes—for example, using a red circle for the calibrated first coordinate, a blue square for the calibrated second coordinate, and a green triangle for the captured nodes—the markers are clearly distinguished on the timeline, and the generated flight path diagram can also be displayed. During video playback, each marker can be overlaid onto the video frame in real-time according to the style used when inserted into the timeline. Different styles allow for differentiation of the marker's origin, and each marker corresponds one-to-one with the marker type in its corresponding SRT entry. Furthermore, because SRT entries correspond to markers, users can quickly jump to the corresponding video frame and SRT entry when clicking on the timeline, achieving a synchronized display of the timeline, SRT entries, and video frame.

[0153] The visualization module's user interface also provides an SRT file export button to allow users to download SRT files. Simultaneously, the interface offers an SRT file upload interface, enabling the uploading of external SRT files for parsing and synchronous display. Combined with the aforementioned SRT entry editing operations, this achieves a closed loop of "import—synchronize—edit—export".

[0154] For example, such as Figure 4 As shown, the video is played in the video display area of ​​the visualization module's operation interface, and real-time playback, pause, fast forward, and rewind control functions are provided. The associated flight path timeline is displayed below the video display area. Manually marked points, AI-recognized marked points, and photo nodes are inserted on the flight path timeline, and the flight path map is displayed. At the same time, SRT file operations and secondary editing of marks can be performed through function buttons such as exporting SRT files, uploading SRT files, and editing marks.

[0155] The video tagging method provided in this invention utilizes the feature point set of historical video frames and the feature point set of the target video frame to determine the transformation relationship between the image coordinate system of the historical video frame and the image coordinate system of the target video frame. Then, based on the transformation relationship, the first and second coordinates of the target in the historical video frame are transformed to the target video frame. The first and second coordinates of the target in the target video frame are calibrated based on the transformed first and second coordinates. By employing frame-level target localization technology, frame-level target localization and real-time offset correction of the tag points are achieved, reducing the spatial matching error between the tag points and the video frame. This enables the tag points to accurately anchor the target, thereby improving the matching accuracy between the tag points and the video frame.

[0156] The calibrated first and second coordinates of the target within each video frame, along with the drone's photo capture node, are inserted into the timeline. This establishes a linkage between the flight path timeline and the playback progress of the video stream to be marked. A flight path map is generated based on the trajectory data, and each first coordinate, each second coordinate, and the photo capture node are marked on the flight path map. This achieves a deep correlation between the flight path and the marker points, breaks down data silos, and forms a complete spatiotemporal data chain, allowing users to obtain complete information without switching between multiple interfaces.

[0157] Different colors or icons are set for the calibrated first and second coordinates of targets within each target video frame inserted into the coordinate axis, as well as for the captured nodes, so that each marker point is displayed differently on the time axis. The SRT entries corresponding to each marker point are also defined. The first and second coordinates of targets within each target video frame are aligned with the drone's flight segments in the time domain, and the correspondence between each first and second coordinate and each flight segment is determined. This associates the marker points, SRT files, and flight segments, realizing the spatiotemporal association of the video stream to be marked. This is beneficial for improving scene tracing capabilities, allowing users to quickly trace the background of marker generation and the corresponding aerial scene, thus improving data analysis efficiency.

[0158] The system generates SRT files to represent the timeline of flight paths. These SRT files support synchronized playback with video and are cross-platform compatible, adapting to various scenarios such as on-site operations, video editing, and offline analysis. They also support file export, enhancing the user's visualization experience and improving the ease of use of SRT files. Furthermore, the generated SRT files conform to industry standards and can be directly imported into mainstream video players and editing software, meeting the needs of cross-system data sharing, secondary video creation, and compliant archiving, making them suitable for a wide range of scenarios.

[0159] The video tagging method provided in this invention supports a variety of mainstream aerial video formats and common resolution specifications, has a wide range of applications, and can meet the application needs of multiple fields such as surveying and exploration, environmental monitoring, and power line inspection.

[0160] This invention also provides a video playback method, which includes synchronously playing a video to be played and a video subtitle file corresponding to the video to be played, wherein the video subtitle file is pre-marked using any of the video marking methods described above.

[0161] The video playback method provided in this invention utilizes the feature point sets of historical video frames and the feature point set of the target video frame to determine the transformation relationship between the image coordinate systems of the historical video frames and the target video frame. Then, based on this transformation relationship, the first and second coordinates of the target within the historical video frames are transformed to those within the target video frame. The first and second coordinates of the target within the target video frame are then calibrated based on the transformed first and second coordinates, achieving frame-level target localization of marker points. This reduces the spatial matching error between marker points and video frames, thereby improving the matching accuracy between marker points and video frames. By inserting the calibrated first and second coordinates of the targets within each target video frame, along with the drone's image capture node, into the timeline, a linkage relationship is established between the flight path timeline and the playback progress of the video stream to be marked. This achieves spatiotemporal correlation of the video stream to be marked, which is beneficial for improving scene tracing capabilities and data analysis efficiency. A video subtitle file representing the flight path timeline is generated, enhancing the user's visualization experience. Furthermore, by synchronously playing the video subtitle files corresponding to the video to be played and the video to be played, a deep correlation between flight trajectory and marker points is achieved, breaking down data silos and forming a complete spatiotemporal data chain. Complete information can be obtained without switching between multiple interfaces, which is conducive to improving scene tracing capabilities. This allows users to quickly trace the background of marker generation and the corresponding aerial photography scene, thereby improving data analysis efficiency.

[0162] This invention also provides a video tagging system, such as... Figure 5 As shown, the video tagging system includes:

[0163] The video stream acquisition module 501 is used to acquire the video stream to be labeled obtained by drone aerial photography and the drone's photo data;

[0164] The marker generation module 502 is used to identify the coordinates of the target in each target video frame in the video stream to be marked as the first coordinate, and to obtain the coordinates of the target in the manually marked target video frame as the second coordinate;

[0165] The labeling and calibration module 503 is used to determine the transformation relationship between the image coordinate system of the historical video frame and the image coordinate system of the target video frame for each target video frame in the video stream to be labeled, based on the feature point set of the historical video frame and the feature point set of the target video frame; based on the transformation relationship, the first and second coordinates of the target in the historical video frame are transformed to the target video frame to obtain the transformed first and second coordinates; and the first and second coordinates of the target in the target video frame are calibrated based on the transformed first and second coordinates respectively.

[0166] The flight path trajectory timeline generation module 504 is used to insert the calibrated first coordinate and calibrated second coordinate of the target in each target video frame into the timeline according to the timestamp of each target video frame, and to insert the UAV's photo-taking nodes into the timeline according to the photo-taking data to obtain the flight path trajectory timeline.

[0167] The subtitle file generation module 505 is used to generate video subtitle files that represent the timeline of the flight path.

[0168] Optionally, the calibration module 503 is specifically used to calibrate the transformed first coordinate and the transformed second coordinate based on the attitude compensation factor, the time decay factor, and the first and second coordinates of the target within the target video frame, respectively, to obtain a first calibrated coordinate and a second calibrated coordinate; the attitude compensation factor is used to characterize the attitude change of the UAV from the acquisition of the historical video frame to the adoption of the target video frame; the time decay factor is used to characterize the duration of the period; the distortion correction is performed on the first calibrated coordinate and the second calibrated coordinate respectively to obtain a first corrected coordinate and a second corrected coordinate as the calibrated first coordinate and calibrated second coordinate of the target within the target video frame.

[0169] Optionally, the video tagging system further includes a feature acquisition module.

[0170] The feature acquisition module is used to extract corner points located within the target area of ​​the target in the historical video frame; extract edge points of the target in the historical video frame; track the corner points and edge points of the historical video frame and the next video frame using optical flow method, and add corner points and edge points that meet the displacement threshold to the feature point set of the historical video frame, wherein the next video frame is the next video frame that is temporally adjacent to the historical video frame.

[0171] Optionally, the video marking system further includes a flight path generation module.

[0172] The flight path generation module is used to acquire the trajectory data of the UAV; according to the timestamp of each target video frame, align the first coordinate and the second coordinate of the target in each target video frame with the flight segment of the UAV in the time domain to obtain the correspondence between each first coordinate and each second coordinate and the flight segment; based on the trajectory data, generate a flight path map, and based on the correspondence between each first coordinate and each second coordinate and the flight segment and the photographed data, mark each first coordinate, each second coordinate and the photographed node on the flight path map.

[0173] Optionally, the video subtitle file is an SRT file; the subtitle file generation module 505 is specifically used for:

[0174] Whenever an SRT entry is generated, the checksum of the temporary SRT file is calculated. If the checksum matches the cached checksum, the SRT entry is appended to the temporary SRT file. If the checksum does not match the cached checksum, the temporary SRT file is reverted to the state where the checksum last matched the corresponding cached checksum, and the SRT entry is regenerated.

[0175] When the number of SRT entries in the temporary SRT file meets a preset threshold or when the video stream to be tagged is tagged, the SRT entries in the temporary SRT file are persistently stored.

[0176] When the time axis interval of an SRT entry added to the temporary SRT file overlaps with that of an SRT entry already stored in the temporary SRT file, the SRT entry already stored in the temporary SRT file with overlapping time axis intervals is updated to the added SRT entry.

[0177] Determine whether the GPS trajectory data corresponding to the SRT entry appended to the temporary SRT file matches the GPS trajectory segment of the video stream to be labeled. If they do not match, mark the SRT entry appended to the temporary SRT file as an abnormal entry.

[0178] Optionally, the video stream to be labeled is a segment of the target video stream, and the subtitle file generation module 505 is further used for:

[0179] When an aerial photography mission or video marking mission is restarted or resumed after an interruption, the generation status of the SRT entries for each segment is identified. The generation status includes not started, generating, completed, and abnormal.

[0180] For a fragment in the generation state, the checksum of the temporary SRT file of the fragment is checked. If the checksum of the temporary SRT file of the fragment is consistent with the corresponding cache checksum, the SRT entries are appended and stored starting from the latest stored SRT entry. Otherwise, the temporary SRT file of the fragment is deleted and the SRT entries of the fragment are regenerated.

[0181] For shards with an abnormal generation status, determine whether the abnormality has been recovered. If the abnormality has been recovered, regenerate the SRT entry for the shard. If the abnormality has not been recovered, mark the shard as pending.

[0182] Once the aerial photography or video marking task is completed, sort all completed SRT files in ascending order according to the identifier of each segment, and rearrange all SRT entries in the rearranged SRT files to verify the consistency between the timeline interval of each SRT entry and the playback progress of each segment.

[0183] If any abnormal entries are found, they shall be marked as abnormal.

[0184] According to the SRT standard format, all SRT entries are merged to generate a complete SRT file corresponding to the complete video stream of each segment, and all temporary SRT files are deleted.

[0185] Optionally, the video stream to be labeled is a segment of the target video stream, and the subtitle file generation module 505, after persistently storing the SRT entries in the temporary SRT file, is further used for:

[0186] In response to a visual operation command, the SRT entries in the temporary SRT file are displayed synchronously with the video stream to be labeled;

[0187] In response to a file export command, the SRT entries in the temporary SRT file are exported;

[0188] In response to an edit command for a target SRT entry, the edited SRT entry is updated in the temporary SRT file, and after the editing of the SRT entries in the temporary SRT file is completed, the updated temporary SRT file is persistently stored.

[0189] This invention also provides a video playback system, which includes a video playback module;

[0190] The video playback module is used to synchronously play the video to be played and the corresponding video subtitle file of the video to be played, wherein the video subtitle file is pre-marked using any of the video marking methods described above.

[0191] This invention also provides an electronic device, such as... Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.

[0192] Memory 603 is used to store computer programs;

[0193] When the processor 601 executes the program stored in the memory 603, it implements any of the above methods to achieve the same technical effect.

[0194] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0195] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0196] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0197] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0198] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements any of the above methods to achieve the same technical effect.

[0199] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the methods described above to achieve the same technical effect.

[0200] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0201] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0202] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system / electronic device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0203] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A video tagging method, characterized in that, The method includes: Acquire the video stream to be labeled obtained from drone aerial photography and the photo data taken by the drone; For each target video frame in the video stream to be labeled, the coordinates of the target in the target video frame are identified as the first coordinates, and the manually labeled coordinates of the target in the target video frame are obtained as the second coordinates. For each target video frame in the video stream to be labeled, based on the feature point set of historical video frames and the feature point set of the target video frame, a transformation relationship between the image coordinate system of the historical video frames and the image coordinate system of the target video frame is determined; based on the transformation relationship, the first and second coordinates of the target in the historical video frames are transformed to the target video frame to obtain the transformed first and second coordinates; and the first and second coordinates of the target in the target video frame are calibrated based on the transformed first and second coordinates respectively. According to the timestamp of each target video frame, the calibrated first coordinate and calibrated second coordinate of the target in each target video frame are inserted into the time axis, and the photo-taking node of the UAV is inserted into the time axis according to the photo-taking data to obtain the flight path time axis; Generate a video subtitle file to represent the timeline of the flight path.

2. The method according to claim 1, characterized in that, The calibration of the first and second coordinates of the target within the target video frame based on the transformed first and second coordinates respectively includes: Based on the attitude compensation factor, the time decay factor, and the first and second coordinates of the target within the target video frame, the transformed first coordinates and the transformed second coordinates are calibrated respectively to obtain the first calibrated coordinates and the second calibrated coordinates; the attitude compensation factor is used to characterize the attitude change of the UAV from the acquisition of the historical video frame to the adoption of the target video frame; the time decay factor is used to characterize the duration of the period. Distortion corrections are performed on the first calibration coordinates and the second calibration coordinates respectively to obtain the first corrected coordinates and the second corrected coordinates as the calibrated first coordinates and the calibrated second coordinates of the target within the target video frame.

3. The method according to claim 1, characterized in that, The method further includes: Extract the corner points located within the target area in the historical video frames; Extract the edge points of the target in the historical video frames; The corner and edge points of the historical video frame and the next video frame are tracked using optical flow. Corner and edge points that meet the displacement threshold are added to the feature point set of the historical video frame. The next video frame is the next video frame that is temporally adjacent to the historical video frame.

4. The method according to claim 1, characterized in that, The method further includes: Acquire the trajectory data of the drone; Based on the timestamps of each target video frame, the first coordinates and second coordinates of the target within each target video frame are aligned with the flight segments of the UAV in the time domain to obtain the correspondence between each first coordinate and each second coordinate and the flight segments. Based on the trajectory data, a route map is generated, and based on the correspondence between each of the first coordinates and each of the second coordinates and the flight segment, as well as the photographed data, each of the first coordinates, each of the second coordinates, and the photographed node are marked on the route map.

5. The method according to claim 1, characterized in that, The video subtitle file is an SRT file; The generation of the video subtitle file for representing the timeline of the flight path includes: Whenever an SRT entry is generated, the checksum of the temporary SRT file is calculated. If the checksum matches the cached checksum, the SRT entry is appended to the temporary SRT file. If the checksum does not match the cached checksum, the temporary SRT file is reverted to the state where the checksum last matched the corresponding cached checksum, and the SRT entry is regenerated. When the number of SRT entries in the temporary SRT file meets a preset threshold or when the video stream to be tagged is tagged, the SRT entries in the temporary SRT file are persistently stored. When the time axis interval of an SRT entry added to the temporary SRT file overlaps with that of an SRT entry already stored in the temporary SRT file, the SRT entry already stored in the temporary SRT file with overlapping time axis intervals is updated to the added SRT entry. Determine whether the GPS trajectory data corresponding to the SRT entry appended to the temporary SRT file matches the GPS trajectory segment of the video stream to be labeled. If they do not match, mark the SRT entry appended to the temporary SRT file as an abnormal entry.

6. The method according to claim 5, characterized in that, The video stream to be labeled is a segment of the target video stream, and the method further includes: When an aerial photography mission or video marking mission is restarted or resumed after an interruption, the generation status of the SRT entries for each segment is identified. The generation status includes not started, generating, completed, and abnormal. For a fragment in the generation state, the checksum of the temporary SRT file of the fragment is checked. If the checksum of the temporary SRT file of the fragment is consistent with the corresponding cache checksum, the SRT entries are appended and stored starting from the latest stored SRT entry. Otherwise, the temporary SRT file of the fragment is deleted and the SRT entries of the fragment are regenerated. For shards with an abnormal generation status, determine whether the abnormality has been recovered. If the abnormality has been recovered, regenerate the SRT entry for the shard. If the abnormality has not been recovered, mark the shard as pending. Once the aerial photography or video marking task is completed, sort all completed SRT files in ascending order according to the identifier of each segment, and rearrange all SRT entries in the rearranged SRT files to verify the consistency between the timeline interval of each SRT entry and the playback progress of each segment. If any abnormal entries are found, they shall be marked as abnormal. According to the SRT standard format, all SRT entries are merged to generate a complete SRT file corresponding to the complete video stream of each segment, and all temporary SRT files are deleted.

7. The method according to claim 5, characterized in that, After persistently storing the SRT entries in the temporary SRT file, the method further includes: In response to a visual operation command, the SRT entries in the temporary SRT file are displayed synchronously with the video stream to be labeled; In response to a file export command, the SRT entries in the temporary SRT file are exported; In response to an edit command for a target SRT entry, the edited SRT entry is updated in the temporary SRT file, and after the editing of the SRT entries in the temporary SRT file is completed, the updated temporary SRT file is persistently stored.

8. A video playback method, characterized in that, The method includes: The video to be played and the corresponding video subtitle file are played synchronously, wherein the video subtitle file is pre-marked using any one of the video marking methods described in claims 1-7.

9. A video tagging system, characterized in that, The video tagging system includes: The video stream acquisition module is used to acquire the video stream to be labeled obtained by drone aerial photography and the photo data of the drone; The marker generation module is used to identify the coordinates of the target in each target video frame in the video stream to be marked as the first coordinate, and to obtain the manually marked coordinates of the target in the target video frame as the second coordinate; The labeling and calibration module is used to determine, for each target video frame in the video stream to be labeled, the transformation relationship between the image coordinate system of the historical video frames and the image coordinate system of the target video frames based on the feature point set of the historical video frames and the feature point set of the target video frames; based on the transformation relationship, the first and second coordinates of the target in the historical video frames are transformed to the target video frames to obtain the transformed first and second coordinates; and the first and second coordinates of the target in the target video frames are calibrated based on the transformed first and second coordinates. The flight path trajectory timeline generation module is used to insert the calibrated first coordinates and calibrated second coordinates of the target within each target video frame into the timeline according to the timestamp of each target video frame, and to insert the UAV's photo-taking nodes into the timeline according to the photo-taking data, so as to obtain the flight path trajectory timeline. The subtitle file generation module is used to generate video subtitle files that represent the timeline of the flight path.

10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-8.