Real-time intelligent marking method and system for aerial photography of unmanned aerial vehicle and storage medium
By employing a real-time intelligent labeling method based on multi-channel parallel analysis and decision fusion models, the problem of low efficiency in manual screening of drone aerial videos was solved. This method enables automatic identification and semantic annotation of highlight clips, thereby improving material utilization and editing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-10
AI Technical Summary
Existing drone aerial videos rely on manual screening and lack real-time labeling and intelligent analysis, resulting in low material utilization and poor editing efficiency. Furthermore, existing intelligent analysis methods cannot comprehensively consider event semantics and visual aesthetics, leading to discrepancies between the selection results and human subjective perception.
A multi-channel parallel analysis method is used to extract event features and aesthetic features. A decision fusion model is used to generate a highlight confidence score. Highlight video segments are automatically delineated by combining threshold comparison and semantic tags are generated to achieve real-time intelligent tagging.
It enables real-time identification and marking of potential highlights during drone aerial photography, improving the accuracy and real-time performance of highlight extraction, reducing manual screening workload, and increasing material utilization and editing efficiency.
Smart Images

Figure CN121640311A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of unmanned aerial vehicles (UAVs), and more specifically, to a real-time intelligent tagging method, system, and storage medium for UAV aerial photography. Background Technology
[0002] With the rapid development of drone technology and the widespread application of image recognition algorithms, drone aerial photography has been widely used in news reporting, film and television production, emergency reconnaissance, sports events, and tourism promotion. Existing drone aerial photography systems typically possess high-definition video capture and storage capabilities, and some models can achieve basic automatic exposure, target tracking, or image stabilization control. However, the identification and annotation of "highlights" or "exciting moments" in aerial video content still largely relies on manual playback and selection, which is inefficient and highly subjective.
[0003] In practical applications, drone-captured video data is massive, and aerial filming often lasts from tens of minutes to several hours. Operators need to search for valuable clips frame by frame from this vast amount of footage, a highly demanding task that easily leads to missing crucial shots. While some existing intelligent analysis solutions can achieve automatic filtering based on target recognition or motion detection, they typically rely on a single feature, such as "detecting a person or vehicle" or "dramatic scene changes," to determine if a clip is compelling. These methods fail to comprehensively consider event semantics and visual aesthetics, resulting in selections that deviate from human subjective perception and ultimately failing to meet the demands of high-quality aerial video production. Summary of the Invention
[0004] The main purpose of this application is to provide a real-time intelligent tagging method, system, and storage medium for drone aerial photography, so as to solve the problems of low material utilization and poor editing efficiency caused by the reliance on manual screening and lack of real-time tagging and intelligent analysis mechanisms in existing drone aerial videos.
[0005] To achieve the above objectives, according to one aspect of this application, a real-time intelligent tagging method, system, and storage medium for drone aerial photography are provided.
[0006] A real-time intelligent tagging method for drone aerial photography according to this application includes the following steps: S1. Receive video stream data collected in real time by the drone's onboard camera unit; S2. Perform multi-channel parallel analysis on the video stream data to extract at least one event feature and at least one aesthetic feature; S3. Based on the extracted event features and aesthetic features, feature fusion is performed through a decision fusion model to generate a time-varying highlight confidence score. S4. Based on the dynamic comparison between the highlight confidence score and the preset threshold, define one or more highlight video segments; S5. Generate and associate at least one semantic tag with the highlight video clip. The semantic tag is generated based on the semantic description of event features and aesthetic features, and is stored in the form of metadata associated with the video clip.
[0007] Furthermore, the event features include at least one of the following: the macro-scene type determined by the scene semantic recognition model; the specific action or group behavior identified by the dynamic event detection model; and the specific key target or subject detected by the target recognition model.
[0008] Furthermore, aesthetic features include at least one of the following: special lighting and reflection effects detected by the light and shadow quality analysis model; the rationality of the image composition evaluated by the composition quality analysis model; and video jitter and smoothness evaluated by the image stability analysis model.
[0009] Furthermore, the decision fusion model in step S3 adopts a joint logic of weighted fusion and rule-based judgment; wherein, the weighting coefficients are adaptively adjusted according to the confidence and temporal continuity of different feature sources; and, when multiple different types of features co-occur within the same time window, the weight gain of the highlight confidence score is adjusted.
[0010] Furthermore, the process of defining the highlight video segment in step S4 includes: When the highlight confidence score changes from below the preset threshold to above the preset threshold, that moment is determined as the segment start point; When the score changes from above the preset threshold to below the preset threshold, that moment is determined as the end point of the segment; A time buffer is set at the boundary of the preset threshold to filter out misjudgments caused by transient fluctuations.
[0011] Furthermore, the semantic labels generated in step S5 are structured labels, the contents of which include: event category, aesthetic feature category and corresponding confidence information.
[0012] Furthermore, semantic tags and highlight fragment association information are embedded in the original video file in the form of metadata, or stored in a separate tag file that is synchronously associated with the original video file through timestamps.
[0013] A real-time intelligent tagging system for drone aerial photography, used to implement the above method, includes: The data acquisition module is used to receive real-time video stream data from the UAV's onboard camera unit; The feature analysis module is used to perform multi-channel parallel analysis on video stream data to extract event features and aesthetic features; The decision-making module is used to fuse event features and aesthetic features through a decision fusion model to generate a high-gloss confidence score. The segment and tag management module is used to delineate highlight video segments based on highlight confidence scores and generate associated semantic tags; The data storage module is used to store video data and corresponding tag metadata.
[0014] Furthermore, the feature analysis module, decision-making module, and fragment and tag management module are deployed on the UAV's onboard computing unit.
[0015] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned real-time intelligent tagging method for drone aerial photography.
[0016] In this embodiment, video stream data acquired in real time by the onboard camera unit of the UAV is received, and multi-channel parallel analysis is performed on the video stream to extract event features and aesthetic features. By using a decision fusion model to weightedly fuse and dynamically determine multi-source features, a time-varying highlight confidence score is generated. Combined with threshold comparison, highlight segments are automatically identified and semantic labels are generated, achieving the goal of real-time identification and labeling of potential highlights during drone flight photography. This enables intelligent screening and semantic management of aerial footage, improving the accuracy and real-time performance of highlight extraction. It also solves the problems of low footage utilization and poor editing efficiency caused by reliance on manual screening and lack of real-time labeling and intelligent analysis mechanisms in existing drone aerial videos. Attached Figure Description
[0017] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings: Figure 1 This is a flowchart illustrating a real-time intelligent tagging method for drone aerial photography according to an embodiment of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0020] In this application, the terms "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are primarily for the purpose of better describing the invention and its embodiments, and are not intended to limit the indicated device, element, or component to having a specific orientation, or to be constructed and operated in a specific orientation.
[0021] Furthermore, in addition to indicating direction or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in certain situations to indicate a dependency or connection. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0022] Furthermore, the terms "installation," "setup," "equipped with," "connection," "linking," and "socketing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral structure; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] like Figure 1 As shown, this application relates to a real-time intelligent labeling method for drone aerial photography, applicable to multi-rotor or fixed-wing drone systems equipped with onboard computing units. This method achieves intelligent recognition and semantic labeling of highlight segments during aerial photography through real-time video stream analysis.
[0025] In step S1, the UAV's onboard camera unit continuously acquires video stream data and inputs the video stream to the onboard computing unit in real time via the video transmission module. The video data frame rate is set to 30fps, the resolution is 4K, and the computing unit uses an embedded platform with GPU acceleration capabilities to support real-time feature extraction and judgment.
[0026] In step S2, the system performs multi-channel parallel analysis on the received video stream data. The parallel channels include an event feature extraction channel and an aesthetic feature extraction channel.
[0027] The event feature extraction channel calls scene semantic recognition models (such as DeepLab or SegFormer structures) to perform semantic segmentation and macro-scene classification on each frame, outputting scene types such as "city", "mountain", and "coast". At the same time, dynamic event detection models (such as ActionNet based on spatiotemporal convolutional networks) are used to identify specific actions or group behaviors, such as "crowd gathering", "running", and "waving". In addition, object detection models (such as YOLOv8) detect specific subjects in the scene, such as key targets such as "people", "vehicles", and "animals".
[0028] The aesthetic feature extraction channel includes a lighting and shadow analysis model, a composition analysis model, and a stability analysis module. The lighting and shadow analysis model identifies high-light, strong reflection, or backlighting effects by calculating brightness histograms and local contrast. The composition analysis model evaluates the rationality of the composition using image symmetry and subject position distribution features. The stability analysis module calculates the degree of video jitter based on changes in optical flow between adjacent frames. The results of these feature extractions are output as feature vectors, along with their respective confidence values.
[0029] In step S3, the decision fusion model performs a fusion calculation on event features and aesthetic features. This model employs a joint logical structure of weighted fusion and rule-based decision-making. First, it determines adaptive weighting coefficients based on the confidence and temporal continuity of each feature source. When multiple features of different types co-occur within the same time window (e.g., within 1 second), the highlight confidence score for that time window is adjusted by weight gain. For example, when "special lighting effects" and "target person waving event" occur simultaneously, the system increases the highlight confidence by approximately 15%. The final output highlight confidence score dynamically changes over time, forming a continuous temporal signal.
[0030] In step S4, the system dynamically compares the highlight confidence score with a preset threshold to define highlight video segments. When the score changes from below the threshold to above the threshold, this moment is recorded as the segment start point; when the score drops back below the threshold, it is determined as the segment end point. To avoid misjudgments caused by instantaneous fluctuations, a time buffer, such as a 0.5-second delay confirmation zone, is set at the threshold boundary to smooth the confidence curve and improve marking accuracy.
[0031] In step S5, for each highlight video clip, the system generates semantic tags based on the corresponding event features and aesthetic features. The semantic tags are in a structured form and include fields such as: "Event category = waving; Aesthetic feature = backlighting; Confidence level = 0.92".
[0032] The tags are either embedded in the video file header as metadata or saved as a separate JSON format tag file, and are synchronized precisely with the original video through timestamps.
[0033] Through the above steps, this embodiment enables the automatic identification and labeling of visually attractive or semantically significant segments during the acquisition of drone aerial videos, reducing the workload of manual screening in the later stages and improving the efficiency of material processing.
[0034] Example 2:
[0035] This embodiment is essentially the same as Embodiment 1, differing only in the system deployment architecture and computing location. In this embodiment, the feature analysis module, decision-making module, and segment and tag management module are all deployed on the UAV's onboard computing unit, while the data storage module can be located at a ground receiving station. Real-time video streams are transmitted via a 5GHz wireless link, and highlight segment tagging data is uploaded synchronously in the form of lightweight metadata. This architecture allows for real-time push of highlight segments to the cloud for direct use by subsequent editing software while analyzing data during flight. This design achieves real-time tagging functionality through collaborative onboard computing and ground storage, reducing the pressure on data backhaul.
[0036] Example 3:
[0037] When drones perform complex aerial photography tasks (such as live sports broadcasts or documentary filming), a dynamic threshold adjustment mechanism can be further introduced. The system adaptively adjusts the highlight confidence threshold based on the real-time scene complexity and ambient lighting conditions. For example, in low-light environments, the weight of "light and shadow features" in aesthetic features is increased; in fast-moving scenes, the weight of stability analysis features is enhanced. Through this dynamic adjustment mechanism, the system can maintain high labeling accuracy and robustness in various shooting scenarios.
[0038] Example 4:
[0039] This embodiment provides a real-time intelligent tagging system for UAV aerial photography that complements the above method. The system includes a data acquisition module, a feature analysis module, a decision-making module, a segment and tag management module, and a data storage module.
[0040] The data acquisition module is responsible for acquiring video stream data from the onboard camera unit; the feature analysis module adopts a multi-model parallel architecture to perform event and aesthetic feature extraction; the decision-making module has a built-in fusion algorithm to generate a highlight confidence score; the segment and tag management module completes the highlight segment delineation and tag generation based on this score; and the data storage module synchronously saves or embeds video and tag metadata. All modules can be connected via an internal bus or Ethernet to form an overall system structure.
[0041] By combining the above systems and methods, a real-time processing mechanism can be achieved for drones to collect, identify, and label aerial footage during flight, thereby enabling automatic labeling and semantic description of highlight segments and significantly improving the retrieval and editing efficiency of video content.
[0042] Obviously, those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps as a single integrated circuit module. Thus, the present invention is not limited to any particular hardware and software combination.
[0043] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A real-time intelligent marking method for unmanned aerial vehicle aerial photography, characterized in that, The method comprises the following steps: S1, receiving real-time video stream data collected by a UAV onboard camera unit; S2, performing multi-channel parallel analysis on the video stream data to extract at least one event feature and at least one aesthetic feature; S3, performing feature fusion based on the extracted event feature and aesthetic feature through a decision fusion model to generate a high-light confidence score that changes over time; S4, delimiting one or more high-light video clips according to dynamic comparison of the high-light confidence score with a preset threshold; S5, generating and associating at least one semantic label for the high-light video clip, the semantic label being generated based on semantic description of the event feature and aesthetic feature and being stored in the form of metadata in association with the video clip.
2. The real-time intelligent tagging method of claim 1, wherein, The event feature comprises at least one of the following: a macro scene type determined through a scene semantic recognition model; a specific action or group behavior identified through a dynamic event detection model; and a specific key target or subject detected through a target recognition model. The aesthetic feature comprises at least one of the following: a special lighting and reflection effect detected through a light and shadow quality analysis model; a picture composition rationality evaluated through a picture composition quality analysis model; and video jitter and fluency evaluated through a picture stability analysis model.
3. The real-time smart labeling method of claim 1, wherein, The decision fusion model in step S3 adopts a joint logic of weighted fusion and rule determination; wherein the weighting coefficients are self-adaptively adjusted according to the confidence of different feature sources and time continuity; and when multiple different types of features coexist in the same time window, the high-light confidence score is adjusted with weight gain.
4. The real-time smart labeling method of claim 1, wherein, The process of delimiting the high-light video clip in step S4 comprises:
5. The real-time smart labeling method of claim 1, wherein, When the high-light confidence score changes from below the preset threshold to above the preset threshold, the time point is determined as the starting point of the clip; When the score changes from above the preset threshold to below the preset threshold, the time point is determined as the ending point of the clip; Wherein, a time buffer is set at the boundary of the preset threshold to filter false positives caused by transient fluctuations. The semantic label generated in step S5 is a structured label, and its content includes: event category, aesthetic feature category and corresponding confidence information.
6. The real-time smart labeling method of claim 1, wherein, The semantic label and high-light clip association information are embedded in the original video file in the form of metadata, or stored in an independent marking file that is synchronously associated with the original video file through a time stamp.
7. The real-time smart labeling method of claim 1, wherein, The method comprises:
8. A real-time intelligent marking system for UAV aerial photography, for implementing the method according to any one of claims 1 to 7, characterized in that, a data acquisition module for receiving real-time video stream data of a UAV onboard camera unit; a feature analysis module for performing multi-channel parallel analysis on the video stream data to extract event features and aesthetic features; a decision determination module for performing fusion on the event features and aesthetic features through a decision fusion model to generate a high-light confidence score; a clip and label management module for delimiting a high-light video clip according to the high-light confidence score and generating an associated semantic label; a data storage module for storing video data and corresponding marking metadata. The feature analysis module, decision determination module and clip and label management module are deployed on a UAV onboard computing unit.
9. The real-time smart labeling system of claim 8, wherein, 10. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the real-time intelligent marking method for UAV aerial photography according to any one of claims 1 to 7.