Panoramic video acquisition unmanned aerial vehicle navigation method and system based on eye movement tracking modeling
By using eye-tracking modeling to record users' gaze behavior in VR, and optimizing drone shooting strategies, the problem of drone shooting being unable to accurately capture areas of interest for users was solved, achieving efficient and personalized panoramic video acquisition of bamboo forests.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INT CENT FOR BAMBOO & RATTAN
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-17
AI Technical Summary
Existing drone-captured panoramic videos of bamboo forests cannot accurately capture areas of interest to users, resulting in a lack of information density in the video content and an inability to dynamically adapt to the differences in attention among different users.
An eye-tracking modeling approach is adopted to record user eye movement data through a VR headset, generate gaze heatmaps and region access frequency maps, optimize drone aerial photography paths and shooting strategies, and build a user attention preference model by combining deep learning and graph neural network structures to adaptively adjust flight altitude, speed and angle.
It achieves a high degree of matching between drone-captured content and user needs, improves the information density and personalized adaptability of video, enhances the intelligence and autonomy of drones in complex environments, and is suitable for intelligent content acquisition in various natural or man-made environments.
Smart Images

Figure CN121876973A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) control technology, and relates to a navigation method and system for panoramic video acquisition UAVs based on eye-tracking modeling. Background Technology
[0002] In recent years, the rise of panoramic video (360° video) and virtual reality (VR) technologies has provided new pathways for immersive presentation of the natural environment, showing great potential, especially in ecotourism, science education, and visual inspection.
[0003] However, panoramic bamboo forest video shooting faces several key challenges: First, drone-generated panoramic videos often follow fixed trajectories or preset paths, making it difficult to accurately capture areas of genuine user interest. This results in some video content lacking information density, while important areas may not be adequately presented. Second, existing shooting processes rely heavily on the photographer's experience or static terrain analysis, failing to incorporate end-user attention and creating a significant disconnect between video content and user needs. Finally, in scenarios such as ecological education, immersive learning, or targeted monitoring, different users exhibit significantly different focuses within the scene, and existing shooting methods cannot dynamically adapt to these varying attentional needs.
[0004] Therefore, there is an urgent need for a mechanism that can learn the regions of interest from users' actual attention behavior and optimize the drone shooting path and content acquisition strategy in reverse, so as to realize a closed loop of "machine acquisition" driven by "user perception".
[0005] Eye tracking, as a crucial technique for studying human visual cognition and interest behavior, has been widely applied in fields such as psychology, user experience evaluation, and advertising. Introducing eye-tracking data into panoramic video analysis of bamboo forests in a VR environment can capture user attention patterns in a spatial dimension, revealing their preferences for scene content. If this attention modeling result can guide drone shooting, it holds promise for a paradigm shift from "passive data collection" to "cognitive-driven active data collection," significantly enhancing the intelligence level of forestry visualization systems. Summary of the Invention
[0006] The purpose of this invention is to solve the problems existing in the prior art and to provide a panoramic video acquisition drone navigation method and system based on eye-tracking modeling.
[0007] To achieve the above objectives, the technical solution adopted by this invention is: a panoramic video acquisition UAV navigation method based on eye-tracking modeling, comprising the following steps: The drone is equipped with a panoramic camera to capture video of the target area; The collected videos are stitched together into a 360° panoramic video, and the stitching result is transformed by spherical projection to unify the spatial coordinate system. Users view a panoramic scene of the target area through a VR headset, while simultaneously recording the user's eye movement data; Eye-tracking data is mapped onto spherical coordinates of panoramic video and preprocessed to generate individual gaze heatmaps and region access frequency maps. Based on gaze heatmaps and region access frequency maps, model user attention regions and predict areas that users may pay attention to; The optimal aerial photography path for the drone is generated based on the prediction results, and the drone's flight altitude, flight speed, and shooting angle are adjusted.
[0008] Preferably, the gaze heatmap reflects the distribution of user attention to each area during viewing by accumulating and spatially smoothing all gaze points on the spherical image; the area access frequency map is based on the image division into areas and counts the number of user visits to quantify the overall attractiveness and behavioral coverage characteristics of the content area.
[0009] Preferably, the method for modeling user-focused regions includes: A deep learning-based semantic segmentation model is used to perform pixel-level content parsing on panoramic video images, extracting key region information with target scene semantics from them; Based on the gaze heatmap and access frequency map, establish the matching relationship between user gaze and semantic regions, and generate a set of attention region labels; A user attention preference model is constructed based on attention mechanisms or graph neural structures. The model takes the semantic region features of the image as input and the user's gaze heatmap and attention region labels as supervision signals. It learns the distribution of the user's interest intensity in the content of the region under different contexts through an attention weighting mechanism, thereby obtaining the user's individualized visual attention preference expression.
[0010] Preferably, the semantic segmentation model is based on an improved image convolutional network, a visual Transformer structure, or a pre-trained multi-scale perceptual network, supporting semantic understanding and structural region segmentation of spherical images.
[0011] Preferably, generating the optimal drone aerial photography path based on the prediction results specifically includes: Based on the predicted distribution of areas of concern Extract the center point location of each region of interest from the map space. And construct the cost function To optimize the path The goal is to minimize the total path length, energy consumption, and execution time, while maximizing coverage of the high-interest region. The optimization objective function is as follows: ; in, Represents the Euclidean distance of a path segment; Indicates the total energy consumption or execution time of the path; For the first User attention weight in each region; and These are adjustable weighting parameters used to balance task cost and content value; A graph search optimization method is used to solve the problem, and the optimal path that satisfies both spatial constraints and flight boundary conditions is output. : .
[0012] Preferably, the drone's flight altitude is adaptively adjusted in the following manner: ; in, Indicates the first The spatial coverage diameter of the region, The scaling factor. Minimum altitude for safe flight; image resolution The following are inversely proportional to height: .
[0013] Preferably, the drone's flight speed is determined based on the complexity of the target scene's texture. With exposure time Adjustment: ; in To adjust the factor, For maximum flight speed, It can be measured by image entropy or edge density.
[0014] Preferably, the shooting angle is determined by a viewpoint optimization algorithm to find the most suitable perspective for displaying content of interest to the user, and a lens shooting direction vector is set for each waypoint. The calculation is based on the fixation prediction point and the target center: ; in, This indicates the current coordinates of the drone.
[0015] Preferably, if the area is a point of interest from multiple angles, then the following approach is adopted: ; Includes the time spent filming. This constitutes a delayed multi-angle acquisition strategy.
[0016] The present invention also provides a panoramic video acquisition UAV navigation system based on eye-tracking modeling. The system includes a computer program, which, when executed, performs the steps of the panoramic video acquisition UAV navigation method based on eye-tracking modeling.
[0017] This invention establishes a human-centered panoramic content acquisition method by combining user eye-tracking modeling with UAV navigation control. It achieves closed-loop optimization of the entire process from user-perceived interest to aerial mission planning. Compared with existing technologies, this invention has the following advantages: First, this invention significantly improves the content matching accuracy and information density of panoramic videos. By analyzing users' gaze behavior towards panoramic videos in a virtual reality environment, the system can accurately identify the areas and content that users are truly interested in, avoiding a large number of invalid images present in traditional fixed-path drone shooting, thereby improving the utilization efficiency and visual value of video resources; Secondly, this invention possesses a high degree of personalization adaptability. The interest model built based on user attention data can be used to predict the differences in scene preferences among different user groups (such as forestry experts, ecotourists, and educational audiences), and dynamically generate differentiated drone shooting strategies accordingly, effectively meeting the personalized content acquisition needs under multiple scenarios and multiple objectives; Furthermore, this invention enhances the intelligence and autonomy of UAV aerial photography. By focusing on prediction results to generate refined flight paths, shooting angles, and time scheduling strategies, the system significantly improves the targeting and efficiency of UAV mission execution in complex natural environments, reduces human intervention, and improves operational quality. Furthermore, this invention possesses excellent scalability and versatility. This method is not only applicable to forest ecosystems such as bamboo groves, but can also be extended to other natural or man-made environments such as farmland inspections, cultural heritage site photography, and post-disaster surveys, enabling broad-based intelligent content acquisition and visual navigation control. Attached Figure Description
[0018] Figure 1 This is a flowchart of the UAV navigation method for panoramic video acquisition of bamboo forests based on eye-tracking modeling in the implementation of this invention; Figure 2 This is a flowchart of the drone navigation method for panoramic video acquisition of bamboo forests based on eye-tracking modeling in the implementation of this invention. Detailed Implementation
[0019] To facilitate understanding of the present invention, the specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and specific examples. The following examples or drawings are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0020] This embodiment uses panoramic video capture of a bamboo forest as an example to illustrate the method of the present invention in detail, such as... Figure 1 and Figure 2 As shown, the specific steps of this method are as follows: I. Preparation stage for panoramic video acquisition of bamboo forest 1.1 The drone, equipped with a panoramic camera, flew and collected video of the bamboo forest area: A drone platform equipped with a panoramic camera was used to capture video of the target bamboo forest area. The panoramic camera uses a multi-view image acquisition structure, capable of acquiring multi-directional video data covering 360 degrees in a single flight. The drone flew stably along a preset flight path, continuously acquiring multi-channel image frame sequences during flight, and simultaneously recording GPS positioning information and flight attitude parameters (including yaw angle, pitch angle, and roll angle) for subsequent image and spatial alignment processing.
[0021] First, the drone is equipped with a panoramic camera to perform autonomous aerial photography missions. The panoramic camera consists of multiple wide-angle lenses facing different directions, which theoretically can cover the entire spatial field of view in each frame captured. During the acquisition of each image frame, the system simultaneously records the global positioning information corresponding to the current frame. and flight attitude information These correspond to the time of the drone. The position and yaw, pitch, and roll angles. These pose parameters will be used in subsequent image stitching and coordinate projection for image alignment and transformation matrix calculation.
[0022] 1.2 Perform image stitching on the video sequence to generate a 360° panoramic video of the bamboo forest: The acquired image sequence undergoes panoramic stitching. Specifically, geometric correction, feature point extraction and matching are performed on multiple images to estimate the spatial transformation relationship between each channel. Then, image stitching techniques such as weighted fusion or multi-band fusion are employed to generate a high-resolution, seamless 360-degree panoramic image sequence. This process can be implemented using image processing algorithms and supports automatic evaluation and correction of stitching quality.
[0023] The acquired multi-channel image frame sequence There is some overlap in the field of view between the two images. Spatial registration between the images is achieved through feature point extraction and matching. Let's assume two images... Corresponding perspectives The key point set is obtained through the SIFT feature extraction algorithm. Then, RANSAC is used to estimate the homography matrix between the two graphs. ,satisfy: ; in, For image The pixel coordinates on For image The corresponding points are constructed by building a global reference frame. Other images are mapped onto this reference frame to form a complete visual sphere. To ensure smooth image seams, a multi-band image fusion strategy is adopted. Each image is decomposed into multi-scale versions through pyramid decomposition and then synthesized in the frequency domain to avoid the accumulation of color difference and overlap errors.
[0024] 1.3 Perform a spherical projection transformation on the stitched result to unify the spatial coordinate system: The stitched panoramic image sequence is then mapped using spherical projection to construct a unified spherical coordinate system. This spherical projection employs an equidistant cylindrical projection method, mapping the pixel coordinates of the two-dimensional images to latitude and longitude coordinates in spherical space. This achieves coordinate alignment between the image and physical space, providing a consistent reference framework for subsequent spatial mapping of eye-tracking data and semantic modeling of regions of interest.
[0025] The stitched image is in equirectangular format, which is an equidistant cylindrical projection (ERP) with latitude and longitude reference. This projection represents points on the sphere. Mapped to two-dimensional coordinates on the image ,in, and These represent the width and height of the ERP, respectively. Their mapping relationship is as follows: ; The reverse mapping is: ; If you need to convert it to a unit vector in a three-dimensional spherical coordinate system Then we have: .
[0026] This coordinate transformation ensures that every pixel in the panoramic image has a clear spherical orientation, which facilitates subsequent eye-tracking point projection mapping, spatial region of interest labeling, and viewpoint optimization navigation.
[0027] The output of this stage is a sequence of 360° panoramic image frames that are spatially uniform, have known poses, and can be mapped to coordinates, denoted as: ; This structure serves as the spatial visual foundation for subsequent eye-tracking modeling, user preference analysis, and navigation control.
[0028] II. User Viewing and Eye-Tracking Data Processing Stage 2.1 Users view a panoramic view of the bamboo forest through a VR headset: After completing the stitching and spherical projection of the panoramic video of the bamboo forest, the processed panoramic video of the bamboo forest will be... Loaded into virtual reality devices to enable immersive playback and interactive browsing. Virtual reality devices can be VR headsets with head posture tracking and eye tracking capabilities, such as head-mounted displays equipped with eye-tracking sensors. After wearing the VR device, users can freely rotate their viewpoint in three-dimensional space and view the bamboo forest scene. The system renders images in real time that match the user's perspective, creating an immersive experience of the bamboo forest environment.
[0029] Complete the panoramic video of the bamboo forest After splicing and spherical projection processing, the system will Loaded into virtual reality devices with head posture tracking and eye tracking capabilities, it achieves an immersive 3D visual experience. The HTC VivePro Eye VR device, with its built-in infrared camera and depth sensing module, can track the user's head posture and gaze direction in real time. After wearing the VR headset, the user can freely rotate their viewpoint on the spherical panoramic image, and the system tracks the head posture angle... The rendered image is adjusted accordingly to achieve dynamic synchronization between the user's head movement and the direction of their field of vision.
[0030] During the user's viewing experience, the system collects the user's eye movement data in real time, recording their gaze trajectory on the spherical image. Each eye movement record can be represented as a quadruple: ; in, For the first User gaze point on frame image Two-dimensional pixel coordinates; Indicates the duration of the gaze; For this gaze point The corresponding timestamp for the panoramic video of the bamboo forest.
[0031] 2.2 Projecting and mapping eye-tracking data onto the spherical coordinates of the bamboo forest panoramic video: The gaze coordinates in the user's eye-tracking data are converted to their corresponding positions in a spherical projection coordinate system. Since the bamboo forest panoramic video uses equirectangular projection, the two-dimensional pixel coordinates of each frame can be directly mapped to latitude and longitude coordinates on a sphere. Therefore, the screen coordinates in the eye-tracking data can be converted to spherical coordinates based on the image size and projection rules. This enables accurate positioning of the viewpoint on the spherical image.
[0032] Since the panoramic image is an equirectangular projection, the system will use two-dimensional gaze coordinates. Mapped to spherical orientation coordinates The mapping formula is: ; .
[0033] Furthermore, the gaze point is converted into a unit spherical vector form: .
[0034] 2.3 Perform time-series registration with the corresponding frame image: To ensure consistency between eye-tracking data and panoramic video frames of the bamboo forest, the system performs temporal registration of gaze point timestamps with the video frame sequence. By comparing the recording time of the eye-tracking data with the playback time of the video frames, the system establishes the video frame number corresponding to each gaze point and binds its projection position to the corresponding frame image, thereby achieving spatiotemporally unified gaze data mapping.
[0035] To achieve accurate registration between eye-tracking data and panoramic video frames of the bamboo forest, the system is based on the time series of video frames. With sampling timestamp A mapping function is established using linear interpolation or frame number matching: ,satisfy Minimum; This makes each gaze vector Each corresponds to a unique video frame. The image perspective within that frame.
[0036] 2.4 Generate individual gaze heatmaps and region visit frequency maps: After completing spatial mapping and temporal alignment, the system constructs individual gaze heatmaps and region access frequency maps based on eye-tracking trajectories. The gaze heatmap reflects the distribution of user attention to different regions during viewing by accumulating and spatially smoothing all gaze points on a spherical image. The region access frequency map, based on image region division, counts the number of user visits to quantify the overall attractiveness and behavioral coverage characteristics of content regions. These intermediate representations serve as input features for subsequent user attention preference modeling and interest region prediction, further guiding the generation of optimized UAV content acquisition strategies.
[0037] Subsequently, according to all Constructing an individual gaze heatmap using vectors and gaze duration Heatmaps reflect the density of user attention on a spherical image. Gaussian kernel smoothing implementation for points: ; in, It is a two-dimensional Gaussian kernel. for Point to the corresponding position.
[0038] In addition, based on image spatial partitioning, the system counts the number of times user "gaze" visits each region and generates a region access frequency map. It serves as an auxiliary feature in subsequent modeling of the region of interest.
[0039] Ultimately, the eye-tracking data is organized into a structured sequence as follows: .
[0040] This eye-tracking dataset will serve as the core behavioral signal for subsequent heatmap construction, interest modeling, and navigation optimization, reflecting users' true visual attention preferences for bamboo forest scenes in the virtual environment.
[0041] III. Focus on the Regional Modeling and Interest Prediction Stage After completing the spatial mapping and visualization of user eye-tracking data, the next stage is attention region modeling and interest prediction. This stage aims to extract semantically meaningful content regions from the panoramic bamboo forest video and, in conjunction with the user's gaze behavior, establish a personalized attention preference model to predict the distribution of shooting areas that the user may be interested in in new scenes.
[0042] 3.1 Extracting semantic regions from the panorama using a semantic segmentation model: The system employs a deep learning-based semantic segmentation model to perform pixel-level content parsing of panoramic bamboo forest video images, extracting key regional information with forestry scene semantics. This includes, but is not limited to: different types of bamboo varieties, vegetation areas with disease spots or yellowing, areas visible through sunlight, ground paths, forest clearings, and water bodies. The semantic segmentation model can be based on improved image convolutional networks, visual Transformer structures, or pre-trained multi-scale perceptual networks, supporting semantic understanding and structural region segmentation of spherical images.
[0043] First, the system processes the panoramic image. Semantic segmentation is performed to extract representative semantic regions from the forestry scene. This semantic segmentation model is denoted as: ; in, This refers to semantic segmentation networks based on deep learning (such as SegFormer). Number of categories (e.g., lesions, sunlight, ground, etc.) Each pixel in Represents pixels The semantic category to which it belongs.
[0044] 3.2 Constructing a set of tags for regions of interest: After semantic region identification, the system, combining the aforementioned gaze heatmap and access frequency map, establishes a matching relationship between user gaze and semantic regions, generating a set of attention region labels. This label set is used to identify which types or attributes of regions the user shows a higher tendency to pay attention to during viewing, serving as supervisory labels for subsequent preference modeling.
[0045] The system then displays the completed gaze heatmap. With semantic regions Perform cross-statistics to construct a set of tags for areas of interest. Each of them Corresponding to a semantic category Rather than accumulating gaze intensity in the heatmap: .
[0046] By normalizing, the relative attention weight of each type of region can be obtained: .
[0047] Tag set It can be used as a supervisory signal for user visual preferences to train user attention models.
[0048] 3.3 Establishing a user attention preference model based on the attention mechanism: The system constructs a user attention preference model based on an attention mechanism or graph neural network structure. This model takes image semantic region features as input and the user's gaze heatmap and attention tags as supervision signals. Through an attention weighting mechanism, it learns the distribution of user interest intensity towards regional content under different contextual conditions, thereby obtaining a personalized expression of the user's visual attention preferences.
[0049] Next, the system constructs a user preference modeling network based on an attention mechanism. Let each semantic region... The region feature vector is (Derived from the aggregation of Transformer features within the region), the system constructs an attention-weighted model to learn the user's interest response to different regions: ; ; in, This is a user preference query vector, which can be generated from historical gaze features or individual embeddings. For learnable parameter matrix, Indicates attention weights. This is a weighted preference representation used to reflect the overall attention structure of users across all regions of the image.
[0050] 3.4 Predicting the distribution of shooting areas that users may be interested in: The model can predict the distribution of shooting areas that a user might be interested in in new, unobserved scenes, based on contextual information such as the semantic features of the input region, spatial location, and ambient lighting. This prediction can guide drones to prioritize covering areas of high user interest in subsequent aerial photography, achieving optimized content acquisition based on cognitive preferences.
[0051] In the prediction phase, for new images The system extracts its regional semantic feature set. And input it into the attention model above: ; Obtain the predicted weight distribution of regions that the user may be interested in in the image. .
[0052] Based on the Top-K screening strategy, the predicted set of areas of interest can be obtained: .
[0053] The results are used to guide drones to focus on capturing or dynamically plan routes for predicted areas of interest, thereby enabling panoramic content acquisition driven by user cognitive preferences.
[0054] IV. UAV Navigation Strategy Generation Stage After completing the modeling and interest prediction of the user's area of interest, the system enters the UAV navigation strategy generation stage. This stage aims to adaptively optimize the UAV's flight path and shooting parameters based on the user preference prediction results, thereby realizing a personalized content collection task guided by user interests.
[0055] 4.1 Generate the optimal aerial photography path based on the predicted area of interest: The system first generates aerial photography paths covering the optimal areas of interest based on the predicted distribution of high-interest areas, combined with terrain data and flight constraints. During path generation, the system employs a goal-driven path optimization algorithm to maximize coverage of the user's high-interest areas while avoiding redundant or low-value shooting areas. It also considers factors such as spatial continuity, energy minimization, and execution efficiency to ensure that the generated paths strike a balance between flight safety and mission efficiency.
[0056] First, the system predicts the distribution of areas of interest. Extract the center point location of each region of interest from the map space. And construct the cost function To optimize the path This minimizes the total path length, energy consumption, and execution time, while maximizing coverage of high-interest regions. The optimized cost function takes the following form: ; in: Represents the Euclidean distance of a path segment; Indicates the total energy consumption or execution time of the path; For the first User attention weight in each region; and This is an adjustable weighting parameter used to balance task cost and content value.
[0057] This path optimization problem is solved using a graph search optimization method (such as Dijkstra's algorithm), which outputs the optimal path that satisfies both spatial constraints and flight boundary conditions. : .
[0058] 4.2 Plan the flight altitude, speed, shooting angle, and path: The system further performs joint planning of key parameters for the flight mission, including flight altitude, flight speed, shooting angle, and lens field of view. Flight altitude is adjusted based on the spatial density of the target area and the desired image resolution; flight speed is dynamically adjusted according to the complexity of the area and lens exposure requirements; and the shooting angle is determined by a viewpoint optimization algorithm to find the most suitable perspective for displaying content of interest to the user. When necessary, a "staying shot" strategy can be introduced to delay or capture images from multiple angles over specific high-interest areas to enhance the expressiveness of the video content.
[0059] In obtaining the optimal path Based on this, the system further jointly plans flight parameters, including flight altitude. Flight speed Camera tilt angle Yaw angle This is done to ensure optimal image acquisition. The flight altitude can be adaptively adjusted as follows: ; in, Indicates the first The spatial coverage diameter of the region, The scaling factor. Minimum altitude for safe flight. (Image resolution) The following are inversely proportional to height: .
[0060] The shooting speed depends on the complexity of the target scene texture. With exposure time Adjustment: ; in, To adjust the factor, For maximum flight speed, It can be measured by image entropy or edge density.
[0061] At the same time, the system sets the camera shooting direction vector for each waypoint. The calculation is based on the predicted focus and the target center: ; in, This indicates the current coordinates of the drone.
[0062] If the area is a point of interest from multiple angles, then the following approach will be adopted: ; Includes the time spent filming. This constitutes a delayed multi-angle acquisition strategy.
[0063] 4.3 Output complete mission instructions to guide the drone to perform the shooting mission again: The system packages the generated flight path and shooting parameters into a complete task instruction set, including waypoint sequences, flight attitude parameters, and shooting control commands, and sends them to the aircraft through the UAV ground control system. After receiving the task instructions, the UAV can autonomously complete the aerial photography task based on user-focused modeling optimization, achieving closed-loop optimization of the human-centered panoramic bamboo forest content acquisition process.
[0064] The system packages the path Π* and its corresponding joint parameter configuration into a task instruction set. The format is as follows: .
[0065] This instruction set is transmitted from the ground control system to the UAV flight control module, which then controls the UAV to complete high-quality panoramic aerial photography missions according to user-focused optimization strategies. The entire process achieves closed-loop adaptive optimization from "user visual modeling" to "machine shooting execution," significantly improving the targeting, efficiency, and user value of content acquisition.
[0066] 5. Personalized content collection and feedback stage 5.1 The drone re-flies and takes pictures based on the navigation results. After generating and issuing an optimized navigation strategy, the drone re-executes the aerial photography mission according to the strategy, completing personalized content collection targeting the user's areas of interest. During flight, the drone automatically adjusts its flight altitude, speed, and shooting angle based on previously generated waypoint paths and control commands, focusing on covering areas of predicted user interest to acquire image or video data with high information density. The system supports advanced control strategies such as multiple shots, multi-angle acquisition, and localized hovering to enhance the richness and interactivity of expression for areas of high interest.
[0067] After completing the navigation strategy generation and command issuance, the UAV follows the mission instruction set. The aerial photography mission resumed, entering the personalized content collection phase based on user preferences. During the flight, the system controlled the drone to follow the planned path. Automatically adjust flight altitude Flight speed With shooting direction Ensure that at each focus prediction point High-quality image acquisition was completed at all locations. For multi-angle or key areas, the system employs a multi-angle dwell strategy to control the UAV to perform operations in that area. Each shoot will last for [number] hours. This forms a multi-angle acquisition sequence: .
[0068] All image frame data and their pose information are synchronously packaged into: .
[0069] This dataset serves as the content collection result for high-interest areas, providing raw materials for subsequent video generation.
[0070] 5.2 Generate optimized, high-profile panoramic videos of the bamboo forest: After completing the aerial photography mission, the system performs panoramic stitching and spherical mapping on the newly acquired video to generate an optimized, high-profile panoramic video of the bamboo forest. Compared to the initial version, this video is more aligned with user preferences in terms of content expression, has higher viewing value and information utilization, and is suitable for various downstream scenarios such as teaching demonstrations, ecological tours, or forestry analysis.
[0071] After the task is completed, the system performs image stitching and spherical mapping on the image sequence I. The process is the same as the initial stage, but the input is a densely sampled image of the region of interest, and the stitched output image I′ has higher spatial redundancy and content focus density. The system uses spherical projection mapping: ; .
[0072] Generate a new video frame sequence The final, highly anticipated panoramic video of the bamboo forest is generated by compressing the video using a video encoder (such as H.265). The video significantly improves the overlap rate with the user's focus point in terms of spatial distribution. Its average gaze overlap can be calculated using the following metrics: ; in, A gaze heatmap generated during user viewing reflects the intensity of the user's visual attention. To generate a true viewpoint mask in the video, The larger the value, the better the video matches the user's interests.
[0073] 5.3 Supports user re-viewing and iterative optimization of modeling (closed-loop feedback mechanism): Meanwhile, to achieve continuous collaboration between user behavior and drone data collection, the system further supports users to view the optimized panoramic bamboo forest video a second time and collect eye-tracking data. By collecting and comparing the user's new gaze behavior in the optimized video, the system can update the user attention model, achieve iterative optimization of preference expression and adaptive evolution of system performance, forming a closed-loop optimization mechanism of "viewing—modeling—prediction—collection—feedback", and continuously improve the intelligence level and user experience of the personalized aerial photography system.
[0074] To enable continuous system optimization and personalized learning, the system supports users in... The system collects re-viewing and fixation data to generate a new round of eye movement recordings (𝔊′). The system compares the new eye movement data with the predicted values from the previous model to construct the attention error: ; in, To predict heatmaps for the model, Generate a heatmap for new viewing, error Loss function used to update preference modeling networks : ; in, and To control the hyperparameters of its weights, the system continuously iterates on the user interest modeling network. Thus, the system forms a complete cognitive loop of "viewing → modeling → shooting → optimizing → re-viewing," continuously improving the accuracy and robustness of personalized data collection and achieving the synergistic evolution of human-like perception and UAV navigation systems.
Claims
1. A panoramic video acquisition drone navigation method based on eye tracking modeling, characterized in that, Includes the following steps: The drone is equipped with a panoramic camera to capture video of the target area; The collected videos are stitched together into a 360° panoramic video, and the stitching result is transformed by spherical projection to unify the spatial coordinate system. Users view a panoramic scene of the target area through a VR headset, while simultaneously recording the user's eye movement data; Eye-tracking data is mapped onto spherical coordinates of panoramic video and preprocessed to generate individual gaze heatmaps and region access frequency maps. Based on gaze heatmaps and region access frequency maps, model user attention regions and predict areas that users may pay attention to; The optimal aerial photography path for the drone is generated based on the prediction results, and the drone's flight altitude, flight speed, and shooting angle are adjusted. 2.The panoramic video acquisition drone navigation method based on eye movement tracking modeling according to claim 1, wherein, The gaze heatmap reflects the distribution of user attention to different areas during viewing by accumulating and spatially smoothing all gaze points on a spherical image; the area access frequency map is based on the division of the image into regions and counts the number of times users access the content area, which is used to quantify the overall attractiveness and behavioral coverage characteristics of the content area.
3. The panoramic video acquisition UAV navigation method based on eye-tracking modeling according to claim 1, characterized in that, Methods for modeling user-focused regions include: A deep learning-based semantic segmentation model is used to perform pixel-level content parsing on panoramic video images, extracting key region information with target scene semantics from them; Based on the gaze heatmap and access frequency map, establish the matching relationship between user gaze and semantic regions, and generate a set of attention region labels; A user attention preference model is constructed based on attention mechanisms or graph neural structures. The model takes the semantic region features of the image as input and the user's gaze heatmap and attention region labels as supervision signals. It learns the distribution of the user's interest intensity in the content of the region under different contexts through an attention weighting mechanism, thereby obtaining the user's individualized visual attention preference expression.
4. The panoramic video acquisition UAV navigation method based on eye-tracking modeling according to claim 3, characterized in that, The semantic segmentation model described above is based on an improved image convolutional network, a visual Transformer structure, or a pre-trained multi-scale perceptual network, and supports semantic understanding and structural region segmentation of spherical images.
5. The panoramic video acquisition UAV navigation method based on eye-tracking modeling according to claim 1, characterized in that, The method for generating the optimal aerial photography path for drones is as follows: Based on the predicted distribution of areas of concern Extract the center point location of each region of interest from the map space. And construct the cost function To optimize the path This minimizes the total path length, energy consumption, and execution time, while maximizing coverage of the high-interest region; the cost function is as follows: ; in, Represents the Euclidean distance of a path segment; Indicates the total energy consumption or execution time of the path; For the first User attention weight in each region; and These are adjustable weighting parameters used to balance task cost and content value; A graph search optimization method is used to solve the problem, and the optimal path that satisfies both spatial constraints and flight boundary conditions is output. : 。 6. The panoramic video acquisition UAV navigation method based on eye-tracking modeling according to claim 5, characterized in that, The drone's flight altitude is adaptively adjusted as follows: ; in, Indicates the first The spatial coverage diameter of the region, The scaling factor. Minimum altitude for safe flight; image resolution The following are inversely proportional to height: 。 7. The panoramic video acquisition UAV navigation method based on eye-tracking modeling according to claim 5, characterized in that, The drone's flight speed depends on the complexity of the target scene texture. With exposure time Adjustment: ; in To adjust the factor, For maximum flight speed, Measured by image entropy or edge density.
8. The panoramic video acquisition UAV navigation method based on eye-tracking modeling according to claim 5, characterized in that, The shooting angle is determined by a viewpoint optimization algorithm to find the most suitable perspective for displaying content that the user is interested in, and a lens shooting direction vector is set for each waypoint. The calculation is based on the fixation prediction point and the target center: ; in, This indicates the current coordinates of the drone.
9. The UAV navigation method for panoramic video acquisition of bamboo forests based on eye-tracking modeling according to claim 8, characterized in that, If the area is a point of interest from multiple angles, then the following approach will be adopted: ; Includes the time spent filming. This constitutes a delayed multi-angle acquisition strategy.
10. A panoramic video acquisition UAV navigation system based on eye-tracking modeling, characterized in that, The system includes a computer program that, when executed, performs the steps of the method as described in any one of claims 1-9.