Multimodal-based unmanned aerial vehicle ski field high-precision spatiotemporal data acquisition method and system

CN122835375APending Publication Date: 2026-09-29CHINA INST OF SPORT SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611086178.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0006]本发明提供一种基于多模态的无人机雪场高精度时空数据采集方法及系统,其通过多传感器融合和特定场景优化,解决雪场环境中定位精度低、数据采集不完整的技术问题,能够在复杂雪场环境中对滑雪者进行高精度定位与数据采集

Benefits of technology

第一、本发明通过多传感器融合与雪场专用图像增强处理,配合闭环曝光控制和视觉惯性紧耦合位姿估计,在雪地高反光、低纹理环境下实现厘米级定位精度,生成标准化4D时空数据包,兼顾了高精度与强环境适应性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835375A_ABST
    Figure CN122835375A_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned plane ski field high-precision space-time data acquisition method and system based on multi-modal, it is related to unmanned plane technical field.The method of the present application synchronously collects data by binocular vision camera, IMU inertial measurement unit and RTK-GPS module carried on unmanned plane electronic nacelle, exposure parameter is adjusted in closed loop during acquisition process, special image enhancement processing is performed to image, visual feature point is extracted and visual inertial tight coupling pose estimation is performed, skier is detected and human skeleton key point containing foot key point is extracted, three-dimensional coordinates are calculated using binocular parallax and space-time consistency matching is carried out, and the 4D space-time data package containing microsecond level time stamp and skeletal motion posture is generated by inputting multi-source data into fusion filter.The present application mainly solves the problems of low positioning accuracy and incomplete data acquisition in ski field environment, and can perform high-precision positioning and data acquisition on skiers in complex ski field environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology. More specifically, this invention relates to a multimodal UAV-based method and system for high-precision spatiotemporal data acquisition in ski resorts. Background Technology

[0002] In outdoor sports settings such as ski resorts, precise positioning and data collection of athletes are of significant application value. These scenarios typically require obtaining accurate three-dimensional motion trajectories and posture information of skiers for motion analysis, technical training, or safety monitoring. Drones, with their aerial maneuverability and flexible perspective, have made data collection at ski resorts possible. Currently, drone low-altitude remote sensing and visual-inertial fusion sensing technologies are widely used in outdoor spatial information acquisition, moving target detection, and posture calculation. By working collaboratively with multiple sensors such as visual cameras, inertial sensors, and satellite positioning, target image acquisition, carrier pose calculation, and motion state analysis can be achieved, making it a primary technical means for data collection at ski resorts.

[0003] However, the unique environmental characteristics of ski resorts present numerous practical challenges to this type of data collection, leading to significant limitations in existing technological solutions. Traditional GPS positioning suffers from weak signals and low accuracy in ski resort environments, making it difficult to meet the precise positioning needs of athletes. Therefore, the industry often uses visual positioning combined with inertial sensing to compensate for the shortcomings of GPS. However, the highly reflective and large-area, low-texture characteristics of snow surfaces make it difficult for conventional visual feature extraction algorithms to obtain a sufficient number of effective feature points, significantly reducing the stability of visual positioning. This is one of the core reasons why existing vision-based positioning systems generally suffer from insufficient positioning accuracy in the specific and complex environment of ski resorts.

[0004] During actual data collection at ski resorts, skiers move at high speeds and follow unpredictable trajectories, often involving sharp turns and jumps, placing high demands on the real-time performance of data acquisition and the stability of tracking. Existing conventional UAV data acquisition solutions mostly employ single visual sensors or simple multi-sensor combinations, lacking specific optimization for the high-speed dynamics of ski resorts and effective multi-sensor synchronization mechanisms. This results in significant temporal discrepancies in various sensor data, making it difficult to achieve real-time and stable tracking of skiers' movements. Furthermore, the global brightness and localized shadows of snow make conventional camera exposure strategies unsuitable, easily leading to overexposure or loss of detail. Traditional image enhancement methods are also not optimized for ski resort scenarios, failing to suppress imaging interference from snow reflections, further impacting positioning accuracy and data acquisition quality.

[0005] Currently, existing technical solutions generally suffer from insufficient positioning accuracy, poor data acquisition stability, and low real-time performance when applied in ski resort scenarios. In particular, under conditions such as lack of texture in snow, low lighting, and high dynamic motion blur, the performance of traditional visual algorithms will drop significantly, and even tracking will be interrupted. They cannot reliably obtain accurate three-dimensional motion trajectory and posture information of skiers, making it difficult to meet the actual needs of ski resort motion analysis, technical training, and safety monitoring. Summary of the Invention

[0006] This invention provides a method and system for high-precision spatiotemporal data acquisition in ski resorts based on multimodal drones. By using multi-sensor fusion and specific scene optimization, it solves the technical problems of low positioning accuracy and incomplete data acquisition in ski resort environments, enabling high-precision positioning and data acquisition of skiers in complex ski resort environments.

[0007] To achieve these objectives and other advantages according to the present invention, a method for high-precision spatiotemporal data acquisition of ski resorts using unmanned aerial vehicles (UAVs) based on multimodal perception is provided, comprising the following steps: Step A: Using a binocular vision camera, an IMU inertial measurement module, and an RTK-GPS module mounted on the UAV's electronic pod, the original image sequence, acceleration and angular velocity data, and geographic location data are collected synchronously according to a unified hardware clock. The electronic pod has a built-in edge computing unit that adjusts the exposure parameters of the binocular vision camera in a closed loop based on the real-time calculated snow surface brightness statistics during the acquisition process. Step B: Perform ski resort-specific image enhancement processing on the acquired original image sequence, including at least: extracting non-snow areas in the image as regions of interest, and performing contrast-limited adaptive histogram equalization on these regions to obtain enhanced left and right eye images; Step C: Extract visual feature points from the enhanced left and right eye images and perform binocular matching. Combine acceleration and angular velocity data to perform visual-inertial tightly coupled pose estimation and calculate the visual odometry pose of the UAV in the world coordinate system. Step D: Detect skier targets in the enhanced left and right eye images, extract the two-dimensional coordinates of human skeletal key points including foot key points in the image coordinate system, calculate the three-dimensional coordinates of skeletal key points in the camera coordinate system using binocular parallax, and match and assign identity markers to skeletal key points between adjacent frames based on spatiotemporal consistency. Step E: Input the visual odometry pose, acceleration and angular velocity data, geographic location data, and the 3D coordinates of the skeletal key points in the camera coordinate system into the multi-source sensor fusion filter, perform timestamp alignment and coordinate system unification, and output the 3D spatial coordinates of the UAV and each skier in the world coordinate system and the skeletal motion posture data, generating a 4D spatiotemporal data package including microsecond-level timestamps, 3D spatial coordinates, and skeletal motion posture data.

[0008] Preferably, the image enhancement process in step B further includes: The camera motion blur radius and direction are estimated using the acceleration and angular velocity data, the point spread function is reconstructed, and deconvolution processing is performed using Wiener filtering.

[0009] Preferably, step C further includes a multimodal failure compensation step based on health assessment: Visual modal health and IMU modal health are calculated in real time. Visual modal health is calculated based on the number of successfully matched feature points and their distribution variance in the image plane. IMU modal health is calculated based on the IMU pre-integration residual and temperature drift rate. When the visual modal health is below the first threshold and the IMU modal health is not below the second threshold, virtual visual observations are generated using acceleration and angular velocity data and injected into the visual-inertial tightly coupled pose estimation process. When the IMU modal health is lower than the second threshold and the visual modal health is not lower than the first threshold, virtual IMU measurements are generated using the triangulation results of visual feature points in consecutive frames and injected into the visual-inertial tightly coupled pose estimation process. When the visual modal health is below the first threshold and the IMU modal health is below the second threshold, the centroid motion data of the key points of the human skeleton are extracted, kinematic constraints are constructed in combination with the prior knowledge of human movement biomechanics, the pose result obtained by short-time integration of the IMU is corrected, and short-time stable pose estimation is maintained until any sensor modality is recovered.

[0010] Preferably, step C further includes an active structured light-assisted feature extraction step: The electronic pod also includes a near-infrared VCSEL structured light projector and a narrowband filter for the corresponding wavelength; When the visual modality health is below the weak texture preset threshold, the near-infrared VCSEL structured light projector is activated. The density of the projection grating is adaptively adjusted according to the real-time flight altitude of the drone to project random speckle or coded grid patterns onto the snow surface. Structured light images are acquired using narrowband filters, and speckle centers or grid corners are extracted as visual feature points for binocular matching and pose estimation. The emission power of the near-infrared VCSEL structured light projector is adaptively adjusted based on ambient light intensity and image contrast.

[0011] Preferably, the key points of the human skeleton in step D include the tip of the nose, neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, left and right heels, left and right toes, and the outer sides of left and right heels, totaling 20 key points.

[0012] Preferably, the model used in step D for detecting skier targets and extracting key points of human skeleton is a lightweight deep learning model, which is trained through cloud knowledge distillation. The Transformer model is used as a high-parameter teacher network, pre-trained on a dataset including skiing postures and snow obstacles, and high-dimensional feature response maps are extracted. The lightweight deep learning model is used as a student network, with the same image as input and the corresponding feature layer output is extracted. Training is completed by minimizing the KL divergence between the corresponding feature layer output of the student network and the high-dimensional feature response map of the teacher network.

[0013] Preferably, it also includes an active feedforward gimbal control step based on human skeletal kinematics prediction: Using multi-frame human skeleton keypoint data stored in history, a simplified kinematic model of the skier, including center of mass position, velocity, and acceleration, is established. State estimation of a skier’s simplified kinematic model is performed using unscented Kalman filtering to predict the skier’s three-dimensional position, velocity, and acceleration at the next moment. Based on the pre-integration results of the IMU in the electronic pod and the motor speed data fed back by the UAV flight control, the attitude and rotational angular velocity of the UAV at the next moment are predicted. Transform the skier's predicted position and the drone's predicted position to the same world coordinate system, calculate the relative position vector between the two, and calculate the gimbal's expected pointing angle and feedforward compensation angular velocity based on the relative position vector. The feedforward compensation angular velocity is superimposed on the output of the PID feedback control loop of the gimbal motor.

[0014] A high-precision spatiotemporal data acquisition system for ski resorts based on multimodal perception, used to execute the method described above, the system comprising: Unmanned aerial vehicle (UAV) platform; An electronic pod is mounted on the UAV platform. The electronic pod includes a sensor unit and an edge computing unit. The sensor unit includes a binocular vision camera, an IMU inertial measurement module, and an RTK-GPS module. The edge computing unit is electrically connected to the sensor unit. The sensor unit is used to synchronously acquire raw image sequences, acceleration and angular velocity data, and geographic location data according to a unified hardware clock, and transmit the raw image sequences, acceleration and angular velocity data, and geographic location data to the edge computing unit; The edge computing unit includes an image enhancement module, an exposure control module, a visual feature extraction and matching module, a visual-inertial tightly coupled pose estimation module, a target detection and skeleton extraction module, a spatiotemporal consistency matching module, and a multi-source sensor fusion module. The image enhancement module is used to perform ski resort-specific image enhancement processing on the original image sequence, including extracting non-snow areas in the image as regions of interest and performing contrast-limited adaptive histogram equalization on the regions, and outputting enhanced left and right eye images; The exposure control module is connected to the image enhancement module and the binocular vision camera respectively. It is used to receive the snow surface brightness statistics calculated based on the original image sequence from the image enhancement module, and generate an exposure adjustment command based on the statistics and output it to the binocular vision camera to adjust the exposure parameters of the binocular vision camera in a closed loop. The visual feature extraction and matching module is connected to the image enhancement module, receives the enhanced left and right eye images, and is used to extract visual feature points from the enhanced left and right eye images and perform binocular matching. The visual-inertial tightly coupled pose estimation module is connected to the visual feature extraction and matching module and the IMU inertial measurement module. It receives the matching results of visual feature points as well as acceleration and angular velocity data, and is used to perform visual-inertial tightly coupled pose estimation to calculate the visual odometry pose of the UAV in the world coordinate system. The target detection and skeleton extraction module is connected to the image enhancement module, receives the enhanced left and right eye images, and is used to detect skier targets in the enhanced left and right eye images, extract the two-dimensional coordinates of human skeleton key points including foot key points in the image coordinate system, and use binocular parallax to calculate the three-dimensional coordinates of skeleton key points in the camera coordinate system. The spatiotemporal consistency matching module is connected to the target detection and skeleton extraction module, and receives the skeletal key points of adjacent frames, which are used to match the skeletal key points of adjacent frames and assign identity identifiers based on spatiotemporal consistency. The multi-source sensor fusion module is connected to the visual-inertial tightly coupled pose estimation module, the IMU inertial measurement module, the RTK-GPS module, and the target detection and skeleton extraction module. It receives visual odometry pose, acceleration and angular velocity data, geographic location data, and three-dimensional coordinates of skeletal key points. It is used to perform timestamp alignment and coordinate system unification, output the three-dimensional spatial coordinates of the UAV and each skier in the world coordinate system, as well as the skeletal motion posture data, and generate a 4D spatiotemporal data packet including microsecond-level timestamps, three-dimensional spatial coordinates, and skeletal motion posture data. A ground control station, which is communicatively connected to the electronic pod, is used to receive and display real-time data sent by the electronic pod. The cloud server is communicatively connected to the ground control station and / or the electronic pod, and is used to receive 4D spatiotemporal data packets, perform big data analysis and model optimization, and send the updated model to the edge computing unit.

[0015] Preferably, the edge computing unit further includes a motion blur compensation module and a multimodal failure compensation module; The motion blur compensation module is connected to the sensor module and the image enhancement module. It is used to receive acceleration and angular velocity data, estimate the camera motion blur radius and direction, reconstruct the point spread function, perform deconvolution processing using Wiener filtering, and output the processed image sequence to the image enhancement module. The multimodal failure compensation module is connected to the visual-inertial tightly coupled pose estimation module, the visual feature extraction and matching module, the IMU inertial measurement module, and the target detection and skeleton extraction module. It is used to calculate visual modal health and IMU modal health in real time. Visual modal health is calculated based on the number of successfully matched feature points and their distribution variance in the image plane. IMU modal health is calculated based on the IMU pre-integration residual and the temperature drift rate. When the visual modal health is below the first threshold and the IMU modal health is not below the second threshold, virtual visual observations are generated using acceleration and angular velocity data and injected into the visual-inertial tightly coupled pose estimation module. When the IMU modal health is lower than the second threshold and the visual modal health is not lower than the first threshold, virtual IMU measurements are generated using the triangulation results of visual feature points in consecutive frames and injected into the visual-inertial tightly coupled pose estimation module. When the visual modal health is below the first threshold and the IMU modal health is below the second threshold, the centroid motion data of key points of the human skeleton are extracted, kinematic constraints are constructed by combining human movement biomechanical priors, the pose results obtained by short-time integration of the IMU are corrected, and short-time stable pose estimation is maintained until any sensor modality is recovered.

[0016] Preferably, the electronic pod also includes a near-infrared VCSEL structured light projector, a narrowband filter of the corresponding wavelength, a gimbal servo control module, and an environmental sensor; The near-infrared VCSEL structured light projector is connected to the edge computing unit and is used to receive the projection control command output by the edge computing unit to project random speckle or coded grid patterns onto the snow surface. The narrowband filter is installed at the front of the lens of the binocular vision camera to filter ambient light interference and acquire structured light images. The gimbal servo control module is connected to the edge computing unit and the three-axis motor of the electronic pod, and is used to receive gimbal control commands output by the edge computing unit and drive the three-axis motor to rotate. The environmental sensor is connected to the edge computing unit and is used to synchronously collect ski resort environmental parameters, including at least one of temperature, humidity, wind speed, snow depth and visibility. The edge computing unit also includes an active structured light-assisted feature extraction module, an active feedforward gimbal control module, and an environmental data integration module; The active structured light-assisted feature extraction module is connected to the near-infrared VCSEL structured light projector, the binocular vision camera, and the visual feature extraction and matching module. When the visual modality health is lower than the weak texture preset threshold, the module activates the near-infrared VCSEL structured light projector, adaptively adjusts the projection grating density according to the real-time flight altitude of the UAV, extracts speckle centers or grid corners as visual feature points, performs binocular matching and pose estimation, and adaptively adjusts the emission power of the near-infrared VCSEL structured light projector according to the ambient light intensity and image contrast. The active feedforward gimbal control module is connected to the target detection and skeleton extraction module, the IMU inertial measurement module, the flight control system of the UAV platform, and the gimbal servo control module. It is used to establish a simplified kinematic model of the skier, including the center of mass position, velocity, and acceleration, using multi-frame human skeleton key point data stored in history. The state of the simplified kinematic model of the skier is estimated by unscented Kalman filtering to predict the skier's three-dimensional position, velocity, and acceleration at the next moment. Based on the IMU pre-integration results in the electronic pod and the motor speed data fed back by the UAV flight control, the pose and self-rotation angular velocity of the UAV at the next moment are predicted. The predicted position of the skier and the predicted position of the UAV are transformed to the same world coordinate system, and the relative position vector between the two is calculated. The desired pointing angle and feedforward compensation angular velocity of the gimbal are calculated based on the relative position vector. The feedforward compensation angular velocity is superimposed on the output of the PID feedback control loop of the gimbal motor. The environmental data integration module is connected to the environmental sensor and the multi-source sensor fusion module, and is used to transmit the ski resort environmental parameters to the multi-source sensor fusion module and integrate them into the 4D spatiotemporal data package.

[0017] The present invention has at least the following beneficial effects: First, this invention achieves centimeter-level positioning accuracy in high-reflectivity, low-texture environments like snowfields by combining multi-sensor fusion with snowfield-specific image enhancement processing, along with closed-loop exposure control and visual-inertial tightly coupled pose estimation. It generates standardized 4D spatiotemporal data packets, balancing high precision with strong environmental adaptability.

[0018] Secondly, by introducing multimodal health assessment and hierarchical virtual observation reconstruction, this invention maintains continuous pose output even when visual or inertial sensors temporarily fail due to snow fog or impact saturation. Furthermore, by creating artificial textures through active near-infrared structured light projection, it eradicates the problem of tracking loss caused by weak textures in snow and significantly improves the robustness of the system.

[0019] Third, this invention trains a lightweight model through cloud-based knowledge distillation, achieving high-precision skier detection and skeleton extraction at the edge. Furthermore, by using feedforward gimbal control based on skeletal kinematics prediction, tracking is transformed from post-event feedback correction to pre-event prediction guidance, effectively reducing tracking latency in high-speed dynamic scenarios.

[0020] Fourth, this invention integrates the UAV platform, electronic pod, ground control station and cloud server into a closed-loop data processing system through a modular system architecture, which supports flexible hardware expansion and online algorithm updates, facilitating actual deployment and continuous optimization at ski resorts.

[0021] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0022] Figure 1 This is a system architecture diagram of one technical solution of the present invention; Figure 2 This is a configuration diagram of an electronic pod according to one technical solution of the present invention; Figure 3 This is a flowchart illustrating a method processing technique of the present invention. Figure 4 This is a 4D data packet format diagram of one technical solution of the present invention. Detailed Implementation

[0023] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description. It should be understood that terms such as "having," "comprising," and "including" as used herein do not exclude the presence or addition of one or more other elements or combinations thereof. It should be noted that the experimental methods described in the following embodiments, unless otherwise specified, are conventional methods, and the reagents and materials, unless otherwise specified, are commercially available and therefore should not be construed as limiting the present invention.

[0024] like Figure 3 As shown, this invention provides a method for high-precision spatiotemporal data acquisition of ski resorts using unmanned aerial vehicles (UAVs) based on multimodal perception, comprising the following steps: Step A: The electronic pod, which houses all the sensors, is connected to the UAV fuselage. Through the binocular vision camera, IMU inertial measurement module, and RTK-GPS module mounted on the pod, it synchronously acquires raw image sequences, acceleration and angular velocity data, and geographic location data according to a unified hardware clock. Specifically, this includes, but is not limited to, the following methods: A high-precision temperature-compensated crystal oscillator is used as the master clock source. Pulse signals are simultaneously sent to the frame synchronization input pin of the binocular camera and the external trigger pin of the IMU via a trigger line. The rising edge of each pulse latches a sampling moment and generates a corresponding timestamp tag. The RTK-GPS module injects its absolute time reference into the synchronization network through second pulse signals and serial time messages, aligning the timestamps of all sensors to the same time coordinate system with microsecond-level accuracy. The electronic pod has a built-in edge computing unit that continuously performs brightness statistical analysis on the raw images during acquisition. Based on the real-time calculated snow surface brightness statistics, the exposure parameters of the binocular vision camera are adjusted in a closed loop. Specifically, this includes, but is not limited to, the following methods: The average brightness of an image is calculated in parallel using FPGAs or GPUs in edge computing units. L avg and brightness variance s 2 When detected s 2 Smaller and L avg At extremely high brightness levels, the scene is determined to be a large area of ​​snow. In this case, the camera sensor's analog gain and exposure time are automatically adjusted to adjust the target brightness value. L target The formula for adjusting the exposure time to shift towards the center of the histogram is: in, T exp_old This is the current exposure time. T exp_new The adjusted exposure time. L target The target brightness value is set. α It is a special overexposure protection factor for ski resorts, with a value range of 0.8 to 0.9, which makes the exposure adjustment tend to be slightly lower than the normal exposure level, so that richer highlight texture details can be obtained in snow scenes; The adjustment process forms a closed-loop feedback. The brightness statistics of the previous frame determine the exposure parameters of the next frame, and continuously and dynamically converge to the exposure configuration suitable for the current snowfield lighting conditions, so as to avoid the pixel values ​​in the highlight area reaching the saturation limit and losing texture information. Step B: Large areas of snow exhibit high brightness and low contrast in the visible light band. Performing conventional enhancement processing on the entire image would amplify sensor noise while increasing the contrast of the snow-covered area, and the increased computational overhead would be detrimental to subsequent real-time processing. Therefore, snow-specific image enhancement processing is performed on the acquired original image sequence, including at least: extracting non-snow-covered areas as regions of interest and performing contrast-constrained adaptive histogram equalization on these regions to obtain enhanced left and right eye images. Specifically, this includes, but is not limited to, the following methods: The non-snowy areas in the image are extracted as regions of interest using a fast saliency detection algorithm. Non-snowy areas typically include objects with actual texture features, such as skiers, trees, and ski slope edges. The gray-level histogram of the region of interest is then calculated. Within the region of interest, it is divided into several rectangular sub-blocks. The size of each sub-block is set according to the image resolution and the desired local enhancement granularity, usually set to 8×8 or 16×16 pixels. The grayscale histogram of each sub-block is calculated separately. Amplification processing is performed on the histogram, and a truncation threshold is set. The threshold is related to the total number of pixels in the sub-block and the number of gray levels in the histogram. The frequency part exceeding the threshold is truncated and evenly distributed to each gray level of the histogram to prevent excessive noise amplification due to excessive local contrast. The cumulative distribution function is calculated for the histogram after amplification to obtain the gray-level mapping curve. In order to eliminate the boundary effect between sub-blocks, the mapping values ​​of adjacent sub-blocks are fused through bilinear interpolation. Finally, the enhanced left and right eye images are output. Step C: Extract visual feature points from the enhanced left and right eye images and perform binocular matching. Combine acceleration and angular velocity data to perform visual-inertial tightly coupled pose estimation and calculate the visual odometry pose of the UAV in the world coordinate system. Specifically, this includes, but is not limited to, the following methods: Run feature extraction algorithms such as ORB or SIFT on the enhanced left and right images respectively. Locate the location with significant local structure by detecting the gradient distribution characteristics in the neighborhood of each pixel in the image. Specifically, including but not limited to the following methods: calculate the corner response function of each pixel, or construct the image scale space and detect local extreme points as feature points. For each detected feature point, generate a description vector based on the gray-level distribution of its surrounding pixels. The dimension of the description vector can be selected according to the trade-off between matching speed and discriminability. Binocular matching is performed between the feature point sets of the left and right eye images. During matching, for the feature point of the left eye, the feature point with the smallest description vector distance is searched on the corresponding epipolar line of the right eye image, and a bidirectional consistency check is applied, that is, the reverse matching result from the right eye point to the left eye point should also be consistent, so as to eliminate ambiguous matching. For each pair of successfully matched feature points, the pixel coordinates of the left and right eyes and the baseline length and focal length of the binocular camera are known. The depth coordinates of the spatial point in the camera coordinate system are calculated by the triangulation formula. Visual-inertial tightly coupled pose estimation (VIS) jointly processes visual observations and IMU measurements within a unified optimization framework. Optimization variables include the UAV's position, attitude, and velocity at multiple time points, as well as the IMU's bias parameters. Attitude can be represented using quaternions or rotation matrices. The cost function includes two types of residual terms: the first is the visual reprojection residual, which calculates the pixel distance between the actual projected position of each spatial point observed at multiple time points and the predicted projected position based on the currently estimated camera pose and spatial point position in each frame; the second is the IMU pre-integration residual, which pre-integrates the acceleration and gyroscope measurements between two adjacent time points to obtain a... The residual is the difference between the relative pose increment and velocity increment, which is the difference between the pre-integrated increment and the nominal increment calculated based on the current estimated state at two time points. The residual vector includes the deviations of the three components: position, velocity, and attitude. The optimization solution adopts conventional iterative methods such as the Gauss-Newton method or the Levenberg-Marquardt method. In each iteration, the nonlinear cost function is linearized and approximated near the current estimation point. After solving the incremental equation, the optimization variables are updated until the residual converges. Through this tightly coupled process, visual information corrects the drift of the IMU integral accumulated over time, while IMU information provides high-frequency motion priors for vision. They complement each other to obtain high-precision UAV pose estimation. Step D: Detect the skier target in the enhanced left and right eye images, extract the two-dimensional coordinates of the human skeletal key points (including foot key points) in the image coordinate system, calculate the three-dimensional coordinates of the skeletal key points in the camera coordinate system using binocular parallax, and match and assign identity markers to the skeletal key points between adjacent frames based on spatiotemporal consistency. Specifically, this includes, but is not limited to, the following methods: Skier target detection can be achieved using conventional convolutional neural networks such as YOLO and Faster R-CNN. The convolutional neural network slides or densely samples on the enhanced image and outputs the bounding box coordinates and confidence score of each potential target region. The detected regions inside the bounding boxes are fed into conventional skeletal keypoint extraction networks such as OpenPose and HRNet. The skeletal keypoint extraction network classifies the image pixels within the bounding box by part or directly regresses the coordinates of each keypoint. The output human skeletal keypoints include the positions of the main joints of the body and the foot keypoints. The foot keypoints at least cover the ankle and toe positions, which is convenient for subsequent skiing posture analysis. After obtaining the two-dimensional coordinates of the skeletal key points of each skier in the left and right eye images, the disparity is calculated for each pair of left and right eye skeletal key points with the same name using the left and right eye feature matching parameters and camera calibration data completed in step C. The disparity is the difference in the horizontal pixel coordinates of the key point in the left and right eye images. The disparity value is substituted into the triangulation formula to obtain the three-dimensional coordinates of the key point in the camera coordinate system. The combination of the three-dimensional coordinates of each key point constitutes the spatial representation of the skier's skeletal pose in the current frame. The matching of skeletal key points and the assignment of identity labels between frames solve the problem of continuous tracking in multiple frames. For two adjacent frames, the matching cost between each tracked skier in the previous frame and each detected skier in the current frame is calculated. The cost can be based on the sum of the spatial Euclidean distances between the corresponding skeletal key points. A bipartite graph matching problem is constructed, and the Hungarian algorithm is used to solve for the minimum cost. In this process, the prior knowledge of human skeletal structure is used for constraint verification. For example, the length of each bone segment should remain approximately constant between two adjacent frames. If a matching pair has a drastic change in bone length, the cost of the matching is increased by a penalty, so that it is excluded from the optimization process. After matching, the identity labels of consecutively appearing skiers remain unchanged, and newly appearing targets are assigned new labels. Step E: Input the visual odometry pose, acceleration and angular velocity data, geographic location data, and the 3D coordinates of skeletal keypoints in the camera coordinate system into the multi-source sensor fusion filter, perform timestamp alignment and coordinate system unification, and output the 3D spatial coordinates of the UAV and each skier in the world coordinate system, as well as the skeletal motion posture data, to generate a 4D spatiotemporal data package including microsecond-level timestamps, 3D spatial coordinates, and skeletal motion posture data. Specifically, this includes, but is not limited to, the following methods: Using the processing time of the fusion filter as a reference, interpolation is performed on the two frames of data in each data stream whose timestamps are closest to the reference time. The interpolation can be linear interpolation or prediction interpolation based on the motion model to obtain the alignment value of each data source at the reference time. The coordinate system is implemented through predefined and calibrated spatial transformation relationships between various coordinate systems. Between the left eye coordinate system of the binocular vision camera, the IMU coordinate system, the RTK-GPS antenna reference point coordinate system, and the UAV body coordinate system, there exist rotation matrices and translation vectors determined by mechanical installation drawings and offline calibration. The world coordinate system can be selected as a local Cartesian coordinate system with a fixed point at the ski resort as the origin, the east-west direction as the X-axis, the north-south direction as the Y-axis, and the sky direction as the Z-axis. The origin of this coordinate system can be calibrated by RTK-GPS positioning of a fixed marker point at the ski resort. All input data are transformed to the world coordinate system through their respective transformation chains. The fusion filter is implemented using a recursive estimation framework, such as an extended Kalman filter. The specific algorithm is wave or error state Kalman filter. The filter maintains a state vector, including the position, velocity and attitude of the UAV in the world coordinate system, as well as the position and velocity of each tracked skier. The prediction step uses the acceleration and angular velocity data of the IMU to recursively predict the state of the UAV. The update step uses the observation data to correct the state. The pose output by the visual odometry is used as the observation of the UAV's state, the global position output by RTK-GPS is used as the absolute observation of the UAV's position, and the three-dimensional coordinates of the skeletal key points are used as the observation of the skier's state. The filter outputs the corrected state estimate in each cycle, that is, the three-dimensional spatial coordinates of the UAV and each skier in the world coordinate system and the skeletal motion posture. The output data is encapsulated into 4D spatiotemporal data packets. The first field of each record is a timestamp with microsecond precision, which is based on a unified clock source. The three-dimensional spatial coordinate field gives the position of the target in the global coordinate system at that moment. The skeletal motion posture data field records the three-dimensional coordinates of each key point in array form. The data packet can be attached with a frame check sequence to ensure transmission integrity. It is transmitted to the ground station and cloud server in real time through a wireless data link, completing the end-to-end processing from raw sensing data to high-value structured spatiotemporal information.

[0025] In the above technical solution, multiple sensors are integrated into the UAV electronic pod, and multi-source data synchronous acquisition is achieved through a unified hardware clock. The sensor saturation problem caused by the high reflectivity of snow is overcome through closed-loop exposure control. The stable positioning problem in low-texture environments is solved through ski resort-specific enhancement and visual-inertial tight coupling pose estimation. Structured tracking of high-speed skiers is achieved through skeleton extraction with foot key points and spatiotemporal consistency matching. 4D spatiotemporal data that can be directly used for motion analysis is generated through multi-source fusion and standardized data encapsulation. It is adapted to the complex environmental characteristics of ski resorts with high reflectivity, low texture, and high-speed skier movement, and improves the temporal synchronization, imaging adaptability, and pose calculation stability of multi-source data acquisition. It can stably meet the basic usage requirements of ski resort analysis, on-site monitoring, and spatial information acquisition.

[0026] In actual ski resort flights, drones need to frequently perform high-angular-velocity yaw or pitch maneuvers to follow skiers at high speeds. When the camera rotates relative to the scene during the exposure time, linear motion blur (i.e., motion blur) will appear in the image along the direction of motion. This blur will prevent the subsequent visual feature extraction module from obtaining clear feature points, thereby reducing the accuracy of pose estimation and even causing tracking loss. In another technical solution, the image enhancement processing in step B also includes: The camera motion blur radius and direction are estimated using the acceleration and angular velocity data, the point spread function is reconstructed, and Wiener filtering is used to perform deconvolution processing. Specifically; The edge computing unit obtains a three-axis angular velocity data sequence corresponding to the start and end times of exposure of the image frame from the IMU inertial measurement module, and processes the sequence during the exposure time. t exp Integrating the vectors, we obtain the cumulative rotation vector of the camera during the exposure period, and then calculate the average camera angular velocity during the exposure period. oh The rotation vector is transformed into the camera coordinate system using the calibrated extrinsic rotation matrix between the camera and the IMU, thus obtaining the camera's rotation around itself. X shaft and Y The rotational component of the axis, and the direction in which these two components combine on the image plane, represent the trajectory direction of the motion blur in that frame of the image. e ; Fuzzy radius R blur From camera angular velocity oh mold length, exposure time t exp Camera effective focal length f and pixel size p The fuzzy radius is determined jointly. R blur for: R blur = (f / p)×|ω|×t exp ; When the calculation result is not an integer, it is rounded to the nearest integer value, in pixels. In obtaining fuzzy direction e and fuzzy radius R blur Then, the point spread function of the image degradation process in that frame is reconstructed. H ( u , v During reconstruction, assuming the relative rotational motion during exposure can be approximated as uniform, the point spread function in the spatial domain appears as a line along the direction... e Length is 2 R blurA line segment with +1 pixel weights, where all points on the segment have equal weights and their sum is normalized to 1, is represented in the frequency domain by its point spread function after Fourier transform. H ( u , v ); After reconstructing the point spread function, the system uses Wiener filtering to deconvolve the blurred image. The frequency domain expression of Wiener filtering is: in, H *( u , v )for H ( u , v The complex conjugate of ) , | H ( u , v )| 2 The power spectrum of the point spread function. K This is a constant with a small positive value, used to balance the sharpening effect of inverse filtering with noise amplification. During processing, the local region of the blurred image and the reconstructed point spread function are converted to the frequency domain by fast Fourier transform, and the frequency domain data of the image is multiplicatively filtered point by point according to the above formula. Then, the filtering result is subjected to inverse Fourier transform to obtain the real part, and the deconvolutioned spatial domain image is obtained. After deconvolution processing, the high-frequency edge features in the image that were originally degraded due to high-speed maneuvering are restored, and the outline of the skier and the edge texture of the ski slope become clear again. The restored image then enters the contrast-limited adaptive histogram equalization stage, and continues to perform snow-specific contrast enhancement processing, finally generating enhanced left and right eye images with significantly improved quality.

[0027] In the above technical solution, motion blur parameters are estimated online using IMU angular velocity measurements and the point spread function is reconstructed. The blurred high-frequency edge information is efficiently recovered through frequency domain Wiener filtering, which enhances the adaptability of the cascaded enhancement process in dynamic scenes. This ensures that subsequent visual feature extraction, matching, and pose estimation remain stable and reliable during high-maneuverability flight. Since the entire deblurring process directly obtains motion parameters from the IMU data, there is no need to iteratively search and estimate the blur kernel on the image itself. The computation time is extremely low, making it suitable for real-time execution on edge computing units.

[0028] In a ski resort setting, binocular cameras may experience image blurring due to sudden snow fog, large areas of pure white with weak texture, or violent movement, resulting in the instantaneous loss of effective feature points. Meanwhile, the IMU (Inertial Measurement Unit) may experience instantaneous saturation due to acceleration exceeding its range when a skier is cornering at high speed or landing from the air, or bias drift due to rapid temperature changes. These failure modes are uncertain in both time and severity. In another technical solution, step C also includes a multimodal failure compensation step based on health assessment. Visual modal health and IMU modal health are calculated in real time. Visual modal health is calculated based on the number of successfully matched feature points and their distribution variance in the image plane. IMU modal health is calculated based on the IMU pre-integration residual and temperature drift rate. The calculation of visual modality health relies on two metrics. The first metric is the total number of feature points successfully extracted in the current frame and matched with those in the previous frame or local map. N feat The second metric is the uniformity of the spatial distribution of these matched feature points on the image plane, which can be measured by the eigenvalues ​​of the covariance matrix of the two-dimensional coordinates of the feature points or the variance of the distribution of the number of points in each quadrant, denoted as . s 2 dist Visual modal health H vis The overall calculation method is as follows: in, N th To reference a threshold for the number of feature points, the sigmoid function maps the ratio of feature point counts to the interval between 0 and 1, exp(- s 2 dist ) is the spatial distribution uniformity factor. When feature points are more concentrated in a small region of the image, exp(- s 2 dist When the feature points are uniformly distributed, exp(-) approaches 0. s 2 dist It approaches 1. H vis It can comprehensively reflect the richness of visual features and the strength of spatial geometric constraints, when H vis When the value is less than 0.2 (example value, which can be adjusted according to the actual scenario), the current visual modality is determined to be invalid, corresponding to a scenario where large-area feature loss is caused by snow fog occlusion or motion blur. The metrics for assessing IMU modal health are pre-integration residuals and temperature drift rate. H imuThe calculation method is as follows: between two adjacent keyframes, the acceleration and gyroscope measurements are integrated in the IMU body coordinate system to obtain a relative pose increment and velocity increment. At the same time, based on the state of the two keyframes estimated by the current filter, a nominal relative increment can also be calculated. The difference between the two is the pre-integrated residual vector. d z When the IMU is working properly, d z The magnitude should fluctuate randomly around zero. When the accelerometer exhibits clipping distortion due to saturation or the gyroscope produces abnormal drift, [the following occurs]. d z The non-linear growth indicates that the IMU measurement is inaccurate, and the temperature drift rate... b drift The bias estimate is obtained by reading the temperature sensor built into the IMU and calculating the rate of change of the bias estimate with temperature per unit time. This is done when a sustained high angular velocity is detected, causing the integrator to saturate or... d z When the growth is nonlinear, the IMU mode is determined to be in failure, that is, the IMU mode health is lower than the second threshold. When the visual modal health is below the first threshold and the IMU modal health is not below the second threshold, the visual feature matching result is unusable due to insufficient feature points or extremely poor distribution. Forcibly incorporating its reprojection error term into the optimization will introduce erroneous information into the filter, utilizing acceleration. a and angular velocity oh Data, the effective pose output from the filter at the previous time step. x t-1 Starting from the beginning, the predicted pose at the current moment is generated through kinematic integrals: Where ⊕ represents the pose addition operation, the predicted pose is combined with the spatial feature points in the world coordinate system, the predicted projection coordinates on the image plane are calculated by the camera projection model, virtual visual observations are generated, and the visual-inertial tightly coupled pose estimation process is injected to forcibly constrain the visual reprojection error term. When the IMU modal health is below the second threshold and the visual modal health is not below the first threshold, the IMU measurement is unreliable due to saturation or drift. However, this failure is usually instantaneous, and the visual information remains reliable. Virtual IMU measurements are generated using the triangulation results of visual feature points in consecutive frames, and the linear acceleration of the UAV is then deduced. and angular acceleration Virtual IMU measurements are generated by spline interpolation fitting to compensate for the prediction step in Kalman filtering. At the moment a skier lands, the violent impact causes the accelerometer to saturate instantaneously, and the rising snow fog obscures the lens. At this point, data from both types of direct sensors are unreliable. This occurs when the visual modality health falls below a first threshold and the IMU modality health falls below a second threshold, for example... The system initiates protective estimation based on human skeletal kinematics, extracts the centroid motion data of key points in the human skeleton, and calculates the pixel displacement Δ of the skier's torso center in the image plane. u Virtual observations of the center of mass are extracted, and kinematic constraints are constructed by combining prior knowledge of human biomechanics. An upper limit for the acceleration along the center of mass line is set. a max This value typically does not exceed three times the acceleration due to gravity, and is also the upper limit of the rate of change of joint angles. Construct a cost function for nonlinear optimization compensation: in, The pose is obtained by short-time integration of the IMU. The pose is calculated based on the displacement of the skeletal center of mass. For the currently estimated linear acceleration, l c This is the penalty coefficient; Solve the constrained least squares problem, correct the pose result obtained by short-time integration of the IMU, inject the corrected pose into the filter as a virtual observation, maintain stable pose estimation for a short time (within about 0.5 seconds) until any sensor mode recovers, and once any health score rises above the threshold, the system automatically exits the final compensation and switches back to the normal tightly coupled working state.

[0029] In the above technical solution, by introducing health level classification and physical model reconstruction of virtual observation, the problem of tracking interruption caused by the complete failure of sensors in extreme ski resort environments is solved, and the robustness of the system is significantly improved.

[0030] To address the fundamental problem that the lack of texture in large areas of snow in ski resort environments prevents visual odometry from extracting feature points, another technical solution includes step C, which further incorporates an active structured light-assisted feature extraction step. The electronic pod also includes a near-infrared VCSEL structured light projector and a narrowband filter of the corresponding wavelength. VCSEL is a vertical cavity surface-emitting laser. Snow reflects the solar spectrum extremely strongly in the visible light band. If visible light structured light is used, the projected pattern will be submerged by the sunlight reflected from the snow surface, and it will not be able to form an effective contrast in the image. The wavelength of the VCSEL structured light projector is selected in the near-infrared band of 850nm or 940nm, which can effectively avoid the interference band of the strong solar spectrum of the snow surface. At the same time, the snow surface has strong diffuse reflection characteristics of near-infrared light, which can make the projected structured light pattern form a clear image on the sensor. The narrowband filter is strictly matched with the wavelength of the projector, and only receives the structured light pattern emitted by the projector, filtering out the snow reflection and ambient stray light in the visible light band. When the visual modality health is lower than the preset threshold for weak texture, it indicates that the current scene is a large area of ​​snow with weak texture, and natural feature extraction can no longer support pose estimation. The near-infrared VCSEL structured light projector is then activated. The density of the projection grating is adaptively adjusted according to the real-time flight altitude of the UAV, and random speckle or coded grid patterns are projected onto the snow surface. When the flight altitude is high, the coverage area of ​​a single structured spot on the ground increases, and the system projects a sparser coded grid pattern to provide a sufficient number of feature points in a larger field of view. When the flight altitude is low, the system projects a finer random speckle pattern to obtain higher local depth detail resolution. After the structured light pattern is projected onto the snow surface, since snow is a Lambertian scatterer, it will be uniformly scattered in all directions after being irradiated with near-infrared light, forming a high-contrast, clear speckle or grid pattern on the camera sensor. The binocular camera acquires the structured light image through a narrow-band filter. At this time, the visual odometry method no longer relies on natural texture, but extracts the speckle center or grid corner point of the structured light projection as a highly robust visual feature point. The speckle center can be extracted by local gray-level extreme value detection or Gaussian fitting, and the grid corner point can be extracted by a checkerboard corner point detection algorithm for binocular matching and pose estimation. To prevent near-infrared light from causing discomfort to the human eye or violating laser safety regulations, the emission power of this near-infrared VCSEL structured light projector is adaptively adjusted based on ambient light intensity and image contrast. P laser The adjustment formula is: in, P base As the reference transmit power, I sun The current sunlight intensity is measured by an ambient light sensor. I max Δ is the upper limit of the reference for sunlight intensity. P contrastThis is the amount of power fine-tuning based on the structured light contrast feedback in the image; When the ambient light is strong, the reflection from the snow surface already provides some illumination, so the projection power can be reduced accordingly. When the ambient light is weak or the image contrast is insufficient, the emission power should be increased appropriately to ensure the quality of feature point extraction and to ensure that while providing sufficient texture contrast, the eye safety level (Class 1) is maintained.

[0031] In the above technical solution, by introducing near-infrared active structured light projection and using a narrowband filter for wavelength-selective imaging, the problem of visual SLAM failure in snowy, low-texture environments is solved. The actively projected structured light creates artificial textures on the snow surface, providing sub-pixel-level feature matching accuracy. Combined with binocular parallax, the measurement accuracy of local terrain undulations on the snow surface can be improved to the millimeter level. The selection of the near-infrared band and the cooperation with the narrowband filter make the system unaffected by visible light snow reflections, achieving stable operation in all weather conditions. This ensures that the system provides effective feature assistance while always meeting human eye safety standards.

[0032] In another technical solution, the key points of the human skeleton in step D include 20 key points, and the skeletal lines connecting these key points constitute the complete human movement: Head (1 point): The tip of the nose has relatively stable visual recognizability and serves as a reference point for the spatial position of the head. Trunk (9 points): Neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips. The neck connects the head and trunk and is the key node for describing the head's posture relative to the trunk. The left and right shoulders and left and right hips form the four corner points of the trunk, expressing the trunk's twisting, forward tilting and side tilting postures. The left and right elbows and left and right wrists describe the movement of the upper limbs and can capture the skier's upper limb movements such as arm swinging and pole tapping. Lower limbs (10 points): left and right knees, left and right ankles, left and right heels, left and right toes, and the outer sides of left and right heels. Conventional human skeletal models usually only define up to the ankle joint. This model adds left and right heels, left and right toes, and the outer sides of left and right heels. The left and right toes are located at the base of the big toe and are used to mark the position of the front of the foot. The left and right heels are located at the most prominent position on the back of the calcaneus and are used to mark the position of the back of the foot. The outer sides of left and right heels are located in the outer area of ​​the calcaneus and are used to help determine the inversion and supination posture of the foot, accurately capturing the skier's foot force posture and center of gravity distribution on the skis.

[0033] In the above technical solution, by expanding the key points of the human skeleton to include 20 detailed key points of the feet, the capture of the skier's posture extends from the macroscopic movement of the torso and limbs to the details of the feet, providing refined kinematic data for analyzing the skier's center of gravity transfer and ski control posture.

[0034] Under the limited computing power of the edge computing unit of the drone, it is necessary to obtain a target detection and skeletal key point extraction model with both high accuracy and low latency. In another technical solution, the model used in step D to detect the skier target and extract human skeletal key points is a lightweight deep learning model, which is trained by cloud knowledge distillation. The training dataset includes images of various skiing postures and snow obstacles. The skiing postures cover typical actions such as straight downhill skiing, snowplow turns, parallel turns, and jumps. The snow obstacles include gateposts, snow mounds, trees, etc. Each image is simultaneously labeled with the skier's bounding box and the coordinates of multiple skeletal key points. The Transformer model is used as a high-parameter teacher network (such as ViT-Large). The Transformer model captures long-distance pixel dependencies in the image through a self-attention mechanism and has a strong ability to perceive blurred contours and partial occlusion. It is pre-trained on a dataset that includes skiing postures and snow obstacles, so that its output can accurately locate the skier's bounding box and each skeleton key point. The high-dimensional feature response map is extracted. The high-dimensional feature response map is a multi-channel tensor generated by the intermediate or output layer of the network. Each channel selectively activates a specific semantic pattern and includes rich encoding information such as target category, contour, and texture. A lightweight deep learning model is used as the student network. The backbone of the student network adopts an improved depthwise separable convolutional network, which includes four downsampling stages and a total of about 18 layers. The depthwise separable convolution decomposes the standard convolution into two steps: channel-wise convolution and pointwise convolution, which significantly reduces the number of parameters and computational cost. A lightweight feature pyramid structure is introduced into the neck network to fuse multi-scale features with less computational overhead, thereby enhancing the adaptability to skier scale at different distances. The network input resolution is set to 320×320 pixels. This size achieves a balance between detection accuracy and inference speed. At this input size, the single-frame inference computation of the student network is on the order of several GFLOPS, and a processing speed of no less than 60 frames per second can be achieved on the edge computing units. During distillation training, the same image is input to both the teacher and student networks. The teacher network outputs a high-dimensional feature response map during forward propagation, denoted as... f teacher The feature output of the corresponding layer of the student network is denoted as... f student The corresponding feature layer output is extracted, and training is completed by minimizing the KL divergence between the corresponding feature layer output of the student network and the high-dimensional feature response map of the teacher network. The following optimization objective function is minimized: in, L CE For classification cross-entropy loss, y For accurate labeling, This loss term, representing the prediction output of the student network, guarantees the student network's basic prediction accuracy for the target category and keypoint coordinates. L KD The distillation loss is specifically calculated by analyzing the high-dimensional feature response map of the teacher network. f teacher and the characteristic output of the student network f student After temperature coefficient T After softening with the softmax function, the KL divergence between the two probability distributions is calculated. T This is the temperature coefficient, usually set to 4.0. T 2 This is the square factor of the temperature coefficient, used to compensate for the effect of temperature softening on the gradient magnitude during gradient backpropagation. l This is a balancing factor used to adjust the relative weights between classification loss and distillation loss. minimize L KD The process essentially involves gradually approximating the feature response distribution of the student network to that of the teacher network. Through multiple rounds of iterative training, the student network gradually learns the teacher network's ability to generalize and recognize difficult samples such as blurred skier outlines and partial occlusions, without adding any extra computational overhead during inference. After training, the cloud server distributes the model parameters of the student network to the edge computing unit of the drone for use by the target detection and skeleton extraction modules. Under the condition that the single frame latency is controlled within 15 milliseconds, detection and skeleton extraction effects close to the accuracy level of large models are obtained.

[0035] In the above technical solution, the feature representations learned by the large-parameter Transformer teacher network on massive skiing scene data are transferred to the lightweight student network suitable for real-time inference at the edge through cloud knowledge distillation training. This overcomes the contradiction between edge computing power limitations and model accuracy requirements, and enables the accuracy of skier target detection and skeletal key point extraction to meet practical requirements in terms of both inference speed and accuracy.

[0036] In conventional visual servo control, the gimbal generates feedback control based on the deviation of the skier's position from the center of the image frame. This results in significant lag when the skier is turning or jumping at high speed, easily causing the skier to deviate from the center of the image or even run completely out of the field of view. Another technical solution also includes an active feedforward gimbal control step based on human skeletal kinematics prediction. By utilizing multi-frame human skeletal keypoint data stored historically, motion information of the skier's torso and lower limbs is extracted to establish a simplified kinematic model of the skier, including center of mass position, velocity, and acceleration. Specifically, the system retains the most recent... T Frame skeletal keypoint data,T Typically, 10 frames are used, in this... T In the frame data, we focus on the position of the midpoint of the left and right hips as the center of mass of the torso, as well as the position changes of key points of the feet such as the left and right ankles, toes and heels. Based on these historical data sequences, we establish a simplified kinematic model with the position, velocity and acceleration of the center of mass as state variables, and regard the skier's motion as the motion of a point mass in three-dimensional space. At the same time, we implicitly cover the uncertainty of nonlinear maneuvers such as turning and jumping by setting state noise. Because skiing involves nonlinear characteristics such as turning and jumping, a simplified kinematic model of the skier is estimated using the Unscented Kalman Filter (UKF). The UKF selects a set of Sigma sampling points near the state mean through an unscented transformation. These sampling points are then propagated through a nonlinear state transition function, and weighted statistically analyzed to obtain the mean and covariance of the predicted state. A prediction-update loop is executed after each frame of skeletal data arrives, and the prediction is recursively pushed forward by a time step Δ. t Predicting skiers t+ Δ t Three-dimensional position at time ,speed and acceleration For nonlinear motions such as turning and jumping, UKF has higher prediction accuracy compared to linear Kalman filtering; The drone's own motion prediction is performed simultaneously. Based on the IMU pre-integration results in the electronic pod and the motor speed data fed back by the drone's flight control, specifically, the pre-integration compresses multiple frames of IMU sampled values ​​into relative pose and velocity increments between two adjacent moments. The edge computing unit obtains motor speed feedback data from the drone's flight control system and, combined with the IMU pre-integration results and motor speed data, predicts the drone's motion. t+ Δ t Position at any moment and its own rotational angular velocity ; Predicted skier location With drone-predicted location Transform to the same world coordinate system, calculate the relative position vectors of the two, and calculate the desired gimbal pointing angle. Align the camera's optical axis with the predicted position of the skier: Calculate the compensated relative angular velocity : in, i cur Δ represents the current actual angle of the gimbal. t To predict the time step, The predicted angular velocity of the drone's own rotation; Will As a feedforward control quantity, it is directly superimposed on the output of the PID feedback control loop of the gimbal motor, so that the gimbal can rotate in the correct direction in advance when the skier is about to change position.

[0037] In the above technical solution, by establishing a nine-dimensional state vector of the skier's center of mass position, velocity, and acceleration, the unscented Kalman filter is used to perform nonlinear state estimation and prediction on the uniformly accelerated motion model. At the same time, the IMU pre-integration and motor speed feedback are combined to predict the UAV's own motion. The feedforward compensation angular velocity is calculated in a unified world coordinate system and superimposed on the gimbal servo control loop. High-speed ski tracking changes from feedback lag correction to feedforward prediction guidance, which significantly reduces dynamic tracking delay and ensures that the skier always stays stably near the center of the image.

[0038] like Figure 1 As shown, a high-precision spatiotemporal data acquisition system for ski resorts based on multimodal perception is used to execute the method described above. The system includes: The drone platform's internal flight control system interacts with the electronic pod via a serial data interface, including sending flight status data to the electronic pod and receiving gimbal control commands. The drone platform's power output provides a stable DC power supply to the electronic pod. Electronic pods, mounted on the aforementioned drone platform, such as Figure 2 As shown, the electronic pod includes a sensor unit and an edge computing unit. The sensor unit includes a binocular vision camera, an IMU inertial measurement module, and an RTK-GPS module. The edge computing unit is electrically connected to the sensor unit. The electronic pod is connected to the UAV fuselage via a shock-absorbing bracket and an electrical interface. The binocular vision camera consists of two industrial cameras calibrated with intrinsic and extrinsic parameters, installed parallel to each other at a fixed baseline distance. The IMU inertial measurement module integrates a three-axis accelerometer and a three-axis gyroscope. The RTK-GPS module provides centimeter-level global coordinates by receiving dual-frequency signals from a base station and satellites. The edge computing unit is electrically connected to each sensor of the sensor unit via a high-speed data bus. The data bus can use interface protocols such as MIPI, USB, or PCIe, and is responsible for transmitting raw image data and sensor measurement values. The sensor unit is used to synchronously acquire raw image sequences, acceleration and angular velocity data, and geographic location data according to a unified hardware clock, and transmit the raw image sequences, acceleration and angular velocity data, and geographic location data to the edge computing unit. Specifically, a temperature-compensated crystal oscillator is set in the electronic pod as the master clock source. Pulse signals are sent to the frame synchronization pin of the binocular camera and the external trigger pin of the IMU simultaneously through the trigger line. The sampling time is latched on the rising edge of each pulse and a timestamp tag is generated. The RTK-GPS module injects the absolute time reference into the synchronization network through the second pulse signal and serial port time message. The timestamps of each sensor are aligned to the same time coordinate system with an alignment accuracy of up to microsecond level. After the synchronous acquisition is completed, the sensor unit pushes the data stream to the data receiving buffer of the edge computing unit. The edge computing unit includes an image enhancement module, an exposure control module, a visual feature extraction and matching module, a visual-inertial tightly coupled pose estimation module, a target detection and skeleton extraction module, a spatiotemporal consistency matching module, and a multi-source sensor fusion module. The image enhancement module is used to perform ski resort-specific image enhancement processing on the original image sequence. This includes extracting non-snow areas in the image as regions of interest and performing contrast-limited adaptive histogram equalization on these regions to output enhanced left and right eye images. Specifically, a fast saliency detection algorithm is first used to extract non-snow areas in the image. Effective target areas such as skiers, trees, and ski slope edges are segmented to generate masks. Within the regions of interest covered by the masks, the image is divided into 8×8 or 16×16 pixel rectangular sub-blocks. A gray-level histogram is calculated for each sub-block separately and amplitude limiting is performed. Frequency components exceeding the threshold are truncated and evenly distributed to each gray level. The cropped histogram is mapped using a cumulative distribution function. Adjacent sub-blocks are fused using bilinear interpolation to eliminate boundary effects. Finally, enhanced left and right eye images are output. The exposure control module is connected to both the image enhancement module and the binocular vision camera. It receives statistical values ​​of snow surface brightness calculated based on the original image sequence from the image enhancement module and generates an exposure adjustment command based on these values, which is then output to the binocular vision camera to adjust its exposure parameters in a closed-loop manner. Specifically, the exposure control module obtains the average brightness of the current frame image from the image enhancement module. L avg and brightness variance s 2 When detected s 2 Smaller and L avgWhen the brightness is extremely high, the current scene is determined to be a large area of ​​snow. The new exposure time and analog gain value are calculated according to the principle of shifting the target brightness value to the middle of the histogram. The instruction is written to the control register of the camera sensor through the I2C or SPI communication interface. The adjusted exposure parameters make the brightness of the next frame image return to a reasonable range. The statistical value of the previous frame determines the parameters of the next frame, forming a closed-loop dynamic convergence. The visual feature extraction and matching module is connected to the image enhancement module. It receives the enhanced left and right eye images and extracts visual feature points from the enhanced left and right eye images and performs binocular matching. Specifically, the visual feature extraction and matching module runs feature detection algorithms on the enhanced left and right eye images respectively. It locates corner points or spot features by calculating the gradient distribution of the pixel neighborhood and generates a description vector for each feature point. During binocular matching, it uses epipolar constraints to search for the most matching feature point on the corresponding horizontal scan line of the right eye image and applies a bidirectional consistency check to eliminate ambiguous matches. The successfully matched feature point pairs will be output to the subsequent pose estimation module. The visual-inertial tightly coupled pose estimation module is connected to the visual feature extraction and matching module and the IMU inertial measurement module. It receives the matching results of visual feature points as well as acceleration and angular velocity data, and is used to perform visual-inertial tightly coupled pose estimation to calculate the visual odometry pose of the UAV in the world coordinate system. Specifically, the visual-inertial tightly coupled pose estimation module constructs an optimization problem including the UAV's position, attitude, velocity, and IMU bias. The visual reprojection error and the IMU pre-integration residual are used as the joint cost function. The visual reprojection error measures the pixel deviation between the actual projection and the predicted projection of the spatial feature points on the image. The IMU pre-integration residual measures the difference between the IMU measurement integral and the nominal increment between adjacent time steps. By iteratively solving this nonlinear optimization problem, the six-degree-of-freedom pose of the UAV in the global coordinate system is obtained. The target detection and skeleton extraction module is connected to the image enhancement module. It receives the enhanced left and right eye images and uses them to detect skier targets in the enhanced left and right eye images. It extracts the two-dimensional coordinates of human skeleton key points, including foot key points, in the image coordinate system, and calculates the three-dimensional coordinates of the skeleton key points in the camera coordinate system using binocular parallax. Specifically, the target detection and skeleton extraction module runs a lightweight deep learning model. First, it detects the bounding boxes of each skier in the image. Then, it locates predefined human skeleton key points in each bounding box, including the tip of the nose, neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, left and right heels, left and right toes, and the outer sides of left and right heels. After obtaining the pixel coordinates of the same key points in the left and right eyes, it uses the calibrated binocular camera parameters to calculate the depth value of each key point in the camera coordinate system using the parallax triangulation formula, thereby obtaining its three-dimensional coordinates. The spatiotemporal consistency matching module is connected to the target detection and skeleton extraction module. It receives skeletal key points from adjacent frames and uses them to match and assign identity markers to skeletal key points between adjacent frames based on spatiotemporal consistency. Specifically, the spatiotemporal consistency matching module calculates the spatial distance cost matrix between the sets of skeletal key points of each skier in adjacent frames and uses the Hungarian algorithm to solve the global optimal matching with the minimum cost. During the matching process, the prior knowledge of constant human skeleton length is introduced as a constraint. If the length of a bone segment in a matching pair changes beyond a preset proportion, a penalty cost is added to exclude it. After matching, skiers that appear consecutively retain their original identity markers, newly appearing targets are assigned new markers, and temporarily missing targets have their markers entered into a cache queue to wait for recovery. The multi-source sensor fusion module is connected to the visual-inertial tightly coupled pose estimation module, the IMU inertial measurement module, the RTK-GPS module, and the target detection and skeleton extraction module. It receives visual odometry pose, acceleration and angular velocity data, geographic location data, and the 3D coordinates of skeletal key points. It performs timestamp alignment and coordinate system unification, outputting the 3D spatial coordinates of the UAV and each skier in the world coordinate system, as well as skeletal motion posture data. It also generates a 4D spatiotemporal data package including microsecond-level timestamps, 3D spatial coordinates, and skeletal motion posture data. Specifically, the multi-source sensor fusion module... The block interpolates and aligns each input data stream based on the fusion processing time. Using pre-calibrated rotation matrices and translation vectors between various coordinate systems such as camera-IMU, IMU-body, and GPS antenna-body, all data are uniformly transformed into the world coordinate system. The filter adopts extended Kalman filter or error state Kalman filter. The state vector includes the position, velocity, attitude, and IMU bias of the UAV, as well as the position and velocity of each tracked skier. Each processing cycle performs state prediction and observation update, outputs the corrected global pose and skeleton pose of the UAV and skiers, and encapsulates it into a standard format 4D spatiotemporal data packet. The ground control station is connected to the electronic pod and is used to receive and display the real-time data sent by the electronic pod. The display interface presents the drone's flight trajectory, the spatial position and skeletal posture of each skier in real time in a three-dimensional view and two-dimensional projection, so that the operator can monitor the system's operating status and the skiers' movement. The cloud server is communicatively connected to the ground control station and / or the electronic pod, and is used to receive 4D spatiotemporal data packets, perform big data analysis and model optimization, use the accumulated massive data to perform knowledge distillation training or incremental fine-tuning on the target detection and skeleton extraction models, send the optimized model parameters to the edge computing unit through the data link to realize online model updates, and send the updated model to the edge computing unit.

[0039] In the above technical solution, the system uses a four-level architecture of UAV platform, electronic pod, ground control station and cloud server to organically integrate multi-source sensor data acquisition, real-time edge processing, ground monitoring and cloud continuous optimization into a closed loop, which is suitable for the high-precision spatiotemporal data acquisition needs of complex ski resort environments.

[0040] In another technical solution, the edge computing unit of the UAV high-precision spatiotemporal data acquisition system for ski resorts based on multimodal perception further includes a motion blur compensation module and a multimodal failure compensation module. The motion blur compensation module is connected to the sensor module and the image enhancement module. It receives acceleration and angular velocity data, estimates the camera motion blur radius and direction, reconstructs the point spread function, performs deconvolution processing using Wiener filtering, and outputs the processed image sequence to the image enhancement module. Specifically, the motion blur compensation module obtains the three-axis angular velocity data corresponding to the current frame exposure time from the IMU inertial measurement module. Combined with the camera's exposure time, effective focal length, and pixel size, it calculates the motion blur radius and blur trajectory direction of the frame image. Under the assumption that the camera movement is approximately uniform during exposure, it reconstructs the point spread function describing the image degradation process. This function is represented as a line segment along the blur direction in the spatial domain. The module performs frequency domain transformation on the original image and the point spread function respectively. In the frequency domain, it performs inverse filtering using a Wiener filter. A small constant is used to balance the relationship between the inverse filtering sharpening degree and noise amplification. After inverse transformation, a clear image after deconvolution is obtained and output to the image enhancement module for further processing. The multimodal failure compensation module is connected to the visual-inertial tightly coupled pose estimation module, the visual feature extraction and matching module, the IMU inertial measurement module, and the target detection and skeleton extraction module. It is used to calculate visual modal health and IMU modal health in real time. Visual modal health is calculated based on the number of successfully matched feature points and their distribution variance in the image plane. IMU modal health is calculated based on the IMU pre-integration residual and the temperature drift rate. When the visual modal health is below the first threshold and the IMU modal health is not below the second threshold, the visual feature matching result is unreliable due to too few feature points or extremely poor distribution. The multimodal failure compensation module obtains the current effective acceleration and angular velocity measurements from the IMU inertial measurement module. Using the acceleration and angular velocity data, and taking the effective pose output by the filter at the previous moment as the starting point, the module calculates the predicted pose at the current moment through kinematic integration. The module then combines the predicted pose with the spatial feature points to calculate the image projection coordinates, generates virtual visual observations, and injects them into the visual-inertial tightly coupled pose estimation module to constrain the reprojection error term in the optimization. When the IMU modal health is below the second threshold and the visual modal health is not below the first threshold, the IMU measurement value is unreliable due to saturation or drift, but the visual information remains valid. The multimodal failure compensation module generates virtual IMU measurement values ​​by using the triangulation results of visual feature points in consecutive frames. That is, by using the triangulated three-dimensional coordinates of visual feature points in two consecutive frames, the equivalent linear acceleration and angular acceleration of the UAV are deduced from the spatial point displacement between frames. After spline interpolation fitting, a virtual measurement sequence aligned with the IMU sampling time is generated and injected into the visual-inertial tightly coupled pose estimation module. The pre-integration stage replaces the failed real IMU data to complete the state prediction. When the visual modality health is below the first threshold and the IMU modality health is below the second threshold, both the visual and IMU modalities fail simultaneously. The multimodal failure compensation module extracts the center of mass motion data of key points of the human skeleton, and combines the prior knowledge of human movement biomechanics to construct and set the upper limit of the center of mass acceleration and the upper limit of the joint angle change rate as kinematic hard constraints. It constructs a constrained pose correction problem, and finds the optimal pose correction amount under the premise of satisfying kinematic rationality. It corrects the pose result obtained by short-time integration of the IMU and maintains short-time stable pose estimation until any sensor modality recovers.

[0041] In the above technical solution, by adding a motion blur compensation module and a multimodal failure compensation module to the edge computing unit, motion blur degradation and short-term sensor failure in extreme environments are included in the fault tolerance protection scope. The motion blur compensation module restores image clarity from the data source to ensure the quality of subsequent feature extraction. The multimodal failure compensation module maintains the continuity of pose estimation in the event of single-modal or dual-modal failure through hierarchical progressive compensation, thereby enhancing the robustness and environmental adaptability of the system.

[0042] In another technical solution, the electronic pod also includes a near-infrared VCSEL structured light projector, a narrowband filter of the corresponding wavelength, a gimbal servo control module, and an environmental sensor; The near-infrared VCSEL structured light projector is connected to the edge computing unit and installed in the middle between the binocular cameras. It selects the 850nm or 940nm band to avoid the interference band of strong visible light reflection from the snow. It is used to receive the projection control command output by the edge computing unit. The near-infrared VCSEL structured light projector projects random speckle or coded grid patterns onto the snow surface below, and uses the Lambertian scattering characteristics of the snow surface to form a clear high-contrast speckle or grid image on the camera sensor. The narrowband filter is installed at the front of the lens of the binocular vision camera. The center wavelength of the narrowband filter is strictly matched with the output wavelength of the VCSEL structured light projector, allowing only near-infrared light of the corresponding band to pass through. It is used to filter ambient light interference, physically filter out snow reflection and ambient stray light in the visible light band, and acquire structured light images. When the binocular vision camera acquires images, the sensor only receives the return signal of the structured light projection pattern and is not affected by changes in natural lighting conditions. The gimbal servo control module is connected to the edge computing unit and the three-axis motor of the electronic pod, and is used to receive gimbal control commands output by the edge computing unit, drive the three-axis motor to rotate, and control the rotation angle and angular velocity of the pitch axis, yaw axis and roll axis motors in the electronic pod. The environmental sensor is connected to the edge computing unit and is used to synchronously collect ski resort environmental parameters, including at least one of temperature, humidity, wind speed, snow depth and visibility, to provide environmental context information for 4D spatiotemporal data packets; The edge computing unit also includes an active structured light-assisted feature extraction module, an active feedforward gimbal control module, and an environmental data integration module; The active structured light-assisted feature extraction module is connected to the near-infrared VCSEL structured light projector, the binocular vision camera, and the visual feature extraction and matching module. The active structured light-assisted feature extraction module continuously monitors the visual modality health. When the visual modality health is lower than the preset threshold for weak texture, it indicates that the current scene is a large area of ​​snow with weak texture, and natural features can no longer support pose estimation. The near-infrared VCSEL structured light projector is activated, and the density of the projection grating is adaptively adjusted according to the real-time flight altitude of the UAV. When flying at high altitude, a sparse coded grid is projected to cover a large field of view, and a dense random speckle is projected when flying at low altitude to obtain high detail resolution. The module instructs the visual feature extraction and matching module to extract speckle centers or grid corners from the image acquired by the narrowband filter as visual feature points for binocular matching and pose estimation. The emission power of the near-infrared VCSEL structured light projector is adaptively adjusted according to the ambient light intensity and image contrast. When the ambient light is strong, the power is reduced, and when the image contrast is insufficient, the power is appropriately increased. Under the premise of ensuring the quality of feature extraction, the laser output is maintained within the Class 1 human eye safety level range. The active feedforward gimbal control module is connected to the target detection and skeleton extraction module, the IMU inertial measurement module, the flight control system of the UAV platform, and the gimbal servo control module. The active feedforward gimbal control module reads multi-frame human skeleton keypoint data stored in the target detection and skeleton extraction module, establishes a simplified kinematic model of the skier including the center of mass position, velocity, and acceleration, performs state estimation on the simplified kinematic model using unscented Kalman filtering, selects sampling points near the state mean through unscented transformation and propagates the data, accurately captures the probability distribution characteristics of nonlinear motions such as turns and jumps, and predicts the skier's three-dimensional position and velocity at the next moment. Synchronously with acceleration, the active feedforward gimbal control module predicts the UAV's pose and rotational angular velocity at the next moment based on the IMU pre-integration results in the electronic pod and the motor speed data fed back by the UAV flight control. It transforms the predicted position of the skier and the predicted position of the UAV to the same world coordinate system, calculates the relative position vector between the two, calculates the desired pointing angle and feedforward compensation angular velocity of the gimbal based on the relative position vector, and superimposes the feedforward compensation angular velocity onto the output of the PID feedback control loop of the gimbal motor, so that the gimbal starts to rotate in the predicted direction before the skier's position changes. This transforms post-feedback correction into pre-prediction guidance, effectively reducing tracking delay in high-speed dynamic scenarios. The environmental data integration module is connected to the environmental sensor and the multi-source sensor fusion module. It transmits ski resort environmental parameters such as temperature, humidity, wind speed, snow depth, and visibility to the multi-source sensor fusion module. After aligning the data with the pose data and skeletal data timestamps, it integrates these parameters into the 4D spatiotemporal data package. Figure 4 As shown, each spatiotemporal data record is accompanied by a corresponding environmental context.

[0043] In the above technical solution, by expanding the active structured light projection and narrowband filtering hardware, gimbal servo control module and multi-parameter environmental sensor for the electronic pod, and deploying active structured light assisted feature extraction module, active feedforward gimbal control module and environmental data integration module, significant enhancements were achieved in three dimensions: weak texture adaptability, high-speed dynamic tracking accuracy and spatiotemporal data richness. Active structured light fundamentally solves the visual failure problem caused by weak texture in snow at the hardware level, feedforward gimbal control greatly reduces tracking latency, and environmental data integration provides complete environmental background information for motion analysis.

[0044] The number of devices and processing scale described herein are for the purpose of simplifying the description of the invention. Applications, modifications, and variations of the invention will be readily apparent to those skilled in the art.

[0045] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A method for high-precision spatiotemporal data acquisition of ski resorts using unmanned aerial vehicles (UAVs) based on multimodal perception, characterized in that: Includes the following steps: Step A: Using a binocular vision camera, an IMU inertial measurement module, and an RTK-GPS module mounted on the UAV's electronic pod, the original image sequence, acceleration and angular velocity data, and geographic location data are collected synchronously according to a unified hardware clock. The electronic pod has a built-in edge computing unit that adjusts the exposure parameters of the binocular vision camera in a closed loop based on the real-time calculated snow surface brightness statistics during the acquisition process. Step B: Perform ski resort-specific image enhancement processing on the acquired original image sequence, including at least: extracting non-snow areas in the image as regions of interest, and performing contrast-limited adaptive histogram equalization on these regions to obtain enhanced left and right eye images; Step C: Extract visual feature points from the enhanced left and right eye images and perform binocular matching. Combine acceleration and angular velocity data to perform visual-inertial tightly coupled pose estimation and calculate the visual odometry pose of the UAV in the world coordinate system. Step D: Detect skier targets in the enhanced left and right eye images, extract the two-dimensional coordinates of human skeletal key points including foot key points in the image coordinate system, calculate the three-dimensional coordinates of skeletal key points in the camera coordinate system using binocular parallax, and match and assign identity markers to skeletal key points between adjacent frames based on spatiotemporal consistency. Step E: Input the visual odometry pose, acceleration and angular velocity data, geographic location data, and the 3D coordinates of the skeletal key points in the camera coordinate system into the multi-source sensor fusion filter, perform timestamp alignment and coordinate system unification, and output the 3D spatial coordinates of the UAV and each skier in the world coordinate system and the skeletal motion posture data, generating a 4D spatiotemporal data package including microsecond-level timestamps, 3D spatial coordinates, and skeletal motion posture data.

2. The method for high-precision spatiotemporal data acquisition of ski resorts by UAV based on multimodal perception according to claim 1, characterized in that, The image enhancement process in step B also includes: The camera motion blur radius and direction are estimated using the acceleration and angular velocity data, the point spread function is reconstructed, and deconvolution processing is performed using Wiener filtering.

3. The method for high-precision spatiotemporal data acquisition of ski resorts by UAV based on multimodal perception according to claim 1, characterized in that, Step C also includes a multimodal failure compensation step based on health assessment: Visual modal health and IMU modal health are calculated in real time. Visual modal health is calculated based on the number of successfully matched feature points and their distribution variance in the image plane. IMU modal health is calculated based on the IMU pre-integration residual and temperature drift rate. When the visual modal health is below the first threshold and the IMU modal health is not below the second threshold, virtual visual observations are generated using acceleration and angular velocity data and injected into the visual-inertial tightly coupled pose estimation process. When the IMU modal health is lower than the second threshold and the visual modal health is not lower than the first threshold, virtual IMU measurements are generated using the triangulation results of visual feature points in consecutive frames and injected into the visual-inertial tightly coupled pose estimation process. When the visual modal health is below the first threshold and the IMU modal health is below the second threshold, the centroid motion data of the key points of the human skeleton are extracted, kinematic constraints are constructed in combination with the prior knowledge of human movement biomechanics, the pose result obtained by short-time integration of the IMU is corrected, and short-time stable pose estimation is maintained until any sensor modality is recovered.

4. The method for high-precision spatiotemporal data acquisition of ski resorts by UAV based on multimodal perception according to claim 3, characterized in that, Step C also includes an active structured light-assisted feature extraction step: The electronic pod also includes a near-infrared VCSEL structured light projector and a narrowband filter for the corresponding wavelength; When the visual modality health is below the weak texture preset threshold, the near-infrared VCSEL structured light projector is activated. The density of the projection grating is adaptively adjusted according to the real-time flight altitude of the drone to project random speckle or coded grid patterns onto the snow surface. Structured light images are acquired using narrowband filters, and speckle centers or grid corners are extracted as visual feature points for binocular matching and pose estimation. The emission power of the near-infrared VCSEL structured light projector is adaptively adjusted based on ambient light intensity and image contrast.

5. The method for high-precision spatiotemporal data acquisition of ski resorts by UAV based on multimodal perception according to claim 1, characterized in that, The key points of the human skeleton in step D include the tip of the nose, neck, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, left and right heels, left and right toes, and the outer sides of left and right heels, totaling 20 key points.

6. The method for high-precision spatiotemporal data acquisition of ski resorts by UAV based on multimodal perception according to claim 1, characterized in that, In step D, the model used to detect skier targets and extract key points of the human skeleton is a lightweight deep learning model, which is trained through cloud knowledge distillation. The Transformer model is used as a high-parameter teacher network, and it is pre-trained on a dataset that includes skiing postures and snow obstacles to extract high-dimensional feature response maps. The lightweight deep learning model is used as a student network, with the same image as input and the corresponding feature layer output is extracted. The training is completed by minimizing the KL divergence between the corresponding feature layer output of the student network and the high-dimensional feature response map of the teacher network.

7. The method for high-precision spatiotemporal data acquisition of ski resorts by UAV based on multimodal perception according to claim 1, characterized in that, It also includes active feedforward gimbal control steps based on human skeletal kinematics prediction: Using multi-frame human skeleton keypoint data stored in history, a simplified kinematic model of the skier, including center of mass position, velocity, and acceleration, is established. State estimation of a skier’s simplified kinematic model is performed using unscented Kalman filtering to predict the skier’s three-dimensional position, velocity, and acceleration at the next moment. Based on the pre-integration results of the IMU in the electronic pod and the motor speed data fed back by the UAV flight control, the attitude and rotational angular velocity of the UAV at the next moment are predicted. Transform the skier's predicted position and the drone's predicted position to the same world coordinate system, calculate the relative position vector between the two, and calculate the gimbal's expected pointing angle and feedforward compensation angular velocity based on the relative position vector. The feedforward compensation angular velocity is superimposed on the output of the PID feedback control loop of the gimbal motor.

8. A high-precision spatiotemporal data acquisition system for ski resorts based on multimodal perception, characterized in that: The system for performing the method of claim 1, comprising: Unmanned aerial vehicle (UAV) platform; An electronic pod is mounted on the UAV platform. The electronic pod includes a sensor unit and an edge computing unit. The sensor unit includes a binocular vision camera, an IMU inertial measurement module, and an RTK-GPS module. The edge computing unit is electrically connected to the sensor unit. The sensor unit is used to synchronously acquire raw image sequences, acceleration and angular velocity data, and geographic location data according to a unified hardware clock, and transmit the raw image sequences, acceleration and angular velocity data, and geographic location data to the edge computing unit; The edge computing unit includes an image enhancement module, an exposure control module, a visual feature extraction and matching module, a visual-inertial tightly coupled pose estimation module, a target detection and skeleton extraction module, a spatiotemporal consistency matching module, and a multi-source sensor fusion module. The image enhancement module is used to perform ski resort-specific image enhancement processing on the original image sequence, including extracting non-snow areas in the image as regions of interest and performing contrast-limited adaptive histogram equalization on the regions, and outputting enhanced left and right eye images; The exposure control module is connected to the image enhancement module and the binocular vision camera respectively. It is used to receive the snow surface brightness statistics calculated based on the original image sequence from the image enhancement module, and generate an exposure adjustment command based on the statistics and output it to the binocular vision camera to adjust the exposure parameters of the binocular vision camera in a closed loop. The visual feature extraction and matching module is connected to the image enhancement module, receives the enhanced left and right eye images, and is used to extract visual feature points from the enhanced left and right eye images and perform binocular matching. The visual-inertial tightly coupled pose estimation module is connected to the visual feature extraction and matching module and the IMU inertial measurement module. It receives the matching results of visual feature points as well as acceleration and angular velocity data, and is used to perform visual-inertial tightly coupled pose estimation to calculate the visual odometry pose of the UAV in the world coordinate system. The target detection and skeleton extraction module is connected to the image enhancement module, receives the enhanced left and right eye images, and is used to detect skier targets in the enhanced left and right eye images, extract the two-dimensional coordinates of human skeleton key points including foot key points in the image coordinate system, and use binocular parallax to calculate the three-dimensional coordinates of skeleton key points in the camera coordinate system. The spatiotemporal consistency matching module is connected to the target detection and skeleton extraction module, and receives the skeletal key points of adjacent frames, which are used to match the skeletal key points of adjacent frames and assign identity identifiers based on spatiotemporal consistency. The multi-source sensor fusion module is connected to the visual-inertial tightly coupled pose estimation module, the IMU inertial measurement module, the RTK-GPS module, and the target detection and skeleton extraction module. It receives visual odometry pose, acceleration and angular velocity data, geographic location data, and three-dimensional coordinates of skeletal key points. It is used to perform timestamp alignment and coordinate system unification, output the three-dimensional spatial coordinates of the UAV and each skier in the world coordinate system, as well as the skeletal motion posture data, and generate a 4D spatiotemporal data packet including microsecond-level timestamps, three-dimensional spatial coordinates, and skeletal motion posture data. A ground control station, which is communicatively connected to the electronic pod, is used to receive and display real-time data sent by the electronic pod. The cloud server is communicatively connected to the ground control station and / or the electronic pod, and is used to receive 4D spatiotemporal data packets, perform big data analysis and model optimization, and send the updated model to the edge computing unit.

9. The UAV high-precision spatiotemporal data acquisition system for ski resorts based on multimodal perception according to claim 8, characterized in that, The edge computing unit also includes a motion blur compensation module and a multimodal failure compensation module; The motion blur compensation module is connected to the sensor module and the image enhancement module. It is used to receive acceleration and angular velocity data, estimate the camera motion blur radius and direction, reconstruct the point spread function, perform deconvolution processing using Wiener filtering, and output the processed image sequence to the image enhancement module. The multimodal failure compensation module is connected to the visual-inertial tightly coupled pose estimation module, the visual feature extraction and matching module, the IMU inertial measurement module, and the target detection and skeleton extraction module. It is used to calculate visual modal health and IMU modal health in real time. Visual modal health is calculated based on the number of successfully matched feature points and their distribution variance in the image plane. IMU modal health is calculated based on the IMU pre-integration residual and the temperature drift rate. When the visual modal health is below the first threshold and the IMU modal health is not below the second threshold, virtual visual observations are generated using acceleration and angular velocity data and injected into the visual-inertial tightly coupled pose estimation module. When the IMU modal health is lower than the second threshold and the visual modal health is not lower than the first threshold, virtual IMU measurements are generated using the triangulation results of visual feature points in consecutive frames and injected into the visual-inertial tightly coupled pose estimation module. When the visual modal health is below the first threshold and the IMU modal health is below the second threshold, the centroid motion data of key points of the human skeleton are extracted, kinematic constraints are constructed by combining human movement biomechanical priors, the pose results obtained by short-time integration of the IMU are corrected, and short-time stable pose estimation is maintained until any sensor modality is recovered.

10. The UAV high-precision spatiotemporal data acquisition system for ski resorts based on multimodal perception according to claim 8, characterized in that, The electronic pod also includes a near-infrared VCSEL structured light projector, a narrowband filter of the corresponding wavelength, a gimbal servo control module, and an environmental sensor. The near-infrared VCSEL structured light projector is connected to the edge computing unit and is used to receive the projection control command output by the edge computing unit to project random speckle or coded grid patterns onto the snow surface. The narrowband filter is installed at the front of the lens of the binocular vision camera to filter ambient light interference and acquire structured light images. The gimbal servo control module is connected to the edge computing unit and the three-axis motor of the electronic pod, and is used to receive gimbal control commands output by the edge computing unit and drive the three-axis motor to rotate. The environmental sensor is connected to the edge computing unit and is used to synchronously collect ski resort environmental parameters, including at least one of temperature, humidity, wind speed, snow depth and visibility. The edge computing unit also includes an active structured light-assisted feature extraction module, an active feedforward gimbal control module, and an environmental data integration module; The active structured light-assisted feature extraction module is connected to the near-infrared VCSEL structured light projector, the binocular vision camera, and the visual feature extraction and matching module. When the visual modality health is lower than the weak texture preset threshold, the module activates the near-infrared VCSEL structured light projector, adaptively adjusts the projection grating density according to the real-time flight altitude of the UAV, extracts speckle centers or grid corners as visual feature points, performs binocular matching and pose estimation, and adaptively adjusts the emission power of the near-infrared VCSEL structured light projector according to the ambient light intensity and image contrast. The active feedforward gimbal control module is connected to the target detection and skeleton extraction module, the IMU inertial measurement module, the flight control system of the UAV platform, and the gimbal servo control module. It is used to establish a simplified kinematic model of the skier, including the center of mass position, velocity, and acceleration, using multi-frame human skeleton key point data stored in history. The state of the simplified kinematic model of the skier is estimated by unscented Kalman filtering to predict the skier's three-dimensional position, velocity, and acceleration at the next moment. Based on the IMU pre-integration results in the electronic pod and the motor speed data fed back by the UAV flight control, the pose and self-rotation angular velocity of the UAV at the next moment are predicted. The predicted position of the skier and the predicted position of the UAV are transformed to the same world coordinate system, and the relative position vector between the two is calculated. The desired pointing angle and feedforward compensation angular velocity of the gimbal are calculated based on the relative position vector. The feedforward compensation angular velocity is superimposed on the output of the PID feedback control loop of the gimbal motor. The environmental data integration module is connected to the environmental sensor and the multi-source sensor fusion module, and is used to transmit the ski resort environmental parameters to the multi-source sensor fusion module and integrate them into the 4D spatiotemporal data package.