Over-the-horizon hovercar grid occupation prediction method and device
By enhancing the image of flying cars and implementing a three-dimensional grid prediction model, the problem of insufficient accuracy in predicting environmental changes of flying cars in existing technologies has been solved, achieving high-precision environmental perception and decision-making capabilities.
Patent Information
- Application Number
- CN202510755340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
When existing technologies directly apply occupancy grid prediction methods for ground vehicle or robot scenes to flying cars, the accuracy of environmental change prediction is affected.
A pre-trained prediction model is used to enhance the captured images. The enhanced images are projected onto a three-dimensional occupancy grid in a unified coordinate system using an improved voxel rendering mechanism and spatial alignment strategy. Combined with the flight status of the flying car, environmental changes at future moments are predicted.
It significantly improves the accuracy of environmental change prediction for flying cars in complex high-altitude environments and provides precise spatial perception and decision-making capabilities.
Smart Images

Figure CN120635862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving perception technology, and in particular to a method and device for predicting a grid occupied by a flying car beyond visual range. Background Art
[0002] Flying cars, a key development direction for future urban transportation, are attracting widespread attention. Occupancy grid prediction technology is crucial for sensing the surrounding environment. It allows flying cars to understand the occupancy of their surroundings in real time, providing an accurate basis for route planning and decision-making. However, this field still faces many challenges.
[0003] Related technologies use sensors like lidar and cameras to acquire three-dimensional environmental information, enabling predictions of future environmental changes. However, these technologies are primarily designed for use in ground vehicles or robotics, and direct application to flying cars presents incompatibility issues, impacting the accuracy of their predictions. Summary of the Invention
[0004] This invention provides a method and device for predicting the grid occupancy of a flying car beyond visual range (BVR). This method addresses the problem that direct application of related technologies to flying cars can affect the accuracy of their predictions of environmental changes. The technical solution is as follows:
[0005] In one aspect, a method for predicting a grid occupied by a flying car beyond visual range is provided, the method comprising:
[0006] Obtaining a captured image of the flying car at the current moment, and inputting the captured image into a pre-trained prediction model;
[0007] The prediction model is used to process the acquired image as follows:
[0008] Performing image enhancement on the collected image to obtain an enhanced image;
[0009] Using an improved voxel rendering mechanism and spatial alignment strategy, the enhanced image is projected onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying car, so as to construct the three-dimensional occupancy grid at the current moment;
[0010] Using the flying car's flight status at the current moment and several historical moments, as well as the current moment's three-dimensional occupancy grid, predict environmental changes in the future;
[0011] Obtain the prediction result output by the prediction model.
[0012] In another aspect, a device for predicting a grid occupied by a flying car beyond visual range is provided, the device comprising:
[0013] A first acquisition unit is used to acquire a captured image of the flying car at a current moment and input the captured image into a pre-trained prediction model;
[0014] A prediction unit is configured to process the captured image using the prediction model as follows: performing image enhancement on the captured image to obtain an enhanced image; utilizing an improved voxel rendering mechanism and a spatial alignment strategy to project the enhanced image onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying vehicle, thereby constructing the three-dimensional occupancy grid at the current moment; and utilizing the flight state of the flying vehicle at the current moment and at several historical moments, as well as the three-dimensional occupancy grid at the current moment, to predict environmental changes at future moments;
[0015] The second acquisition unit is used to obtain the prediction result output by the prediction model.
[0016] On the other hand, a computer device is provided, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of the above-mentioned method for predicting the occupancy grid of a flying car beyond visual range.
[0017] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for predicting the occupancy grid of a flying car beyond visual range are implemented.
[0018] On the other hand, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for predicting the occupancy grid of a flying car beyond visual range.
[0019] The technical solution provided by the present invention can at least bring the following beneficial effects:
[0020] By pre-training a prediction model, images are captured at each current moment, fed into the prediction model, and processed using the prediction model to output prediction results. During this process, the prediction model first enhances the captured images to effectively improve their texture clarity and semantic integrity, providing high-quality input data for the subsequent construction of a three-dimensional occupancy grid. Then, using an improved voxel rendering mechanism and spatial alignment strategy, the enhanced images are projected onto a three-dimensional occupancy grid in a unified coordinate system based on the vehicle's current flight state. This incorporation of the vehicle's flight state effectively restores the environment's three-dimensional structural information, enabling the vehicle to exhibit strong stability and adaptability from a high-altitude perspective and providing accurate spatial perception. Finally, using the current and several historical flight states, as well as the current 3D occupancy grid, future environmental changes are predicted, thereby enhancing the vehicle's decision-making capabilities in beyond-visual-range environments. This solution significantly improves the accuracy of environmental change predictions for vehicles in complex, high-altitude environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 This is a flow chart of a method for predicting a grid occupied by a flying car beyond visual range, provided by one embodiment of the present invention;
[0023] Figure 2 This is a structural diagram of a beyond-visual-range flying car occupancy grid prediction device provided by one embodiment of the present invention;
[0024] Figure 3 This is a hardware architecture diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0026] Please refer to Figure 1An embodiment of the present invention provides a method for predicting a grid occupied by a flying car beyond visual range, the method comprising:
[0027] Step 100: Acquire a captured image of the flying car at the current moment, and input the captured image into a pre-trained prediction model;
[0028] Step 102: The prediction model is used to process the captured image as follows: performing image enhancement on the captured image to obtain an enhanced image; utilizing an improved voxel rendering mechanism and spatial alignment strategy, projecting the enhanced image onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying vehicle, thereby constructing a three-dimensional occupancy grid at the current moment; and utilizing the current flight state of the flying vehicle and several historical moments, as well as the current three-dimensional occupancy grid, to predict environmental changes at future moments.
[0029] Step 104: Obtain the prediction result output by the prediction model.
[0030] In an embodiment of the present invention, a prediction model is pre-trained to capture images at each current moment. The captured images are then fed into the prediction model and processed using the prediction model to output prediction results. During the prediction model processing, the captured images are first enhanced to effectively improve texture clarity and semantic integrity, providing high-quality input data for the subsequent construction of a three-dimensional occupancy grid. Then, an improved voxel rendering mechanism and spatial alignment strategy are used to project the enhanced images onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying vehicle. This incorporates the flying vehicle's flight state, effectively restoring the three-dimensional structural information of the environment. This allows the flying vehicle to exhibit strong stability and adaptability from a high-altitude perspective, providing the flying vehicle with accurate spatial perception. Finally, the flight state at the current moment and several historical moments, as well as the current three-dimensional occupancy grid, are used to predict environmental changes at future moments, thereby enhancing the flying vehicle's decision-making capabilities in beyond-visual-range environments. This solution can significantly improve the accuracy of a flying vehicle's prediction of environmental changes in complex, high-altitude environments.
[0031] Described below Figure 1 How to perform the steps shown.
[0032] First, for steps 100 and 104, the flying car acquires collected images in real time while driving. The collected images may include multi-view two-dimensional graphics acquired by the camera and three-dimensional point cloud data acquired by the lidar; and the collected images acquired at the current moment are input into the pre-trained prediction model, which processes the collected images to output prediction results.
[0033] Then, the processing of the prediction model in step 102 is described respectively.
[0034] With respect to step 1020 , image enhancement is performed on the collected image to obtain an enhanced image.
[0035] Taking into account the problems of image blur, detail loss and information attenuation caused by factors such as long distance, complex lighting and meteorological interference in the beyond-visual-range environment of flying cars, in an embodiment of the present invention, the prediction model may include an image enhancement module for performing image enhancement processing on the collected image to obtain an enhanced image.
[0036] In an embodiment of the present invention, the image enhancement module may be implemented using the AutoLUT-Aero image enhancement method of a learnable lookup table (LUT). Specifically, step 1020 may include:
[0037] Step 10200: For each pixel in the captured image, perform the following steps: using the RGB value of the pixel as an index, and using a pre-acquired lookup table and the index to find a mapping value corresponding to the index, and using the mapping value as a new RGB value of the pixel;
[0038] Step 10202: Calculate enhancement values based on new RGB values of pixels in the captured image to obtain an enhanced image.
[0039] AutoLUT-Aero, an image enhancement method that uses a learnable lookup table (LUT), relies on the fast mapping capabilities of the LUT. Using RGB values as indices, it maps the enhanced RGB values to the LUT, thus avoiding lengthy deep convolution calculations. Furthermore, in step 10202, trilinear interpolation can be used to calculate the enhanced value based on the new RGB values of the pixel.
[0040] Furthermore, considering that enhancing all pixels in the acquired image not only requires a large amount of computation but may also cause over-processing of some texture areas, in one embodiment of the present invention, an automatic sampling mechanism is introduced to sample the pixels of the acquired image. Specifically, before performing image enhancement on the acquired image through the automatic sampling method, the following steps may be further included:
[0041] Dividing the acquired image into regions and calculating a region score using image parameters within the region; the image parameters include at least one of texture intensity, saliency, and information entropy;
[0042] Each region is automatically sampled using the region score of the acquired image to obtain sampling pixels, so as to perform the image enhancement on the sampling pixels; wherein the number of sampling pixels is proportional to the corresponding region score value.
[0043] In one implementation, the region division can be performed by evenly dividing the image into multiple regions, or by dividing multiple adjacent pixels whose RGB values are within a specified range into one region according to the RGB values of the pixels; in another implementation, each pixel is treated as a region to select pixels with high information density during automatic sampling.
[0044] By performing regional division, the image parameters within the region can be dynamically analyzed, and the regional score can be calculated using the scoring function, so that the enhancement resources can be focused on the most informative areas in the image.
[0045] In one implementation, the scoring function may be:
[0046]
[0047] Among them, S(i,j) is the score value of the pixel at the (i,j) position, is the image gradient, which is used to reflect the texture change E(I i,j ) is the local entropy Sal(I i,j ) is the saliency map value, and λ1, λ2, and λ3 are all weight coefficients.
[0048] In this way, image enhancement processing can be performed only on the sampled pixels, thereby reducing the computational cost.
[0049] Furthermore, to compensate for the limitations of the lookup table in processing structural edges and detailed textures, a lightweight residual network can be introduced into the image enhancement module to use a shallow convolutional neural network to learn the lost high-frequency components in the image and restore the structural edges. Specifically, after calculating the region score, the following steps can also be included:
[0050] Using a residual network, the edge of the collected image is repaired based on the regions corresponding to the multiple maximum region scores to obtain a residual map;
[0051] The enhanced image is superimposed on the residual image to obtain an updated enhanced image, and the three-dimensional grid occupancy is constructed using the updated enhanced image.
[0052] It can be seen that the image enhancement module and the residual network used for edge restoration are a parallel architecture, which enables the prediction model to maintain the rapid enhancement capability of the lookup table while enhancing the fidelity of structural details, thereby obtaining sharper, clearer, and more structurally coherent images in complex high-altitude flight scenes.
[0053] In order to achieve end-to-end training from inputting the captured image to obtaining the enhanced image, a joint loss can be used for training. This can include:
[0054]
[0055] in, is the pixel-level error loss, is the perceptual loss, The regularized loss of the residual, α, β, and γ are all hyperparameters, To enhance the image, I gt is the real reference image, To enhance the semantic features of the image, φ(I gt ) is the semantic feature of the real reference image.
[0056] In an embodiment of the present invention, by superimposing the residual map and the enhanced image, a high-quality image with complete structure and clear semantics can be obtained, and during the prediction process, the image enhancement module can achieve extremely low latency, relying on the fast retrieval capability of the lookup table, and can achieve real-time image processing capabilities of more than 30fps on the edge computing device without relying on the GPU. In addition, the AutoLUT-Aero module can not only improve the flying car's image perception ability of distant blurred targets, but also provide a reliable visual foundation for subsequent three-dimensional occupancy grid modeling and state prediction. Especially in complex high-altitude environments, it can effectively alleviate image degradation problems caused by long distances, weather changes, sparse textures, etc., so that subsequent models can also have clear and stable environmental understanding capabilities in beyond-visual-range scenarios, building a key "seeing clearly" capability pillar in the entire perception system.
[0057] For step 1022, the enhanced image is projected into a three-dimensional occupancy grid in a unified coordinate system based on the flight state of the flying car at the current moment using an improved voxel rendering mechanism and spatial alignment strategy to construct the three-dimensional occupancy grid at the current moment.
[0058] In an embodiment of the present invention, it is necessary to accurately map the two-dimensional enhanced image into a dense, clearly structured three-dimensional occupancy grid (3D Occupancy Grid) to achieve spatial reconstruction of the structure of the flying car's surrounding environment. This process not only needs to consider the mapping accuracy from the image to the voxel space, but also needs to dynamically adapt to the impact of flight altitude, pitch angle and viewing angle changes on perception accuracy during flight. Based on this, in an embodiment of the present invention, the prediction model can include a three-dimensional modeling module, which is used to utilize an improved voxel rendering mechanism and spatial alignment strategy so that the modeling method of the image center can maintain structural consistency and spatial accuracy in a highly dynamic flight environment.
[0059] Specifically, this step 1022 may include:
[0060] The depth prediction network lightweight model is used to predict the depth value of each pixel from the enhanced image to obtain a dense depth map corresponding to the enhanced image.
[0061] Using the camera intrinsic parameter K and the flying car's attitude matrix T cam→world , map each pixel (u,v) to a three-dimensional point (x,y,z);
[0062] Using the current flight state of the flying car, the transformation parameters between the image and the real space are estimated to construct a perspective compensation matrix, and the perspective compensation matrix is used to perform rotation compensation on the three-dimensional point;
[0063] Project the compensated 3D points to the 3D occupancy grid v∈{0,1} in the unified coordinate system X×Y×Z , to indicate whether each grid in the space is occupied.
[0064] Among them, the attitude matrix of the flying car can be provided by the inertial navigation system.
[0065] The method of mapping pixels to three-dimensional points can be achieved by the following formula:
[0066]
[0067] Furthermore, considering that the images captured by the flying car in a complex high-altitude environment have problems of tilt and field of view distortion, an embodiment of the present invention provides a spatial alignment module (Alignment Head). Before projecting the three-dimensional points onto the occupied grid, the flying car's current flight state (including flight altitude h, pitch angle θ, yaw angle φ) is used to estimate the transformation parameters between the image and the real space to construct a perspective compensation matrix. The three-dimensional points are rotated and compensated, so that the alignment transformation in the projection stage can be achieved. Among them, the perspective compensation matrix It is used to represent the rigid transformation matrix that aligns a 3D point in the camera coordinate system or image frame to the world coordinate system or a unified space coordinate system. The entire matrix belongs to SE(3), thus ensuring that data from different viewpoints or different time frames are mapped to a unified 3D space. SE(3) is used to represent all rigid body transformations in 3D space, including rotations and translations. Each element in SE(3) can be represented by a 4×4 matrix.
[0068] Furthermore, the prediction model can also include a semantic fusion module that performs semantic segmentation on the enhanced image, obtains the category label and boundary information of each pixel, and writes them into the 3D occupancy grid during the projection process, combining them with the dense depth map to form a 3D structure with semantic attributes. The resulting 3D occupancy grid not only contains spatial geometry information (occupied / free) but also includes semantic descriptions such as categories (such as roads, buildings, pedestrians, and obstacles), providing richer spatial perception input for subsequent modules.
[0069] In one implementation, the occupancy grid of an embodiment of the present invention is represented as a continuous probability form, and each voxel stores its occupancy probability [0, 1] and a semantic vector c∈R^C, which facilitates subsequent differentiable optimization and uncertainty modeling.
[0070] In an embodiment of the present invention, step 1022 is combined with step 1020. On the one hand, the enhanced image can be used to improve the accuracy of depth and semantic estimation. On the other hand, the uncertainty of high-altitude modeling can be constrained through state guidance, so that the reconstruction of the entire three-dimensional occupancy grid can still maintain a stable and reliable reconstruction effect in complex flight scenarios such as urban airspace, forests, and mountainous areas. After these two steps, the flying car not only has a clear "image" but also has a structured "spatial perception", building a basic picture for understanding the surrounding three-dimensional world, providing spatial support for subsequent predictions and decisions.
[0071] With respect to step 1024, the flight status of the flying car at the current moment and several historical moments and the three-dimensional occupancy grid at the current moment are used to predict environmental changes at future moments.
[0072] Considering that the three-dimensional occupancy grid at the current moment alone cannot fully support path planning and obstacle avoidance requirements, in an embodiment of the present invention, the prediction model includes a 4D occupancy prediction module, which is used to realize dynamic modeling and deduction of environmental changes at future moments, thereby having beyond-visual-range cognition and prediction capabilities.
[0073] Specifically, the three-dimensional grid O occupied by the flying car at the current moment can be t ∈R X×Y×Z×C As the input of the 4D occupancy prediction module, the flight status of the combined flying car at the current moment and several historical moments (such as speed acceleration attitude quaternion q, etc.), predict the changes in the environment in the next several time steps, that is, {O t+1 ,O t+2 ,...,O t+n}, thus forming a temporal four-dimensional occupancy expression.
[0074] The 4D occupancy prediction module can use a Transformer model based on a temporal encoder-decoder structure as the core temporal modeling module. The 4D occupancy prediction module receives the three-dimensional occupancy grid sequence and the flight state sequence of the previous T frames, first encoding the voxel grid of each frame into a high-dimensional semantic vector, and then integrating the dynamic state information of the aircraft into the time series modeling process through the state condition fusion module. This fusion strategy can guide the network to consider the inertial trend of the system when making spatial predictions, so that the prediction results not only reflect the movement of the environment itself (such as the forward direction of dynamic obstacles), but also reflect the impact of the perspective change brought about by the flying car's own movement.
[0075] The optimization goal of the 4D occupancy prediction module is to minimize the difference between the predicted voxel distribution and the true future occupancy state. The loss function L OCC for:
[0076]
[0077] Where BCE(·) represents the binary cross entropy loss at the voxel level, TV(·) is the spatial total variation regularization term of the predicted voxels, which is used to encourage the smoothness and structural consistency of the prediction results, and λ smooth is the adjustment coefficient, and the three-dimensional point (x, y, z) corresponds to a voxel in the three-dimensional grid If the occupancy probability of the voxel is close to 1, it means that the position may be occupied by an obstacle. If the occupancy probability of the voxel is close to 0, it means that the position is empty. For the real occupation label;
[0078] To further improve the spatiotemporal consistency of predictions, a motion prior constraint module based on optical flow can be introduced. By tracking the motion trajectory of voxels in the three-dimensional occupancy grid sequence, its change trend can be estimated and used as an auxiliary supervision signal to optimize the coherence of occupancy boundary changes.
[0079] In practical applications, the 4D occupancy prediction module supports multi-scale time step prediction, and can adaptively output multi-step prediction results such as 0.2s, 0.5s, and 1.0s according to the flight speed and control frequency to meet the forward-looking needs of different tasks. For example, in high-speed flight, more attention can be paid to the movement trend of distant areas, while in low-speed hovering scenarios, the changing details of the local environment are emphasized. In order to adapt to the modeling of three-dimensional occupancy grids at different resolutions, the 4D occupancy prediction module also adopts a multi-resolution voxel fusion mechanism (Hierarchical Voxel Decoder) to hierarchically reconstruct occupancy maps of different accuracies in the prediction stage, and optimize them through cross-scale consistency loss.
[0080] Furthermore, to account for the common occlusions and blind spots in high-altitude flight environments, the 4D occupancy prediction module incorporates an occlusion-aware refinement head (Occlusion-Aware Refinement Head). This leverages temporal consistency and state-guided results to rationally infer invisible areas during prediction. This mechanism is particularly advantageous in scenarios such as high-altitude flight over urban canyons and high-speed passage through woodlands, effectively reducing the risk of "missing visibility."
[0081] By introducing 4D occupancy prediction based on flight status, the prediction model not only possesses current 3D mapping capabilities but also further models environmental evolution trends, providing forward-looking perception capabilities for flying cars. In conjunction with the previous two modules (the image enhancement module and the 3D modeling module), this 4D occupancy prediction module achieves a closed-loop perception and reasoning loop from image to space and from space to time, laying a solid foundation for safe flight, path planning, and autonomous decision-making.
[0082] In order to achieve high-precision prediction using the above prediction model, it is necessary to use training samples to obtain a prediction model with strong robustness and generalization ability. However, there is currently a lack of large-scale 3D graphic label data. In view of the practical problems of high cost of 3D occupancy labeling, complex occlusion, and large labeling deviation in the flying car mission, in one embodiment of the present invention, the prediction model can be trained as follows:
[0083] Acquire a pseudo-label voxel set; the pseudo-label voxel set includes a plurality of first samples; the first samples are generated by mapping semantic information and depth information of a two-dimensional sample image to a three-dimensional voxel space to form the first samples;
[0084] Acquire a plurality of second samples; the second samples are fine-label three-dimensional label data;
[0085] A dual-path semi-supervised loss structure is adopted. Several second samples are used on the main path to calculate the fine supervision loss, and multiple first samples are used on the auxiliary path to calculate the pseudo supervision loss. A confidence gating mechanism is used to select voxel locations with confidence higher than a set value to participate in the training process, so as to complete the training of the prediction model using the total loss function.
[0086] The total loss function includes at least fine supervision loss, pseudo supervision loss and consistency constraint loss; the consistency constraint loss is used to represent the consistency constraint of the prediction results in different perspectives and states for the same scene.
[0087] The first sample is generated by using a pre-trained semantic segmentation network to obtain a pixel-level semantic probability map for the input 2D sample image. A depth map is then obtained using a monocular depth estimation network. Using camera intrinsics and the vehicle's pose matrix, the 2D pixels are mapped to a 3D point cloud, which is then projected into voxel space to form the first sample. This mapping process incorporates an uncertainty modeling mechanism to weightedly mask low-confidence regions in the depth estimate to prevent erroneous projections from interfering with training.
[0088] It should be noted that the consistency constraint loss is constructed by combining temporal consistency loss and perspective invariance loss, which can effectively enhance the stability of the prediction model under complex flight dynamics.
[0089] During model training, the high-quality images output by the image enhancement module AutoLUT-Aero are integrated as input to the perception frontend, further improving the discriminability of pseudo-labels. Furthermore, to avoid cumulative errors caused by semantic drift in pseudo-labels, a self-distillation mechanism is employed, using the current model's high-confidence predictions as the source of pseudo-labels for future training, continuously optimizing pseudo-label quality over multiple training cycles.
[0090] Furthermore, to improve the model's adaptability in sparse scenes and long-tail categories, an adversarial learning mechanism can be introduced. This mechanism uses a voxel-level discriminator to distinguish occupancy patterns from real annotations and pseudo-labels. During the training phase, adversarial targets are used to optimize the representation capabilities of the generator network, improving recognition and generalization performance in complex scenarios. This module is particularly suitable for low-sample scenarios common in flying cars, such as mountainous canyons, urban elevated roads, and woodland edges.
[0091] This semi-supervised training approach not only significantly reduces reliance on costly 3D annotated data, but also combines strategies such as 2D label mapping, state prior guidance, and image enhancement optimization to establish a scalable, efficient, and realistic training solution for flight scenarios. This provides a feasible engineering path for large-scale deployment of the entire model. Ultimately, the embodiment of the present invention achieves a closed-loop cross-modal modeling process from sparse labeling to high-precision predictions, demonstrating excellent prediction accuracy and environmental adaptability in real-world flight experiments.
[0092] Please refer to Figure 2 An embodiment of the present invention provides a beyond-visual-range flying car occupancy grid prediction device, the device comprising:
[0093] A first acquisition unit 200 is used to acquire a captured image of the flying car at a current moment and input the captured image into a pre-trained prediction model;
[0094] The prediction unit 202 is configured to process the captured image using the prediction model as follows: perform image enhancement on the captured image to obtain an enhanced image; utilize an improved voxel rendering mechanism and spatial alignment strategy to project the enhanced image onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying vehicle to construct the current three-dimensional occupancy grid; and utilize the current flight state of the flying vehicle and several historical moments, as well as the current three-dimensional occupancy grid, to predict environmental changes at future moments.
[0095] The second acquiring unit 204 is configured to acquire the prediction result output by the prediction model.
[0096] In one embodiment of the present invention, the prediction model includes an image enhancement module, a 3D modeling module, and a 4D occupancy prediction module;
[0097] The image enhancement module is used to perform the image enhancement on the collected image by automatic sampling to obtain the enhanced image, specifically including:
[0098] For each pixel in the captured image, the following steps are performed: using the RGB value of the pixel as an index, and using a pre-acquired lookup table and the index to find a mapping value corresponding to the index, and using the mapping value as a new RGB value of the pixel;
[0099] An enhancement value is calculated based on the new RGB values of the pixels in the collected image to obtain an enhanced image.
[0100] In one embodiment of the present invention, before the image enhancement module performs image enhancement on the acquired image by automatic sampling, it is also used to: divide the acquired image into regions, and calculate regional scores using image parameters within the regions; the image parameters include at least one of texture intensity, saliency and information entropy; automatically sample each region using the regional score of the acquired image to obtain sampling pixel points, so as to perform the image enhancement on the sampling pixel points; wherein the number of sampling pixel points is proportional to the corresponding regional score value.
[0101] In one embodiment of the present invention, after calculating the region scores, the image enhancement module is further used to use a residual network to perform edge repair on the acquired image based on the regions corresponding to several maximum region scores to obtain a residual map; the enhanced image is superimposed with the residual map to obtain an updated enhanced image, so as to use the updated enhanced image to execute the construction of the three-dimensional grid occupancy.
[0102] In one embodiment of the present invention, the three-dimensional modeling module is used to: use a lightweight depth prediction network model to predict the depth value of each pixel from the enhanced image to obtain a dense depth map corresponding to the enhanced image; use the posture matrix of the flying car in the camera to map each pixel into a three-dimensional point; use the flight state of the flying car at the current moment to estimate the transformation parameters between the image and the real space to construct a perspective compensation matrix, and use the perspective compensation matrix to perform rotation compensation on the three-dimensional points; and project the compensated three-dimensional points into a three-dimensional occupancy grid in a unified coordinate system.
[0103] In one embodiment of the present invention, the prediction model is trained in the following manner:
[0104] Acquire a pseudo-label voxel set; the pseudo-label voxel set includes a plurality of first samples; the first samples are generated by mapping semantic information and depth information of a two-dimensional sample image to a three-dimensional voxel space to form the first samples;
[0105] Acquire a plurality of second samples; the second samples are fine-label three-dimensional label data;
[0106] A dual-path semi-supervised loss structure is adopted. Several second samples are used on the main path to calculate the fine supervision loss, and multiple first samples are used on the auxiliary path to calculate the pseudo supervision loss. A confidence gating mechanism is used to select voxel locations with confidence higher than a set value to participate in the training process, so as to complete the training of the prediction model using the total loss function.
[0107] The total loss function includes at least fine supervision loss, pseudo supervision loss and consistency constraint loss; the consistency constraint loss is used to represent the consistency constraint of the prediction results in different perspectives and states for the same scene.
[0108] It should be noted that the above-described embodiments of the device for predicting the occupancy grid of a beyond-visual-range flying vehicle are merely illustrative of the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be distributed among different functional modules as needed, i.e., the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the above-described embodiments of the device for predicting the occupancy grid of a beyond-visual-range flying vehicle and the embodiments of the method for predicting the occupancy grid of a beyond-visual-range flying vehicle share the same concept. The specific implementation process is detailed in the method embodiments and will not be further elaborated here.
[0109] The embodiment of the present application also provides a computer device, please refer to Figure 3The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the beyond-visual-range flying car occupancy grid prediction method provided by each of the above-mentioned method embodiments.
[0110] An embodiment of the present application also provides a computer-readable storage medium, on which is stored at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement the beyond-visual-range flying car occupancy grid prediction method provided by the above-mentioned method embodiments.
[0111] An embodiment of the present application also provides a computer program product, which includes a computer program. A processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the beyond-visual-range flying car occupancy grid prediction method described in any of the above embodiments.
[0112] For the convenience of description, the above systems or devices are described as being divided into various modules or units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0113] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0114] Finally, it should be noted that, in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0115] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for predicting the grid occupied by a flying car beyond visual range, characterized in that: The method comprises: Obtaining a captured image of the flying car at the current moment, and inputting the captured image into a pre-trained prediction model; The prediction model is used to process the acquired image as follows: Performing image enhancement on the collected image to obtain an enhanced image; Using an improved voxel rendering mechanism and spatial alignment strategy, the enhanced image is projected onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying car, so as to construct the three-dimensional occupancy grid at the current moment; Using the flying car's flight status at the current moment and several historical moments, as well as the current moment's three-dimensional occupancy grid, predict environmental changes in the future; Obtain the prediction result output by the prediction model.
2. The method according to claim 1, characterized in that The step of performing image enhancement on the collected image by automatic sampling to obtain an enhanced image includes: For each pixel in the captured image, the following steps are performed: using the RGB value of the pixel as an index, and using a pre-acquired lookup table and the index to find a mapping value corresponding to the index, and using the mapping value as a new RGB value of the pixel; An enhancement value is calculated based on the new RGB values of the pixels in the collected image to obtain an enhanced image.
3. The method according to claim 2, characterized in that Before performing image enhancement on the acquired image by automatic sampling, the method further includes: Dividing the acquired image into regions and calculating a region score using image parameters within the region; the image parameters include at least one of texture intensity, saliency, and information entropy; Each region is automatically sampled using the region score of the acquired image to obtain sampling pixels, so as to perform the image enhancement on the sampling pixels; wherein the number of sampling pixels is proportional to the corresponding region score value.
4. The method according to claim 3, characterized in that After calculating the regional score, the following steps are further included: Using a residual network, the edge of the collected image is repaired based on the regions corresponding to the multiple maximum region scores to obtain a residual map; The enhanced image is superimposed on the residual image to obtain an updated enhanced image, and the three-dimensional grid occupancy is constructed using the updated enhanced image.
5. The method according to claim 1, wherein The method of projecting the enhanced image onto a three-dimensional occupancy grid in a unified coordinate system using an improved voxel rendering mechanism and a spatial alignment strategy comprises: Predicting the depth value of each pixel from the enhanced image using a lightweight depth prediction network model to obtain a dense depth map corresponding to the enhanced image; Using the attitude matrix of the flying car in the camera, each pixel is mapped to a three-dimensional point; Using the current flight state of the flying car, the transformation parameters between the image and the real space are estimated to construct a perspective compensation matrix, and the perspective compensation matrix is used to perform rotation compensation on the three-dimensional point; The compensated 3D points are projected onto the 3D occupancy grid in a unified coordinate system.
6. The method according to any one of claims 1 to 5, characterized in that: The prediction model is trained as follows: Acquire a pseudo-label voxel set; the pseudo-label voxel set includes a plurality of first samples; the first samples are generated by mapping semantic information and depth information of a two-dimensional sample image to a three-dimensional voxel space to form the first samples; Acquire a plurality of second samples; the second samples are fine-label three-dimensional label data; A dual-path semi-supervised loss structure is adopted. Several second samples are used on the main path to calculate the fine supervision loss, and multiple first samples are used on the auxiliary path to calculate the pseudo supervision loss. A confidence gating mechanism is used to select voxel locations with confidence higher than a set value to participate in the training process, so as to complete the training of the prediction model using the total loss function. The total loss function includes at least fine supervision loss, pseudo supervision loss and consistency constraint loss; the consistency constraint loss is used to represent the consistency constraint of the prediction results in different perspectives and states for the same scene.
7. A beyond-visual-range flying car occupancy grid prediction device, characterized in that: The device comprises: A first acquisition unit is used to acquire a captured image of the flying car at a current moment and input the captured image into a pre-trained prediction model; A prediction unit is configured to process the captured image using the prediction model as follows: performing image enhancement on the captured image to obtain an enhanced image; utilizing an improved voxel rendering mechanism and a spatial alignment strategy to project the enhanced image onto a three-dimensional occupancy grid in a unified coordinate system based on the current flight state of the flying vehicle, thereby constructing the three-dimensional occupancy grid at the current moment; and utilizing the flight state of the flying vehicle at the current moment and at several historical moments, as well as the three-dimensional occupancy grid at the current moment, to predict environmental changes at future moments; The second acquisition unit is used to obtain the prediction result output by the prediction model.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of any one of the methods described in claims 1-6.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.