Unmanned aerial vehicle autonomous air refueling target segmentation and detection method simulating raptor vision

By drawing on the biological mechanism of Raptor Vision and combining deep learning technology, an autonomous drone aerial refueling target segmentation and detection method that imitates Raptor Vision has been established, which solves the accuracy of target segmentation and detection of drones in complex environments, and achieves efficient autonomous navigation and positioning.

CN120126136APending Publication Date: 2025-06-10BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510076806.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

During autonomous aerial refueling, drones face complex environmental and background interference, which affects the accuracy of target segmentation and detection, especially when targets and backgrounds are deeply integrated, light changes, motion blur and similarity interference.

Method used

Drawing on the biological mechanism of Raptor Vision, a deep learning mechanism that imitates Raptor Vision is established, combined with the spatial and temporal information integration mechanism and the cross-attention mechanism, is used for the segmentation and detection of autonomous aerial refueling targets of drones. Specifically, it includes establishing a basic SAM model framework, introducing a time-series fusion mask model (TFMM) that mimics the spatiotemporal information integration mechanism of Raptor, and a memory-first affinity model (MPAM) with cross-repression characteristics.

Benefits of technology

It improves the segmentation and detection capabilities of the drone under the target-background deep coupling, enhances the success rate and safety of autonomous navigation and positioning, and realizes accurate autonomous docking and refueling in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126136A_ABST
    Figure CN120126136A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle autonomous air refueling target segmentation and detection method simulating raptor vision. The method comprises the following steps: 1, establishing a basic SAM model framework; 2, establishing a time sequence fusion mask model (TFMM) of a spatio-temporal information integration mechanism imitating the brauer vision; 3, establishing a memory priority affinity model (MPAM) simulating a brauer visual cross suppression characteristic; and 4, training and testing an unmanned aerial vehicle camouflage object target detection method imitating raptor vision in a real scene. The method can adapt to various complex combat environments and extreme conditions of target-background deep coupling; according to the method, a cross attention mechanism is combined in a mask embedding space to collect space-time information, the cross attention mechanism is combined in the mask embedding space to collect the space-time information to enhance information flow, feature representation is optimized, and the robustness of a model is enhanced; the method shows strong segmentation detection performance under the condition of less parameter quantity, is suitable for being carried on an unmanned aerial vehicle, and has wide military application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is a method for segmenting and detecting the target of an unmanned aerial vehicle (UAV) for autonomous in-air refueling by imitating the vision of raptors, belonging to the field of computer vision. Background Art

[0002] With the rapid development of UAV technology, UAVs are being used more and more widely in military, civilian, and commercial fields. Especially in the military field, the autonomous in-air refueling ability of UAVs has become an important means to improve combat effectiveness. However, in actual operation, when UAVs perform autonomous in-air refueling, they face complex environmental and background interferences, which makes target segmentation and detection particularly difficult. During the in-air refueling process of UAVs, it is usually necessary to autonomously identify and track the tanker or the receiver. However, factors such as the complexity and variability of the background environment, changes in illumination, motion blur, and similarity interference reduce the resolution between the UAV and the background, affecting the accuracy of target segmentation and detection. The purpose of the present invention is to propose a method for segmenting and detecting the target of UAV autonomous in-air refueling that is simple in principle, efficient in combat, direct, and flexible, enabling the UAV to effectively segment and detect the target UAV in a complex environment and achieve precise docking and refueling.

[0003] In recent years, researchers have made remarkable progress in the field of computer vision. However, existing target detection methods still face many challenges when dealing with extreme situations where the main body and the background are difficult to distinguish. First, during the in-air refueling process of UAVs, the target usually blends in with the background (such as the sky, the ground, etc.), and its boundary is blurred and difficult to distinguish. This makes traditional target detection models perform poorly in identifying these objects because they usually rely on clear object boundaries and features. In addition, in a complete video sequence, target segmentation and detection not only need to identify the target in the current frame but also need to maintain temporal consistency. Existing methods (such as video object segmentation (VOS) and video motion segmentation (VMS)) have difficulties in capturing object motion and maintaining inter-frame consistency, especially when the object is moving, and the blurriness of the boundary makes motion estimation unreliable.

[0004] With the development of computer vision, a basic deep learning-based image segmentation model, "Segment Anything" (SAM), has gained wide recognition for its excellent performance. It aims to generate high-quality segmentation masks when dealing with image segmentation in various scenarios. Based on the SAM model, after performing image segmentation on the target in each frame of the video sequence and then conducting detection processing, it provides a feasible solution to the problem of target segmentation and detection when the UAV is deeply fused with the background and the main body is difficult to distinguish.

[0005] To address the challenges in visual tasks, more and more researchers are choosing to seek inspiration from nature. The biological visual systems of raptors, represented by them, are highly developed. The unique physiological structures and information processing mechanisms of their eyeballs and midbrains are the key determinants of the strong vision and intelligent perception of raptors. The broad vision and sharp eyesight of raptors have attracted much attention in research fields such as target detection, feature extraction, and target tracking. Mapping the relevant biological mechanisms of raptor vision into deep learning mechanisms and embedding them into the method for detecting hard-to-distinguish object targets can significantly improve the success rate of target segmentation and detection in the case of target-background depth coupling, thus achieving precise autonomous docking refueling.

[0006] In summary, the present invention proposes a method for target segmentation and detection of unmanned aerial vehicle (UAV) autonomous in-air refueling imitating raptor vision, which maps the biological mechanism of raptor vision into a deep learning mechanism and is used for the segmentation and detection algorithm of UAVs in the case of target-background depth coupling. The method is simple, efficient, has good real-time performance, conforms to the actual combat scenario, and has certain reference significance. Summary of the Invention

[0007] The object of the present invention is to provide a method for target segmentation and detection of UAV autonomous in-air refueling imitating raptor vision, aiming to improve the segmentation and detection ability of UAVs in the case of target-background depth coupling, enhance the success rate and safety of its autonomous navigation and positioning, and thus play an important role in applications such as military, surveillance, and environmental monitoring. By drawing on the biological mechanism of raptor vision in nature, an information processing mechanism based on raptor vision is established and embedded into the method for target segmentation and detection of UAV autonomous in-air refueling, providing a new solution idea for the problems of UAV autonomous navigation and precise segmentation and detection.

[0008] The present invention proposes a method for target segmentation and detection of UAV autonomous in-air refueling imitating raptor vision, and the specific implementation steps are as follows:

[0009] Step 1: Establish a basic SAM model framework

[0010] The basic SAM framework takes a video clip containing the target-background depth coupling situation as input and generates a series of pixel-level binary masks corresponding to each frame in the video. Suppose the video frame consists of T frames, denoted as

[0011] S11. Extract image features using an image encoder

[0012] The present invention uses a pre-trained masked autoencoder MAE and Vision Transformer for image feature extraction. The output of the image encoder SAM Enc is the image embedding of the original image, with a size 16 times smaller than the original image.

[0013] S12. Generate sparse embeddings using a prompt encoder

[0014] The prompt encoder SAM ProEnc Encode the position information of the input points and boxes to generate sparse embeddings, while the input mask prompts are encoded as dense embeddings by the mask encoder SAM. The sparse embeddings and dense embeddings together form the prompt embeddings. MaskEnc The sparse embeddings and dense embeddings together form the prompt embeddings.

[0015] S13. Generate an output mask using a mask decoder

[0016] The mask decoder SAM Dec Map the image embeddings and prompt embeddings to the output mask.

[0017] Step 2: Establish a temporal fusion mask model (TFMM) for the visual spatio-temporal information integration mechanism of a falcon-like model

[0018] To incorporate the cross-attention mechanism in the mask embedding space to collect spatio-temporal information, the present invention introduces a temporal fusion mask model (TFMM, Temporal Fusion Mask Module) for the visual spatio-temporal information integration mechanism of a falcon-like model.

[0019] S21. Model the information extraction mechanism of the high-resolution vision of a falcon-like model

[0020] Falcon vision has extremely high resolution and can clearly see details in the distance. By analogy with the characteristics of falcon vision, first, add the positional embedding PE TFMM to the image embeddings of the current frame , apply layer normalization LN and use the query head Q TFMM of TFMM to obtain the query vector Q of the TFMM module, which represents the attention of the falcon-like vision to the current frame; similarly, add the positional embedding to the connected image embeddings from memory , apply layer normalization LN, and use the key head K TFMM of TFMM to obtain the key vector K of the TFMM module; finally, apply layer normalization to the connected mask embeddings and use the value head V TFMM of TFMM to obtain the value vector V of the TFMM module. K and V represent the memory of the falcon-like vision for the information of previous frames, that is:

[0021]

[0022] where Q TFMM and Q represent the query head function and the query vector of the TFMM module respectively, K TFMM and K represent the key head function and the key vector of the TFMM module respectively, and V TFMMand V represent the value head function and the value vector of the TFMM module respectively, LN represents the layer normalization function, and PE TFMM represents the positional embedding of the TFMM module. represents the image embedding of the current frame, that is, the feature representation of the current frame image processed by the image encoder SAM Enc ; represents the connected image embedding from memory, that is, the set of image embeddings from the (i - n + 1)-th frame to the i-th frame in the time series; represents the connected mask embedding from memory, that is, the set of mask embeddings from the (i - n + 1)-th frame to the i-th frame in the time series.

[0023] S22. Modeling of the spatio-temporal information integration mechanism of the visual cross-attention mechanism imitating raptors

[0024] Raptors can quickly identify prey in complex environments and ignore irrelevant background information. This visual selection mechanism can be mapped to the cross-attention mechanism in deep learning, that is, by calculating the similarity between the query and the key, the image segmentation model can selectively focus on the information relevant to the current task and ignore the unimportant parts.

[0025] (1) Calculate the attention output value

[0026] After obtaining the query vector Q, key vector K, and value vector V of the TFMM module, calculate the attention output value O through the following formula TFMM , that is, the information weighted by the attention mechanism, representing the similarity between the query and the key:

[0027] O TFMM = Attention AFMM (Q, K, V)

[0028] where Attention AFMM is a function of the cross-attention mechanism that calculates the relationship between Q, K, and V, expressed as:

[0029]

[0030] where the activation function softmax converts the attention scores into a probability distribution, d k is the dimension of the key vector, and the superscript T is the matrix transpose.

[0031] (2) Attention output processing imitating the multi-cortex mechanism of the raptor visual midbrain

[0032] The visual system of raptors includes multiple cortical regions, which are responsible for processing different types of visual information. For example, the primary visual cortex (edges, orientations, motion) is responsible for extracting basic visual features, the secondary visual cortex (shape, color, depth) processes more complex visual information, and the higher-order visual cortex (object recognition, spatial localization) is responsible for more advanced visual cognitive tasks. The midbrain of raptors can integrate information from different cortical regions to form a complete visual image. This integration ability enables raptors to quickly identify prey, judge distance and speed, and make corresponding hunting decisions.

[0033] Similar to the multi-cortical mechanism of the raptor visual midbrain, the present invention employs a multi-layer perceptron to process the attention output information. A multi-layer perceptron (MLP) typically consists of multiple fully connected layers, and a non-linear activation function is applied between each layer to enhance the model's expressive power. The input to the multi-layer perceptron is the output value O of the attention. TFMM , and the output of the MLP can be expressed as:

[0034] MLP(O TFMM ) = σ{W L ·ReLU[W L-1 ·ReLU(...ReLU(W 1 ·O TFMM + b 1 ) + b 2 …)+b L-1 + b L}

[0035] Where L is the number of layers of the MLP; W i (i = 1,..., L) is the weight matrix of the i-th layer, which is used to transform the input vector into the feature representation of the next layer, and its dimension depends on the number of input and output features; b i (i = 1,..., L) is the bias vector of the i-th layer, which is used to adjust the output of each neuron and help the model better fit the data; the activation function ReLU = max(0, x) represents the rectified linear unit, which is used to introduce non-linearity so that the model can learn complex features; σ is the activation function of the output layer.

[0036] The attention output O TFMM_Attention after being processed by the multi-layer perceptron is expressed as:

[0037] O TFMM_Attention = MLP(O TFMM ) + O TFMM

[0038] Where MLP represents the multi-layer perceptron function, and O TFMM is the attention output value.

[0039] (3) Obtain the final output through layer normalization

[0040] After using the MLP to enhance the model's expressive power, layer normalization is applied to improve the model's stability and performance. Layer normalization is a regularization technique aimed at improving the training process of deep learning models. It standardizes the output of each layer so that the output of each neuron has the same mean and variance, thereby reducing internal covariate shift, which helps to accelerate training and improve the model's convergence.

[0041] The final masked embedding output Output TFMM is expressed as:

[0042] Output TFMM = LN(O TFMM_Attention )

[0043] where LN represents the layer normalization operation, ensuring the stability and consistency of the output; O TFMM_Attention and Output TFMM are the attention output value processed by the multi-layer perceptron and the final masked embedding output, respectively.

[0044] In summary, the TFMM model uses the image embedding of the current frame the connected image embedding from the memory and the connected masked embedding from the memory as inputs, performs spatio-temporal cross-attention, and outputs the masked embedding Output containing spatio-temporal information TFMM . The above process is denoted as:

[0045]

[0046]

[0047] where TFMM represents the processing process of the temporal fusion masked model simulating the raptor visual spatio-temporal information integration mechanism. Step 3: Establish a memory-prior affinity model (MPAM) with raptor-like visual cross-suppression characteristics

[0048] To combine the output of the temporal fusion masked module TFMM with the image embedding of the current frame to enhance the information flow, the present invention introduces a memory-prior affinity model (MPAM, Memory Prior Affinity Module) with raptor-like visual cross-suppression characteristics.

[0049] S31. Modeling the memory information transfer mechanism with raptor-like visual cross-suppression characteristics

[0050] Raptor vision has the ability to accurately obtain the current position of a target based on memory information from a highly complex hunting background, and this ability stems from the visual cross-inhibition characteristics of raptors. Similar to the visual systems of most birds, in the visual system of raptors, the regions for visual information transmission and processing have a well-defined hierarchical structure and connection relationships, and the coordinated work among various tissues enables the visual system of eagles to have the ability to quickly capture and respond to targets. The connection relationship between the retina and the brain is a key factor for raptors to complete various visual tasks and is also crucial for ensuring that they can distinguish targets from complex backgrounds.

[0051] The MPAM module aims to improve the generation process of mask embeddings. Traditional methods usually rely on the current image or stitch the current frame with previous masks, while the MPAM module extracts the output values from the TFMM, providing a generalized representation of the mask embedding. This method utilizes the existing SAM encoder, eliminating the necessity of training an additional encoder usually required when stitching images and masks in the original image space.

[0052] First, generate the query vector and key vector of the MPAM module:

[0053]

[0054] where q and k are the query vector and key vector of the MPAM module respectively, LN represents the layer normalization operation, and PE MPAM is the position embedding of the MPAM module, providing information about the position of the current frame; is the image embedding of the current frame, representing the encoded features of the current frame; is the concatenation of the image embeddings from the i-th frame to the i + n - 1-th frame, representing the features of the previous n frames as the output of the TFMM module; represents the concatenation operation, which concatenates the image embedding of the current frame with the output of the TFMM.

[0055] S32. Modeling the Retina-Brain Information Projection Mechanism Imitating Raptor Vision

[0056] In the visual physiological structure of raptors, the connection relationship and information projection mechanism between the retina and the brain are crucial. Among them, the optic tectum is an important visual center in the visual system, receiving a large number of optic nerve stimulations and projections. The optic tectum has an obvious hierarchical and partitioned structure. Each optic tectum structure can receive contralateral visual stimulus signals and can control the movement of the contralateral eye. The precise conjugate eye movement control of binocular vision can be mediated through the thalamic pathway and strongly projected back to the optic tectum. The inhibitory effect between raptor visual cells enables raptors to exhibit functional characteristics such as visual lateral inhibition, which is an important support for raptors to learn rich feature representations of targets and ignore irrelevant information.

[0057] By analogy with the visual retina-brain information projection mechanism of raptors, the present invention maps the query and key vectors to different spaces respectively, enabling the model to learn richer feature representations. This mapping allows the model to process information more flexibly in the attention mechanism. Specifically, linear transformations are applied to the query vector q and key vector k of the MPAM module to generate the key vector qk, mk and value vector qv, mv after linear transformation respectively, that is:

[0058] qk,qv = Linear1(q),Linear2(q)

[0059] mk,mv = Linear3(k),Linear4(k)

[0060] Among them, the linear transformations Linear1 and Linear2 are used to transform the query vector q of the MPAM module into the key vector qk and value vector qv after linear transformation, and the linear transformations Linear3 and Linear4 are used to transform the key vector k of the MPAM module into the key vector mk and value vector mv after linear transformation; the key vectors qk and mk after linear transformation help the model to match with other vectors when calculating the attention weights, and the value vectors qv and mv after linear transformation provide the final output information after calculating the attention.

[0061] Similar to TFMM, the output value V of the final attention is calculated through the cross-attention function MPAM , that is:

[0062]

[0063] Among them, qk is the key vector after the linear transformation of the query vector q of the MPAM module, mk is the key vector after the linear transformation of the key vector k of the MPAM module, and mv is the value vector after the linear transformation of the key vector k of the MPAM module, is the splicing operation, Attention MPAM is the cross-attention calculation function, expressed as:

[0064]

[0065] Among them, the activation function softmax converts the attention scores into a probability distribution, d mk is the dimension of the key vector mk.

[0066] After that, a dense embedding is generated through a multi-layer perceptron to generate the final feature representation, expressed as:

[0067]

[0068] MLP(V MPAM ) = σ'{WL' · ReLU[W L'-1 · ReLU(…ReLU(W 1 · V MPAM + b 1 ) + b 2 '…)+ b L-1 ) + b L}

[0069] where L' is the number of layers of the MLP; W i (i = 1, ..., L') is the weight matrix of the i-th layer,; b i (i = 1, ..., L') is the bias vector of the i-th layer; the activation function ReLU = max(0, x) represents the rectified linear unit; σ' is the activation function of the output layer.

[0070] Finally, pass the dense embedding and the image embedding of the current frame to the decoder SAM Dec to generate the predicted mask M i+1 .

[0071] In summary, MPAM splices the mask embedding Output TFMM containing spatio-temporal information output by TFMM with the image embedding of the current frame and the image embeddings from the i-th frame to the i + n - 1-th frame to perform an affinity operation, thereby enhancing the temporal information and generating the dense embedding of the current frame The above process is denoted as:

[0072]

[0073] where MPAM represents the processing process of the memory-prior affinity model that emulates the visual cross-suppression characteristics of raptors.

[0074] Generally speaking, the visualization structures of the temporal fusion mask model (TFMM) that emulates the visual spatio-temporal information integration mechanism of raptors and the memory-prior affinity model (MPAM) that emulates the visual cross-suppression characteristics of raptors proposed in the present invention are as Figure 1 shown. The flowchart of the method for autonomous aerial refueling target segmentation and detection of drones that emulates the vision of raptors proposed in the present invention is as shown in the appendix Figure 2 and is described as follows:

[0075] (1) For the given j-frame images I j and j-frame masks M j with highly coupled target-background, calculate the image embedding and the mask embedding as follows:

[0076]

[0077] Among them, SAM Enc and SAM MaskEnc are an image encoder and a mask encoder, respectively.

[0078] (2) Given the (i + 1)-th frame image I i+1 , in order to predict the mask M i+1 of the (i + 1)-th frame, the present invention first embeds the image of the current frame as well as the image embeddings of the previous n frames and the corresponding mask embeddings and passes them to the "Temporal Fusion Mask Model (TFMM) of Raptor-like Visual Spatiotemporal Information Integration Mechanism". Then, the output Output TFMM of the TFMM model, as well as the image embedding of the current frame and the image embeddings from the i-th frame to the (i + n - 1)-th frame are passed to the Memory Priority Affinity Model (MPAM) with Raptor-like Visual Cross-Suppression Characteristics to generate the dense embedding of the current frame That is

[0079]

[0080] Among them, TFMM and MPAM respectively represent the processing processes of the TFMM model and the MPAM model. After that, the dense embedding of the current frame and the image embedding are passed to the decoder to obtain the current mask prediction M i+1 That is

[0081]

[0082] Among them, SAM Dec is the mask decoder.

[0083] Process (2) will be repeated throughout the video sequence. During this process, the memory will be updated by adding new images and dense embeddings, and at the same time, the outdated embeddings will be removed when the memory limit is reached.

[0084] Step Four: Train the Raptor-like Visual UAV Autonomous In-air Refueling Target Segmentation and Detection Method, and output the final loss function

[0085] S41. Configure the virtual environment of the Raptor-like Visual UAV Autonomous In-air Refueling Target Segmentation and Detection Method

[0086] The present invention configures the python environment using the virtual environment management software Anaconda3. A new environment is created in the Conda Prompt command-line window, and the software packages such as pytorch, torchvision, torchaudio, cuda, opencv, pycocotools, matplotlib, onnxruntime, onnx, xformers, lightning, scikit-image, gdown, wandb, and omegaconf are installed.

[0087] S42. Deploy the training set of the method for target segmentation and detection of autonomous in-air refueling of drones with raptor-like vision

[0088] This method uses MoCA-Mask-Train, COD10K-v3, and Camouflaged-Animal-Dataset as the training set, which can be downloaded by running the command line bash dataset_download.sh.

[0089] S43. Load the pre-trained model and conduct two-stage training

[0090] In the program, the pre-trained model is loaded for two-stage training. In the first stage, the number of training epochs is set to 10, and the learning rate is 0.0004; in the second stage, the number of training epochs is set to 140, and the learning rate is 0.0005. That is, the total number of training epochs is 150.

[0091] S44. Output the classification loss, mask generation loss, bounding box regression loss, and spatio-temporal consistency loss

[0092] The present invention designs a total of four loss functions to ensure that the model can correctly classify objects, accurately generate masks, accurately regress bounding boxes, and maintain spatio-temporal consistency.

[0093] (1) The classification loss is designed as

[0094]

[0095] where N is the number of samples, y i is the true label, represents the probability value predicted by the model, indicating the probability that the sample belongs to an object.

[0096] (2) The bounding box regression loss is designed as

[0097]

[0098] The SmoothL1 loss is defined as

[0099]

[0100] Among them, N is the number of samples, and t i is the coordinate of the true bounding box, and t i is the coordinate of the bounding box predicted by the model.

[0101] (3) The spatio-temporal consistency loss is defined as

[0102]

[0103] Among them, T is the number of time frames, N is the number of samples, and f t (i) is the feature representation of the i-th sample in the time frame t, and f t-1 (i) is the feature representation of the i-th sample in the time frame t-1. ||·|| represents the L2 norm, which is used to calculate the distance between features.

[0104] (4) The mask generation loss is defined as

[0105]

[0106] Among them, N is the number of samples, and m i is the true label, represents the probability value predicted by the model, indicating the probability that the sample belongs to an object.

[0107] (5) The final comprehensive loss function is defined as

[0108] L = λ cls L cls + λ bbox L bbox + λ temporal L temporal + λ mask L mask

[0109] Among them, λ cls , λ bbox , λ temporal and λ mask are the weight coefficients of the classification loss, bounding box regression loss, spatio-temporal consistency loss, and mask generation loss, respectively.

[0110] Step Five: Test the UAV autonomous in-air refueling target segmentation and detection method for raptor-like vision under the UAV23 dataset and the self-made dataset

[0111] The present invention uses the UAV video sequences in the UAV123 dataset and the self-made dataset as the test sets for the proposed target segmentation and detection method. The two datasets contain video frames with a high coupling of the target UAV and the background, which is highly similar to the real scenario of autonomous in-air refueling of UAVs. Therefore, the present invention conducts tests on these two datasets and selects some pictures for visualization. The visualization results of the method for segmenting and detecting the target of autonomous in-air refueling of UAVs imitating raptor vision in real scenarios are as shown in Figure 3 , Figure 4 and Figure 5 . Among them, Figure 3 is the original image of the test dataset, Figure 4 shows the mask generation process, Figure 5 shows the target segmentation and detection results. It can be seen that the method for segmenting and detecting the target of autonomous in-air refueling of UAVs imitating raptor vision proposed by the present invention has a high success rate of segmentation and detection for the case of deep coupling between the target UAV and the background; in addition, this method shows strong segmentation and detection performance with fewer parameters, is suitable for being carried on UAVs, and has broad military application potential. Description of the Drawings

[0112] Figure 1 is the visualization structure of the temporal fusion mask model (TFMM) of the raptor vision-inspired spatio-temporal information integration mechanism and the memory-priority affinity model (MPAM) of the raptor vision-inspired cross-inhibition characteristics.

[0113] Figure 2 is the flow chart of the method for segmenting and detecting the target of autonomous in-air refueling of UAVs imitating raptor vision.

[0114] Figure 3 is the original image of the test dataset.

[0115] Figure 4 is the visualization result of the method for segmenting and detecting the target of autonomous in-air refueling of UAVs imitating raptor vision in real scenarios - mask generation.

[0116] Figure 5 is the visualization result of the method for segmenting and detecting the target of autonomous in-air refueling of UAVs imitating raptor vision in real scenarios - target segmentation and detection. Detailed Embodiment

[0117] Step 1: Establish the basic SAM model framework

[0118] The basic SAM framework takes video segments containing the case of deep coupling between the target and the background as input and generates a series of pixel-level binary masks corresponding to each frame in the video. Suppose the video frames consist of T frames, denoted as

[0119] S11. Extract image features using an image encoder

[0120] The present invention uses a pre-trained Masked Autoencoder (MAE) and Vision Transformer for image feature extraction. The image encoder SAM Enc outputs an image embedding of the original image, which is 16 times smaller in size than the original image.

[0121] S12. Generate a sparse embedding using a prompt encoder

[0122] The prompt encoder SAM ProEnc encodes the position information of the input point / box to generate a sparse embedding, while the input mask prompt is encoded into a dense embedding by the mask encoder SAM MaskEnc The sparse embedding and the dense embedding together form the prompt embedding.

[0123] S13. Generate an output mask using a mask decoder

[0124] The mask decoder SAM Dec maps the image embedding and the prompt embedding to the output mask.

[0125] Step 2: Establish a Temporal Fusion Mask Model (TFMM) that emulates the raptor's visual spatio-temporal information integration mechanism

[0126] S21. Model the information extraction mechanism of the raptor's visual high-resolution

[0127] By analogy with the characteristics of raptor vision, first, add the positional embedding PE TFMM to the image embedding of the current frame , apply layer normalization LN, and use the query head Q TFMM of TFMM to obtain the query vector Q of the TFMM module; similarly, add the positional embedding to the concatenated image embedding from memory , apply layer normalization LN, and use the key head K TFMM of TFMM to obtain the key vector K of the TFMM module; finally, apply layer normalization to the concatenated mask embedding and use the value head V TFMM of TFMM to obtain the value vector V of the TFMM module, that is:

[0128]

[0129] where Q TFMM and Q respectively represent the query head function and the query vector of the TFMM module, K TFMM and K respectively represent the key head function and the key vector of the TFMM module, V TFMM and V respectively represent the value head function and the value vector of the TFMM module, LN represents the layer normalization function, and PETFMM Represents the positional embedding of the TFMM module. Represents the image embedding of the current frame, i.e., the feature representation of the current frame image processed by the image encoder SAM Enc ; Represents the concatenated image embeddings from memory, i.e., the set of image embeddings from the (i - n + 1)-th frame to the i-th frame in the time series; Represents the concatenated mask embeddings from memory, i.e., the set of mask embeddings from the (i - n + 1)-th frame to the i-th frame in the time series.

[0130] S22. Modeling of the spatio-temporal information integration mechanism imitating the raptor visual cross-attention mechanism

[0131] (1) Calculate the attention output value

[0132] After obtaining the query vector Q, key vector K, and value vector V of the TFMM module, calculate the attention output value O through the following formula TFMM , which is the information weighted by the attention mechanism and represents the similarity between the query and the key.

[0133] O TFMM = Attention AFMM (Q, K, V)

[0134] where Attention AFMM is a function of the cross-attention mechanism that calculates the relationship between Q, K, and V, expressed as

[0135]

[0136] where the activation function softmax converts the attention scores into a probability distribution, and d k is the dimension of the key vector.

[0137] (2) Attention output processing imitating the raptor visual midbrain multi-cortex mechanism

[0138] The input of the multi-layer perceptron is the attention output value O TFMM , and the output of the MLP can be expressed as

[0139] MLP(O TFMM ) = σ{W L ·ReLU[W L-1 ·ReLU(…ReLU(W 1 ·O TFMM + b 1 ) + b 2 …)+ b L-1 + b L}

[0140] Among them, L is the number of layers of the MLP; W i (i = 1, ..., L) is the weight matrix of the i-th layer, which is used to convert the input vector into the feature representation of the next layer, and its dimension depends on the number of input and output features; b i (i = 1, ..., L) is the bias vector of the i-th layer, which is used to adjust the output of each neuron to help the model better fit the data; the activation function ReLU = max(0, x) represents the rectified linear unit, which is used to introduce non-linearity so that the model can learn complex features; σ is the activation function of the output layer.

[0141] The attention output O after being processed by the multi-layer perceptron TFMM_Attention is expressed as

[0142] O TFMM_Attention = MLP(O TFMM ) + O TFMM

[0143] Among them, MLP represents the multi-layer perceptron function, and O TFMM is the attention output value.

[0144] (3) Obtain the final output through layer normalization

[0145] After using the MLP to enhance the model's expressive ability, layer normalization is applied to improve the stability and performance of the model. Layer normalization is a regularization technique aimed at improving the training process of deep learning models. It standardizes the output of each layer so that the output of each neuron has the same mean and variance, thereby reducing internal covariate shift and helping to accelerate training and improve the convergence of the model.

[0146] The final mask embedding output Output TFMM is expressed as

[0147] Output TFMM = LN(O TFMM_Attention )

[0148] Among them, LN represents the layer normalization operation to ensure the stability and consistency of the output; O TFMM_Attention and Output TFMM are respectively the attention output value after being processed by the multi-layer perceptron and the final mask embedding output.

[0149] In summary, the TFMM model uses the image embedding of the current frame the connected image embedding from the memory and the connected mask embedding from the memory as inputs, performs spatio-temporal cross-attention, and outputs the mask embedding Output TFMM containing spatio-temporal information. The above process is denoted as

[0150]

[0151] Among them, TFMM represents the processing process of the temporal fusion mask model of the raptor-like visual spatio-temporal information integration mechanism.

[0152] Step 3: Establish a memory-priority affinity model (MPAM) for raptor-like visual cross-suppression characteristics

[0153] S31. Modeling the memory information transfer mechanism of raptor-like visual cross-suppression characteristics

[0154] The MPAM module aims to improve the generation process of mask embedding. Traditional methods usually rely on the current image or stitch the current frame with the previous mask, while the MPAM module extracts the output value from TFMM and provides a generalized representation of mask embedding. This method uses the existing SAM encoder, eliminating the need to train an additional encoder usually required when stitching images and masks in the original image space.

[0155] First, generate the query vector and key vector of the MPAM module

[0156]

[0157] Among them, q and k are the query vector and key vector of the MPAM module respectively, LN represents the layer normalization operation, and PE MPAM is the position embedding of the MPAM module, providing information about the position of the current frame; is the image embedding of the current frame, representing the encoded features of the current frame; is the concatenation of the image embeddings from the i-th frame to the i + n - 1-th frame, indicating that the features of the previous n frames are the output of the TFMM module; represents the concatenation operation, which concatenates the image embedding of the current frame with the output of TFMM.

[0158] S32. Modeling the retina-brain information projection mechanism of raptor-like vision

[0159] Analogous to the raptor-like visual retina-brain information projection mechanism, the present invention maps the query and key vectors to different spaces respectively, enabling the model to learn richer feature representations. This mapping enables the model to process information more flexibly in the attention mechanism. Specifically, linear transformations are used on the query vector q and key vector k of the MPAM module to generate the key vectors qk, mk and value vectors qv, mv after linear transformation respectively, that is:

[0160] qk, qv = Linear1(q), Linear2(q)

[0161] mk, mv = Linear3(k), Linear4(k)

[0162] Among them, the linear transformations Linear1 and Linear2 are used to transform the query vector q of the MPAM module into the linearly transformed key vector qk and value vector qv, and the linear transformations Linear3 and Linear4 are used to transform the key vector k of the MPAM module into the linearly transformed key vector mk and value vector mv; the linearly transformed key vectors qk and mk help the model match with other vectors when calculating the attention weights, and the linearly transformed value vectors qv and mv provide the final output information after calculating the attention.

[0163] Similar to TFMM, the final attention output value V is calculated through the cross-attention function MPAM , that is

[0164]

[0165] Among them, qk is the key vector linearly transformed from the query vector q of the MPAM module, mk is the key vector linearly transformed from the key vector k of the MPAM module, and mv is the value vector linearly transformed from the key vector k of the MPAM module. is the concatenation operation, Attention MPAM is the cross-attention calculation function, expressed as

[0166]

[0167] Among them, the activation function softmax converts the attention scores into a probability distribution, d mk is the dimension of the key vector mk.

[0168] After that, a dense embedding is generated through a multi-layer perceptron to generate the final feature representation, expressed as

[0169]

[0170] MLP(V MPAM ) = σ'{W L' ·ReLU[W L'-1 ·ReLU(…ReLU(W 1 ·V MPAM +b 1 ')+b 2 '…)+b L-1 ']+b L '}

[0171] Among them, L' is the number of layers of the MLP; W i (i = 1,..., L') is the weight matrix of the i-th layer,; bi The bias vector for the i-th layer is (i = 1,..., L'); the activation function ReLU = max(0, x) represents the rectified linear unit; and σ' is the activation function of the output layer.

[0172] Finally, the dense embedding and the image embedding of the current frame are passed to the decoder SAM Dec to generate the predicted mask M i+1 .

[0173] In summary, MPAM splices the mask embedding Output TFMM containing spatio-temporal information output by TFMM with the image embedding of the current frame and the image embeddings from the i-th frame to the (i + n - 1)-th frame to perform an affinity operation, thereby enhancing the temporal information and generating the dense embedding of the current frame The above process is denoted as:

[0174]

[0175] where MPAM represents the processing procedure of the memory-prioritized affinity model that emulates the visual cross-suppression characteristics of raptors.

[0176] Generally speaking, the temporal fusion mask model (TFMM) of the raptor-inspired visual spatio-temporal information integration mechanism and the memory-prioritized affinity model (MPAM) of the raptor-inspired visual cross-suppression characteristics proposed in the present invention are described as follows:

[0177] (1) For the given j-frame images I j and the j-frame mask M j , the image embedding and the mask embedding are calculated as follows:

[0178]

[0179] where SAM Enc and SAM MaskEnc are the image encoder and the mask encoder, respectively.

[0180] (2) Given the (i + 1)-th frame image I i+1 , in order to predict the mask M i+1 of the (i + 1)-th frame, the present invention first passes the image embedding of the current frame and the image embeddings of the previous n frames and the corresponding mask embeddings to the "temporal fusion mask model (TFMM) of the raptor-inspired visual spatio-temporal information integration mechanism", and then passes the output Output of the TFMM modelTFMM and the image embedding of the current frame and the image embeddings from the i-th frame to the (i + n - 1)-th frame are passed to the Memory-Prioritized Affinity Model (MPAM) with Raptor-like visual cross-suppression characteristics to generate the dense embedding of the current frame That is

[0181]

[0182] where TFMM and MPAM respectively represent the processing procedures of the TFMM model and the MPAM model. After that, the dense embedding of the current frame and the image embedding are passed to the decoder to obtain the current mask prediction M i+1 , that is

[0183]

[0184] where SAM Dec is the mask decoder.

[0185] Process (2) is repeated throughout the video sequence. During this process, the memory is updated by adding new images and dense embeddings, and at the same time, outdated embeddings are removed when the memory limit is reached.

[0186] Step Four: Training the UAV autonomous in-air refueling target segmentation and detection method with Raptor-like vision

[0187] S41. Configure the virtual environment for the UAV autonomous in-air refueling target segmentation and detection method with Raptor-like vision

[0188] The present invention uses the virtual environment management software Anaconda3 to configure the python environment. A new environment is created in the Conda Prompt command line window, and software packages such as pytorch, torchvision, torchaudio, cuda, opencv, pycocotools, matplotlib, onnxruntime, onnx, xformers, lightning, scikit-image, gdown, wandb, and omegaconf are installed.

[0189] S42. Deploy the training set for the UAV autonomous in-air refueling target segmentation and detection method with Raptor-like vision

[0190] This method uses MoCA-Mask-Train, COD10K-v3, and Camouflaged-Animal-Dataset as the training set, which can be downloaded by running the command line bash dataset_download.sh.

[0191] S43. Load the pre-trained model and perform two-stage training

[0192] Load the pre-trained model in the program and perform two-stage training. In the first stage, set the number of training epochs to 10 and the learning rate to 0.0004; in the second stage, set the number of training epochs to 140 and the learning rate to 0.0005. That is, the total number of training epochs is 150.

[0193] S44. Output the classification loss, mask generation loss, bounding box regression loss, and spatio-temporal consistency loss

[0194] The present invention designs a total of four loss functions to ensure that the model can correctly classify objects, accurately generate masks, accurately regress bounding boxes, and maintain spatio-temporal consistency.

[0195] (1) The classification loss is designed as

[0196]

[0197] where N is the number of samples, y i is the true label, represents the probability value predicted by the model, indicating the probability that the sample belongs to an object.

[0198] (2) The bounding box regression loss is designed as

[0199]

[0200] The SmoothL1 loss is defined as

[0201]

[0202] where N is the number of samples, t i is the true bounding box coordinate, and t i is the bounding box coordinate predicted by the model.

[0203] (3) The spatio-temporal consistency loss is defined as

[0204]

[0205] where T is the number of time frames, N is the number of samples, and f t (i) is the feature representation of the i-th sample in the time frame t, and f t-1(i) is the feature representation of the i-th sample in time frame t-1, and ||·|| represents the L2 norm, which is used to calculate the distance between features.

[0206] (4) The mask generation loss is defined as

[0207]

[0208] where N is the number of samples, m i is the true label, represents the probability value predicted by the model, indicating the probability that the sample belongs to the object.

[0209] (5) The final comprehensive loss function is defined as

[0210] L = λ cls L cls + λ bbox L bbox + λ temporal L temporal + λ mask L mask

[0211] where λ cls 、λ bbox 、λ temporal and λ mask are the weight coefficients of the classification loss, bounding box regression loss, spatio-temporal consistency loss, and mask generation loss, respectively.

[0212] Step Five: Test the UAV autonomous in-air refueling target segmentation and detection method for raptor-like vision under the UAV23 dataset and the self-made dataset

[0213] The present invention uses the UAV video sequences in the UAV123 dataset and the self-made dataset as the test sets for the proposed target segmentation and detection method. The two datasets contain video frames with a high coupling of the target UAV and the background, which is highly similar to the real UAV autonomous in-air refueling scenario. Therefore, the present invention conducts tests on these two datasets and selects some pictures for visualization display.

Claims

1. A method for segmenting and detecting targets for autonomous aerial refueling of unmanned aerial vehicles (UAVs) based on raptor vision, characterized by: The method comprises the following steps: Step 1: Establish the basic SAM model framework; The basic SAM framework takes a video clip containing the object-background depth coupling as input and generates a series of pixel-level binary masks corresponding to each frame in the video; first, an image encoder is used to extract image features, and its output is an image embedding of the original image, which is 16 times smaller than the original image; then a hint encoder is used to generate a sparse embedding, and at the same time, the input mask hint is encoded into a dense embedding through the mask encoder, and the sparse embedding and the dense embedding together constitute the hint embedding; finally, a mask decoder is used to generate an output mask, that is, the image embedding and the hint embedding are mapped to the output mask; Step 2: Establish a temporal fusion mask model TFMM that simulates the raptor's visual spatiotemporal information integration mechanism; First, the position embedding is added to the image embedding of the current frame, layer normalization is applied, and the query header of TFMM is used to obtain the query vector of the TFMM module; similarly, the position embedding is added to the connected image embedding from the memory, layer normalization is applied, and the key header of TFMM is used to obtain the key vector of the TFMM module; secondly, layer normalization is applied to the connected mask embedding, and the value header of TFMM is used to obtain the value vector of the TFMM module; after obtaining the query vector, key vector and value vector of the TFMM module, the cross-attention mechanism function is used to calculate the information weighted by the attention mechanism to represent the similarity between the query and the key, and the attention output is processed by the multi-cortical mechanism of the raptor-like visual midbrain to enhance the model's expressiveness; finally, the final output is obtained by layer normalization; Step 3: Establish the memory priority affinity model MPAM that imitates the visual cross-inhibition characteristics of birds of prey; First, the query vector and key vector of the MPAM module are generated using the output of TFMM. Then, linear transformation is applied to the query vector and key vector of the MPAM module to generate the key vector and value vector after linear transformation respectively. Similar to TFMM, the output value of the final attention is calculated through the cross attention function. Then, similar to TFMM, the output value of the final attention is calculated through the cross attention function. Finally, the dense embedding and the image embedding of the current frame are passed to the decoder to generate the predicted mask; Step 4: Train the raptor-like vision-based target segmentation and detection method for autonomous aerial refueling of UAVs.

2. The method according to claim 1, characterized in that: The method further comprises: Step 5: Test the UAV autonomous aerial refueling target segmentation and detection method based on raptor-like vision on the UAV23 dataset and the self-made dataset.

3. The method according to claim 1 or 2, characterized in that: The specific process of step 2 is as follows: S21. Modeling of high-resolution information extraction mechanism based on raptor vision By analogy with the characteristics of raptor vision, first, the position is embedded in PE TFMM Image embed added to the current frame In the above example, we apply layer normalization LN and use the query head Q of TFMM. TFMM Get the query vector Q of the TFMM module, which is the attention of the raptor vision on the current frame; similarly, add the position embedding to the connected image embedding from the memory In the above example, we apply layer normalization LN and use the key K of TFMM TFMM Get the key vector K of the TFMM module; finally, apply layer normalization to the concatenated mask embedding and use the value header V of the TFMM TFMM Get the value vector V of the TFMM module; K and V represent the raptor's visual memory of the previous frame information, that is: Among them, Q TFMM and Q represent the query header function and the query vector of the TFMM module respectively, K TFMM and K represent the key header function and the key vector of the TFMM module respectively, V TFMM and V represent the value header function and the value vector of the TFMM module, LN represents the layer normalization function, PE TFMM Represents the position embedding of the TFMM module; Represents the image embedding of the current frame, that is, through the image encoder SAM Enc Feature representation of the processed current frame image; represents the connected image embeddings from memory, i.e., the set of image embeddings from the i-n+1th frame to the i-th frame in the time series; represents the connected mask embedding from memory, i.e., the set of mask embeddings from frame i-n+1 to frame i in the time series; S22. Modeling the spatiotemporal information integration mechanism of raptor-like visual cross-attention mechanism After obtaining the query vector Q, key vector K and value vector V of the TFMM module, the output value O of the attention is calculated by the following formula TFMM , which is the information weighted by the attention mechanism, representing the similarity between the query and the key; O TFMM =Attention AFMM (Q,K,V) Among them, Attention AFMM is the function of the cross-attention mechanism for calculating the relationship between Q, K and V, expressed as: Among them, the activation function softmax converts the attention score into a probability distribution, d k is the dimension of the key vector, and the superscript T is the matrix transpose; Then, the attention output is processed by a multi-layer perceptron MLP that imitates the multi-cortical mechanism of the raptor visual midbrain. The output of MLP is expressed as: MLP(O TFMM )=σ{W L ·ReLU[W L-1 ·ReLU(…ReLU(W1·O TFMM +b1)+b2…)+b L-1 ]+b L } Where L is the number of layers of MLP; W i (i=1,...,L) is the weight matrix of the i-th layer, which is used to convert the input vector into the feature representation of the next layer. Its dimension depends on the number of input and output features; b i (i=1,...,L) is the bias vector of the i-th layer, which is used to adjust the output of each neuron to help the model fit the data better; the activation function ReLU=max(0,x) represents the rectified linear unit, which is used to introduce nonlinearity so that the model can learn complex features; σ is the activation function of the output layer; Attention output O after multi-layer perceptron processing TFMM_Attention It is expressed as; EITHER TFMM_Attention =MLP(O TFMM )+O TFMM Among them, MLP represents the multi-layer perceptron function, O TFMM Output value for attention; Finally, the final output is obtained by layer normalization, and the final mask embedding output Output TFMM Expressed as Output TFMM =LN(O TFMM_Attention ) Among them, LN represents the layer normalization operation to ensure the stability and consistency of the output; TFMM_Attention and Output TFMM They are the attention output value after processing by the multi-layer perceptron and the final mask embedding output respectively.

4. The method according to claim 1 or 2, characterized in that: The specific process of step three is as follows: S31. Establishment of a memory priority affinity model MPAM that mimics the visual cross-inhibition characteristics of birds of prey First, generate the query vector and key vector of the MPAM module: Among them, q and k are the query vector and key vector of the MPAM module, LN represents the layer normalization operation, and PE MPAM Provides information about the current frame position for the MPAM module position embedding; The image embedding of the current frame represents the encoded features of the current frame; is the connection of the image embedding from the i-th frame to the i+n-1-th frame, indicating that the features of the previous n frames are the output of the TFMM module; represents the splicing operation, which splices the image embedding of the current frame with the output of TFMM; S32. Modeling of retina-brain information projection mechanism of raptor-like vision The query vector q and key vector k of the MPAM module are linearly transformed to generate the key vectors qk, mk and value vectors qv, mv after linear transformation, namely: qk,qv=Linear1(q),Linear2(q) mk,mv=Linear3(k),Linear4(k) Among them, linear transformations Linear1 and Linear2 are used to transform the query vector q of the MPAM module into the key vector qk and value vector qv after linear transformation, and linear transformations Linear3 and Linear4 are used to transform the key vector k of the MPAM module into the key vector mk and value vector mv after linear transformation; the key vectors qk and mk after linear transformation help the model match other vectors when calculating the attention weights, and the value vectors qv and mv after linear transformation provide the final output information after calculating the attention; Similar to TFMM, the final attention output value V is calculated by the cross attention function MPAM ,Right now: Among them, qk is the key vector after linear transformation of the query vector q of the MPAM module, mk is the key vector after linear transformation of the key vector k of the MPAM module, and mv is the value vector after linear transformation of the key vector k of the MPAM module. For the splicing operation, Attention MPAM is the cross attention calculation function, expressed as: Among them, the activation function softmax converts the attention score into a probability distribution, d mk is the dimension of the key vector mk; Afterwards, a dense embedding is generated through a multi-layer perceptron Generate the final feature representation, expressed as: MLP(V MPAM )=σ'{W L' ReLU[W L'-1 ·ReLU(…ReLU(W1·V MPAM +b1')+b2'…)+b L-1 ']+b L '}Where L' is the number of layers of MLP; W i (i=1,...,L') is the weight matrix of the i-th layer; b i (i=1,...,L') is the bias vector of the i-th layer; the activation function ReLU=max(0,x) represents the rectified linear unit; σ' is the activation function of the output layer; Finally, the dense embedding and the image embedding of the current frame Passed to decoder SAM Dec , to generate the predicted mask M i+1 .

5. The method according to claim 1 or 2, characterized in that: The specific process of step 4 is as follows: S41. Virtual environment for target segmentation and detection of autonomous aerial refueling of UAVs using raptor-like vision S42, Training set of target segmentation and detection methods for autonomous aerial refueling of UAVs using raptor-like vision S43. Load the pre-trained model and perform two-stage training S44, output classification loss, mask generation loss, bounding box regression loss, and spatiotemporal consistency loss.

6. The method according to claim 5, characterized in that: The classification loss is designed as: Where N is the number of samples, y i is the true label, Represents the probability value predicted by the model, indicating the probability that the sample belongs to the object; The bounding box regression loss is designed as: SmoothL1 loss is defined as: Where N is the number of samples, t i is the real bounding box coordinate, t i The bounding box coordinates predicted by the model; The spatiotemporal consistency loss is defined as: Where T is the number of time frames, N is the number of samples, and f t (i) is the feature representation of the i-th sample in time frame t, f t-1 (i) is the feature representation of the i-th sample in time frame t-1, ||·|| represents the L2 norm, which is used to calculate the distance between features; The mask generation loss is defined as: Where N is the number of samples, m is i is the true label, Represents the probability value predicted by the model, indicating the probability that the sample belongs to the object; The final comprehensive loss function is defined as: L=λ cls L cls +λ bbox L bbox +λ temporal L temporal +λ mask L mask Among them, λ cls , bbox , temporal and λ mask They are the weight coefficients for classification loss, bounding box regression loss, spatiotemporal consistency loss, and mask generation loss, respectively.

Citation Information

Patent Citations

  • Arbitrary angle target detection method based on coarse mask smooth label supervision

    CN115393710A

  • Lightweight method of image segmentation model SAM

    CN118334344A