A space target detection method based on a Sim-YOLOv5 model

CN117975288BActive Publication Date: 2026-09-29NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410093858.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2026-09-29
Estimated Expiration
2044-01-23

AI Technical Summary

Technical Problem

一阶段目标检测算法通过直接提取图像特征来预测目标的类别以及在图像中的位置,优势是检测速度非常快,适合做实时检测任务,但是效果可能较差;二阶段目标检测算法是第一步先进行候选框区域的生成,然后第二步对候选区域中的目标进行分类以及位置的回归,效果通常不错,但是检测速度慢

Benefits of technology

[0021](1)针对公共可用空间数据集非常有限的问题,本发明使用Blender构建模拟空间目标数据集,用于模型的训练、验证和测试;另外,生成模拟空间卫星图像时随机采用地球模型和星空背景,并且使用了运动模糊、相机畸变和色散等数据增强技术,有效地提高了训练后模型的泛化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117975288B_ABST
    Figure CN117975288B_ABST
Patent Text Reader

Abstract

The application discloses a space target detection method based on a Sim-YOLOv5 model, and an embedded platform carried on a micro-nano satellite has limitations in computing power, storage and power consumption, therefore, a target detection algorithm should be as light as possible to adapt to the micro-nano satellite platform. The application constructs a simulated space target data set based on a satellite model, and proposes a Sim-YOLOv5 network model, wherein a RepGhost module, a SimSPPF module and a GSConv module are introduced, the YOLOv5 model is improved in a certain lightness, the parameter quantity and the calculation quantity of the model are greatly reduced, meanwhile, the network model can guarantee a certain detection precision and speed when performing target detection, and the demand of the micro-nano satellite embedded platform application is met. When the application performs target detection on the embedded platform, TensorRT inference acceleration is also used, and the target detection speed can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision technology, specifically relating to a spatial target detection method based on the Sim-YOLOv5 model. Background Technology

[0002] In recent years, microsatellites and nanosatellites have developed rapidly in the field of space-based on-orbit servicing due to their low R&D costs, short development cycles, and flexible launch methods, gradually becoming a research hotspot in the aerospace field worldwide. In spacecraft on-orbit servicing technology, the detection of space targets is a crucial link. The prerequisite for satellite operations on space targets is the accurate identification of the servicing target, a function performed by the target detection system onboard the satellite. Due to the limitations of computing power, storage, and power consumption of microsatellite embedded platforms, it is necessary to design lightweight target detection algorithms to meet these hardware constraints.

[0003] Currently, most mainstream object detection algorithms are based on deep learning, possessing good learning capabilities and detection accuracy. They are divided into two types: one-stage algorithms and two-stage algorithms. One-stage object detection algorithms directly extract image features to predict the object's category and location within the image. Their advantage is very fast detection speed, making them suitable for real-time detection tasks, but their performance may be poor. Two-stage object detection algorithms first generate candidate bounding boxes, and then classify and regress the objects within those boxes. Their performance is usually good, but their detection speed is slow.

[0004] Deploying deep learning-based object detection algorithms on embedded platforms requires consideration of limitations in computing resources, storage space, and power consumption, while also ensuring the performance of the algorithms in practical applications, such as detection accuracy and real-time performance. Summarizing current research, one-stage object detection algorithms are primarily used on embedded platforms, with the YOLO series being the most popular. However, for implementing space object detection on micro-nano satellite embedded platforms, certain lightweight improvements are needed while maintaining the algorithm's object detection performance. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention proposes a spatial target detection method based on the Sim-YOLOv5 model, which can realize spatial target detection on resource-constrained embedded platforms, while ensuring a certain level of detection accuracy and speed while making the algorithm model lightweight.

[0006] The technical solution to achieve the purpose of this invention is: a spatial target detection method based on the Sim-YOLOv5 model, comprising the following steps:

[0007] Step 1: Using 3D graphics software and a 3D satellite model, create several simulated images of space targets against an Earth model background and a starry sky background. Generate YOLO-format labels for these simulated space target images and save them in a .txt file. All the images and their corresponding labels constitute the dataset. Divide the dataset into training, validation, and test sets according to a ratio of 70%, 20%, and 10%. Add real-life satellite model images to the test set. Proceed to Step 2.

[0008] Step 2: Improve the YOLOv5 model by combining the RepGhost module, the SimSPPF module, and the GSConv module to replace and improve some hierarchical structures in the network model, resulting in the Sim-YOLOv5 model. The specific operations are as follows:

[0009] Step 2.1: The RepGhost module, referring to the Ghost module, divides several convolutional layers in the YOLOv5 model into two parts: one part is ordinary convolution, which is used to generate original features; the other part uses the original features to generate reused features through a series of inexpensive operations. The RepGhost module uses the Add operation instead of the Concat operation to fuse the original features and reused features together, and modifies the module structure to meet the structural reparameterization rules.

[0010] The RepGhost bottleneck module is built based on the RepGhost module. The RepGhost module is introduced into the YOLOv5 model, and the RepGhost bottleneck replaces the CSP bottleneck in Backbone.

[0011] Step 2.2: Use the SimSPPF module as the last layer in the Sim-YOLOv5 model backbone, where the convolutional layer uses ReLU activation function, calculated as follows:

[0012]

[0013] Where x represents the input and max represents the maximum value.

[0014] Step 2.3: Select the GSConv module to build the Neck part of the lightweight model;

[0015] Based on the GSConv module, the VoV-GSCSP module is constructed. In this invention, the GSConv module and the VoV-GSCSP module are used in the Neck part of the YOLOv5 model to satisfy the narrow neck design paradigm, and then proceed to step 3.

[0016] Step 3: Place the Sim-YOLOv5 model into the deep learning environment, add module functions, and adjust the model training configuration parameters; then train the model using the images and labels from the training set.

[0017] During training, the target localization loss, confidence loss, and target classification loss gradually decrease and approach 0, while the precision P, recall R, and mean average precision mAP gradually increase and then stabilize. After the model training is completed, the optimal model weight best.pt is obtained, and then the optimal Sim-YOLOv5 model is obtained, and then proceed to step 4.

[0018] Step 4: Input the validation set to validate the model performance. This will give you the specific performance parameters of the optimal Sim-YOLOv5 model, including precision P, recall R, and mean average precision mAP. Proceed to step 5.

[0019] Step 5: Deploy the optimal Sim-YOLOv5 model to an embedded platform and accelerate it using TensorRT inference. Then, input the model into the test set for detection to obtain the target detection performance parameters of the model on the embedded platform, including detection accuracy and detection speed.

[0020] Compared with the prior art, the significant advantages of this invention are:

[0021] (1) In view of the problem that the publicly available space datasets are very limited, this invention uses Blender to construct a simulated space target dataset for model training, verification and testing; in addition, when generating simulated space satellite images, the Earth model and star background are randomly adopted, and data augmentation techniques such as motion blur, camera distortion and dispersion are used to effectively improve the generalization ability of the trained model.

[0022] (2) This invention introduces the RepGhost module into the YOLOv5 model for the first time, realizes implicit feature reuse through structural reparameterization, uses the Add operation to replace the Concat concatenation operation, and can directly fuse the original features with its inexpensive reused features. On the one hand, the inexpensive linear operation greatly reduces the number of model parameters and computational load, and on the other hand, the Add operation has higher hardware efficiency than the Concat operation.

[0023] (3) In this invention, the SimSPPF module replaces the SPPF module in the original YOLOv5 model. Specifically, SimSPPF replaces the activation function in the SPPF convolutional layer with ReLU. ReLU is faster to compute and can improve the inference speed of the model.

[0024] (4) The present invention uses a thin-neck design paradigm in the Neck layer of the model and introduces a lightweight convolutional module GSConv and its constructed VoV-GSCSP module, which greatly reduces the number of parameters and computation of the network model. The GSConv module can also enhance the nonlinear expression capability of the network, thereby reducing the inference time of the model while maintaining a certain detection accuracy. Attached Figure Description

[0025] Figure 1 This is a flowchart of the spatial target detection method based on the Sim-YOLOv5 model of the present invention.

[0026] Figure 2 This is a schematic diagram of the network structure of YOLOv5s, the basic model used in this invention.

[0027] Figure 3 This is a detailed flowchart of the steps in an embodiment of the present invention.

[0028] Figure 4 This is a graph showing the training parameters of the model in an embodiment of the present invention. Detailed Implementation

[0029] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0030] Combination Figure 1 and Figure 2 The spatial target detection method based on the Sim-YOLOv5 model described in this invention has the following specific steps:

[0031] Step 1: This invention uses the 3D graphics software Blender in conjunction with a 3D satellite model to create a large number of simulated space target images against an Earth model background and a starry sky background. Simultaneously, data augmentation techniques such as motion blur, camera distortion, and lens aberration are used to improve the diversity of image samples. Then, the Labelimg image annotation tool is used to generate YOLO-format labels for each simulated space target image and save them in a .txt file. All the images and their corresponding labels form a complete dataset. Finally, the dataset is divided into training, validation, and test sets according to a ratio of 70%, 20%, and 10%, respectively. Real-world images of the satellite model are also added to the test set. Proceed to Step 2.

[0032] Step 2: Modify the YOLOv5 model by combining the RepGhost module, the SimSPPF module, and the GSConv module to replace and improve some hierarchical structures in the network model, resulting in the improved Sim-YOLOv5 model. The specific steps are as follows:

[0033] Step 2.1: The YOLOv5 model contains several convolutional layers, and the output feature maps of these convolutional layers often have a lot of redundancy. The RepGhost module, referencing the Ghost module, divides the original convolutional layers into two parts: one part is the original ordinary convolution, used to generate the original features; the other part uses the original features to generate reused features through a series of inexpensive operations. However, the RepGhost module uses the Add operation instead of the Concat operation to fuse the original and reused features together, and modifies the module structure to meet the structural reparameterization rules. The Add operation has higher hardware efficiency than the Concat operation.

[0034] Assumption This represents the input data that needs to be processed and reused, i.e., the m original feature maps generated after ordinary convolution processing. This represents a feature map with n channels; Φ i (X) represents a series of inexpensive linear operations used to generate (nm) shadow feature maps. Without sacrificing generality, the RepGhost module's feature reuse through the Add operation is represented as follows:

[0035] Y = Add(X, Φ1(X), ..., Φ n-m (X))

[0036] Where h represents the height of the feature channel, w represents the width of the feature channel, and i represents the sequence number of the cheap operation.

[0037] This invention introduces the RepGhost module into the YOLOv5 model for the first time, and replaces the CSP bottleneck in the Backbone with the RepGhost bottleneck to reduce the overall number of parameters and computational load of the model, and achieve faster inference speed during object detection.

[0038] Step 2.2: The last layer in the original YOLOv5 backbone uses the SPPF module, where the convolutional layer uses the SiLU activation function, calculated as follows:

[0039]

[0040] This invention uses the SimSPPF module as the last layer in the model backbone, where the convolutional layer uses ReLU activation function, calculated as follows:

[0041]

[0042] In the above two formulas, x represents the input, e represents the exponentiation operation, and max represents taking the maximum value.

[0043] From the calculation formulas of SiLU and ReLU, it can be seen that e in SiLU x Power operations have high computational costs, so using the ReLU activation function instead of the SiLU activation function can speed up computation and thus improve the inference speed of the model.

[0044] Step 2.3: Replacing standard convolutions with depthwise separable convolutions can significantly reduce the number of model parameters and computational cost, making it very suitable for embedded platforms with limited storage and computing resources. However, its feature extraction and fusion capabilities are much lower than standard convolutions, so it is not suitable for this invention. Nevertheless, the lightweight convolution GSConv built using depthwise separable convolutions achieves a good balance between model accuracy and speed. Therefore, this invention chooses to use the GSConv module to construct the Neck part of the lightweight model.

[0045] Assumption This represents the input data for depthwise separable convolutions in the GSConv module, including Input feature mapping for each channel, This indicates that the GSConv module has c2 channels of output feature mapping, D s (I) represents the standard convolution operation. Without sacrificing generality, the feature output O of the GSConv module is represented as:

[0046]

[0047] Where h represents the height of the input feature map, w represents the width of the input feature map, and s represents the channel number for standard convolution; Shuffle represents the shuffling operation, used to shuffle the order of the feature set; Concat represents the concatenation operation, used to concatenate the features into a set.

[0048] Based on the GSConv module, the VoV-GSCSP module is constructed. This invention uses the GSConv module and the VoV-GSCSP module in the Neck part of the YOLOv5 model to meet the narrow neck design paradigm, thereby reducing the number of network parameters and computational cost, and enhancing the nonlinear expressive ability of the network. This reduces the inference time of the model while maintaining a certain level of detection accuracy.

[0049] The above three steps describe the improvement of the YOLOv5 model. In this invention, the improved model is named Sim-YOLOv5. Proceed to step 3.

[0050] Step 3: Place the improved Sim-YOLOv5 model in the configured deep learning environment on the PC, add the corresponding module functions to the algorithm, and modify some related model training configuration parameters, such as the number of training epochs, batch size, input image resolution (img-size), learning rate (lr), optimizer momentum (momnetum), optimizer weight decay coefficient (weight_decay), etc. Combine the images and labels from the training set, and then run train.py to start model training. During training, the terminal output shows that the target localization loss, confidence loss, and target classification loss gradually decrease and approach 0, while precision (P), recall (R), and mean average precision (mAP) gradually increase and then stabilize.

[0051] Target classification loss L CE The cross-entropy loss function is calculated using the following formula:

[0052]

[0053] p represents the probability distribution of the true label values, q represents the probability distribution of the predicted label values, k represents the label index, and l represents the total number of labels.

[0054] This invention replaces the target localization loss with the range intersection-union loss L. DIoU The calculation formula is as follows:

[0055]

[0056] Where IoU represents the intersection-over-union ratio of the predicted bounding box and the ground truth bounding box, and b represents the center position of the predicted bounding box. gt ρ represents the center position of the ground truth box, ρ represents the Euclidean distance, and c represents the diagonal length of the minimum closure rectangle region between the predicted box and the ground truth box.

[0057] Compared to the general intersection-union ratio (GIoU) ​​loss used in the original YOLOv5 algorithm, the distance intersection-union ratio (DIoU) loss takes into account the center distance between the predicted bounding box and the real target, which can provide more accurate target localization information.

[0058] After the model training is completed, the optimal model weights best.pt are obtained, which in turn yields the optimal Sim-YOLOv5 model. Proceed to step 4.

[0059] Step 4: Input the validation set to validate the model performance. This will give you the specific performance parameters of the optimal Sim-YOLOv5 model, including precision P, recall R, and mean average precision mAP. Proceed to step 5.

[0060] Step 5: Deploy the optimal Sim-YOLOv5 model to an embedded platform and accelerate it using TensorRT inference. Then, input the model into the test set for detection to obtain the target detection performance parameters of the model on the embedded platform, including detection accuracy and detection speed.

[0061] Example 1

[0062] Combination Figure 3 The spatial target detection method based on the Sim-YOLOv5 model described in this invention has the following specific implementation steps:

[0063] Step 1: Using the 3D graphics software Blender, combined with a BeiDou-3 satellite model, 5000 simulated images of space targets were created against an Earth model background and a starry sky background. An additional 1000 simulated images of space targets were created using data augmentation techniques such as motion blur, camera distortion, and lens aberration. Then, the Labelimg image annotation tool was used to generate YOLO-format labels for each image and save them in a .txt file. All the images and their corresponding labels constituted a set of space satellite target datasets. Finally, the training set included 4200 images, the validation set included 1200 images, and the test set included 600 images.

[0064] Step 2: Modify the YOLOv5 model to construct an improved Sim-YOLOv5 model. Specifically, introduce the RepGhost, SimSPPF, and GSConv modules into the YOLOv5 algorithm. Replace all C3 layers in the Backbone with C3Ghost layers, replace the CSP bottleneck in the Backbone with the RepGhost bottleneck, and replace the SPP layers with SimSPPF layers. Then, add the GSConv module to the Neck section, replacing all Conv layers with GSConv layers and all C3 layers with VoV-GSCSP layers. The improved Sim-YOLOv5 model is significantly smaller in size, computational cost, and parameter count compared to the original model.

[0065] Step 3: Place the improved Sim-YOLOv5 model in the configured deep learning environment on the PC, add the corresponding module functions to the algorithm, and then modify some related model training configuration parameters in the algorithm. The training epochs are 300, the training batch size is 4, the input image resolution size is 640, the initial learning rate lr0 is 0.01, the optimizer momentum is 0.937, and the optimizer weight decay coefficient is 0.0005.

[0066] Using the 4200 images and their corresponding labels from the training set, we directly run `train.py` to begin training the model. Figure 4 As can be seen, with the increase of training rounds, the localization loss, confidence loss, and classification loss gradually decrease and approach 0, but the precision P, recall R, and mean precision mAP gradually increase and then tend to stabilize.

[0067] Step 4: Input the validation set to validate the model performance. This will give you the specific performance parameters of the optimal Sim-YOLOv5 model, including precision (P), recall (R), and mean average precision (mAP).

[0068] Step 5: Deploy the improved lightweight object detection model Sim-YOLOv5 obtained in Step 2 and the optimal model weights best.pt obtained after training in Step 3 to the Jetson TX2 embedded platform. Configure the deep learning environment required to run the object detection algorithm on the Jetson TX2 embedded platform, and verify the performance of the improved Sim-YOLOv5 model on the Jetson TX2 embedded platform using a test set. To achieve faster detection speed and improve real-time performance, TensorRT inference acceleration is used. Input the test image into the detection network to output the detection results.

[0069] In summary, the improved Sim-YOLOv5 network model achieves a certain level of detection accuracy while reducing model size. When applied to the Jetson TX2 embedded platform, it can also utilize TensorRT to accelerate object detection inference, thereby improving the real-time performance of object detection.

[0070] The following description, based on experimental results from this embodiment, further illustrates the effectiveness of the present invention. The experiment primarily compares the performance parameters of the original YOLOv5 and the improved Sim-YOLOv5 in this invention, using YOLOv5s as the baseline network. Experimental verification is then conducted on a generated simulated spatial target dataset. The following are the objective data regarding the experimental environment and results:

[0071] 1. Experimental hardware and software environment

[0072] operating system Windows 10 CPU model Intel Core i7-8700 CPU cores 6 cores GPU model NVIDIA GeForce GTX 1660 GPU memory 6.0GB CUDA 11.6 Python 3.8.16 Pytorch 1.13.1

[0073] 2. Comparative Analysis of Experimental Results

[0074]

[0075] The data above shows that the improved Sim-YOLOv5s achieves a reduction in model size, number of parameters, and computational cost compared to YOLOv5s, while maintaining almost the same mAP as the original model. The detection speed also decreases only slightly, by 1.4%. In summary, Sim-YOLOv5s achieves a significant reduction in the weight of the original YOLOv5s network model, with negligible losses in detection accuracy and speed, resulting in a good overall performance.

[0076] The improved Sim-YOLOv5s algorithm and network model were imported into the Jetson TX2 embedded platform used in this embodiment. The `detect.py` file was run to detect images on the prepared test set, with an average detection time of 465.5ms per image. Then, the optimal weights `best.pt` were transformed accordingly, and TensorRT inference acceleration was used to accelerate detection on the same test set, resulting in an average detection time of 35.4ms per image. As can be seen, TensorRT inference acceleration significantly improves the object detection speed on embedded platforms.

[0077] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A spatial target detection method based on the Sim-YOLOv5 model, characterized in that, The steps are as follows: Step 1: Using 3D graphics software and a 3D satellite model, create several simulated images of space targets against an Earth model background and a starry sky background. Generate YOLO-format labels for these simulated space target images and save them in a .txt file. All the images and their corresponding labels constitute the dataset. Divide the dataset into training, validation, and test sets according to a ratio of 70%, 20%, and 10%. Add real-life satellite model images to the test set. Proceed to Step 2. Step 2: Improve the YOLOv5 model by combining the RepGhost module, the SimSPPF module, and the GSConv module to replace and improve some hierarchical structures in the network model, resulting in the Sim-YOLOv5 model. The specific operations are as follows: Step 2.1: The RepGhost module, referring to the Ghost module, divides several convolutional layers in the YOLOv5 model into two parts: one part is ordinary convolution, which is used to generate original features; the other part uses the original features to generate reused features through a series of inexpensive operations. The RepGhost module uses the Add operation instead of the Concat operation to fuse the original features and reused features together, and modifies the module structure to meet the structural reparameterization rules. The RepGhost bottleneck module is built based on the RepGhost module. The RepGhost module is introduced into the YOLOv5 model, and the RepGhost bottleneck replaces the CSP bottleneck in Backbone. Step 2.2: Use the SimSPPF module as the last layer in the Sim-YOLOv5 model backbone, where the convolutional layer uses ReLU activation function, calculated as follows: Where x represents the input and max represents the maximum value; Step 2.3: Select the GSConv module to build the Neck part of the lightweight model; In the Neck section of the YOLOv5 model, the GSConv module and VoV-GSCSP module are used to satisfy the narrow neck design paradigm, and then proceed to step 3. Step 3: Place the Sim-YOLOv5 model into the deep learning environment, add module functions, and adjust the model training configuration parameters; then train the model using the images and labels from the training set. During training, the target localization loss, confidence loss, and target classification loss gradually decrease and approach 0, while the precision P, recall R, and mean average precision mAP gradually increase and then stabilize. After the model training is completed, the optimal model weight best.pt is obtained, and then the optimal Sim-YOLOv5 model is obtained. Proceed to step 4. Step 4: Input the validation set to validate the model performance and obtain the specific performance parameters of the optimal Sim-YOLOv5 model, including precision P, recall R and mean average precision mAP, and then proceed to step 5. Step 5: Deploy the optimal Sim-YOLOv5 model to an embedded platform and accelerate it using TensorRT inference. Then, input the model into the test set for detection to obtain the target detection performance parameters of the model on the embedded platform, including detection accuracy and detection speed.

2. The spatial target detection method based on the Sim-YOLOv5 model according to claim 1, characterized in that, In step 1, data augmentation is performed on the images of the simulated space target that exhibit motion blur, camera distortion, and lens aberration. The augmented images are then used to replace the original images before further processing.

3. The spatial target detection method based on the Sim-YOLOv5 model according to claim 1, characterized in that, Step 2.1 is as follows: Assumption This represents the input data that needs to be processed and reused, i.e., the m original feature maps generated after ordinary convolution processing. This represents a feature map with n channels; Φ i (X) represents a series of inexpensive linear operations used to generate (nm) shadow feature maps. Without sacrificing generality, feature reuse via the Add operation can be represented as follows: Y=Add(X,Φ1(X),...,Φ n-m (X)) Where h represents the height of the feature channel, w represents the width of the feature channel, and i represents the sequence number of the cheap operation. The RepGhost bottleneck module is built based on the RepGhost module. The RepGhost module is introduced into the YOLOv5 model, and the RepGhost bottleneck replaces the CSP bottleneck in Backbone.

4. The spatial target detection method based on the Sim-YOLOv5 model according to claim 3, characterized in that, In step 2.3, the GSConv module is selected to construct the Neck part of the lightweight model, as follows: Assumption I represents the input data for depthwise separable convolutions in the GSConv module, including The GSConv module has c2 channels of input feature mapping and c2 channels of output feature mapping. D s (I) represents the standard convolution operation. Without sacrificing generality, the feature output O of the GSConv module is represented as: Where h represents the height of the input feature map, w represents the width of the input feature map, and s represents the channel number for standard convolution; Shuffle represents the shuffling operation, used to shuffle the order of the feature set; Concat represents the concatenation operation, used to concatenate the features into a set.

5. The spatial target detection method based on the Sim-YOLOv5 model according to claim 4, characterized in that, In step 3, the model training configuration parameters include training epochs, training batch size, input image resolution (img-size), learning rate (lr), optimizer momentum (momnetum), and optimizer weight decay coefficient (weight_decay).

6. The spatial target detection method based on the Sim-YOLOv5 model according to claim 5, characterized in that, In step 3, the classification loss L of the target CE The cross-entropy loss function is calculated using the following formula: p is the probability distribution of the true label values, q is the probability distribution of the predicted label values, k represents the label number, and l represents the total number of labels; Change the target localization loss to use the distance intersection and reunification loss L DIoU The calculation formula is as follows: Where IoU represents the intersection-over-union ratio of the predicted bounding box and the ground truth bounding box, and b represents the center position of the predicted bounding box. gt ρ represents the center position of the ground truth box, ρ represents the Euclidean distance, and c represents the diagonal length of the minimum closure rectangle region between the predicted box and the ground truth box.