Method and system for detecting target object easy to reflect light in weak light environment in nuclear industry

By combining data enhancement and multi-channel attention mechanism residual neural network in the nuclear industrial environment for object detection, the target detection error and accuracy problems caused by low light and easy reflection in the nuclear industrial environment are solved, and a more efficient and safe target recognition and detection effect is achieved.

CN119992055APending Publication Date: 2025-05-13HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510086099.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The low-light and easy-reflective characteristics in the nuclear industrial environment lead to large errors in robot target detection and poor accuracy, limiting the application range of robots in nuclear environments and may cause safety hazards.

Method used

Using a residual neural network combining data augmentation and multi-channel attention mechanism, a square overexposure canvas is generated through Gaussian distribution and the target image is superimposed for data augmentation, and the spatial and channel attention mechanisms are used to adjust the degree of attention of the model, and finally the fusion feature is generated through a multi-path feature aggregation pyramid network for object detection.

Benefits of technology

It significantly improves the robot's ability to identify target objects in a complex nuclear industrial environment, enhances the efficiency and safety of maintenance work, reduces the risk of manual intervention, and ensures the reliability and safe operation of nuclear facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992055A_ABST
    Figure CN119992055A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for detecting a target object easy to reflect light in a weak light environment in the nuclear industry. The core lies in a residual neural network combining data enhancement and a multi-channel attention mechanism. The method comprises the following steps: generating an overexposure canvas by using Gaussian distribution, and superposing the overexposure canvas with target image pixels for data enhancement; complex environment features are extracted through a residual neural network; parallel processing is carried out by adopting a space and channel attention module, and the correlation and attention of the features are enhanced; a multi-path feature aggregation pyramid network is utilized to generate fusion features; finally, the category probability is calculated through a full connection layer and an activation function of the prediction module, and accurate recognition is achieved. According to the method, through decoupling design and block construction of the optimization model, the target identification capability of the robot in a complex nuclear industry environment is remarkably improved, the maintenance efficiency and safety are enhanced, the manual intervention risk is reduced, and reliable and safe operation of nuclear facilities is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of nuclear energy technology, in particular to the field of indoor target detection in a nuclear environment, and in particular to a method and system for detecting reflective target objects in a weak-light environment in the nuclear industry. Background Art

[0002] In the nuclear industry environment, accurate target detection can timely detect and locate potential faults or damage to prevent accidents; and accurately identifying and locating objects and obstacles in the environment can enable robots to efficiently plan paths and perform tasks, avoiding collisions and misoperations. The weak light and reflective environmental object characteristics in nuclear industry hot rooms lead to large robot object recognition errors and poor detection target accuracy, which limits the application scope of robots in nuclear environments and may also cause safety hazards and endanger the safe operation of nuclear facilities. Therefore, there is an urgent need for an accurate and efficient target detection system to improve the robot's target detection capabilities.

[0003] Existing target detection methods can be roughly divided into two categories:

[0004] 1) Object detection method based on artificial features. By artificially constructing feature extraction standards and rules (such as prominent appearance, angular edges, etc.), the features of the target object in the image are extracted, and the extracted information is converted into feature information with greater discrimination (such as gradient, grayscale). By matching the extracted features, object detection is achieved.

[0005] 2) Object detection method based on traditional deep learning. By using the powerful feature extraction ability of deep learning, the image information is adaptively converted into a high-dimensional feature vector with significant discrimination. By formulating a reasonable loss function and optimizer, the weight parameters of the network are continuously updated, and finally the classifiers such as SoftMax are used to calculate the probability of each classification to obtain the best object detection result.

[0006] The process of target detection method based on artificial features is as follows Figure 1 First, the image information is acquired through the camera, and the feature extraction standards and rules (such as prominent appearance, angular edges, etc.) are artificially constructed according to the input data. The extracted information is converted into feature information with greater discrimination (such as gradient, grayscale), and the target object features in the image are extracted. By matching the extracted features, target detection is achieved.

[0007] The disadvantages of the target detection method based on artificial features are mainly the following two points:

[0008] 1) Artificially formulated feature extraction standards and rules are based on experience and understanding of specific problems. This feature selection may not fully capture the complex and diverse patterns in the data. When the target is far away from the light source, the target appears too dim and difficult to distinguish due to insufficient light; when the target is close to the light source, overexposure around the target may cause the target to be submerged, increasing the risk of missed detection. Even considering the noise, lighting changes, rotation and scale changes in the image, there are still problems such as insufficient robustness and poor detection effect.

[0009] 2) The computational complexity of artificial feature extraction is high. When processing high-resolution images or real-time applications, the computational cost is high and it is not suitable for resource-constrained devices.

[0010] The process of target detection method based on deep learning is as follows Figure 2 As shown in the figure, by using the powerful feature extraction capability of deep learning, the image information is adaptively converted into a high-dimensional feature vector with significant discrimination. By formulating a reasonable loss function and optimizer, the weight parameters of the network are continuously updated. Finally, the MLP calculates the probability of each classification to obtain the best target detection result.

[0011] The shortcomings of traditional deep learning detection methods are mainly the following two points:

[0012] 1) When dealing with complex environments such as low light and easy reflection, the data enhancement method based on traditional deep learning has limited effect in improving detection accuracy through means such as brightness adjustment and contrast adjustment.

[0013] 2) The convolutional feature extraction method based on traditional deep learning cannot focus on important areas in complex environments, resulting in the neglect of key features. Summary of the invention

[0014] In view of the shortcomings of the existing technology, the present invention proposes a target object detection method and system of a residual neural network that combines data enhancement and multi-channel attention mechanism in view of the weak light and reflective environment characteristics of the nuclear industry. Through specific data processing and neural network structure design, this method significantly improves the robot's ability to recognize target objects in complex nuclear industrial environments, thereby enhancing the efficiency and safety of maintenance work and reducing the risk of manual intervention, which is of great significance to ensuring the reliable and safe operation of nuclear facilities.

[0015] In order to solve the above technical problems, the technical solution adopted by the present invention is: a method for detecting reflective target objects in a weak light environment of the nuclear industry, comprising the following steps:

[0016] (1) Generate a square overexposed canvas through Gaussian distribution and superimpose it with the target image pixels in a complex environment to achieve data enhancement. Use a residual neural network to perform convolution operations on the enhanced data to extract target features in a complex environment. The features output by the residual neural network are used as the input of the attention module. The spatial attention module and the channel attention module assign weights to different positions and channels in parallel, adjust the model's attention level, and obtain target aggregation features, ensuring the channel correlation and spatial correlation of the features. The aggregated features are input into the multi-path feature aggregation pyramid network to generate fusion features.

[0017] (2) The fused features are passed to the prediction module, which calculates the final category probability distribution through a series of fully connected layers and activation functions to achieve accurate recognition of targets in complex environments.

[0018] The method of the present invention not only performs data enhancement for image data in complex environments, but also guides the model to focus on important areas through the attention mechanism, ensuring that the network will not miss detections or make false detections, thereby improving the effect of feature extraction and processing efficiency, and ensuring the accuracy of target object detection in weak light and reflective environments in the nuclear industry.

[0019] The specific implementation process of step (1) includes:

[0020] 1) Input image data;

[0021] 2) Perform data enhancement on the image;

[0022] 3) performing convolution processing on the data augmented image to extract target input features;

[0023] 4) Performing attention fusion processing on the input features to obtain aggregated features;

[0024] 5) performing cross-level fusion on the aggregated features to obtain fused level features;

[0025] It can be seen from the above process that the present invention enhances the data in weak light and reflective environments, thereby improving the feature extraction capability of the target.

[0026] The attention mechanism is used before the multi-path feature aggregation pyramid network to guide the model to focus on important areas and ignore irrelevant background, thereby improving the effect of feature extraction and the final task performance.

[0027] The specific implementation process of image data enhancement includes: traversing the labeled targets, deciding whether to perform overexposure enhancement with a certain probability, using uniform distribution to randomly sample the target object annotation box to obtain a center point, and generating a square canvas based on the weighted Gaussian distribution in multi-dimensional space with the center point as the center. The data enhancement operation is completed by superimposing the expanded dimension with the original image pixels.

[0028] In step 2), the target data enhancement operation is expressed as follows:

[0029]

[0030] Among them, x and y are the position coordinates of a pixel point in the canvas, v is the grayscale value of a pixel point in the canvas on the original image, α and β are two hyperparameters, α is used to control the influence of grayscale space on the generated canvas, and β is used to control the influence of distance space on the generated canvas.

[0031] In step 3), CSP_ResNet52 is used to perform convolution processing on the image. CSP_ResNet52 is based on the ResNet residual network, and CSPNet is added to enrich the network gradient flow propagation and reduce the network calculation amount. The network structure integrates CBL, residual unit and CSP modules. The CBL module contains convolution layer, batch normalization layer and Leaky ReLU activation function. The Res Unit module implements residual connection through the CBL layer and performs element addition operation. The CSPX module is composed of CBL layer and several residual units, and splicing operation is performed at the end.

[0032] In step 4), the attention mechanism is composed of two attention modules: channel attention and spatial attention. The channel attention mechanism assigns different weights to different channels, so that the model can better focus on channels containing important information. The spatial attention mechanism assigns different weights to each position of the feature map, so that the model can better focus on positions containing important information. The input features are respectively passed through the two attention modules to obtain the first-stage module features, the first-stage module features are superimposed in the feature dimension direction to obtain the first-stage connection features, the first-stage connection features are input into the two attention modules to obtain the second-stage module features, the second-stage module features are superimposed in the feature dimension direction to obtain the second-stage connection features, and the final aggregated features are obtained. The spatial attention module and the channel attention module operate in parallel to assign weights to different positions and channels, ensuring the channel correlation and spatial correlation of the features.

[0033] In step 5), a multi-path feature aggregation pyramid network is selected for cross-level fusion, and the network includes two feature fusion loops. The first bottom-up feature fusion loop of the network increases the deep semantic information of high-resolution features; the second top-down feature fusion loop of the network enables low-resolution feature maps to have higher positioning accuracy; the network aggregates features from different pyramid levels through multiple parallel paths to enhance the recognition ability of the target detection model for complex scenes and multi-scale targets.

[0034] The specific implementation process of step (2) includes:

[0035] 1) Mapping the fusion level features to the prediction module to calculate the category probability distribution and obtain the final detection result;

[0036] The present invention uses the fused hierarchical features output by the feature extraction network as the input of the prediction module, and uses the prediction module to convert the fused hierarchical features into classification probability distribution to obtain the final detection result.

[0037] The prediction module uses a detection method based on anchor boxes. The anchor boxes are predefined boxes. A set of fixed anchor boxes are used as reference boxes. Adjustments are made according to the characteristics of the data set. A series of bounding boxes of different proportions and sizes are generated around the boxes to adapt to various target sizes and shapes, and finally the target detection results are obtained.

[0038] Accordingly, the present invention also provides a system for detecting reflective target objects in a weak-light environment in the nuclear industry, comprising a computer device; the computer device is programmed or configured to be used in the steps of the method of the present invention.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The present invention enhances data based on weak light and reflective complex environments, thereby improving the feature extraction capability of the target.

[0041] 2. The present invention guides the model to focus on important areas through the attention mechanism before the multi-path feature aggregation pyramid network, ignoring irrelevant background, thereby improving the effect of feature extraction and the final target detection accuracy.

[0042] 3. The feature extraction network incorporates the CSPNet module based on the ResNet residual network to enrich the gradient propagation of the network and reduce the amount of network calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of the target detection method based on artificial features;

[0044] Figure 2 This is a flow chart of the target detection method based on traditional deep learning;

[0045] Figure 3 This is a schematic diagram of the architecture of the nuclear industry's low-light environment reflective target object detection system;

[0046] Figure 4 This is a comparison chart of the input image data enhancement effect (the left side is the original image, and the right side is the enhanced image);

[0047] Figure 5 This is the CSP_ResNet52 network structure diagram;

[0048] Figure 6This is the network structure diagram of the attention mechanism;

[0049] Figure 7 A pyramid network structure diagram for multi-path feature aggregation;

[0050] Figure 8 Comparison between the present invention and the open source algorithm YOLOV8 (YOLOV8 on the left and the present invention on the right);

[0051] Fig. 9 This is a schematic diagram of the core process of the present invention. DETAILED DESCRIPTION

[0052] The system architecture of the present invention is as follows Figure 3 As shown. Image information is obtained through the camera; a square overexposed canvas is generated through Gaussian distribution, and data enhancement is achieved by superimposing it with the target image pixels in a complex environment. A residual neural network is used to perform convolution operations on the enhanced data to extract target features in a complex environment. The features output by the residual neural network are used as the input of the attention module. The spatial attention module and the channel attention module assign weights to different partial positions and channels in parallel, adjust the degree of attention of the model, and obtain the target aggregation features, which ensures the channel correlation and spatial correlation of the features. The aggregated features are input into the multi-path feature aggregation pyramid network to generate fused features. Then, these fused features are passed to the prediction module, which calculates the final category probability distribution through a series of fully connected layers and activation functions to achieve accurate recognition of targets in complex environments.

[0053] The feature extraction module construction process is shown in the figure. The main method steps of this module are as follows:

[0054] (1) Input image data;

[0055] The image data of the present invention is obtained by a binocular camera, and a light source is arranged at a specified position to illuminate the environment to obtain images under weak light and easy reflection conditions.

[0056] (2) Perform target data enhancement on the image;

[0057] First, the labeled objects are traversed, and a decision is made with a certain probability whether to perform overexposure enhancement. A center point is obtained by randomly sampling the labeled box of the target object using uniform distribution. A square canvas is generated based on the weighted Gaussian distribution of multidimensional space with the center point as the center. The overexposure data enhancement operation is completed by superimposing the expanded dimension with the original image pixels. The data enhancement effect is as follows: Figure 4 shown.

[0058] (3) performing convolution processing on the data augmented image to extract target input features;

[0059] Use CSP_ResNet52 to perform convolution processing on the image. CSP_ResNet52 is based on the ResNet residual network. CSPNet is added to enrich the network gradient flow propagation and reduce the network calculation amount. The network structure combines CBL, residual unit and CSP modules. The CBL module contains convolution layer, batch normalization layer and Leaky ReLU activation function. The Res Unit module implements residual connection through the CBL layer and performs element addition operation. The CSPX module is composed of CBL layer and several residual units, and splicing operation is performed at the end. The network structure of CSP_ResNet52 is as follows Figure 5 shown.

[0060] (4) performing attention mechanism fusion processing on the input features to obtain aggregated features;

[0061] The attention mechanism consists of two attention modules: channel attention and spatial attention. The channel attention mechanism assigns different weights to different channels, so that the model can better focus on channels containing important information. The spatial attention mechanism assigns different weights to each position in the feature map, so that the model can better focus on positions containing important information. The input features are respectively passed through the two attention modules to obtain the first-stage module features. The first-stage module features are superimposed in the feature dimension direction to obtain the first-stage connection features. The first-stage connection features are input into the two attention modules to obtain the second-stage module features. The second-stage module features are superimposed in the feature dimension direction to obtain the second-stage connection features to obtain the final aggregated features. The spatial attention module and the channel attention module operate in parallel to assign weights to different positions and channels, ensuring the channel correlation and spatial correlation of the features. The network structure of the attention mechanism is as follows: Figure 6 shown.

[0062] (5) performing cross-level fusion on the aggregated features to obtain fused level features;

[0063] Cross-level fusion selects a multi-path feature aggregation pyramid network, which contains two feature fusion loops. The first bottom-up feature fusion loop of the network increases the deep semantic information of high-resolution features; the second top-down feature fusion loop of the network enables low-resolution feature maps to have higher positioning accuracy; the network aggregates features from different pyramid levels through multiple parallel paths to improve the recognition ability of the target detection model for complex scenes and multi-scale targets. The structure of the multi-path feature aggregation pyramid network is as follows Figure 7 shown.

[0064] The fused hierarchical features finally obtained by the feature extraction module are input into the prediction module. The prediction module uses a detection method based on anchor boxes. The anchor boxes are predefined boxes. A set of fixed anchor boxes are used as reference boxes. They are adjusted according to the characteristics of the data set. Bounding boxes of different proportions and sizes are generated around the boxes to adapt to various target sizes and shapes, and finally the target detection results are obtained.

[0065] The core process of the present invention is as follows Fig. 9 shown.

[0066] Comparative experiment: The present invention establishes a hot chamber target data set based on lighting angle and intensity for training and evaluating the network model. A module peeling test is designed to verify the effectiveness of each module of the present invention. In addition, the present invention is compared with other methods to prove the detection accuracy of the present invention.

[0067] (1) Nuclear industry hot cell target data set

[0068] Nuclear industry hot rooms are typical non-general scenarios. Existing public data sets lack data sets under weak lighting and reflective conditions in nuclear environment hot rooms. To verify the effectiveness of the model, the present invention simulates the motion path of the robotic arm and captures the target from different positions to obtain 33,000 frames of image data to form a hot room target data set. The data set consists of two parts: a pre-training data set and a refined training data set. The annotation box accuracy of the pre-training data set is slightly lower than that of the refined training data set. Considering the size and production time of the data set, it is divided into two parts. At the same time, the pre-training data set also plays a role in enhancing the robustness of the model. It uses various color gamut color lights to enhance the model's generalization ability for light colors. The final data set sample classification is shown in Table 1.

[0069] Table 1. Sample distribution of nuclear industry hot room target dataset

[0070]

[0071] (2) Model building and training

[0072] The model building platform of the present invention is a 64-bit Ubuntu 20.04 system, using an Intel i7 processor and NVIDIA RTX3050. The Python 3.8 programming language and the pytorch 1.9 deep learning training framework are used to write the model program, and the NVIDIA CUDA 11.8.0 and cuDNN 8.9.2 deep learning accelerators are used to improve the model operation speed. The model contains a total of 52 convolutional layers.

[0073] (3) Table 2 Model training hyper parameters

[0074]

[0075] (3) Data enhancement model stripping experiment

[0076] The constructed network model is taken as the baseline model. Overexposed data enhancement is added to the baseline model as a comparison model. The evaluation results on the overexposed data set are used to analyze the model overexposed data enhancement algorithm. The experimental model indicators are shown in Table 3. Through analysis, it is found that after adding overexposed data enhancement, the indicators of the model on the overexposed data set are greatly improved, map50 is improved by 3%, and map5095 is improved by 10%.

[0077] Table 3 Data enhancement model stripping experiment

[0078]

[0079] In this invention, map refers to the average of the average precision (ap) of all categories. This invention is a single-category detection, so map is the same as ap. Map50 and map5095 refer to the map values ​​under different IOU thresholds. Map5095 represents the average of 10 map values ​​with an intersection-over-union (IOU) from 0.5 to 0.95 with a step size of 0.05. Compared with map50, this indicator can better reflect the model's prediction position accuracy and the generalization performance of the model.

[0080] (4) Attention mechanism model peeling experiment

[0081] The constructed network model is used as the baseline model, and the attention mechanism is added to the baseline model as a comparison model. The specific indicators of the experiment are shown in Table 4. After adding the attention mechanism, the two evaluation indicators of the model are greatly improved. Compared with the basic network model, the multi-attention mechanism designed in this paper has increased to 90.39% on map50 and 61.58% on map5095, which shows the effectiveness of the overall structure of the multi-attention mechanism.

[0082] Table 4 Attention mechanism model stripping experiment

[0083]

[0084]

[0085] (5) Comparison of existing methods

[0086] The present invention is compared with the open source algorithm YOLOV8 to verify the accuracy of the target detection task in complex environments. Table 5 shows the comparison results. Figure 8 The experimental parameter settings of the two are the same. In terms of data enhancement, both use a random flip probability of 0.5 and a Mosaic data enhancement probability of 1.0.

[0087] Table 5 Comparison of existing methods

[0088]

Claims

1. A method and system for detecting reflective target objects in a weak light environment in the nuclear industry, characterized in that: The following steps are involved: (1) Generate a square overexposed canvas through Gaussian distribution and superimpose it with the target image pixels in a complex environment to achieve data enhancement. Use a residual neural network to perform convolution operations on the enhanced data to extract target features in a complex environment. The features output by the residual neural network are used as the input of the attention module. The spatial attention module and the channel attention module assign weights to different positions and channels in parallel, adjust the model's attention level, and obtain target aggregation features, ensuring the channel correlation and spatial correlation of the features. The aggregated features are input into the multi-path feature aggregation pyramid network to generate fusion features. (2) The fused features are passed to the prediction module, which calculates the final category probability distribution through a series of fully connected layers and activation functions to achieve accurate recognition of targets in complex environments.

2. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 1 is characterized in that: The specific implementation process of step (1) includes: 1) Input image data; 2) Perform data enhancement on the image; 3) performing convolution processing on the data augmented image to extract target input features; 4) Performing attention fusion processing on the input features to obtain aggregated features; 5) Perform cross-level fusion on the aggregated features to obtain fused hierarchical features.

3. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 2, characterized in that: The specific implementation process of image data enhancement includes: traversing the labeled targets, deciding whether to perform overexposure enhancement with a certain probability, using uniform distribution to randomly sample the target object annotation box to obtain a center point, and generating a square canvas based on the weighted Gaussian distribution in multidimensional space with the center point as the center. The overexposure data enhancement operation is completed by superimposing the expanded dimension with the original image pixels.

4. The method and system for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 3, characterized in that: In step 2), the target data enhancement operation is expressed as follows: Among them, x and y are the position coordinates of a pixel point in the canvas, v is the grayscale value of a pixel point in the canvas on the original image, α and β are two hyperparameters, α is used to control the influence of grayscale space on the generated canvas, and β is used to control the influence of distance space on the generated canvas.

5. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 2, characterized in that: In step 3), CSP_ResNet52 is used to perform convolution processing on the image. CSP_ResNet52 is based on the ResNet residual network, and CSPNet is added to enrich the network gradient flow propagation and reduce the network calculation amount. The network structure integrates CBL, residual unit and CSP modules. The CBL module contains convolution layer, batch normalization layer and Leaky ReLU activation function. The Res Unit module implements residual connection through the CBL layer and performs element addition operation. The CSPX module is composed of CBL layer and several residual units, and splicing operation is performed at the end.

6. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 2, characterized in that: In step 4), the attention mechanism is composed of two attention modules: channel attention and spatial attention. The input features are respectively passed through the two attention modules to obtain the first-stage module features. The first-stage module features are superimposed in the feature dimension direction to obtain the first-stage connection features. The first-stage connection features are input into the two attention modules to obtain the second-stage module features. The second-stage module features are superimposed in the feature dimension direction to obtain the second-stage connection features to obtain the final aggregated features.

7. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 2, characterized in that: In step 5), a multi-path feature aggregation pyramid network is selected for cross-level fusion, and the network includes two feature fusion loops. The first bottom-up feature fusion loop of the network increases the deep semantic information of high-resolution features; the second top-down feature fusion loop of the network enables low-resolution feature maps to have higher positioning accuracy; the network aggregates features from different pyramid levels through multiple parallel paths to enhance the recognition ability of the target detection model for complex scenes and multi-scale targets.

8. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 1, characterized in that: The specific implementation process of step 2) includes: 6) Mapping the fusion level features to the prediction module to calculate the category probability distribution and obtain the final detection result.

9. The method for detecting reflective target objects in a weak light environment in the nuclear industry according to claim 8, characterized in that: In step 6), the prediction module uses a detection method based on anchor boxes. Anchor boxes are predefined boxes, and a set of fixed anchor boxes are used as reference boxes. According to the characteristics of the data set, bounding boxes of different proportions and sizes are generated around the reference box to adapt to various target sizes and shapes, and finally the target detection result is obtained.

10. A nuclear industry low-light environment reflective target object detection system, characterized in that: The method comprises a computer device, wherein the computer device is programmed or configured to execute the steps of the method according to any one of claims 1 to 9.