A target detection method based on random uncertainty

By constructing an object detection model, using adaptive feature alignment and general distribution modeling of detection box coordinates, the position and category prediction of detection boxes are optimized, and the random uncertainty problem in deep learning object detection is solved, and the detection accuracy and box quality representation are improved.

CN117058476BActive Publication Date: 2025-08-26UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310778187.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2025-08-26
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

The existing deep learning-based object detection methods have problems with random uncertainty, including spatial uncertainty and semantic uncertainty, which leads to inaccurate boundary and category prediction of detection boxes and incomplete quality representation of detection boxes.

Method used

The object detection model is constructed, through the adaptive feature alignment module, the general distribution model of detection box coordinates, and the prediction box weighted average module, the probability distribution and random sampling technology are used to optimize the position and category prediction of the detection box, and the FocalLoss and GIoULoss loss functions are trained.

Benefits of technology

It significantly improves the accuracy of the position and category prediction of the detection box, improves the detection accuracy of the model in complex scenarios, and provides high-quality detection box information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058476B_ABST
    Figure CN117058476B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of target detection, and in particular to a target detection method based on random uncertainty. This method inputs an image to be identified into a constructed target detection model to obtain the category and coordinates of the object in the image. The training of the target detection model includes the following steps: extracting original classification features and original regression features from training data and inputting them into an adaptive feature alignment module to obtain optimized classification features; calculating the general distribution of detection frame coordinates and the determined values ​​of detection frame coordinates based on the original regression features; inputting the original regression features, optimized classification features, and the determined values ​​of detection frame coordinates into a prediction frame weighted average module to obtain optimized detection frame coordinates; inputting the optimized classification features and the general distribution of detection frame coordinates into a target category prediction network to obtain an optimized category score; and training the target detection model according to a loss function. The present invention can improve detection accuracy in complex scenarios and predict high-quality detection frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a target detection method based on random uncertainty. Background Art

[0002] Object detection, a key task in computer vision, is widely used in fields such as autonomous driving and object tracking. In recent years, deep learning-based object detection methods have significantly improved model accuracy and inference speed. Mainstream object detection methods consist of two modules: a feature extraction module and a detector module. The detector module typically consists of a classification branch and a regression branch. Most deep learning-based object detection methods propose deterministic object detection models, representing the coordinates of the detection box as fixed values ​​and modeling the convolutional sampling process of the detector classification branch as a deterministic process.

[0003] However, due to factors related to the observed data itself, such as signal acquisition noise and data annotation errors, deep learning methods are subject to random uncertainty. Object detection methods based on deep learning also suffer from random uncertainty. Based on the regression and classification tasks in object detection, random uncertainty can be further divided into spatial uncertainty and semantic uncertainty.

[0004] First, for regression tasks, the boundaries of the detection box are uncertain due to problems such as object truncation, occlusion, and blurred input images. This means that object detection tasks have spatial uncertainty. Second, for classification tasks, the shape of each object in the input image is random, while the convolutional receptive field of the detector classification branch is fixed. The convolutional features are not aligned with the object's position, resulting in uncertainty in the object's category. This means that object detection tasks have semantic uncertainty, which ultimately leads to inaccurate object category predictions.

[0005] Secondly, the parallel structure of the classification and regression branches of the target detector will also lead to spatial prediction misalignment, affecting the detection performance of the model.

[0006] Finally, mainstream object detection methods only use the category score as a measure of the quality of the detection frame, while ignoring the position quality of the detection frame. This inaccurately represents the quality of the detection frame, leading to the inadvertent deletion of high-quality detection frames during post-processing. This results in inaccurate and incomplete object detection results. Quality refers to the accuracy and reliability of the detection frame. A high-quality detection frame is one that is accurately positioned, appropriately sized, and accurately predicts the target object category and confidence level. Summary of the Invention

[0007] In order to solve the above problems, the present invention provides a target detection method based on random uncertainty.

[0008] This method builds a target detection model, inputs the image to be identified into the target detection model, and outputs the category and coordinates of the object in the image. The training of the target detection model includes the following steps:

[0009] Step 1: Prepare image data for target category and category score annotation, detection frame coordinate annotation, and pre-process the annotated image as training data;

[0010] Step 2: Input the training data into the feature extraction network to extract its spatial semantic features;

[0011] Step 3: Input the spatial semantic features into the classification branch feature extraction network and the regression branch feature extraction network to obtain the original classification feature X cls With the original regression feature X reg ;

[0012] Step 4: convert the original classification feature X cls With the original regression feature X reg Input to the adaptive feature alignment module to obtain optimized classification features

[0013] Step 5: Based on the original regression feature X reg Calculate the general distribution of detection box coordinates and the determined value y of the detection box coordinates dtrmd ;

[0014] Step 6: Regress the original feature X reg , optimize classification features The determined value y of the detection box coordinate dtrmd , input to the prediction box weighted average module to obtain the optimized detection box coordinate r refine ;

[0015] Step 7: Optimize classification features and the general distribution of detection box coordinates Input to the target category prediction network to obtain the optimized category score;

[0016] Step 8: Train the target detection model according to the classification loss function FocalLoss and the regression loss function GIoULoss until the preset training completion conditions are met.

[0017] Furthermore, step two specifically includes inputting the training data into a convolutional feature extraction network to obtain multi-layer convolutional features, and inputting the multi-layer convolutional features into a spatial semantic feature enhancement network to obtain spatial semantic features.

[0018] Furthermore, the convolutional feature extraction network is ResNet-50 or ResNet-101.

[0019] Furthermore, the spatial semantic feature enhancement network is a multi-level feature pyramid network FPN.

[0020] Furthermore, step four specifically includes:

[0021] The original regression feature X reg Input to the convolution layer to generate a random offset P;

[0022] The random offset P and the original classification feature X cls Perform random sampling to obtain the aligned classification features X align :

[0023]

[0024] Among them, m is the number of convolution sampling points, p i Indicates the location of the current convolution kernel center point, R is the set of sampling positions of the convolution on the feature map, p m For each position on R, Δp m Indicates p m The offset learned by the position, w(p m ) represents the convolution kernel p m The weight of the position;

[0025] The original classification feature X cls and aligned categorical features X align Fusion to obtain optimized classification features

[0026]

[0027] Among them, α represents the original classification feature coefficient.

[0028] Furthermore, step five specifically includes:

[0029] The general distribution approximation model for the detection box coordinates is defined as Among them, y i The distance between the feature point position of the current detection frame and the detection frame boundary is i, P() is the probability density function, and n represents the number of discrete values ​​of the general distribution;

[0030] According to the general distribution approximation model of the detection frame coordinates, the original regression feature X reg Input into a layer of convolutional network to obtain a feature map;

[0031] Input the feature map into the Softmax activation function to obtain the general distribution of the detection box coordinates

[0032] The general distribution of the detection box coordinates Input into the mathematical expectation calculation module to obtain the determined value y of the detection frame coordinate dtrmd .

[0033] Furthermore, step six specifically includes:

[0034] The original regression feature X reg and optimizing classification features Splicing is performed on the channel dimension to obtain the fusion feature X concat ;

[0035] The fused feature X concat Input into a convolutional network to generate the detection box position sampling offset O;

[0036] The detection frame position sampling offset O and the determined value y of the detection frame coordinates dtrmd Input into a deformable convolutional network to obtain the optimized detection box coordinate r refine :

[0037]

[0038] Among them, j represents the sequence of the current deformable convolution sampling points, l represents the number of deformable convolution sampling points, r represents the original prediction box coordinate value, x and y represent the horizontal and vertical coordinates of the current point respectively, Δx j and Δy j They represent the horizontal and vertical offsets of the current point respectively, and k represents the number of the detection box coordinates.

[0039] Furthermore, step seven specifically includes:

[0040] The classification features will be optimized Input into a convolutional neural network to obtain a logical operator, and input the logical operator into the Sigmoid activation function to obtain the category score;

[0041] From the general distribution of detection box coordinates Extract the three largest probability values, mean, and variance and input them into the probability guidance module to obtain the position quality estimate;

[0042] Multiply the location quality estimate by the class score to obtain the optimized class score.

[0043] Furthermore, step eight trains the target detection model based on the classification loss function FocalLoss and the regression loss function GIoULoss, specifically including:

[0044] The input of the classification loss function FocalLoss is the target category annotated by the training data, the category score annotated by the training data, and the optimized category score obtained in step 7;

[0045] The input of the regression loss function GIoULoss is the detection box coordinates of the training data and the optimized detection box coordinates r refine .

[0046] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0047] 1. To address the spatial uncertainty in target detection tasks, this paper designs a detection box coordinate modeling module based on general distribution. The detection box coordinates are modeled in the form of probability distribution. This can capture boundary uncertainty information in complex scenes, significantly improve the position quality of the detection box, alleviate the spatial uncertainty problem in target detection tasks, and improve the accuracy of detection box position prediction.

[0048] 2. In response to the semantic uncertainty in target detection tasks, the present invention designs an adaptive feature alignment module based on random sampling, which can adaptively learn the optimal offset for each sampling position of the convolution operation, align all feature points on the entire feature map, align classification features, improve the accuracy of category prediction, alleviate the semantic uncertainty problem of target detection tasks, and improve category prediction accuracy.

[0049] 3. To address the spatial prediction misalignment problem of target detectors, this paper designs a prediction box weighted averaging module based on random sampling, which uses the higher-quality surrounding detection box coordinates to optimize the detection box coordinates of the current feature point, improve the position quality of the detection box, and enhance the model accuracy.

[0050] 4. To address the quality representation problem of detection frames in target detection tasks, the present invention designs a probability guidance module that uses the position information contained in the general distribution of detection frame coordinates to obtain position quality estimation, thereby optimizing the quality representation of the detection frame and improving model accuracy.

[0051] In summary, the method proposed in this paper can improve the detection accuracy in complex scenes, predict high-quality detection frames, and provide more accurate location information for downstream decision-making of target detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A schematic diagram of a target detection method based on random uncertainty provided by an embodiment of the present invention;

[0053] Figure 2 A diagram of a detector network structure based on random uncertainty provided by an embodiment of the present invention;

[0054] Figure 3 A structural diagram of an adaptive feature alignment module based on random sampling provided by an embodiment of the present invention;

[0055] Figure 4A network diagram of a detection frame coordinate modeling module based on general distribution provided by an embodiment of the present invention;

[0056] Figure 5 A structural diagram of a prediction box weighted averaging module based on random sampling provided by an embodiment of the present invention;

[0057] Figure 6 This is a diagram of the target category prediction network structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0058] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. Before describing in detail the technical solutions of each embodiment of the present invention, the nouns and terms involved are explained. In this specification, components with the same name or the same number represent similar or identical structures and are for illustrative purposes only.

[0059] Aiming at the spatial uncertainty and semantic uncertainty problems existing in target detection scenarios in current practical applications, this paper proposes a high-performance target detection method based on random uncertainty, such as Figure 1 As shown, the method constructs a target detection model, inputs the image to be identified into the target detection model, and outputs the category and coordinates of the object in the image. The target detection model includes a feature extraction network and a detector network based on random uncertainty. Based on the theory of random uncertainty, the present invention introduces probability distribution and random sampling methods into the detector network based on random uncertainty. The method includes the following steps:

[0060] 1. Data Preparation

[0061] 1.1 Dataset Preparation

[0062] First, image data in autonomous driving scenarios is collected. Then, pedestrians, cars, bicycles, traffic lights, trees and other objects are labeled with target categories and category scores, and detection box coordinates. Category score annotation refers to the confidence level that the object in the image data belongs to the labeled category. Finally, the training set and test set are divided into a ratio of 9:1, which are used for model training and testing respectively.

[0063] 1.2 Data Augmentation

[0064] Data augmentation is required for input images during training and testing to adapt to the varying input image sizes in real-world scenarios and enhance the robustness of the object detection model. During training, multi-scale image augmentation is used to randomly scale the short side to a range of 480-960 pixels, and the long side is scaled proportionally to a maximum of 1333 pixels. One scale is then input into the object detection model for training. During testing, multi-scale image augmentation is still used, with the short side scaled to 480, 600, 720, 840, and 960 pixels, respectively, and the long side scaled proportionally to a maximum of 1333 pixels. Images at all five scales are fed into the object detection model, and the model's detection results are then fed simultaneously into the post-processing module to yield the final test results.

[0065] 2. Extracting spatial semantic features of the input image

[0066] like Figure 1 As shown in FIG, the image that has been data-enhanced in step 1.2 is input into the feature extraction network to extract the spatial semantic features of the input image.

[0067] The feature extraction network includes a skeleton network and a spatial semantic feature enhancement network. First, the image is input into the skeleton network to obtain the multi-layer convolutional features of the image. The skeleton network uses commonly used convolutional feature extraction networks such as ResNet-50 and ResNet-101; then the above multi-layer convolutional features are input into the spatial semantic feature enhancement network to obtain the spatial semantic features of the image. The spatial semantic feature enhancement network is a multi-level feature pyramid network FPN, which is used to obtain low-level spatial features and high-level semantic features of the image and optimize the detection accuracy of multi-scale targets.

[0068] 3. Obtaining target detection results based on random uncertainty

[0069] like Figure 1 As shown in FIG, the spatial semantic features obtained in step 2 are input into the detector network based on random uncertainty to obtain the detection results, which include the target category and its score, and the detection box coordinates.

[0070] like Figure 2 As shown in the figure, the detector network based on random uncertainty includes a classification branch and a regression branch. The classification branch includes a classification branch feature extraction network, an adaptive feature alignment module based on random sampling, a probability guidance module, and a target category prediction network. The regression branch includes a regression branch feature extraction network, a detection box coordinate modeling module based on general distribution, and a prediction box weighted average module based on random sampling.

[0071] 3.1 Extracting classification features and regression features

[0072] like Figure 2As shown in the figure, the spatial semantic features obtained in step 2 are used as the input feature map, which are input into the classification branch feature extraction network and the regression branch feature extraction network respectively to output the original classification feature X cls With the original regression feature X reg ,The classification branch feature extraction network and the regression branch feature extraction network are two parallel branches, both composed of four-layer convolutional networks.

[0073] 3.2 Optimizing Classification Features Based on Random Sampling

[0074] like Figure 2 As shown, the original classification feature X cls With the original regression feature X reg As input, it is also input into the adaptive feature alignment module based on random sampling, and the output is the optimized classification feature Using the original regression feature X reg The adaptive feature alignment module based on random sampling consists of a layer of ordinary convolutional network, a layer of deformable convolutional network and a point-to-point accumulation operation.

[0075] The specific process is as follows Figure 3 As shown, it is divided into three steps:

[0076] First, the original regression feature X reg Input into the ordinary convolution to generate a random offset P = Conv(X reg ), where Conv(*) represents a normal convolution operation;

[0077] Next, the random offset P and the original classification feature X cls Input into deformable convolution for random sampling operation to obtain aligned classification features Wherein, m is the number of convolution sampling points, which is set to 9 in the embodiment of the present invention, and p i Indicates the location of the center point of the current convolution kernel, R is the sampling position set of ordinary convolution on the feature map, p m For each position on R, Δp m Indicates p m The offset learned by the position, w(p m ) represents the convolution kernel p m The weight of the position;

[0078] Finally, the original classification feature X cls and aligned categorical features X align Fusion to obtain optimized classification features It is used for the final classification task, where α represents the original classification feature coefficient, which is set to 0.3 in the embodiment of the present invention.

[0079] 3.3 Modeling Detection Box Coordinates Based on General Distribution

[0080] like Figure 2 As shown, the original regression feature X reg The input is sent to the detection box coordinate modeling module based on general distribution, and the general distribution of the detection box coordinates and the determined value of the detection box coordinates are output in sequence. Modeling the detection box coordinates in the form of probability distribution can capture the boundary uncertainty information in complex scenes, which is used to alleviate the spatial uncertainty problem of the target detection task. The detection box coordinate modeling module based on general distribution consists of a layer of ordinary convolutional network, Softmax activation function and mathematical expectation calculation module.

[0081] The specific process is as follows Figure 4 As shown, it is divided into three steps:

[0082] (1) Modeling detection box coordinates based on general distribution

[0083] The general distribution model of the detection box coordinates is expressed as Among them, P() is the probability density function of the general distribution, q represents the distance from the feature point position of the current detection frame to the detection frame boundary, and then the minimum value y0 of the upper and lower limits of the integral is set to 0 and the maximum value y n is 16, n is the number of general distribution discrete values, n = 16, and finally the continuous random distribution is discretized. The interval of the discrete random variable is 1, then the number of discrete values ​​is n + 1, and the discrete values ​​are {0, 1, 2, ..., 16}. The general distribution approximate model of the detection box coordinates is expressed as Among them, y i Indicates that the distance from the feature point position of the current detection frame to the detection frame boundary is i, P(y i ) represents the discrete value y i The corresponding probability;

[0084] (2) General distribution of output detection box coordinates

[0085] According to the general distribution approximate expression of the detection frame coordinates derived in step (1), first the original regression feature X reg Input into a convolutional network to obtain a feature map with n+1 output channels, and then input the feature map into the Softmax activation function to obtain the general distribution of the detection box coordinates. The Softmax activation function is used to ensure that the sum of the probabilities of all discrete random variables predicted by the network is 1, that is,

[0086] (3) Calculate the mathematical expectation of the general distribution to obtain the determined value of the detection frame coordinates

[0087] The general distribution of the detection frame coordinates obtained in step (2) is input into the mathematical expectation calculation module to obtain the determined value y of the detection frame coordinates dtrmd , the mathematical expectation is approximately calculated according to the general distribution approximation model representation formula of the detection box coordinates in step (1).

[0088] 3.4 Optimizing detection box coordinates based on random sampling

[0089] like Figure 2 As shown, the original regression feature X reg , optimize classification features The determined value y of the detection box coordinate dtrmd , input to the prediction box weighted average module based on random sampling, and output the optimized detection box coordinates; using the optimized classification features The target category information and original regression features X contained in reg The position-related information contained in adaptively captures the higher-quality surrounding detection box coordinates of each feature point, and then optimizes the detection box coordinates of the current feature point to alleviate the problem of spatial prediction misalignment; the prediction box weighted average module based on random sampling consists of a splicing fusion operation, a layer of ordinary convolutional network, and a layer of deformable convolutional network.

[0090] The specific process is as follows Figure 5 As shown, it is divided into three steps:

[0091] First, the original regression feature X reg and optimizing classification features Splicing is performed on the channel dimension to obtain fusion features concat(*) represents the feature map concatenation operation;

[0092] Then the fusion feature X concat Input into a layer of ordinary convolutional network to generate the detection box position sampling offset O = Conv(X concat );

[0093] Finally, the detection frame position sampling offset O and the determined value y of the detection frame coordinates are dtrmd Input into a deformable convolutional network, perform weighted averaging on the high-quality detection frames around the current detection frame, and obtain the optimized detection frame coordinates Among them, j represents the sequence of the current deformable convolution sampling points, l represents the number of deformable convolution sampling points, r represents the original prediction box coordinate value, x and y represent the horizontal and vertical coordinates of the current point respectively, Δx j and Δy jThey represent the horizontal and vertical offsets of the current point respectively, k represents the number of the detection frame coordinates, and in this embodiment of the present invention, the set of k is {0, 1, 2, 3}.

[0094] 3.5 Optimizing category scores using the general distribution of detection boxes

[0095] like Figure 2 As shown, the classification features will be optimized and the general distribution of detection box coordinates Input into the target category prediction network and output the optimized category score; use the general distribution of detection box coordinates The position information contained in is used to obtain the position quality estimation, and then the quality representation of the detection frame is optimized to obtain a more accurate detection frame quality representation, which is used to alleviate the phenomenon that high-quality detection frames are mistakenly deleted in the post-processing process of target detection; the target category prediction network consists of a layer of ordinary convolutional network, Sigmoid activation function, probability guidance module and point multiplication operation.

[0096] The specific process is as follows Figure 6 As shown, it is divided into three steps:

[0097] First, optimize the classification features Input into a convolutional neural network to get the logical operator, and then input the logical operator into the Sigmoid activation function to get the category score;

[0098] Then from the general distribution of detection box coordinates The three largest probability values, mean, and variance are extracted as one-dimensional statistics and input into the probability guidance module to obtain the position quality estimate. The probability guidance module consists of a fully connected layer and a Sigmoid activation function.

[0099] Finally, the location quality estimate is multiplied by the category score to obtain the final optimized category score.

[0100] In summary, the output of the current model is the optimized detection box coordinate r predicted in step 3.4 refine , and the final optimized class scores predicted in step 3.5.

[0101] 4. Loss Function

[0102] The loss function of the present invention consists of two parts, namely the classification loss function FocalLoss for training the classification branch and the regression loss function GIoULoss for training the regression branch. The above loss functions are all universal loss functions in target detection tasks. The input of the classification loss function FocalLoss is the target category and its score marked in step 1.1 and the category score finally optimized in step 3.5. The input of the regression loss function GIoULoss is the detection box coordinates marked in step 1.1 and the optimized detection box coordinates predicted in step 3.4. refine .

[0103] 5. Train the target detection model based on the loss function

[0104] like Figure 1 As shown in the figure, during the training process, the input image is first enhanced according to the data enhancement method of the training link described in step 1.2, and then the data-enhanced image is input into the feature extraction network and the detector network based on random uncertainty to obtain the prediction results of the model, namely the category score and the detection box coordinates; then the loss function described in step 4 is used to calculate the loss of the model, and the SGD optimizer is used to update the network parameters. The initial learning rate is set to 0.01, the momentum is set to 0.9, and the Warmup training strategy is used. After 12 cycles of training, a trained target detection model is obtained.

[0105] 6. Putting the model into practical use

[0106] like Figure 1 As shown, during the test process, the input image is first enhanced according to the test link data enhancement method described in step 1.2, and then input into the trained target detection model to obtain the model's prediction results, i.e., category scores and detection frame coordinates; finally, after target detection post-processing, redundant detection frames are removed to obtain accurate category scores and high-quality detection frame positions. The commonly used post-processing model NMS for target detection is used, and the IoU threshold is set to 0.6 in this embodiment of the present invention.

[0107] The above-described embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A target detection method based on random uncertainty constructs a target detection model. The image to be identified is input into the target detection model, and the class and coordinates of the object in the image are output. The training of the target detection model includes the following steps: Step 1: Prepare image data for target category and category score annotation, detection frame coordinate annotation, and pre-process the annotated images as training data; Step 2: Input the training data into the feature extraction network to extract its spatial semantic features; Step 3: Input the spatial semantic features into the classification branch feature extraction network and the regression branch feature extraction network to obtain the original classification features. With the original regression features ; Step 4: The original classification features With the original regression features Input to the adaptive feature alignment module to obtain optimized classification features , specifically including: The original regression features Input to the convolution layer to generate random offsets ; The random offset and the original classification features Perform random sampling operations to obtain aligned classification features : ; in, is the number of convolution sampling points, Indicates the location of the current convolution kernel center point, is the set of sampling positions of the convolution on the feature map, express For each position on express The offset learned by the position, Represents the convolution kernel The weight of the position; The original classification features and aligned categorical features Fusion to obtain optimized classification features : ; in, represents the original classification feature coefficient; Step 5: Based on the original regression features Calculate the general distribution of detection box coordinates and the determined values ​​of the detection frame coordinates ; Step 6: The original regression features , optimize classification features , the determined value of the detection frame coordinates , input to the prediction box weighted average module to obtain the optimized detection box coordinates ; Step 7: Optimize classification features and the general distribution of detection box coordinates Input to the target category prediction network to obtain the optimized category score; Step 8: Train the target detection model according to the classification loss function FocalLoss and the regression loss function GIoULoss until the preset training completion conditions are met.

2. The target detection method based on random uncertainty according to claim 1, characterized in that: Step 2 specifically includes inputting the training data into a convolutional feature extraction network to obtain multi-layer convolutional features, and inputting the multi-layer convolutional features into a spatial semantic feature enhancement network to obtain spatial semantic features.

3. The target detection method based on random uncertainty according to claim 2, characterized in that: The convolutional feature extraction network is ResNet-50 or ResNet-101.

4. The target detection method based on random uncertainty according to claim 2, characterized in that: The spatial semantic feature enhancement network is a multi-level feature pyramid network FPN.

5. The target detection method based on random uncertainty according to claim 1, characterized in that: Step 5 specifically includes: The general distribution approximation model of the detection box coordinates is defined as ,in, Indicates that the distance from the feature point position of the current detection frame to the detection frame boundary is , is the probability density function, Represents the number of discrete values ​​of a general distribution; According to the general distribution approximation model of the detection frame coordinates, the original regression features Input into a layer of convolutional network to obtain a feature map; Input the feature map into the Softmax activation function to obtain the general distribution of the detection box coordinates ; The general distribution of the detection box coordinates Input into the mathematical expectation calculation module to get the determined value of the detection frame coordinates .

6. The target detection method based on random uncertainty according to claim 1, characterized in that: Step six specifically includes: The original regression features and optimizing classification features Splicing is performed on the channel dimension to obtain fusion features ; The fusion features Input into a convolutional network to generate the detection box position sampling offset ; Sampling the detection frame position offset and the determined values ​​of the detection frame coordinates Input into a deformable convolutional network to obtain the optimized detection box coordinates : ; in, Represents the sequence of current deformable convolution sampling points, Represents the number of deformable convolution sampling points, Represents the original prediction box coordinate value, and Respectively represent the horizontal and vertical coordinates of the current point, and Respectively represent the horizontal and vertical offsets of the current point. A number indicating the coordinates of the detection box.

7. The target detection method based on random uncertainty according to claim 1, characterized in that: Step seven specifically includes: The classification features will be optimized Input into a convolutional neural network to obtain a logical operator, and input the logical operator into the Sigmoid activation function to obtain the category score; From the general distribution of detection box coordinates Extract the three largest probability values, mean, and variance and input them into the probability guidance module to obtain the position quality estimate; Multiply the location quality estimate by the class score to obtain the optimized class score.

8. The target detection method based on random uncertainty according to claim 1, characterized in that: Step 8: Training the target detection model based on the classification loss function FocalLoss and the regression loss function GIoULoss, specifically including: The input of the classification loss function FocalLoss is the target category annotated by the training data, the category score annotated by the training data, and the optimized category score obtained in step 7; The input of the regression loss function GIoULoss is the detection box coordinates of the training data annotation and the optimized detection box coordinates .

Citation Information

Patent Citations

  • Image target detection optimization method and device, electronic equipment and storage medium

    CN111860494A

  • Ship target detection method based on adaptive feature extraction and decoupling prediction head

    CN115272701A