A Novel Brain-Inspired Target Detection Method for Occluded Targets in Complex Environments

By combining the brain-inspired model and the DETR model, the bilateral neural circuit online learning and reference point aggregation convergence method are used to solve the problems of insufficient accuracy of occluded target detection and slow training speed in complex environments, and more efficient object detection and faster training process are achieved.

CN116935196BActive Publication Date: 2025-06-13CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310963809.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2025-06-13
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

The existing neural networks have insufficient detection accuracy and slow training speed for occluded targets in complex environments.

Method used

A new brain-inspired object detection method for obstructed targets in complex environments is adopted, combined with brain-inspired model and DETR model, and the detection accuracy and training efficiency are improved through bilateral neural circuit online learning and reference point aggregation convergence method.

Benefits of technology

It improves the detection accuracy of the obstructed target in complex environments, shortens the training time, and enhances the model's adaptability and adjustment speed in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935196B_ABST
    Figure CN116935196B_ABST
Patent Text Reader

Abstract

The present invention relates to a novel brain-inspired target detection method for occluded targets in complex environments, which includes constructing a target detection model. The target detection model includes a brain-inspired model and a DETR model. The image to be detected is input into the brain-inspired model to obtain several prediction boxes and class labels for the target, and then input into the DETR model. The DETR model preset reference points for the image data, calibrates the side distance and box position of the prediction box with respect to the reference points, and gives query objects to achieve class label matching through bipartite matching. By calculating the side distance and box position of the prediction box and iterating until it is within the set offset amplitude threshold range, then the category and coordinates corresponding to the highest confidence of the target in the prediction box are selected. In the present invention, the brain-inspired model can continuously extract knowledge during learning to better and faster learn new tasks, interpret and store the feature data of existing training samples and deconstruct and release them, so that the chaotic data stream can be simulated as a stable dynamic data stream, alleviating the forgetting problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly relates to an intelligent detection method for occluded targets. Background Art

[0002] In recent years, with the rapid development of artificial intelligence, object detection has become one of the key research directions in the field of computer vision. Since the proposal of object detection methods such as DETR single-stage, in order to solve the problems of insufficient detection accuracy and slow training speed for occluded targets existing in the working process of its network structure, researching advanced neural network training methods, optimizing the network structure and training learning environment are the main ways to improve its detection accuracy and learning efficiency.

[0003] Due to the fact that in complex environments, such as when there are occluders in the environment where the detection target is located, the object detection technology still has great difficulties. Neuroscience has been playing an indispensable role in the development of artificial intelligence. The memory mechanism of the brain has extremely high learning efficiency for the feature memories extracted during the learning process, especially for situations where there are many feature points in occluded scenes. And during the rapid learning process, organisms can continuously extract knowledge during learning to better and faster learn new tasks without forgetting. The brain-inspired neural network learning method interprets and stores the feature data of existing training samples and deconstructs and releases them, enabling chaotic data streams to be simulated as stable dynamic data streams, thereby alleviating the forgetting problem. At present, there are a large number of relevant literature studies at home and abroad on applying brain-inspired and optimization learning methods to object detection.

[0004] 1. In the article titled "Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object Detection", the authors Ghosh A et al. aimed to solve the problem of insufficient detection ability for small and distant objects. They proposed a training prior method of two-plane perspective prior, which combines rough geometric constraints from 3D scene interpretations of 2D images, improving the detection performance of object detection for small and distant objects. This method implements end-to-end learning and can improve the performance of the detector at a lower scale. However, for complex scenes, its vanishing point estimation is difficult to be accurate. Due to the overly large computational amount of its geometric constraints and the strict requirements of the method for the image perspective, this method is only applicable to static scenes such as cameras and has poor generalization ability.

[0005] 2. In the article titled "Semi-DETR: Semi-Supervised Object Detection with Detection Transformers", the authors Hua W et al. proposed Semi-DETR, an end-to-end semi-supervised object detector based on the transformer structure and a cost-based pseudo-label module, to solve the problems of incorrect matching and long training time during the learning process of the DETR network. The detector includes a phased hybrid matching learning method that combines the advantages of one-to-many assignment and one-to-one assignment strategies, solves the problem of incorrect matching during the training process, improves the training efficiency, and ensures the applicability of consistent regularization. However, while ensuring the learning efficiency, this method cannot enhance the detection ability of the network structure for occluded targets and the generalization ability for different environments.

[0006] 3. In the article titled "DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection", the authors Jiawei Ma et al. proposed a new DiGeo training decision method to solve the problems of insufficient discriminative feature learning and lack of generalization ability for each category in existing networks, for learning geometry-aware features of between-class separation and within-class compactness. To tighten the clustering of each class, this method incorporates an adaptive class-specific margin into the classification loss and encourages features close to the class center. This learning method effectively improves the generalization to new classes without affecting the detection accuracy of the base classes. However, this training learning method has poor learning efficiency for small samples and occluded targets, and it is difficult to avoid the problem of feature data forgetting during the learning process. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a new brain-inspired object detection method for occluded targets in complex environments, to solve the technical problems such as insufficient detection accuracy and slow training speed of existing neural networks for occluded targets in complex environments.

[0008] The new brain-inspired object detection method of the present invention for occluded targets in complex environments includes the following steps:

[0009] 1) Construct a target detection model, and the target detection model includes a brain-inspired model and a DETR model;

[0010] The brain-inspired model is composed of an interacting generator, discriminator, and classifier. Train the brain-inspired model to determine the hyperparameters of the brain-inspired model and the internal parameters of the discriminator and classifier, and obtain a well-trained generator, discriminator, and classifier;

[0011] 2) Input the image to be detected into the generator of the brain-inspired model. During the target detection process, online update the internal parameters of the discriminator through online learning in the bilateral neural circuit, and determine whether to adopt the output of the discriminator or the output of the classifier based on the confidence judgment. Then, decode the adopted output through the prediction box decoder to obtain several prediction boxes and class labels of the target, and input the prediction boxes and class labels of the target into the DETR model;

[0012] 3) The DETR model preset reference points for the image data, calibrate the lateral distance and box position of the prediction box with respect to the reference points, and give query objects to achieve class label matching through bipartite matching;

[0013] 4) The DETR model calculates the lateral distance and box position of the prediction box to obtain the offset of the prediction box, and iterates between steps 2) and 4). According to the obtained box prediction coordinates, iterate until within the set offset amplitude threshold range, and select the category corresponding to the highest score among the target confidence levels in the prediction boxes. This category is the category to which the target belongs, and the corresponding coordinates are the position information of the target in the image.

[0014] Furthermore, training the brain-inspired model in step 1) includes:

[0015] Before training the current task t, the generator G generates the generated dataset of the old task x 0:t-1 、z 0:t-1 and construct the replay dataset S t as the training dataset for task t;

[0016] During the training of task t: The generator G randomly generates data labeled c using the random noise vector z of the current task t z is Gaussian p z = N(0,1), and the sampling distribution of all Y t classes in the task is uniform p c = U{1,Y t}; Both the discriminator D and the classifier C receive the replay dataset S′; The input of the discriminator D is x, and the discriminator D models the probability P(y|x) of the output being y, constructs an auxiliary classifier D′, uses the discriminator D to be responsible for discriminating false / true examples, uses the auxiliary classifier D′ to be responsible for predicting the class label, and transforms the probability problem into optimizing the internal parameters {θ D ,θ D′ ,θ C ,θ G} of the discriminator and the classifier;

[0017] To stabilize the training process, the loss function L of the auxiliary classifier D′The loss function L of the auxiliary classifier follows the following formula D′ consists of a cross-entropy term and a regularization term, and the expression is as follows:

[0018]

[0019]

[0020] where is the empirical parameter learned by the auxiliary classifier from the old task, represents the binary distribution operation of (x, y c ) with respect to S′, y c represents the number of class labels in the classifier, λ D′ represents the integrated impedance parameter in the auxiliary classifier, L CE (θ D′ ) uses cross-entropy operation and is calculated from the results of the auxiliary classifier classifying the true labels of the training data of the current task and the data generated by the previous task; D′(x) represents is the class distribution probability parameter; is the regularization term;

[0021] During the training process, the loss functions L D 、L C and θ D 、θ C are updated as follows:

[0022]

[0023]

[0024]

[0025]

[0026] D(G(z, c)) represents the false / true probability parameter of the data generated by the generator G in the discriminator D. In formulas (1) and (4), F represents the Fisher information integration operation, and the subscripts of F represent the operation environment and serial number respectively. The loss function L of the classifier C includes a cross-entropy term and a regularization term. Another expression format of the cross-entropy term is The cross-entropy term minimizes the difference between P C (y|x) and P D′ (y|x) on the replay dataset S′, and transfers the empirical knowledge from D′ to C. Because the regularization term in formula (2) penalizes the gap between P D′ (y|x) and P l (y|x), Pl (y|x) is the parameter of the l-th layer for auxiliary classifier regularization;

[0027] The loss function of the generator G is updated as follows:

[0028]

[0029]

[0030]

[0031]

[0032] Among them, R M is the sparse regularizer of the attention mask, is the attention weight of the l-th layer on task t, initialized to 0.5, where s is the proportionality factor, is the mask embedding matrix, σ is the sigmoid function, where N i is the number of parameters of the l-th layer;

[0033] During the training process, the prediction made by the discriminator D or the classifier C is determined in the following way:

[0034] D′ first estimates P D′ (y = k|x) and P C (y = k|x), the probability that a certain input data x belongs to a certain class in class k, where [k] = {0, 1,..., k - 1}; then the prediction result to be adopted is determined according to the confidence, and the confidence formula is as follows:

[0035]

[0036] Furthermore, in step 2), the internal parameters of the discriminator are updated online through bilateral neural circuit online learning. The bilateral neural circuit online learning consists of internal neural circuit training and external neural circuit training. The internal neural circuit training is multi-mini-batch gradient descent training, and the external neural circuit training is multi-batch maintenance training, including:

[0037] Randomly extract n old training samples from the extra storage space and mix them with the new samples received during the object detection process, and train through multiple mini-batch gradient descents to update the internal parameters of the discriminator D; introduce an attention coefficient s to adjust the learning rate of the bilateral neural circuit model on the new samples, so as to adjust the attention of the bilateral neural circuit model to the new samples according to the number of replay training samples; for the outer neural circuit, that is, the multi-batch new data received by the model, for a single batch of data, based on the parameter update of the bilateral neural circuit model on multiple mini-batches constructed in the inner neural circuit, use the meta-learning rate parameter λ for θ D Perform the final empirical update; the internal parameters θ of the discriminator D The update formula is as follows:

[0038]

[0039]

[0040] Among them, i and k represent the serial numbers of the batches, j and l represent the serial numbers of the mini-batches. The mini-batch is the inner neural circuit, and the batch is the outer neural circuit. And B and b respectively represent the numbers of the batch and the mini-batch, L D (x ij ,y ij ) represents the loss obtained in the brain-inspired learning model. Q is the inner product between the gradients of different samples of the bilateral neural circuit model. When Q is greater than zero, it means that migration occurs between the training samples.

[0041] Furthermore, in step 3), the DETR model preset reference points for the image data, calibrated the lateral distance and box position of the prediction box with respect to the reference points, and gave query objects to achieve class label matching through bipartite matching; specifically including:

[0042] Flatten the image data into a fixed grid area, initialize the upper left corner of the image as the reference point, and the 2D coordinates of the reference point r = {x, y} ∈ [0, 1] 2 , and instantiate a learnable 4D offset distance s = {l, t, r, b} ∈ [0, 1] from the reference point to the prediction box for each object query 4 , l, t, r, b respectively represent the distances from the reference point to the four sides of the prediction box, in the order of upper left, upper right, lower left, and lower right. The object query reference is q = {e, r, s}, where e ∈ R d is the content embedding of d dimensions; the model directly supervises the four-dimensional offset of the four sides of the prediction box to the reference point; the final box prediction formula is: They are the corner coordinates of the corresponding direction box and its four-dimensional offset respectively; after fixing the reference point {x, {y}} = {x, y}, prediction updates are performed layer by layer, and the prediction update calculation for each decoder layer is as follows:

[0043] ΔS l = BoxHead l (s l-1 , e l-1 , r l-1 ),

[0044]

[0045]

[0046] where σ and σ -1 are the sigmoid and inverse sigmoid function operations respectively; Δs l represents the lateral offset prediction; and are the lateral distance, reference point, and box position predicted from the l-th layer decoder respectively; BoxHead l is the prediction head after the layer-l decoder, and in the model setting, it is independent between different decoder layers;

[0047] During the detection process, each query is only allowed to predict the bounding boxes that overlap with its reference region, and this rule is applied to the bipartite matching process through the internal matching cost Linner; given N query objects Q = {q 1 , q 2 , ···, q n} and M ground truth objects G = {g 1 , g 2 , ···, g M}, the step function L inner (g i , q i ) penalizes the reference point r i of q i when r i is outside the prediction box of g i , and the penalty value is the penalty cost k, with a default value of 95; the final permutation formula for label assignment is:

[0048]

[0049]

[0050] where L match is the original pairwise matching cost, including the classification cost and the localization cost; is the bipartite matching permutation set of N elements.

[0051] Further, the offset amplitude threshold in step 4) is s grid , and the side distance {0~100} and the box position of the prediction box b are calculated through PointHead based on point-by-point features and s-type function activation to obtain the predicted offsets Δr i ′ and Δr l . The movement update calculation of the reference point is as follows:

[0052] If then Δr k ′ = PointHead k (s l-1 , e l-1 , r k-1 ),

[0053]

[0054]

[0055] where, Δr l ′ and Δr k are the predicted offsets from r k to r k-1 and r 0 respectively before the s-type function activation σ. is the confidence of the feature point, is the detach function in pytorch.

[0056] Advantages of the present invention:

[0057] 1. The present invention is a novel brain-inspired object detection method for occluded objects in complex environments. The brain-inspired method is applied to the training and learning method of the neural network model. A brain-inspired model is designed by emulating the brain memory mechanism. It has extremely high learning efficiency for the feature memories extracted during the learning process, especially for the case where there are many feature points in the occlusion scene. And during the fast learning process, the brain-inspired model can continuously extract knowledge during learning to better and faster learn new tasks without forgetting. The brain-inspired model interprets and stores the feature data of the existing training samples and deconstructs and releases them, so that the chaotic data stream can be simulated as a stable dynamic data stream, thereby slowing down the forgetting problem.

[0058] 2. The present invention is a novel brain-inspired object detection method for occluded objects in complex environments. It designs a novel reference point aggregation and convergence method, which can effectively accelerate the convergence speed time. The aggregation function has optimized parameter settings, and compared with the traditional convergence mode used, it has achieved a faster response speed and accurate prediction effect.

[0059] 3. To address the problem of insufficient generalization ability of the DEPT model in complex environments, the present invention designs a bidirectional neural circuit online learning method. The bidirectional neural circuit online learning method can smooth and stabilize the disordered feature data stream, thereby improving the adaptability and adjustment speed of the model in different environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram of the brain memory system;

[0061] Figure 2 It is a schematic diagram of the architecture and training of the brain-inspired model;

[0062] Figure 3 It is a schematic diagram of the DETR network structure;

[0063] Figure 4 It is a schematic diagram of the detection process of occluded targets in complex environments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The present invention will be further described below with reference to the drawings and embodiments.

[0065] The novel brain-inspired target detection method for occluded targets in complex environments in this embodiment includes the following steps:

[0066] 1) Construct a target inspection model, and the target detection model includes a brain-inspired model and a DETR model.

[0067] The brain-inspired model is composed of an interacting generator, discriminator, and classifier. Train the brain-inspired model to determine the hyperparameters of the brain-inspired model and the internal parameters of the discriminator and classifier, and obtain a well-trained generator, discriminator, and classifier.

[0068] 2) Input the image to be detected into the generator of the brain-inspired model. During the target detection process, online update the internal parameters of the discriminator through bilateral neural circuit online learning, and determine whether to adopt the output of the discriminator or the output of the classifier through confidence judgment. Then, decode the adopted output through the prediction box decoder to obtain several prediction boxes and class labels of the target, and input the prediction boxes and class labels of the target into the DETR model.

[0069] 3) The DETR model preset reference points for the image data, calibrates the lateral distance and box position of the prediction box with respect to the reference points, and gives query objects to achieve class label matching through bipartite matching.

[0070] 4) The DETR model obtains the offset of the predicted bounding box by calculating the lateral distance and box position of the predicted bounding box, iterates between steps 2) and 4), iterates to within the set offset amplitude threshold range according to the obtained box prediction coordinates, selects the category corresponding to the highest score of the object confidence in the predicted bounding box, and this category is the category to which the object belongs, and the corresponding coordinates are the position information of the object in the image.

[0071] During the process of brain learning and memory, the hippocampus reduces interference through neurogenesis and neural inhibition mechanisms. Neurogenesis generates additional neurons to create space for incoming experiences. Inhibitory neurons act simultaneously to inhibit irrelevant parts of the network, while the neocortex encodes more general memories and consolidates memories through gradually strengthened synaptic connections. In the neocortex, the prefrontal cortex develops a discrimination mechanism to regulate the encoding of specific memories in the hippocampus and integrates the sensory cortex in the neocortex to encode general memories. Individually consolidated memories complement each other in the interaction of the hippocampus, prefrontal cortex, and sensory cortex to avoid catastrophic forgetting. This detection method for occluded targets in complex environments constructs a brain-inspired model based on this brain memory strategy, applies the idea of generative adversarial networks, and models the functions of the hippocampus, prefrontal cortex, and sensory cortex as the combined actions of a generator, discriminator, and classifier respectively. The generator learns the distribution of the training data, similar to the specific form of memories encoded in the hippocampus. The generated data is replayed into the discriminator, and during memory replay, the classifier learns new tasks under the supervision of the discriminator. To regularize the difference between the generated data and the old training data, both the classifier and the discriminator adopt a weight consolidation algorithm. The training of the brain-inspired model in step 1) includes:

[0072] Before training the current task t, the generator G generates a generated dataset of the old tasks x 0:t-1 、z 0:t-1 and constructs a replay dataset S t as the training dataset for task t;

[0073] During the process of training task t: The generator G randomly generates data labeled c using the random noise vector z of the current task t z is Gaussian p z = N(0,1), and the sampling distribution of all Y t classes in the task is uniform p c = U{1,Y t}; Both the discriminator D and the classifier C receive the replay dataset S′; The input of the discriminator D is x. The discriminator D models the probability P(y|x) of the output being y, constructs an auxiliary classifier D′, uses the discriminator D to be responsible for discriminating false / true examples, and uses the auxiliary classifier D′ to be responsible for predicting the class label, transforming the probability problem into optimizing the internal parameters {θ D , θ D′ , θ C , θ G}.

[0074] To stabilize the training process, the loss function L D′ of the auxiliary classifier follows the following formula. The loss function L D′ of the auxiliary classifier consists of a cross-entropy term and a regularization term, and the expression is as follows:

[0075]

[0076]

[0077] In the formula, is the empirical parameter learned by the auxiliary classifier from the old task, represents the binary distribution operation of (x, y c ) with respect to S′, y c represents the number of class labels in the classifier, λ D′ represents the integrated impedance parameter in the auxiliary classifier, L CE (θ D′ ) adopts the cross-entropy operation and is calculated from the results of the auxiliary classifier classifying the true labels of the current task training data and the data generated by the previous task; D′(x) represents is the class distribution probability parameter; is the regularization term.

[0078] During the training process, the loss functions L D , L C of the discriminator D and the classifier C, and θ D , θ C are updated as follows:

[0079]

[0080]

[0081]

[0082]

[0083] D(G(z,c)) represents the false / true probability parameter of the discriminator D with respect to the data generated by the generator G. In formulas (1) and (4), F represents the Fisher information integration operation, and the subscripts of F represent the operation environment and sequence number respectively. The loss function L of the classifier C includes a cross-entropy term and a regularization term. The cross-entropy term has another expression format as The cross-entropy term minimizes the difference between P C (y|x) and P D′ (y|x) on the replay dataset S′, transferring empirical knowledge from D′ to C. Because the regularization term in formula (2) penalizes the difference between P D′ (y|x) and P l (y|x). P l (y|x) is the parameter of the l-th layer for auxiliary classifier regularization. The squared difference of the two-norm part in formulas (3) to (6) is the weight consolidation function.

[0084] The loss function of the generator G is updated as follows:

[0085]

[0086]

[0087]

[0088] where R M is the sparse regularizer of the attention mask, is the attention weight of the l-th layer on task t, initialized to 0.5, where s is the proportionality factor, is the mask embedding matrix, σ is the sigmoid function, where N i is the number of parameters of the l-th layer;

[0089] During the training process, it is decided whether to adopt the prediction made by the discriminator D or the classifier C in the following way:

[0090] D′ first estimates p D′ (y = k|x) and P C (y = k|x) for the probability that a certain input data x belongs to a certain class in class k, where [k] = {0, 1,..., k - 1}; then decides the prediction result to be adopted according to the confidence. The confidence formula is as follows:

[0091]

[0092] In step 2) of this embodiment, the internal parameters of the discriminator are updated online through bilateral neural circuit online learning. The bilateral neural circuit online learning consists of internal neural circuit training and external neural circuit training. The internal neural circuit training is multi-mini-batch gradient descent training, and the external neural circuit training is multi-batch maintenance training, including:

[0093] Randomly extract n old training samples from the extra storage space and mix them with the new samples received during the target detection process, and update the internal parameters of the discriminator D through multiple mini-batch gradient descent trainings; introduce an attention coefficient s to adjust the learning rate of the bilateral neural circuit model on the new samples, so as to adjust the attention of the bilateral neural circuit model to the new samples according to the number of replay training samples; for the external neural circuit, that is, the multi-batch new data received by the model, for a single batch of data, based on the parameter update of the bilateral neural circuit model on multiple mini-batches constructed by the internal neural circuit, use the meta-learning rate parameter λ for θ D to perform the final empirical update; the internal parameter θ of the discriminator D The update formula is as follows:

[0094]

[0095]

[0096] where i, k represent the serial numbers of batches, j, l represent the serial numbers of mini-batches, mini-batch is the internal neural circuit, batch is the external neural circuit, and B and b respectively represent the numbers of batches and mini-batches, L D (x ij ,y ij ) represents the loss obtained in the brain-inspired learning model, Q is the inner product between the gradients of the bilateral neural circuit model on different samples. When Q is greater than zero, it means that migration occurs between training samples. The bilateral neural circuit online learning method can not only make θ D approach the minimum value of "loop training", enhance the model's resistance to the forgetting problem, but also enhance the migration between samples and improve the model's generalization ability.

[0097] In step 3) of this embodiment, the DETR model preset reference points for the image data, calibrated the lateral distance and box position of the prediction box with respect to the reference points, and gave the query object to achieve class label matching through bipartite matching; specifically including:

[0098] Lay the image data flat into a fixed grid area, initialize the upper left corner of the image as the reference point, and the 2D coordinates of the reference point r = {x, y} ∈ [0, 1] 2, and instantiate a learnable 4D offset distance s = {l, t, r, b} ∈ [0, 1] from the reference point to the predicted box for each object query instance 4 , where l, t, r, and b represent the distances from the reference point to the four sides of the predicted box, in the order of top - left, top - right, bottom - right, bottom - left. The object query reference is q = {e, r, s}, where e ∈ R d is the content embedding of d dimensions; the model directly supervises the four - dimensional offset of the four sides of the predicted box to the reference point; the final box prediction formula is: are the box - corner coordinates and their four - dimensional offsets in the corresponding directions respectively; after fixing the reference point {x, {y}} = {x, y}, the prediction update is performed layer by layer, and the prediction update calculation for each decoder layer is as follows:

[0099] ΔS l = BoxHead l (s l-1 , e l-l , r l-1 ),

[0100]

[0101]

[0102] where σ and σ -1 are sigmoid and inverse sigmoid function operations respectively; Δs l represents the lateral offset prediction; and are the lateral distance, reference point, and box position predicted from the l - th layer decoder respectively; BoxHead l is the prediction head after the layer - l decoder, and in the model setting, it is independent between different decoder layers.

[0103] During the detection process, each query is only allowed to predict a predicted box that overlaps with its reference region. This rule is applied to the bipartite matching process through the internal matching cost Linner; given N query objects Q = {q 1 , q 2 , ···, q n} and M ground - truth objects G = {g 1 , g 2 , ···, g M}, the step function L inner (g i , q i ), when r i is outside the predicted box of g i , for the reference point r i of q iA penalty is imposed, and its value is the penalty cost k, with a default value of 95; the final permutation formula for label assignment is:

[0104]

[0105]

[0106] where L match is the original pairwise matching cost, including classification cost and localization cost; is the bipartite matching permutation set of N elements.

[0107] Due to the sparsity of the fixed reference points, when there are no reference points inside the detection target, small and slender detection targets may not be distinguishable. Although bipartite matching forces each object to be assigned to an object query, since the reference points are outside the specified prediction box, the positive query cannot accurately regress to the predicted values with a distance between 0 and 1 on each side. A direct solution is to adjust the positions of the reference points within the ground truth prediction box to ensure that each object can be detected by the internal reference points. However, this full-point regression inevitably expands the search space because a large number of variables are determined, resulting in the final reference points being trapped in an unexpected corner of the prediction box. To narrow the training search space, in this embodiment, the range of the offset value is restricted by scaling the offset amplitude of the points within a specific grid area, thus avoiding a large search space. In this embodiment, the offset amplitude threshold in step 4) is s grid , and the side distance {0~100} and the box position of the prediction box b are calculated through PointHead based on point-by-point features and sigmoid function activation to obtain the predicted offsets Δr l ′ and Δr l , and the movement update calculation of the reference points is as follows:

[0108] If then Δr l ′ = PointHead l (s l-1 , e l-1 , r l-1 ),

[0109]

[0110]

[0111] where, Δr k ′ and Δr l are the predicted offsets from r l to r l-1 and r 0 respectively before the sigmoid function activation σ, is the confidence of feature points, is the separation function in PyTorch.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A novel brain-inspired object detection method for occluded objects in complex environments, Characterized in that: It includes the following steps: 1) Construct an object detection model, which includes a brain-inspired model and a DETR model; The brain-inspired model is composed of an interacting generator, discriminator, and classifier. Train the brain-inspired model to determine the hyperparameters of the brain-inspired model and the internal parameters of the discriminator and classifier, and obtain a well-trained generator, discriminator, and classifier; 2) Input the image to be detected into the generator of the brain-inspired model. During the object detection process, online update the internal parameters of the discriminator through online learning of the bilateral neural circuit, and determine whether to adopt the output of the discriminator or the classifier through confidence judgment. Then, decode the adopted output through the prediction box decoder to obtain several prediction boxes and class labels of the object, and input the prediction boxes and class labels of the object into the DETR model; In this step, the internal parameters of the discriminator are updated online through online learning of the bilateral neural circuit. The online learning of the bilateral neural circuit consists of internal neural circuit training and external neural circuit training. The internal neural circuit training is multi-mini-batch gradient descent training, and the external neural circuit training is multi-batch maintenance training, including: Randomly extract n old training samples from the extra storage space and mix them with the new samples received during the object detection process, and train through multiple mini-batch gradient descents to update the internal parameters of the discriminator D; introduce an attention coefficient s to adjust the learning rate of the bilateral neural circuit model on the new samples, so as to adjust the attention of the bilateral neural circuit model to the new samples according to the number of replay training samples; for the outer neural circuit, that is, the multi-batch new data received by the model, for a single batch of data, based on the parameter update of the bilateral neural circuit model on multiple mini-batches constructed in the inner neural circuit, use the meta-learning rate parameter λ for θ D to perform the final empirical update; the internal parameters θ D The update formula is as follows: Among them, i and k represent the serial numbers of batches, j and l represent the serial numbers of mini - batches. The mini - batch is the internal neural circuit, and the batch is the external neural circuit. Also, B and b represent the numbers of batches and mini - batches respectively, and L D (x ij ,y ij ) represents the loss obtained in the brain - inspired learning model. Q is the inner product between different sample gradients of the bilateral neural circuit model. When Q is greater than zero, it indicates that migration occurs between training samples; 3) The DETR model preset reference points for the image data, calibrates the lateral distance and box position of the prediction box with respect to the reference points, and gives query objects to achieve class label matching through bipartite matching; 4) The DETR model calculates the lateral distance and box position of the prediction box to obtain the offset of the prediction box. Iterate between steps 2) and 4). According to the obtained box prediction coordinates, iterate until within the set offset amplitude threshold range, select the category corresponding to the highest score of the object confidence in the prediction box. This category is the category to which the object belongs, and the corresponding coordinates are the position information of the object in the image.

2. The novel brain-inspired object detection method for occluded objects in complex environments according to claim 1, Characterized in that: The training of the brain-inspired model in step 1) includes: Before training the current task t, the generator G generates a generated dataset for the old task x 0:t-1 , z 0:t-1 and constructs a replay dataset S t as the training dataset for task t; During the training task t: The generator G randomly generates data labeled c using the random noise vector z of the current task t H 0 and W 0 are the image length and width information respectively, z is Gaussian p z = N(0, 1), and the sampling distribution of all Y t classes in the task is uniform p c = U{1, Y t}; Both the discriminator D and the classifier C receive the replay dataset S'; The input of the discriminator D is x. The discriminator D models the probability P(y|x) with the output y, constructs an auxiliary classifier D', uses the discriminator D to be responsible for discriminating false / true examples, uses the auxiliary classifier D' to be responsible for predicting the class label, and transforms the probability problem into optimizing the internal parameters {θ D , θ D′ , θ C , θ G} of the discriminator and the classifier To stabilize the training process, the loss function L of the auxiliary classifier D′ follows the following formula. The loss function L of the auxiliary classifier D′ is composed of a cross-entropy term and a regularization term, and the expression is as follows: Wherein, is the empirical parameter learned by the auxiliary classifier from the old task, represents the binary distribution operation of (x, y c ) with respect to S′, and y c represents the number of class labels in the classifier, and λ D′ represents the integrated impedance parameter in the auxiliary classifier, is calculated by using the cross-entropy operation and is obtained from the result of classifying the true labels of the training data of the current task and the generated data of the previous task by the auxiliary classifier; D′(x) represents is the class distribution probability parameter; is the regularization term; The loss functions \(L\) of the discriminator \(D\) and the classifier \(C\) during the training process D and \(L\) C and \(\theta\) D and \(\theta\) C are updated as follows: D(G(z,c)) represents the false / true probability parameter of the data generated by the generator G in the discriminator D. In formulas (1) and (4), F represents the Fisher information integration operation, and the subscripts of F represent the operation environment and serial number respectively; The loss function of the generator G is updated as follows: where R M is the sparse regularizer of the attention mask, is the attention weight of the l-th layer on task t, where s is a proportionality factor, is the mask embedding matrix, σ is the sigmoid function, where N i is the number of parameters of the l-th layer; During the training process, the following method is used to determine whether to adopt the prediction made by the discriminator D or the classifier C: D' first estimates P D′ (y = k|x) and P C the probability that a certain input data x in (y = k|x) belongs to a certain class in class K, where [K] = {0, 1,..., k - 1}; then determines the prediction result to be adopted according to the confidence level, and the confidence level formula is as follows:

3. The novel brain-inspired object detection method for occluded objects in complex environments according to claim 1, Characterized in that: In step 3), the DETR model preset reference points for the image data, calibrates the lateral distance and box position of the prediction box with respect to the reference points, and gives query objects to achieve class label matching through bipartite matching; specifically including: Flatten the image data into a fixed grid region, initialize the top-left corner of the image as the reference point, and the 2D coordinates of the reference point are r = {x, y} ∈ [0, 1] 2 , and instantiate a learnable 4D offset distance s = {l, t, r, b} ∈ [0, 1] from the reference point to the predicted box for each object query 4 , where l, t, r, and b represent the distances from the reference point to the four sides of the predicted box, in the order of top-left, top-right, bottom-right, and bottom-left. The object query reference is q = {e, r, s}, where e ∈ R d is the content embedding of dimension d; the model directly supervises the four-dimensional offset of the four sides of the predicted box to the reference point; the final box prediction formula is: are the box corner coordinates in the corresponding directions and their four-dimensional offsets respectively; after fixing the reference point {x, {y}} = {x, y}, prediction updates are performed layer by layer, and the prediction update calculation for each decoder layer is as follows: ΔS l = BoxHead l (s l-1 ,e l-1 ,r l-1 ), Among them, σ and σ -1 are s-type and inverse s-type function operations respectively; Δs l represents the side offset prediction; and are the side distance, reference point, and box position predicted from the l-th layer decoder respectively; BoxHead l is the prediction head after the layer-l decoder, which is independent among different decoder layers in the model setting; During the detection process, each query is only allowed to predict bounding boxes that overlap with its reference region, and this rule is applied to the bipartite matching process through the internal matching cost Linner; given N query objects Q = {q 1 , q 2 , ···, q n} and M ground truth objects G = {g 1 , g 2 , ···, g M}, the step function L inner (g i , q i ), when r i is outside the bounding box of g i , the reference point r i of q i is penalized, and its value is the penalty cost k, with a default value of 95; the final permutation formula for label assignment is: where L match is the original pairwise matching cost, including the classification cost and the positioning cost; is the bipartite matching permutation set of N elements.

4. The novel brain-inspired object detection method for occluded objects in complex environments according to claim 3, Characterized in that: The offset amplitude threshold in step 4) is s grid , and the side distance {0~100} of the prediction box b and the box position are calculated by the PointHead based on point-by-point features and the s-type function activation to obtain the predicted offset Δr l ' and Δr l . The movement update calculation of the reference point is as follows: If then Δr l ′ = PointHead l (s l-1 , e l-1 , r l-1 ), Among them, Δr l ′ and Δr l are the predicted offsets from r l to r l-1 and r 0 respectively before the s-shaped function activation σ, is the confidence of the feature point, is the detach function in PyTorch.

Citation Information

Patent Citations

  • Target detection method and device based on lifelong learning

    CN115620099A

  • High-voltage line nest detection method, model training method, device and equipment

    CN116524357A