Target robust detection method based on brain cognitive model

By adopting a brain cognitive model-based method in object detection, using modules such as simulated visual information and spatial attention, memory functions, etc. to extract and fuse image features, the problem of insufficient accuracy and robustness of traditional object detection methods in complex environments is solved, and more efficient object detection and stronger network interpretability are achieved.

CN120198655AActive Publication Date: 2025-06-24NORTHWESTERN POLYTECHNICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510669224.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Traditional deep neural network-based object detection methods have insufficient accuracy and robustness in complex environments.

Method used

The target robust detection method based on brain cognitive model is adopted, including an image feature extraction module that simulates the attention perception function of visual information, an image feature extraction module that simulates the attention function of visual space, a standard visual memory bank module that simulates the visual memory function, a perceptual-standard feature association fusion module that simulates the visual and memory association function, and a prediction and reasoning module that simulates the visual prediction and reasoning function. Through the combination of these modules, image features can be extracted and fused to achieve object detection.

Benefits of technology

It improves the accuracy and robustness of object detection in complex environments, while enhancing the interpretability of neural networks, reducing the computational complexity, and improving the network's ability to identify effective features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198655A_ABST
    Figure CN120198655A_ABST
Patent Text Reader

Abstract

The invention discloses a target robust detection method based on a brain cognitive model, and relates to the technical field of target detection, and the method comprises the steps: firstly inputting a real image and a standard image which are paired; secondly, an image feature extraction module with a visual information attention perception simulation function is used for extracting primary perception features from the input real image; then, an image feature extraction module with a visual space attention simulation function is used for extracting target perception features from the primary perception features; a standard visual memory library module with a visual memory simulating function is used for extracting target memory features of the standard image; inputting the target perception features and the target memory features into a perception-standard feature association fusion module which simulates a visual and memory association function, and performing interactive fusion learning on the target perception features and the target memory features by using a domain adaptation technology to form fusion features; and finally, inputting the fused features into a predictive reasoning module with a simulated visual predictive reasoning function, and completing position prediction of the target by using Kalman filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and particularly to a method for robust target detection based on a brain cognitive model. Background Art

[0002] In recent years, the deep integration and mutual reference of brain and cognitive science, artificial intelligence, and computer science have become an important international trend in the field of scientific research. Especially inspired by the human brain cognitive mechanism, the artificial intelligence system has been injected with new impetus and become a research hotspot for relevant domestic institutions.

[0003] Research on the latest progress in the field of artificial intelligence has found that improving the cognitive performance and interpretability of algorithms by imitating the neural function modules and information flow architectures of brain cognition is the forefront hot direction in this field and a new form of brain-inspired artificial intelligence. For example, the whole-brain reference architecture method proposed by Hiroshi Yamakawa et al. in Japan realizes general artificial intelligence and has been successfully applied in the field of robot navigation. The autonomous driving team at Xi'an Jiaotong University draws on the cognitive mechanism of the human brain's perception-motor loop to study brain-inspired and neuroscience-inspired methods for traffic scene understanding, scenario prediction, and driving decision-making, achieving safe and reliable autonomous driving.

[0004] The working mode of the brain is fundamentally different from that of traditional unexplainable deep neural networks. It operates quickly and efficiently by combining and interconnecting different brain function modules to form neural circuits. And the brain-inspired artificial intelligence constructed according to this brain working mode has natural interpretability in the sense of bionics. At the same time, compared with traditional artificial intelligence, it also has higher cognitive efficiency and lower power consumption. The visual cognitive function architecture is as Figure 1 shown. The main research results of domestic scholars in this field are summarized as follows: Based on the research results of cognitive psychology or neurophysiology of the brain, the human brain's visual attention mechanism is introduced into the target recognition and tracking in complex scenes, which is conducive to realizing a more stable recognition and tracking algorithm closer to the human cognitive mechanism.The literature "A robust visual tracking based on selective attention shift" regards the visual object tracking problem as a process of the transfer of human visual attention. The saliency map is calculated through Itti's basic model. The early attention selection process is used to locate salient objects, and the late attention process is used to transfer attention to achieve the purpose of identifying the lost target that has been tracked; the literature "Residual Attention Network for Image Classification" applies visual attention to object classification and recognition. This method combines bottom-up pre-attention features and top-down context information, and uses the attention mechanism to assist in the content area of the image, which can enhance effective information while suppressing invalid information, and shows excellent performance in tasks such as image recognition and classification; the literature "TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering" proposes a spatio-temporal attention model for the actual visual field. Inspired by the human brain's visual attention mechanism, this system extracts and integrates the charge-discharge mechanism and lateral interaction mechanism that combine local features (pixel level) and global features (objects) with permanent responses to accurately identify the image content of key frames and the main semantic regions of the image; the literature "Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition" proposes a recurrent neural network to model the temporal transfer process of gaze points. The modeled gaze point transfer is not simply a transfer of two-dimensional image positions, but a transfer of the semantic information of the area corresponding to the gaze point. Considering that the visual saliency model based on images / videos and the gaze point transfer model based on tasks are complementary in modeling methods, a hybrid network architecture is proposed to unify the two complementary models, and the gaze point prediction performance has been significantly improved compared with existing methods; the literature "Show, Attend and Tell: Neural Image Caption Generation with Visual Attention" combines bottom-up attention and top-down attention for robot visual tracking. Bottom-up divides the scene into a ground area and a salient area, and under the guidance of top-down attention, the target area and the obstacle area are separated from the salient area; however, most domestic scholars only focus on starting from one or a few local cognitive functions to explore how to imitate certain specific functions of the brain.This method often still remains within the framework of traditional artificial neural networks in overall architecture design, lacking a macroscopic understanding and mapping of the overall structure and function of the brain's cognitive process. In addition, most current research on brain-inspired intelligence mostly simulates and optimizes specific neural activity patterns or single cognitive functions, and has not been able to comprehensively and systematically reflect the overall cognitive process of the brain. Therefore, the research on brain-inspired intelligence urgently needs to break through the limitations of traditional artificial neural networks, construct a more comprehensive and systematic cognitive function mapping model, and explore a more complex, flexible and highly adaptable algorithm framework.

[0005] In summary, traditional object detection methods based on deep neural networks have insufficient accuracy and robustness in object detection in complex environments. Therefore, there is an urgent need for a target robust detection method based on a brain cognitive model, so that the cognitive mechanism can improve the recognition ability in complex combat environments, thereby overcoming the problems existing in the prior art. Summary of the Invention

[0006] The object of the present invention is to provide a target robust detection method based on a brain cognitive model, which solves the problem that traditional object detection methods based on deep neural networks in the prior art have insufficient accuracy and robustness in object detection in complex environments.

[0007] To achieve the above object, the present invention provides a target robust detection method based on a brain cognitive model, including the following steps: Step 1, input paired real images and standard images; Step 2, use an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real image; Step 3, use an image feature extraction module that simulates the visual spatial attention function to extract target perception features from the primary perception features; Step 4, use a standard visual memory library module that simulates the visual memory function to extract target memory features from the standard image; Step 5, input the target perception features and the target memory features into a perception-standard feature association and fusion module that simulates the visual and memory association function, and use domain adaptation technology to perform interactive fusion learning on the target perception features and the target memory features to form fusion features, forcing the network to strengthen the adaptation to regions with more key information; Step 6, input the fusion features into a prediction and reasoning module that simulates the visual prediction and reasoning function, and use Kalman filtering to complete the position prediction of the target.

[0008] Preferably, the process of using an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real image in Step 2 is as follows: S21, for the input real image , where is the number of image channels, is the image width, is the image height. After passing through a global average pooling layer to compress the spatial dimension, it is then input into a fully connected layer and the ReLU activation function. The expression is as follows: ; Among them, , represents the intermediate feature, represents the ReLU activation function, represents the fully connected layer, represents the global average pooling layer; S22. Based on the intermediate features obtained in S21, the corresponding attention weights are obtained through the position branch, channel branch, and filter branch respectively based on the Sigmoid function. The calculation expression is as follows: ; In the formula, , , represent the attention weights of the position branch, channel branch, and filter branch respectively, is the convolution kernel size, is the number of input channels, is the number of output channels, represents the position multi-layer perceptron, represents the channel multi-layer perceptron, represents the filter multi-layer perceptron; S23. Based on the branch attention weights, channel branch attention weights, filter branch attention weights, and the real image, the primary perception feature is obtained. The calculation expression is as follows: ; In the formula, represents the weight coefficient.

[0009] Preferably, in step 3, the process of extracting the target perception feature from the primary perception feature using the image feature extraction module that simulates the visual spatial attention function is as follows: S31. Use a large-scale convolution kernel to highlight the effective features of the original feature and suppress the invalid features, and then use the Sigmoid function to obtain the spatial block feature , so as to allow the network to focus on more critical effective regions. The calculation expression is as follows: ; In the formula, represents the feature after grouped convolution, Indicates a grouped convolution operation; S32. Fuse the spatial block features with the original features using average pooling to obtain spatial point features , and the calculation expression is as follows: ; S33. Reconstruct the channel features with a fully connected layer to obtain the channel-level importance of all feature maps, and the calculation expression is as follows: ; In the formula, represents the obtained channel features, represents the activation function, represents the fully connected channel layer; S34. Multiply the obtained in S33 with the original features to obtain the target perception features , and the expression is as follows: .

[0010] Preferably, the expression for extracting the target memory features from the standard image using the standard visual memory library module that simulates the visual memory function in step 4 is as follows: ; In the formula, is the input standard image, is the feature extractor, are the decoding structures for category and location respectively, are the obtained category and location predictions respectively, are the target category memory features and target location memory features respectively.

[0011] Preferably, in step 5, the target perception features and the target memory features are input into the perception-standard feature association and fusion module that simulates the visual and memory association function, and the process of interactively fusing and learning the target perception features and the target memory features using domain adaptation technology is as follows: S51. Input the target perception features extracted in step 3 into the location and category feature decoupler to obtain the target category perception features and the target location perception features , and the calculation expression is as follows: ; S52. Take the average value of the target category perception features and the target category memory features along the channel dimension, and respectively multiply them with the foreground mask Multiply them to obtain the foreground response map, and the calculation expression is as follows: ; In the formula, represents the target category perception response feature, represents the target category memory response feature, is the number of channels; S53. Subtract the two obtained foreground response maps and normalize them to (0, 1) to obtain the spatial reweighting map , and the expression is as follows: ; In the formula, represents the operation of normalizing the mapping to 0 - 1; S54. Send the difference between the target category perception feature and the target category memory feature into the C - R module to obtain the channel reweighting vector , and the expression is as follows: ; In the formula, respectively represent average pooling, fully connected layer, and softmax operations; S55. Multiply the channel reweighting vector and the spatial reweighting map to obtain the domain adaptation weight , and the expression is as follows: ; S56. Multiply the domain adaptation weight by the target location perception feature to obtain the final fused feature.

[0012] Preferably, in step 6, the fused feature is input into the prediction and inference module that simulates the visual prediction and inference function, and the process of using Kalman filtering to complete the target position prediction is as follows: S61. Obtain the current image detection result; input the fused feature into the detection module to obtain the current image detection result ; S62. Initialize the trajectory pipeline and perform position prediction; create its corresponding trajectory pipeline for the result detected in the first image; initialize the motion variables of the Kalman filter, and predict its corresponding position through the prediction equation of the Kalman filter; at this time, the state of the trajectory pipeline is marked as undetermined; the prediction equation of the Kalman filter is: ; Among them, is The predicted state vector at a moment, is the state transition matrix, is the estimated state vector at a moment, is the control input matrix, is the control input vector, is the process noise; S63. IOU matching and cost matrix calculation; The results of the current image object detection and the positions predicted by the trajectory pipeline for the previous image are subjected to IOU matching, and then the cost matrix is calculated based on the results of the IOU matching ; Assume that there are detection results in the current image, and there are positions predicted by the trajectory pipeline for the previous image. Calculate a cost matrix , using as the cost, and the calculation formula is as follows: ; Among them, represents and the intersection over union ratio, represents the cost between the th detection result and the th predicted position; The areas of two bounding boxes and are and respectively, the intersection area is , and the union area is ; represents the bounding box corresponding to the rd detection result in the current frame, represents the bounding box corresponding to the th predicted position in the previous frame, , ; S64. All the cost matrices obtained in S63 As the input of the Hungarian algorithm, a linear matching result is obtained; the matching result includes three cases: track pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track pipelines. Among them, when the matching result is track pipeline mismatch, the mismatched track pipeline (at this time, the track pipeline is in an uncertain state) is directly deleted; when the matching result is detection mismatch, it is initialized as a new track pipeline; when the matching result is successful pairing of detection and predicted track pipelines, it indicates that the current image and the previous image are successfully tracked, and the corresponding detection updates the corresponding track pipeline variable through the update equation of the Kalman filter. The update expression of the Kalman filter is as follows: ; Wherein, is the estimated state vector at time is the Kalman gain, is the measurement value at time is the measurement matrix; S65. Predict the positions corresponding to the track pipelines in the confirmed state and the track pipelines in the unconfirmed state through the prediction equation of the Kalman filter; cascade-match the predicted position of the track pipeline in the confirmed state and the detection result ; S66. Obtain two results of the cascade match, namely: track pipeline match and detection and track pipeline mismatch; among them, when the matching result is track pipeline match, update the corresponding track pipeline variable through the update equation of the Kalman filter; when the matching result is detection and track pipeline mismatch, perform IOU matching on the track pipeline in the unconfirmed state of S65 and the mismatched track pipeline together with the detection results that have not been successfully matched, and then calculate its cost matrix , and the calculation method is the same as that in S63; S67. Use the obtained in S66 as the input of the Hungarian algorithm to obtain a linear matching result. At this time, the matching result still includes three cases: track pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track pipelines, and the corresponding processing methods are the same as those in S64; S68. Repeatedly loop through S65 to S67 until the entire image sequence ends.

[0013] Therefore, the present invention adopts the above-mentioned target robust detection method based on the brain cognitive model, and has the following beneficial effects: (1) Improve the accuracy of object detection: By means of an image feature extraction module that simulates the visual information attention perception function, the attention weights of different dimensions of the convolutional kernel are calculated in parallel, enhancing the weak structural features of small objects and suppressing background interference, thereby improving the accuracy of object detection. (2) Enhance the interpretability of the neural network: Drawing on the working mode of the brain, a brain-inspired artificial intelligence is constructed, making the neural network inherently interpretable. Compared with traditional deep neural networks, it can better understand its decision-making process and basis. (3) Reduce the computational complexity: The block-aware channel attention method that simulates the visual spatial attention function uses a large-scale convolutional kernel to sample each channel, emphasizing the importance of spatial blocks in the feature map. By means of pooling operations, effective features are extracted and transformed into spatial point features, reducing the size of the feature map and the computational complexity. (4) Improve the network's ability to identify effective features: The block-aware channel attention module independently models the feature maps of each channel, analyzes them separately in a targeted manner, obtains the channel importance and assigns weights, performs channel fusion, and then fuses the channel importance and spatial importance of all feature maps, greatly improving the network's ability to identify effective features. (5) Achieve robust detection and tracking of objects in complex environments: The prediction and inference module that simulates the visual prediction and inference function combines motion and appearance information, as well as motion compensation and a more accurate Kalman filter state vector, enabling robust detection and tracking of objects in complex environments.

[0014] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0015] Figure 1 It is the visual cognitive function architecture diagram in the background art; Figure 2 It is the overall flowchart of a method for robust object detection based on a brain cognitive model of the present invention; Figure 3 It is a schematic diagram of an image feature extraction module that simulates the visual information attention perception function in an embodiment of the present invention; Figure 4 It is a schematic diagram of an image feature extraction module that simulates the visual spatial attention function in an embodiment of the present invention; Figure 5 It is a schematic diagram of a perception-criterion feature association and fusion module that simulates the visual and memory association function in an embodiment of the present invention; Figure 6 It is a schematic diagram of a prediction and inference module that simulates the visual prediction and inference function in an embodiment of the present invention. Specific Embodiments

[0016] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0017] Please refer to Figures 1-6 , a target robust detection method based on a brain cognitive model, comprising the following steps: Step 1, input paired real images and standard images; Step 2, use an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real image; the image feature extraction module that simulates the visual information attention perception function can parallelly calculate the attention weights of different dimensions of the convolution kernel, while enhancing the weak structural features of small targets and suppressing background interference, introducing a relatively low additional computational amount; the specific process is as follows: S21, for the input real image , where is the number of image channels, is the image width, is the image height, compress the spatial dimension through a global average pooling layer, and then input it into a fully connected layer and a ReLU activation function, and the expression is as follows: ; Among them, , represents the intermediate feature, represents the ReLU activation function, represents the fully connected layer, represents the global average pooling layer; S22, based on the Sigmoid function, obtain the corresponding attention weights for the intermediate feature obtained in S21 through the position branch, channel branch, and filter branch respectively, and the calculation expression is as follows: ; In the formula, , , respectively represent the attention weights of the position branch, channel branch, and filter branch, is the convolution kernel size, is the number of input channels, is the number of output channels, represents the position multi-layer perceptron, represents the channel multi-layer perceptron, represents the filter multi-layer perceptron; S23: Obtain primary perceptual features based on branch attention weights, channel branch attention weights, filter branch attention weights and real images , the calculation expression is as follows: ; In the formula, Represents the weight coefficient.

[0018] Step 3, use the image feature extraction module that simulates the visual spatial attention function to extract the target perceptual features from the primary perceptual features; the image feature extraction module that simulates the visual spatial attention function can eliminate the interference of negative features. By independently modeling the feature map of each channel, the image feature extraction module can perform targeted and separate analysis on the feature maps of different channels. On this basis, the image feature extraction module uses a large-scale convolution kernel to sample each channel, abandoning the strategy of the traditional attention mechanism that over-considers the microscopic relationship between pixels. On the contrary, the image feature extraction module emphasizes the importance of spatial blocks in the feature map to highlight the relatively important parts in the spatial domain, that is, the effective features of the object. Secondly, the pooling operation is used to extract these effective features from the spatial blocks and convert them into spatial point features to reduce the size of the feature map and reduce the computational complexity; then, the image feature extraction module obtains the channel importance and assigns weights, and performs channel fusion through the fully connected layer. Finally, the channel importance and spatial importance of all the obtained feature maps are fused; the image feature extraction module greatly improves the ability of the network to identify effective features; the specific process is as follows: S31. Use large-scale convolution kernels to highlight original features The effective features are used to suppress invalid features, and then the Sigmoid function is used to obtain the spatial block features. , thereby allowing the network to focus on more critical effective areas, the calculation expression is as follows: ; In the formula, represents the features after group convolution, Represents a grouped convolution operation; S32, the spatial block feature With the original features Fusion, using average pooling Get spatial point features , the calculation expression is as follows: ; S33. Use the fully connected layer to reconstruct the channel features and obtain the channel-level importance of all feature maps. The calculation expression is as follows: ; In the formula, represents the obtained channel features, represents an activation function, represents a fully connected channel layer; S34. Multiply the obtained in S33 by the original features to obtain the target perception features , and the expression is as follows: .

[0019] Step 4. Use the standard visual memory library module that simulates the visual memory function to extract the target memory features from the standard images; inspired by the semantic memory function of the anterior temporal lobe in the visual perception process, a standard visual memory library module is introduced in the network architecture to simulate visual memory. The standard visual memory library module takes the standard images with sufficient and complete information as the input and maintains the same architecture as the image perception feature extraction module; the specific expression is as follows: ; In the formula, is the input standard image, is the feature extractor, are the decoding structures for the category and location respectively, are the obtained category and location predictions respectively, are the target category memory features and target location memory features respectively.

[0020] Step 5. Input the target perception features and target memory features into the perception-standard feature association and fusion module that simulates the visual and memory association function, and use domain adaptation technology to perform interactive fusion learning on the target perception features and target memory features to form fusion features, forcing the network to strengthen the adaptation to the regions with more key information; inspired by the visual feature and memory integration function of the inferior longitudinal fasciculus, an association and fusion module based on domain adaptation technology is introduced in the network architecture to integrate the perception features and standard memory features. During the training process, the perception features encoded by the image perception feature extraction module are adaptively mapped to the memory features generated by the memory feature generation module through domain adaptation, so as to establish feature associations; among them, the process of performing interactive fusion learning on the target perception features and target memory features to form fusion features is as follows: S51. Input the target perception features extracted in step 3 into the location and category feature decoupler to obtain the target category perception feature and the target location perception feature , and the calculation expression is as follows: ; S52. Take the average value of the target category perception feature and the target category memory feature along the channel dimension, and respectively multiply them with the foreground mask Multiply to obtain the foreground response map, and the calculation expression is as follows: ; In the formula, represents the target category perception response feature, represents the target category memory response feature, is the number of channels; S53. Subtract the two obtained foreground response maps and normalize them to (0, 1) to obtain the spatial re-weighting map , and the expression is as follows: ; In the formula, represents the operation of normalizing the mapping to 0-1; S54. Send the difference between the target category perception feature and the target category memory feature into the C-R module to obtain the channel re-weighting vector , and the expression is as follows: ; In the formula, respectively represent average pooling, fully connected layer, and softmax operations; S55. Multiply the channel re-weighting vector and the spatial re-weighting map to obtain the domain adaptation weight , and the expression is as follows: ; S56. Multiply the domain adaptation weight by the target location perception feature to obtain the final fused feature.

[0021] Step 6. Input the fused feature into the prediction and inference module that simulates the visual prediction and inference function, and use Kalman filtering to complete the position prediction of the target; inspired by the prefrontal visual perception and inference function in the high-level visual cortex, a perception and inference module is added to the network model to simulate the inference function based on visual memory and semantic knowledge to simulate the visual inference function of the prefrontal lobe in the high-level visual cortex; combining the advantages of motion and appearance information, as well as motion compensation and a more accurate Kalman filter state vector, to achieve robust detection and tracking of the target in complex environments; the specific process is as follows: S61. Obtain the detection result of the current picture; input the fused feature into the detection module to obtain the detection result of the current picture ; S62. Initialize the trajectory pipeline and position prediction; create its corresponding trajectory pipeline for the result detected in the first picture; set the motion variable of the Kalman filter Initialization: Predict its corresponding position through the prediction equation of the Kalman filter; at this time, the state of the trajectory pipeline is marked as undetermined; the prediction equation of the Kalman filter is: ; Wherein, is the predicted state vector at time is the state transition matrix, is the estimated state vector at time is the control input matrix, is the control input vector, is the process noise; S63, IOU matching and cost matrix calculation; Perform IOU matching on the result of the current image object detection and the position predicted by the trajectory pipeline in the previous image, and then calculate its cost matrix through the result of the IOU matching; Assume that there are detection results in the current image and predicted positions by the trajectory pipeline in the previous image, calculate a cost matrix , use as the cost, and the calculation formula is as follows: ; Wherein, represents and the intersection over union of, represents the cost between the th detection result and the th predicted position; The areas of two bounding boxes and are and respectively, the intersection area is , and the union area is ; represents the bounding box corresponding to the th detection result in the current frame, represents the bounding box corresponding to the th predicted position in the previous frame, , ; S64. All the cost matrices As the input of the Hungarian algorithm, a linear matching result is obtained; the matching result includes three cases: track pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track pipeline. Among them, when the matching result is track pipeline mismatch, the mismatched track pipeline (at this time, the track pipeline is in an uncertain state) is directly deleted; when the matching result is detection mismatch, it is initialized as a new track pipeline; when the matching result is successful pairing of detection and predicted track pipeline, it indicates that the current image and the previous image are successfully tracked, and the corresponding detection updates the corresponding track pipeline variable through the update equation of the Kalman filter. The update expression of the Kalman filter is as follows: ; Wherein, is the estimated state vector at time is the Kalman gain, is the measurement value at time is the measurement matrix; S65. Predict the positions corresponding to the confirmed track pipeline and the unconfirmed track pipeline through the prediction equation of the Kalman filter; cascade-match the predicted position of the confirmed track pipeline and the detection result ; S66. Obtain two results of the cascade match, namely: track pipeline match and detection and track pipeline mismatch; among them, when the matching result is track pipeline match, update the corresponding track pipeline variable through the update equation of the Kalman filter; when the matching result is detection and track pipeline mismatch, the unconfirmed track pipeline in S65 and the mismatched track pipeline are matched with the detection results that have not been successfully matched through IOU, and then the cost matrix is calculated according to the result of the IOU match, and the calculation method is the same as that in S63; S67. Use the obtained in S66 as the input of the Hungarian algorithm to obtain a linear matching result. At this time, the matching result still includes three cases: track pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track pipeline, and the corresponding processing methods are the same as those in S64; S68. Repeatedly loop through S65 to S67 until the entire image sequence ends.

[0022] Therefore, the present invention adopts the above-mentioned target robust detection method based on a brain cognitive model. Starting from the whole process of the biological brain processing natural environment visual images, an image feature extraction module simulating the visual information attention perception function, an image feature extraction module simulating the visual spatial attention function, a standard visual memory library module simulating the visual memory function, a perception-standard feature association and fusion module simulating the visual and memory association function, and a prediction and reasoning module simulating the visual prediction and reasoning function are respectively designed. Compared with the traditional target detection method based on a deep neural network, the present invention effectively improves the accuracy and robustness of target detection in complex environments. At the same time, based on the biological brain structure, the neural network has a certain interpretability.

[0023] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A target robust detection method based on a brain cognitive model, characterized in that, It includes the following steps: Step 1, input paired real images and standard images; Step 2, use an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real images; Step 3, use an image feature extraction module that simulates the visual spatial attention function to extract target perception features from the primary perception features; Step 4, use a standard visual memory bank module that simulates the visual memory function to extract target memory features from the standard images; Step 5, input the target perception features and target memory features into a perception-standard feature association and fusion module that simulates the visual and memory association function, and use domain adaptation technology to perform interactive fusion learning on the target perception features and target memory features to form fusion features; Step 6, input the fusion features into a prediction and reasoning module that simulates the visual prediction and reasoning function, and use Kalman filtering to complete the position prediction of the target.

2. The target robust detection method based on a brain cognitive model according to claim 1, wherein The process of using the image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real images in Step 2 is as follows: S21. For the input real image , where is the number of image channels, is the image width, is the image height. After passing through a global average pooling layer to compress the spatial dimension, it is then input into a fully connected layer and the ReLU activation function. The expression is as follows: ; Among them, , represents the intermediate feature, represents the ReLU activation function, represents the fully connected layer, represents the global average pooling layer; S22, based on the Sigmoid function, obtain the corresponding attention weights for the intermediate features obtained in S21 through the position branch, channel branch, and filter branch respectively. The calculation expression is as follows: ; Wherein, , , represent the position branch attention weight, the channel branch attention weight, and the filter branch attention weight respectively, is the convolution kernel size, is the number of input channels, is the number of output channels, represents the position multi-layer perceptron, represents the channel multi-layer perceptron, represents the filter multi-layer perceptron; S23. Obtain the primary perception features based on the branch attention weight, channel branch attention weight, filter branch attention weight, and real image, and the calculation expression is as follows: , as follows: ; In the formula, represents the weight coefficient.

3. The target robust detection method based on the brain cognitive model according to claim 2, wherein, The process of using the image feature extraction module that simulates the visual spatial attention function to extract target perception features from the primary perception features in Step 3 is as follows: S31. Use a large-scale convolutional kernel to highlight the effective features of the original features, suppress the ineffective features, and then use the Sigmoid function to obtain the spatial block features The calculation expression is as follows: ​ ; In the formula, represents the feature after grouped convolution, represents the grouped convolution operation; S32. Fuse the spatial block feature with the original feature using average pooling to obtain the spatial point feature , and the calculation expression is as follows: ; S33, reconstruct the channel features with a fully connected layer to obtain the channel-level importance of all feature maps. The calculation expression is as follows: ; In the formula, represents the obtained channel feature, represents the activation function, represents the fully connected channel layer; S34. Multiply the obtained in S33 by the original feature to obtain the target perception feature . The expression is as follows: 。 4. A target robust detection method based on a brain cognitive model according to claim 3, characterized in that, The expression for using the standard visual memory bank module that simulates the visual memory function to extract target memory features from the standard images in Step 4 is as follows: ; In the formula, is the input standard image, is the feature extractor, are the decoding structures for the class and location respectively, are the obtained class and location predictions respectively, are the target class memory feature and the target location memory feature respectively.

5. The target robust detection method based on a brain cognitive model according to claim 4, wherein: The process of inputting the target perception features and target memory features into a perception-standard feature association and fusion module that simulates the visual and memory association function, and using domain adaptation technology to perform interactive fusion learning on the target perception features and target memory features to form fusion features in Step 5 is as follows: S51. Input the target perception features extracted in step 3 into the position and class feature decoupler to obtain the target class perception feature and the target position perception feature . The calculation expression is as follows: ; S52. Take the average value of the target category perception feature and the target category memory feature along the channel dimension, and multiply them with the foreground mask respectively to obtain the foreground response map. The calculation expression is as follows: ; In the formula, represents the target category perception response feature, represents the target category memory response feature, is the number of channels; S53. Subtract the two obtained foreground response maps and normalize them to (0, 1) to obtain the spatial re-weighting map , and the expression is as follows: ; In the formula, represents the operation of normalizing the mapping to 0-1; S54. Feed the difference between the target category perception feature and the target category memory feature into the C-R module to obtain the channel reweighting vector . The expression is as follows: ; In the formula, represent average pooling, fully connected layer, and softmax operation respectively; S55. Multiply the channel reweighting vector and the spatial reweighting map to obtain the domain adaptation weight , and the expression is as follows: ; S56. Multiply the domain adaptation weight with the target location-aware feature to obtain the final fused feature.

6. The target robust detection method based on a brain cognitive model according to claim 5, characterized in that: The process of inputting the fusion features into a prediction and reasoning module that simulates the visual prediction and reasoning function, and using Kalman filtering to complete the position prediction of the target in Step 6 is as follows: S61. Obtain the detection result of the current image; input the fused feature into the detection module , and obtain the detection result of the current image ; S62, initialize the trajectory pipeline and perform position prediction; The results detected from the first picture Create its corresponding trajectory pipeline ; The motion variables of the Kalman filter are initialized, and their corresponding positions are predicted through the prediction equation of the Kalman filter; at this time, the status of the trajectory pipeline is marked as undetermined; the prediction equation of the Kalman filter is: ; Among them, is the predicted state vector at a moment, is the state transition matrix, is the estimated state vector at a moment, is the control input matrix, is the control input vector, is the process noise; S63, IOU matching and cost matrix calculation; the results of object detection in the current image and the positions predicted by the trajectory pipeline in the previous image are subjected to IOU matching, and then the cost matrix is calculated based on the results of IOU matching ; assume that there are detection results in the current image, and there are positions predicted by the trajectory pipeline in the previous image, and calculate a cost matrix , using as the cost, and the calculation formula is as follows: ; Among them, represents and 's intersection over union, represents the cost between the th detection result and the th predicted position; the areas of two bounding boxes and are and respectively, the intersection area is and the union area is ; represents the bounding box corresponding to the th detection result in the current frame, represents the bounding box corresponding to the th predicted position in the previous frame, , ; S64. All the cost matrices obtained in S63 are used as the input of the Hungarian algorithm to obtain a linear matching result. The matching result includes three cases: track-pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track-pipeline. Among them, when the matching result is track-pipeline mismatch, the mismatched track-pipeline is directly deleted; when the matching result is detection mismatch, it is initialized as a new track-pipeline; when the matching result is successful pairing of detection and predicted track-pipeline, the corresponding detection updates the corresponding track-pipeline variable through the update equation of the Kalman filter. The update expression of the Kalman filter is as follows: ; Among them, is the estimated state vector at a moment, is the Kalman gain, is the measured value at a moment, is the measurement matrix; S65. Predict the trajectory pipelines in the confirmed state and the trajectory pipelines in the unconfirmed state through the prediction equation of the Kalman filter and the corresponding positions; cascade and match the predicted positions of the trajectory pipelines in the confirmed state with the detection results ; ​ S66. Obtain two results of cascade matching, namely: trajectory pipeline matching and detection and trajectory pipeline mismatch. Among them, when the matching result is trajectory pipeline matching, update the corresponding trajectory pipeline variables through the update equation of Kalman filtering. When the matching result is detection and trajectory pipeline mismatch, the unconfirmed trajectory pipeline in S65 and the mismatched trajectory pipeline are jointly subjected to IOU matching with the detection results that have not been successfully matched, and then calculate their cost matrix based on the results of IOU matching ; S67. Use the result obtained in S66 as the input of the Hungarian algorithm to obtain a linear matching result. At this time, the matching result still includes three cases: trajectory pipeline mismatch, detection mismatch, and successful pairing of the detected and predicted trajectory pipelines; S68, repeatedly loop through S65 to S67 until the entire image sequence ends.

Citation Information

Patent Citations

  • Video perception-fused multi-task synergetic recognition method and system

    CN108846384A

  • Multi-visual memory unit-based glancing path prediction method

    CN116563524A

  • Target tracking method based on gated attention mechanism and space-time memory network

    CN119131085A

  • RGB-t multispectral pedestrian detection method based on target perception fusion policy

    WO2024197762A1