A target robust detection method based on a brain cognitive model

By simulating the visual information and predictive reasoning function of brain cognitive processes, combined with Kalman filtering, the problem of insufficient accuracy and robustness of traditional object detection methods in complex environments is solved, and high-precision and robust object detection are achieved.

CN120198655BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510669224.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-22
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Traditional deep neural network-based object detection methods are insufficient in complex environments and cannot fully and systematically reflect the overall cognitive process of the brain.

Method used

A robust target detection method based on brain cognitive model is adopted, including image feature extraction that simulates visual information attention perception function, feature extraction that simulates visual spatial attention function, feature extraction that simulates visual memory function, and predictive reasoning module, combined with Kalman filtering to predict target position, and simulate the cognitive process of the brain.

Benefits of technology

It improves the accuracy and robustness of object detection, reduces the computational complexity, enhances the interpretability of neural networks, and can achieve robust detection and tracking of targets in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198655B_ABST
    Figure CN120198655B_ABST
Patent Text Reader

Abstract

The present invention discloses a target robust detection method based on a brain cognitive model, which relates to the technical field of target detection. First, paired real images and standard images are input; secondly, a primary perception feature is extracted from the input real image by an image feature extraction module that simulates the visual information attention perception function; then, a target perception feature is extracted from the primary perception feature by an image feature extraction module that simulates the visual spatial attention function; subsequently, a target memory feature is extracted from the standard image by a standard visual memory bank module that simulates the visual memory function; then, the target perception feature and the target memory feature are input into a perception-standard feature association and fusion module that simulates the visual and memory association function, and using domain adaptation technology, the target perception feature and the target memory feature are interactively fused and learned to form a fusion feature; finally, the fusion feature is input into a prediction and inference module that simulates the visual prediction and inference function, and the position prediction of the target is completed using Kalman filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a method for robust target detection based on a brain cognitive model. Background Art

[0002] In recent years, the deep integration and mutual reference of brain and cognitive science, artificial intelligence, and computer science have become an important international trend in the field of scientific research. In particular, inspired by the human brain cognitive mechanism, the artificial intelligence system has been injected with new impetus and has become a research hotspot for relevant domestic institutions.

[0003] Research on the latest progress in the field of artificial intelligence has found that improving the cognitive performance and interpretability of algorithms by imitating the neural function modules and information flow architectures of brain cognition is a cutting-edge hot direction in this field and a new form of brain-inspired artificial intelligence. For example, the whole-brain reference architecture method proposed by Hiroshi Yamakawa et al. in Japan realizes general artificial intelligence and has been successfully applied in the field of robot navigation. The autonomous driving team at Xi'an Jiaotong University draws on the cognitive mechanism of the human brain's perception-motor loop to study brain-inspired and neuroscience-inspired methods for traffic scene understanding, scenario prediction, and driving decision-making, achieving safe and reliable autonomous driving.

[0004] The working mode of the brain is fundamentally different from that of traditional unexplainable deep neural networks. It operates quickly and efficiently by combining and interconnecting different brain function modules to form neural circuits. Brain-inspired artificial intelligence constructed based on this brain working mode has natural interpretability in the sense of bionics, and at the same time has higher cognitive efficiency and lower power consumption compared with traditional artificial intelligence. The visual cognitive function architecture is as Figure 1 shown. The main research results of domestic scholars in this field are summarized as follows:

[0005] Based on the research results of cognitive psychology or neurophysiology, the human brain's visual attention mechanism is introduced into the target recognition and tracking in complex scenes, which is conducive to realizing a more stable recognition and tracking algorithm closer to the human cognitive mechanism.The literature "A robust visual tracking based on selective attention shift" regards the visual object tracking problem as a process of the transfer of human visual attention. The saliency map is calculated through Itti's basic model. The early attention selection process is used to locate salient objects, and the late attention process is used to transfer attention in order to achieve the purpose of recognizing the lost tracked object. The literature "Residual Attention Network for Image Classification" applies visual attention to object classification and recognition. This method combines bottom-up pre-attention features and top-down context information, and uses the attention mechanism to assist in the content area of the image, which can enhance effective information while suppressing invalid information, and shows excellent performance in tasks such as image recognition and classification. The literature "TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering" proposes a spatio-temporal attention model for the actual visual field. Inspired by the human brain's visual attention mechanism, this system extracts and integrates the charge-discharge mechanism and lateral interaction mechanism that combine local features (pixel level) and global features (objects) with permanent responses to accurately identify the image content of key frames and the main semantic regions of images. The literature "Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition" proposes a recurrent neural network to model the temporal transfer process of fixation points. The modeled fixation point transfer is not simply the transfer of two-dimensional image positions, but the transfer of semantic information in the region corresponding to the fixation point. Considering that the visual saliency model based on images / videos and the fixation point transfer model based on tasks are complementary in modeling methods, a hybrid network architecture is proposed to unify the two complementary models, and the fixation point prediction performance has been significantly improved compared with existing methods. The literature "Show, Attend and Tell: Neural Image Caption Generation with Visual Attention" combines bottom-up attention and top-down attention for robot visual tracking. Bottom-up divides the scene into a ground area and a salient area, and under the guidance of top-down attention, the target area and the obstacle area are separated from the salient area. However, most domestic scholars only focus on starting from one or a few local cognitive functions to explore how to imitate certain specific functions of the brain.This method often still remains within the framework of traditional artificial neural networks in overall architecture design, lacking a macroscopic understanding and mapping of the overall structure and function of the brain's cognitive process. In addition, most current research on brain-inspired intelligence mostly simulates and optimizes specific neural activity patterns or single cognitive functions, and has not been able to comprehensively and systematically reflect the overall cognitive process of the brain. Therefore, the research on brain-inspired intelligence urgently needs to break through the limitations of traditional artificial neural networks, construct a more comprehensive and systematic cognitive function mapping model, and explore a more complex, flexible and highly adaptable algorithm framework.

[0006] In summary, traditional object detection methods based on deep neural networks have insufficient accuracy and robustness in object detection in complex environments. Therefore, there is an urgent need for a robust object detection method based on a brain cognitive model to enable the cognitive mechanism to improve the recognition ability in complex combat environments, thereby overcoming the problems existing in the prior art. Summary of the Invention

[0007] The object of the present invention is to provide a robust object detection method based on a brain cognitive model, which solves the problem that traditional object detection methods based on deep neural networks in the prior art have insufficient accuracy and robustness in object detection in complex environments.

[0008] To achieve the above object, the present invention provides a robust object detection method based on a brain cognitive model, including the following steps:

[0009] Step 1, input paired real images and standard images;

[0010] Step 2, use an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real images;

[0011] Step 3, use an image feature extraction module that simulates the visual spatial attention function to extract object perception features from the primary perception features;

[0012] Step 4, use a standard visual memory library module that simulates the visual memory function to extract object memory features from the standard images;

[0013] Step 5, input the object perception features and object memory features into a perception-standard feature association and fusion module that simulates the visual and memory association function, and use domain adaptation technology to perform interactive fusion learning on the object perception features and object memory features to form fusion features, forcing the network to strengthen the adaptation to regions with more key information;

[0014] Step 6, input the fusion features into a prediction and reasoning module that simulates the visual prediction and reasoning function, and use Kalman filtering to complete the position prediction of the object.

[0015] Preferably, in step 2, the process of extracting primary perception features from the input real image by the image feature extraction module that simulates the visual information attention perception function is as follows:

[0016] S21. For the input real image , where is the number of image channels, is the image width, is the image height. After passing through a global average pooling layer to compress the spatial dimension, it is then input into a fully connected layer and a ReLU activation function. The expression is as follows:

[0017] ;

[0018] Among them, , represents the intermediate feature, represents the ReLU activation function, represents the fully connected layer, represents the global average pooling layer;

[0019] S22. Based on the intermediate feature obtained in S21, the corresponding attention weights are obtained through the position branch, channel branch, and filter branch respectively based on the Sigmoid function. The calculation expression is as follows:

[0020] ;

[0021] In the formula, , , respectively represent the attention weight of the position branch, the attention weight of the channel branch, and the attention weight of the filter branch, is the convolution kernel size, is the number of input channels, is the number of output channels, represents the position multi-layer perceptron, represents the channel multi-layer perceptron, represents the filter multi-layer perceptron;

[0022] S23. Based on the branch attention weight, channel branch attention weight, filter branch attention weight, and real image, the primary perception feature is obtained. The calculation expression is as follows:

[0023] ;

[0024] In the formula, represents the weight coefficient.

[0025] Preferably, in step 3, the process of extracting target perception features from primary perception features using an image feature extraction module that simulates visual spatial attention function is as follows:

[0026] S31. Use large-scale convolution kernels to highlight original features The effective features are used to suppress invalid features, and then the Sigmoid function is used to obtain the spatial block features. , thereby allowing the network to focus on more critical effective areas, the calculation expression is as follows:

[0027] ;

[0028] In the formula, represents the features after group convolution, Represents a grouped convolution operation;

[0029] S32, the spatial block feature With the original features Fusion, using average pooling Get spatial point features , the calculation expression is as follows:

[0030] ;

[0031] S33. Use the fully connected layer to reconstruct the channel features and obtain the channel-level importance of all feature maps. The calculation expression is as follows:

[0032] ;

[0033] In the formula, represents the obtained channel features, represents the activation function, represents the fully connected channel layer;

[0034] S34, the data obtained in S33 Multiply with the original feature to get the target perception feature , the expression is as follows:

[0035] .

[0036] Preferably, in step 4, the expression for extracting the target memory feature from the standard image using the standard visual memory library module simulating the visual memory function is as follows:

[0037] ;

[0038] In the formula, is the standard input image, is a feature extractor, are the decoding structures of categories and positions respectively, They are the obtained category and location predictions, which are the target category memory feature and the target location memory feature respectively.

[0039] Preferably, in step 5, the target perception feature and the target memory feature are input into the perception-criterion feature association and fusion module that simulates the visual and memory association functions. Using domain adaptation technology, the process of interactively fusing and learning the target perception feature and the target memory feature to form a fusion feature is as follows:

[0040] S51. Input the target perception feature extracted in step 3 into the location and category feature decoupler to obtain the target category perception feature and the target location perception feature . The calculation expressions are as follows:

[0041] ;

[0042] S52. Take the average value of the target category perception feature and the target category memory feature along the channel dimension, and multiply them with the foreground mask respectively to obtain the foreground response map. The calculation expressions are as follows:

[0043] ;

[0044] In the formula, represents the target category perception response feature, represents the target category memory response feature, is the number of channels;

[0045] S53. Subtract the two obtained foreground response maps and normalize them to (0, 1) to obtain the spatial reweighting map . The expression is as follows:

[0046] ;

[0047] In the formula, represents the operation of normalizing the mapping to 0 - 1;

[0048] S54. Send the difference between the target category perception feature and the target category memory feature into the C-R module to obtain the channel reweighting vector . The expression is as follows:

[0049] ;

[0050] In the formula, Represent average pooling, fully connected layer and softmax operation respectively;

[0051] S55, multiply the channel weight vector and the spatial weight map to obtain the domain adaptation weight , the expression is as follows:

[0052] ;

[0053] S56. Domain adaptation weights and target position perception features Multiply them together to get the final fusion feature.

[0054] Preferably, in step 6, the fusion feature is input into the prediction and reasoning module simulating the visual prediction and reasoning function, and the process of using Kalman filtering to complete the position prediction of the target is as follows:

[0055] S61, obtaining the current image detection result; inputting the fusion feature into the detection module , get the current image detection result ;

[0056] S62, trajectory pipeline initialization and position prediction; the results detected in the first picture Create its corresponding trajectory pipeline ; Kalman filter motion variables Initialization, the corresponding position is predicted by the prediction equation of the Kalman filter; at this time, the state of the trajectory pipeline is marked as undetermined; the prediction equation of the Kalman filter is:

[0057] ;

[0058] in, yes The predicted state vector at time , is the state transition matrix, yes The estimated state vector at time , is the control input matrix, is the control input vector, is the process noise;

[0059] S63, IOU matching and cost matrix calculation; the result of the current image target detection And the position predicted by the trajectory pipeline of the previous image Perform IOU matching, and then calculate the cost matrix based on the results of IOU matching ; Assume that the current image has detection results, the position predicted by the trajectory pipeline for the previous image is , calculate one The cost matrix , using as the cost, the calculation formula is as follows:

[0060] ;

[0061] Among them, represents and the intersection over union of represents the cost between the th detection result and the th predicted position; the areas of the two bounding boxes and are and respectively, the intersection area is and the union area is ; represents the bounding box corresponding to the th detection result in the current frame, represents the bounding box corresponding to the th predicted position in the previous frame, , ;

[0062] S64. Use all the cost matrices obtained in S63 as the input of the Hungarian algorithm to get a linear matching result; the matching result includes three cases: track pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track pipeline. Among them, when the matching result is track pipeline mismatch, directly delete the mismatched track pipeline (at this time, the track pipeline is in an uncertain state); when the matching result is detection mismatch, initialize it as a new track pipeline; when the matching result is successful pairing of detection and predicted track pipeline, it means that the current image and the previous image are successfully tracked, and update the corresponding track pipeline variables of the detection through the update equation of the Kalman filter. The update expression of the Kalman filter is as follows:

[0063] ;

[0064] Among them, is the estimated state vector at time is the Kalman gain, is the measurement value at time is the measurement matrix;

[0065] S65. Predict the positions corresponding to the confirmed track pipeline and the unconfirmed track pipeline through the prediction equation of the Kalman filter; the predicted positions of the confirmed track pipeline and the detection results perform cascaded matching;

[0066] S66. Obtain two results of cascaded matching, namely: trajectory pipeline matching and detection and trajectory pipeline mismatch; among them, when the matching result is trajectory pipeline matching, update the corresponding trajectory pipeline variables through the update equation of Kalman filtering; when the matching result is detection and trajectory pipeline mismatch, the trajectory pipeline in the unconfirmed state of S65 and the mismatched trajectory pipeline are together matched with the detection results that have not been successfully matched by IOU, and then the cost matrix is calculated through the result of IOU matching , and the calculation method is the same as S63;

[0067] S67. Use the obtained in S66 as the input of the Hungarian algorithm to obtain a linear matching result. At this time, the matching result still includes three situations: trajectory pipeline mismatch, detection mismatch, and successful pairing of detection and predicted trajectory pipeline. The corresponding processing methods are the same as S64;

[0068] S68. Repeatedly loop S65 to S67 until the entire image sequence ends.

[0069] Therefore, the present invention adopts the above-mentioned target robust detection method based on a brain cognitive model, and has the following beneficial effects:

[0070] (1) Improve the accuracy of target detection: By parallelly calculating the attention weights of different dimensions of the convolution kernel in the image feature extraction module that simulates the visual information attention perception function, enhance the weak structural features of small targets, suppress background interference, and thus improve the accuracy of target detection;

[0071] (2) Enhance the interpretability of the neural network: Draw on the working mode of the brain to construct a brain-like artificial intelligence, making the neural network have natural interpretability. Compared with traditional deep neural networks, it can better understand its decision-making process and basis;

[0072] (3) Reduce the computational complexity: The block-aware channel attention method that simulates the visual spatial attention function samples each channel with a large-scale convolution kernel, emphasizes the importance of spatial blocks in the feature map, extracts effective features through pooling operations and transforms them into spatial point features, reduces the size of the feature map, and reduces the computational complexity;

[0073] (4) Improve the ability of the network to identify effective features: The block-aware channel attention module independently models the feature map of each channel, analyzes it separately in a targeted manner, obtains the channel importance and assigns weights, performs channel fusion, and then fuses the channel importance and spatial importance of all feature maps, greatly improving the ability of the network to identify effective features;

[0074] (5) Implement robust target detection and tracking in complex environments: The prediction and inference module that simulates the visual prediction and inference function, combined with motion and appearance information, as well as motion compensation and a more accurate Kalman filter state vector, can achieve robust target detection and tracking in complex environments.

[0075] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0076] Figure 1 is the visual cognitive function architecture diagram in the background art;

[0077] Figure 2 is the overall flowchart of a method for robust target detection based on a brain cognitive model of the present invention;

[0078] Figure 3 is a schematic diagram of an image feature extraction module that simulates the visual information attention perception function in an embodiment of the present invention;

[0079] Figure 4 is a schematic diagram of an image feature extraction module that simulates the visual spatial attention function in an embodiment of the present invention;

[0080] Figure 5 is a schematic diagram of a perception-criterion feature association and fusion module that simulates the visual and memory association function in an embodiment of the present invention;

[0081] Figure 6 is a schematic diagram of a prediction and inference module that simulates the visual prediction and inference function in an embodiment of the present invention. Detailed Embodiments

[0082] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0083] Please refer to Figures 1-6 , a method for robust target detection based on a brain cognitive model, comprising the following steps:

[0084] Step 1, input paired real images and standard images;

[0085] Step 2: Use the image feature extraction module that simulates the attention perception function of visual information to extract primary perception features from the input real image. The image feature extraction module that simulates the attention perception function of visual information can calculate the attention weights of different dimensions of the convolution kernel in parallel, enhancing the weak structural features of small targets and suppressing background interference while introducing a relatively low additional computational cost. The specific process is as follows:

[0086] S21. For the input real image , where is the number of image channels, is the image width, is the image height. After passing through a global average pooling layer to compress the spatial dimension, it is then input into a fully connected layer and a ReLU activation function. The expression is as follows:

[0087] ;

[0088] Among them, , represents the intermediate feature, represents the ReLU activation function, represents the fully connected layer, represents the global average pooling layer;

[0089] S22. Based on the intermediate feature obtained in S21, obtain the corresponding attention weights through the position branch, channel branch, and filter branch respectively based on the Sigmoid function. The calculation expressions are as follows:

[0090] ;

[0091] In the formula, , , respectively represent the attention weight of the position branch, the attention weight of the channel branch, and the attention weight of the filter branch, is the convolution kernel size, is the number of input channels, is the number of output channels, represents the position multi-layer perceptron, represents the channel multi-layer perceptron, represents the filter multi-layer perceptron;

[0092] S23. Based on the branch attention weight, channel branch attention weight, filter branch attention weight, and real image, obtain the primary perception feature , and the calculation expression is as follows:

[0093] ;

[0094] In the formula, Indicates the weight coefficient.

[0095] Step 3: Use the image feature extraction module that simulates the visual spatial attention function to extract the target perception features from the primary perception features; the image feature extraction module that simulates the visual spatial attention function can eliminate the interference of negative features. By independently modeling the feature maps of each channel, the image feature extraction module can perform targeted individual analysis on the feature maps of different channels. On this basis, the image feature extraction module uses a large-scale convolutional kernel to sample each channel, abandoning the strategy of the traditional attention mechanism that over-considers the microscopic relationship between pixels. Instead, the image feature extraction module emphasizes the importance of spatial blocks in the feature map to highlight the relatively important parts in the spatial domain, that is, the effective features of the object. Secondly, pooling operations are used to extract these effective features from the spatial blocks and convert them into spatial point features, reducing the size of the feature map and the computational complexity; then, the image feature extraction module obtains the channel importance and assigns weights, and performs channel fusion through a fully connected layer. Finally, the channel importance and spatial importance of all the obtained feature maps are fused; the image feature extraction module greatly improves the network's ability to identify effective features; the specific process is as follows:

[0096] S31. Use a large-scale convolutional kernel to highlight the effective features of the original features, suppress the invalid features, and then use the Sigmoid function to obtain the spatial block features so as to allow the network to focus on more critical effective regions, and the calculation expression is as follows:

[0097] ;

[0098] In the formula, represents the features after grouped convolution, represents the grouped convolution operation;

[0099] S32. Fuse the spatial block features with the original features and use average pooling to obtain the spatial point features , and the calculation expression is as follows:

[0100] ;

[0101] S33. Reconstruct the channel features with a fully connected layer to obtain the channel-level importance of all feature maps, and the calculation expression is as follows:

[0102] ;

[0103] In the formula, represents the obtained channel features, represents the activation function,​ Indicates the fully connected channel layer;

[0104] S34. Multiply the result obtained in S33 by the original feature to obtain the target perception feature , and the expression is as follows:

[0105] .

[0106] Step 4. Use the standard visual memory library module that simulates the visual memory function to extract the target memory features from the standard image; inspired by the semantic memory function of the anterior temporal lobe in the visual perception process, a standard visual memory library module is introduced in the network architecture to simulate visual memory. The standard visual memory library module takes the standard image with sufficient and complete information as the input and maintains the same architecture as the image perception feature extraction module; the specific expression is as follows:

[0107] ;

[0108] In the formula, is the input standard image, is the feature extractor, are the decoding structures for the category and location respectively, are the category and location predictions obtained respectively, are the target category memory feature and the target location memory feature respectively.

[0109] Step 5. Input the target perception feature and the target memory feature into the perception-standard feature association and fusion module that simulates the visual and memory association function, and use the domain adaptation technology to perform interactive fusion learning on the target perception feature and the target memory feature to form a fusion feature, forcing the network to strengthen the adaptation to the regions with more key information; inspired by the visual feature and memory integration function of the inferior longitudinal fasciculus, an association and fusion module based on the domain adaptation technology is introduced in the network architecture to integrate the perception feature and the standard memory feature. During the training process, the perception feature encoded by the image perception feature extraction module is adaptively mapped to the memory feature generated by the memory feature generation module through domain adaptation, so as to establish a feature association; among them, the process of performing interactive fusion learning on the target perception feature and the target memory feature to form a fusion feature is as follows:

[0110] S51. Input the target perception feature extracted in step 3 into the position and category feature decoupler , and obtain the target category perception feature and the target position perception feature , and the calculation expression is as follows:

[0111] ;

[0112] S52. Take the average of the target category perception feature and the target category memory feature along the channel dimension, and multiply them with the foreground mask respectively to obtain the foreground response map. The calculation expression is as follows:

[0113] ;

[0114] In the formula, represents the target category perception response feature, represents the target category memory response feature, is the number of channels;

[0115] S53. Subtract the two obtained foreground response maps and normalize them to (0, 1) to obtain the spatial reweighting map , and the expression is as follows:

[0116] ;

[0117] In the formula, represents the operation of normalizing the mapping to 0 - 1;

[0118] S54. Feed the difference between the target category perception feature and the target category memory feature into the C - R module to obtain the channel reweighting vector , and the expression is as follows:

[0119] ;

[0120] In the formula, respectively represent average pooling, fully connected layer, and softmax operations;

[0121] S55. Multiply the channel reweighting vector and the spatial reweighting map to obtain the domain adaptation weight , and the expression is as follows:

[0122] ;

[0123] S56. Multiply the domain adaptation weight with the target location perception feature to obtain the final fused feature.

[0124] Step 6: Input the fused features into the prediction and inference module that simulates the visual prediction and inference function, and use Kalman filtering to complete the position prediction of the target; inspired by the prefrontal visual perception and inference function in the high-level visual cortex, a perception and inference module is added to the network model to simulate the inference function based on visual memory and semantic knowledge, so as to simulate the visual inference function of the prefrontal lobe in the high-level visual cortex; combining the advantages of motion and appearance information, as well as motion compensation and a more accurate Kalman filter state vector, to achieve robust detection and tracking of the target in complex environments; the specific process is as follows:

[0125] S61. Obtain the detection result of the current image; input the fused features into the detection module , and obtain the detection result of the current image ;

[0126] S62. Initialize the trajectory pipeline and position prediction; create its corresponding trajectory pipeline for the result detected in the first image ; initialize the motion variables of the Kalman filter , and predict its corresponding position through the prediction equation of the Kalman filter; at this time, the state of the trajectory pipeline is marked as undetermined; the prediction equation of the Kalman filter is:

[0127] ;

[0128] where is the predicted state vector at time , is the state transition matrix, is the estimated state vector at time , is the control input matrix, is the control input vector, is the process noise;

[0129] S63. IOU matching and cost matrix calculation; perform IOU matching on the result of the current image target detection and the position predicted by the trajectory pipeline in the previous image , and then calculate its cost matrix ; assume that there are detection results in the current image, and there are positions predicted by the trajectory pipeline in the previous image, calculate a cost matrix of , and use as the cost, then the calculation formula is as follows:

[0130] ;

[0131] where​​ denote and the intersection over union, denote the cost between the -th detection result and the -th predicted position; the areas of two bounding boxes and are and respectively, the intersection area is and the union area is ; denote the bounding box corresponding to the -th detection result in the current frame, denote the bounding box corresponding to the -th predicted position in the previous frame, , ;

[0132] S64. Take all the cost matrices obtained in S63 as the input of the Hungarian algorithm to obtain a linear matching result; the matching result includes three cases: track pipeline mismatch, detection mismatch, and successful pairing of detection and predicted track pipeline. Among them, when the matching result is track pipeline mismatch, directly delete the mismatched track pipeline (at this time, the track pipeline is in an uncertain state); when the matching result is detection mismatch, initialize it as a new track pipeline; when the matching result is successful pairing of detection and predicted track pipeline, it means that the current image and the previous image are successfully tracked, and update the corresponding track pipeline variable of the detection through the update equation of the Kalman filter. The update expression of the Kalman filter is as follows:

[0133] ;

[0134] Among them, is the estimated state vector at time, is the Kalman gain, is the measurement value at time,

[0135] S65. Predict the positions corresponding to the confirmed track pipeline and the unconfirmed track pipeline through the prediction equation of the Kalman filter; cascade-match the predicted position of the confirmed track pipeline and the detection result ;

[0136] S66. Obtain two cascade matching results, namely: trajectory pipeline matching and detection and trajectory pipeline mismatch; among them, when the matching result is trajectory pipeline matching, update the corresponding trajectory pipeline variables through the update equation of Kalman filter; when the matching result is detection and trajectory pipeline mismatch, the unconfirmed trajectory pipeline in S65 and the mismatched trajectory pipeline are subjected to IOU matching with the detection results that have not been successfully matched, and then calculate their cost matrix based on the results of the IOU matching , and the calculation method is the same as that in S63;

[0137] S67. Use the result obtained in S66 as the input of the Hungarian algorithm to obtain a linear matching result. At this time, the matching result still includes three situations: trajectory pipeline mismatch, detection mismatch, and successful pairing of detection and predicted trajectory pipelines. The corresponding processing methods are the same as those in S64;

[0138] S68. Repeatedly loop from S65 to S67 until the entire image sequence ends.

[0139] Therefore, the present invention adopts the above-mentioned target robust detection method based on the brain cognitive model. Starting from the entire process of the biological brain processing natural environment visual images, an image feature extraction module that simulates the visual information attention perception function, an image feature extraction module that simulates the visual spatial attention function, a standard visual memory library module that simulates the visual memory function, a perception-standard feature association and fusion module that simulates the visual and memory association function, and a prediction and reasoning module that simulates the visual prediction and reasoning function are respectively designed; compared with the traditional target detection method based on deep neural networks, the present invention effectively improves the accuracy and robustness of target detection in complex environments, and at the same time, based on the biological brain structure, makes the neural network have a certain interpretability.

[0140] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A target robust detection method based on a brain cognitive model, characterized in that, It includes the following steps: Step 1, input paired real images and standard images; Step 2, use an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real images; Step 3, use an image feature extraction module that simulates the visual spatial attention function to extract target perception features from the primary perception features; Step 4, use a standard visual memory library module that simulates the visual memory function to extract target memory features from the standard images; Step 5, input the target perception features and target memory features into a perception-standard feature association and fusion module that simulates the visual and memory association function, and use domain adaptation technology to perform interactive fusion learning on the target perception features and target memory features to form fusion features; Step 6, input the fusion features into a prediction and reasoning module that simulates the visual prediction and reasoning function, and use Kalman filtering to complete the position prediction of the target; The expression for using the standard visual memory library module that simulates the visual memory function to extract target memory features from the standard images in Step 4 is as follows: ; In the formula, is the input standard image, is the feature extractor, are the decoding structures for the class and location respectively, are the obtained class and location predictions respectively, are the target class memory feature and the target location memory feature respectively; The process of inputting the target perception features and target memory features into a perception-standard feature association and fusion module that simulates the visual and memory association function, and using domain adaptation technology to perform interactive fusion learning on the target perception features and target memory features to form fusion features in Step 5 is as follows: S51. Input the target perception features extracted in step 3 into the position and category feature decoupler to obtain the target category perception feature and the target position perception feature . The calculation expression is as follows: ; S52. Take the average value of the target category perception feature and the target category memory feature along the channel dimension, and multiply them with the foreground mask respectively to obtain the foreground response map. The calculation expression is as follows: ; In the formula, represents the target category perception response feature, represents the target category memory response feature, is the number of channels; S53. Subtract the two obtained foreground response maps and normalize them to (0, 1) to obtain the spatial re-weighting map , and the expression is as follows: ; Wherein, represents an operation of normalizing the mapping to 0-1; S54. Feed the difference between the target category perception feature and the target category memory feature into the C-R module to obtain the channel reweighting vector . The expression is as follows: ; In the formula, represent average pooling, fully connected layer, and softmax operation respectively; S55. Multiply the channel re-weight vector and the spatial re-weight map to obtain the domain adaptation weight , and the expression is as follows: ; S56. Multiply the domain adaptation weight with the target location-aware feature to obtain the final fused feature.

2. The target robust detection method based on a brain cognitive model according to claim 1, wherein, The process of using an image feature extraction module that simulates the visual information attention perception function to extract primary perception features from the input real images in Step 2 is as follows: S21. For the input real image , where is the number of image channels, is the image width, is the image height. After passing through a global average pooling layer to compress the spatial dimension, it is then input into a fully connected layer and a ReLU activation function. The expression is as follows: ; Among them, , represents an intermediate feature, represents the ReLU activation function, represents a fully connected layer, represents a global average pooling layer; S22, based on the Sigmoid function, obtain the corresponding attention weights for the intermediate features obtained in S21 through the position branch, channel branch, and filter branch respectively. The calculation expression is as follows: ; Wherein, , , respectively represent the position branch attention weight, the channel branch attention weight, and the filter branch attention weight, is the convolution kernel size, is the number of input channels, is the number of output channels, represents the position multi-layer perceptron, represents the channel multi-layer perceptron, represents the filter multi-layer perceptron; S23. Obtain the primary perception features based on the branch attention weight, channel branch attention weight, filter branch attention weight, and real image. , and the calculation expression is as follows: ; In the formula, represents the weight coefficient.

3. A target robust detection method based on a brain cognitive model according to claim 2, characterized in that, The process of using an image feature extraction module that simulates the visual spatial attention function to extract target perception features from the primary perception features in Step 3 is as follows: S31. Use a large-scale convolutional kernel to highlight the effective features of the original features, suppress the invalid features, and then use the Sigmoid function to obtain the spatial block features The calculation expression is as follows: The calculation expression is as follows: ; In the formula, represents the feature after grouped convolution, represents the grouped convolution operation; S32. Fuse the spatial block features with the original features using average pooling to obtain the spatial point features . The calculation expression is as follows: ; S33, reconstruct the channel features with a fully connected layer to obtain the channel-level importance of all feature maps. The calculation expression is as follows: ; In the formula, represents the obtained channel feature, represents the activation function, represents the fully connected channel layer; S34. Multiply the obtained in S33 by the original feature to obtain the target perception feature . The expression is as follows: 。 4. The target robust detection method based on a brain cognitive model according to claim 3, wherein: The process of inputting the fusion features into a prediction and reasoning module that simulates the visual prediction and reasoning function, and using Kalman filtering to complete the position prediction of the target in Step 6 is as follows: S61. Obtain the detection result of the current image; input the fused feature into the detection module , and obtain the detection result of the current image ; S62, initialize the trajectory pipeline and predict the position; The results detected from the first picture Create its corresponding trajectory pipeline ; The motion variables of the Kalman filter are initialized, and their corresponding positions are predicted through the prediction equation of the Kalman filter; at this time, the status of the trajectory pipeline is marked as undetermined; the prediction equation of the Kalman filter is: ; Among them, is the predicted state vector at a moment, is the state transition matrix, is the estimated state vector at a moment, is the control input matrix, is the control input vector, is the process noise; S63, IOU matching and cost matrix calculation; the results of the current image object detection and the positions predicted by the trajectory pipeline for the previous image are subjected to IOU matching, and then the cost matrix is calculated based on the results of the IOU matching ; assume that there are detection results in the current image, and there are positions predicted by the trajectory pipeline for the previous image. Calculate a cost matrix , using as the cost, and the calculation formula is as follows: ; Among them, represents and the intersection over union, represents the cost between the th detection result and the th predicted position; the areas of two bounding boxes and are and respectively, the intersection area is and the union area is ; represents the bounding box corresponding to the th detection result in the current frame, represents the bounding box corresponding to the th predicted position in the previous frame, , ; S64. Take all the cost matrices obtained in S63 as the input of the Hungarian algorithm to obtain a linear matching result; the matching result includes three cases: trajectory-pipeline mismatch, detection mismatch, and successful pairing of detection and predicted trajectory pipelines. Among them, when the matching result is a trajectory-pipeline mismatch, directly delete the mismatched trajectory pipeline; when the matching result is a detection mismatch, initialize it as a new trajectory pipeline; when the matching result is successful pairing of detection and predicted trajectory pipelines, update the corresponding trajectory pipeline variables of the detection through the update equation of the Kalman filter. The update expression of the Kalman filter is as follows: ; Among them, is the estimated state vector at a moment, is the Kalman gain, is the measured value at a moment, is the measurement matrix; S65. Predict the trajectory pipelines in the confirmed state and the trajectory pipelines in the unconfirmed state through the prediction equation of the Kalman filter and the corresponding positions; cascade and match the predicted positions of the trajectory pipelines in the confirmed state with the detection results ; ​ S66. Obtain two results of cascade matching, namely: trajectory pipeline matching and detection and trajectory pipeline mismatch. Among them, when the matching result is trajectory pipeline matching, update the corresponding trajectory pipeline variables through the update equation of Kalman filtering; when the matching result is detection and trajectory pipeline mismatch, the trajectory pipeline in the unconfirmed state of S65 and the mismatched trajectory pipeline are jointly subjected to IOU matching with the detection results that have not been successfully matched, and then calculate their cost matrix based on the results of IOU matching ; S67. Use the result obtained in S66 as the input of the Hungarian algorithm to obtain a linear matching result. At this time, the matching result still includes three cases: trajectory pipeline mismatch, detection mismatch, and successful pairing of the detected and predicted trajectory pipelines; S68, repeatedly loop through S65 to S67 until the entire image sequence ends.

Citation Information

Patent Citations

  • Video perception-fused multi-task synergetic recognition method and system

    CN108846384A

  • Multi-visual memory unit-based glancing path prediction method

    CN116563524A