Dynamic object recognition brain-like modeling method based on visual dual-channel
By simulating the dual-path mechanism of the human vision system, combined with VideoSwin Transformer and ResNet networks, the problem of insufficient recognition capabilities of computer vision models in dynamic scenarios is solved, and more efficient and interpretable dynamic object recognition is achieved, providing a neuroscientific basis.
Patent Information
- Application Number
- CN202510490215.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
The existing computer vision models are poor in handling dynamic scenarios, and it is difficult to simulate the dual-path mechanism of human vision systems, and there are problems such as high data and computing costs and poor interpretability.
A dynamic object recognition brain-like modeling method based on visual dual pathways is adopted. By simulating the ventral and dorsal pathways of the human visual system, the spatiotemporal motion characteristics and static characteristics are extracted using VideoSwin Transformer and ResNet convolutional neural networks, and feature fusion is performed through a cross-modal neural fusion module, combining task-driven bionic training strategies and cortical topology regularization constraints, model training is optimized.
It improves the accuracy of object recognition in dynamic scenarios, enhances the interpretability of the model, and provides a neuroscientific basis for the design of brain-like intelligent systems.
Smart Images

Figure CN120375153A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision and artificial intelligence, and particularly relates to a brain-inspired modeling method for dynamic object recognition based on visual dual pathways. Background Art
[0002] Vision is the main sensory channel for humans to obtain external information, and object recognition is one of the core functions of the visual system. In the past few decades, research on the primate visual cortex has revealed the "dual pathway" theory of visual information processing: the ventral pathway ("what" pathway) mainly processes features such as the shape and color of objects, while the dorsal pathway ("where / how" pathway) focuses on processing spatial location and motion information. These two pathways converge in the parietal lobe region of the brain to integrate and form a complete perception of dynamic objects.
[0003] Existing computer vision models mainly focus on single-modal feature extraction. Although deep learning techniques, especially convolutional neural networks (CNNs) and the recent Vision Transformer (ViT), have made significant progress in static image recognition tasks, they still have obvious deficiencies in processing dynamic scenes. These models are difficult to effectively simulate the functional specific processing mechanism of the human visual system for dynamic objects, resulting in a decline in recognition accuracy in complex and rapidly changing scenes, especially for targets with specific motion patterns, such as fast-moving humans or faces with changing expressions. At the same time, traditional dynamic object recognition models such as Two-Stream Networks and 3D CNNs, although introducing the time dimension, lack biological rationality and cannot truly reflect the information processing mechanism of the cerebral cortex. In addition, these models usually require a large amount of training data, have high computational overhead, and are difficult to explain their internal decision-making processes, limiting their application in scenarios that require high interpretability.
[0004] Therefore, there is a need to develop a model architecture that can effectively recognize dynamic objects and has biological rationality to better simulate the working mechanism of the human visual system and improve the object recognition performance in dynamic scenes. Summary of the Invention
[0005] To solve the above problems in the prior art, that is, the poor ability of existing computer vision models to process dynamic scenes, the difficulty of single-modal models in simulating the human visual mechanism, the lack of biological rationality in traditional dynamic recognition models, and the common problems of high data and computational costs and poor interpretability, the present invention provides a brain-inspired modeling method for dynamic object recognition based on visual dual pathways, which includes the following steps:
[0006] Step S1, obtaining a video frame segment for dynamic object recognition, and performing frame extraction on the video frame segment to obtain single-frame static images and temporally continuous optical flow images;
[0007] Step S2: Input the single-frame static image and the temporally continuous optical flow images into the trained bionic two-stream architecture network to obtain the recognition result of the dynamic object;
[0008] The bionic two-stream architecture network includes a dorsal pathway module, a ventral pathway module, a two-branch interaction module, and a cross-modal neural fusion module;
[0009] The dorsal pathway module is configured to use the hierarchically optimized VideoSwin Transformer for the temporally continuous optical flow images to extract spatio-temporal motion features from local to global, and generate a spatio-temporal dynamic representation similar to the motion energy map of the MT area;
[0010] The ventral pathway module is configured to use the ResNet convolutional neural network for the single-frame static image to simulate the hierarchical selective representation characteristics of the IT area for static shapes and extract static features;
[0011] The two-branch interaction module is configured to use both the spatio-temporal motion features and the static features as original features; perform dimensionality transformation on each original feature, and after dimensionality transformation, send it to the Block block in the opposite pathway for processing to obtain the interactively generated features; fuse each interactively generated feature with the corresponding original feature to obtain the residually fused features, and the residually fused features include residually fused static features and residually fused spatio-temporal dynamic features;
[0012] The cross-modal neural fusion module is configured to fuse the residually fused static features and the residually fused spatio-temporal dynamic features respectively to obtain fused features; process the fused features through a multi-layer perceptron to obtain the recognition result of the dynamic object.
[0013] In some preferred embodiments, the method for frame extraction of the video frame segment to obtain a single-frame static image and temporally continuous optical flow images is as follows:
[0014] Randomly extract one frame from the video frame segment as the single-frame static image;
[0015] Perform frame extraction on the video frame segment through an optical flow estimation algorithm to obtain temporally continuous optical flow images; the optical flow estimation algorithm includes the Lucas-Kanade optical flow estimation algorithm.
[0016] In some preferred embodiments, the hierarchically optimized Video Swin Transformer is a Transformer encoder including multiple stages, and each stage includes at least two Swin-Transformer blocks.
[0017] In some preferred embodiments, the method for obtaining the residually fused features is as follows:
[0018] Perform dimensionality transformation on the spatio-temporal motion features and the static features through a multi-layer perceptron to obtain the spatio-temporal motion features X' after dimensionality transformation swin , static features X' resnet , and the expression is:
[0019] X' swin = MLP1(X swin )
[0020] X' resnet = MLP2(X resnet )
[0021] where X swin and X resnet respectively represent the extracted spatio-temporal motion features and static features; X' swin is the feature after dimensionality transformation of the spatio-temporal motion features; X' resnet is the feature after dimensionality transformation of the static features; the dimension of X' swin is the same as the dimension of X resnet , and the dimension of X' resnet is the same as the dimension of X swin ;
[0022] Send the spatio-temporal motion features X' swin and static features X' resnet after dimensionality transformation into the Blcok block of the contralateral pathway for processing to obtain the interactively generated features:
[0023]
[0024]
[0025] where represents the interactively generated spatio-temporal motion features, represents the interactively generated static features;
[0026] Fuse each interactively generated feature with the corresponding original feature to obtain the residually fused features:
[0027]
[0028] where Y swin represents the residually fused spatio-temporal motion features, and Y resnet represents the residually fused static features.
[0029] In some preferred embodiments, fuse the residually fused static features and the residually fused spatio-temporal dynamic features to obtain the fused features, and the method is as follows:
[0030] The residual fusion static features and the residual fusion spatio-temporal dynamic features are respectively input into a fully connected layer to be mapped to the same dimension, and global average pooling is performed on the features mapped to the same dimension to obtain fused features.
[0031] In some preferred embodiments, the training method of the bionic two-stream architecture network is as follows:
[0032] It is trained using a task-driven bionic training strategy, and the training strategy includes a functional decoupling training protocol and a cortical topology regularization constraint.
[0033] In some preferred embodiments, the bionic two-stream architecture network is trained using a functional decoupling training protocol, and the method is as follows:
[0034] The first stage: Freeze the dorsal pathway module and only optimize the parameters of the ventral pathway module;
[0035] The second stage: After the first stage is completed, unfreeze the frozen network parameters, add adversarial samples to the training dataset, and train the bionic two-stream architecture network; the adversarial samples include noise samples with motion noise opposite to the main motion direction added to video frame segments in the training dataset and samples with static occlusions.
[0036] In some preferred embodiments, the cortical topology regularization constraint is specifically: introducing a cortical similarity loss term into the loss function to constrain the geometric structure consistency between the model parameters and the fMRI responses of the primate visual cortex through contrastive learning;
[0037] The loss function includes a classification loss function and a cortical similarity loss function.
[0038] In some preferred embodiments, the loss function is:
[0039] L = (L cls + L cortex ) / 2;
[0040] In the formula, L cls is the classification loss function, and L cortex is the cortical similarity loss function.
[0041] In some preferred embodiments, the cortical similarity loss function:
[0042]
[0043] r model (i, j) = Rank soft (D model (i, j); {D model(m, n)};
[0044] r brain (i, j) = Rank soft (D brain (i, j); {D brain (m, n)};
[0045] Among them, M represents the number of sample pairs; rmodel(i, j) represents the soft percentile value of the distance Dmodel(i, j) between sample i and sample j in the model feature space, which is calculated based on the set of distances {Dmodel(m, n)} of all sample pairs in the model feature space; rbrain(i, j) represents the soft percentile value of the distance Dbrain(i, j) between sample i and sample j in the brain visual system signal space, which is calculated based on the set of distances {Dbrain(m, n)} of all sample pairs in the brain visual system signal space.
[0046] Advantages of the present invention:
[0047] By simulating the ventral pathway and dorsal pathway of the human visual system, the dorsal pathway adopts a hierarchical VideoSwin-Transformer to input optical flow images, which can effectively extract spatio-temporal motion features and generate relevant spatio-temporal dynamic representations. The ventral pathway uses ResNet to input single-frame static images, which can extract static features and simulate the representation characteristics of the IT area for static shapes. Then, through a biologically inspired neural fusion mechanism, that is, feature interaction and cross-modal interaction, a connection with the real visual cortex response is established. Using a task-driven bionic training strategy, that is, the experimental data of fMRI of the human brain to guide the training and optimization of the model, the feature space extracted by the model is made closer to the brain, aiming to improve the accuracy of video analysis, enhance the interpretability of the model, and provide a neuroscience basis for the design of brain-like intelligent systems. Description of the Drawings
[0048] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more obvious:
[0049] Figure 1 is a module diagram of the brain-like modeling method for dynamic object recognition based on the visual dual pathway of the present invention.
[0050] Figure 2 is a bionic two-stream architecture model diagram of the brain-like modeling method for dynamic object recognition based on the visual dual pathway of the present invention.
[0051] Figure 3 is a cross-modal neural fusion module diagram of the brain-like modeling method for dynamic object recognition based on the visual dual pathway of the present invention. Detailed Embodiments
[0052] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the relevant invention are shown in the drawings.
[0053] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0054] For a clearer description of the brain-inspired modeling method for dynamic object recognition based on the visual dual pathway of the present invention, the following will be combined with Figures 1 to 3 Each step in the embodiments of the present invention will be described in detail.
[0055] The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway in the first embodiment of the present invention, refer to Figure 1 , the method includes the following steps:
[0056] Step S1, obtain a video frame segment for dynamic object recognition, and perform frame extraction on the video frame segment to obtain a single-frame static image and a temporally continuous optical flow image;
[0057] In this embodiment, the method for performing frame extraction on the video frame segment to obtain a single-frame static image and a temporally continuous optical flow image is as follows:
[0058] Randomly extract one frame from the video frame segment as the single-frame static image;
[0059] Perform frame extraction on the video frame segment through an optical flow estimation algorithm to obtain a temporally continuous optical flow image; the optical flow estimation algorithm includes the Lucas-Kanade optical flow estimation algorithm;
[0060] The sources of video frame segments are, on the one hand, the video sequence data collected, which contains various dynamic objects, and on the other hand, from multiple public databases, such as The Weizmann Dataset, training_lib_KTH, IAS-Lab_Action_Dataset, Ixmas, XD145_actions, Jester gesture dataset, and HMDB51 dataset. To ensure the reliability of data quality and the consistency of data format, the video data is preprocessed to obtain video frame segments, and the video frame segments are frame-sampled to obtain single-frame static images and temporally continuous optical flow images. Among them, the preprocessing includes video frame extraction, spatio-temporal normalization, and illumination correction. Specifically, considering that the video data contains two key information, spatial form and temporal dynamics (spatial form information represents the static appearance characteristics of objects, and temporal dynamics information reflects the motion patterns of objects), the length of the video segment is set to 16 frames. For the case where the number of frames in the original video segment exceeds 16 frames, a uniform sampling strategy is used to extract 16 frames; if the original number of frames is less than 16 frames, the cyclic padding method is used to meet the length requirement; each frame of the image is uniformly adjusted to a resolution of 224×224 pixels and normalized to ensure the consistency of data in terms of quality and specifications, laying a good data foundation for subsequent model training.
[0061] To verify the biological interpretability of the model, functional magnetic resonance imaging (fMRI) brain imaging data of primates when watching the same video stimuli are further collected to construct a neural response database. The data collected mainly involves the blood oxygenation level-dependent (BOLD) signals in the visual cortex regions of V1, V2, V4, and MT areas.
[0062] Step S2, input the single-frame static image and the temporally continuous optical flow image into the trained bionic two-stream architecture network to obtain the recognition result of the dynamic object.
[0063] See Figure 2 , the bionic two-stream architecture network includes a dorsal pathway module, a ventral pathway module, a two-branch interaction module, and a cross-modal neural fusion module.
[0064] The dorsal pathway module is configured to extract spatio-temporal motion features from local to global for the temporally continuous optical flow image using a hierarchically optimized VideoSwin Transformer, and generate a spatio-temporal dynamic representation similar to the motion energy map in the MT area.
[0065] In this embodiment, the hierarchically optimized Video Swin Transformer is a Transformer encoder including multiple stages, and each stage includes at least two Swin-Transformer blocks.
[0066] Specifically, since the dorsal pathway is designed to process motion information, the temporally continuous optical flow images are input into a hierarchically optimized Video Swin-Transformer encoder, which consists of four stages of Transformer encoders connected in sequence. The connection method between stages is sequential, that is, the output of the previous stage is used as the input of the next stage. Each stage contains 2, 2, 6, and 2 Swin-Transformer blocks respectively. After the optical flow images are input, they are first processed by the patch embedding layer, which converts the images into serialized feature vectors. Subsequently, the window coverage range is dynamically adjusted by means of the window shift mechanism for feature extraction. In this process, through the sequential operations of the Transformer encoders in each stage, spatio-temporal motion features are gradually extracted from local to global, and finally a spatio-temporal dynamic representation similar to the motion energy map of the MT area is generated;
[0067] The temporally continuous optical flow images are a video frame optical flow sequence of T×H×W×3, where T represents the number of time frames of the video frames, H is the image height, W is the image width, and 3 represents the number of image channels;
[0068] The ventral pathway module is configured to use a ResNet convolutional neural network for the single-frame static image to simulate the hierarchical selective representation characteristics of the IT area for static shapes and extract static features;
[0069] Since the ventral pathway focuses on processing static shape information, a ResNet convolutional neural network is used for the single-frame static image to simulate the hierarchical selective representation characteristics of the IT area for static shapes and extract static features. Specifically, the single-frame static image is input into the ResNet convolutional neural network as the basic encoder, and the single-frame image is subjected to feature extraction through convolutional operations, batch normalization processing, and residual connections to obtain static features;
[0070] The single-frame static image is a single-frame static image randomly extracted from the video frame segment with a size of 1×H×W×3, where 1 represents a single frame; H and W respectively represent the height and width of the image; 3 represents the number of image channels;
[0071] To achieve effective information exchange and fusion between two different encoder (architecture) visual processing streams, a design is made to bridge the dorsal pathway based on Video Swin Transformer and the ventral pathway based on ResNet, enhancing the model's representation ability through a feature interaction mechanism. Specifically, a dual-branch interaction module is set up. The dual-branch interaction module is configured to use both the spatio-temporal motion features and the static features as original features; perform dimensionality transformation on each original feature after each stage of the encoder, and send the transformed features into the Block of the opposite pathway for processing to obtain the interactively generated features; fuse each interactively generated feature with the corresponding original feature to obtain the residually fused features, where the residually fused features include residually fused static features and residually fused spatio-temporal dynamic features.
[0072] In this embodiment, the method for obtaining the residually fused features is as follows:
[0073] Perform dimensionality transformation on the spatio-temporal motion features and the static features through a multi-layer perceptron to obtain the spatio-temporal motion features X′ swin and static features X′ resnet , and the expression is:
[0074] X′ swin = MLP1(X swin )
[0075] X′ resnet = MLP2(X resnet )
[0076] where X swin and X resnet respectively represent the extracted spatio-temporal motion features and static features; X′ swin is the feature after dimensionality transformation of the spatio-temporal motion features; X′ resnet is the feature after dimensionality transformation of the static features; the dimension of X′ swin is the same as the dimension of X resnet , and the dimension of X′ resnet is the same as the dimension of X swin .
[0077] Send the dimensionality-transformed spatio-temporal motion features X′ swin and static features X′ resnet into the Block of the opposite pathway for processing to obtain the interactively generated features:
[0078]
[0079]
[0080] where Represents the spatio-temporal motion features generated by interaction, Represents the static features generated by interaction;
[0081] Fuse each interaction-generated feature with the corresponding original feature to obtain the residual-fused feature:
[0082]
[0083] Among them, Y swin Represents the residual-fused spatio-temporal motion feature, and Y resnet Represents the residual-fused static feature;
[0084] This residual fusion mechanism enables each encoder branch to effectively absorb and integrate complementary information from the other branch while maintaining its inherent feature representation ability, thereby enhancing the richness and expressiveness of the overall feature representation;
[0085] To reproduce the functional integration between cortical regions, an adaptive gating attention mechanism is introduced in the parietal simulation layer through a cross-modal neural fusion module, and bidirectional cross-modal interaction of motion-shape features is achieved through dynamic weight allocation to reproduce the functional integration mechanism between cortical regions. Specifically, see Figure 3 The cross-modal neural fusion module is configured to fuse the residual-fused static feature and the residual-fused spatio-temporal dynamic feature respectively to obtain a fused feature; the fused feature is processed through a multi-layer perceptron to obtain the recognition result of the dynamic object;
[0086] In this embodiment, the method for fusing the residual-fused static feature and the residual-fused spatio-temporal dynamic feature to obtain a fused feature is as follows:
[0087] The residual-fused static feature and the residual-fused spatio-temporal dynamic feature are respectively input into a fully connected layer to be mapped to the same dimension, and global average pooling is performed on the features mapped to the same dimension to obtain a fused feature;
[0088] The training method of the bionic two-stream architecture network is as follows:
[0089] It is trained using a task-driven bionic training strategy, and the training strategy includes a functional decoupling training protocol and a cortical topology regularization constraint;
[0090] The method for training the bionic two-stream architecture network using the functional decoupling training protocol is as follows:
[0091] The first stage (15 epochs): Freeze the dorsal pathway module and only optimize the parameters of the ventral pathway module, and set the learning rate to 1e-4;
[0092] Second stage (35 epochs): After the first stage is completed, unfreeze the frozen network parameters, add adversarial samples to the training dataset, and train the bionic two-stream architecture network; the adversarial samples include noise samples with motion noise added in the video frame segments in the training dataset that is opposite to the main motion direction, and samples with static occlusion added.
[0093] The cortical topology regularization constraint is specifically: The cortical topology regularization constraint introduces a cortical similarity loss term into the loss function, and realizes aligning the model with neural signals by constraining the geometric structure consistency between the model parameters and the fMRI responses of the primate visual cortex through contrastive learning.
[0094] The loss function includes a classification loss function and a cortical similarity loss function.
[0095] The classification loss function uses the cross-entropy loss function, and the formula is as follows:
[0096]
[0097] In the formula, Lcls represents the classification loss, n represents the number of samples, i represents the value of the sample index i that varies from 1 to n, yi represents the true label of the i-th sample, p(yi = 1) represents the probability that the model predicts the i-th sample as the positive class, log(p(yi = 1)) represents the logarithmic probability that the i-th sample is predicted as the positive class by the model, (1 - yi) represents the complement of the true label of the i-th sample, and log(1 - p(yi = 1)) represents the logarithmic probability that the i-th sample is predicted as the negative class by the model.
[0098] Cortical Similarity Loss represents an innovative neural network training paradigm. Through a carefully designed contrastive learning framework, this method aims to constrain the consistency between the hidden layer features of the neural network and the functional magnetic resonance imaging (fMRI) responses of the primate visual cortex, and realizes aligning the model with neural signals. The theoretical basis of this method comes from the cross-field of cognitive neuroscience and computer vision, and attempts to bridge the functional differences between artificial intelligence systems and biological vision systems. By introducing biologically inspired constraints, this method can guide the model to form a feature representation space that is more in line with the characteristics of the primate visual system. The core goal of this method is to gradually make the feature representation space of the computational model approach the feature extraction space of the human brain visual system. The human brain visual system has formed a highly efficient and robust visual processing mechanism through millions of years of evolutionary optimization. By simulating the information processing characteristics of this system, we can construct a video processing model that is closer to the human brain visual system in terms of computational efficiency, generalization ability, and response to natural stimuli. This biologically inspired model design concept is expected to bring a paradigm-breaking breakthrough in the field of computer vision.
[0099] During the actual training process, this method adopts a batch-based computing strategy to achieve an effective alignment between the model feature space and the representational structure of the biological visual system. For each training batch, first calculate the Euclidean distances between the feature vectors of all samples at the last layer of the model, so as to obtain the relative distance distribution of these samples in the model feature space. This distribution can be formalized as a distance matrix, denoted as D_model, which intuitively shows how the model distinguishes and organizes different visual inputs;
[0100] Meanwhile, through a pre-designed precise experimental paradigm, the brain fMRI signals of primates when viewing the same visual stimulus samples are recorded. Based on these neuroimaging data, the relative distance distribution of these samples in the brain visual system signal space can also be calculated, expressed as a matrix D_brain. This matrix captures the discrimination pattern of the primate visual cortex for different visual stimuli, and truly reflects the internal representational structure and processing mechanism of the biological visual system;
[0101] When designing the loss function, a more flexible and robust method is adopted. Specifically, it is not required that the two distance matrices D_model(i,j) and D_brain(i,j) are exactly the same in absolute values, but rather that the relative percentile structures they embody are similar. That is to say, for any pair of samples (i,j), if the distance between these two samples is smaller in the brain data (for example, located in a lower percentile), then in the model feature space, this pair of samples should also fall into the same percentile interval as much as possible. Such a matching strategy can effectively reduce the dependence on scale and non-linear transformation, and focuses more on capturing the consistency of the overall structure and order;
[0102] To achieve this goal, the concept of a soft percentile function is introduced. For a certain distance d (for example, a value in D_model(i,j)), first define its "percentile" in the set X of all distances in this batch (for example, X = {D_model(m,n), for all m < n}). The formula of the percentile function is as follows:
[0103]
[0104] where X = {D_model(m,n), for all m < n}) is the set of all distances, d is the distance, σ(z) is the Sigmoid function, γ > 0 is the smoothing parameter. When γ is large, σ(z) approaches the step function, making rank soft approach the true percentile value;
[0105] Based on the above-mentioned soft percentile concept function, calculate the soft percentile values rmodel(i,j) and rbrain(i,j) in the model and brain data for each pair of samples respectively:
[0106] r model (i, j) = Rank soft (D model (i,j); {D model (m,n)})
[0107] r brain (i, j) = Rank soft (D brain (i,j); {D brain (m,n)})
[0108] Based on the soft percentile values rmodel(i,j) and rbrain(i,j), calculate the cortical similarity loss function:
[0109] Finally, take the average of the classification loss function and the cortical similarity loss function to obtain the loss function:
[0110] L = (L cls + L cortex ) / 2;
[0111] Where M represents the number of sample pairs; rmodel(i,j) represents the soft percentile value of the distance Dmodel(i,j) between sample i and sample j in the model feature space, calculated based on the distance set {Dmodel(m,n)} of all sample pairs in the model feature space; rbrain(i,j) represents the soft percentile value of the distance Dbrain(i,j) between sample i and sample j in the brain visual system signal space, calculated based on the distance set {Dbrain(m,n)} of all sample pairs in the brain visual system signal space;
[0112] By minimizing this loss function, the model will gradually adjust its internal representation to make it closer to the organization of the biological visual system, thereby achieving a higher degree of biological inspiration in terms of structure and function;
[0113] The training uses the AdamW optimizer, with preset weight decay (preferably set to 0.01), batch size (preferably set to 32), and uses the cosine annealing learning rate scheduling strategy during training, where the initial learning rate (preferably set to 0.0001) and the minimum learning rate (preferably set to 1e-6) are preset;
[0114] Apply the trained model to a standard dynamic object recognition dataset, and comprehensively evaluate the recognition results using preset metrics; the preset metrics include classification performance metrics, spatio-temporal feature evaluation metrics (temporal consistency metric (Temporal Consistency) and spatial consistency metric (Spatial Consistency), which measure the model's perception ability of object temporal changes and spatial forms respectively), bio-interpretability metric (the correlation coefficient between model features and cortical responses (Cortical Correlation Coefficient, CCC), which quantifies the consistency between the representations of each layer of the network and the activation patterns of the corresponding brain regions. The closer the CCC is to 1, the higher the bio-interpretability of the model), and anti-interference metric (test the model performance under different degrees of noise, blur, and occlusion conditions, draw a robustness curve, and evaluate the model's adaptability in complex environments). Among them, the classification performance metrics include accuracy, precision, recall, and F1-score, which are visually presented through a confusion matrix.
[0115] Although the steps in the above embodiments are described in the above sequential order, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.
[0116] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0117] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.
[0118] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to those process, method, article, or apparatus / device.
[0119] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A brain-inspired modeling method for dynamic object recognition based on the visual dual-pathway, characterized in that, The method includes the following steps: Step S1: Obtain a video frame segment to be recognized for dynamic objects, perform frame extraction on the video frame segment to obtain a single-frame static image and temporally continuous optical flow images; Step S2: Input the single-frame static image and the temporally continuous optical flow images into a trained bionic two-stream architecture network to obtain the recognition result of the dynamic object; The bionic two-stream architecture network includes a dorsal pathway module, a ventral pathway module, a two-branch interaction module, and a cross-modal neural fusion module; The dorsal pathway module is configured to use a hierarchically optimized Video SwinTransformer for the temporally continuous optical flow images to extract spatio-temporal motion features from local to global, and generate a spatio-temporal dynamic representation similar to the motion energy map of the MT area; The ventral pathway module is configured to use a ResNet convolutional neural network for the single-frame static image to simulate the hierarchical selective representation characteristics of the IT area for static shapes and extract static features; The two-branch interaction module is configured to use both the spatio-temporal motion features and the static features as original features; perform dimensionality transformation on each original feature, and after dimensionality transformation, send it to the Blcok block in the opposite pathway for processing to obtain interactively generated features; fuse each interactively generated feature with the corresponding original feature to obtain residue-fused features, and the residue-fused features include residue-fused static features and residue-fused spatio-temporal dynamic features; The cross-modal neural fusion module is configured to fuse the residue-fused static features and the residue-fused spatio-temporal dynamic features respectively to obtain fused features; process the fused features through a multi-layer perceptron to obtain the recognition result of the dynamic object.
2. The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway according to claim 1, characterized in that The method for performing frame extraction on the video frame segment to obtain a single-frame static image and temporally continuous optical flow images is as follows: Randomly extract one frame from the video frame segment as the single-frame static image; Perform frame extraction on the video frame segment through an optical flow estimation algorithm to obtain temporally continuous optical flow images; the optical flow estimation algorithm includes the Lucas-Kanade optical flow estimation algorithm.
3. The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway according to claim 1, characterized in that, The hierarchically optimized Video Swin Transformer is a Transformer encoder including multiple stages, and each stage includes at least two Swin-Transformer blocks.
4. The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway according to claim 1, characterized in that The method for obtaining the residue-fused features is as follows: Perform dimensionality transformation on the spatio-temporal motion features and the static features through a multi-layer perceptron to obtain the spatio-temporal motion features X' after dimensionality transformation swin and static features X' resnet , and the expression is: X′ swin = MLP1(X swin ) X′ resnet = MLP2(X resnet ) Among them, X swin and X resnet respectively represent the extracted spatio-temporal motion features and static features; X′ swin is the feature after dimensionality transformation of the spatio-temporal motion features; X′ resnet is the feature after dimensionality transformation of the static features. X′ swin has the same dimension as X resnet and X′ resnet has the same dimension as X swin ; Send the spatio-temporal motion feature X′ after dimensional transformation swin and the static feature X′ resnet into the Block of the contralateral path for processing to obtain the interactively generated feature: Among them, represents the spatio-temporal motion features generated by interaction, represents the static features generated by interaction; Fuse each interactively generated feature with the corresponding original feature to obtain residue-fused features: Among them, Y swin represents the residual fusion spatio-temporal motion feature, and Y resnet represents the residual fusion static feature.
5. The brain-inspired modeling method for dynamic object recognition based on visual dual pathways according to claim 1, wherein The method for fusing the residue-fused static features and the residue-fused spatio-temporal dynamic features to obtain fused features is as follows: Respectively input the residue-fused static features and the residue-fused spatio-temporal dynamic features into a fully connected layer to map them to the same dimension, and perform overall global average pooling on the features mapped to the same dimension to obtain fused features.
6. The brain-inspired modeling method for dynamic object recognition based on visual dual pathways according to claim 1, wherein The training method of the bionic two-stream architecture network is as follows: Perform training using a task-driven bionic training strategy, and the training strategy includes a functional decoupling training protocol and a cortical topology regularization constraint.
7. The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway according to claim 6, wherein The method for training the bionic two-stream architecture network using the functional decoupling training protocol is as follows: The first stage: Freeze the dorsal pathway module and only optimize the parameters of the ventral pathway module; The second stage: After the first stage is completed, unfreeze the frozen network parameters, add adversarial samples to the training dataset, and train the bionic two-stream architecture network; the adversarial samples include noise samples with motion noise added in the video frame segments of the training dataset in the opposite direction of the main motion direction and samples with static occlusion added.
8. The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway according to claim 6, wherein The cortical topology regularization constraint is specifically: Introduce a cortical similarity loss term into the loss function, and constrain the geometric structure consistency between the model parameters and the fMRI responses of the primate visual cortex through contrastive learning; The loss function includes a classification loss function and a cortical similarity loss function.
9. The brain-inspired modeling method for dynamic object recognition based on the visual dual pathway according to claim 8, wherein The loss function is: L = (L cls + L cortex ) / 2; where, L cls is the classification loss function, and L cortex is the cortical similarity loss function.
10. The brain-inspired modeling method for dynamic object recognition based on visual dual pathways according to claim 8, wherein The cortical similarity loss function: r model (i, j) = Rank soft (D model (i, j); {D model (m, n)}); r brain (i, j) = Rank soft (D brain (i, j); {D brain (m, n)}); Where M represents the number of sample pairs; rmodel(i,j) represents the soft percentile value of the distance Dmodel(i,j) between samples i and j in the model feature space, calculated based on the set of distances {Dmodel(m,n)} of all sample pairs in the model feature space; rbrain(i,j) represents the soft percentile value of the distance Dbrain(i,j) between samples i and j in the brain visual system signal space, calculated based on the set of distances {Dbrain(m,n)} of all sample pairs in the brain visual system signal space.
Citation Information
Cited By
Brain-like calculation driven visual information instant analysis method and system
CN120564006A
Brain-inspired computing driven visual information instant analysis method and system
CN120564006B