Brain-like intelligent imaging target detection method

By building a brain-like intelligent multi-module model and simulating the collaborative cooperation of the brain vision system, the problem of missed detection of traditional object detection methods in complex scenarios and imaging equipment motion is solved, and fast and accurate object detection is achieved, reducing computing costs and hardware requirements.

CN120259743APending Publication Date: 2025-07-04HOHAI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317917.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional object detection methods are difficult to adapt to in the body movement of imaging equipment and complex scenarios, resulting in mis-detection or missed detection. The deep learning methods are computationally costly and have high hardware requirements, making it difficult to deal with interference factors such as lighting changes and occlusion.

Method used

Build a brain-like intelligent multi-module model, including visual information reception, motion perception, color perception, feature binding and feature identification modules, simulate the collaborative cooperation of the brain's visual system, extract scene information through color perception and motion perception dual neural channels, suppress background noise and highlight the target.

Benefits of technology

It realizes fast and accurate object detection in dynamic scenarios, reduces computing costs and hardware requirements, has significant real-time and accuracy, and can effectively suppress background noise and enhance object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259743A_ABST
    Figure CN120259743A_ABST
Patent Text Reader

Abstract

The invention discloses a brain-like intelligent imaging target detection method, and the method achieves the accurate detection of a target under the motion condition of an imaging device body in a complex scene through building a color vision and motion perception dual-neural channel model, and carrying out the modeling of a neural activity process in which different brain regions and a dopamine system cooperate in a visual target detection process. The brain-like intelligent model designed by the invention can extract the motion and color information of the target, and perceives the scene target through bionic fusion visual information. Compared with the prior art, under the condition that the detection accuracy is guaranteed, the problems of over-fitting and under-fitting which are often caused by deep learning are effectively solved, the advantage of lightweight is achieved, and the training time and the hardware cost are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a brain-inspired intelligent imaging target detection method by simulating the process of visual nerve activity, and belongs to the technical field of imaging data target detection. Background Art

[0002] In the current fields of computer vision and artificial intelligence, target detection is a key technical challenge. It requires the system to be able to quickly and accurately identify moving objects or regions in complex dynamic scenes. Traditional target detection methods often rely on fixed feature extraction and classifier design, and it is difficult to adapt to the challenges brought by complex and changeable scenes and the movement of the imaging device itself. For example, in the case of the movement of the imaging device itself, the background information may change, resulting in false detections or missed detections by traditional target detection methods. Although the subsequent emerging methods based on machine learning or deep learning have improved the accuracy of target detection to a certain extent, they are often accompanied by limitations such as high computational costs and high hardware requirements. In addition, factors such as illumination changes, occlusion, and noise in complex scenes will also interfere with target detection, further increasing the difficulty of target detection.

[0003] In order to overcome the above technical problems, researchers have begun to explore bionic vision-based intelligent target detection methods. Brain-inspired intelligence is an artificial intelligence technology that simulates the working principle of the biological brain. By simulating the connection and information transmission process between brain neurons, more intelligent and adaptive target detection can be achieved. Summary of the Invention

[0004] Object of the Invention: The present invention proposes a brain-inspired intelligent imaging target detection method. For the target detection of imaging information, this method constructs a dual neural channel of color vision and motion perception, comprehensively simulates the cooperation between different regions of the brain and the dopamine system, and constructs a system-level neural network model. This model can make full use of the advantages of the visual system to achieve fast and accurate target detection in dynamic scenes. Moreover, the model constructed by the present invention has high flexibility and adaptability and can be widely applied to fields such as intelligent monitoring, autonomous driving, and human-computer interaction.

[0005] Technical Solution: A brain-inspired intelligent imaging target detection method includes: constructing a brain-inspired intelligent multi-module model, and the brain-inspired intelligent multi-module model includes a visual information receiving module, a motion perception module, a color perception module, a feature binding module, a mushroom body module, and a feature identification module.

[0006] The visual information receiving module is used to simulate the retinal function of biological vision, input the images obtained from imaging, and capture the time-varying information of the scene; the motion perception module extracts motion information, and the color perception module extracts color information; the feature binding module and the mushroom body module fuse multiple visual features; the feature identification module outputs the target detection result of the final image information, and calibrates the target area and the noise area.

[0007] The visual information receiving module captures the time-varying information of the scene under the condition of the body's movement by simulating the receptive field characteristics of biological vision and filtering out interference factors such as dynamic noise:

[0008] P t =|f t H -f t |

[0009] where P t is the frame difference, f t is the frame at time t, f t H is the transformation result of the frame f t-1 at t-1, and the transformation process is as follows:

[0010]

[0011] where H is the transformation matrix, representing the matching correspondence between the frame at time t and the frame at time t-1.

[0012] In the feature extraction link, a dual neural channel model of color vision and motion perception is constructed: through these two neural channels, the color and motion information in the scene are extracted respectively. The motion perception module extracts motion features, and the color perception module extracts color features.

[0013] The motion perception module takes the time-varying information P t of the scene as the input. This module is mainly composed of three neural layers: the lamina, the medulla, and the lobula. In the lamina layer of the motion perception module, the accumulation of inhibitory signals comes from the responses of lateral neurons:

[0014]

[0015] where, I t (x,y) is the inhibition corresponding to the neuron located at the point (x, y), and ω is an inhibition template matrix of size q×s composed of elements 0 and 1. For example, when the left half of the template matrix is filled with 1 and the right half is filled with 0, the left direction is emphasized. The direction selectivity is achieved through the morphological change of the ω matrix in the motion perception module:

[0016]

[0017] The above four morphologies ωu , ω d , ω l , ω r correspond to the selection of the upper, lower, left, and right direction sensitivities respectively.

[0018] In the motion perception module, the medulla layer simulates the direction selectivity perception of T cells, extracts motion cues along the discriminative direction, and the motion intensities in the four directions of left, right, up, and down can be expressed as:

[0019]

[0020] Among them, P t (x, y) is the time-varying information of the scene at the coordinate (x, y), W is the Gaussian kernel of each outer neuron group, and and are the inhibitory signals in the four directions of left, right, up, and down, * is the half-wave rectification calculation, which is a monotonically increasing function of the scene contrast. The half-wave rectification calculation is represented by the sigmoid function as follows:

[0021]

[0022] where η is the regulator that constrains the zero-crossing gradient.

[0023] The lobule layer in the motion perception module simulates the integration function of the tangential cells of the lobule plate, and integrates the motion information OUT in all directions t p :

[0024]

[0025] The color perception module takes the color information of the spatial transformation result f t H of the visual information receiving module as the input, and this module is mainly composed of two neural layers, the medulla and the lobule.

[0026] In the medulla layer, the visual information received within the unit time t is decomposed into three independent processing channels: red (R), green (G), and blue (B), as follows:

[0027] P t R = (f t H - βP t G - γP t B ) / α, P t G = (f tH -αP t R -γP t B ) / β, P t B = (f t H -αP t R -βP t G ) / γ, where P t R 、P t G 、P t B correspond to the information components decomposed into the red, green, and blue channels respectively, while α, β, and γ represent the sensitivity of the brain to different color stimuli.

[0028] Subsequently, the channel information is preprocessed by Gaussian filtering and divided into ON and OFF channels. One input forms the ON pathway and the other forms the OFF pathway. Taking the red channel as an example, the separation process is as follows:

[0029] ON(x, y, t) = (P t R + |P t R |) / 2

[0030] OFF(x, y, t) = |(P t R - |P t R |)| / 2

[0031] The input on the ON or OFF pathway is divided into excitatory (E) and inhibitory (I) signals. Specifically, in the ON pathway, the excitatory signal is directly transmitted by the ON pathway input ON(x, y, t), denoted as E on , as follows:

[0032] E on (x, y, t) = ON(x, y, t)

[0033] The inhibitory signal I in the ON pathway comes from the delayed excitation effect around the convolution process, denoted as I on . The calculation process is as follows:

[0034]

[0035] where r represents the radius of the convolution kernel (the size of the inhibitory area), usually set to 1. W i is the convolution kernel, as follows:

[0036]

[0037] In the convolutional kernel Wi, the four nearest neighbor units exhibit relatively large weights and short latencies. D on represents a delayed signal with an excitatory signal delay of tens to hundreds of milliseconds, D on has a first-order low-pass filtering cooperative relationship with the excitatory signal ON, as follows:

[0038]

[0039] where τ s is a dynamic time parameter that can vary in the range of tens to hundreds of milliseconds.

[0040] Compared with the delay information of the ON channel, the excitatory current in the OFF channel is delayed relative to the inhibitory current. That is, the excitatory signal E in the OFF channel off is:

[0041]

[0042] In the OFF pathway, the inhibitory signal is directly transmitted by the OFF pathway input OFF(x, y, t), denoted as I off :

[0043] I off (x,y,t) = OFF(x,y,t)

[0044] Then, the inhibitory signals on each pathway and the biases w1 and w2 linearly integrate to respectively inhibit the excitatory signals E on each pathway:

[0045] S on (x,y,t) = E on (x,y,t) - w1I on (x,y,t)

[0046] S off (x,y,t) = E off (x,y,t) - w2I off (x,y,t)

[0047] where w1 and w2 are the bias coefficients for excitation and inhibition.

[0048] Finally, the shunt signals of the ON and OFF paths interact in a superlinear (multiplicative and linear) manner to generate the ON / OFF path processing result of the red channel, that is, the integrated signal S R :

[0049] S R(x, y, t) = θ1S on (x, y, t) + θ2S off (x, y, t) + θ3S on (x, y, t)S off (x, y, t)

[0050] Where θ1, θ2, and θ3 represent the interaction control variables between the ON and OFF channels.

[0051] The integrated signal can reflect the edge information of the target to a certain extent. Subsequently, the lobule layer integrates the information S of different color channels R 、S G 、S B according to the sensitivity of the brain to different colors (i.e., α, β, γ), and then jointly considers the constraints using the intra-frame and inter-frame correspondence relationships, so as to simulate the function of the tangential cells of the lobule plate, filter out the background clutter, and leave the key pixel points OUT i C ,and extract the color information of the target. The specific process is as follows:

[0052] S(x, y, t) = αS R (x, y, t) + βS G (x, y, t) + γS B (x, y, t)

[0053] D(x, y, t) = |S(x, y, t) - S(x, y, t - 1)|

[0054]

[0055] Where T is the automatically calculated threshold, simulating the activation threshold of the tangential cells of the lobule layer.

[0056] The feature binding module, simulating the dopamine system, receives the outputs from the motion perception module and the color perception module and determines whether to activate the mushroom body module. The specific process is as follows:

[0057]

[0058] Where r is the number of pixel points in the local area. OUT i P and OUT i C represent the pixel points of the outputs of the motion perception module and the color perception module. and It represents the mean value of the pixel points in the local area of the motion perception module and the color perception module. Q represents the determination threshold of dopamine concentration release. In the present invention, it is tentatively set that when Q < 0.5, it indicates that the intensity of the current scene color information is low, dopamine release is inhibited, and the mushroom body is not activated; when Q > 0.5, it indicates that the intensity of the current scene color information is high, dopamine is released, and the mushroom body is activated.

[0059] The constructed mushroom body module: The mushroom body only participates in the work when it is in the activated state. When the mushroom body is activated, the output of the i-th pixel point of the mushroom body module is as follows:

[0060]

[0061] where w m and w n are the synaptic weights from the i-th unit of the motion perception module and the color perception module to the i-th unit of the feature binding module.

[0062] The feature identification module simulates the function of the central complex, selects the outputs of the mushroom body module and the motion perception module, and fully displays the valid pixel points OUT i in the target area, and generates the final target detection result. The selection of valid pixel points is based on the activation state of the mushroom body. When the mushroom body is not activated, the color clues are discarded, and only the motion clues are relied on to select the valid pixel points. When the mushroom body is activated, it represents that the color clues are strong, and the pixel points after fusing the motion and color clues are selected. Thus, the target area is displayed, and the clutter in the non-target area is completely filtered out. The specific process is as follows:

[0063]

[0064] where γ is the number of neuron groups outside the target area recognized by the mushroom body or the position module; i is the neuron group corresponding to the target area; w3 and w4 indicate the activity state of the mushroom body. When the mushroom body is activated, w3 = 1 and w4 = 0. When the mushroom body is not activated, w3 = 0 and w4 = 1.

[0065] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the steps of the brain-like intelligent imaging target detection method as described above.

[0066] A computer-readable storage medium stores a computer program for executing the brain-like intelligent imaging target detection method as described above.

[0067] Beneficial effects: Compared with the prior art, the present invention imitates the principle of the brain's visual system based on the visual spatio-temporal regularity, and uses the continuous correlation mechanism of motion and color perception and the visual convergence mechanism to suppress background noise and highlight the target state; it gets rid of training and background modeling, can quickly and effectively detect moving targets in complex dynamic scenes, realizes effective perception of the scene, and has remarkable real-time performance, effectiveness and accuracy. Brief Description of the Drawings

[0068] Figure 1 It is a schematic diagram of the flow of each module of the brain-inspired intelligent imaging target detection method of the present invention and a schematic diagram of neurons in the brain visual area;

[0069] Figure 2 It is a schematic diagram of the neural layer in the motion perception module of the brain-inspired intelligent imaging target detection method of the present invention;

[0070] Figure 3 It is a schematic diagram of the neural layer in the color perception module of the brain-inspired intelligent imaging target detection method of the present invention;

[0071] Figure 4 It is an example illustration of the brain-inspired intelligent imaging target detection method of the present invention. Detailed Embodiments

[0072] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention fall within the scope defined by the appended claims of this application.

[0073] The following describes in detail the application principle of the present invention with reference to the accompanying drawings.

[0074] Embodiment For example Figure 4 As shown, the present invention designs a brain-inspired intelligent imaging target detection method. The proposal of the present invention mainly relies on two biological discoveries of visual perception neural networks and two visual perception mechanisms: ① Direction-sensitive neurons can detect the motion cues of the target; ② Light-sensitive and color-sensitive neurons detect color cues; ③ The information integration and convergence mechanism at the end of the biological visual nerve is used to fuse the information in two aspects of motion and color; ④ Use the continuous correlation mechanism of visual perception to be independent of background noise and enhance the target. According to discovery ①, in the video scene captured by a moving camera, the motion cues of the target can be well detected (such as Figure 4 (b)). According to discovery ②, the color cues of the target can be effectively detected (such as Figure 4(c)). According to mechanism ③, the feature binding module and the mushroom body module cooperate to select, judge, and fuse multiple cues to strengthen the target information. According to mechanism ④, this mechanism can suppress the excitation of irrelevant background noise after the fusion of the two models to obtain the final target area (such as Figure 4 (d)).

[0075] First, according to the brain-inspired visual information processing process, a multi-module network is constructed, as shown in Figure 1 . For the visual information receiving module: There are many regularly arranged heterogeneous photoreceptors on the surface of the retina of organisms. These heterogeneous photoreceptors, as optical sensors, have extremely high sensitivity to the spectrum and can effectively respond to intensity changes. Functionally, this module simulates the retina to receive the original image data from the dynamic scene, including information such as brightness and shape:

[0076] P t =|f t H -f t |

[0077] where P t is the frame difference, f t is the frame at time t, and f t H is transformed from the f t-1 frame at time t-1.

[0078] For the motion perception module: As shown in Figure 2 , it is composed of three nerve layers: the lamina, the medulla, and the lobula. Taking the scene time-varying information P t as the input, the lamina layer receives the time motion information and has strong direction selectivity sensitivity to one of the four basic directions, and can filter out noise / disturbances.

[0079]

[0080] where I t (x,y) is the inhibition corresponding to the neuron located at the point (x, y), and ω is an inhibition template matrix of size q×s composed of elements 0 and 1. The direction selectivity in the motion perception module is achieved through the morphological changes of the ω matrix:

[0081]

[0082] The above four morphologies respectively correspond to the four direction sensitivity selections of up, down, left, and right.

[0083] After extracting the inhibitions in the four directions, calculate the motion information in each direction:

[0084]

[0085] Among them, W is the Gaussian kernel of each outer neuron group, and and are inhibitory signals in the four directions of left, right, up, and down. * is the half-wave rectification calculation:

[0086]

[0087] where η is the regulator that constrains the zero-crossing gradient.

[0088] The lobular layer integrates the motion signals from all directions in the medullary layer to obtain the synthetic motion perception OUT t p , as follows:

[0089]

[0090] The finally obtained result is as Figure 4 (b) shown, and the motion information of the moving target 'fly' is effectively extracted.

[0091] For the color perception module: as Figure 3 shown, it is composed of two neural layers, the medulla and the lobule. Taking the color information of the spatial transformation result f t H of the visual information receiving module as the input. According to the different sensitivities of different channels to different colors, the information is extracted by the RGB three-color channels respectively.

[0092] P t R = (f t H - βP t G - γP t B ) / α, P t G = (f t H - αP t R - γP t B ) / β, P t B = (f t H - αP t R - βP t G ) / γ, where P t R , P t G , P t BCorrespond to the information components decomposed into the red, green, and blue channels respectively, while α, β, and γ characterize the sensitivity of the brain to different color stimuli.

[0093] Subsequently, the information is preprocessed by Gaussian filtering and divided into parallel channels. One input forms the ON pathway, and the other forms the OFF pathway. Taking the red channel as an example, the separation process is as follows:

[0094] ON(x,y,t) = (P t R + |P t R |) / 2

[0095] OFF(x,y,t) = |(P t R - |P t R |)| / 2

[0096] The inputs on the ON or OFF pathways are divided into excitatory (E) and inhibitory (I) signals. Specifically, in the ON pathway, the excitatory signal E is directly transmitted by the ON pathway input ON(x, y, t), denoted as E on . The inhibitory signal I in the ON pathway comes from the delayed excitation effect around the convolution process, denoted as I on . The calculation process is as follows:

[0097] E on (x,y,t) = ON(x,y,t)

[0098]

[0099] where r represents the radius of the convolution kernel (the size of the inhibitory area), usually set to 1. The convolution kernel Wi is:

[0100]

[0101] In the convolution kernel Wi, the four nearest neighbor units exhibit relatively large weights and short delays. D on represents the delayed signal of the excitatory signal delayed by dozens to hundreds of milliseconds. D on has a first-order low-pass filtering cooperative relationship with the excitatory signal ON, as follows:

[0102]

[0103] where τ s is a dynamic time parameter that can vary in the range of dozens to hundreds of milliseconds.

[0104] Similarly, in the OFF pathway, the inhibitory signal I is directly transmitted by the OFF pathway OFF(x, y, t), denoted as I off . The excitation results from the delayed inhibitory effect around the convolution, denoted as E off . The calculation is as follows:

[0105] I off (x,y,t) = OFF(x,y,t)

[0106]

[0107] Then, the inhibitory signals I on each pathway are linearly integrated with the biases w1 and w2 to respectively inhibit the excitatory signals E on each pathway:

[0108] S on (x,y,t) = E on (x,y,t) - w1I on (x,y,t)

[0109] S off (x,y,t) = E off (x,y,t) - w2I off (x,y,t)

[0110] Finally, the shunt signals of the ON and OFF paths interact with each other in a superlinear (multiplicative and linear) manner to generate the processing result of the ON / OFF unit of the red channel, that is, the integrated signal S R :

[0111] S R (x,y,t) = θ1S on (x,y,t) + θ2S off (x,y,t) + θ3S on (x,y,t)S off (x,y,t)

[0112] Among them, the coefficient combination of the terms {θ1, θ2, θ3} represents the interaction between the ON and OFF pathways.

[0113] The integrated signal can reflect the edge information of the target to a certain extent. Subsequently, the lobular layer integrates the information S R 、S G 、S B of different color channels according to the sensitivities of different colors in the brain (i.e., α, β, γ), and then jointly considers the constraints using the correspondence within and between frames, so as to simulate the function of the tangential cells of the lobular plate, filter out the background clutter, and leave the key pixel points OUT i C , and extract the color information of the target. The specific process is as follows: S(x,y,t) = αSR (x, y, t) + βS G (x, y, t) + γS B (x, y, t)

[0114] D(x, y, t) = |S(x, y, t) - S(x, y, t - 1)|

[0115]

[0116] where T is an automatically calculated threshold, simulating the condition for the activation of tangential cells in the lobule plate.

[0117] The finally obtained result is as Figure 4 (c) shows that the color information OUT of the moving target 'fly' at time t t C is effectively extracted.

[0118] For the feature binding module: When faced with cue conflicts, this module takes the outputs of the motion perception module and the color perception module as inputs and determines whether to activate the mushroom body module, so as to help the saliency system select the most appropriate features.

[0119]

[0120] Q represents the dopamine release determination threshold. In the present invention, it is tentatively assumed that when Q < 0.5, it indicates that the intensity of the color information in the current environment is low, dopamine release is inhibited, and the mushroom body is not activated. The feature recognition module cannot access color features. When Q > 0.5, the color information is regarded as high intensity. Dopamine is released and the mushroom body is activated. r is the number of pixel points in the region. OUT i P and OUT i C represent the pixel points of the outputs of the motion perception module and the color perception module. and represent the means of the pixel points in the local regions of the motion perception module and the color perception module.

[0121] For the mushroom body module: This module receives the output of the feature binding module and only participates in the work when it is in the activated state. Moreover, only when the dopamine concentration increases, the mushroom body is activated, performs lateral inhibition on multiple features, and generates a fused feature map. When the dopamine concentration does not change, the mushroom body does not participate in neural activities. When the mushroom body is activated, the output of the i-th pixel point of the mushroom body module is as follows:

[0122]

[0123] where w m and w nIt is the synaptic weight from the i-th unit of the motion perception module and the color perception module to the i-th unit of the feature binding module.

[0124] For the feature identification module: This module integrates all visual features of the previous modules, performs neural refinement on the feature map, and fully displays the valid pixel points OUT within the target area. i Then, the final target detection result is generated. The selection of valid pixel points is based on the activation state of the mushroom body. When the mushroom body is not activated, color cues are discarded, and only motion cues are relied on to select valid pixel points. When the mushroom body is activated, it means that the color cues are strong, and the pixel points after fusing motion and color cues are selected. The specific process is as follows:

[0125]

[0126] Among them, γ is the number of neuron groups outside the target area recognized by the mushroom body or the position module; i is the neuron group corresponding to the target area; w3 and w4 indicate the activity state of the mushroom body. When the mushroom body is activated, w3 = 1 and w4 = 0. When the mushroom body is not activated, w3 = 0 and w4 = 1.

[0127] So far, the valid pixel points regarding the target are extracted, the fusion and enhancement of multiple cues are completed, and the target detection of the video captured by the moving camera is completed, as shown in Figure 4 (d).

[0128] Obviously, those skilled in the art should understand that each step of the above-mentioned brain-inspired intelligent imaging target detection method of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be made into individual integrated circuit modules respectively, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. A brain-inspired intelligent imaging target detection method, characterized in that Comprising: Construct a brain-inspired intelligent multi-module model, which includes a visual information reception module, a motion perception module, a color perception module, a feature binding module, a mushroom body module, and a feature identification module. The visual information reception module is used to simulate the retinal function of biological vision, input the image obtained by imaging, and capture the time-varying information of the scene; the motion perception module extracts motion information, and the color perception module extracts color information; the feature binding module and the mushroom body module fuse multi-visual features; The feature identification module outputs the target detection result of the final image information, and calibrates the target area and the noise area.

2. The imaging target detection method for brain-like intelligence according to claim 1, wherein The visual information reception module captures the time-varying information of the scene under the condition of the body movement by simulating the receptive field characteristics of biological vision and filtering out interference factors such as dynamic noise: P t = |f t H - f t | Where P t is the inter-frame difference, and f t is the frame at time t, and f t H is the frame f at t - 1 t-1 The transformation result is as follows: Where H is the transformation matrix, representing the matching correspondence between the frame at time t and the frame at time t-1.

3. The imaging target detection method for brain-like intelligence according to claim 1, characterized in that The motion perception module uses the scene time-varying information P t as input. This module is mainly composed of three neural layers: the lamina, the medulla, and the lobule. In the lamina layer of the motion perception module, inhibitory signal accumulation comes from the responses of lateral neurons: Among them, I t (x, y) is the inhibition corresponding to the neuron located at the point (x, y), and ω is an inhibition template matrix of size q×s composed of elements 0 and 1; in the motion perception module, direction selectivity is achieved through the morphological change of the ω matrix: The above four forms ω u 、ω d 、ω l 、ω r respectively correspond to the selection of four directional sensitivities of up, down, left, and right; In the motion perception module, the medulla layer simulates the direction selectivity perception of T cells and extracts motion cues along the discriminative direction. The motion intensity along the four directions of left, right, up, and down is expressed as: where P t (x, y) is the time-varying information of the scene at coordinates (x, y), W is the Gaussian kernel of each outer neuron group, and and are the inhibitory signals in the four directions of left, right, up, and down, * is the half-wave rectification calculation, which is a monotonically increasing function of the scene contrast. The half-wave rectification calculation is represented by the sigmoid function as follows: Where η is a regulator that constrains the zero-crossing gradient; In the motion perception module, the integration function of the tangential cells of the lobular lamella in the lobular layer simulates the integration of all motion information in all directions OUT t p :

4. The imaging target detection method for brain-like intelligence according to claim 1, wherein, The color perception module uses the color information of the spatial transformation result f of the visual information receiving module as input, and this module is mainly composed of two nerve layers, namely the medulla and the lobule; t H ​ In the medulla layer, the visual information received within unit time t is decomposed into three independent processing channels: red (R), green (G), and blue (B), as follows: P t R =(f t H -βP t G -γP t B ) / α, P t G =(f t H -αP t R -γP t B ) / β, P t B =(f t H -αP t R -βP t G ) / γ, where P t R , P t G , P t B correspond to the information components decomposed into the red, green, and blue channels respectively, while α, β, and γ represent the sensitivity of the brain to different color stimuli; Subsequently, the channel information is preprocessed by Gaussian filtering and divided into ON and OFF channels; one forms the input of the ON pathway, and the other forms the input of the OFF pathway; for the red channel, the separation process is as follows: ON(x,y,t) = (P t R + |P t R |) / 2 OFF(x,y,t) = |(P t R - |P t R |)| / 2 Inputs on the ON or OFF pathways are classified into excitatory (E) and inhibitory (I) signals; specifically, in the ON pathway, the excitatory signal is directly transmitted by the ON pathway input ON(x, y, t), denoted as E on , as follows: E on (x, y, t) = ON(x, y, t) The inhibitory signal I in the ON path comes from the delayed excitation effect around the convolution process, denoted as I on ; The calculation process is as follows: where r represents the radius of the convolution kernel, and W i is the convolution kernel, as shown below: In the convolutional kernel Wi, four nearest neighboring units exhibit relatively large weights and short delays; D on represents a delayed signal with an excitatory signal delay of tens to hundreds of milliseconds, D on has a first-order low-pass filtered cooperative relationship with the excitatory signal ON, as follows: where τ s is a dynamic time parameter; Compared with the delay information of the ON channel, the excitatory current of the OFF channel is delayed relative to the inhibitory current; that is, the excitatory signal E in the OFF channel off is as follows: In the OFF pathway, the inhibitory signal is directly transmitted by the OFF pathway input OFF(x, y, t), denoted as I off : I off (x, y, t) = OFF(x, y, t) Then, the inhibitory signals on each pathway and the biases w1 and w2 respectively inhibit the excitatory signals E on each pathway through linear integration: S on (x,y,t) = E on (x,y,t) - w1I on (x,y,t) S off (x,y,t) = E off (x,y,t) - w2I off (x,y,t) Where w1 and w2 are the bias coefficients of excitation and inhibition; Finally, the shunt signals of the ON and OFF paths interact with each other in a superlinear manner to generate the ON / OFF path processing result of the red channel, that is, the integrated signal S R : S R (x,y,t) = θ1S on (x,y,t) + θ2S off (x,y,t) + θ3S on (x,y,t)S off (x,y,t) Where θ1, θ2, θ3 represent the interaction control variables between the ON and OFF pathways; Subsequently, the lobular layer integrates the information S of different color channels R , S G , S B according to the sensitivities of different colors in the brain, and then jointly considers the constraints by using the correspondence relationships within and between frames, so as to simulate the function of the tangential cells of the lobular plate, filter out the background clutter, and leave the key pixel points OUT i C , and extract the color information of the target; the specific process is as follows: S(x,y,t) = αS R (x,y,t) + βS G (x,y,t) + γS B (x,y,t) D(x,y,t) = |S(x,y,t) - S(x,y,t-1)| Where T is an automatically calculated threshold, simulating the activation threshold of the tangential cells in the lobula layer.

5. The imaging target detection method for brain-like intelligence according to claim 1, characterized in that, The feature binding module simulates the dopamine system, receives the outputs from the motion perception module and the color perception module, and determines whether to activate the mushroom body module; the specific process is as follows: Among them, r is the number of pixel points in the local area; OUT i P and OUT i C represent the pixel points of the outputs of the motion perception module and the color perception module; and represent the means of the pixel points in the local areas of the motion perception module and the color perception module; Q represents the dopamine concentration release determination threshold.

6. The imaging target detection method for brain-like intelligence according to claim 1, characterized in that, The construction of the mushroom body module: The mushroom body will only participate in the work when it is in the activated state. When the mushroom body is activated, the output of the i-th pixel point of the mushroom body module is as follows: where w m and w n are the synaptic weights from the i-th unit of the motion perception module and the color perception module to the i-th unit of the feature binding module.

7. The imaging target detection method for brain-like intelligence according to claim 1, characterized in that, The feature recognition module simulates the function of the central complex, selects the outputs of the mushroom body module and the motion perception module, and completely displays the effective pixel points OUT in the target area i , and generates the final target detection result; selects the effective pixel points according to the activation state of the mushroom body; when the mushroom body is not activated, discards the color clues and only relies on the motion clues to select the effective pixel points; when the mushroom body is activated, it means that the color clues are strong, and the pixel points after fusing the motion and color clues are selected; thus, the target area is displayed, and the clutter in the non-target area is completely filtered out; the specific process is as follows: Where γ is the number of neuron groups outside the target area recognized by the mushroom body or the position module; i is the neuron group corresponding to the target area; w3 and w4 indicate the activity state of the mushroom body, w3 = 1 and w4 = 0 when the mushroom body is activated, and w3 = 0 and w4 = 1 when the mushroom body is not activated.

8. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the steps of the imaging target detection method of brain-inspired intelligence as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the imaging target detection method of brain-inspired intelligence as described in any one of claims 1-7.