A target detection method, device and equipment based on bionic vision and a medium

By simulating the biological sensing mechanism of mantis shrimp, a biomimetic detection model for targets was constructed, which solved the problems of high light saturation and noise interference in complex environments by traditional light intensity detection. This improved the accuracy and anti-interference ability of target detection, and adopted multimodal information fusion and adaptive processing technology.

CN120599238BActive Publication Date: 2025-11-11CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114213.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-11
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Traditional light intensity detection faces challenges such as high light saturation, contrast decay, and noise interference in complex environments, making it difficult to effectively identify target materials and resulting in insufficient target detection accuracy and anti-interference capability.

Method used

Simulating the biological perception mechanism of mantis shrimp, a target biomimetic detection model based on neural networks is constructed. By acquiring the target's spectral and polarization information, multimodal fusion features are generated. The target feature tensor map is generated using recursive fractal partitioning and cross-scale fusion attention mechanism. The feature representation and decoding are combined with residual skip connections and manifold learning methods to improve information extraction efficiency and anti-interference ability.

Benefits of technology

It improves the accuracy and anti-interference ability of target detection in complex environments, realizes efficient fusion and adaptive processing of multimodal information, and enhances the information extraction efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599238B_ABST
    Figure CN120599238B_ABST
Patent Text Reader

Abstract

This application discloses a target detection method, apparatus, device, and medium based on biomimetic vision, relating to the field of artificial intelligence technology. The method includes: acquiring collected target spectral information and target polarization information maps; constructing a spectral feature pyramid corresponding to the target spectral information and determining the target polarization features corresponding to the target polarization information maps; generating target multimodal fusion features based on the spectral feature pyramid and target polarization features; generating a target spatial interest map based on the target multimodal fusion features and generating a target feature tensor map; generating corresponding target enhancement features based on the target spatial interest map and target feature tensor map; constructing a corresponding target feature representation space based on manifold learning methods, Lie group operations, target enhancement features, and the target spatial interest map; and decoding the target feature representation space to obtain the corresponding target detection results. In this way, this application can improve the accuracy and anti-interference capability of target detection in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a target detection method, apparatus, device, and medium based on biomimetic vision. Background Technology

[0002] In recent years, airborne ground and sea detection technologies have made significant progress in various fields such as ship detection, environmental monitoring, and disaster early warning, providing important technical support for target identification under complex sea and atmospheric conditions. However, the traditional mode of relying solely on light intensity detection still faces several significant bottlenecks in complex environments: on the one hand, specular reflection from the ocean and strong solar flares often cause high light saturation of the detector, obscuring key details such as ship shadows and wakes, severely limiting the dynamic range; on the other hand, atmospheric scattering leads to attenuation of the contrast between the target and the background, reducing detection accuracy; in addition, time-varying noise generated by multipath reflection from the sea surface and dynamic cloud cover can cause drastic fluctuations in the signal-to-noise ratio, easily leading to misjudgment; more importantly, single light intensity data lacks spectral and polarization information, making it difficult to suppress stray light from the sea surface or identify target materials through extinction effects, severely limiting the ability to resist interference and accurately identify targets in unknown environments.

[0003] In conclusion, improving the accuracy and anti-interference capability of target detection in complex environments is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a target detection method, device, equipment, and medium based on biomimetic vision, which can improve the accuracy and anti-interference capability of target detection in complex environments. The specific solution is as follows:

[0005] Firstly, this application provides a target detection method based on biomimetic vision, including:

[0006] By simulating the biological perception mechanism of mantis shrimp and constructing a target biomimetic detection model based on neural networks, the target spectral information and target polarization information map are acquired during the process of using the target biomimetic detection model to detect targets in the environment.

[0007] Construct a spectral feature pyramid corresponding to the target spectral information, determine the target polarization feature corresponding to the target polarization information map, and generate target multimodal fusion features based on the spectral feature pyramid and the target polarization feature using a preset feature concatenation method;

[0008] A target spatial attention map is generated based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and a target feature tensor map is generated based on a cross-scale fusion attention mechanism.

[0009] The target enhancement features are generated by using residual skip connections and channel attention mechanisms, and based on the target space attention map and the target feature tensor map.

[0010] Based on the manifold learning method, Lie group operation, the target enhancement features, and the target space attention graph, a corresponding target feature representation space is constructed, and the target feature representation space is decoded to obtain the corresponding target detection results.

[0011] Optionally, constructing the spectral feature pyramid corresponding to the target spectral information and determining the target polarization features corresponding to the target polarization information map includes:

[0012] The local spectral features of each band in the target spectral information are extracted using a receptive field convolution kernel of a first preset size, and a spectral feature tensor is generated based on the channel attention mechanism and the local spectral features.

[0013] Multi-scale convolution, batch normalization, and nonlinear activation operations are performed on the spectral feature tensor to generate the spectral feature pyramid corresponding to the target spectral information.

[0014] Perform multi-angle convolution operation on the target polarization information map to extract the polarization angle and polarization degree features of the target polarization information map;

[0015] The channel attention mechanism is used to adaptively weight the polarization angle, the polarization degree feature, and the target polarization information map to generate the target polarization feature corresponding to the target polarization information map.

[0016] Optionally, the step of generating target multimodal fusion features based on the spectral feature pyramid and the target polarization features using a preset feature concatenation method includes:

[0017] The spectral feature pyramid and the target polarization feature are channel-cascaded to obtain a first multimodal fusion feature, and the channels corresponding to the first multimodal fusion feature are compressed using a convolution kernel of a second preset size to obtain a second multimodal fusion feature.

[0018] The second multimodal fusion feature is fused based on a multi-head attention mechanism to generate the target multimodal fusion feature.

[0019] Optionally, the step of generating a target spatial attention map based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and generating a target feature tensor map based on a cross-scale fusion attention mechanism, includes:

[0020] Based on the recursive segmentation of the target multimodal fusion features using an image pyramid, target fractal windows of several scales are obtained. The pixel features corresponding to each target fractal window are then reduced in dimension and aggregated to obtain the target cluster vector corresponding to each target fractal window. The target fractal window is a local region of the target multimodal fusion features.

[0021] The target cluster vector is divided into an initial vector subset based on the scale of the target fractal window corresponding to the target cluster vector, and the target cluster vector in the initial vector subset is divided into a target vector subset based on the color channel corresponding to the target cluster vector;

[0022] The self-attention weights of the target vector subsets are calculated using preset parallel attention heads, and the contribution of the preset parallel attention heads is dynamically adjusted based on a first preset learnable scaling parameter. The self-attention weights are then weighted and summed based on the contribution to obtain the target mask tensor.

[0023] Based on the target cluster vector and the target mask tensor, a target spatial attention map corresponding to the target multimodal fusion features is generated, and based on the cross-scale fusion attention mechanism and the target mask tensor, a target feature tensor map is generated.

[0024] Optionally, the step of generating corresponding target enhancement features based on the target space attention map and the target feature tensor map through residual skip connections and channel attention mechanisms includes:

[0025] The target feature tensor map is translated, and the difference feature map between the target feature tensor map before translation and the target feature tensor map after translation is determined;

[0026] The difference feature map is subjected to convolution, batch normalization and nonlinear activation operations to generate residual mapping features corresponding to the difference feature map, and the first feature is generated based on the residual mapping features and the target feature tensor map before translation using the preset feature concatenation method.

[0027] The target feature tensor map before translation is subjected to local comparison processing to obtain a local comparison map, and the local comparison map is enhanced based on a preset convolutional block. The enhanced local comparison map is then processed through the channel attention mechanism to obtain a second feature.

[0028] Attention weights are generated based on the target space attention graph, and the first feature and the second feature are fused according to the second preset learnable scaling parameter, the residual skip connection and the attention weights to obtain the third feature. The target enhancement feature is generated based on the gating mechanism and the third feature.

[0029] Optionally, the step of constructing a corresponding target feature representation space based on manifold learning methods, Lie group operations, the target enhancement features, and the target space attention graph, and decoding the target feature representation space to obtain the corresponding target detection result, includes:

[0030] Obtain scene perception information in the current environment, and perform distribution estimation on the target enhancement feature and the target spatial attention map respectively to obtain the first distribution estimation information corresponding to the target enhancement feature and the second distribution estimation information corresponding to the target spatial attention map;

[0031] The target Lie group space is constructed through the Lie group operation and based on the scene perception information, the first distribution estimation information, and the second distribution estimation information.

[0032] Based on the manifold learning method, a target manifold structure corresponding to the first distribution estimation information and the second distribution estimation information is established in the target Lie group space;

[0033] The target feature representation space is constructed based on the target Lie group space and the target manifold structure, and the target feature representation space is decoded to obtain the corresponding target detection result.

[0034] Optionally, the step of constructing the target feature representation space based on the target Lie group space and the target manifold structure, and decoding the target feature representation space to obtain the corresponding target detection result, includes:

[0035] The target interaction features are obtained by interacting the target augmentation features, the first distribution estimation information, and the second distribution estimation information using the target Lie group space and the target manifold structure;

[0036] The target interaction features are subjected to graph convolution operation, and the target interaction features after graph convolution are reconstructed to obtain the target feature representation space;

[0037] The target feature representation space is decoded using a preset decoder to obtain the corresponding target detection result.

[0038] Secondly, this application provides a target detection device based on biomimetic vision, comprising:

[0039] The information acquisition module is used to construct a target biomimetic detection model based on a neural network by simulating the biological perception mechanism of mantis shrimp. During the process of using the target biomimetic detection model to detect targets in the environment, the module acquires the collected target spectral information and target polarization information map.

[0040] The target multimodal fusion feature generation module is used to construct a spectral feature pyramid corresponding to the target spectral information, determine the target polarization feature corresponding to the target polarization information map, and generate target multimodal fusion features based on the spectral feature pyramid and the target polarization feature using a preset feature concatenation method;

[0041] The target feature tensor graph generation module is used to generate a target spatial attention graph based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and to generate a target feature tensor graph based on a cross-scale fusion attention mechanism.

[0042] The target enhancement feature generation module is used to generate corresponding target enhancement features based on the target space attention map and the target feature tensor map through residual skip connections and channel attention mechanisms.

[0043] The target detection result determination module is used to construct a corresponding target feature representation space based on manifold learning methods, Lie group operations, the target enhancement features, and the target spatial attention graph, and to decode the target feature representation space to obtain the corresponding target detection result.

[0044] Thirdly, this application provides an electronic device, comprising:

[0045] Memory, used to store computer programs;

[0046] A processor is used to execute the computer program to implement the aforementioned target detection method based on biomimetic vision.

[0047] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned target detection method based on bionic vision.

[0048] In this application, a target biomimetic detection model is first constructed by simulating the biological perception mechanism of mantis shrimp and based on a neural network. During the detection of targets in the environment using this model, the collected target spectral information and target polarization information map are acquired. Then, a spectral feature pyramid corresponding to the target spectral information is constructed, and the target polarization features corresponding to the target polarization information map are determined. A multimodal fusion feature of the target is generated based on the spectral feature pyramid and the target polarization features using a preset feature concatenation method. Subsequently, a target spatial attention map is generated based on a preset recursive fractal partitioning mechanism and the target multimodal fusion feature, and a target feature tensor map is generated based on a cross-scale fusion attention mechanism. Then, corresponding target enhancement features are generated based on the target spatial attention map and the target feature tensor map using residual skip connections and channel attention mechanisms. Finally, a corresponding target feature representation space is constructed based on manifold learning methods, Lie group operations, the target enhancement features, and the target spatial attention map, and the target feature representation space is decoded to obtain the corresponding target detection results. As can be seen, this application first constructs a target biomimetic detection model based on a neural network architecture by simulating the biological perception mechanism of mantis shrimp. In the process of conducting environmental detection tasks using a biomimetic target detection model, the first step is to acquire the collected target spectral information and target polarization information map. Then, a spectral feature pyramid corresponding to the target spectral information is constructed. Simultaneously, the polarization features corresponding to the target polarization information map are determined. Using a pre-defined feature concatenation method, the spectral feature pyramid and target polarization features are fused to generate a multimodal fusion feature of the target. Subsequently, based on a pre-defined recursive fractal partitioning mechanism and the generated multimodal fusion feature of the target, a target spatial attention map is generated; and a target feature tensor map is generated using a cross-scale fusion attention mechanism. Next, residual skip connections and channel attention mechanisms are introduced, and combined with the generated target spatial attention map and target feature tensor map, corresponding target enhancement features are generated. Then, based on manifold learning methods, Lie group operations, the generated target enhancement features, and the target spatial attention map, a corresponding target feature representation space is constructed. Finally, the target feature representation space is decoded to obtain the corresponding target detection results. In this way, this application can simulate the biological perception mechanism of mantis shrimp and construct a target biomimetic detection model with environmental adaptability based on neural networks. Meanwhile, this application addresses the problem of information redundancy in multimodal input data during network inference in existing technologies, improves the efficiency of model information extraction, and performs parallel processing on the collected multimodal information. As a result, this application can enhance the accuracy and anti-interference capability of target detection in complex environments, and promote optical detection technology from a single dimension to a new paradigm of bio-inspired multimodal fusion. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0050] Figure 1 A flowchart of a target detection method based on biomimetic vision is provided for this application;

[0051] Figure 2 A flowchart illustrating a specific network parameter control method provided in this application;

[0052] Figure 3 A flowchart illustrating the workflow of a specific channel-wise depthwise convolution module provided in this application;

[0053] Figure 4 A flowchart illustrating a specific fractal multi-head latent attention module provided in this application;

[0054] Figure 5 A flowchart illustrating the workflow of a specific side-suppression fusion module provided in this application;

[0055] Figure 6 A schematic diagram of a target detection device based on biomimetic vision is provided for this application;

[0056] Figure 7 This application provides a structural diagram of an electronic device. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] In recent years, airborne ground and sea detection technologies have made significant progress in various fields such as ship detection, environmental monitoring, and disaster early warning, providing important technical support for target identification under complex sea and atmospheric conditions. However, traditional methods relying solely on light intensity detection still face several significant bottlenecks in complex environments: on the one hand, specular reflection from the ocean and strong solar flares often cause high light saturation in detectors, obscuring key details such as ship shadows and wakes, severely limiting the dynamic range; on the other hand, atmospheric scattering leads to attenuation of the contrast between the target and the background, reducing detection accuracy; furthermore, time-varying noise generated by multipath reflection from the sea surface and dynamic cloud cover can cause drastic fluctuations in the signal-to-noise ratio, easily leading to misjudgments; more importantly, single light intensity data lacks spectral and polarization information, making it difficult to suppress stray light from the sea surface or identify target materials through extinction effects, severely limiting the anti-interference and accurate identification capabilities in unknown environments. To address this, this application provides a target detection scheme based on biomimetic vision, which can improve the accuracy and anti-interference capabilities of target detection in complex environments.

[0059] See Figure 1 As shown, this embodiment of the invention discloses a target detection method based on biomimetic vision, which may include:

[0060] Step S11: By simulating the biological perception mechanism of mantis shrimp and constructing a target biomimetic detection model based on a neural network, the target spectral information and target polarization information map are acquired during the process of using the target biomimetic detection model to detect targets in the environment.

[0061] In this embodiment, a target biomimetic detection model is constructed based on a neural network by simulating the biological perception mechanism of mantis shrimps. This model includes a channel-specific deep convolution module, a fractal multi-head latent attention module, a lateral inhibition fusion module, and a deep feature integration manifold module. It should be noted that this embodiment aims to simulate the synergistic effect of the mantis shrimp's innate phylogenetic memory and acquired experiential learning-modulated amine substances, achieving dynamic regulation and optimization of the feature manifold. Specifically, the target biomimetic detection model incorporates the mantis shrimp's innate characteristic phylogenetic memory, combined with the biomimetic acquired experiential learning-modulated amine substance regulation mechanism under different environments. Through the process of establishing and integrating basic features, verifying the correctness of feature integration, establishing the feature manifold, and optimizing the manifold, the model simulates the mantis shrimp's ability to dynamically regulate the overall excitability of neurons by modulating receptor sensitivity and ion channel opening probability through common modulating molecules such as serotonin and dopamine, under its inherent genetic information integration mechanism. This allows it to adapt to complex changes in spectral, polarization, and motion information in the environment. In this embodiment, features are directly embedded into the model, mimicking the characteristic phylogenetic memory pattern of mantis shrimp. This provides effective guidance in the early stages of feature extraction and integration, reducing reliance on large-scale data and long-term gradient training, thereby improving training efficiency and robustness. See also Figure 2As shown, the core objective of the first stage, establishing and integrating basic features, is to construct the network's genetic-like information features, enabling it to generate a data distribution that most directly and simply approximates the expected output, given a pre-existing input. To eliminate all interference from "acquired information," this embodiment sets an ideal scenario: when the input data itself is not distorted, i.e. Then the mapping problem becomes simple and clear. In this case, the optimal mapping operator should satisfy... Therefore, it can be deduced that when the input is... At times, there should also be This achieves the ideal effect of information perception. Distance metric is introduced. Specifically, the drift error can be expressed as:

[0062] ;

[0063] in, For model parameters, This is a function identifier.

[0064] The second stage involves verifying the correctness of feature integration. Guided by innate "genetic" information, the model encounters and perceives external input for the first time. The core task at this stage is to learn a parameterizable mapping operator. This enables the distorted observation signal z to reproduce the representation of the clean sample x as closely as possible in the output space after mapping, i.e., let and Maintaining high consistency within the embedding spaces of Lie groups and manifold learning allows for the integration and construction of key information. Ideally, this process is equivalent to minimizing the difference between their distributions while satisfying predefined optimization constraints. Therefore, the error term of the initial empirical learning can be formalized as:

[0065] ;

[0066] The third and fourth stages primarily focus on further optimizing output characteristics, aiming to simulate the mantis shrimp's information integration mechanism based on its genes. This involves utilizing common modulating molecules such as serotonin and dopamine to dynamically regulate the overall excitability of neurons by adjusting receptor sensitivity and ion channel opening probability. Therefore, this embodiment sets a neuron-like excitation threshold, modulating the upper and lower bound thresholds of the neural network, leading to the third and fourth optimization constraints:

[0067] ;

[0068] in, For parameters that are frozen, gradient backpropagation is not performed. This is a function identifier.

[0069] ;

[0070] Understandably, traditional deep learning mainly relies on a single gradient descent method for weight updates, which often struggles to adapt to the variability of spectral, polarization, and motion information in the environment. In this embodiment, the process of modulating molecules to regulate receptor sensitivity and ion channel opening probability is simulated, enabling the model to flexibly adjust according to the external environment and achieve rapid adaptation and effective integration of multimodal information.

[0071] Step S12: Construct the spectral feature pyramid corresponding to the target spectral information, determine the target polarization feature corresponding to the target polarization information map, and generate target multimodal fusion features based on the spectral feature pyramid and the target polarization feature using a preset feature concatenation method.

[0072] In this embodiment, see Figure 3 As shown, the multi-channel deep convolution module of the target biomimetic detection model adopts a deep neural network architecture with multi-branch parallel processing, which can extract and fuse features for two types of heterogeneous information: spectral and polarization. The module consists of two parts: a spectral channel processing unit and a polarization channel processing unit. In the middle, feature cascading and channel attention mechanisms are used to achieve cross-compensation and enhancement of multimodal information, thereby simulating the biological preprocessing and response strategy of the mantis shrimp compound eye to light signals of different wavelengths and polarization states.

[0073] It should be noted that the above-mentioned construction of the spectral feature pyramid corresponding to the target spectral information may include: firstly, using a convolutional kernel with a receptive field of a first preset size to extract local spectral features of each band in the target spectral information, and generating a spectral feature tensor based on the channel attention mechanism and the local spectral features; then, performing multi-scale convolution, batch normalization, and nonlinear activation operations on the spectral feature tensor to generate the spectral feature pyramid corresponding to the target spectral information. Specifically, the spectral channel processing unit is structurally divided into a low-level spectral response layer and a high-level spectral fusion layer. First, the low-level spectral response layer uses a small receptive field convolutional kernel and combines it with a channel attention mechanism to perform preliminary filtering and feature compression on the channels of the input target spectral information to extract the specific response of each band, and finally outputs the spectral feature tensor of each channel. Then, the high-level spectral fusion layer constructs the spectral feature pyramid by performing stacked multi-scale convolution, batch normalization, and nonlinear activation operations on the extracted spectral feature tensor to achieve information interaction and global feature integration between different bands.

[0074] In this embodiment, determining the target polarization feature corresponding to the target polarization information map can include: firstly, performing a multi-angle convolution operation on the target polarization information map to extract the polarization angle and polarization degree features; then, through the channel attention mechanism, adaptively weighting the polarization angle, the polarization degree features, and the target polarization information map to generate the target polarization feature corresponding to the target polarization information map. Specifically, the polarization channel processing unit is structurally divided into two parts: a direction-sensitive convolutional layer and a polarization feature enhancement layer. First, the direction-sensitive convolutional layer uses multi-angle oriented convolutional kernels to filter the input target polarization information map in parallel to finely extract the polarization angle and polarization degree features. Then, the polarization feature enhancement layer adaptively weights the polarization angle and polarization degree features with the target polarization information map through the channel attention mechanism, wherein the target polarization information map forms a global polarization map after one convolution transformation. Then, the global polarization map is combined to enhance the discrimination ability of polarization information, outputting a target polarization feature with direction sensitivity and polarization coupling characteristics.

[0075] It should be noted that the above-mentioned method of generating target multimodal fusion features based on the spectral feature pyramid and the target polarization features using a preset feature concatenation method may include: firstly, concatenating the spectral feature pyramid and the target polarization features through channels to obtain a first multimodal fusion feature, and then compressing the channels corresponding to the first multimodal fusion feature using a convolution kernel of a second preset size to obtain a second multimodal fusion feature; subsequently, fusing the second multimodal fusion feature based on a multi-head attention mechanism to generate the target multimodal fusion feature. Specifically, the channel-specific deep convolution module concatenates the output features of the spectral and polarization channels along the channel dimension in the middle part, then performs channel compression via a 1×1 convolution kernel, and introduces a multi-head attention mechanism for deep fusion, ultimately generating a unified feature representation with spectral-polarization-intensity coupling characteristics, i.e., the target multimodal fusion feature, providing a highly robust multimodal information foundation for subsequent target detection tasks.

[0076] Step S13: Generate a target space attention map based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and generate a target feature tensor map based on a cross-scale fusion attention mechanism.

[0077] In this embodiment, see Figure 4As shown, the fractal multi-head latent attention module of the target biomimetic detection model adopts a deep network architecture that combines recursive fractals with MLA (Multi-head Latent Attention) to simulate the high-speed windowing scanning mechanism of the mantis shrimp's compound eyes, enabling efficient "block-based" information capture and low-dimensional clustering representation of static images. This module consists of three parts: a fractal window partitioning unit, a multi-head latent attention unit, and a multi-scale fusion attention unit.

[0078] It should be noted that the above-mentioned generation of a target spatial attention map based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and generation of a target feature tensor map based on a cross-scale fusion attention mechanism, may include: firstly, recursively segmenting the target multimodal fusion features based on an image pyramid to obtain target fractal windows of several scales, and then performing dimensionality reduction and aggregation on the pixel features corresponding to each target fractal window to obtain a target cluster vector corresponding to each target fractal window; wherein, the target fractal window is a local region of the target multimodal fusion features; then, partitioning the target cluster vector based on the scale of the target fractal window corresponding to the target cluster vector to obtain an initial vector subset. The target vector subset is obtained by dividing the target cluster vector in the initial vector subset based on the color channel corresponding to the target cluster vector. Then, the self-attention weights of the target vector subset are calculated using a preset parallel attention head, and the contribution of the preset parallel attention head is dynamically adjusted based on a first preset learnable scaling parameter. The self-attention weights are then weighted and summed based on the contribution to obtain the target mask tensor. Finally, the target spatial attention map corresponding to the target multimodal fusion feature is generated based on the target cluster vector and the target mask tensor, and the target feature tensor map is generated based on the cross-scale fusion attention mechanism and the target mask tensor. Specifically, the fractal window partitioning unit consists of a recursive fractal generation layer and an information cluster mapping layer. The recursive fractal generation layer, based on image pyramid theory, recursively partitions the input target multimodal fusion features into several target fractal windows of different scales, ensuring that the color and texture distribution within each window is compactly presented in a low-dimensional subspace. The information cluster mapping layer performs dimensionality reduction and aggregation on the pixel features within each target fractal window, generating corresponding superpixel cluster vectors, i.e., target cluster vectors, for subsequent attention calculations. This significantly reduces the reliance on pixel-by-pixel high-precision processing, improving scanning speed and computational efficiency. Subsequently, the multi-head latent attention unit uses parallel attention heads containing multiple latent attention heads to calculate the self-attention weights of target vector subsets at different fractal scales and colors, capturing long-range dependencies between and within windows. Simultaneously, adaptive scaling is employed within the attention heads, introducing learnable scaling parameters to the output of each attention head to dynamically adjust the contribution of each attention head's output to the final feature representation. This simulates the mantis shrimp's rapid focusing mechanism on key color regions, and the self-attention weights are weighted and summed based on their contributions to obtain the target mask tensor. Subsequently, the multi-scale fusion attention unit adaptively weights the data according to the channel dimension based on global fractal window statistics, highlighting the spectral-polarization features that are highly distinguishable from the target.Finally, based on the target cluster vector and the target mask tensor, a spatial attention map is generated at the full-scale graph to guide the network to focus on regions with rich fractal color textures. A cross-scale fusion attention mechanism is introduced, and the target mask tensor is cascaded and fused through cross-level skip connections and multi-head fusion layers to obtain a target feature tensor map with global and fine-grained features, so as to ensure the coordinated expression of detailed and global information.

[0079] Step S14: Generate corresponding target enhancement features based on the target space attention map and the target feature tensor map by using residual skip connections and channel attention mechanisms.

[0080] In this embodiment, see Figure 5 As shown, the lateral inhibition fusion module of the target biomimetic detection model adopts a deep fusion network architecture based on residual coupling and local contrast strategies. The lateral inhibition fusion module includes the IGAB (Illumination-Guided Attention Block), designed to simulate the biological lateral inhibition effect of amplified signal differences between adjacent receptors in the mantis shrimp ganglion, enhancing the extraction of target edge and motion features. The IGAB module consists of LN (Layer Normalization), MLA (Multi-head Latent Attention), LN, and FFN (Feed-Forward Network). This module comprises three parts: a residual contrast mapping unit, a local contrast response unit, and an attention-gated inhibition unit. Efficient interaction and adaptive fusion of cross-layer, multi-channel information are achieved through residual skip connections and channel attention mechanisms.

[0081] It should be noted that the above-mentioned generation of corresponding target enhancement features based on the target space attention map and the target feature tensor map through residual skip connections and channel attention mechanisms may include: firstly, translating the target feature tensor map and determining the difference feature map between the target feature tensor map before and after the translation; then, performing convolution, batch normalization, and nonlinear activation operations on the difference feature map to generate residual mapping features corresponding to the difference feature map, and using the preset feature concatenation method based on the residual mapping features and the target feature tensor map before and after the translation. The target feature tensor map is used to generate a first feature; then, the target feature tensor map before translation is subjected to local comparison processing to obtain a local comparison map, and the local comparison map is enhanced based on a preset convolutional block, and the enhanced local comparison map is processed through the channel attention mechanism to obtain a second feature; finally, attention weights are generated based on the target space attention map, and the first feature and the second feature are fused according to a second preset learnable scaling parameter, the residual skip connection and the attention weights to obtain a third feature, and the target enhanced feature is generated based on the gating mechanism and the third feature. Specifically, the residual contrast mapping unit incorporates a neighborhood difference residual block. This block calculates the difference between the target feature tensor before and after spatial translation, yielding a difference feature map. Convolution, batch normalization, and nonlinear activation operations are then applied to this difference feature map to generate residual mapping features, amplifying signal differences between adjacent receptors. Simultaneously, residual fusion is applied, concatenating the residual mapping features with the target feature tensor before translation along the channel dimension. A 1×1 convolution kernel is used to reassemble the information, preserving the global context of the original signal, resulting in the first feature. Next, the local contrast response unit employs a local statistical strategy while windowing. Within a fixed-size window, it calculates the difference between the center pixel of the target feature tensor before translation and the average feature of its neighborhood, generating a local contrast map that highlights edge and detail changes. The local contrast response unit also includes a contrast feature enhancement method. This method inputs the local contrast map into a series of lightweight convolutional blocks and, combined with a channel attention mechanism, adaptively amplifies key signals to obtain the second feature. Finally, the attention-gated inhibition unit employs channel inhibition gating. Based on the mutual exclusion relationship between channels, a gating mechanism is used to apply negative inhibition weights to highly correlated channels, achieving information competition and lateral inhibition. Simultaneously, to achieve adaptability to diverse and complex environments, dynamic weighted fusion is adopted. The first and second features are weighted and fused using learnable scaling parameters and attention weights generated from the target space attention map, outputting inhibition-enhancing features that combine fine-grained variation with global consistency, thus obtaining the target enhancement features. By generating target enhancement features, differences in features similar to the environment can be highlighted, achieving adaptive inhibition enhancement.

[0082] Step S15: Construct a corresponding target feature representation space based on manifold learning method, Lie group operation, the target enhancement features and the target space attention graph, and decode the target feature representation space to obtain the corresponding target detection results.

[0083] In this embodiment, the deep feature integration manifold module of the target biomimetic detection model adopts a deep network architecture that combines manifold learning and graph convolution. It aims to simulate the ability of deep neurons in the mantis shrimp's optic lobe to converge and cross-integrate multimodal features in the intermediate layer over long distances, and to further perform high-dimensional spatial mapping and information fusion processing on the target enhancement features obtained above.

[0084] It should be noted that the above-mentioned construction of a corresponding target feature representation space based on manifold learning methods, Lie group operations, the target enhancement features, and the target spatial attention map, and decoding of the target feature representation space to obtain the corresponding target detection results, may include: firstly, acquiring scene perception information in the current environment, and performing distribution estimation on the target enhancement features and the target spatial attention map respectively to obtain first distribution estimation information corresponding to the target enhancement features and second distribution estimation information corresponding to the target spatial attention map; then, constructing a target feature representation space based on the scene perception information, the first distribution estimation information, and the second distribution estimation information using the Lie group operation. The target Lie group space is then used. Based on the manifold learning method, a target manifold structure corresponding to the first and second distribution estimation information is established in the target Lie group space. Then, the target Lie group space and the target manifold structure are used to interact with the target enhancement features, the first and second distribution estimation information to obtain target interaction features. Afterwards, graph convolution is performed on the target interaction features, and feature reconstruction is performed on the graph-convolved target interaction features to obtain the target feature representation space. Finally, a preset decoder is used to decode the target feature representation space to obtain the corresponding target detection result. Specifically, the scene perception information in the current environment is obtained from the gradient backpropagation results during the network optimization stage and can be directly obtained from network training. The deep feature integration manifold module performs distribution estimation on the target enhancement features and the target space attention map respectively to obtain the first distribution estimation information corresponding to the target enhancement features and the second distribution estimation information corresponding to the target space attention map. Then, information interaction and fusion operations for different feature sources are implemented in the Lie group space and the manifold structure, thereby establishing a more expressive information feature representation system. This process involves leveraging deep feature integration to incorporate built-in information interaction terms within the manifold module. The network-extracted target enhancement features are mapped and fused sequentially with the first distribution estimation information to obtain target interaction features. Simultaneously, the target space interest graph and the second distribution estimation information are mapped and fused sequentially to obtain target interaction features. Then, graph convolution operations are performed on these target interaction features, and feature reconstruction is conducted to construct a feature representation space with Lie group structure constraints and manifold continuity, resulting in the target feature representation space. Finally, a decoder is used to decode the target feature representation space to obtain the corresponding target detection results.

[0085] As can be seen from the above, this embodiment first simulates the biological perception mechanism of mantis shrimp and constructs a target biomimetic detection model based on a neural network. During the detection of targets in the environment using this biomimetic detection model, the collected target spectral information and target polarization information map are acquired. Then, a spectral feature pyramid corresponding to the target spectral information is constructed, and the target polarization features corresponding to the target polarization information map are determined. A preset feature concatenation method is used to generate target multimodal fusion features based on the spectral feature pyramid and the target polarization features. Subsequently, a target spatial attention map is generated based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and a target feature tensor map is generated based on a cross-scale fusion attention mechanism. Then, corresponding target enhancement features are generated based on the target spatial attention map and the target feature tensor map through residual skip connections and channel attention mechanisms. Finally, a corresponding target feature representation space is constructed based on manifold learning methods, Lie group operations, the target enhancement features, and the target spatial attention map, and the target feature representation space is decoded to obtain the corresponding target detection results. As can be seen from the above, this embodiment first simulates the biological perception mechanism of mantis shrimp and constructs a target biomimetic detection model based on a neural network architecture. In the process of conducting environmental detection tasks using a target biomimetic detection model, the first step is to acquire the collected target spectral information and target polarization information map. Then, a spectral feature pyramid corresponding to the target spectral information is constructed. Simultaneously, the polarization features corresponding to the target polarization information map are determined. Using a pre-defined feature concatenation method, the spectral feature pyramid and target polarization features are fused to generate a target multimodal fusion feature. Subsequently, based on a pre-defined recursive fractal partitioning mechanism and the generated target multimodal fusion feature, a target spatial attention map is generated; and a target feature tensor map is generated using a cross-scale fusion attention mechanism. Next, residual skip connections and channel attention mechanisms are introduced, and combined with the generated target spatial attention map and target feature tensor map, corresponding target enhancement features are generated. Then, based on manifold learning methods, Lie group operations, the generated target enhancement features, and the target spatial attention map, a corresponding target feature representation space is constructed. Finally, the target feature representation space is decoded to obtain the corresponding target detection results. In this way, this embodiment can simulate the biological perception mechanism of mantis shrimp and construct a target biomimetic detection model with environmental adaptability based on neural networks. Meanwhile, this embodiment solves the problem of information redundancy in multimodal input data during network inference in existing technologies, improves the information extraction efficiency of the model, and performs parallel processing on the collected multimodal information. As a result, this embodiment can improve the accuracy and anti-interference capability of target detection in complex environments, and promote optical detection technology from a single dimension to a new paradigm of bio-inspired multimodal fusion.

[0086] It should be noted that the target biomimetic detection model includes a multi-channel deep convolution module, a fractal multi-head latent attention module, a lateral inhibition fusion module, and a deep feature integration manifold module. The fractal multi-head latent attention module and the lateral inhibition fusion module will be described in detail in this embodiment.

[0087] In this embodiment, the fractal multi-head latent attention module of the target biomimetic detection model constructs a deep network structure that can efficiently simulate the rapid scanning of the compound eyes of a mantis shrimp by combining recursive fractal partitioning with a multi-head latent attention mechanism. Unlike traditional global or multi-scale sliding window self-attention, which requires maintaining independent convolutional kernels and attention weights for each scale or position, the fractal multi-head latent attention module significantly reduces the scale and computational cost through recursive fractal partitioning and parameter sharing. Specifically, let the input target feature tensor map be... If a traditional multi-scale window attention mechanism configures k×k convolutional kernels at L scales, the number of parameters is:

[0088] ;

[0089] By adopting a fractal structure, all scales share the same set of k×k convolutional kernels and attention heads, reducing the number of parameters to:

[0090] ;

[0091] Since the fractal partition is generated at the i-th level Each window contains [number] windows. With a pixel count, the attention computation is reduced to:

[0092] ;

[0093] This reduces the computational load from multiple channels in the original multimodal approach to approximately 1.33 channels, significantly improving computational efficiency and real-time performance while maintaining multi-scale, long-range dependency capture. The fractal multi-head latent attention module not only reproduces the "block-like" clustering perception pattern of the mantis shrimp's compound eye's second-level multi-point scanning but also achieves extremely low parameter count and computational complexity for miniaturized deployment by relying on recursive fractals and parameter sharing strategies. Finally, the multi-scale fusion attention unit of the fractal multi-head latent attention module uses global fractal window statistics to adaptively weight the channels to highlight the most discriminative spectral-polarization features and generate a spatial attention map to guide the network to focus on key regions rich in fractal texture. Based on this, through cross-scale skip connections and multi-head fusion layers, attention results from various fractal scales are cascaded and fused, ensuring that detailed and global information are co-expressed in the same representation, providing a rich and robust multimodal feature foundation for subsequent target recognition or reconstruction tasks.

[0094] In this embodiment, the lateral inhibition fusion module of the target biomimetic detection model is built on a deep network framework of residual coupling and local contrast, aiming to reproduce the lateral inhibition effect of the contrast gain of adjacent receptors in the mantis shrimp ganglion, thereby enhancing the response of edge and motion features. Specifically, the residual contrast mapping unit maps the input target feature tensor F with a version translated by spatial offsets Δx and Δy. Calculate the difference The residual mapping feature R is generated in the convolution-batch normalization-activation structure. The residual mapping feature R is concatenated with the original target feature tensor F along the channel dimension, and then subjected to a 1×1 convolution. Achieve feature recombination:

[0095] ;

[0096] Here, [F,R] represents the concatenation of the residual mapping feature R with the original target feature tensor F along the channel dimension. This design amplifies subtle differences between adjacent receptors while preserving global semantic information, providing a significant signal basis for subsequent contrast processing. The local contrast response unit of the lateral suppression fusion module uses a fixed-size m×m window to calculate the local contrast mapping between the center pixel and its neighborhood average value. The resulting local contrast map is then fed into a series of lightweight convolutional blocks and channel attention modules to adaptively enhance significant contrast features.

[0097] In this way, this embodiment not only significantly reduces hardware complexity and computational burden, but also enables real-time, high-precision, and interference-resistant multi-scenario target detection on miniaturized carriers such as UAVs and miniature avionics, significantly improving the reliability and adaptability of detection under complex sea conditions and variable atmospheric conditions.

[0098] Accordingly, see Figure 6 As shown in the embodiments of this application, a target detection device based on bionic vision is also provided, which may include:

[0099] The information acquisition module 11 is used to construct a target biomimetic detection model by simulating the biological perception mechanism of mantis shrimp and based on a neural network. During the process of using the target biomimetic detection model to detect targets in the environment, the module acquires the collected target spectral information and target polarization information map.

[0100] The target multimodal fusion feature generation module 12 is used to construct the spectral feature pyramid corresponding to the target spectral information, determine the target polarization feature corresponding to the target polarization information map, and generate target multimodal fusion features based on the spectral feature pyramid and the target polarization feature using a preset feature concatenation method;

[0101] The target feature tensor graph generation module 13 is used to generate a target spatial attention graph based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and to generate a target feature tensor graph based on a cross-scale fusion attention mechanism.

[0102] The target enhancement feature generation module 14 is used to generate corresponding target enhancement features based on the target space attention map and the target feature tensor map through residual skip connections and channel attention mechanisms.

[0103] The target detection result determination module 15 is used to construct a corresponding target feature representation space based on the manifold learning method, Lie group operation, the target enhancement features and the target spatial attention graph, and to decode the target feature representation space to obtain the corresponding target detection result.

[0104] As can be seen from the above, this application first simulates the biological perception mechanism of mantis shrimp and constructs a target biomimetic detection model based on a neural network. During the detection of targets in the environment using this biomimetic detection model, the collected target spectral information and target polarization information map are acquired. Then, a spectral feature pyramid corresponding to the target spectral information is constructed, and the target polarization features corresponding to the target polarization information map are determined. A preset feature concatenation method is used to generate target multimodal fusion features based on the spectral feature pyramid and the target polarization features. Subsequently, a target spatial attention map is generated based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and a target feature tensor map is generated based on a cross-scale fusion attention mechanism. Then, corresponding target enhancement features are generated based on the target spatial attention map and the target feature tensor map through residual skip connections and channel attention mechanisms. Finally, a corresponding target feature representation space is constructed based on manifold learning methods, Lie group operations, the target enhancement features, and the target spatial attention map, and the target feature representation space is decoded to obtain the corresponding target detection results. As can be seen from the above, this application first simulates the biological perception mechanism of mantis shrimp and constructs a target biomimetic detection model based on a neural network architecture. In the process of conducting environmental detection tasks using a biomimetic target detection model, the first step is to acquire the collected target spectral information and target polarization information map. Then, a spectral feature pyramid corresponding to the target spectral information is constructed. Simultaneously, the polarization features corresponding to the target polarization information map are determined. Using a pre-defined feature concatenation method, the spectral feature pyramid and target polarization features are fused to generate a multimodal fusion feature of the target. Subsequently, based on a pre-defined recursive fractal partitioning mechanism and the generated multimodal fusion feature of the target, a target spatial attention map is generated; and a target feature tensor map is generated using a cross-scale fusion attention mechanism. Next, residual skip connections and channel attention mechanisms are introduced, and combined with the generated target spatial attention map and target feature tensor map, corresponding target enhancement features are generated. Then, based on manifold learning methods, Lie group operations, the generated target enhancement features, and the target spatial attention map, a corresponding target feature representation space is constructed. Finally, the target feature representation space is decoded to obtain the corresponding target detection results. In this way, this application can simulate the biological perception mechanism of mantis shrimp and construct a target biomimetic detection model with environmental adaptability based on neural networks. Meanwhile, this application addresses the problem of information redundancy in multimodal input data during network inference in existing technologies, improves the efficiency of model information extraction, and performs parallel processing on the collected multimodal information. As a result, this application can enhance the accuracy and anti-interference capability of target detection in complex environments, and promote optical detection technology from a single dimension to a new paradigm of bio-inspired multimodal fusion.

[0105] In some specific embodiments, the target multimodal fusion feature generation module 12 may include:

[0106] The spectral feature tensor generation unit is used to extract local spectral features of each band in the target spectral information using a receptive field convolution kernel of a first preset size, and to generate a spectral feature tensor based on the channel attention mechanism and the local spectral features.

[0107] The spectral feature pyramid generation unit is used to perform multi-scale convolution, batch normalization and nonlinear activation operations on the spectral feature tensor to generate the spectral feature pyramid corresponding to the target spectral information.

[0108] The polarization degree feature extraction unit is used to perform multi-angle convolution operation on the target polarization information map to extract the polarization angle and polarization degree features of the target polarization information map;

[0109] The target polarization feature generation unit is used to adaptively weight the polarization angle, the polarization degree feature, and the target polarization information map through the channel attention mechanism to generate the target polarization feature corresponding to the target polarization information map.

[0110] In some specific embodiments, the target multimodal fusion feature generation module 12 may include:

[0111] The second multimodal fusion feature determination unit is used to cascade the spectral feature pyramid and the target polarization feature to obtain the first multimodal fusion feature, and to compress the channel corresponding to the first multimodal fusion feature using a convolution kernel of the second preset size to obtain the second multimodal fusion feature.

[0112] The target multimodal fusion feature generation unit is used to fuse the second multimodal fusion feature based on a multi-head attention mechanism to generate the target multimodal fusion feature.

[0113] In some specific embodiments, the target feature tensor map generation module 13 may include:

[0114] The target cluster vector determination unit is used to recursively segment the target multimodal fusion feature based on the image pyramid to obtain target fractal windows of several scales, and to perform dimensionality reduction and aggregation on the pixel features corresponding to each target fractal window to obtain the target cluster vector corresponding to each target fractal window; wherein, the target fractal window is a local region of the target multimodal fusion feature;

[0115] The target vector subset determination unit is used to divide the target cluster vector based on the scale of the target fractal window corresponding to the target cluster vector to obtain an initial vector subset, and to divide the target cluster vector in the initial vector subset based on the color channel corresponding to the target cluster vector to obtain a target vector subset;

[0116] The target mask tensor determination unit is used to calculate the self-attention weights of the target vector subsets using preset parallel attention heads, dynamically adjust the contribution of the preset parallel attention heads based on a first preset learnable scaling parameter, and perform a weighted summation of the self-attention weights based on the contribution to obtain the target mask tensor.

[0117] The target feature tensor graph generation unit is used to generate a target spatial attention graph corresponding to the target multimodal fusion features based on the target cluster vector and the target mask tensor, and to generate the target feature tensor graph based on the cross-scale fusion attention mechanism and the target mask tensor.

[0118] In some specific embodiments, the target enhancement feature generation module 14 may include:

[0119] The difference feature map determination unit is used to translate the target feature tensor map and determine the difference feature map between the target feature tensor map before translation and the target feature tensor map after translation.

[0120] The first feature generation unit is used to perform convolution, batch normalization and nonlinear activation operations on the difference feature map to generate residual mapping features corresponding to the difference feature map, and to generate a first feature based on the residual mapping features and the target feature tensor map before translation using the preset feature concatenation method.

[0121] The second feature determination unit is used to perform local comparison processing on the target feature tensor map before translation to obtain a local comparison map, perform feature enhancement on the local comparison map based on a preset convolutional block, and process the enhanced local comparison map through the channel attention mechanism to obtain a second feature.

[0122] The target enhancement feature generation unit is used to generate attention weights based on the target space attention map, perform feature fusion on the first feature and the second feature according to the second preset learnable scaling parameter, the residual skip connection and the attention weights to obtain a third feature, and generate the target enhancement feature based on the gating mechanism and the third feature.

[0123] In some specific embodiments, the target detection result determination module 15 may include:

[0124] The distribution estimation information determination submodule is used to acquire scene perception information in the current environment, and to perform distribution estimation on the target enhancement feature and the target spatial attention map respectively, so as to obtain the first distribution estimation information corresponding to the target enhancement feature and the second distribution estimation information corresponding to the target spatial attention map;

[0125] The target Lie group space construction submodule is used to construct the target Lie group space through the Lie group operation and based on the scene perception information, the first distribution estimation information and the second distribution estimation information;

[0126] The target manifold structure establishment submodule is used to establish the target manifold structure corresponding to the first distribution estimation information and the second distribution estimation information in the target Lie group space based on the manifold learning method.

[0127] The target detection result determination submodule is used to construct the target feature representation space based on the target Lie group space and the target manifold structure, and decode the target feature representation space to obtain the corresponding target detection result.

[0128] In some specific implementations, the target detection result determination submodule may include:

[0129] The target interaction feature determination unit is used to obtain target interaction features by interacting the target enhancement features, the first distribution estimation information and the second distribution estimation information using the target Lie group space and the target manifold structure;

[0130] The target feature representation space determination unit is used to perform graph convolution operation on the target interaction features and to reconstruct the target interaction features after graph convolution to obtain the target feature representation space.

[0131] The target detection result determination unit is used to decode the target feature representation space using a preset decoder to obtain the corresponding target detection result.

[0132] Furthermore, embodiments of this application also disclose an electronic device, Figure 7 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the target detection method based on bionic vision disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0133] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0134] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0135] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the bionic vision-based target detection method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0136] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned biomimetic vision-based target detection method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0137] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0138] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0139] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0140] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0141] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A target detection method based on biomimetic vision, characterized in that, include: By simulating the biological perception mechanism of mantis shrimp and constructing a target biomimetic detection model based on neural networks, the target spectral information and target polarization information map are acquired during the process of using the target biomimetic detection model to detect targets in the environment. Construct a spectral feature pyramid corresponding to the target spectral information, determine the target polarization feature corresponding to the target polarization information map, and generate target multimodal fusion features based on the spectral feature pyramid and the target polarization feature using a preset feature concatenation method; A target spatial attention map is generated based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and a target feature tensor map is generated based on a cross-scale fusion attention mechanism. The target enhancement features are generated by using residual skip connections and channel attention mechanisms, and based on the target space attention map and the target feature tensor map. Based on the manifold learning method, Lie group operation, the target enhancement features, and the target space attention graph, a corresponding target feature representation space is constructed, and the target feature representation space is decoded to obtain the corresponding target detection results.

2. The target detection method based on bionic vision according to claim 1, characterized in that, The construction of the spectral feature pyramid corresponding to the target spectral information and the determination of the target polarization features corresponding to the target polarization information map include: The local spectral features of each band in the target spectral information are extracted using a receptive field convolution kernel of a first preset size, and a spectral feature tensor is generated based on the channel attention mechanism and the local spectral features. Multi-scale convolution, batch normalization, and nonlinear activation operations are performed on the spectral feature tensor to generate the spectral feature pyramid corresponding to the target spectral information. Perform multi-angle convolution operation on the target polarization information map to extract the polarization angle and polarization degree features of the target polarization information map; The channel attention mechanism is used to adaptively weight the polarization angle, the polarization degree feature, and the target polarization information map to generate the target polarization feature corresponding to the target polarization information map.

3. The target detection method based on bionic vision according to claim 1, characterized in that, The method of generating target multimodal fusion features based on the spectral feature pyramid and the target polarization features using a preset feature concatenation method includes: The spectral feature pyramid and the target polarization feature are channel-cascaded to obtain a first multimodal fusion feature, and the channels corresponding to the first multimodal fusion feature are compressed using a convolution kernel of a second preset size to obtain a second multimodal fusion feature. The second multimodal fusion feature is fused based on a multi-head attention mechanism to generate the target multimodal fusion feature.

4. The target detection method based on bionic vision according to claim 1, characterized in that, The step of generating a target spatial attention map based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and generating a target feature tensor map based on a cross-scale fusion attention mechanism, includes: Based on the recursive segmentation of the target multimodal fusion features using an image pyramid, target fractal windows of several scales are obtained. The pixel features corresponding to each target fractal window are then reduced in dimension and aggregated to obtain the target cluster vector corresponding to each target fractal window. The target fractal window is a local region of the target multimodal fusion features. The target cluster vector is divided into an initial vector subset based on the scale of the target fractal window corresponding to the target cluster vector, and the target cluster vector in the initial vector subset is divided into a target vector subset based on the color channel corresponding to the target cluster vector; The self-attention weights of the target vector subsets are calculated using preset parallel attention heads, and the contribution of the preset parallel attention heads is dynamically adjusted based on a first preset learnable scaling parameter. The self-attention weights are then weighted and summed based on the contribution to obtain the target mask tensor. Based on the target cluster vector and the target mask tensor, a target spatial attention map corresponding to the target multimodal fusion features is generated, and based on the cross-scale fusion attention mechanism and the target mask tensor, a target feature tensor map is generated.

5. The target detection method based on biomimetic vision according to claim 1, characterized in that, The process of generating corresponding target enhancement features based on the target space attention map and the target feature tensor map through residual skip connections and channel attention mechanisms includes: The target feature tensor map is translated, and the difference feature map between the target feature tensor map before translation and the target feature tensor map after translation is determined; The difference feature map is subjected to convolution, batch normalization and nonlinear activation operations to generate residual mapping features corresponding to the difference feature map, and the first feature is generated based on the residual mapping features and the target feature tensor map before translation using the preset feature concatenation method. The target feature tensor map before translation is subjected to local comparison processing to obtain a local comparison map, and the local comparison map is enhanced based on a preset convolutional block. The enhanced local comparison map is then processed through the channel attention mechanism to obtain a second feature. Attention weights are generated based on the target space attention graph, and the first feature and the second feature are fused according to the second preset learnable scaling parameter, the residual skip connection and the attention weights to obtain the third feature. The target enhancement feature is generated based on the gating mechanism and the third feature.

6. The target detection method based on biomimetic vision according to any one of claims 1 to 5, characterized in that, The process of constructing a corresponding target feature representation space based on manifold learning, Lie group operations, the target enhancement features, and the target space attention graph, and then decoding the target feature representation space to obtain the corresponding target detection results, includes: Obtain scene perception information in the current environment, and perform distribution estimation on the target enhancement feature and the target spatial attention map respectively to obtain the first distribution estimation information corresponding to the target enhancement feature and the second distribution estimation information corresponding to the target spatial attention map; The target Lie group space is constructed through the Lie group operation and based on the scene perception information, the first distribution estimation information, and the second distribution estimation information. Based on the manifold learning method, a target manifold structure corresponding to the first distribution estimation information and the second distribution estimation information is established in the target Lie group space; The target feature representation space is constructed based on the target Lie group space and the target manifold structure, and the target feature representation space is decoded to obtain the corresponding target detection result.

7. The target detection method based on bionic vision according to claim 6, characterized in that, The step of constructing the target feature representation space based on the target Lie group space and the target manifold structure, and decoding the target feature representation space to obtain the corresponding target detection result, includes: The target interaction features are obtained by interacting the target augmentation features, the first distribution estimation information, and the second distribution estimation information using the target Lie group space and the target manifold structure; The target interaction features are subjected to graph convolution operation, and the target interaction features after graph convolution are reconstructed to obtain the target feature representation space; The target feature representation space is decoded using a preset decoder to obtain the corresponding target detection result.

8. A target detection device based on biomimetic vision, characterized in that, include: The information acquisition module is used to construct a target biomimetic detection model based on a neural network by simulating the biological perception mechanism of mantis shrimp. During the process of using the target biomimetic detection model to detect targets in the environment, the module acquires the collected target spectral information and target polarization information map. The target multimodal fusion feature generation module is used to construct a spectral feature pyramid corresponding to the target spectral information, determine the target polarization feature corresponding to the target polarization information map, and generate target multimodal fusion features based on the spectral feature pyramid and the target polarization feature using a preset feature concatenation method; The target feature tensor graph generation module is used to generate a target spatial attention graph based on a preset recursive fractal partitioning mechanism and the target multimodal fusion features, and to generate a target feature tensor graph based on a cross-scale fusion attention mechanism. The target enhancement feature generation module is used to generate corresponding target enhancement features based on the target space attention map and the target feature tensor map through residual skip connections and channel attention mechanisms. The target detection result determination module is used to construct a corresponding target feature representation space based on manifold learning methods, Lie group operations, the target enhancement features, and the target spatial attention graph, and to decode the target feature representation space to obtain the corresponding target detection result.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, which is loaded and executed by the processor to implement the target detection method based on bionic vision as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the target detection method based on bionic vision as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cluster fixed-wing aircraft trajectory estimation method based on time sequence similar feature information

    CN116152295A

  • Optical communication filter appearance defect detection method and related equipment

    CN120177498A