Biological recognition method and system based on monocular depth guided multi-modal fusion
Through the multimodal fusion method guided by monocular depth, the problem of difficulty in biological camouflage recognition is solved, and efficient and accurate biological category recognition is achieved.
Patent Information
- Application Number
- CN202510788636.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
In nature, biological camouflage phenomenon leads to difficulties in biometric identification, and the prior art is difficult to effectively improve the recognition efficiency of camouflage creatures.
A multimodal fusion method based on monocular depth guidance is adopted to carry out depth estimation through the monocular depth model, multi-level features are extracted, sampling position prediction and fusion are performed, direction-specific features are extracted and enhanced, and finally pixel-level prediction is performed to achieve biological category recognition.
It improves the recognition accuracy and efficiency of camouflage organisms, adapts to biological needs of different shapes, and enhances the accuracy of detailed feature extraction and recognition.
Smart Images

Figure CN120299100A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and particularly to a biometric method and system based on monocular depth-guided multimodal fusion. Background Art
[0002] In lightweight deployment scenarios such as biodiversity monitoring and smart agriculture, identifying the organisms and their corresponding categories in the environment helps with further protection or governance. However, in nature, there are often phenomena of biological camouflage, such as camouflaged animals and hidden plants, which pose great difficulties for biological identification. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a biometric method and system based on monocular depth-guided multimodal fusion to solve the problem of how to improve the recognition efficiency of camouflaged organisms. The specific technical solutions are as follows: In the first aspect of the embodiments of the present application, a biometric method based on monocular depth-guided multimodal fusion is first provided. The method includes: Obtain an image of the organism to be recognized, where the organism to be recognized is an animal or a plant; Perform depth estimation on the image of the organism to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; Extract features of multiple resolutions from the pseudo-depth map and the image of the organism to be recognized to obtain multi-level features, where in the multi-level features, different-level features correspond to different resolutions; For each level of features, predict the sampling position through a convolutional layer, and perform sampling and fusion according to the predicted sampling position to obtain multi-scale features; extract and enhance direction-specific features through the image features and the depth features to obtain enhanced features; Perform upsampling and boundary feature enhancement according to the multi-scale features and the enhanced features to obtain decoded features; Perform pixel-level prediction according to the decoded features to obtain the category of the organism to be recognized.
[0004] In a possible implementation, the step of, for each level of features, predicting the sampling position through a convolutional layer and performing sampling and fusion according to the predicted sampling position to obtain multi-scale features includes: For each level of features, predict dynamic sampling offsets through a convolutional layer; adjust the sampling position according to the dynamic sampling offsets, and perform sampling according to the adjusted sampling position to obtain multiple sampled features; For each sampling feature, perform global average pooling and multi-layer perception to obtain corresponding attention weights; Modify according to the multiple sampling features and corresponding attention weights to obtain multiple modified sampling features; Fuse the multiple modified sampling features to obtain the multi-scale features.
[0005] In a possible implementation manner, the extracting of direction-specific features and feature enhancement through the image features and the depth features to obtain enhanced features includes: Extract direction-specific features from the image features and the depth features; Calculate attention weights according to the extracted direction-specific features; According to the calculated attention weights, the image features and the depth features, Calculate to obtain the enhanced features.
[0006] In a possible implementation manner, the upsampling and boundary feature enhancement according to the multi-scale features and the enhanced features to obtain decoded features includes: Through the first-level decoder in multiple decoders, perform upsampling on the multi-scale features and the enhanced features to obtain an initial prediction result; calculate a first gating feature according to the initial prediction result; calculate an output result of the first-level decoder according to the initial prediction result and the first gating feature; Through the remaining levels of decoders in multiple decoders, perform upsampling according to the output result of the previous level to obtain a prediction result of this level, where the remaining levels of decoders in the multiple decoders are the remaining levels of decoders except the first-level decoder; calculate a gating feature of this level according to the prediction result of this level; calculate an output result of this level of decoder according to the prediction result of this level and the gating feature of this level, and use the output result of the last level as the decoded features.
[0007] In a possible implementation manner, the pixel-level prediction according to the decoded features to obtain the category of the to-be-identified organism includes: Through a convolutional layer and a classifier, calculate the probability of each pixel value corresponding to a preset category according to the decoded features; Calculate the category of the to-be-identified organism according to the probability of each pixel value corresponding to the preset category.
[0008] In the second aspect of the embodiments of the present application, a biometric recognition system based on monocular depth-guided multi-modal fusion is provided. The system includes: An image acquisition module for acquiring an image of a biological object to be recognized, where the biological object to be recognized is an animal or a plant; A depth estimation module for estimating the depth of the image of the biological object to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; A feature extraction module for extracting features of multiple resolutions from the pseudo-depth map and the image of the biological object to be recognized to obtain multi-level features, where among the multi-level features, different-level features correspond to different resolutions; A position prediction module for respectively predicting the sampling positions through a convolutional layer for each level of features, and sampling and fusing according to the predicted sampling positions to obtain multi-scale features; extracting and enhancing direction-specific features through the image features and the depth features to obtain enhanced features; A feature decoding module for performing upsampling and enhancing boundary features according to the multi-scale features and the enhanced features to obtain decoded features; A biological classification module for performing pixel-level prediction according to the decoded features to obtain the category of the biological object to be recognized.
[0009] In a possible implementation manner, the position prediction module is specifically configured to respectively predict dynamic sampling offsets through a convolutional layer for each level of features; adjust the sampling positions according to the dynamic sampling offsets, and sample according to the adjusted sampling positions to obtain a plurality of sampled features; respectively perform global average pooling and multi-layer perception on each sampled feature to obtain corresponding attention weights; correct according to the plurality of sampled features and the corresponding attention weights to obtain a plurality of corrected sampled features; fuse the plurality of corrected sampled features to obtain the multi-scale features.
[0010] In a possible implementation manner, the position prediction module is specifically configured to extract direction-specific features from the image features and the depth features; calculate attention weights according to the extracted direction-specific features; calculate the enhanced features according to the calculated attention weights, the image features, and the depth features.
[0011] In a possible implementation manner, the feature decoding module is specifically configured to perform upsampling on the multi-scale features and the enhanced features through the first-level decoder among the multiple decoders to obtain an initial prediction result; calculate a first gating feature according to the initial prediction result; calculate an output result of the first-level decoder according to the initial prediction result and the first gating feature; perform upsampling on the output result of the previous level through the remaining levels of decoders among the multiple decoders to obtain a prediction result of this level, where the remaining levels of decoders among the multiple decoders are the remaining levels of decoders except the first-level decoder; calculate a gating feature of this level according to the prediction result of this level; calculate an output result of the decoder of this level according to the prediction result of this level and the gating feature of this level, and use the output result of the last level as the decoded feature.
[0012] In a possible implementation manner, the biological classification module is specifically configured to calculate the probability of each pixel value corresponding to a preset category according to the decoded feature through a convolutional layer and a classifier; calculate the category of the biological to be recognized according to the probability of each pixel value corresponding to the preset category.
[0013] On the other hand, an embodiment of the present application further provides an electronic device, including: A memory for storing a computer program; A processor, when executing the program stored on the memory, implements any of the above-mentioned biometric recognition methods based on monocular depth-guided multi-modal fusion.
[0014] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any of the above-mentioned biometric recognition segmentation methods based on monocular depth-guided multi-modal fusion is implemented.
[0015] On the other hand, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the above-mentioned biometric recognition methods based on monocular depth-guided multi-modal fusion.
[0016] Advantages of the embodiments of the present application: The embodiments of the present application provide a biometric recognition method and system based on monocular depth-guided multi-modal fusion. The method includes: obtaining an image of a biological object to be recognized, where the biological object to be recognized is an animal or a plant; performing depth estimation on the image of the biological object to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; extracting features of multiple resolutions from the pseudo-depth map and the image of the biological object to be recognized to obtain multi-level features, where among the multi-level features, different-level features correspond to different resolutions; respectively for each level of features, predicting the sampling positions through a convolutional layer, and performing sampling and fusion according to the predicted sampling positions to obtain multi-scale features; extracting direction-specific features and enhancing features through the image features and the depth features to obtain enhanced features; performing upsampling and enhancing boundary features according to the multi-scale features and the enhanced features to obtain decoded features; performing pixel-level prediction according to the decoded features to obtain the category of the biological object to be recognized. Through the solution of the embodiments of the present application, not only can the depth of the image be estimated and multi-level features be extracted, but also the prediction and sampling of the sampling positions can be performed to meet the needs of biological objects of different shapes, improving the accuracy of sampling. Moreover, direction-specific features can be extracted and features can be enhanced to achieve the extraction of more detailed features, thereby improving the accuracy and efficiency of biological category prediction through pixel-level prediction.
[0017] Of course, implementing any product or method of the present application does not necessarily require achieving all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.
[0019] Figure 1 It is a schematic flowchart of a biometric recognition method based on monocular depth-guided multi-modal fusion provided by the embodiments of the present application; Figure 2 It is a schematic flowchart of obtaining multi-scale features provided by the embodiments of the present application; Figure 3 It is a schematic flowchart of deformable convolution provided by the embodiments of the present application; Figure 4 It is a schematic flowchart of obtaining enhanced features provided by the embodiments of the present application; Figure 5 It is an example diagram of a biometric recognition method based on monocular depth-guided multi-modal fusion provided by the embodiments of the present application; Figure 6 This is a schematic structural diagram of a biometric device based on monocular depth-guided multi-modal fusion provided by an embodiment of the present application; Figure 7 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0021] In the first aspect of the embodiments of the present application, first, a biometric recognition method based on monocular depth-guided multi-modal fusion is provided. Refer to Figure 1 , Figure 1 This is a schematic flowchart of a biometric recognition method based on monocular depth-guided multi-modal fusion provided by an embodiment of the present application. The method includes: Step S11: Obtain an image of a biological object to be recognized, where the biological object to be recognized is an animal or a plant; Step S12: Perform depth estimation on the image of the biological object to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; Step S13: Extract features of multiple resolutions from the pseudo-depth map and the image of the biological object to be recognized to obtain multi-level features, where in the multi-level features, different-level features correspond to different resolutions; Step S14: For each level of features, predict the sampling positions through a convolutional layer, and perform sampling and fusion according to the predicted sampling positions to obtain multi-scale features; extract direction-specific features and perform feature enhancement through the image features and depth features to obtain enhanced features; Step S15: Perform upsampling and boundary feature enhancement according to the multi-scale features and the enhanced features to obtain decoded features; Step S16: Perform pixel-level prediction according to the decoded features to obtain the category of the biological object to be recognized.
[0022] Corresponding to the above step S11, the biological object to be recognized in the embodiments of the present application can be either an animal or a plant. And to meet the actual needs, the solution of the present application can be applied to the recognition of camouflaged organisms. Then the biological object to be recognized can be an animal disguised as a plant or other objects, etc. In the actual use process, when acquiring the image of the biological object to be recognized, an image in a specified format and rule can be acquired. For example, the corresponding resolution can be preset in advance. In one example, to acquire the image of the biological object to be recognized, a single RGB (Red, Green, Blue) image can be acquired. The format of this image is: 3 channels, and the resolution is H (height) × W (width) × 3.
[0023] Corresponding to the above step S12, when performing depth estimation on the image of the biological object to be recognized through a monocular depth model, this monocular depth model can predict the distance from each pixel point in the scene to the camera through a single RGB image, and obtain and output a pseudo-depth map. Among them, the pseudo-depth map includes image features and depth features. In one example, when performing depth estimation on the image of the biological object to be recognized through a monocular depth model, a pre-trained MDE (Monocular Depth Estimation Model) can be called to generate a pseudo-depth map (single channel, H×W×1). This pseudo-depth map can include multi-modal input tensors. Specifically, it can include: image features corresponding to red, green, and blue + Depth (depth features corresponding to spatial distance) bimodal data.
[0024] Corresponding to the above step S13, when extracting features of multiple resolutions from the pseudo-depth map and the image of the biological object to be recognized, PVTv2 (Pyramid Vision Transformer v2, a pyramid structure model) can be sampled to extract features of multiple resolutions. The features of multiple resolutions extracted can include both detailed features and semantic features. Among them, in the multi-level features, different-level features correspond to different resolutions. In one example, 4-level features {F1, F2, F3, F4} can be extracted. Among them, F1 is high-resolution detailed features (in one example, the size of F1 is H / 2×W / 2×C1, C1 = 64), retaining low-level information such as edges and textures; F4 is low-resolution semantic features (in one example, the size of F4 is H / 16×W / 16×C4, C4 = 512), capturing the overall contour of the target.
[0025] Corresponding to the above step S14, the sampling position is predicted through a convolutional layer, and sampling and fusion are performed according to the predicted sampling position to obtain multi-scale features. The dynamic sampling offset can be predicted through the convolutional layer, so that the sampling position can be adjusted according to the predicted dynamic sampling offset, and then sampling is performed through the adjusted sampling position. Finally, the sampling results of each layer are fused to obtain multi-scale features. In the actual use process, due to the large differences in the body shapes of different organisms, such as being either slender or oval, therefore, by adjusting the sampling position, it is possible to adapt to the body shapes of different organisms, thereby improving the accuracy and efficiency of sampling features. Through image features and depth features, the extraction and enhancement of direction-specific features can be carried out. Direction-specific features can be extracted by performing convolution on image features and depth features. Finally, the prediction of attention weights is carried out according to the extracted direction-specific features, and feature enhancement is performed according to the predicted weights to obtain enhanced features. In the actual use process, the direction-specific features represent the direction of animal fur texture, the depth gradient direction of plant leaves, etc. Therefore, by extracting the direction-specific features, the recognition efficiency of organisms can be improved.
[0026] Corresponding to the above step S15, according to the multi-scale features and enhanced features, upsampling and the enhancement of boundary features are carried out. Upsampling can be performed through multi-level cascaded decoding, and during the upsampling process, the GGA (Gate Guided Attention) mechanism is introduced to enhance the attention to the object boundary and details, thereby obtaining decoded features.
[0027] Corresponding to the above step S16, pixel-level prediction is performed according to the decoded features. Fine-grained modeling and inference can be carried out for each pixel in the image to generate a probability map, where each pixel value represents the probability of belonging to the target category. Thus, the category of the organism to be recognized can be obtained through this probability map.
[0028] It can be seen that through the solution of the embodiment of the present application, not only can the depth of the image be estimated and multi-level features be extracted, but also the needs of organisms with different shapes can be met by predicting and sampling the sampling position, improving the accuracy of sampling. Moreover, the extraction and enhancement of direction-specific features can be carried out to realize the extraction of more detailed features, thereby improving the accuracy and efficiency of biological category prediction through pixel-level prediction.
[0029] In a possible implementation manner, referring to Figure 2 , for each level of features respectively, the sampling position is predicted through a convolutional layer, and sampling and fusion are performed according to the predicted sampling position to obtain multi-scale features, including: Step S21: For each hierarchical feature, predict the dynamic sampling offset through a convolutional layer; adjust the sampling position according to the dynamic sampling offset, and sample according to the adjusted sampling position to obtain multiple sampled features; Step S22: For each sampled feature, perform global average pooling and multi-layer perception to obtain the corresponding attention weight; Step S23: Modify according to the multiple sampled features and the corresponding attention weights to obtain multiple modified sampled features; Step S24: Fuse the multiple modified sampled features to obtain multi-scale features.
[0030] In this application, for each hierarchical feature, the dynamic sampling offset is predicted through a convolutional layer; the sampling position is adjusted according to the dynamic sampling offset, and sampling is performed according to the adjusted sampling position, which can be achieved by deformable convolutional dynamic sampling. Specifically, for each layer of feature F i , the dynamic sampling offset Δp(x) is predicted through a convolutional layer, and then the sampling position of the convolutional kernel is adjusted according to the offset to adapt to the target morphology (such as curved limbs, overlapping leaves): , where p(k) is the standard sampling point; Δp(x) is the dynamic offset, predicted by Fi through 1×1 convolution; x is the pixel value of the input feature map F i , that is, the original feature tensor to be sampled; k is the index of the sampling point in the convolutional kernel, corresponding to the predefined position of the standard convolutional kernel; is the weight parameter of the convolutional kernel; y(x) is the sampling result. For each sampled feature, performing global average pooling and multi-layer perception to obtain the corresponding attention weight, and modification according to the multiple sampled features and the corresponding attention weights can be achieved through cross-hierarchical channel attention. Specifically, the channel attention weight α c can be generated through GAP (global average pooling) and MLP (multi-layer perception) to suppress the background channels and enhance the target features: α c =MLP(GAP(F i )), Fi′ = Fi × α c . Fusing the multiple modified sampled features to obtain multi-scale features can be achieved through multi-scale feature fusion. Specifically, {F1′, F2′, F3′, F4′} can be upsampled to the F1 resolution (H / 2×W / 2), and multi-scale feature F cup is generated through bottom-up fusion. In one example, the convolutional kernel in this application can be referred to Figure 3, including an input feature map; sampling region 1 of a deformable convolution kernel (wherein, the deformable convolution kernel can adaptively adjust the size of the convolution kernel); the output of sampling region 1 of the deformable convolution kernel; and sampling region 2 of the deformable convolution kernel (wherein, the deformable convolution kernel can also adaptively adjust the size of the convolution kernel); the output of sampling region 2 of the deformable convolution kernel; finally, obtaining the output of the deformable convolution kernel. In the traditional solution, the convolution kernel sampling is generally fixed.
[0031] In a possible implementation, referring to Figure 4 , through image features and depth features, extracting direction-specific features and enhancing features to obtain enhanced features, including: Step S41: Extracting direction-specific features from the image features and depth features; Step S42: Calculating attention weights according to the extracted direction-specific features; Step S43: Calculating enhanced features according to the calculated attention weights, image features, and depth features.
[0032] Extracting direction-specific features from the image features and depth features can be achieved through direction-sensitive grouped convolution. Specifically, 8 groups of 3×3 convolutions can be performed on the RGB and depth features respectively to extract direction-specific features (such as the direction of animal fur texture, the depth gradient direction of plant leaves), and the output direction-specific features include: F rgb-dir , F d-dir , where F rgb-dir is the animal fur texture direction feature, and F d-dir is the plant leaf depth gradient direction feature. Calculating attention weights according to the extracted direction-specific features can be achieved through a channel attention mask. Specifically, an attention weight β can be generated through a compression-expansion structure to amplify the target channel feature: β = Sigmoid(MLP(AvgPool(F rgb-dir ⊕F d-dir) ))); F fused = β ⊙ (F rgb-dir + F d-dir ); where AvgPool represents average pooling; Sigmoid represents an activation function; and MLP represents a multi-layer perceptron. Calculating enhanced features according to the calculated attention weights, image features, and depth features can be achieved through detail hallucination recovery. Specifically, for the occluded region, high-frequency edge features F hf can be extracted through 3×3 convolution, and details can be restored by combining a reverse mask : F enhanced = F fused +(M inv ⊙ Fhf ) , where F fused represents the fused feature; F enhanced represents the enhanced feature; F hf represents the high-frequency edge feature; M inv represents the inverse mask.
[0033] In a possible implementation, based on the multi-scale feature and the enhanced feature, upsampling and enhancement of the boundary feature are performed to obtain the decoded feature, including: through the first-level decoder in multiple decoders, upsampling the multi-scale feature and the enhanced feature to obtain an initial prediction result; calculating a first gating feature according to the initial prediction result; calculating an output result of the first-level decoder according to the initial prediction result and the first gating feature; through the remaining levels of decoders in multiple decoders, upsampling according to the output result of the previous level to obtain the prediction result of this level, where the remaining levels of decoders in multiple decoders are the remaining levels of decoders except the first-level decoder; calculating the gating feature of this level according to the prediction result of this level; calculating the output result of the decoder of this level according to the prediction result of this level and the gating feature of this level, and taking the output result of the last level as the decoded feature.
[0034] Specifically, multi-level feature fusion can be performed, such as upsampling (bilinear interpolation, factor = 2) step by step through a 4-level cascaded decoder, and each level of decoder fuses the feature from the encoder and the decoded feature of the previous level. And a GGA (Gate Guided Attention) module is introduced in each level of decoding, and the feature response is adaptively adjusted through the gating feature F gate to enhance the attention to the target boundary and details. For example: in the first-level decoding, the GGA module generates a gating signal based on the initial prediction result to optimize the expression of the low-level feature; in the subsequent cascades, the GGA module combines the intermediate prediction result of upsampling to iteratively refine the feature representation. Each level of decoder can retain the original feature information through a residual block and refine the target boundary by combining the gating attention mechanism to suppress background noise.
[0035] In a possible implementation, pixel-level prediction is performed according to the decoded feature to obtain the category of the biological to be recognized, including: through a convolutional layer and a classifier, calculating the probability corresponding to each pixel value for a preset category according to the decoded feature; calculating the category of the biological to be recognized according to the probability corresponding to each pixel value for the preset category. Specifically, finally, the output convolutional layer and the classifier generate a probability map P ∈ [0, 1]H×W, where each pixel value represents the probability of belonging to the target category. Then, the category of the biological to be recognized is determined through this probability map. Specifically, the category of the biological to be recognized can be determined through the maximum probability method, binary classification method, etc. through this probability map.
[0036] In an example, see Figure 5, after obtaining the RGB image, a pseudo-depth map is generated through monocular depth estimation, and then feature extraction is performed through a backbone network. The backbone network includes a pyramid structure composed of F1, F2, F3, and F4. After multi-scale integration through the CUP module, image features and depth features are obtained. Then, using the fusion and detail hallucination mechanism, the fused features X1, X2, X3, and X4 are obtained after processing. For X1, X2-feed is obtained through decoding, and then through upsampling, X2-gate is obtained; X2 and X2-gate are fused to obtain X3-feed; for X3-feed, through upsampling, X3-gate is obtained; X3 and X3-gate are fused to obtain X4-feed; for X4-feed, through upsampling, X4-gate is obtained; X4 and X4-gate are fused, and through preliminary coarse prediction, then convolution, bilinear interpolation, and further fine prediction are performed to generate a pixel-level camouflage prediction map. The input data in the embodiments of the present application is RGB+Depth (image features corresponding to red, green, and blue + depth features corresponding to spatial distance). This input data can fuse two-dimensional vision and three-dimensional geometric information, replace traditional hardware input, generate Depth through monocular estimation, overcome the dependence on dedicated devices, and reduce costs. The backbone network can be a { } pyramid structure. This backbone network can achieve multi-resolution feature extraction, and this backbone network can be a hierarchical Transformer architecture, which has a better detection effect on small target features than traditional CNN (Convolutional Neural Networks). The output of the CUP (Collaborative Unified Pyramid) module can be multi-scale fused feature F cup , which can dynamically adapt to the target morphology (deformable convolution + channel attention), and the number of deformable sampling points is better than that of traditional fixed convolution. The output of the ODEM (Occlusion-Aware Detail Enhancement and Fusion Module) module can be cross-modal enhanced feature F enhanced , which can separate the RGB texture and Depth gradient features and restore the occlusion details. The direction-sensitive grouped convolution (8 groups) in it is more sensitive to direction-specific information than traditional single-group convolution and is more suitable for applications in scenarios such as animal camouflage and plant concealment. The output result can be a segmentation map P (single channel), which can achieve pixel-level target probability prediction. Among them, Sigmoid activation is used instead of threshold segmentation, which supports subsequent adaptive threshold adjustment and improves the segmentation accuracy. Among them, the data flow and hierarchical relationship can be: Input stage: RGB → Monocular depth estimation → Pseudo-depth map; Feature extraction: PVTv2 generates 4-level pyramid features ( , the resolution is halved step by step, and the semantic abstraction degree increases); Core processing: The CUP and ODEM modules process multi-level features sequentially, and the multi-scale fusion features output by the CUP are transmitted into the ODEM module for further cross-modal enhancement features; Decoding output: The cascaded decoder gradually fuses the multi-scale features output by the CUP and ODEM modules. Through bilinear upsampling (factor = 2) and cross-level feature aggregation, combined with the GGA gating attention module to dynamically screen key information, after optimizing the boundary details through the residual block, a pixel-level probability segmentation map is generated.
[0037] The solution of this application not only achieves a breakthrough in three-dimensional information fusion: In the technical solution, a pseudo-depth map is generated through a monocular depth estimation model (such as Depth Anything V2). This data structure innovation enables the system to obtain three-dimensional geometric clues in the scene. Compared with the traditional method that only relies on RGB, on the animal dataset, the structural metric of the present invention is improved from 0.856 to 0.880, and the mean absolute error is reduced from 0.065 to 0.039. The principle is that depth information can assist in distinguishing the spatial distance between the target and the background. Especially for the camouflaged animals with blurred contours, the edges are located through the depth gradient, resulting in a significant improvement in the segmentation accuracy. Moreover, it achieves cross-modal feature enhancement: The ODEM module realizes the efficient fusion of cross-modal features through a three-level architecture of direction decoupling-channel enhancement-dynamic compensation. The constructed hallucination generator extracts high-frequency details and dynamically injects them into the low-confidence regions through the confidence mask, reducing the missed detection rate in complex scenes such as camouflage. This module achieves a double breakthrough in cross-modal spatial alignment and occlusion detail restoration with a small increase in computational cost. And it can achieve dynamic morphological adaptation: The CUP module uses deformable convolution to dynamically adjust the sampling offset, combined with cross-level channel attention, to adaptively process the morphological differences of different species. In the scene of overlapping plant leaves, the contour integrity rate is improved. The principle is that the deformable convolution dynamically expands the receptive field according to the target morphology, and at the same time, the channel attention mechanism enhances the target feature channels and weakens the background noise channels.
[0038] In a specific example, the hardware in the inventor's training stage is as follows: GPU acceleration: Supports NVIDIA RTX3090 / 4090 graphics cards, and uses CUDA (Compute Unified Device Architecture) parallel computing to accelerate monocular depth estimation and feature extraction. Distributed training: Supports 8-card synchronous training, with a batch size of 32. The number of training rounds for plant concealment detection, plant camouflage detection, and animal camouflage detection are 30, 120, and 100 respectively, adjusted according to the dataset scale, modal complexity, and task characteristics. The video memory occupancy ≤ 20GB.
[0039] The inventors' research found that traditional methods are mostly designed for a single species (such as animals or plants), and the missed detection rate exceeds 25% in mixed scenarios. Through the dynamic morphological adaptation of the CUP module and the cross-modal feature fusion of the ODEM module, the present invention realizes the unified detection of animal and plant camouflage for the first time. In the mixed test set containing animals and plants, the detection performance is effectively improved. The core principle is that deformable convolution and direction-sensitive convolution can adapt to the morphological and texture differences of different species, and multi-modal data fusion provides a richer information dimension for cross-species feature modeling, enabling the model to have stronger generalization ability.
[0040] In the second aspect of the embodiments of the present application, a biometric recognition system based on monocular depth-guided multi-modal fusion is provided. Refer to Figure 6 , the system includes: An image acquisition module 601, configured to acquire an image of a biometric to be recognized, where the biometric to be recognized is an animal or a plant; A depth estimation module 602, configured to perform depth estimation on the image of the biometric to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; A feature extraction module 603, configured to extract features of multiple resolutions from the pseudo-depth map and the image of the biometric to be recognized to obtain multi-level features, where among the multi-level features, different-level features correspond to different resolutions; A position prediction module 604, configured to respectively for each level of features, predict the sampling position through a convolutional layer, and perform sampling and fusion according to the predicted sampling position to obtain multi-scale features; through the image features and depth features, extract direction-specific features and perform feature enhancement to obtain enhanced features; A feature decoding module 605, configured to perform upsampling and boundary feature enhancement according to the multi-scale features and the enhanced features to obtain decoded features; A biometric classification module 606, configured to perform pixel-level prediction according to the decoded features to obtain the category of the biometric to be recognized.
[0041] In a possible implementation manner, the position prediction module is specifically configured to respectively for each level of features, predict the dynamic sampling offset through a convolutional layer; adjust the sampling position according to the dynamic sampling offset, and perform sampling according to the adjusted sampling position to obtain multiple sampled features; respectively for each sampled feature, perform global average pooling and multi-layer perception to obtain the corresponding attention weights; correct according to the multiple sampled features and the corresponding attention weights to obtain multiple modified sampled features; fuse the multiple modified sampled features to obtain multi-scale features.
[0042] In a possible implementation, the position prediction module is specifically configured to extract direction-specific features from the image features and depth features; calculate attention weights based on the extracted direction-specific features; and calculate enhanced features based on the calculated attention weights, image features, and depth features.
[0043] In a possible implementation, the feature decoding module is specifically configured to upsample the multi-scale features and enhanced features through the first-level decoder in multiple decoders to obtain an initial prediction result; calculate a first gating feature based on the initial prediction result; calculate the output result of the first-level decoder based on the initial prediction result and the first gating feature; upsample according to the output result of the previous level through the remaining levels of decoders in multiple decoders to obtain the prediction result of this level, where the remaining levels of decoders in multiple decoders are the remaining levels of decoders except the first-level decoder; calculate the gating feature of this level based on the prediction result of this level; calculate the output result of the decoder of this level based on the prediction result of this level and the gating feature of this level, and use the output result of the last level as the decoded feature.
[0044] In a possible implementation, the biological classification module is specifically configured to calculate the probability of each pixel value corresponding to a preset category according to the decoded feature through a convolutional layer and a classifier; and calculate the category of the biological to be recognized according to the probability of each pixel value corresponding to the preset category.
[0045] It can be seen that through the solution of the embodiments of the present application, not only can the depth of the image be estimated and multi-level features be extracted, but also the prediction and sampling of the sampling position can be performed to adapt to the needs of organisms with different shapes, improving the accuracy of sampling, and the extraction of direction-specific features and feature enhancement can be performed to realize the extraction of more detailed features, thereby improving the accuracy and efficiency of biological category prediction through pixel-level prediction.
[0046] The embodiments of the present application also provide an electronic device, as Figure 7 shown, including: A memory 701 for storing a computer program; A processor 702, when executing the program stored on the memory 701, implements the following steps: Obtain an image of the biological to be recognized, where the biological to be recognized is an animal or a plant; Perform depth estimation on the image of the biological to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; Extract features of multiple resolutions from the pseudo-depth map and the image of the biological to be recognized to obtain multi-level features, where among the multi-level features, different-level features correspond to different resolutions; For each hierarchical feature, the sampling position is predicted through a convolutional layer, and sampling and fusion are performed according to the predicted sampling position to obtain multi-scale features; through image features and depth features, the extraction of direction-specific features and feature enhancement are carried out to obtain enhanced features; According to the multi-scale features and the enhanced features, upsampling and boundary feature enhancement are performed to obtain decoded features; Pixel-level prediction is performed according to the decoded features to obtain the category of the biological object to be recognized.
[0047] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0048] The communication interface is used for communication between the above electronic device and other devices.
[0049] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0050] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0051] In another embodiment provided by this application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above biological recognition methods based on monocular depth-guided multi-modal fusion are implemented.
[0052] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which, when running on a computer, causes the computer to execute any one of the biometric identification methods based on monocular depth-guided multimodal fusion in the above embodiments.
[0053] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a solid-state disk (SSD), etc.
[0054] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0055] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system, electronic device, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0056] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A biometric recognition method based on monocular depth-guided multimodal fusion, characterized in that, The method includes: Obtaining an image of a biological object to be recognized, where the biological object to be recognized is an animal or a plant; Performing depth estimation on the image of the biological object to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; Extracting features of multiple resolutions from the pseudo-depth map and the image of the biological object to be recognized to obtain multi-level features, where in the multi-level features, different-level features correspond to different resolutions; For each level of features respectively, predicting the sampling positions through a convolutional layer, and performing sampling and fusion according to the predicted sampling positions to obtain multi-scale features; extracting and enhancing direction-specific features through the image features and the depth features to obtain enhanced features; Performing upsampling and enhancing boundary features according to the multi-scale features and the enhanced features to obtain decoded features; Performing pixel-level prediction according to the decoded features to obtain the category of the biological object to be recognized.
2. The method according to claim 1, wherein The step of, for each level of features respectively, predicting the sampling positions through a convolutional layer, and performing sampling and fusion according to the predicted sampling positions to obtain multi-scale features includes: For each level of features respectively, predicting dynamic sampling offsets through a convolutional layer; adjusting the sampling positions according to the dynamic sampling offsets, and performing sampling according to the adjusted sampling positions to obtain multiple sampled features; Performing global average pooling and multi-layer perception on each of the sampled features respectively to obtain corresponding attention weights; Correcting according to the multiple sampled features and the corresponding attention weights to obtain multiple modified sampled features; Fusing the multiple modified sampled features to obtain the multi-scale features.
3. The method according to claim 1, wherein The step of extracting and enhancing direction-specific features through the image features and the depth features to obtain enhanced features includes: Extracting direction-specific features from the image features and the depth features; Calculating attention weights according to the extracted direction-specific features; According to the calculated attention weights, the image features and the depth features, Calculating to obtain the enhanced features.
4. The method according to claim 1, characterized in that, The step of performing upsampling and enhancing boundary features according to the multi-scale features and the enhanced features to obtain decoded features includes: Performing upsampling on the multi-scale features and the enhanced features through the first-level decoder among multiple decoders to obtain an initial prediction result; calculating a first gating feature according to the initial prediction result; calculating an output result of the first-level decoder according to the initial prediction result and the first gating feature; Performing upsampling through the remaining levels of decoders among multiple decoders according to the output result of the previous level to obtain a prediction result of this level, where the remaining levels of decoders among multiple decoders are the remaining levels of decoders except the first-level decoder; calculating a gating feature of this level according to the prediction result of this level; calculating an output result of the decoder of this level according to the prediction result of this level and the gating feature of this level, and taking the output result of the last level as the decoded features.
5. The method according to claim 1, wherein Performing pixel-level prediction according to the decoded features to obtain the category of the biological object to be recognized includes: Calculating the probabilities of the preset categories corresponding to each pixel value according to the decoded features through a convolutional layer and a classifier; Calculating the category of the biological object to be recognized according to the probabilities of the preset categories corresponding to each pixel value.
6. A biometric recognition system based on monocular depth-guided multimodal fusion, characterized in that, The system includes: An image acquisition module for acquiring an image of a biological object to be recognized, where the biological object to be recognized is an animal or a plant; A depth estimation module for performing depth estimation on the image of the biological object to be recognized through a monocular depth model to obtain a pseudo-depth map, where the pseudo-depth map includes image features and depth features; A feature extraction module for extracting features of multiple resolutions from the pseudo-depth map and the image of the biological object to be recognized to obtain multi-level features, where among the multi-level features, different-level features correspond to different resolutions; A position prediction module for, respectively for each level of features, predicting the sampling positions through a convolutional layer, and performing sampling and fusion according to the predicted sampling positions to obtain multi-scale features; extracting direction-specific features and enhancing features through the image features and the depth features to obtain enhanced features; A feature decoding module for performing upsampling and enhancing boundary features according to the multi-scale features and the enhanced features to obtain decoded features; A biological classification module for performing pixel-level prediction according to the decoded features to obtain the category of the biological object to be recognized.
7. The system according to claim 6, wherein The position prediction module is specifically configured to, respectively for each level of features, predict dynamic sampling offsets through a convolutional layer; adjust the sampling positions according to the dynamic sampling offsets, and perform sampling according to the adjusted sampling positions to obtain a plurality of sampled features; Performing global average pooling and multi-layer perception respectively for each sampled feature to obtain corresponding attention weights; correcting according to the plurality of sampled features and the corresponding attention weights to obtain a plurality of modified sampled features; Fusing the plurality of modified sampled features to obtain the multi-scale features.
8. The system according to claim 6, wherein The position prediction module is specifically configured to extract direction-specific features from the image features and the depth features; calculate attention weights according to the extracted direction-specific features; calculate the enhanced features according to the calculated attention weights, the image features, and the depth features.
9. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for, when executing the program stored on the memory, implementing the method according to any one of claims 1-5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Non-contact multi-modal fusion biological recognition system, method and device
CN119007252A
Online car-hailing passenger position rapid positioning method based on voice interaction and visual perspective
CN119399845A
Joint target detection and tracking method based on spatial-temporal feature aggregation of multi-sensor data
CN119693758A
Systems and methods for biometric identification
WO2012106728A1
Three-dimensional object detection method based on multi-modal fusion and deep attention mechanism
WO2024217115A1
Cited By
High-temperature steam ventilation supercavity interface oscillation analysis method combined with artificial intelligence model
CN120931604A