Visual identification method and system
Through multimodal data fusion and adaptive decision-making engine, multimodal feature tensors with space-time consistency are generated, which solves the robustness and accuracy problems of traditional visual recognition methods in complex environments, and achieves high robustness and high accuracy visual recognition effects.
Patent Information
- Application Number
- CN202510586295.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Traditional visual recognition methods are difficult to achieve high robustness and high accuracy target recognition in complex environments, and are severely affected by factors such as light changes, occlusion and environmental noise.
A multimodal data fusion network is used to generate a space-time and consistent multimodal heterogeneous feature tensor, combined with an adversarial generation network and an adaptive decision engine, and the confidence score is fusion through the self-correction module to output the final identification result.
It realizes high robustness and high accuracy visual recognition in complex environments, improves the system's adaptability to light changes and occlusion, and enhances the stability and accuracy of recognition.
Smart Images

Figure CN120495603A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual recognition technology, and in particular to a visual recognition method and system. Background Art
[0002] With the rapid development of applications such as industrial automation, intelligent monitoring, and autonomous driving, visual recognition technology is playing an increasingly important role in target detection, behavior analysis, and environmental perception. Traditional visual recognition methods, which primarily rely on single-spectrum imagery, struggle to cope with complex environmental interference and occlusion, resulting in reduced recognition accuracy. Furthermore, factors such as illumination variations, occlusion, and environmental noise significantly impact recognition performance, making achieving robust and accurate target recognition in real-world scenarios a challenge. Summary of the Invention
[0003] The purpose of the present invention is to provide a visual recognition method and system to address the deficiencies in the prior art and to achieve visual recognition with high robustness and high accuracy.
[0004] An embodiment of the present application provides a visual recognition method, the method comprising:
[0005] Obtain visible light images, infrared thermal images, and depth point cloud data of the target scene, perform cross-domain alignment on the multimodal data using a heterogeneous feature fusion network, and generate a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features from different spectral bands.
[0006] Inputting the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstructing a feature map through an adaptive high-frequency enhancement filter to generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity;
[0007] Based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occluded samples, and a recognition feature vector with enhanced adversarial robustness is generated in combination with a temporal consistency constraint, wherein the temporal consistency constraint is enforced to retain the continuity of temporal features through an optical flow field loss function;
[0008] Inputting the identified feature vector into an adaptive decision engine, dynamically adjusting the classification boundaries through a reinforcement learning-driven graph neural network to generate dynamic decision parameters that are adaptive to the environment, wherein the graph neural network automatically activates subgraph structures of different depths based on the complexity of the real-time scene;
[0009] A multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through a self-correction module to output the final recognition result, wherein the self-correction module adopts an asymmetric loss function to eliminate cross-layer feature conflicts.
[0010] Optionally, the visible light image, infrared thermal imaging, and depth point cloud data of the target scene are acquired, and cross-domain alignment of the multimodal data is performed through a heterogeneous feature fusion network to generate a spatiotemporally consistent multimodal heterogeneous feature tensor, wherein the heterogeneous feature fusion network adopts a self-attention mechanism to fuse edge response features of different spectral bands, including:
[0011] Based on the frame synchronization signal of the visible light image and the timestamp of the infrared thermal image, a dynamic interpolation compensation algorithm is used to align the sampling interval of the depth point cloud data to generate a trimodal data stream with strict time axis synchronization;
[0012] Perform multispectral edge detection on visible light images, use the temperature gradient of infrared thermal imaging to correct visible light edge breakage areas, and extract the curvature mutation features of the depth point cloud to generate a cross-modal edge response map;
[0013] The cross-modal edge response map is input into a dual-path self-attention network, and weights are dynamically assigned through a visible light-infrared feature cross-calibration module to generate a spectrally consistent edge feature matrix.
[0014] The fused edge feature matrix is subjected to three-dimensional spatiotemporal convolution, and combined with the spatial topological constraints of the depth point cloud, a multimodal heterogeneous feature tensor with spatiotemporal consistency is output.
[0015] Optionally, the multimodal heterogeneous feature tensor is input into a spatial frequency-aware optimizer, and a feature map is reconstructed through an adaptive high-frequency enhancement filter to generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the illumination intensity, including:
[0016] According to the brightness histogram distribution of the visible light image, a dynamic threshold segmentation algorithm is used to extract the scene illumination intensity level and generate a quantitative illumination intensity index.
[0017] Perform two-dimensional wavelet decomposition on the multimodal heterogeneous feature tensor to separate the high-frequency detail component and the low-frequency contour component, and dynamically adjust the convolution kernel weight ratio of the high-frequency detail component based on the light intensity quantization index;
[0018] Adaptive inverse wavelet transform is used to fuse the weighted high-frequency detail components and low-frequency contour components, and residual connection is used to compensate for the reconstruction error to generate an optimized feature matrix for detail enhancement.
[0019] The optimized feature matrix is subjected to local contrast analysis, and a nonlinear filtering algorithm is used to suppress artifact noise caused by sudden changes in illumination, and the final optimized feature matrix after noise suppression is output.
[0020] Optionally, based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occlusion samples, and a temporal consistency constraint is combined to generate a recognition feature vector with enhanced adversarial robustness, wherein the temporal consistency constraint is forced to retain the continuity of temporal features through an optical flow field loss function, including:
[0021] The optimized feature matrix is input into the morphological erosion generator, and the edge area of the feature map is eroded by random structural elements to synthesize a multi-scale occlusion sample set;
[0022] Based on the optimized feature matrix of adjacent frames, the dense optical flow estimation algorithm is used to extract the motion trajectory of the sample occluded area and generate the spatiotemporal continuity constraint vector;
[0023] Combining the spatiotemporal continuity constraint vector of the optical flow trajectory with the multi-scale occlusion sample set, we design a directional consistency loss function to force the adversarial generative network to preserve temporal motion characteristics when synthesizing samples.
[0024] The generator after adversarial training is used to perform feature perturbation enhancement on the optimized feature matrix that retains the temporal motion features, and the recognition feature vector with enhanced adversarial robustness is extracted through contrastive learning of the Siamese network.
[0025] Optionally, the identification feature vector is input into an adaptive decision engine, and the classification boundary is dynamically adjusted through a reinforcement learning-driven graph neural network to generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths according to the complexity of the real-time scene, including:
[0026] According to the spatial distribution density of the identified feature vectors, the node density clustering algorithm is used to calculate the real-time scene complexity index;
[0027] A deep routing strategy based on the scene complexity index triggers the graph neural network, automatically selecting shallow local subgraphs or deep global subgraphs for feature aggregation;
[0028] Dynamically adjust the classification decision boundaries of shallow local subgraphs or deep global subgraph nodes through a reinforcement learning reward mechanism, where the reward function is determined by both the classification confidence and the density of feature distribution;
[0029] The optimized classification decision boundary is Gaussian smoothed, and the dynamic decision parameter matrix for environment adaptation is output in combination with the scene complexity index.
[0030] Optionally, constructing a multi-scale verification pyramid based on the dynamic decision parameters, fusing the confidence scores of each level through a self-correction module, and outputting a final recognition result, wherein the self-correction module uses an asymmetric loss function to eliminate cross-layer feature conflicts, including:
[0031] According to the scale sensitivity of the dynamic decision parameter matrix, a stratified random sampling strategy is used to construct a multi-scale verification pyramid to generate a hierarchical confidence benchmark;
[0032] Perform bidirectional attention alignment on the confidence scores of each level of the pyramid, eliminate the score deviation between scales through feature similarity measurement, and generate a consistent calibrated confidence distribution;
[0033] Designing a directional constrained asymmetric loss function to perform differential weighted fusion of high-level semantic confidence and low-level detail confidence to suppress cross-layer feature conflicts;
[0034] The fused confidence distribution is input into a multi-level weighted voter, which dynamically selects the optimal level result based on the preset threshold, and outputs the final recognition result and a traceable decision path.
[0035] Optionally, the multi-scale verification pyramid is constructed using a stratified random sampling strategy based on the scale sensitivity of the dynamic decision parameter matrix to generate a hierarchical confidence benchmark, including:
[0036] According to the gradient distribution characteristics of the dynamic decision parameter matrix, the ripple diffusion algorithm is used to analyze the sensitivity of parameter changes at different scales and generate a three-dimensional distribution map containing the scale-sensitivity mapping relationship;
[0037] Based on the three-dimensional distribution map, the dynamic registration algorithm is used to extract the characteristic basis vectors at each scale, and the feature similarity measurement is used to filter the redundant basis vectors to generate a scale orthogonal basis vector set.
[0038] The scale-orthogonal basis vector set is input into the adaptive sliding window generator, and a pyramid hierarchical structure is constructed through the nonlinear relationship between window size and scale sensitivity. At the same time, a neighborhood smoothing algorithm is used to eliminate noise interference.
[0039] Cross-scale confidence propagation is performed on each pyramid level, and the level confidence benchmark and error tolerance threshold are generated by combining the statistical distribution characteristics of historical recognition results.
[0040] Another embodiment of the present application provides a visual recognition system, the system comprising:
[0041] An acquisition module is used to acquire visible light images, infrared thermal images, and depth point cloud data of the target scene, and to perform cross-domain alignment of the multimodal data using a heterogeneous feature fusion network to generate a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features from different spectral bands.
[0042] an enhancement module, configured to input the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstruct a feature map through an adaptive high-frequency enhancement filter, and generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity;
[0043] A generation module is configured to synthesize multi-scale occluded samples using a generative adversarial network based on the optimized feature matrix, and generate a recognition feature vector with enhanced adversarial robustness in combination with a temporal consistency constraint, wherein the temporal consistency constraint is enforced to preserve the continuity of temporal features through an optical flow field loss function;
[0044] An adjustment module, configured to input the identified feature vector into an adaptive decision engine, dynamically adjust the classification boundaries through a reinforcement learning-driven graph neural network, and generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths based on the complexity of the real-time scene;
[0045] The recognition module is used to construct a multi-scale verification pyramid based on the dynamic decision parameters, fuse the confidence scores of each level through a self-correction module, and output the final recognition result, wherein the self-correction module adopts an asymmetric loss function to eliminate cross-layer feature conflicts.
[0046] Yet another embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any of the above methods when run.
[0047] Yet another embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any of the above methods.
[0048] Compared with the existing technology, the visual recognition method provided by the present invention obtains visible light images, infrared thermal imaging and depth point cloud data of the target scene to generate a spatiotemporally consistent multimodal heterogeneous feature tensor; inputs the multimodal heterogeneous feature tensor into a spatial frequency perception optimizer to generate an optimized feature matrix with enhanced details; based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occlusion samples to generate a recognition feature vector with enhanced adversarial robustness; the recognition feature vector is input into an adaptive decision engine to generate dynamic decision parameters for environmental adaptation; a multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through a self-correction module to output the final recognition result, thereby achieving highly robust and accurate visual recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1A hardware structure block diagram of a computer terminal for a visual recognition method provided by an embodiment of the present invention;
[0050] Figure 2 A flowchart of a visual recognition method provided by an embodiment of the present invention;
[0051] Figure 3 A schematic diagram of the structure of a visual recognition system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention.
[0053] The embodiment of the present invention first provides a visual recognition method, which can be applied to electronic devices such as computer terminals, specifically ordinary computers.
[0054] The following describes it in detail by taking running on a computer terminal as an example. Figure 1 The hardware structure block diagram of a computer terminal for a visual recognition method provided by an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0055] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the visual recognition methods.
[0056] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0057] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any visual recognition method.
[0058] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0059] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0060] See also Figure 2 , an embodiment of the present invention provides a visual recognition method, which may include the following steps:
[0061] S201, obtaining visible light images, infrared thermal images, and depth point cloud data of a target scene, performing cross-domain alignment on the multimodal data through a heterogeneous feature fusion network, and generating a spatiotemporally consistent multimodal heterogeneous feature tensor, wherein the heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features of different spectral bands;
[0062] This step builds a 3D environmental perception system through collaborative multi-sensor acquisition, employing a self-attention mechanism to address inherent differences in spectral response and spatial resolution between modal data. A heterogeneous feature fusion network maps visible light texture, infrared thermal radiation, and deep geometric features into a unified high-dimensional tensor space through frequency domain transformation. This overcomes the data limitations of a single modality in traditional visual recognition. Cross-domain feature alignment significantly improves environmental perception under complex lighting conditions, providing standardized feature inputs that are consistent across time and space for subsequent processing.
[0063] S202: Inputting the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstructing a feature map through an adaptive high-frequency enhancement filter, and generating an optimized feature matrix for detail enhancement, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity;
[0064] This step innovatively combines image frequency domain analysis with deep learning, achieving feature enhancement through a dynamic filtering mechanism based on illumination perception. The adaptive high-frequency enhancement filter uses a differentiable wavelet transform kernel to adjust frequency band weights in real time based on local illumination conditions, enhancing high-frequency details while preserving low-frequency contours. This addresses the performance degradation of traditional feature extraction methods in scenarios with sudden changes in illumination. Through frequency domain optimization guided by physical priors, the expressiveness of key features such as edges and textures is significantly improved.
[0065] S203, based on the optimized feature matrix, using a generative adversarial network to synthesize multi-scale occlusion samples, combined with a temporal consistency constraint, to generate a recognition feature vector with enhanced adversarial robustness, wherein the temporal consistency constraint is enforced to preserve the continuity of temporal features through an optical flow field loss function;
[0066] This step constructs a dynamic defense system through generative adversarial learning, leveraging optical flow constraints to ensure the spatiotemporal plausibility of synthesized samples. A morphological erosion generator simulates realistic occlusion scenarios, while a temporal consistency loss function forces the model to learn motion continuity features, significantly improving the system's robustness against dynamic occlusion and adversarial attacks. Joint spatiotemporal optimization enables the model to continuously learn and adapt to complex scene changes.
[0067] S204: Input the identified feature vector into an adaptive decision engine, dynamically adjust the classification boundary through a reinforcement learning-driven graph neural network, and generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths according to the complexity of the real-time scene;
[0068] This step builds a dynamic decision-making architecture based on reinforcement learning. The graph neural network automatically switches between shallow local reasoning and deep global reasoning modes by sensing scene complexity. The reward mechanism combines classification confidence with the density of feature distribution to achieve optimal adjustment of the decision boundary. This overcomes the limitations of fixed-structure neural networks and enables the system to adaptively balance computational efficiency and recognition accuracy based on environmental complexity.
[0069] S205, constructing a multi-scale verification pyramid based on the dynamic decision parameters, fusing the confidence scores of each level through a self-correction module, and outputting a final recognition result, wherein the self-correction module uses an asymmetric loss function to eliminate cross-layer feature conflicts.
[0070] This step employs a hierarchical verification mechanism to address cross-scale feature conflicts. An asymmetric loss function differentiates high-level semantics from low-level detail features. A bidirectional attention mechanism enables cross-layer confidence calibration, and a multi-level voter provides an interpretable decision path, significantly reducing the risk of mismatches in multi-scale analysis. This traceable hierarchical verification process improves the system's decision reliability in open environments.
[0071] Specifically, the method acquires visible light images, infrared thermal images, and depth point cloud data of the target scene, performs cross-domain alignment on the multimodal data through a heterogeneous feature fusion network, and generates a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features of different spectral bands, including:
[0072] Based on the frame synchronization signal of the visible light image and the timestamp of the infrared thermal image, a dynamic interpolation compensation algorithm is used to align the sampling interval of the depth point cloud data to generate a trimodal data stream with strict time axis synchronization;
[0073] At the hardware level, raw data from visible light cameras (such as the Sony IMX477, 30fps), infrared thermal imagers (FLIR A35, 15fps), and depth sensors (Intel RealSense D455, 10Hz point cloud sampling rate) experience temporal misalignment due to sampling rate differences. The core of the dynamic interpolation compensation algorithm is to achieve multimodal synchronization through timestamp alignment and data reconstruction.
[0074] Timestamp alignment:
[0075] The frame synchronization signals of visible light and infrared are synchronized in microseconds through hardware trigger lines (such as GPIO), and the deviation is controlled within ±0.5ms.
[0076] The timestamps of the depth point cloud data are aligned with the host computer clock via NTP (Network Time Protocol) with an accuracy of ±2ms. For 10Hz depth data, linear interpolation is used to insert nine virtual point cloud frames every 100ms, increasing the equivalent sampling rate to 100Hz, matching the visible light / infrared frame rate. For example, when the real point cloud frames are at t = 100ms and 200ms, interpolated frames are generated at t = 110ms, 120ms, and so on.
[0077] Data reconstruction:
[0078] Visible light image: original resolution 1920×1080, using bicubic interpolation to maintain edge sharpness;
[0079] Infrared thermal imaging: resolution 640×480, temperature contrast enhanced by adaptive histogram equalization;
[0080] Deep point cloud: The interpolated point cloud uses KD-Tree spatial indexing to accelerate the nearest neighbor search and fill in missing areas (such as holes caused by occlusion).
[0081] The final trimodal data stream is strictly aligned based on the timestamp. Each frame of data contains:
[0082] Visible light RGB image (30fps);
[0083] Infrared temperature matrix (15fps interpolated to 30fps);
[0084] Depth point cloud (10Hz interpolated to 30fps, approximately 300,000 points per frame).
[0085] Perform multispectral edge detection on visible light images, use the temperature gradient of infrared thermal imaging to correct visible light edge breakage areas, and extract the curvature mutation features of the depth point cloud to generate a cross-modal edge response map;
[0086] Cross-modal edge detection improves the robustness of edge detection by fusing the complementary information of visible light, infrared and depth data.
[0087] Multispectral edge detection:
[0088] Visible light edges: An improved Canny algorithm (Gaussian kernel σ = 1.2, high and low threshold ratio 1:3) was used to extract edges with gradient intensity > 50;
[0089] Infrared correction: Calculate the temperature gradient of adjacent pixels. If the temperature difference is greater than 3°C, mark it as a potential edge. If the visible light edge is broken at the same location (e.g. due to reflection), fill it with the infrared edge.
[0090] Depth Curvature: Perform local surface fitting on the point cloud (5cm radius neighborhood) and calculate the average curvature. Regions with a sudden change in curvature > 0.05mm-1 are marked as geometric edges.
[0091] Edge fusion logic:
[0092] Logic and fusion: When visible light and infrared detect edges at the same location, visible light edge details are retained;
[0093] Logic or fusion: In low-light areas (visible light gradient <20), infrared or depth edge is preferred;
[0094] Conflict resolution: If the spatial position deviation of the three types of edges is greater than 5 pixels, a voting mechanism is initiated (if at least two types are consistent, the edge is retained).
[0095] The cross-modal edge response map is stored in a probabilistic form, with each pixel value representing the confidence level (0 to 1) of the edge. For example, in a reflective metal area, where the visible light edge is broken but the infrared temperature difference is significant, the final confidence level is revised from 0.3 to 0.85.
[0096] The cross-modal edge response map is input into a dual-path self-attention network, and weights are dynamically assigned through a visible light-infrared feature cross-calibration module to generate a spectrally consistent edge feature matrix.
[0097] The Dual-path Self-Attention Network (DSAN) consists of a visible light branch, an infrared branch, and a cross-calibration module. The structure is as follows:
[0098] Visible light branch:
[0099] Input: cross-modal edge response map (512×512×1);
[0100] Processing: 4 layers of convolution (kernel size 3×3, number of channels 16→32→64→128), ReLU activation, output high-level semantic feature map (64×64×128).
[0101] Infrared branch:
[0102] Input: infrared temperature matrix (upsampled to 512×512);
[0103] Processing: 4 layers of convolution with the same structure, outputting feature maps of the same size.
[0104] Cross calibration module:
[0105] Self-attention mechanism: Query, key, and value matrices are calculated for the visible light and infrared feature maps, with dimensions of 64×64×128. Attention weights are calculated by scaling the dot product attention (scaling factor 1 / √128);
[0106] Dynamic weight allocation: Softmax is used to generate visible light weight α and infrared weight β (α + β = 1). For example, β can reach 0.8 in low-light areas.
[0107] Feature fusion: fusion feature = α × visible light feature + β × infrared feature.
[0108] Output processing:
[0109] Deconvolution is performed on the fused features (kernel size 4×4, step size 2) to gradually restore them to the original resolution;
[0110] The final output is a spectrally consistent edge feature matrix (512×512×1), with an edge continuity error of <1 pixel.
[0111] The fused edge feature matrix is subjected to three-dimensional spatiotemporal convolution, and combined with the spatial topological constraints of the depth point cloud, a multimodal heterogeneous feature tensor with spatiotemporal consistency is output.
[0112] Three-dimensional spatiotemporal convolution is used to fuse temporal and spatial information, and the spatial constraints of the depth point cloud are combined to enhance feature consistency.
[0113] Spatiotemporal convolution structure:
[0114] Input: edge feature matrix sequence (30 frames / second, each frame 512×512×1);
[0115] Convolutional layer: 3D convolution kernel (3×3×3, 3 frames in the temporal dimension and 3×3 in the spatial dimension), with the number of channels expanded from 1 to 32;
[0116] Activation function: LeakyReLU (negative slope 0.1);
[0117] Pooling layer: max pooling (2×2×1), downsampling to 256×256×32.
[0118] Depth space constraints:
[0119] Point cloud projection: Project the depth point cloud onto the 2D image plane to generate a depth map (512×512);
[0120] Spatial mask: Apply weight attenuation (weight × 0.2) to areas with a depth greater than 5m (such as the background) to focus on near-field objects;
[0121] Topology optimization: Construct a point cloud topology network based on Delaunay triangulation, and constrain the convolution kernel to propagate features between neighboring nodes.
[0122] Output tensor:
[0123] Dimensions: 256 × 256 × 32 × 30 (width × height × channels × time);
[0124] Spatiotemporal consistency index: the difference in features between adjacent frames is <0.05 (cosine similarity).
[0125] For example, in a pedestrian detection scenario, this tensor can simultaneously represent the pedestrian's outline (visible light edge), body temperature characteristics (infrared), and three-dimensional position (depth), with the spatiotemporal offset error controlled within 3 pixels.
[0126] Specifically, the multimodal heterogeneous feature tensor is input into a spatial frequency perception optimizer, and a feature map is reconstructed through an adaptive high-frequency enhancement filter to generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity, including:
[0127] According to the brightness histogram distribution of the visible light image, a dynamic threshold segmentation algorithm is used to extract the scene illumination intensity level and generate a quantitative illumination intensity index.
[0128] The brightness histogram of a visible light image reflects the statistical characteristics of the scene's illumination distribution. First, the input image is grayscaled and its brightness histogram (scale 0 to 255) is calculated. The dynamic threshold segmentation algorithm uses a modified Otsu method to determine the optimal segmentation threshold by maximizing the inter-class variance. For example, in an indoor scene with uniform illumination, the histogram exhibits a unimodal distribution, and the segmentation threshold may be around 128. In contrast, in a high-contrast outdoor scene, the histogram may exhibit a bimodal distribution, and the segmentation threshold may be reduced to 64.
[0129] Generation of light intensity quantitative index:
[0130] Dynamic range division: divide the histogram into 3 intervals (low light: 0-85, medium light: 86-170, high light: 171-255);
[0131] Weight calculation: Count the pixel proportions in each interval, for example, the low-light area accounts for 15%, the medium-light area accounts for 60%, and the high-light area accounts for 25%;
[0132] Exponential mapping: The weights are mapped to a quantitative index between 0 and 1 through a piecewise linear function (e.g., a low light weight of 0.15 corresponds to an index of 0.3, a medium light weight of 0.6 corresponds to an index of 0.7, and a high light weight of 0.25 corresponds to an index of 0.9).
[0133] For example, in an indoor surveillance scenario, the quantization index is 0.7, triggering the conservative mode of subsequent high-frequency enhancement; in a backlit scene, the index may drop to 0.4, triggering the strong high-frequency compensation mode.
[0134] Perform two-dimensional wavelet decomposition on the multimodal heterogeneous feature tensor to separate the high-frequency detail component and the low-frequency contour component, and dynamically adjust the convolution kernel weight ratio of the high-frequency detail component based on the light intensity quantization index;
[0135] The multimodal heterogeneous feature tensor contains the fusion result of visible light, infrared and depth features (e.g., size 256×256×64). The two-dimensional wavelet decomposition uses the Daubechies 4 wavelet basis (DB4) and is decomposed into 4 subbands:
[0136] Low-frequency subband (LL): preserves contour information (size 128×128×64);
[0137] High frequency sub-bands (LH, HL, HH): contain horizontal, vertical and diagonal details (128×128×64 each).
[0138] Dynamic adjustment of convolution kernel weights:
[0139] Weight scaling function: define the weight scaling factor α = 1.5 × light quantization index (α = 1.05 when the index is 0.7);
[0140] High-frequency enhancement strategy: Apply 3×3 convolution kernels to the LH, HL, and HH subbands respectively, and scale the kernel weights proportionally. For example, the original kernel weight matrix is [[0.1, 0.2, 0.1], [0.2, 0.3, 0.2], [0.1, 0.2, 0.1]], which becomes [[0.105, 0.21, 0.105], [0.21, 0.315, 0.21], [0.105, 0.21, 0.105]] after adjustment.
[0141] Low-frequency suppression: A weight attenuation factor β = 0.8 is applied to the LL subband to suppress redundant contour information.
[0142] For example, in a low-light scene (exponent 0.3), α = 0.45, high-frequency components are only slightly enhanced to avoid noise amplification; while in a high-light scene (exponent 0.9), α = 1.35, detailed features are significantly enhanced.
[0143] Adaptive inverse wavelet transform is used to fuse the weighted high-frequency detail components and low-frequency contour components, and residual connection is used to compensate for the reconstruction error to generate an optimized feature matrix for detail enhancement.
[0144] Adaptive inverse wavelet transform uses a dual-channel reconstruction network:
[0145] Low-frequency channel: Perform bilinear interpolation upsampling (2 times) on the attenuated LL subband to restore it to 256×256 resolution;
[0146] High-frequency channel: Deconvolution layers (kernel 3×3, stride 2) are applied to the enhanced LH / HL / HH subbands to restore the detailed structure;
[0147] Feature fusion: The outputs of low-frequency and high-frequency channels are concatenated channel by channel (256×256×128) and compressed to 64 channels through 1×1 convolution.
[0148] Residual connection design:
[0149] Skip connection: adds the original feature tensor to the reconstruction result to compensate for the information loss in the wavelet transform;
[0150] Residual learning: Add a residual block (consisting of three 3×3 convolutional layers, each followed by a ReLU activation) to learn the compensation amount for the reconstruction error.
[0151] Example: In surface reconstruction of industrial parts with complex textures, residual connections improve edge sharpness by 15% and reduce artifacts by 30%.
[0152] The optimized feature matrix is subjected to local contrast analysis, and a nonlinear filtering algorithm is used to suppress artifact noise caused by sudden changes in illumination, and the final optimized feature matrix after noise suppression is output.
[0153] Local contrast analysis uses a sliding window (16×16 pixels) to traverse the feature matrix:
[0154] Contrast calculation: The standard deviation of pixels within the window is used as a contrast index (if the standard deviation is > 25, it is determined to be a high-contrast area);
[0155] Noise detection: Combined with the thermal distribution stability of infrared features (standard deviation <5°C is considered a stable area), it distinguishes true edges from noise.
[0156] Nonlinear filtering algorithm:
[0157] Bilateral filtering: The spatial domain kernel (σ_d = 3 pixels) and the grayscale domain kernel (σ_r = 10 gray levels) work together to preserve edges while smoothing homogeneous areas.
[0158] Adaptive median filtering: Dynamically adjusts the window size (3×3 to 7×7) to suppress salt and pepper noise;
[0159] Lighting balancing: When a local overexposed area is detected (e.g., quantization index > 0.8), the CLAHE algorithm (Contrast-Limited Adaptive Histogram Equalization, tile size 8×8, Clip Limit = 2.0) is applied to balance the brightness.
[0160] For example, in a backlit face image, the original feature matrix has overexposure artifacts in the cheekbone area. After filtering, the artifact intensity is reduced by 60% while retaining pupil details.
[0161] Specifically, based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occluded samples, and a temporal consistency constraint is combined to generate a recognition feature vector with enhanced adversarial robustness, wherein the temporal consistency constraint is forced to retain the continuity of temporal features through an optical flow field loss function, including:
[0162] The optimized feature matrix is input into the morphological erosion generator, and the edge area of the feature map is eroded by random structural elements to synthesize a multi-scale occlusion sample set;
[0163] The core goal of the Morphological Erosion Generator (MEG) is to enhance the model's anti-interference ability by simulating physical occlusions (such as raindrops, dust, and partial object occlusion). Its implementation is divided into three stages:
[0164] Random Structuring Element Generation: Structuring Element (SE) is the core of the corrosion operation. MEG adopts a dynamic random generation strategy:
[0165] Shape: circle (radius 3-15 pixels), rectangle (side length 5-20 pixels), or irregular polygon (number of vertices 3-6), randomly selected in a 1:1:1 ratio;
[0166] Scale: Dynamically adjusted according to the input feature map resolution. For example, for a 512×512 feature map, the maximum SE size is 32×32.
[0167] Density: Control the occlusion area coverage rate (5% to 30%) through Poisson distribution.
[0168] For example, in an autonomous driving scenario, the generator might create a circular SE with a radius of 8 pixels, covering 20% of the edge of the vehicle detection box to simulate raindrop occlusion.
[0169] Multi-scale corrosion operation: perform channel-by-channel corrosion on the optimized feature matrix (assuming the dimension is H×W×C, such as 256×256×64):
[0170] Corrosion kernel sliding: traverse the feature map with a step size of 4 pixels, apply SE to each position for min-filtering to weaken the edge response;
[0171] Multi-scale superposition: Apply SE of different scales (such as 5×5, 9×9, 13×13) to the same feature map in sequence to generate multi-level occlusion samples;
[0172] Channel weighting: Perform attention weighting on the corrosion results of different channels (such as channel 1 weight 0.8, channel 2 weight 0.5) to retain key features.
[0173] Enhanced sample diversity:
[0174] Dynamic occlusion: In a video sequence, the position of the occluded area in consecutive frames is interpolated to simulate motion occlusion (such as a bird flying across the camera);
[0175] Texture fusion: The eroded area is mixed with random noise (Gaussian noise σ = 0.1) or real occlusion texture (such as leaf pattern) to improve the authenticity of the sample.
[0176] For example, in the person re-identification task, when MEG generates occlusion samples, it applies a rectangular SE (15×10 pixels) to the upper body area of the pedestrian, covering the backpack area, and superimposes an irregular polygon SE on the lower body to simulate shadow occlusion, ultimately generating a sample set containing 10 occlusion patterns.
[0177] Based on the optimized feature matrix of adjacent frames, the dense optical flow estimation algorithm is used to extract the motion trajectory of the sample occluded area and generate the spatiotemporal continuity constraint vector;
[0178] The dense optical flow estimation algorithm is used to capture the temporal motion consistency of the occluded area, ensuring that the generated occlusion samples conform to the laws of physical motion:
[0179] Optical flow estimation model: PWC-Net (Pyramidal Warping and Cost Volume Network) is used. Its lightweight architecture (9.4M parameters) is suitable for real-time processing. The input is the optimized feature matrix of two adjacent frames (frame t and frame t+1), and the output is an H×W×2 optical flow field (horizontal and vertical displacement).
[0180] Extraction of motion trajectory of occluded area:
[0181] Forward-backward optical flow consistency check: Calculate the forward optical flow F_forward from t→t+1 and the backward optical flow F_backward from t+1→t. If ||F_forward+F_backward||>1 pixel, it is determined to be an occluded area.
[0182] Trajectory chain construction: For each pixel in the occluded area, a motion trajectory chain is generated through optical flow recursion. For example, the trajectory of a raindrop in 5 frames is {(x1, y1), (x2, y2), ..., (x5, y5)}, with a displacement of Δx = 3 pixels / frame and Δy = 0.5 pixels / frame.
[0183] Generation of space-time continuity constraint vector:
[0184] Motion consistency encoding: The displacement (Δx, Δy) of the trajectory chain is combined with the acceleration (Δ 2 x,Δ 2 y) is encoded as a 4-dimensional vector;
[0185] Time series smoothing filter: Kalman filtering is performed on the trajectory chain (process noise Q = 0.1, observation noise R = 1.0) to eliminate estimation jitter;
[0186] Region aggregation: DBSCAN clustering (eps=5, min_samples=3) is performed on adjacent trajectories (distance < 5 pixels) to generate region-level constraint vectors.
[0187] For example, in a surveillance video, a car is partially occluded, and optical flow estimation shows that an occluder (such as a flying bird) in the hood area is moving toward the upper right at a rate of 15 pixels per second. The constraint vector encodes this motion pattern, ensuring that subsequently generated occlusion samples are temporally continuous.
[0188] Combining the spatiotemporal continuity constraint vector of the optical flow trajectory with the multi-scale occlusion sample set, we design a directional consistency loss function to force the adversarial generative network to preserve temporal motion characteristics when synthesizing samples.
[0189] The Direction Consistency Loss (DCL) function improves the temporal robustness of adversarial training by constraining the motion direction of generated samples to be consistent with the true optical flow:
[0190] Loss function design:
[0191] Cosine similarity constraint: Calculate the cosine similarity between the optical flow direction (θ_gen) of the generated sample occlusion area and the real optical flow direction (θ_real). The loss term is 1-Similarity, forcing the direction to align.
[0192] Amplitude matching constraint: L1 loss L_magnitude is applied to the optical flow amplitude (||F||);
[0193] Total loss: DCL = 0.7 × (1-Similarity) + 0.3 × L_magnitude.
[0194] Adversarial training process:
[0195] Generator (G): Input is optimized feature matrix, output is occlusion sample;
[0196] Discriminator (D): The input is a real sample or a generated sample, and the output is the authenticity probability + optical flow consistency score;
[0197] Training strategy: Alternately optimize G and D, minimize DCL + adversarial loss (Wasserstein distance) in each round of G training, and maximize the difference between the real sample score and the generated sample score in D training.
[0198] Dynamic weight adjustment:
[0199] Initial stage (first 1000 rounds): DCL weight is 0.5, focusing on directional consistency;
[0200] Convergence stage (after 1000 rounds): The DCL weight is reduced to 0.3, focusing on generation quality.
[0201] For example, in a drone tracking task, the leaf occlusion samples synthesized by the generator may initially deviate from the target trajectory (with an error of 10 pixels). After DCL constraints, the error is reduced to within 2 pixels, and the occlusion's motion direction is consistent with the target.
[0202] The generator after adversarial training is used to perform feature perturbation enhancement on the optimized feature matrix that retains the temporal motion features, and the recognition feature vector with enhanced adversarial robustness is extracted through contrastive learning of the Siamese network.
[0203] Feature perturbation enhancement and Siamese network contrastive learning work together to improve the model's robustness to occlusion and dynamic interference:
[0204] Feature perturbation enhancement:
[0205] Occlusion injection: Use the trained generator to insert multi-scale occlusions into the optimized feature matrix, with the perturbation ratio controlled between 15% and 40%;
[0206] Channel random shielding: randomly shield some channels with a probability of 0.2 (for example, shield 13 channels out of 64 channels) to simulate sensor failure;
[0207] Spatiotemporal jittering: Randomly translate (±5 pixels) and rotate (±3°) the feature map to enhance spatial invariance.
[0208] Twin network architecture:
[0209] Backbone network: ResNet-50 (removing the last two fully connected layers) is used to output 1024-dimensional features;
[0210] Contrastive loss: Use InfoNCE (Noise Contrastive Estimation) loss with a temperature parameter τ = 0.07;
[0211] Positive and negative sample pair construction:
[0212] Positive samples: occluded samples and unoccluded samples of the same target;
[0213] Negative samples: feature vectors of different targets, or different occlusion patterns of the same target.
[0214] Training strategy:
[0215] Difficult sample mining: 256 samples are sampled in each batch, including 30% high-difficulty samples (occlusion area > 25%);
[0216] Progressive training: light occlusion (10% coverage) is used in the initial stage and gradually increased to 40%;
[0217] Feature distillation: The student network (lightweight MobileNetV3) is constrained by KL divergence to imitate the feature distribution of the teacher network (ResNet-50).
[0218] For example, in the face recognition task, the adversarially trained generator synthesizes occlusions such as glasses and masks. The Siamese network maps different occlusion states of the same person to neighboring areas in the feature space (cosine similarity > 0.85) through contrastive learning, and the feature distance between different people is > 0.5.
[0219] Specifically, the recognition feature vector is input into the adaptive decision engine, and the classification boundary is dynamically adjusted through the reinforcement learning-driven graph neural network to generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths according to the complexity of the real-time scene, including:
[0220] According to the spatial distribution density of the identified feature vectors, the node density clustering algorithm is used to calculate the real-time scene complexity index;
[0221] The recognition feature vector is a high-dimensional vector (e.g., 512-dimensional) output by the feature extraction network after adversarial robustness enhancement. Its spatial distribution density reflects the density and diversity of target objects in the scene. To quantify the complexity of the real-time scene, a node density clustering algorithm (such as DBSCAN) is used to cluster the feature vectors.
[0222] Parameter settings: neighborhood radius eps = 0.5 (Euclidean distance threshold), minimum number of samples min_samples = 5.
[0223] Density calculation: For each eigenvector, the number of samples in its neighborhood is calculated. For example, if there are 8 samples within a distance of 0.5 around a vector, it is marked as a core point; if there are only 3 samples, it is marked as a noise point.
[0224] Complexity index generation:
[0225] Number of clusters: If DBSCAN detects more than 5 clusters, it is marked as a high complexity scene (index > 0.7);
[0226] Noise point ratio: When the noise ratio exceeds 20% (e.g., there are 200 noise points in a total of 1000 samples), the index is increased by an additional 0.2;
[0227] Example: In an indoor monitoring scenario, 8 clusters are detected and the noise accounts for 15%, then the complexity index is 0.76.
[0228] Basic rule setting
[0229] Cluster number threshold: 5 clusters; noise ratio threshold: 20%; index interval division: low complexity: 0-0.5 (≤3 clusters); medium complexity: 0.5-0.7 (4-5 clusters); high complexity: >0.7 (≥6 clusters);
[0230] Example scenario calculation process
[0231] 8 clusters detected → Exceeding the threshold of 5 clusters → Base index = 0.7 (high complexity starting value);
[0232] Noise accounts for 15% → below the 20% threshold → no additional increase is triggered;
[0233] But 8 clusters are 3 more than the threshold of 5 → each additional cluster increases by 0.02 → 0.7 + (3 × 0.02) = 0.76;
[0234] Complete calculation formula
[0235] Complexity index = base value (determined by the number of clusters) + cluster number compensation (+0.02 for every 1 exceeding the threshold) + noise compensation (+0.05 for every 5% exceeding the threshold).
[0236] Design Principles
[0237] Each newly added cluster represents an additional independent moving target / area in the scene;
[0238] Noise reflects the intensity of environmental interference;
[0239] The segmented accumulation system is used to ensure that the index range is between 0 and 1.
[0240] This design takes into account both the number of targets (clusters) and environmental interference (noise), achieving refined complexity quantification through a linear compensation mechanism. In practice, these parameters can be adjusted and optimized based on specific scenarios using a validation set.
[0241] Dynamic calibration mechanism:
[0242] Light adaptation: In low-light conditions (such as infrared mode at night), the EPS is adjusted to 0.8 to expand the neighborhood range and avoid cluster splitting due to feature ambiguity;
[0243] Motion compensation: A time-decayed weight is applied to the feature vectors of dynamic targets (such as pedestrians), with the weight of the recent frame being 1.0 and the weight 10 frames ago being reduced to 0.3 to reflect the instantaneous complexity of the scene.
[0244] A deep routing strategy based on the scene complexity index triggers the graph neural network, automatically selecting shallow local subgraphs or deep global subgraphs for feature aggregation;
[0245] Graph Neural Network (GNN) uses a dynamic routing strategy to select subgraph structures based on the scene complexity index:
[0246] Shallow local subgraph: Suitable for low-complexity scenarios (index ≤ 0.5), only aggregates the target node and its 1-hop neighbor features. For example, in a simple face recognition scenario, only local features of facial features (such as eyes and nose nodes) are aggregated.
[0247] Example architecture: 2-layer GAT (Graph Attention Network), 4 attention heads per layer, and 256 output dimensions.
[0248] Deep global subgraph: This is used for high-complexity scenarios (index > 0.5), aggregating multi-hop neighbors (e.g., 3 hops) and cross-region correlation features. For example, in traffic monitoring, it is necessary to correlate multiple interactions among vehicles, pedestrians, traffic lights, and so on.
[0249] Example architecture: 5-layer GraphSAGE (number of sampled neighbors [10, 5, 3]), combined with Gated Recurrent Unit (GRU) to model temporal dependencies.
[0250] Routing trigger logic:
[0251] Threshold judgment: If the complexity index is > 0.5, activate the deep subgraph; otherwise use the shallow subgraph.
[0252] Hybrid mode: In the transition area (index 0.4-0.6), adaptive hybrid routing is used, for example, 70% of the traffic goes to the deep layer and 30% goes to the shallow layer.
[0253] Feature aggregation example:
[0254] Local aggregation: For a certain vehicle node, aggregate its direct component features such as tires and lights;
[0255] Global aggregation: further correlates the movement trends of other vehicles and pedestrians in the same lane.
[0256] Dynamically adjust the classification decision boundaries of shallow local subgraphs or deep global subgraph nodes through a reinforcement learning reward mechanism, where the reward function is determined by both the classification confidence and the density of feature distribution;
[0257] The reinforcement learning (RL) framework uses GNN nodes as agents and decision boundary adjustments (such as hyperplane shifts of SVM classifiers) as action spaces:
[0258] State space: node feature vector (512 dimensions), neighbor node feature mean, scene complexity index.
[0259] Action space: Classification boundary offset (-0.1 to +0.1, step size 0.02).
[0260] Reward function design: R = 0.6 × classification confidence + 0.4 × (1-feature distribution density).
[0261] Classification confidence: the maximum probability value output by Softmax (e.g., the probability that a node is classified as a “pedestrian” is 0.92);
[0262] Feature distribution density: the Euclidean distance variance of similar node features (e.g., the feature variance of 10 “vehicle” nodes is 0.3, and the density is 0.7).
[0263] Training process:
[0264] Policy network: uses the PPO (Proximal Policy Optimization) algorithm, learning rate lr = 3e-4, batch size batch_size = 64;
[0265] Exploration mechanism: ε-greedy strategy (ε = 0.1), 10% probability of randomly selecting an action;
[0266] Example decision: In a congested scenario, a "vehicle" node uses RL to adjust the classification boundary from 0.5 to 0.6, reducing the probability of false detection as an "obstacle."
[0267] Dynamic effect verification:
[0268] Low-complexity scenario: The decision boundary adjustment is small (±0.02) to maintain stability;
[0269] High-complexity scenes: Allows larger adjustments (±0.1) to cope with target occlusion and sudden appearance changes.
[0270] The optimized classification decision boundary is Gaussian smoothed, and the dynamic decision parameter matrix for environment adaptation is output in combination with the scene complexity index.
[0271] Gaussian smoothing is used to eliminate mutation noise on the decision boundary and improve model robustness:
[0272] Gaussian kernel parameters: kernel size 5×5, standard deviation σ=1.5, weight matrix generated according to two-dimensional Gaussian distribution;
[0273] Smoothing process: Convolution filtering is performed on the decision boundary parameter matrix. For example, the original value of a certain boundary is [0.5, 0.6, 0.55], and after smoothing it is [0.52, 0.58, 0.56];
[0274] Complexity Weighting: Dynamically adjust the smoothing strength based on the scene complexity index:
[0275] High complexity (index > 0.7): Enhance smoothing (σ = 2.0) to suppress overfitting;
[0276] Low complexity (exponent ≤ 0.3): weak smoothing (σ = 0.5), preserving details.
[0277] Dynamic decision parameter matrix generation:
[0278] Matrix structure: rows represent target categories (e.g. 10 categories), columns represent feature dimensions (512 dimensions), and element values are classification thresholds;
[0279] Example: In complex traffic scenarios, the threshold for the “pedestrian” category is reduced from 0.6 to 0.55 to improve detection recall.
[0280] Real-time deployment optimization:
[0281] Edge computing: Parameter matrices are quantized and compressed using TensorRT, enabling real-time inference on NVIDIA Jetson AGX (latency < 15ms).
[0282] Feedback mechanism: The classification accuracy is counted every 30 seconds. If it drops by more than 5%, model fine-tuning (such as updating Gaussian kernel parameters) is triggered.
[0283] Real-time decision making examples:
[0284] Input: Pedestrians, vehicles, and bicycles appear simultaneously in the surveillance footage;
[0285] Processing: Complexity index 0.75 → Activate deep subgraph → RL adjusts decision boundary → Gaussian smoothing parameter σ = 1.8;
[0286] Output: All objects are classified with >90% accuracy, and frame processing takes 22ms.
[0287] This method achieves high-precision and low-latency visual recognition in complex scenes by tightly coupling density clustering, dynamic routing, reinforcement learning and adaptive smoothing.
[0288] Specifically, a multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are integrated through a self-correction module to output a final recognition result. The self-correction module uses an asymmetric loss function to eliminate cross-layer feature conflicts, including:
[0289] According to the scale sensitivity of the dynamic decision parameter matrix, a stratified random sampling strategy is used to construct a multi-scale verification pyramid to generate a hierarchical confidence benchmark;
[0290] The scale sensitivity of the dynamic decision parameter matrix reflects the stability of the classification boundaries at different resolution levels. To construct a multi-scale verification pyramid, we first use the ripple diffusion algorithm to simulate the gradient distribution characteristics of the parameter matrix. This algorithm simulates the water wave diffusion process and marks sensitive areas at different scales in the parameter matrix:
[0291] Core parameters: diffusion radius (initial value 5 pixels, step size multiplied by scale level), attenuation coefficient (0.8 to 0.95);
[0292] Sensitivity calculation: The rate of change of parameters is recorded during the diffusion process. For example, in target detection scenarios, high-level layers (low resolution) are sensitive to object categories, while low-level layers (high resolution) are sensitive to edge details.
[0293] The dynamic registration algorithm is used to extract feature basis vectors:
[0294] Basis vector extraction: Perform principal component analysis (PCA) on the parameter matrix of each scale level, retain the first K principal components (e.g., K = 3), and generate orthogonal basis vectors;
[0295] Redundancy filtering: We remove highly similar basis vectors by using a cosine similarity threshold (e.g., 0.7) to ensure the independence of features at each scale. For example, in a traffic monitoring scenario, a high-level basis vector may correspond to the overall shape of a vehicle, while a low-level basis vector may correspond to the license plate texture.
[0296] The adaptive sliding window generator dynamically adjusts the window size based on scale sensitivity:
[0297] Window size rule: the higher the sensitivity, the smaller the window size (e.g., a sensitivity index of 0.9 corresponds to a 3×3 window, and 0.6 corresponds to a 7×7 window);
[0298] Neighborhood smoothing: A bilateral filtering algorithm (spatial standard deviation σ_s = 1.5, grayscale standard deviation σ_r = 0.1) is used to suppress noise while preserving edges.
[0299] Cross-scale confidence propagation combined with historical data to optimize the benchmark value:
[0300] Propagation mechanism: High-level confidence is used as a prior to guide low-level calibration. For example, if the high-level layer determines that a certain area is a "vehicle" (confidence 0.85), the low-level layer's confidence in the wheel details in the same area will be increased from 0.7 to 0.78;
[0301] Error tolerance threshold: Dynamically adjusted according to historical recognition accuracy (for example, when the historical accuracy is 90%, the threshold is set to 0.15).
[0302] The resulting pyramid contains four levels (e.g., 32×32, 64×64, 128×128, 256×256), and the confidence benchmark for each level is stored in the form of a probability distribution, for example:
[0303] Level 1 (32×32): vehicle category confidence 0.82, pedestrian 0.15;
[0304] Level 4 (256×256): License plate character confidence is 0.91, and vehicle body color confidence is 0.75.
[0305] Perform bidirectional attention alignment on the confidence scores of each level of the pyramid, eliminate the score deviation between scales through feature similarity measurement, and generate a consistent calibrated confidence distribution;
[0306] The bidirectional attention alignment module consists of two paths: top-down and bottom-up, fusing multi-scale information through a cross-attention mechanism:
[0307] Top-down approach: High-level semantic information guides the calibration of low-level details. For example, a region identified as a "truck" at the high level forces the low-level wheel detection confidence to be improved.
[0308] Bottom-up path: Low-level details correct high-level misjudgments. For example, if a low-level layer detects a "broken car window" but a high-level layer doesn't recognize it, it triggers a reassessment of the high-level layer's confidence.
[0309] The feature similarity measure is calculated by combining the improved KL divergence (Kullback-Leibler Divergence) and the cosine similarity:
[0310] KL Divergence: measures the difference in confidence distributions at different levels. For example, the KL value of the "vehicle" distribution at level 1 and the "truck" distribution at level 4 is 0.3. Calibration is triggered when the value exceeds the threshold of 0.2.
[0311] Cosine similarity: Evaluates the directional consistency of feature vectors. A threshold of 0.8 is set; values below this value are considered to be in conflict between scales.
[0312] Calibration process example (medical imaging scenario):
[0313] High level (low resolution): The lung CT slice is judged as "normal" as a whole, with a confidence level of 0.9;
[0314] Low level (high resolution): local nodule area confidence level 0.65;
[0315] Bidirectional attention alignment: The high-level confidence dropped to 0.7, while the low-level confidence increased to 0.8, triggering further inspection.
[0316] The consistency calibration algorithm is achieved through iterative optimization:
[0317] Initialization: The confidence of each level is loaded according to the original value;
[0318] Attention weighting: Top-Down weight 0.6, Bottom-Up weight 0.4;
[0319] Residual compensation: Gaussian blur smoothing (σ=1.0) is performed on areas with calibration deviation > 0.1.
[0320] Designing a directional constrained asymmetric loss function to perform differential weighted fusion of high-level semantic confidence and low-level detail confidence to suppress cross-layer feature conflicts;
[0321] The fused confidence distribution is input into a multi-level weighted voter, which dynamically selects the optimal level result based on the preset threshold, and outputs the final recognition result and a traceable decision path.
[0322] Asymmetric loss functions design differentiated penalties for different layer features:
[0323] High-level semantic loss: Focusing on category accuracy, Focal Loss (γ=2.0) is used to alleviate category imbalance;
[0324] Bottom-level detail loss: Focus on positioning accuracy, use IoU Loss (intersection over union optimization), and use a weighted coefficient of 0.7;
[0325] Cross-layer conflict penalty: When the confidence of the upper layer and the lower layer are in opposite directions (for example, the upper layer judges "true" and the lower layer judges "false"), an additional penalty (coefficient 0.5) is imposed.
[0326] Weighted fusion example (autonomous driving scenario):
[0327] High level (road type): confidence level 0.9 (highway);
[0328] Low-level (lane line): confidence 0.6 (dashed line);
[0329] Asymmetric weighting: high-level weight 0.7, low-level weight 0.3, and the overall confidence after fusion is 0.81.
[0330] The multi-level weighted voter uses a dynamic threshold mechanism:
[0331] Threshold preset: set according to the task type, such as 0.85 for safety-critical scenarios (medical, transportation), and 0.7 for general scenarios;
[0332] Hierarchical voting: Each level is sorted by confidence, and the top K (e.g., K=2) participate in the voting;
[0333] Decision path tracing: Record the contribution of each level, for example:
[0334] Level 3 (128×128): Contribution 45% (wheels detected);
[0335] Level 4 (256×256): Contribution 35% (license plate detected);
[0336] Level 1 (32×32): Contribution 20% (overall vehicle model matching).
[0337] Final output logic:
[0338] Confidence level meets the standard: directly output the highest score result;
[0339] If the confidence level is close (e.g., the difference is <0.05), multimodal verification is initiated (e.g., secondary verification of infrared features is called);
[0340] If the conflict cannot be resolved, it will be marked as "pending manual review" and all levels of evidence will be saved.
[0341] Application Case (Industrial Quality Inspection):
[0342] Input: Multi-scale image of metal part surface;
[0343] Pyramid levels:
[0344] Level 1: overall shape is acceptable (0.88);
[0345] Level 4: local microcracks (0.65);
[0346] Asymmetric fusion: crack confidence increased to 0.72 (weighting factor 0.6);
[0347] Voting result: Overall confidence level 0.78 < threshold 0.8 → marked as "suspicious", triggering a high-precision 3D scan review.
[0348] Specifically, according to the scale sensitivity of the dynamic decision parameter matrix, a stratified random sampling strategy is adopted to construct a multi-scale verification pyramid to generate a hierarchical confidence benchmark, including:
[0349] According to the gradient distribution characteristics of the dynamic decision parameter matrix, the ripple diffusion algorithm is used to analyze the sensitivity of parameter changes at different scales and generate a three-dimensional distribution map containing the scale-sensitivity mapping relationship;
[0350] The gradient distribution characteristics of the dynamic decision parameter matrix reflect the degree of influence of different scale features on the classification decision. The ripple diffusion algorithm quantifies the propagation range and intensity of parameter changes by simulating the physical process of water wave diffusion. The specific steps are as follows:
[0351] Gradient field construction: Calculate the Sobel gradient for the dynamic decision parameter matrix (e.g., a 512×512 matrix where each element represents the decision weight of a pixel) to generate the horizontal gradient Gx and vertical gradient Gy. The gradient amplitude calculation formula is √(Gx 2 +Gy 2 ), the threshold is set to 0.1 (gradients below this value are considered noise and do not participate in diffusion).
[0352] Ripple initialization: Generate an initial ripple source point at a location where the gradient magnitude is greater than a threshold. For example, if the gradient magnitude of a pixel is 0.3, a source point with a ripple energy value of 0.3 is initialized at that coordinate.
[0353] Diffusion process: Anisotropic diffusion model is used, ripple energy propagates along the gradient direction, the diffusion step is set to 3 pixels, and the attenuation coefficient is 0.8 (that is, the energy decays by 20% for each diffusion step). For example, after 3 steps of diffusion, the energy of a ripple with an initial energy of 0.3 decays to 0.3×0.8. 3 ≈0.153.
[0354] Scale sensitivity calculation: Calculate the cumulative ripple energy value at different scales (such as 1×1, 3×3, and 5×5 windows). For example, if the cumulative energy value in a 5×5 window is 2.5, the sensitivity is 2.5 / 25 = 0.1.
[0355] 3D distribution map generation: Scale (X-axis), spatial position (Y-axis, Z-axis), and sensitivity (color depth) are mapped to 3D volume data, and a Marching Cubes algorithm is used to generate a visual sensitivity distribution map. For example, the sensitivity at the edge of the target can reach 0.8, while the sensitivity in flat areas is only 0.05.
[0356] Technical example: In an industrial parts inspection scenario, the dynamic decision parameter gradient amplitude of a bolt edge is 0.4. The high-sensitivity area (>0.6) generated after ripple diffusion accurately covers the thread structure, while the background area sensitivity is <0.1, effectively distinguishing key features from noise.
[0357] Based on the three-dimensional distribution map, the dynamic registration algorithm is used to extract the characteristic basis vectors at each scale, and the feature similarity measurement is used to filter the redundant basis vectors to generate a scale orthogonal basis vector set.
[0358] The dynamic registration algorithm aims to extract the most representative feature basis vectors from the 3D sensitivity distribution map and eliminate redundant information:
[0359] Multi-scale sampling: In the three-dimensional distribution map, cubic grids (voxel size 0.1×0.1×0.1) are divided according to the scale level (1×1, 3×3, 5×5, etc.), and sensitivity statistical features (mean, variance, peak) are extracted within each grid.
[0360] Basis vector extraction: Principal component analysis (PCA) is performed on the grid features at each scale, and the principal components with a cumulative contribution rate greater than 85% are retained as basis vectors. For example, the feature dimension of the 5×5 scale is reduced from 100 to 15 dimensions.
[0361] Similarity metric: Cosine similarity is used to calculate the correlation between basis vectors. If the similarity is greater than 0.9, the basis vector is considered redundant. For example, if the similarity between a 3×3 basis vector A and a 5×5 basis vector B is 0.92, B is removed.
[0362] Orthogonalization: Use the Gram-Schmidt orthogonalization algorithm to perform orthogonal projection on the retained basis vectors to generate a set of scaled orthogonal basis vectors. For example, after filtering and orthogonalizing the original 20 basis vectors, 12 orthogonal bases are obtained.
[0363] Technical Example: In pedestrian detection, 3×3 basis vectors primarily capture limb contours, while 5×5 basis vectors describe overall posture. A similarity metric found that the similarity between the two basis vectors in the torso region reached 0.88. Therefore, redundant 5×5 basis vectors were removed, while 3×3 basis vectors were retained to reduce computational complexity.
[0364] The scale-orthogonal basis vector set is input into the adaptive sliding window generator, and a pyramid hierarchical structure is constructed through the nonlinear relationship between window size and scale sensitivity. At the same time, a neighborhood smoothing algorithm is used to eliminate noise interference.
[0365] The adaptive sliding window generator dynamically adjusts the window size according to scale sensitivity and constructs a multi-level pyramid structure:
[0366] Window size mapping: defines the nonlinear relationship between window size W and sensitivity S:
[0367] When S<0.3, W=3;
[0368] When 0.3≤S<0.6, W=5;
[0369] When S≥0.6, W=7.
[0370] The mapping is implemented through table lookup method to ensure real-time performance.
[0371] Pyramid level generation:
[0372] Bottom layer (L0): original resolution (e.g. 1024×1024), window size 3×3;
[0373] Middle layer (L1): downsampled to 512×512, window size 5×5;
[0374] High layer (L2): downsampled to 256×256, window size 7×7.
[0375] Neighborhood smoothing algorithm: Bilateral filtering (spatial domain σ = 1.5, grayscale domain σ = 0.1) is used to smooth each window layer, preserving edges while suppressing noise. For example, a sensitivity jump (0.2 → 0.8 → 0.3) caused by uneven illumination within a window is smoothed to (0.3 → 0.7 → 0.4).
[0376] Technical Examples:
[0377] In traffic monitoring scenarios, the vehicle taillight area has a sensitivity of up to 0.9, triggering a large 7×7 window to capture the overall shape. Meanwhile, the license plate area has a sensitivity of 0.5, using a 5×5 window to balance detail and efficiency. Bilateral filtering effectively smooths out jagged noise on the license plate edges.
[0378] Cross-scale confidence propagation is performed on each pyramid level, and the level confidence benchmark and error tolerance threshold are generated by combining the statistical distribution characteristics of historical recognition results.
[0379] Cross-scale confidence propagation improves decision reliability by integrating multi-level information:
[0380] Confidence propagation mechanism:
[0381] Bottom-up propagation: The confidence of local details (such as edge sharpness) of the bottom layer (L0) is weighted and transferred to the top layer (L2) with a weight of 0.6;
[0382] Top-down feedback: The semantic confidence of the high-level (L2) layer (such as vehicle category) reversely corrects the decision of the low-level (L0) layer with a weight of 0.4.
[0383] Historical statistics fusion:
[0384] Create a distribution histogram of historical recognition results (e.g., the frequency of occurrence of the "vehicle" category in the past 100 frames is 30%, and the frequency of occurrence of the "pedestrian" category is 15%).
[0385] Adjust the current confidence using the Bayesian update formula:
[0386]
[0387] Here is what each parameter in the formula means and its technical role:
[0388] P current : Current confidence
[0389] Definition: The initial confidence score of the current frame (or current moment) obtained by the multi-scale verification pyramid, ranging from [0, 1].
[0390] Technical Function: This reflects the model's immediate confidence in the current recognition result. For example, in object detection, if the initial confidence level of a region classified as a "vehicle" is 0.85, it means that the model is 85% confident that the object is a vehicle based on the current image features.
[0391] Calculation method: Usually generated by the softmax output of the neural network or post-processing (such as non-maximum suppression).
[0392] P history : Historical statistical probability
[0393] Definition: The prior probability obtained based on historical data (such as the recognition results of the past N frames), ranging from [0, 1].
[0394] Technical Benefit: This technology introduces contextual information in the temporal dimension to correct misjudgments caused by transient noise (such as sudden changes in illumination and motion blur). For example, if the average probability of the "vehicle" category in historical data is 30%, the current high confidence result (such as 0.9) may be lowered due to scene mismatch (such as detecting a vehicle in the sky).
[0395] Calculation method: Sliding window statistics of the frequency distribution of historical recognition results. For example, the number of occurrences of the "vehicle" category in the past 100 frames is 30, then P history =0.3.
[0396] P new : Updated confidence
[0397] Definition: The posterior probability after fusing the current confidence and the historical probability, ranging from [0,1].
[0398] Technical role: Through Bayesian theorem, real-time observation (P current ) and long-term statistics (P history ) to output a more robust confidence. For example:
[0399] If the current confidence is high (0.9) but the historical probability is low (0.1), the confidence will be significantly reduced after the update (such as 0.47) to avoid false detection; if the two are consistent (such as 0.8), the confidence will remain high after the update (0.8).
[0400] Features: When P history = 0.5 (no historical bias), P new =P current (history does not affect the current result); when P current =1 or 0, it also approaches 1 or 0 (extreme confidence is not affected by history).
[0401] Technical Examples
[0402] Assume that in an industrial quality inspection scenario:
[0403] Current frame: Initial confidence P that a part is classified as a “defect” current =0.8;
[0404] Historical statistics: the frequency P of "defects" in the past 1000 frames history =0.05 (the defect rate on a normal production line is low).
[0405] Then the updated confidence is: P new =0.74.
[0406] Effect: Although the current model gives a high confidence level (0.8), due to the extremely low historical defect rate (5%), the final confidence level was significantly lowered to 17.4%, effectively suppressing false positives.
[0407] Error tolerance threshold calculation:
[0408] Based on the mean square error (σ) of the confidence distribution, the threshold was set as μ ± 2σ (covering the 95% confidence interval);
[0409] Dynamic adjustment mechanism: If the threshold trigger rate is > 90% for 10 consecutive frames, σ is multiplied by 0.8 to tighten the threshold.
[0410] Technical Examples:
[0411] In a drone inspection scenario, a photovoltaic panel crack detection task:
[0412] Bottom layer (L0): edge confidence 0.85 (high detail confidence);
[0413] High-level (L2): semantic confidence 0.75 (slightly lower due to lighting changes);
[0414] After cross-layer propagation, the comprehensive confidence = 0.85×0.6+0.75×0.4=0.81;
[0415] Historical statistics show that the average crack detection rate for sunny scenes is μ = 0.82 and σ = 0.05, so the threshold is set to [0.72, 0.92]. The current result of 0.81 is within the threshold, so the judgment is valid.
[0416] It can be seen that the visible light image, infrared thermal imaging and depth point cloud data of the target scene are obtained to generate a spatiotemporally consistent multimodal heterogeneous feature tensor; the multimodal heterogeneous feature tensor is input into the spatial frequency perception optimizer to generate an optimized feature matrix with enhanced details; based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occlusion samples to generate a recognition feature vector with enhanced adversarial robustness; the recognition feature vector is input into the adaptive decision engine to generate dynamic decision parameters for environmental adaptation; a multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through the self-correction module to output the final recognition result, thereby achieving highly robust and accurate visual recognition.
[0417] Another embodiment of the present invention provides a visual recognition system, see Figure 3 , the system may include:
[0418] Acquisition module 301 is used to acquire visible light images, infrared thermal images, and depth point cloud data of the target scene, and to perform cross-domain alignment of the multimodal data using a heterogeneous feature fusion network to generate a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features of different spectral bands.
[0419] An enhancement module 302 is configured to input the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstruct a feature map using an adaptive high-frequency enhancement filter, and generate an optimized feature matrix for detail enhancement, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the illumination intensity;
[0420] A generation module 303 is configured to synthesize multi-scale occlusion samples using a generative adversarial network based on the optimized feature matrix, and generate a recognition feature vector with enhanced adversarial robustness in combination with a temporal consistency constraint, wherein the temporal consistency constraint is enforced to preserve the continuity of temporal features through an optical flow field loss function;
[0421] An adjustment module 304 is configured to input the identified feature vector into an adaptive decision engine, dynamically adjust the classification boundaries through a reinforcement learning-driven graph neural network, and generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths based on the complexity of the real-time scene;
[0422] The recognition module 305 is used to construct a multi-scale verification pyramid based on the dynamic decision parameters, fuse the confidence scores of each level through the self-correction module, and output the final recognition result, wherein the self-correction module uses an asymmetric loss function to eliminate cross-layer feature conflicts.
[0423] It can be seen that the visible light image, infrared thermal imaging and depth point cloud data of the target scene are obtained to generate a spatiotemporally consistent multimodal heterogeneous feature tensor; the multimodal heterogeneous feature tensor is input into the spatial frequency perception optimizer to generate an optimized feature matrix with enhanced details; based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occlusion samples to generate a recognition feature vector with enhanced adversarial robustness; the recognition feature vector is input into the adaptive decision engine to generate dynamic decision parameters for environmental adaptation; a multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through the self-correction module to output the final recognition result, thereby achieving highly robust and accurate visual recognition.
[0424] An embodiment of the present invention further provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps of any one of the above method embodiments when running.
[0425] Specifically, in this embodiment, the above-mentioned storage medium may be configured to store a computer program for performing the following steps:
[0426] S201, obtaining visible light images, infrared thermal images, and depth point cloud data of a target scene, performing cross-domain alignment on the multimodal data through a heterogeneous feature fusion network, and generating a spatiotemporally consistent multimodal heterogeneous feature tensor, wherein the heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features of different spectral bands;
[0427] S202: Inputting the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstructing a feature map through an adaptive high-frequency enhancement filter, and generating an optimized feature matrix for detail enhancement, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity;
[0428] S203, based on the optimized feature matrix, using a generative adversarial network to synthesize multi-scale occlusion samples, combined with a temporal consistency constraint, to generate a recognition feature vector with enhanced adversarial robustness, wherein the temporal consistency constraint is enforced to preserve the continuity of temporal features through an optical flow field loss function;
[0429] S204: Input the identified feature vector into an adaptive decision engine, dynamically adjust the classification boundary through a reinforcement learning-driven graph neural network, and generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths according to the complexity of the real-time scene;
[0430] S205, constructing a multi-scale verification pyramid based on the dynamic decision parameters, fusing the confidence scores of each level through a self-correction module, and outputting a final recognition result, wherein the self-correction module uses an asymmetric loss function to eliminate cross-layer feature conflicts.
[0431] It can be seen that the visible light image, infrared thermal imaging and depth point cloud data of the target scene are obtained to generate a spatiotemporally consistent multimodal heterogeneous feature tensor; the multimodal heterogeneous feature tensor is input into the spatial frequency perception optimizer to generate an optimized feature matrix with enhanced details; based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occlusion samples to generate a recognition feature vector with enhanced adversarial robustness; the recognition feature vector is input into the adaptive decision engine to generate dynamic decision parameters for environmental adaptation; a multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through the self-correction module to output the final recognition result, thereby achieving highly robust and accurate visual recognition.
[0432] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.
[0433] Specifically, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0434] Specifically, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0435] S201, obtaining visible light images, infrared thermal images, and depth point cloud data of a target scene, performing cross-domain alignment on the multimodal data through a heterogeneous feature fusion network, and generating a spatiotemporally consistent multimodal heterogeneous feature tensor, wherein the heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features of different spectral bands;
[0436] S202: Inputting the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstructing a feature map through an adaptive high-frequency enhancement filter, and generating an optimized feature matrix for detail enhancement, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity;
[0437] S203, based on the optimized feature matrix, using a generative adversarial network to synthesize multi-scale occlusion samples, combined with a temporal consistency constraint, to generate a recognition feature vector with enhanced adversarial robustness, wherein the temporal consistency constraint is enforced to preserve the continuity of temporal features through an optical flow field loss function;
[0438] S204: Input the identified feature vector into an adaptive decision engine, dynamically adjust the classification boundary through a reinforcement learning-driven graph neural network, and generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths according to the complexity of the real-time scene;
[0439] S205, constructing a multi-scale verification pyramid based on the dynamic decision parameters, fusing the confidence scores of each level through a self-correction module, and outputting a final recognition result, wherein the self-correction module uses an asymmetric loss function to eliminate cross-layer feature conflicts.
[0440] It can be seen that the visible light image, infrared thermal imaging and depth point cloud data of the target scene are obtained to generate a spatiotemporally consistent multimodal heterogeneous feature tensor; the multimodal heterogeneous feature tensor is input into the spatial frequency perception optimizer to generate an optimized feature matrix with enhanced details; based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occlusion samples to generate a recognition feature vector with enhanced adversarial robustness; the recognition feature vector is input into the adaptive decision engine to generate dynamic decision parameters for environmental adaptation; a multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through the self-correction module to output the final recognition result, thereby achieving highly robust and accurate visual recognition.
[0441] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings. The above is only a preferred embodiment of the present invention, but the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.
Claims
1. A visual recognition method, characterized in that: The method comprises: Obtain visible light images, infrared thermal images, and depth point cloud data of the target scene, perform cross-domain alignment on the multimodal data using a heterogeneous feature fusion network, and generate a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features from different spectral bands. Inputting the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstructing a feature map through an adaptive high-frequency enhancement filter to generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity; Based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occluded samples, and a recognition feature vector with enhanced adversarial robustness is generated in combination with a temporal consistency constraint, wherein the temporal consistency constraint is enforced to retain the continuity of temporal features through an optical flow field loss function; Inputting the identified feature vector into an adaptive decision engine, dynamically adjusting the classification boundaries through a reinforcement learning-driven graph neural network to generate dynamic decision parameters that are adaptive to the environment, wherein the graph neural network automatically activates subgraph structures of different depths based on the complexity of the real-time scene; A multi-scale verification pyramid is constructed based on the dynamic decision parameters, and the confidence scores of each level are fused through a self-correction module to output the final recognition result, wherein the self-correction module adopts an asymmetric loss function to eliminate cross-layer feature conflicts.
2. The method according to claim 1, characterized in that The method acquires visible light images, infrared thermal images, and depth point cloud data of the target scene, performs cross-domain alignment on the multimodal data through a heterogeneous feature fusion network, and generates a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network adopts a self-attention mechanism to fuse edge response features of different spectral bands, including: Based on the frame synchronization signal of the visible light image and the timestamp of the infrared thermal image, a dynamic interpolation compensation algorithm is used to align the sampling interval of the depth point cloud data to generate a trimodal data stream with strict time axis synchronization; Perform multispectral edge detection on visible light images, use the temperature gradient of infrared thermal imaging to correct visible light edge breakage areas, and extract the curvature mutation features of the depth point cloud to generate a cross-modal edge response map; The cross-modal edge response map is input into a dual-path self-attention network, and weights are dynamically assigned through a visible light-infrared feature cross-calibration module to generate a spectrally consistent edge feature matrix. The fused edge feature matrix is subjected to three-dimensional spatiotemporal convolution, and combined with the spatial topological constraints of the depth point cloud, a multimodal heterogeneous feature tensor with spatiotemporal consistency is output.
3. The method according to claim 2, characterized in that The multimodal heterogeneous feature tensor is input into a spatial frequency perception optimizer, and a feature map is reconstructed through an adaptive high-frequency enhancement filter to generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the illumination intensity, including: According to the brightness histogram distribution of the visible light image, a dynamic threshold segmentation algorithm is used to extract the scene illumination intensity level and generate a quantitative illumination intensity index. Perform two-dimensional wavelet decomposition on the multimodal heterogeneous feature tensor to separate the high-frequency detail component and the low-frequency contour component, and dynamically adjust the convolution kernel weight ratio of the high-frequency detail component based on the light intensity quantization index; Adaptive inverse wavelet transform is used to fuse the weighted high-frequency detail components and low-frequency contour components, and residual connection is used to compensate for the reconstruction error to generate an optimized feature matrix for detail enhancement. The optimized feature matrix is subjected to local contrast analysis, and a nonlinear filtering algorithm is used to suppress artifact noise caused by sudden changes in illumination, and the final optimized feature matrix after noise suppression is output.
4. The method according to claim 3, characterized in that Based on the optimized feature matrix, a generative adversarial network is used to synthesize multi-scale occluded samples, and a temporal consistency constraint is combined to generate a recognition feature vector with enhanced adversarial robustness, wherein the temporal consistency constraint is forced to retain the continuity of temporal features through an optical flow field loss function, including: The optimized feature matrix is input into the morphological erosion generator, and the edge area of the feature map is eroded by random structural elements to synthesize a multi-scale occlusion sample set; Based on the optimized feature matrix of adjacent frames, the dense optical flow estimation algorithm is used to extract the motion trajectory of the sample occluded area and generate the spatiotemporal continuity constraint vector; Combining the spatiotemporal continuity constraint vector of the optical flow trajectory with the multi-scale occlusion sample set, we design a directional consistency loss function to force the adversarial generative network to preserve temporal motion characteristics when synthesizing samples. The generator after adversarial training is used to perform feature perturbation enhancement on the optimized feature matrix that retains the temporal motion features, and the recognition feature vector with enhanced adversarial robustness is extracted through contrastive learning of the Siamese network.
5. The method according to claim 4, characterized in that The identification feature vector is input into the adaptive decision engine, and the classification boundary is dynamically adjusted through the reinforcement learning-driven graph neural network to generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths according to the complexity of the real-time scene, including: According to the spatial distribution density of the identified feature vectors, the node density clustering algorithm is used to calculate the real-time scene complexity index; A deep routing strategy based on the scene complexity index triggers the graph neural network, automatically selecting shallow local subgraphs or deep global subgraphs for feature aggregation; Dynamically adjust the classification decision boundaries of shallow local subgraphs or deep global subgraph nodes through a reinforcement learning reward mechanism, where the reward function is determined by both the classification confidence and the density of feature distribution; The optimized classification decision boundary is Gaussian smoothed, and the dynamic decision parameter matrix for environment adaptation is output in combination with the scene complexity index.
6. The method according to claim 5, characterized in that The multi-scale verification pyramid is constructed according to the dynamic decision parameters, and the confidence scores of each level are integrated through the self-correction module to output the final recognition result, wherein the self-correction module adopts an asymmetric loss function to eliminate cross-layer feature conflicts, including: According to the scale sensitivity of the dynamic decision parameter matrix, a stratified random sampling strategy is used to construct a multi-scale verification pyramid to generate a hierarchical confidence benchmark; Perform bidirectional attention alignment on the confidence scores of each level of the pyramid, eliminate the score deviation between scales through feature similarity measurement, and generate a consistent calibrated confidence distribution; Designing a directional constrained asymmetric loss function to perform differential weighted fusion of high-level semantic confidence and low-level detail confidence to suppress cross-layer feature conflicts; The fused confidence distribution is input into a multi-level weighted voter, which dynamically selects the optimal level result based on the preset threshold, and outputs the final recognition result and a traceable decision path.
7. The method according to claim 6, characterized in that According to the scale sensitivity of the dynamic decision parameter matrix, a stratified random sampling strategy is adopted to construct a multi-scale verification pyramid to generate a hierarchical confidence benchmark, including: According to the gradient distribution characteristics of the dynamic decision parameter matrix, the ripple diffusion algorithm is used to analyze the sensitivity of parameter changes at different scales and generate a three-dimensional distribution map containing the scale-sensitivity mapping relationship; Based on the three-dimensional distribution map, the dynamic registration algorithm is used to extract the characteristic basis vectors at each scale, and the feature similarity measurement is used to filter the redundant basis vectors to generate a scale orthogonal basis vector set. The scale-orthogonal basis vector set is input into the adaptive sliding window generator, and a pyramid hierarchical structure is constructed through the nonlinear relationship between window size and scale sensitivity. At the same time, a neighborhood smoothing algorithm is used to eliminate noise interference. Cross-scale confidence propagation is performed on each pyramid level, and the level confidence benchmark and error tolerance threshold are generated by combining the statistical distribution characteristics of historical recognition results.
8. A visual recognition system, characterized in that: The system comprises: An acquisition module is used to acquire visible light images, infrared thermal images, and depth point cloud data of the target scene, and to perform cross-domain alignment of the multimodal data using a heterogeneous feature fusion network to generate a temporally and spatially consistent multimodal heterogeneous feature tensor. The heterogeneous feature fusion network uses a self-attention mechanism to fuse edge response features from different spectral bands. an enhancement module, configured to input the multimodal heterogeneous feature tensor into a spatial frequency-aware optimizer, reconstruct a feature map through an adaptive high-frequency enhancement filter, and generate an optimized feature matrix with enhanced details, wherein the filter dynamically adjusts the frequency band weight distribution of the convolution kernel according to the light intensity; A generation module is configured to synthesize multi-scale occluded samples using a generative adversarial network based on the optimized feature matrix, and generate a recognition feature vector with enhanced adversarial robustness in combination with a temporal consistency constraint, wherein the temporal consistency constraint is enforced to preserve the continuity of temporal features through an optical flow field loss function; An adjustment module, configured to input the identified feature vector into an adaptive decision engine, dynamically adjust the classification boundaries through a reinforcement learning-driven graph neural network, and generate dynamic decision parameters for environmental adaptation, wherein the graph neural network automatically activates subgraph structures of different depths based on the complexity of the real-time scene; The recognition module is used to construct a multi-scale verification pyramid based on the dynamic decision parameters, fuse the confidence scores of each level through a self-correction module, and output the final recognition result, wherein the self-correction module adopts an asymmetric loss function to eliminate cross-layer feature conflicts.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on adaptive weight learning
CN114187221A
Target detection method based on infrared visible light feature enhancement and fusion
CN119418032A
Cited By
Fusion sensing system and method based on multispectral sensor
CN120742305A
Warehousing checking method based on multi-mode sensing technology, robot and warehousing system
CN120931210A
Visual laser displacement detection system and method based on multimode fusion and intelligent calibration
CN121025972A
Zero sample anomaly detection method and system based on triple perception learning enhanced visual language model
CN121121767A
Fan blade defect detection method and system based on visual feature enhancement
CN121236085A