Robot image recognition system based on convolutional neural network
Through joint training of dynamic retinal fuzzy perception module, differentiable physical defuzzy layer and fuzzy robust CNN backbone network, the motion fuzzy recognition error of industrial robots in dynamic scenarios is solved, and high-precision and robust image recognition are achieved.
Patent Information
- Application Number
- CN202510499113.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-25
AI Technical Summary
When industrial robots recognize images in dynamic scenes, the recognition error problem caused by motion blur is difficult for the prior art to take into account the blur characteristics of different regions, resulting in insufficient recognition robustness.
The dynamic retinal blur perception module is used to estimate the pixel-level dynamic fuzzy kernel parameters, combined with the differentiable physical defuzzy layer and the fuzzy robust CNN backbone network, the fuzzy kernel estimation and recognition process is optimized through the joint training module, forming a closed-loop link, and dynamically adjusting the hollow rate of the convolution layer to adapt to the local fuzzy intensity.
It improves the recognition accuracy and robustness in complex dynamic fuzzy scenarios, improves the generalization ability of real fuzzy data, and solves the recognition bottleneck of traditional methods under drastic changes in local fuzzy intensity.
Smart Images

Figure CN120375154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a robot image recognition system based on a convolutional neural network. Background Art
[0002] When an industrial robot vision system uses a CNN for image recognition, the motion blur interference in a dynamic scene is prominent. A fast-moving workpiece forms a trailing shadow in the image, and the edge texture is mixed with background noise, resulting in the feature extraction layer misjudging the target contour. For example, when a robotic arm grasps a part on a conveyor belt, the blurred area is easily recognized as a foreign object or an incorrect category.
[0003] Existing solutions usually process image restoration and target recognition in stages. The deblurring algorithm focuses on improving visual quality and may over-enhance high-frequency noise, while the recognition module needs to retain specific semantic features. This difference in goals results in the restored image losing key recognition information. For example, artifacts are generated after deblurring the barcode area, directly affecting the accuracy of the classification result.
[0004] Mainstream CNN architectures use fixed convolutional kernels to process the entire image. When there are both static backgrounds and high-speed moving objects in the scene, the unified convolution strategy is difficult to balance the blur characteristics of different regions. The details in the clear area are over-smoothed, while the receptive field in the highly blurred area is insufficient, both of which jointly weaken the recognition robustness of the system.
[0005] Most models rely on synthetic blurred data for training, but the blur patterns in real industrial scenarios are more complex. Phenomena such as non-uniform blur caused by mechanical vibration and light path scattering under multiple light sources are difficult to accurately model in synthetic data. This leads to a significant decline in the system's recognition performance when encountering unseen blur types during actual deployment. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a robot image recognition system based on a convolutional neural network, which solves the problem of robot vision recognition errors caused by motion blur in a dynamic scene.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A robot image recognition system based on a convolutional neural network, comprising: A dynamic retina blur perception module for estimating pixel-level dynamic blur kernel parameters from the input blurred image; A differentiable physical deblurring layer for reconstructing features of the blurred image based on the dynamic blur kernel parameters to generate deblurred image features; A blur-robust CNN backbone network for dynamically adjusting the dilation rate of the convolutional layer according to the dynamic blur kernel parameters to adapt to the local blur intensity, processing the deblurred image features and outputting a recognition result; The joint training module jointly optimizes the blur kernel estimation, feature reconstruction, and recognition processes through an end-to-end loss function, and backpropagates the recognition error to the dynamic retinal blur perception module to correct the blur kernel estimation.
[0008] Preferably, the dynamic retinal blur perception module includes: A spatio-temporal feature extraction unit for receiving an input blurred image and generating pseudo 3D spatio-temporal features; A blur kernel parameter generation unit for outputting pixel-level blur kernel parameters based on the pseudo 3D spatio-temporal features; A dynamic region mask unit for generating a dynamic region mask , and outputting non-zero blur kernel parameters only for the dynamic region.
[0009] Preferably, the differentiable physical deblurring layer includes: An iterative expansion unit for performing multi-step differentiable deconvolution based on the input blur kernel parameters ; A regularization gradient unit for calculating the gradient sparsity constraint and the semantic consistency constraint , and feeding back to the iterative process.
[0010] Preferably, the blur-robust CNN backbone network includes: A blur intensity calculation unit for calculating the local blur intensity according to the input blur kernel parameters ; A dilation rate decision unit for dynamically selecting the dilation rate based on a threshold ; A dynamic dilated convolution unit for processing the deblurred image features at the dilation rate
[0011] Preferably, the formula for calculating the local blur intensity in the blur intensity calculation unit is: ; where and are the horizontal and vertical components of the motion speed, in pixels per frame, and are the standard deviations of the motion paths in the horizontal and vertical directions, in pixels, is the comprehensive index of the local blur intensity, dimensionless.
[0012] Preferably, the joint training module includes: A multi-task loss calculation unit for calculating the end-to-end joint loss; A curriculum learning strategy unit for optimizing the model parameters in two stages; The gradient backpropagation unit backpropagates the classification loss gradient to the dynamic blur kernel parameters through the chain rule.
[0013] Preferably, the end-to-end loss function of the multi-task loss calculation unit is: ; where , and are the weight coefficients of the loss function, is the structural similarity index, measuring the similarity between the deblurred image and the real clear image , is the cross-entropy classification loss, measuring the difference between the predicted label and the real label , is the optical flow continuity physical constraint loss, ensuring the smoothness and physical rationality of the velocity field .
[0014] Preferably, the curriculum learning strategy unit optimizes the model parameters in two stages, including: In the first stage, a synthetic blur dataset is used to fix the dynamic retina blur perception module, and only the differentiable physical deblurring layer and the blur-robust CNN backbone network are optimized; In the second stage, the parameters of the first stage are loaded, and a real blur dataset is used to jointly optimize all modules.
[0015] The present invention provides a robot image recognition system based on a convolutional neural network. It has the following beneficial effects: 1. By adopting the technical solution of combining the dynamic retina blur perception module with the optical flow continuity physical constraint, the present invention realizes the physically reasonable estimation of the motion blur parameters. By introducing the divergence constraint of the velocity field and the velocity gradient smoothing term, the problem of blur kernel distortion caused by traditional pure data-driven methods ignoring physical motion laws is solved. Compared with the "black box" scheme that only relies on CNN to learn the blur kernel in the prior art, the combination of physical prior and data-driven significantly improves the modeling accuracy of complex dynamic blur.
[0016] 2. Based on the design of the CNN backbone network that dynamically adjusts the dilation rate according to the blur intensity, the present invention enables the model to autonomously adapt to the blur intensity differences in different regions. Existing methods using fixed convolution kernels or single deblurring strategies are difficult to handle real scenarios with drastic changes in local blur intensity. The present invention expands the receptive field of high-blur regions while retaining details through the dilation rate decision unit and dynamic dilated convolution, solving the defect of insufficient recognition robustness of traditional schemes under complex blur.
[0017] 3. Through the joint training module, the present invention reversely penetrates the classification error to the blur kernel parameters, forming a closed-loop link of "recognition-driven reconstruction and reconstruction-optimized perception". The prior art's solution of independently training each module in stages severs the association between blur estimation and downstream tasks. Combining with the two-stage curriculum learning strategy, first solidify the physical model and then unlock the parameter joint tuning, which not only avoids unstable initial training but also improves the generalization ability to real blur data. The traditional method has obvious performance bottlenecks due to the lack of global optimization and progressive learning mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a framework diagram of the system of the present invention; Figure 2 is a framework diagram of the dynamic retinal blur perception module of the present invention; Figure 3 is a framework diagram of the differentiable physical deblurring layer of the present invention; Figure 4 is a framework diagram of the blur-robust CNN backbone network of the present invention; Figure 5 is a framework diagram of the joint training module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] Please refer to the attached Figure 1 , the embodiment of the present invention provides a robot image recognition system based on a convolutional neural network, including: Please refer to the attached Figure 2 , a dynamic retinal blur perception module for estimating pixel-level dynamic blur kernel parameters from the input blurred image; The dynamic retinal blur perception module includes: A spatio-temporal feature extraction unit for receiving the input blurred image and generating pseudo-3D spatio-temporal features; A blur kernel parameter generation unit for outputting pixel-level blur kernel parameters based on the pseudo-3D spatio-temporal features; A dynamic region mask unit for generating a dynamic region mask and outputting non-zero blur kernel parameters only for the dynamic region.
[0021] The dynamic retinal blur perception module is the core front-end component of the system, responsible for parsing pixel-level motion parameters from a single-frame blurred image. By fusing spatio-temporal feature extraction and physical motion modeling, this module generates dynamic blur kernel parameters, providing physical constraints for subsequent deblurring and recognition. The output parameters of the module are directly passed to the differentiable physical deblurring layer to form an initial data link for end-to-end optimization.
[0022] In this embodiment, the dynamic retinal blur perception module consists of a spatio-temporal feature extraction unit, a blur kernel parameter generation unit, and a dynamic region mask unit. The specific technical implementation is as follows: Generally, this unit processes the input of a single-frame blurred image and constructs pseudo-3D spatio-temporal features. Specifically, for the input image , bilinear interpolation is used to generate two virtual adjacent images 、 , which are concatenated into a pseudo-3D input .
[0023] In some embodiments, the encoder uses four layers of 3D convolution with a convolution kernel size of 3×3×3, and the number of channels increases layer by layer (16→32→64→128). Strided convolution is used for downsampling to capture multi-scale spatio-temporal motion patterns. The decoder gradually restores the resolution through transposed 3D convolution. The formula is: ; Among them, represents the layer index, 、 are learnable parameters for each layer, and the final output feature , is the output feature map of the -th layer of the decoder.
[0024] In a possible implementation, this unit maps the decoder output to a five-dimensional blur kernel parameter space. Specifically, the number of output channels of the last layer of convolution is 5, and the SoftPlus function is used to ensure non-negativity of the parameters. The expression is: ; Among them 、 are the parameters of the fully connected layer, and the SoftPlus function is defined as .
[0025] The output parameters include: : Horizontal and vertical motion speeds (pixels / frame); : Standard deviations of the motion path in the x / y directions (pixels); : Angle of the main motion direction (radians).
[0026] In some embodiments, the unit detects the motion area through inter-frame difference. Specifically, for the input image sequence { }, the difference between adjacent frames is calculated: ; where is the RGB color vector of the current frame ( ) at the pixel position , is the RGB color vector of the previous frame ( ) at the pixel position , and is the color difference amplitude between the current frame and the previous frame at the pixel position .
[0027] A binary mask is generated: ; where {0, 1} is the binary dynamic area mask, identifying the motion area in the image, and the threshold = 0.1 is set empirically. The blur kernel parameter of the static area ( = 0) is reset to the default value , and the correction formula is: ; As an option, the module introduces the optical flow continuity constraint loss, and its formula is:
[0028] where: , are the divergence terms of the velocity field, forcing the rigid motion assumption to be satisfied; , are the velocity gradient terms, suppressing parameter mutations; = 0.01 is the smoothing term weight coefficient.
[0029] In some embodiments, the spatio-temporal feature extraction unit can use separable 3D convolution to reduce the computational amount. For example, the standard 3D convolution is decomposed into spatial convolution (2D) and temporal convolution (1D), and the formula is: Conv2D Conv1D ; where , are the spatial and temporal convolution kernels respectively, and is the input feature map, is the output feature map of the spatial convolution, is the output feature map of the temporal convolution and is the final result of the separable 3D convolution.
[0030] In a possible implementation, the dynamic region mask can be generated by a pre-trained optical flow network. For example, FlowNet2.0 is used to calculate the optical flow field ( , ), and a mask is generated: ; wherein, = 0.5 is the optical flow amplitude threshold.
[0031] The blur kernel parameters output by this module are directly passed to the downstream differentiable physical deblurring layer. During the training phase, the physical constraint loss and the classification loss jointly participate in the gradient backpropagation, forcing the blur kernel parameters to satisfy both physical rationality and task requirements at the same time. The mask processing of the static region reduces the amount of invalid calculations, ensures that system resources are concentrated on the analysis of dynamic targets, and improves the overall efficiency.
[0032] Please refer to Appendix Figure 3 , the differentiable physical deblurring layer, which reconstructs the features of the blurred image based on the dynamic blur kernel parameters to generate deblurred image features; The differentiable physical deblurring layer is the core reconstruction module of the system. Based on the pixel-level blur kernel parameters provided by the upstream dynamic retina blur perception module, it performs a physically driven feature recovery on the input blurred image. This layer generates deblurred features adapted to the downstream blur-robust CNN backbone network by fusing the differentiable implementation of the traditional iterative deconvolution and the hybrid regularization constraint. The module output is jointly optimized with the recognition task to form a closed-loop link from blur estimation to feature reconstruction.
[0033] In this embodiment, the differentiable physical deblurring layer is composed of an iterative expansion unit and a regularization gradient unit, and its technical implementation and module connection logic are as follows: Generally, this unit expands the Lucy-Richardson deconvolution algorithm into a differentiable multi-step iterative calculation. The input blurred image and the blur kernel parameters are used to perform N = 5 iterations of reconstruction.
[0034] In a possible implementation, the update formula for the th iteration is: ; where: : per-pixel multiplication operator; : Convolution operation, with zero-padding for boundary handling; : Blur kernel 's 180-degree rotation kernel, which is mathematically equivalent to ; = 1×10 −6 : Numerical stability constant to prevent division by zero; = 0.1: Iteration step size to control the update amplitude; = 0.5, = 0.3: Regularization term weight to balance edge sharpening and semantic preservation; Initial value .
[0035] As an option, this unit introduces dual-path regularization constraints that act on the image gradient domain and the semantic feature domain respectively: Gradient sparsity constraint: Use the Sobel operator to extract horizontal and vertical gradients: ; Among them, , is the input image or feature map, is the horizontal Sobel convolution kernel weight for extracting the horizontal direction gradient, is the vertical Sobel convolution kernel weight for extracting the vertical direction gradient, and are the horizontal and vertical direction gradient maps, with single-channel output, is the transpose of the matrix , that is, the result after interchanging the row and column indices.
[0036] Constrain the L1 sparsity of the gradient magnitude: ; Among them, is the RGB vector at the pixel position of the deblurred image at the -th iteration, is the horizontal gradient value of the -th iteration image, is the vertical gradient value of the -th iteration image, is the gradient sparsity regularization term for constraining the edge sharpness of the reconstructed image.
[0037] Constraint: Extract features through the ReLU3-3 layer of the pre-trained VG, G-16 network: ; Among them, is the input image, which is from the intermediate iteration result of the differentiable physical deblurring layer , is the third ReLU activation layer in the third convolutional group (Block3) of the pre-trained VGG-16 neural network, is the spatial resolution of the feature map, which is determined by the VGG-16 network structure, is the number of channels of the feature map, corresponding to the output channels of the third layer convolution in Block3 of the VGG-16 network, is the semantic feature map output by the VGG-16 network, which is used to calculate the semantic consistency loss.
[0038] Calculate the feature similarity loss: ; Among them, is the th iteration deblurred image 's semantic feature map, which is extracted by the ReLU3-3 layer of the pre-trained VGG-16 network, is the semantic feature map of the real clear image extracted by the same VGG-16 network, is the spatial resolution of the feature map, which is determined by three max-pooling operations of the VGG-16 network, is the normalization coefficient, which averages the loss value in the spatial dimension, is the semantic consistency regularization term, which is used to constrain the high-level semantic features of the reconstructed image.
[0039] Upstream input: Blur kernel parameter is provided by the dynamic retina blur perception module, and its calculation depends on the pseudo 3D features output by the spatio-temporal feature extraction unit; Dynamic region mask acts on to ensure that the static region uses the default parameters to reduce invalid calculations.
[0040] Downstream output: Final iteration result is input to the blur-robust CNN backbone network, and its dynamic dilated convolution unit adjusts the receptive field according to the local blur intensity ; In the training stage, the classification loss gradient is backpropagated to the deblurring layer through the chain rule: ; Among them, is the classification loss function, which measures the difference between the predicted class of the output of the fuzzy-robust CNN backbone network and the true label. is the pixel-level blur kernel parameter output by the dynamic retina blur perception module. is after the iterative unfolding unit passes through iterations, the deblurred image output. is the th iteration output for the th iteration input Jacobian matrix. is the th iteration image with respect to the blur kernel parameter partial derivative matrix. is the total number of iterations of the iterative unfolding unit.
[0041] Drives the optimization of the blur kernel parameter estimation module to form an end-to-end closed loop.
[0042] In some embodiments, the iterative process can adopt an adaptive step size strategy. For example, adjust the step size according to the local blur intensity: ; In one possible implementation, the regularization gradient unit can introduce an adversarial loss to improve the visual authenticity of the reconstructed image through the discriminator network D: ; Among them, the discriminator D and the generator (deblurring layer) are alternately trained to enhance the detail generation ability.
[0043] Please refer to Appendix Figure 4 , the fuzzy-robust CNN backbone network, dynamically adjusts the dilation rate of the convolutional layer according to the dynamic blur kernel parameter to adapt to the local blur intensity, processes the deblurred image features and outputs the recognition result; In this embodiment, the fuzzy-robust CNN backbone network is composed of a blur intensity calculation unit, a dilation rate decision unit, and a dynamic dilated convolution unit. The specific technical implementation is as follows: Generally, this unit receives the blur kernel parameter provided by the dynamic retina blur perception module and calculates the local blur intensity index. Specifically, the blur intensity is comprehensively calculated through the velocity component and the path standard deviation: ; Among them, and are the horizontal and vertical direction motion velocity components, with the unit of pixel / frame, and are the horizontal and vertical direction motion path standard deviations, with the unit of pixel, is the local blur intensity comprehensive index, dimensionless.
[0044] As an option, the unit dynamically selects the dilation rate according to a blur intensity threshold Specifically: ; ; = 1.5: Empirical threshold for distinguishing high-blur regions from low-blur regions; ∈ {1, 2}: Dilation rate selection result, controlling the convolution kernel sampling interval.
[0045] In a possible implementation, the unit replaces standard convolution with dynamic dilated convolution, and its output feature is calculated as: ; Where: : Input feature map, from the differentiable physical deblurring layer; : Shared weight matrix, consistent with the standard convolution kernel parameters; : Convolution kernel sampling offset; : The convolution kernel size is 3×3.
[0046] Upstream input: The reconstructed feature output by the deblurring layer As the input to the backbone network; The blur kernel parameters provided by the dynamic retina module Generated by the blur intensity calculation unit .
[0047] Downstream output: The output feature map of the dynamic dilated convolution unit is passed to the subsequent classification layer to generate the recognition result ; In the training stage, the classification loss gradient is backpropagated to the dynamic dilated convolution unit to drive the parameter optimization of blur intensity calculation and dilation rate decision-making.
[0048] In some embodiments, the dilation rate decision-making can adopt a multi-level threshold strategy. For example, three levels of dilation rates are divided according to the blur intensity: ; In a possible implementation, the dynamic dilated convolution unit can integrate a dynamic region mask and only apply dilation rate adjustment to the dynamic region: ; Where, ∈{0,1} is the dynamic region mask, identifying the current position Whether it is a motion region ∈{1,2} is the dynamic hole rate, controlling the convolutional kernel sampling interval is the input feature map, sourced from the output of the differentiable physical deblurring layer , is the shared convolutional kernel weight matrix, compatible with standard convolutional kernel parameters are the position coordinates of the output feature map and are the convolutional kernel sampling offsets, covering the neighborhood of a 3×3 convolutional kernel
[0049] Please refer to Appendix Figure 5 , the joint training module, jointly optimizes the blur kernel estimation, feature reconstruction, and recognition processes through an end-to-end loss function, and backpropagates the recognition error to the dynamic retinal blur perception module to correct the blur kernel estimation
[0050] The joint training module is the optimization core of the system, driving the collaborative learning of the dynamic retinal blur perception module, the differentiable physical deblurring layer, and the blur-robust CNN backbone network through an end-to-end multi-task loss. This module backpropagates the recognition error to the blur kernel parameter estimation process, forming a closed-loop optimization link from perception to reconstruction and then to recognition. The module input covers the outputs of all upstream modules and domesticates the model in stages through a curriculum learning strategy to enhance the generalization ability for synthetic and real data
[0051] In this embodiment, the joint training module includes a multi-task loss calculation unit, a curriculum learning strategy unit, and a gradient backpropagation unit, and its technical implementation is as follows Generally, this unit synthesizes the losses in three aspects: image reconstruction quality, classification accuracy, and physical rationality. The end-to-end loss function is defined as ; Parameter definitions = 0.3, = 0.7, = 0.1: Weight coefficients, balancing the optimization priorities of each task : Structural similarity index, with the calculation formula ; Where , is the local mean (calculated through Gaussian filtering, standard deviation = 1.5); , is the variance is the covariance; , , is the pixel value range.
[0052] : Cross-entropy classification loss, the calculation formula is: ; where, is the number of classes, is the one-hot encoding of the true label, is the predicted probability; : Optical flow continuity loss: ; where, , is the velocity component, 、 is the divergence of the velocity field, 、 is the gradient; : Gradient sparse regularization term: ; where, 、 are calculated through the Sobel convolution kernel; : Semantic consistency regularization term: ; where, (·) is the VGG-16 ReLU3-3 layer feature extractor, =
H / 8
W / 8
[0053] The curriculum learning strategy unit is an option, and this unit optimizes the model parameters in two stages: The first stage (synthetic data pre-training): Input data: Synthetic blurred dataset (such as GoPro), generate blurred kernel parameters through known motion trajectories ; Parameter optimization range: Freeze the parameters of the dynamic retina blur perception module, and only optimize the differentiable physical deblurring layer and the blur-robust CNN backbone network; Optimization objective function: ; Technical effect: Stabilize the initial convergence of the deblurring layer and the classifier, and avoid interference from the early training instability of the blur kernel estimation module.
[0054] The second stage (fine-tuning with real data): Input data: Real blur dataset (such as RealBlur), blur kernel parameters Estimated by the dynamic retina module; Parameter optimization range: Load the parameters of the first stage, unfreeze the dynamic retina module, and jointly optimize all modules; Optimization objective function: ; Technical effect: Improve the adaptability of the model to real complex blur, and strengthen the constraints of physical rationality and semantic consistency.
[0055] In a possible implementation, the unit backpropagates the classification loss gradient to the blur kernel parameters through the chain rule. Specifically, the gradient calculation path is: ; Where: : The deblurred image after 5 iterations, and the iteration formula is: ; is the deblurred image after the th iteration; is the input blurred image; is the 180-degree rotation kernel of the blur kernel ; is the gradient term of the gradient sparsity constraint; is the gradient term of the semantic consistency constraint : Based on the automatic differentiation result of the iteration formula, implemented through PyTorch Autograd or TensorFlow Gradient Tape; Gradient accumulation mechanism: Cover all 5 iteration steps to ensure the integrity of end-to-end optimization.
[0056] Upstream input: Output by the dynamic retina module and the dynamic region mask ; Output by the deblurring layer ; Predicted label output by the CNN backbone network .
[0057] Downstream output: The multi-task loss gradient is backpropagated to all module parameters; The stage training strategy realizes the optimization goal by controlling the parameter freezing / thawing state.
[0058] In some embodiments, the multi-task loss weights can be dynamically adjusted. For example, the uncertainty weighting method is adopted: ; where , , are learnable noise parameters, optimized by backpropagation. In a possible implementation, the curriculum learning strategy can be extended to progressive training, gradually increasing the real data mixing ratio: Mixing ratio epoch); where epoch is the current training epoch, gradually transitioning from synthetic data to real data.
[0059] Working principle: When the system starts, the dynamic retinal blur perception module captures the input blurred image and analyzes the motion cues in the image through the spatio-temporal feature extraction unit. This module simulates the compensation mechanism of the biological retina, expands the single-frame image into a pseudo-3D feature containing a virtual time dimension, and uses multi-layer 3D convolution to mine spatio-temporal correlations. The blur kernel parameter generation unit then outputs the motion speed, path randomness, and direction parameters of each pixel point. The dynamic region mask unit synchronously detects the motion regions in the image, only retains non-zero blur kernel parameters within the dynamic regions, and uses default values in the static regions to reduce computational redundancy.
[0060] After receiving the blur kernel parameters, the differentiable physical deblurring layer starts the iterative unfolding unit for physically constrained feature reconstruction. Each iteration fuses deconvolution operations with regularized gradient correction: the gradient sparsity constraint sharpens the edges through the Sobel operator, and the semantic consistency constraint calls the pre-trained VGG network to compare key features to ensure that the reconstructed image conforms to physical laws and retains semantic information. After multiple iterations, the blurred image is gradually restored to a high-definition feature map.
[0061] The blur-robust CNN backbone network dynamically adjusts the processing strategy according to the local blur intensity. The blur intensity calculation unit synthesizes the speed and path standard deviation to generate a regional blur intensity index; the dilation rate decision unit switches the convolution sampling interval based on a threshold - expanding the receptive field in high-intensity blur regions to capture context information, and maintaining the standard sampling density in low-intensity regions. The dynamic dilated convolution unit operates within the residual structure, maintains the feature dimension through shared weights and 1×1 convolution, and balances blur adaptability and computational efficiency.
[0062] The joint training module drives end-to-end optimization: The multi-task loss unit synchronously evaluates the image reconstruction quality, classification accuracy, and physical rationality. The curriculum learning strategy domesticates the model in stages - initially using synthetic data to stabilize the deblurring and recognition modules, and later introducing real data for joint fine-tuning. During the training process, the classification error penetrates backward to the blur kernel parameters through the chain rule, forcing the dynamic retina module to accurately capture complex motion patterns, forming a closed loop of "perception → reconstruction → recognition → feedback". Finally, the system achieves high-precision real-time recognition in dynamic scenarios.
[0063] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A robot image recognition system based on a convolutional neural network, characterized in that, Comprising: A dynamic retinal blur perception module for estimating pixel-level dynamic blur kernel parameters from an input blurred image; A differentiable physical deblurring layer for reconstructing features of the blurred image based on the dynamic blur kernel parameters to generate deblurred image features; A blur-robust CNN backbone network that dynamically adjusts the dilation rate of convolutional layers according to the dynamic blur kernel parameters to adapt to local blur intensity, processes the deblurred image features and outputs an identification result; A joint training module that jointly optimizes the blur kernel estimation, feature reconstruction, and identification processes through an end-to-end loss function, and backpropagates the identification error to the dynamic retinal blur perception module to correct the blur kernel estimation.
2. The robot image recognition system based on a convolutional neural network according to claim 1, characterized in that The dynamic retinal blur perception module includes: A spatio-temporal feature extraction unit for receiving an input blurred image and generating pseudo-3D spatio-temporal features; A blur kernel parameter generation unit for outputting pixel-level blur kernel parameters based on the pseudo-3D spatio-temporal features; Dynamic region mask unit, generating a dynamic region mask , and outputting non-zero blur kernel parameters only for the dynamic region.
3. The robot image recognition system based on a convolutional neural network according to claim 1, characterized in that, The differentiable physical deblurring layer includes: Iterative expansion unit, based on the input blurred kernel parameters Perform multi-step differentiable deconvolution; Regularized gradient unit for calculating gradient sparsity constraint and semantic consistency constraint , and feedback to the iterative process.
4. The robot image recognition system based on a convolutional neural network according to claim 1, characterized in that, The blur-robust CNN backbone network includes: A blur intensity calculation unit that calculates local blur intensity according to the input blur kernel parameters Calculate the local blur intensity; Void ratio decision unit, based on a threshold Dynamically select the void ratio ; Dynamic dilated convolution unit, with dilation rate Process the deblurred image features.
5. The robot image recognition system based on a convolutional neural network according to claim 4, characterized in that, The calculation formula for local blur intensity in the blur intensity calculation unit is: ; wherein, and the horizontal and vertical velocity components of motion, in pixels per frame, and are the standard deviations of the motion paths in the horizontal and vertical directions, in pixels, is the comprehensive index of local blur intensity, dimensionless.
6. The robot image recognition system based on a convolutional neural network according to claim 1, characterized in that, The joint training module includes: A multi-task loss calculation unit for calculating an end-to-end joint loss; A curriculum learning strategy unit for optimizing model parameters in two stages; A gradient backpropagation unit for backpropagating the classification loss gradient to the dynamic blur kernel parameters through the chain rule.
7. The robot image recognition system based on a convolutional neural network according to claim 6, wherein The end-to-end loss function of the multi-task loss calculation unit is: ; Among them, , and are the weight coefficients of the loss function, is the structural similarity index, which measures the similarity between the deblurred image and the real clear image . is the cross-entropy classification loss, which measures the difference between the predicted label and the real label . is the physical constraint loss of optical flow continuity, which ensures the smoothness and physical rationality of the velocity field .
8. The robot image recognition system based on a convolutional neural network according to claim 6, characterized in that, Optimizing model parameters in two stages in the curriculum learning strategy unit includes: In the first stage, a synthetic blurred dataset is used to fix the dynamic retinal blur perception module, and only the differentiable physical deblurring layer and the blur-robust CNN backbone network are optimized; In the second stage, the parameters of the first stage are loaded, and a real blurred dataset is used to jointly optimize all modules.
Citation Information
Cited By
Aircraft classification method and system based on fuzzy direction prior
CN120808322A
Video stream image optimization method and device
CN121010766A
Video stream image optimization method and device
CN121010766B
Markerless monocular video 3d pose reconstruction method and system
CN122550765A