Image target detection method, device, equipment and medium based on adaptive enhancement
By building a camera degradation fingerprint library and a rule engine for adaptive image processing, combined with Transformer feature fusion and lightweight deployment, the problem of poor imaging quality in old cameras was solved, target detection accuracy was improved, and modification costs were reduced.
Patent Information
- Application Number
- CN202511021992.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-24
AI Technical Summary
The poor imaging quality of old cameras in existing community security systems leads to low target detection accuracy. The high cost of replacing high-definition cameras makes it impossible to effectively address the imaging degradation problem of old equipment.
Image target detection is achieved by building a camera degradation fingerprint library for image diagnosis, using a rule engine to dynamically select image processing combinations, adopting a Transformer-based feature fusion architecture for multi-view feature fusion, and using model compression and dynamic jump inference technology for lightweight deployment.
Without replacing the hardware, the imaging quality and target detection accuracy of old cameras are improved, the cost of intelligent property transformation is reduced, the service life of the equipment is extended, and the accuracy of abnormal event detection and real-time edge processing capabilities are improved.
Smart Images

Figure CN120543995B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image target detection method, device, equipment and medium based on adaptive enhancement. Background Art
[0002] In recent years, with the advancement of intelligent property management, residential security monitoring systems have achieved near-perfect coverage. Current property security systems generally utilize a model combining video surveillance with intelligent analysis, deploying high-definition cameras to enable multi-scenario behavioral analysis. However, these systems face two significant challenges in practical applications. First, existing algorithms are primarily designed for new HD cameras and lack effective solutions for image degradation (such as low resolution, blur, noise, and color distortion) associated with older equipment, resulting in significantly reduced detection performance. Second, the large-scale deployment of new HD cameras incurs high replacement costs, while the need to reuse older equipment is urgent.
[0003] The root causes of the above problems are: traditional image enhancement methods use fixed parameters to process all camera images, which cannot adapt to the degradation characteristics of different devices, resulting in blurred details and residual noise in the enhanced images, affecting the subsequent target detection accuracy; when multiple cameras are collaboratively analyzed, they usually simply splice images and ignore the geometric distortion caused by differences in perspective; when the computing power of edge devices is limited, they often directly reduce the complexity of the algorithm and sacrifice processing accuracy. Especially in typical scenarios such as community garages and areas with insufficient lighting at night, the detection accuracy of existing systems is generally less than 60%, which seriously restricts the implementation of intelligent property management. Therefore, how to solve the problems of poor imaging quality and low target detection accuracy of old cameras in existing community security systems is a technical problem that technicians in this field need to solve urgently. Summary of the Invention
[0004] The embodiments of the present invention provide an image target detection method, apparatus, computer equipment, and medium based on adaptive enhancement, aiming to solve the problems of poor imaging quality and low target detection accuracy of old cameras in existing community security systems.
[0005] In a first aspect, an embodiment of the present invention provides an image target detection method based on adaptive enhancement, comprising:
[0006] Obtain input images captured by cameras with different viewing angles, and perform image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results;
[0007] Based on the image diagnosis result, dynamically selecting an image processing combination using a rule engine, and performing image enhancement processing on the input image using the image processing combination to obtain an enhanced image;
[0008] Using a Transformer-based feature fusion architecture to perform multi-view feature fusion on the enhanced image to obtain a fused image;
[0009] Performing target detection on the fused image to obtain corresponding target detection results, thereby constructing an image target detection model;
[0010] The image target detection model is lightweight deployed by using model compression and dynamic jump reasoning technology, and the deployed image target detection model is used to perform target detection on a specified image.
[0011] In a second aspect, an embodiment of the present invention provides an image object detection device based on adaptive enhancement, comprising:
[0012] An image diagnosis unit is used to obtain input images captured by cameras with different viewing angles, and perform image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results;
[0013] an image enhancement unit, configured to dynamically select an image processing combination using a rule engine based on the image diagnosis result, and perform image enhancement processing on the input image using the image processing combination to obtain an enhanced image;
[0014] An image fusion unit, configured to perform multi-view feature fusion on the enhanced image using a Transformer-based feature fusion architecture to obtain a fused image;
[0015] An image detection unit is used to perform target detection on the fused image to obtain corresponding target detection results, thereby constructing an image target detection model;
[0016] The model deployment unit is used to use model compression and dynamic jump reasoning technology to lightweight deploy the image target detection model, and use the deployed image target detection model to perform target detection on the specified image.
[0017] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the image target detection method based on adaptive enhancement as described in the first aspect is implemented.
[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for image target detection based on adaptive enhancement as described in the first aspect is implemented.
[0019] Embodiments of the present invention provide an image target detection method, apparatus, computer device, and medium based on adaptive enhancement. The method comprises: obtaining an input image captured by cameras with different viewpoints, and performing image diagnosis on the input image using a preset camera degradation fingerprint library to obtain an image diagnosis result; dynamically selecting an image processing combination based on the image diagnosis result using a rule engine, and performing image enhancement processing on the input image using the image processing combination to obtain an enhanced image; performing multi-view feature fusion on the enhanced image using a Transformer-based feature fusion architecture to obtain a fused image; performing target detection on the fused image to obtain a corresponding target detection result, thereby constructing an image target detection model; lightweight deployment of the image target detection model using model compression and dynamic jump reasoning technology, and performing target detection on a specified image using the deployed image target detection model. The embodiment of the present invention provides a scientific basis for subsequent processing through the camera degradation fingerprint library; secondly, image enhancement processing is achieved through the rule engine intelligent combination processing strategy; then, a Transformer-based multi-view feature fusion method is proposed to improve detection performance in complex scenarios; finally, model compression and dynamic jump reasoning technology are used to achieve efficient operation of the algorithm on edge devices. Compared with the existing technology, the embodiments of the present invention improve the accuracy of small target detection while keeping the original camera hardware unchanged, effectively solving the problem of decreased detection performance caused by poor imaging quality of old cameras, improving image quality and target detection accuracy without replacing hardware, reducing the cost of property intelligent transformation, and providing a cost-effective solution for intelligent property management. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of a flow chart of an image target detection method based on adaptive enhancement provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of the principle architecture of an image target detection method based on adaptive enhancement provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of a sub-process of an image target detection method based on adaptive enhancement provided by an embodiment of the present invention;
[0024] Figure 4A schematic block diagram of an image object detection device based on adaptive enhancement provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0026] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0027] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0028] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0029] See below Figure 1 An embodiment of the present invention provides an image target detection method based on adaptive enhancement, which specifically includes: steps S101 to S105.
[0030] Step S101: obtaining input images captured by cameras with different viewing angles, and performing image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results;
[0031] Step S102: Based on the image diagnosis result, dynamically select an image processing combination using a rule engine, and perform image enhancement processing on the input image using the image processing combination to obtain an enhanced image;
[0032] Step S103: using a Transformer-based feature fusion architecture to perform multi-view feature fusion on the enhanced image to obtain a fused image;
[0033] Step S104: performing target detection on the fused image to obtain corresponding target detection results, thereby constructing an image target detection model;
[0034] Step S105: Use model compression and dynamic jump reasoning technology to lightweight deploy the image target detection model, and use the deployed image target detection model to perform target detection on the specified image.
[0035] Combine Figure 2 This embodiment first performs image diagnosis on the input image through the camera degradation fingerprint library to provide a scientific basis for subsequent processing; secondly, it implements image enhancement processing through the rule engine intelligent combination processing strategy; then proposes a multi-view feature fusion method based on Transformer to improve detection performance in complex scenarios; finally, it adopts model compression and dynamic jump reasoning technology to achieve efficient operation of the algorithm on edge devices. Compared with the existing technology, the embodiment of the present invention improves the accuracy of small target detection while keeping the original camera hardware unchanged, effectively solving the problem of reduced detection performance caused by poor imaging quality of old cameras, improving image quality and target detection accuracy without replacing hardware, reducing the cost of intelligent property transformation, and providing an economical and efficient solution for intelligent property management.
[0036] Specifically, this embodiment can achieve adaptive processing of cameras of different models and different degrees of aging by constructing an image diagnosis library of camera degradation fingerprints. This allows for improved image quality without replacing hardware. The cost of renovating a single camera is lower than that of replacing the original, achieving low-cost reuse. Secondly, the rule engine can intelligently select processing strategies based on image features, thereby improving the processing effect by more than 20% compared to traditional fixed-process methods in terms of key indicators such as low-resolution reconstruction and noise reduction. Thirdly, multi-view feature fusion technology can reduce cross-view detection consistency errors, significantly improving detection reliability in complex scenarios. Finally, the lightweight deployment solution reduces the resource usage of the algorithm on edge devices, while improving the processing speed of simple scenarios through a dynamic jump mechanism. These technical effects have directly led to three major improvements in property security systems: first, extending the service life of existing cameras and reducing equipment update costs; second, improving the accuracy of abnormal event detection and reducing false alarms and missed reports; and third, achieving real-time edge processing to avoid privacy and bandwidth issues caused by video data transmission.
[0037] In particular, the method provided in this embodiment can enhance the utilization value of older cameras, reduce equipment replacement costs, and improve property management efficiency and service quality. It is particularly suitable for optimizing image quality and lightweight edge computing deployment for heterogeneous cameras (of varying models and age). By building a camera degradation fingerprint library, a rules engine, and multi-view feature fusion technology, adaptive optimization and intelligent analysis of heterogeneous cameras are achieved.
[0038] In one embodiment, the step of constructing the camera degradation fingerprint library includes:
[0039] Based on a preset standard calibration plate, controlling the camera to collect data from the standard calibration plate under preset conditions;
[0040] Based on the data collection results, the blur characteristic of the camera is estimated using a point spread function;
[0041] A noise model is constructed using a mixed Gaussian-Poisson model, and the noise model is used to describe the noise characteristics of the camera;
[0042] Performing color shift detection on the camera based on the standard calibration plate and the data acquisition result;
[0043] The basic information of the camera is obtained, and the camera degradation fingerprint library is constructed by combining blur characteristics, noise characteristics and color shift using a hierarchical storage structure.
[0044] Furthermore, the step S101 includes:
[0045] Using a lightweight convolutional neural network to predict the probability distribution of the input image;
[0046] The result of the probability distribution prediction is compared and mapped with the camera degradation fingerprint library, and the result of the comparison and mapping is output as the image diagnosis result.
[0047] To achieve adaptive enhancement of heterogeneous cameras, this embodiment first conducts quantitative analysis of the imaging degradation characteristics of different camera models, such as blur, noise, and color distortion, to build a camera degradation fingerprint library. This embodiment establishes a mapping relationship between camera degradation characteristics and physical parameters, providing a scientific basis for subsequent dynamic enhancement. Specifically:
[0048] (1) Degradation parameter extraction stage
[0049] During the degradation parameter extraction phase, this embodiment uses a standard calibration panel (consisting of a checkerboard and a 24-color chart) for data acquisition. Each camera captures images during three typical periods: morning, noon, and evening, to cover imaging characteristics under different lighting conditions. A dark-field image (with the lens completely obscured) is also collected for use in building the noise model.
[0050] For the image captured by the camera, the blur characteristics are first estimated using the point spread function. The point spread function (PSF) is estimated using the improved Richardson-Lucy deconvolution algorithm:
[0051] ;
[0052] in, Represents the two-dimensional space coordinates of image pixels, I k Represents the k-th iteration estimated image (i.e., the deblurred estimated image output by the k-th iteration, represented by I obs and the current PSF estimate is iteratively calculated), I k+1 Indicates the k+1th iteration estimated image, I obs is the observed image (i.e. the calibration plate image actually collected by the camera), represents the convolution operation, PSF T is the transpose of the point spread function. This algorithm can accurately estimate the blur characteristics of the camera.
[0053] Secondly, a mixed Gaussian-Poisson model is used for noise modeling:
[0054] ;
[0055] in, Represents the two-dimensional spatial coordinates of the image pixel, I(i,j) refers to the noise distribution of the normal captured image, α is the mixing coefficient, represents a Gaussian distribution, Represents Poisson distribution. The EM algorithm can be used to estimate the parameters α, μ, and λ, thereby establishing a complete noise characterization.
[0056] Next, perform color shift detection in the CIELAB color space and calculate the color difference ΔE between the standard color card and the actual image:
[0057] ;
[0058] in, It is the brightness component difference, which indicates the brightness difference between the standard color card and the measured image. It is the difference between the red and green components, indicating the color shift in the direction of the red and green axis. It is the difference between the yellow and blue hue components, indicating the color shift in the direction of the yellow-blue axis.
[0059] For example, when ΔE>5, it is determined that there is a significant color shift and color correction is required. Specifically, the color correction matrix M can be fitted based on the least squares method:
[0060] ;
[0061] Among them, I obs is the observed image, I ref is the standard reference image;
[0062] (2) Fingerprint database storage stage
[0063] The camera degradation fingerprint library constructed in this embodiment adopts a hierarchical storage structure. Based on the basic information of the camera, blur characteristics, noise characteristics, color shift and other parameter information, the following core data layers can be constructed:
[0064] Basic information layer: metadata about basic camera information such as camera ID, installation location, and acquisition time;
[0065] Optical property layer: PSF matrix, lens distortion parameters; the PSF matrix can be obtained using the PSF function to represent blur characteristics, and the lens distortion parameters can be directly calculated through calibration plate corner detection + nonlinear optimization (Levenberg-Marquardt algorithm);
[0066] Noise characteristic layer: noise model type and parameters;
[0067] Color characteristic layer: white balance gain, color correction matrix; white balance gain is a coefficient calculated for each of the three RGB channels, used to correct color deviations caused by different light source color temperatures. The color correction matrix can be fitted using the least squares method;
[0068] Environmental adaptation layer: Parameter correction coefficients under different lighting conditions; collect images in the morning, afternoon, and evening, analyze the changes in PSF, noise model, and color parameters under different lighting conditions, and generate correction coefficients.
[0069] In practical applications, the camera degradation fingerprint library can be efficiently accessed through the SQLite database and uses B+ tree indexing to ensure fast query. In addition, the data storage logic of the color feature layer is as follows:
[0070] Step 1: Color difference detection
[0071] The ΔE value of each color block in the 24-color card is calculated in the CIELAB space. If there is a color block with ΔE>5, it is determined that there is a significant color shift.
[0072] Step 2: Prioritization
[0073] White balance priority: If only a slight offset is caused by the color temperature of the light source (such as ΔE < 10 and mainly concentrated in the lightness L channel), white balance gain correction is used first;
[0074] Matrix correction: If complex color deviations exist (such as ΔE>10 or cross-channel deviations), the color correction matrix M is calculated and applied.
[0075] Step 3: Cascade Application
[0076] When actually correcting, first apply the white balance gain, then the color correction matrix;
[0077] (3) Online diagnostic module
[0078] To adapt to changes such as camera aging, this embodiment designs a dynamic update mechanism: when the difference between the online diagnosis result and the fingerprint library record exceeds a threshold, parameter recalibration and fingerprint library update are automatically triggered.
[0079] The online diagnosis module uses a lightweight convolutional neural network to perform real-time analysis. The network structure consists of three convolutional layers and two fully connected layers. The input is a 64×64 image block, and the output is the probability distribution of the degradation type:
[0080] ;
[0081] Where x is the input image tensor, y represents the category label of the network output, W and c are the weight matrix and bias vector of the fully connected layer, respectively. The network achieves inference time of less than 5ms on the NVIDIA V100 platform, meeting real-time requirements.
[0082] The diagnostic results include the following key indicators:
[0083] Blur degree (point spread PSF matrix): 0-1 continuous value, based on image gradient amplitude histogram statistics;
[0084] Noise level (hybrid Gaussian-Poisson model parameters): quantified signal-to-noise ratio (SNR) indicator;
[0085] Color shift: ΔE color difference value;
[0086] Lighting conditions: brightness distribution characteristics.
[0087] By establishing a complete camera degradation diagnosis and fingerprint library system, this embodiment provides a reliable parameter basis for subsequent adaptive image enhancement, effectively solving the technical problem of large differences in imaging quality among heterogeneous cameras.
[0088] In one embodiment, step S102 includes:
[0089] Obtaining a dominant defect type of the input image according to the image diagnosis result, and selecting a corresponding processing module from a preset processing module library according to the dominant defect type; wherein the processing module includes a deblurring module, a noise reduction module, and a super-resolution module;
[0090] Using the selected processing module to perform corresponding processing on the input image to obtain an intermediate image;
[0091] The intermediate image is optimized using a small target optimization mechanism to obtain the enhanced image.
[0092] Specifically, the optimizing the intermediate image using a small target optimization mechanism to obtain the enhanced image includes:
[0093] Generate a small target heat map for the intermediate image using a lightweight U-Net network;
[0094] The small target heat map is compared with a preset probability threshold, and local super-resolution enhancement processing is performed on the area in the small target heat map where the probability value is greater than the threshold.
[0095] To achieve intelligent image quality improvement, this embodiment builds a dynamic enhancement pipeline system. This system intelligently analyzes image features and automatically selects the optimal processing strategy, achieving precise repair and quality improvement for different types of image defects. The system architecture comprises three core components: a processing module library, an intelligent rules engine, and small-target optimization, forming a complete closed loop from diagnosis to processing.
[0096] 1. Processing module library
[0097] The processing module library is the foundation of the system and includes a variety of specialized processing modules for different image defects, such as deblurring, noise reduction, and super-resolution. Each module uses advanced algorithms to ensure efficient processing while preserving image details to the greatest extent possible:
[0098] (1) Deblurring module
[0099] The Wiener filter algorithm based on the estimated PSF kernel is used to fix problems such as motion blur and defocus blur. This algorithm balances noise suppression and detail recovery through frequency domain filtering. Its core formula is:
[0100] ;
[0101] Where H(u,v) is the Fourier transform of the blur function, H*(u,v) is its conjugate, F(u,v) is the Fourier transform of the input image, and SNR is the estimated signal-to-noise ratio.
[0102] (2) Noise reduction module
[0103] Adaptive BM3D algorithm is used to automatically adjust parameters according to the noise model to achieve the best noise reduction effect. The algorithm is divided into two stages:
[0104] In the basic estimation stage, preliminary noise reduction results are generated through collaborative filtering :
[0105] ;
[0106] In the final estimation stage, the basic estimation results are combined for refined processing :
[0107] ;
[0108] Among them, z represents the denoised image finally output by the algorithm, x is the tensor of the input image, is the transformation matrix and ξ is the regularization parameter.
[0109] (3) Super-resolution module
[0110] This example uses ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) as the core model for super-resolution enhancement, supporting ×2 / ×4 resolution increases. Its generator consists of 23 stacked RRDB modules, optimized through adversarial training and perceptual loss, ultimately achieving resolution improvement through a pixel shuffle operation. Compared to traditional super-resolution algorithms, ESRGAN offers significant advantages in restoring texture detail, making it particularly suitable for enhancing low-quality images captured by older cameras.
[0111] ESRGAN is based on the RRDB (Residual-in-Residual Dense Block) structure, and its core unit can be expressed as:
[0112] ;
[0113] Where x is the tensor of the input image, It is a combination of dense connection layer (Dense Layer) and LeakyReLU activation function; β is the residual scaling factor (default 0.2) to avoid training instability.
[0114] The expression of the overall generator G is:
[0115] ;
[0116] Among them, x represents the tensor of the input image, PS represents the Pixel Shuffle upsampling layer; RRDB stack Contains 23 RRDB modules to achieve deep feature extraction.
[0117] ESRGAN introduces the relative discriminator relativistic discriminator, whose loss function for:
[0118] ;
[0119] Among them, x r is a real high-resolution image; x fis low-resolution input; avg represents the mean calculation of batch data.
[0120] 2. Rules Engine
[0121] The rule engine is the core decision-making component of the system, which dynamically selects and combines processing modules based on image diagnosis results. Figure 2 and Figure 3 , its workflow is as follows:
[0122] The rule engine receives image diagnosis results, which mainly include: dominant defect type, such as blur, noise, etc.; defect severity, such as quantitative indicators; and auxiliary parameters, such as PSF kernel estimation and noise model parameters. The rule engine dynamically generates a processing pipeline based on the image diagnosis results. Its decision logic is as follows:
[0123] First, the dominant defect type is determined. If blur is dominant (dominant_type == "blur"), Wiener filtering-based deblurring is prioritized (parameters derived from the PSF kernel estimate), followed by texture-optimized resolution enhancement using the ESRGAN super-resolution network. If noise is dominant (dominant_type == "noise"), the adaptive BM3D algorithm is used to suppress noise (parameters derived from the noise model), followed by edge-preserving super-resolution enhancement.
[0124] General post-processing is then performed. Regardless of defect type, small-target local enhancement is mandatory. A heatmap threshold (heatmap_threshold = 0.7) is used to identify key areas and specifically enhance detail clarity. Finally, parameter adaptation is performed. Parameters for each processing module (such as PSF kernel size, noise model coefficients, and super-resolution mode) are dynamically loaded from the diagnosis results, ensuring that the processing strategy is strictly aligned with the camera's degradation characteristics.
[0125] 3. Small target optimization
[0126] To improve the clarity and recognizability of small objects in images, the system introduces a small object optimization mechanism, which consists of two core components:
[0127] (1) Small target heat map generation
[0128] The small target heat map is generated by the lightweight U-Net network. The network structure can be expressed as:
[0129] ;
[0130] Among them, H map It is a heat map, and each pixel value represents the probability of a small target existing at that location. in Represents the input image, that is, the intermediate image.
[0131] (2) Local super-resolution enhancement
[0132] Perform local ×4 super-resolution enhancement on the areas in the heat map where the probability value is greater than the threshold:
[0133] ;
[0134] in, Represents the two-dimensional spatial coordinates of image pixels, represents the image after local super-resolution enhancement, Represents the original image.
[0135] In addition, a small object discriminator is introduced into the super-resolution network training to improve the quality of small object super-resolution through adversarial learning. The adversarial loss function is:
[0136] ;
[0137] Among them, I hr It is a high-resolution ground truth image containing clear small target details, which is used to train the discriminator to identify real small targets. lr The low-resolution input image (Low-ResolutionInput) is obtained by downsampling the high-resolution image and serves as the input of the super-resolution generator. small is the small target discriminator, G is the overall generator, and by minimizing the loss function, the generator can learn a more effective small target super-resolution strategy.
[0138] This dynamically assembled processing pipeline can perform targeted processing based on the actual situation of the image, which not only ensures the processing effect but also improves the processing efficiency, providing an intelligent and flexible solution for improving image quality.
[0139] In one embodiment, step S103 includes:
[0140] Performing spatial registration on the enhanced images using a hierarchical alignment strategy;
[0141] A shared-weight ResNet-18 network is used as a feature encoder, and the feature encoder is used to extract features from enhanced images of different perspectives to obtain corresponding feature maps;
[0142] Calculating attention weights for the feature maps using a cross-attention mechanism, and performing multi-view adaptive fusion based on the attention weights to obtain an initial fused image;
[0143] The initial fused image is subjected to feature fusion according to different feature levels to obtain the fused features.
[0144] To achieve multi-camera collaborative enhancement, this embodiment designs an innovative feature-level fusion solution for camera groups with overlapping fields of view. This solution, through two core technologies, geometric alignment and intelligent feature fusion, overcomes the resolution limitations of a single camera and significantly improves object detection performance in complex scenes.
[0145] 1. Geometric alignment module
[0146] This embodiment adopts a hierarchical alignment strategy to solve the spatial registration problem of multi-view images and achieve accurate geometric alignment. First, the basic homography matrix H0 is calculated based on the camera calibration parameters to establish the initial spatial mapping relationship. The calculation formula is:
[0147] ;
[0148] Where K1 and K2 are the intrinsic parameter matrices of the two cameras, and R represents the rotation matrix. The matrix H0 provides the basis for the initial alignment between images.
[0149] In order to further improve the alignment accuracy, the improved ORB feature matching algorithm is used for fine registration. The ORB algorithm is used to extract the image feature point pairs of the two perspectives. , x i is the feature point of the first image, y i is the feature point of the second image that matches it, and then the robust optimization method is used to solve the optimal transformation matrix H:
[0150] ;
[0151] Among them, H(x i ) represents the feature point x i Mapped to the predicted position in the second image through the optimal transformation matrix H, represents the true position in the second image, and ρ is the Huber loss function, which can effectively suppress the influence of mismatched points, enhance the robustness of the algorithm to noise and outliers, and ultimately achieve sub-pixel alignment accuracy.
[0152] 2. Intelligent feature fusion technology
[0153] This example proposes a Transformer-based feature fusion architecture to achieve complementary advantages and information aggregation of multi-view images. First, a shared-weight ResNet-18 network is used as a feature encoder to extract features from each view image:
[0154] ;
[0155] in, Represents the encoder network, outputting a 256-dimensional feature map fi , to ensure that each perspective feature has consistent semantic expression capabilities, A and B are two different camera perspectives, I i represents the input image of the i-th view.
[0156] The core cross-attention mechanism achieves adaptive fusion of multi-view information by dynamically calculating attention weights. This mechanism can be formalized as:
[0157] ;
[0158] ;
[0159] ;
[0160] Among them, Q represents the query matrix, which is obtained by transforming the feature map f of view A into A Input to the linear mapping layer Generate, used to find the correlation in the feature space of view B. K and V represent the key matrix and value matrix respectively. By transforming the feature map f B Input to the linear mapping layer The last dimension is split and generated. K is used to calculate the attention weights, and V is used to store the information to be aggregated. attn represents the attention weight matrix. It is generated by performing a dot product operation on the query matrix Q and the transpose of the key matrix K and applying the softmax function. It reflects the degree of attention each position in view A pays to each position in view B. is the scaling factor, where Represents the feature dimension, which is used to alleviate the gradient vanishing problem that may be caused by dot product operations. A The original feature map of view A is added to the attention-weighted result through a residual connection to ensure that the original information is not lost. The output represents the final fused feature, which is the output result after aggregating the information of view B through the attention mechanism while retaining the original features of view A.
[0161] This embodiment dynamically aggregates feature information from different perspectives through the attention weight matrix, and retains the original features through residual connections, effectively integrating the advantageous features of each perspective.
[0162] To fully utilize multi-scale feature information, this embodiment implements feature fusion at multiple feature levels (1 / 4, 1 / 8, and 1 / 16 scales) to form a pyramid-like enhancement structure. The final fused feature can be expressed as:
[0163] ;
[0164] in, , Represents the feature maps extracted from view A and view B on the i-th level of the feature pyramid, where i corresponds to different feature scales (e.g., 1 / 4, 1 / 8, 1 / 16). The feature maps at each scale contain semantic information of different granularity. By implementing feature fusion at multiple feature levels, we can fully utilize multi-view information at different scales to form a pyramid-like enhancement structure, thereby improving the accuracy and robustness of object detection. i is the learnable weight, T i Represents the Transformer module of the layer, which integrates the fusion results of each layer by weighted summation.
[0165] The training of this embodiment uses a multi-task loss function for joint optimization:
[0166] ;
[0167] Among them, L represents the total loss function, which is the super-resolution loss L sr , detection loss L det and consistency loss L consist The weighted sum of is used to jointly optimize the entire multi-view feature fusion system. 、 and They represent the weight coefficients of super-resolution loss, detection loss, and consistency loss, respectively, and are used to balance the contribution of different tasks to the final optimization goal.
[0168] Among them, the super-resolution loss L sr A combination of perceptual loss and MSE is used to balance visual quality and pixel-level accuracy:
[0169] ;
[0170] Represents a pre-trained feature extraction network used to extract high-level semantic features of images. sr Represents the image after super-resolution reconstruction. hr represents a high-resolution ground-truth reference image. η is a coefficient that controls the relative weight of perceptual loss and MSE loss.
[0171] Detection loss L det Use the improved Focal Loss to effectively handle the category imbalance problem in target detection:
[0172] ;
[0173] p t Represents the model's predicted probability for the tth sample. It is a regulation factor used to reduce the weight of easy-to-classify samples and make the model pay more attention to difficult-to-classify samples.
[0174] Consistency loss L consist Then ensure the geometric consistency of multi-view enhancement results:
[0175] ;
[0176] and Respectively represent the projection function of projecting the image to view A and view B. AB Represents the homography transformation matrix from view B to view A. B Represents the original image of view B.
[0177] Through a feature-level fusion solution, this invention effectively solves the registration ambiguity problem of traditional image-level fusion, providing reliable technical support for multi-camera collaborative monitoring. This module can be flexibly enabled based on actual scenario requirements, achieving an optimal balance between computing resources and performance requirements.
[0178] In one embodiment, step S105 includes:
[0179] Using a teacher-student network architecture to transfer knowledge of the image object detection model;
[0180] Performing complexity evaluation on the designated image using a complexity evaluator;
[0181] Based on the result of the complexity evaluation, it is determined whether to enable the image target detection model to perform image enhancement processing on the specified image.
[0182] To achieve efficient operation of complex visual processing systems on edge devices, this embodiment proposes a lightweight deployment and adaptive inference solution. Through model compression, dynamic computing optimization, and hardware-related design, it reduces computing overhead while ensuring algorithm accuracy.
[0183] 1. Model compression technology
[0184] Multi-task knowledge distillation. Using a teacher-student network architecture, we transfer the knowledge of the high-performance super-resolution model ESRGAN to the lightweight model TinyESRGAN. We define the distillation loss function as follows:
[0185] ;
[0186] Among them, L distill Represents the knowledge distillation loss function, which is used to measure the difference in feature representation between the teacher network and the student network. It represents the expectation of all input images I that obey the distribution P(I). is the teacher network feature extractor, It is a feature extractor for the student network, which achieves knowledge transfer by minimizing the difference in feature space.
[0187] Shared encoder design. The super-resolution network and the detection network share the first three convolutional layers to reduce repeated calculations. Let the output of the shared layer be ,in is a shared convolutional layer, I represents the input image. The output feature maps are used for low-resolution feature input of the super-resolution task and initial feature extraction of the detection task.
[0188] 2. Dynamic jump reasoning mechanism
[0189] Complexity estimator. Develop an image complexity evaluation function that uses Shannon entropy and edge density to determine whether a full enhancement process needs to be initiated. The Shannon entropy calculation formula is:
[0190] ;
[0191] Where E represents the Shannon entropy of the image, which is used to quantify image complexity. p(i) is the probability of grayscale distribution of the image. The edge density is detected by the Canny operator to determine the ratio of edge pixels:
[0192] ;
[0193] Among them, ρ edge Represents edge density, that is, the ratio of edge pixels to total pixels. Canny(I) represents the result of edge detection on image I using the Canny operator. Represents the L1 norm (i.e., the sum of absolute values), which is used to calculate the number of edge pixels. and Represents the width and height of the image respectively.
[0194] When H<1.5 and ρ edge When the value is less than 0.1 (simple scenario), the enhancement module can be skipped directly, and the image goes directly to the detection network, which increases the inference speed by 2 times.
[0195] Dynamic computation graph: Based on the complexity evaluation results, a conditional computation graph is constructed:
[0196] ;
[0197] Enhancer is the cascade of super-resolution and small object enhancement modules, Detector is the detection network, which can be a lightweight object detection network improved based on YOLOv5s, and H represents the image Shannon entropy. Conditional branching is used to avoid redundant calculations in simple scenarios.
[0198] 3. Performance indicators
[0199] Processing delay: On a V100 graphics card, the single-frame processing flow (diagnosis + enhancement + detection) meets the following delay requirements:
[0200] ;
[0201] in, Indicates the total latency of single frame processing. Indicates the delay of the complexity diagnosis module. Indicates the delay of the enhancement module. This represents the delay of the detection module. The diagnosis module accounts for 20%, the enhancement module accounts for 55%, and the detection module accounts for 25%. This delay can be reduced to less than 20ms in simple scenarios through the dynamic jump mechanism.
[0202] Improved detection performance: On the COCO dataset, the AP@0.5 improvement in small object detection meets the following requirements:
[0203] ;
[0204] in, Indicates the improvement of the average precision (AP@0.5) of small object detection. It shows the detection performance of the method of this embodiment. Denotes the detection performance of the baseline method.
[0205] This embodiment reduces the amount of computation through model compression, avoids redundant operations through dynamic mechanisms, and optimizes parameter calls through hardware binding, thereby achieving efficient deployment of complex visual systems on edge devices and providing a lightweight solution for real-time visual analysis scenarios.
[0206] Figure 4 A schematic block diagram of an image object detection device 400 based on adaptive enhancement provided by an embodiment of the present invention, the device 400 includes:
[0207] The image diagnosis unit 401 is configured to obtain input images captured by cameras with different viewing angles, and perform image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results.
[0208] An image enhancement unit 402 is configured to dynamically select an image processing combination based on the image diagnosis result using a rule engine, and perform image enhancement processing on the input image using the image processing combination to obtain an enhanced image;
[0209] An image fusion unit 403 is configured to perform multi-view feature fusion on the enhanced image using a Transformer-based feature fusion architecture to obtain a fused image;
[0210] An image detection unit 404 is configured to perform target detection on the fused image to obtain corresponding target detection results, thereby constructing an image target detection model;
[0211] The model deployment unit 405 is used to use model compression and dynamic jump reasoning technology to lightweight deploy the image target detection model, and use the deployed image target detection model to perform target detection on the specified image.
[0212] In one embodiment, the image object detection apparatus 400 based on adaptive enhancement further includes:
[0213] A data acquisition unit, configured to control the camera to acquire data from a preset standard calibration plate under preset conditions based on the preset standard calibration plate;
[0214] a blur characteristic estimation unit, configured to estimate blur characteristics of the camera using a point spread function based on a result of data collection;
[0215] a noise characteristic description unit, configured to construct a noise model using a mixed Gaussian-Poisson model, and describe the noise characteristics of the camera using the noise model;
[0216] A color shift detection unit, configured to perform color shift detection on the camera in combination with the standard calibration plate and the result of data acquisition;
[0217] The fingerprint library construction unit is used to obtain basic information of the camera, and to construct the camera degradation fingerprint library by adopting a hierarchical storage structure in combination with blur characteristics, noise characteristics and color shift.
[0218] In one embodiment, the image diagnosis unit 401 includes:
[0219] A probability prediction unit, configured to perform probability distribution prediction on the input image using a lightweight convolutional neural network;
[0220] A comparison and mapping unit is used to compare and map the result of the probability distribution prediction with the camera degradation fingerprint library, and output the result of the comparison and mapping as the image diagnosis result.
[0221] In one embodiment, the image enhancement unit 402 includes:
[0222] a type acquisition unit, configured to acquire a dominant defect type of the input image based on the image diagnosis result, and select a corresponding processing module from a preset processing module library based on the dominant defect type; wherein the processing module includes a deblurring module, a noise reduction module, and a super-resolution module;
[0223] a module processing unit, configured to perform corresponding processing on the input image using a selected processing module to obtain an intermediate image;
[0224] The small target optimization unit is used to optimize the intermediate image using a small target optimization mechanism to obtain the enhanced image.
[0225] In one embodiment, the small target optimization unit includes:
[0226] A heat map generation unit, configured to generate a small target heat map for the intermediate image using a lightweight U-Net network;
[0227] The local enhancement unit is used to compare the small target heat map with a preset probability threshold, and perform local super-resolution enhancement processing on the area in the small target heat map where the probability value is greater than the threshold.
[0228] In one embodiment, the image fusion unit 403 includes:
[0229] An image alignment unit, configured to perform spatial registration on the enhanced image using a hierarchical alignment strategy;
[0230] A feature extraction unit is configured to use a ResNet-18 network with shared weights as a feature encoder, and to extract features from enhanced images of different viewpoints using the feature encoder to obtain corresponding feature maps;
[0231] a weight calculation unit, configured to calculate attention weights for the feature maps using a cross-attention mechanism, and perform multi-view adaptive fusion according to the attention weights to obtain an initial fused image;
[0232] The feature fusion unit is used to perform feature fusion on the initial fused image according to different feature levels to obtain the fused feature.
[0233] In one embodiment, the model deployment unit 405 includes:
[0234] A knowledge transfer unit, configured to transfer knowledge of the image object detection model using a teacher-student network architecture;
[0235] a complexity evaluation unit, configured to perform complexity evaluation on the designated image using a complexity evaluator;
[0236] The model enabling unit is used to determine whether to enable the image target detection model to perform image enhancement processing on the specified image based on the result of the complexity evaluation.
[0237] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0238] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiment. The storage medium may include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other medium capable of storing program code.
[0239] The present invention also provides a computer device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the computer device may also include various network interfaces, a power supply, and other components.
[0240] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0241] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. An image target detection method based on adaptive enhancement, characterized in that: include: Obtain input images captured by cameras with different viewing angles, and perform image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results; Based on the image diagnosis result, dynamically selecting an image processing combination using a rule engine, and performing image enhancement processing on the input image using the image processing combination to obtain an enhanced image; Using a Transformer-based feature fusion architecture to perform multi-view feature fusion on the enhanced image to obtain a fused image; Performing target detection on the fused image to obtain corresponding target detection results, thereby constructing an image target detection model; Lightweight deployment of the image target detection model is performed using model compression and dynamic jump inference technology, and the deployed image target detection model is used to perform target detection on a specified image; The steps of constructing the camera degradation fingerprint library include: Based on a preset standard calibration plate, controlling the camera to collect data from the standard calibration plate under preset conditions; Based on the data collection results, the blur characteristic of the camera is estimated using a point spread function; A noise model is constructed using a mixed Gaussian-Poisson model, and the noise model is used to describe the noise characteristics of the camera; Performing color shift detection on the camera based on the standard calibration plate and the data acquisition result; The basic information of the camera is obtained, and the camera degradation fingerprint library is constructed by combining blur characteristics, noise characteristics and color shift using a hierarchical storage structure.
2. The image target detection method based on adaptive enhancement according to claim 1, characterized in that: The acquiring of input images acquired by cameras with different viewing angles and performing image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results include: Using a lightweight convolutional neural network to predict the probability distribution of the input image; The result of the probability distribution prediction is compared and mapped with the camera degradation fingerprint library, and the result of the comparison and mapping is output as the image diagnosis result.
3. The image target detection method based on adaptive enhancement according to claim 1, characterized in that: The method of dynamically selecting an image processing combination based on the image diagnosis result using a rule engine, and performing image enhancement processing on the input image using the image processing combination to obtain an enhanced image includes: Obtaining a dominant defect type of the input image according to the image diagnosis result, and selecting a corresponding processing module from a preset processing module library according to the dominant defect type; wherein the processing module includes a deblurring module, a noise reduction module, and a super-resolution module; Using the selected processing module to perform corresponding processing on the input image to obtain an intermediate image; The intermediate image is optimized using a small target optimization mechanism to obtain the enhanced image.
4. The image target detection method based on adaptive enhancement according to claim 3, characterized in that: The step of optimizing the intermediate image using a small target optimization mechanism to obtain the enhanced image includes: Generate a small target heat map for the intermediate image using a lightweight U-Net network; The small target heat map is compared with a preset probability threshold, and local super-resolution enhancement processing is performed on the area in the small target heat map where the probability value is greater than the threshold.
5. The image target detection method based on adaptive enhancement according to claim 1, characterized in that: The method of using a Transformer-based feature fusion architecture to perform multi-view feature fusion on the enhanced image to obtain a fused image includes: Performing spatial registration on the enhanced images using a hierarchical alignment strategy; A shared-weight ResNet-18 network is used as a feature encoder, and the feature encoder is used to extract features from enhanced images of different perspectives to obtain corresponding feature maps; Calculating attention weights for the feature maps using a cross-attention mechanism, and performing multi-view adaptive fusion based on the attention weights to obtain an initial fused image; The initial fused image is subjected to feature fusion according to different feature levels to obtain the fused image.
6. The image target detection method based on adaptive enhancement according to claim 1, characterized in that: The lightweight deployment of the image target detection model using model compression and dynamic jump inference technology, and the use of the deployed image target detection model to perform target detection on a specified image, include: Using a teacher-student network architecture to transfer knowledge of the image object detection model; Performing complexity evaluation on the designated image using a complexity evaluator; Based on the result of the complexity evaluation, it is determined whether to enable the image target detection model to perform image enhancement processing on the specified image.
7. An image target detection device based on adaptive enhancement, characterized in that: include: An image diagnosis unit is used to obtain input images captured by cameras with different viewing angles, and perform image diagnosis on the input images using a preset camera degradation fingerprint library to obtain image diagnosis results; The steps of constructing the camera degradation fingerprint library include: Based on a preset standard calibration plate, controlling the camera to collect data from the standard calibration plate under preset conditions; Based on the data collection results, the blur characteristic of the camera is estimated using a point spread function; A noise model is constructed using a mixed Gaussian-Poisson model, and the noise model is used to describe the noise characteristics of the camera; Performing color shift detection on the camera based on the standard calibration plate and the data acquisition result; Obtaining basic information of the camera, and combining blur characteristics, noise characteristics and color shift, using a hierarchical storage structure to construct the camera degradation fingerprint library; an image enhancement unit, configured to dynamically select an image processing combination using a rule engine based on the image diagnosis result, and perform image enhancement processing on the input image using the image processing combination to obtain an enhanced image; An image fusion unit, configured to perform multi-view feature fusion on the enhanced image using a Transformer-based feature fusion architecture to obtain a fused image; An image detection unit is used to perform target detection on the fused image to obtain corresponding target detection results, thereby constructing an image target detection model; The model deployment unit is used to use model compression and dynamic jump reasoning technology to lightweight deploy the image target detection model, and use the deployed image target detection model to perform target detection on the specified image.
8. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for image target detection based on adaptive enhancement according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image target detection method based on adaptive enhancement according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Camera color calibration method suitable for spherical surface camera array
CN101272513A
Low-quality fluorescent microscopic image reconstruction method based on simulation data and learnable descent algorithm
CN120339109A
Cited By
Camera low-light image enhancement method and system based on dynamic exposure fusion deep learning
CN121616506A