An image processing system based on the Internet of Things

By designing an image processing system combining multi-spectral sensors and edge-cloud computing in IoT devices, the problem of image noise and detail loss in low-light environments is solved, and efficient and real-time image quality improvement is achieved.

CN119788941BActive Publication Date: 2025-05-27SHENZHEN JUXIN IMAGE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272283.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

In low-light environments, the image data collected by IoT devices are easily affected by noise, resulting in reduced image clarity and availability. Traditional noise reduction algorithms are prone to loss of details when processing at the edge, while cloud processing cannot achieve real-time optimization due to transmission delay.

Method used

An image processing system based on the Internet of Things is designed, combining multi-spectral sensor integration module, image feature extraction module, edge computing module, cloud computing guidance module, dynamic task scheduling module and edge-cloud feedback closed-loop module to realize adaptive compression and inference optimization, dynamically optimize noise reduction processing, and improve image quality through cross-modal attention units and noise classification networks.

Benefits of technology

Effectively reduce noise in low-light environments, improve image quality, achieve real-time optimization, and automatically adapt in dynamically changing environments to ensure the efficiency and accuracy of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788941B_ABST
    Figure CN119788941B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing system based on the Internet of Things, which relates to the technical field of image processing and includes a multispectral sensor integration module, an image feature extraction module, an edge computing module, a cloud computing guidance module, a dynamic task scheduling module, and an edge-cloud feedback closed-loop module; the multispectral sensor integration module collects image data through the collaborative work of visible light and infrared sensors; the image feature extraction module is used to extract valuable features from the image data; the edge computing module is used to achieve adaptive compression; the cloud computing guidance module classifies the noise types in the image through a noise classification network; the dynamic task scheduling module includes a high-performance mode, a basic mode, and an emergency power consumption reduction mode; the edge-cloud feedback closed-loop module regularly uploads the image quality indicators after noise reduction to the cloud; the present invention effectively solves the problems of image noise and image quality degradation in low-light environments, while maintaining high computing efficiency and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to an image processing system based on the Internet of Things. Background Art

[0002] With the continuous development of the Internet of Things (IoT) technology, more and more IoT devices (such as night surveillance cameras, environmental monitoring devices, etc.) are widely used in various scenarios. These devices collect image and video data through sensors for further analysis and processing.

[0003] In low-light environments, the image data collected by IoT devices (such as night surveillance cameras) is often severely affected by noise, such as Gaussian noise, color distortion, etc. These noises not only affect the clarity and usability of the images, but may also cause deterioration of the image quality, making it impossible for IoT devices to provide high-quality image data when working at night.

[0004] Traditional noise reduction algorithms usually process directly at the edge side, but this method is prone to detail loss, especially the blurring of edges and textures. If processed through the cloud, although it can provide more powerful computing capabilities, due to the existence of transmission delays, real-time optimization cannot be achieved, thus affecting the efficiency and accuracy of the system. Summary of the Invention

[0005] To solve the defects existing in the prior art, the present invention provides an image processing system based on the Internet of Things.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] The present invention provides an image processing system based on the Internet of Things, including a multispectral sensor integration module, an image feature extraction module, an edge computing module, a cloud computing guidance module, a dynamic task scheduling module, and an edge-cloud feedback closed-loop module;

[0008] The multispectral sensor integration module is used to collect image data through the collaborative work of a visible light sensor and an infrared light sensor in a low-light environment;

[0009] The image feature extraction module is used to extract valuable features from the image data collected by the multispectral sensor;

[0010] The edge computing module is used to achieve adaptive compression and inference optimization, and optimize the utilization of computing resources by switching the model structure when the network load is high;

[0011] The cloud computing guidance module classifies the noise types in the image through a noise classification network, generates optimized noise reduction kernel functions according to different types of noise, and pushes these parameters to the edge side, thereby dynamically optimizing the noise reduction process during the operation of the system;

[0012] The dynamic task scheduling module includes a high-performance mode, a basic mode, and an emergency power consumption reduction mode. The dynamic task scheduling module realizes the switching between the high-performance mode, the basic mode, or the emergency power consumption reduction mode according to environmental light intensity, device power, and network latency factors;

[0013] The edge-cloud feedback closed-loop module regularly uploads the image quality indicators after noise reduction to the cloud, and pushes the new model to the edge side through differential update, thereby continuously optimizing the performance of the system.

[0014] As a preferred technical solution of the present invention, the image feature extraction module includes a high-frequency detail enhancement network, a semantic segmentation network, and a cross-modal attention unit;

[0015] The high-frequency detail enhancement network is used to enhance the thermal radiation features captured by the infrared sensor to highlight the details in low-light environments;

[0016] The semantic segmentation network is used to retain the color and texture information captured by the visible light sensor and enhance the semantic performance of the scene;

[0017] The cross-modal attention unit is used to dynamically allocate weights between sensor channels, automatically adjust the fusion ratio of infrared and visible light data according to environmental light intensity, so as to ensure that the details in low-light conditions are not lost, while retaining the color and texture information of the bright areas.

[0018] As a preferred technical solution of the present invention, the fusion formula of the cross-modal attention unit is:

[0019] ;

[0020] where, F fused represents the fused image features, F IR is the image features of the infrared channel, F VIS is the image features of the visible light channel, and α is the fusion weight, which is dynamically adjusted according to environmental light intensity.

[0021] As a preferred technical solution of the present invention, the calculation formula of the fusion weight α in the cross-modal attention unit is:

[0022] ;

[0023] where, I avg is the average light intensity of the scene.

[0024] As a preferred technical solution of the present invention, the edge computing module reduces unnecessary computing amount through a layer pruning strategy, and only retains channels sensitive to infrared features, with a retention rate greater than 85%.

[0025] As a preferred technical solution of the present invention, the formula of the noise reduction kernel function is as follows:

[0026] ;

[0027] where N filter represents the filtered kernel function after noise reduction, which is used for noise suppression in the image, and N base represents the basic noise reduction kernel function, which is used to handle general noise situations, and β represents the dynamic adjustment weight, based on the noise type and ambient light intensity;

[0028] β is determined by the noise classification network and adjusted in the dynamic autoencoder, and is specifically adjusted through the following formula:

[0029] .

[0030] As a preferred technical solution of the present invention, the noise classification network is based on the ResNet-18 architecture, with the input being the frequency domain features of the low-light image and the output being the noise type label;

[0031] The noise reduction kernel function adjusts the size of the kernel function according to different noise types.

[0032] As a preferred technical solution of the present invention, dual-channel full-resolution fusion is enabled when the light intensity is low, and then enters the high-performance mode;

[0033] When the battery power is low or the light intensity is high, only the key areas of the infrared channel are processed to reduce the computing burden, and then enter the basic mode;

[0034] When the network delay is too high, the local cache noise reduction parameters are adopted, and the cloud backhaul link is closed to save bandwidth and battery consumption, and then enter the emergency power consumption reduction mode.

[0035] As a preferred technical solution of the present invention, the edge-cloud feedback closed-loop module uploads the image quality index after noise reduction to the cloud every 10 seconds, and triggers the cloud retraining operation when the index is lower than the set threshold.

[0036] As a preferred technical solution of the present invention, the cloud computing guidance module includes a noise classification network and a dynamic autoencoder;

[0037] The noise classification network identifies the noise type in the image through a pre-trained model and generates corresponding noise reduction parameters;

[0038] The dynamic autoencoder is used to generate an optimized noise reduction kernel function according to the noise type.

[0039] The beneficial effects of the present invention are as follows:

[0040] 1. In the present invention, the combined information of visible light and infrared light is captured by the multispectral sensor integration module. Combining the adaptive compression and inference optimization strategy of the edge computing module, the image data can be quickly processed at the edge, reducing the cloud computing burden. In addition, the cloud computing module accurately classifies and reduces the noise in the image through the noise classification and dynamic optimization strategy, thereby further improving the image quality.

[0041] 2. Through the continuous optimization of the edge-cloud feedback closed-loop module in the present invention, the system can obtain the noise reduction effect in real time and adjust the optimization strategy according to the image quality index, enabling the system to automatically adapt to the dynamically changing environment and ensuring high-quality image processing effects under various conditions. Overall, the system can effectively solve the problems of image noise and image quality degradation in low-light environments while maintaining high computing efficiency and real-time performance, and is applicable to various practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0043] Figure 1 is a schematic flow chart of the image processing system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention and are not intended to limit the present invention.

[0045] Embodiment 1: As Figure 1 shown, an image processing system based on the Internet of Things includes a multispectral sensor integration module, an image feature extraction module, an edge computing module, a cloud computing guidance module, a dynamic task scheduling module, and an edge-cloud feedback closed-loop module;

[0046] The multi-spectral sensor integration module is used to collect image data through the collaborative work of a visible light sensor and an infrared light sensor in low-light environments. The main function of this module is to collect image data. In low-light environments, a single visible light sensor may not provide sufficient information, while the infrared sensor can complement the deficiencies of the visible light sensor. Therefore, this module enables the visible light sensor (wavelength range: 400–700nm) and the infrared sensor (800–1400nm) to work collaboratively. The two sensors achieve spatio-temporal consistent data acquisition (time error < 1ms) through a synchronous control module (FPGA frequency division trigger), avoiding fusion misalignment in dynamic scenarios, and comprehensively utilizing the advantages of both to achieve efficient acquisition of image data in low-light environments. In this way, the quality of images under low-light conditions can be improved, and more scene information can be provided.

[0047] The image feature extraction module is used to extract valuable features from the image data collected by the multi-spectral sensor. The image feature extraction module is mainly used to extract valuable features from the image data collected by the multi-spectral sensor. These features include but are not limited to color, texture, edge information, etc. To ensure that the extracted features have high usability, this module combines a high-frequency detail enhancement network, a semantic segmentation network, and a cross-modal attention unit.

[0048] The edge computing module is used to achieve adaptive compression and inference optimization, and optimize the utilization of computing resources by switching the model structure when the network load is high. The edge computing module is one of the core modules in this system and is mainly used for real-time computing and processing locally. This module dynamically adjusts the use of computing resources through adaptive compression and inference optimization. Especially when the network load is high, it reduces the computational amount by switching the model structure. To improve efficiency, the edge computing module also combines a layer pruning strategy, reduces unnecessary computations, and only retains channels sensitive to infrared features, ensuring that the retention rate is greater than 85%.

[0049] The cloud computing guidance module classifies the types of noise in the image through a noise classification network, generates optimized noise reduction kernel functions according to different types of noise, and pushes these parameters to the edge side, thereby dynamically optimizing the noise reduction process during the operation of the system. The main function of the cloud computing guidance module is to generate optimized noise reduction kernel functions by classifying the types of image noise and push these parameters to the edge side. This process involves a noise classification network and a dynamic autoencoder.

[0050] The dynamic task scheduling module includes a high-performance mode, a basic mode, and an emergency power-saving mode. The dynamic task scheduling module realizes the switching between the high-performance mode, the basic mode, or the emergency power-saving mode according to environmental light intensity, device power, and network latency factors. The main task of this module is to dynamically adjust the working mode of the system according to environmental changes.

[0051] The edge-cloud feedback closed-loop module regularly uploads the noise-reduced image quality metrics to the cloud and pushes the new model to the edge side through differential update, thereby continuously optimizing the performance of the system. This module is mainly used to achieve feedback and closed-loop optimization between the edge and the cloud. The system will regularly upload the noise-reduced image quality metrics to the cloud. When the metrics are lower than the set threshold, the cloud will trigger a retraining operation. In this way, the system can continuously optimize the noise reduction effect and image quality.

[0052] Among them, through a multi-level processing strategy, combining the advantages of edge computing and cloud computing, the present invention realizes efficient image processing in low-light environments. Specifically, the system of the present invention solves the problem of image noise in low-light conditions and optimizes the image quality through the collaborative work of multiple modules such as a multi-spectral sensor integration module, an image feature extraction module, an edge computing module, a cloud computing guidance module, a dynamic task scheduling module, and an edge-cloud feedback closed-loop module. The high efficiency and flexibility of the system make it have broad prospects in practical applications, especially in the fields of night monitoring, security, and intelligent transportation.

[0053] Furthermore, the image feature extraction module includes a high-frequency detail enhancement network, a semantic segmentation network, and a cross-modal attention unit;

[0054] The high-frequency detail enhancement network is used to enhance the thermal radiation features (such as the temperature difference at the edge of metal scratches) captured by the infrared sensor to highlight the details in low-light environments. In low-light environments, the information provided by the infrared sensor mainly includes the thermal radiation features of objects. Although it can help identify the outline and position of objects, its texture and details are usually not as rich as visible light images. By performing high-frequency detail enhancement processing on infrared images, the system can highlight its detail information, making the infrared images clearer and helping to identify small changes in the scene;

[0055] The semantic segmentation network is used to retain the color and texture information captured by the visible light sensor and enhance the semantic representation of the scene. In image processing, semantic segmentation is an important technology. It can segment an image into different semantic regions to help the system understand the meaning of different parts (such as roads, buildings, people, vehicles, etc.). For visible light images, semantic segmentation can not only retain details but also effectively enhance the semantic information of the image, especially in the significant regions of the image;

[0056] The cross-modal attention unit is used to dynamically allocate weights between sensor channels (such as giving priority to the infrared sensor in dark areas and the visible light sensor in bright areas), automatically adjust the fusion ratio of infrared and visible light data according to the environmental light intensity to ensure that details in low-light conditions are not lost while retaining the color and texture information in bright areas. It can adjust according to the environmental light intensity (Iavg Based on the change of (), automatically adjust the weight distribution of the infrared image and the visible light image during the fusion process to optimize the overall performance of the image;

[0057] During the fusion process, the cross-modal attention unit takes into account the advantages and disadvantages of each sensor. For example, in night or low-light environments, the infrared sensor can provide higher object detection capabilities, while the visible light sensor can provide richer color and texture details in bright areas. The cross-modal attention unit automatically adjusts the weights of the infrared and visible light based on the perception of the environmental light conditions (I avg ), enabling the system to flexibly respond according to the specific scenario and achieve the best fusion effect;

[0058] For example, in low-light environments, the infrared image is given a higher weight to highlight details and object contours; while in high-light environments, the weight of the visible light image increases to ensure that the color and texture performance of the scene is not lost.

[0059] This multi-level image processing technology improves the performance of the system in complex environments, enabling imaging devices in the Internet of Things to adapt to different lighting conditions, provide more accurate and detailed visual information, and be widely used in fields such as intelligent security, autonomous driving, and intelligent monitoring.

[0060] Furthermore, the fusion formula of the cross-modal attention unit is:

[0061] ;

[0062] Among them, F fused represents the fused image feature, which is a weighted combination of the infrared image and the visible light image. F IR is the image feature of the infrared channel, F VIS is the image feature of the visible light channel, and α is the fusion weight, which is dynamically adjusted according to the environmental light intensity and determines the proportion of the infrared image and the visible light image in the fusion;

[0063] The calculation formula for the fusion weight α in the cross-modal attention unit is:

[0064] ;

[0065] Among them, I avg is the average light intensity of the scene, which means that when the light is strong, α will become smaller, and when the light is weak, α will increase;

[0066] I avgIt is an important measure of the ambient light conditions in the image, usually obtained by statistically calculating the brightness of the image area. This value directly reflects the light changes in the environment, and further affects the weighting of different sensor data (such as visible light and infrared) during image fusion;

[0067] In low-light environments, when the light intensity in the environment is weak, the value of log(I avg ) will be small, which makes the value of α increase. At this time, the features of the infrared image will be given a greater weight because the infrared sensor can provide more details and object contours in low-light environments, while the weight of the visible light image (1 - α) will decrease accordingly because the visible light sensor performs poorly under such conditions and cannot provide clear image information;

[0068] In high-light environments, when the light intensity in the environment is strong, the value of log(I avg ) is large, which makes the value of α decrease. At this time, the weight of the visible light image (1 - α) increases because the visible light image can provide richer color and texture information, and the weight of the infrared image decreases because in strong light environments, the detail information of the infrared image is usually limited and will not significantly improve the image quality.

[0069] By dynamically adjusting the fusion weights, the cross-modal attention unit can adaptively select the appropriate sensor data for fusion according to different light conditions, so as to obtain the best image effect. This adaptive fusion mechanism makes the image processing system more adaptable and can provide clear, detailed and semantically rich images in different environments.

[0070] The following are several actual application scenarios:

[0071] Intelligent monitoring system: In dimly lit nights or environments with poor lighting, the system can enhance the details of infrared images to ensure the monitoring effect.

[0072] Autonomous driving: In low-light or extreme weather conditions, infrared images can provide the contour information of objects, while in the daytime or under good lighting conditions, the colors and details provided by visible light images will be more helpful for vehicle navigation and recognition.

[0073] Intelligent security: In complex lighting conditions, the system can adaptively adjust the image fusion strategy to improve the recognition ability of intruders, abnormal behaviors, etc.

[0074] Furthermore, the model structure is switched in real time according to the network load. When the bandwidth resources are limited, the edge computing module reduces unnecessary computational amount through the layer pruning strategy and only retains the channels sensitive to infrared features, and the retention rate is greater than 85%. The channels sensitive to infrared features can be retained by the gradient threshold method.

[0075] Among them, layer pruning refers to reducing the consumption of computing resources and accelerating the inference process by removing some redundant neurons, channels, or layers in the neural network. The specific implementation is as follows:

[0076] Channel pruning: Determine which channels can be removed based on the contribution of each channel in the network to the final output. After pruning, retain those channels that have a greater impact on the target task. This method can not only reduce the computational amount but also save storage space;

[0077] Hierarchical pruning: Selectively remove certain layers according to the contribution or importance of the layers, thereby reducing the scale of the model; Generally speaking, pruning these layers will not have a significant impact on the final accuracy, especially when the model retention rate is more than 85%.

[0078] Through layer pruning and channel pruning, the system can dynamically select the channels sensitive to infrared features for retention, ensuring that the key features of the infrared image are retained. By analyzing the feature importance of the infrared image, the system can evaluate which channels and layers are more important for the extraction of infrared features, so as to retain the part sensitive to infrared information during pruning. The system sets the pruning retention rate to be greater than 85% to ensure that important infrared features will not be lost during the pruning process, while significantly improving the inference speed and efficiency.

[0079] Through the combination of the above technical means, the edge computing module can still maintain good performance in a high-load and resource-constrained environment. The dynamic combination of adaptive compression, inference optimization, model structure switching, and layer pruning strategies enables the system to effectively utilize computing resources in different scenarios and achieve the optimal performance.

[0080] Furthermore, the cloud computing guidance module includes a noise classification network and a dynamic autoencoder, which can solve the contradiction between edge computing power fluctuations and environmental adaptability, and optimize the global processing efficiency through cloud model pre-training and adaptive parameter pushing;

[0081] The noise classification network identifies the noise types in the image through a pre-trained model and generates corresponding noise reduction parameters; The noise classification network is a deep learning-based model responsible for analyzing the noise types in the input image. Through the pre-trained model, the network can learn to identify various noise patterns in the image, such as Gaussian noise, salt-and-pepper noise, image compression noise, etc. Specifically, the working process of the noise classification network includes:

[0082] Pre-training stage: Use a large number of image data labeled with noise types to train the model to identify different noise features. During the training process, the model gradually learns the laws and manifestations of noise by extracting local and global features of the image;

[0083] Noise Identification: When inputting the image to be processed, the noise classification network analyzes the image and outputs the corresponding noise type labels. For example, the system may identify the presence of Gaussian noise or compression artifacts in the image;

[0084] Generate Denoising Parameters: Based on the noise type, the network generates appropriate denoising parameters to provide guidance for subsequent denoising operations.

[0085] The dynamic autoencoder is used to generate an optimized denoising kernel function according to the noise type.

[0086] The dynamic autoencoder generates an optimized denoising kernel function (i.e., denoising filter) according to the noise type. An autoencoder is a neural network architecture, usually consisting of an encoder and a decoder, which can compress the input data (here the noisy image) into a low-dimensional representation and then reconstruct the input data through the decoder. In the task of noise reduction, the working process of the dynamic autoencoder is as follows:

[0087] Noise Type Input: The dynamic autoencoder receives the noise type information from the noise classification network;

[0088] Generate Kernel Function: Based on different noise types, the autoencoder generates an optimized denoising kernel function through the trained network parameters, that is, a filter, which is used to suppress the identified noise type in the image;

[0089] Adaptive Adjustment: The autoencoder can dynamically adjust the parameters of the kernel function according to the noise characteristics of the input image to ensure the best noise suppression effect in different environments.

[0090] Among them, the cloud computing guidance module provides important intelligent support in this image processing system, helping the system dynamically identify the noise type and optimize the denoising process through an adaptive algorithm.

[0091] Furthermore, the design of the denoising kernel function is one of the key steps in the image processing system, which directly affects the noise suppression effect of the system. The formula of the denoising kernel function is as follows:

[0092] ;

[0093] Among them, N filter represents the filtered kernel function after denoising, which is used to suppress noise in the image. N base represents the basic denoising kernel function, which is used to handle general noise situations. It may be a conventional filter. β represents the dynamic adjustment weight, depending on the noise type and the ambient light intensity. F fusedRepresents the image features after fusion through the cross-modal attention module. These features contain the combined information of visible light and infrared sensors, not only having rich detailed information but also containing deep semantic information. During the fusion process, the system can organically combine the image information collected by different sensors through the cross-modal attention mechanism, making the image more complete and expressive;

[0094] In practical applications, visible light images and infrared images often complement each other. Visible light images provide rich colors and details, while infrared images provide temperature information and visibility in low-light environments. Through cross-modal fusion, the quality of the image can be improved, especially in complex environments (such as low light, smoke, haze, etc.);

[0095] β is determined by the noise classification network and adjusted in the dynamic autoencoder. The role of β is mainly reflected in dynamically adjusting the fusion weights. According to different lighting conditions, β will affect how the system fuses the feature information of different sensors (such as visible light images and infrared images), and is specifically adjusted through the following formula:

[0096] ;

[0097] In a strong light environment, the value of I avg is relatively high. At this time, the system usually relies more on the data of the visible light sensor because in sufficient light, the detailed information provided by the visible light image is more accurate. At this time, the value of β will increase, indicating that more attention is paid to the features from the visible light sensor during fusion;

[0098] In a low light environment, the value of I avg is relatively low. The system will rely more on the data of the infrared sensor because the infrared image can provide clearer object contours and heat source information in low light conditions. At this time, the value of β will decrease, thereby reducing the weight of visible light features in the fusion process.

[0099] Furthermore, the noise classification network is based on the ResNet-18 architecture. ResNet is a deep convolutional neural network architecture. By introducing residual connections, the problem of gradient disappearance in deep networks can be effectively avoided. The input is the frequency domain features of the low light image (energy distribution after DCT transformation). For low light images, first, the image is transformed from the spatial domain to the frequency domain through the discrete cosine transform (DCT). DCT represents the information of the image in the form of frequency, thereby separating the low frequency and high frequency of the image. The output is the noise type label (Gaussian / impulse / mixed noise). Using the energy distribution after DCT transformation as the input feature can effectively help the network classify different types of noise;

[0100] The noise reduction kernel function adjusts the size of the kernel function according to different types of noise. For example, for impulse noise, median filtering parameters are generated (the dynamic kernel size ranges from 3×3 to 7×7. A smaller kernel size can precisely preserve details, while a larger kernel size can effectively remove noise).

[0101] The size of the noise reduction kernel function can be dynamically adjusted in the following ways:

[0102] Adaptive algorithm: According to the distribution of noise, the size of the kernel function is automatically adjusted. A dynamic filter size range can be set based on factors such as noise intensity and image features. When processing different types of noise, the network can automatically learn the optimal kernel size through forward propagation;

[0103] Multi-scale analysis: Use multi-scale convolutional filters to perform noise reduction processing at multiple scales. This method can better preserve details in the image while removing noise at different scales, and will not be elaborated here.

[0104] Furthermore, when the light intensity is low (e.g., illuminance Lux < 5), dual-channel full-resolution fusion is enabled, and then it enters the high-performance mode; when the battery power is low or the light intensity is high (e.g., Lux ≥ 5 or battery power < 20%), only the key areas of the infrared channel are processed to reduce the computational burden, and then it enters the basic mode; when the network latency is too high (e.g., network latency > 100ms), local cache noise reduction parameters are adopted, and the cloud backhaul link is closed to save bandwidth and battery consumption, and then it enters the emergency power-saving mode.

[0105] In different modes, the adjustments of other modules are as follows:

[0106] In the high-performance mode:

[0107] Multi-spectral sensor integration module. In this mode, due to the low light intensity, the sensor will maximize the collaborative work of the infrared and visible light sensors to collect more image data. At this time, the system will enable dual-channel full-resolution fusion to ensure that details are not lost and improve the image quality in low-light environments;

[0108] Image feature extraction module. It will preferentially use the high-frequency detail enhancement network and semantic segmentation network to process the image, enhance the detail performance and maintain the semantic information. The cross-modal attention unit dynamically adjusts the fusion ratio of infrared and visible light according to the ambient light to ensure that the details in low-light areas are enhanced while retaining the color and texture of the visible light areas;

[0109] Edge computing module. Adaptive compression and inference optimization are enabled to optimize the computing resources to meet the higher computing requirements. Through model structure switching, when resources are sufficient, the efficiency and accuracy of image processing are guaranteed;

[0110] Cloud computing guidance module. In this mode, due to the sufficient network transmission bandwidth, the cloud can perform real-time optimization and noise reduction processing, generate an optimized noise reduction kernel function according to the noise classification network, and push it to the edge side for real-time noise reduction optimization;

[0111] Dynamic task scheduling module. At this time, the system will be in the high-performance mode, giving priority to ensuring the quality of image processing, and the requirements for power and network latency are relatively loose.

[0112] In the basic mode:

[0113] Multispectral sensor integration module. At this time, the system will focus on processing the data of the key areas of the infrared sensor, avoiding full-scale high-load calculations, only focusing on the features of the infrared image, and reducing the amount of calculation;

[0114] Image feature extraction module. It focuses on processing the infrared data, enhancing the thermal radiation features in low-light environments, reducing complex semantic segmentation and cross-modal attention calculations, and highlighting the details of the infrared channel;

[0115] Edge computing module. By optimizing the inference process, adapting to lower computing resources, and using less computational volume and a simple model structure to process infrared images;

[0116] Cloud computing guidance module. Since the system reduces the amount of calculation for edge-side processing, the cloud can perform more refined noise reduction processing on the image in a timely manner. At this time, the noise reduction strategy will mainly rely on the preset noise reduction kernel function for cloud-guided processing;

[0117] Dynamic task scheduling module. The system is in the basic mode according to the power or network latency situation, and the battery consumption and computational burden are relatively low.

[0118] In the emergency power consumption reduction mode:

[0119] Multispectral sensor integration module. In this mode, in order to reduce the computational burden, the system will only process the key areas of the infrared channel, avoiding full-scale data processing in the case of low power and high network latency. The acquisition of the visible light sensor will be restricted when necessary;

[0120] Image feature extraction module. In order to reduce the processing complexity, it will focus on the feature extraction of the key areas of the infrared channel, and other feature extraction tasks will be delayed or cancelled;

[0121] Edge computing module. At this time, the edge computing module will greatly reduce the amount of calculation through pruning and other strategies, give priority to retaining the channels sensitive to infrared features, reduce unnecessary calculations, and enable the most simplified model to maintain the basic operation of the system;

[0122] Cloud computing guidance module. In the emergency power consumption reduction mode, real-time noise reduction processing is no longer performed on the cloud side. Instead, the model is simply guided by the uploaded image quality metrics only when necessary. The push frequency of the noise reduction kernel function will be reduced, and the network burden will be compressed to the minimum.

[0123] Dynamic task scheduling module. The system reduces the data upload frequency according to factors such as ambient light intensity, device power, and network latency in the emergency power consumption reduction mode, minimizes local computing tasks, and stops the cloud backhaul link simultaneously to avoid additional bandwidth and battery consumption.

[0124] It should be noted that regardless of the mode, the edge-cloud feedback closed-loop module will regularly upload the noise-reduced image quality metrics to the cloud and push the new model to the edge side through differential update. Under different modes, the upload frequency and update mechanism will be adjusted according to the power, network status, and system operation requirements.

[0125] Furthermore, the edge-cloud feedback closed-loop module uploads the noise-reduced image quality metrics to the cloud every 10 seconds and triggers the cloud retraining operation when the metrics are lower than the set threshold.

[0126] Specifically, the image quality metrics can be determined using the structural similarity index (SSIM) or the peak signal-to-noise ratio (PSNR). SSIM is an index to measure the structural similarity between two images, mainly used to measure the similarity between the original image and the compressed or processed image. PSNR is a traditional index to measure image quality, mainly used to evaluate the quality of compressed images.

[0127] In this embodiment, the specific triggering conditions are as follows:

[0128] When SSIM is less than 0.85, it indicates that the structural similarity of the image has decreased and the noise reduction effect is not ideal, and retraining needs to be triggered to improve the model.

[0129] Or when the change in PSNR (ΔPSNR) is greater than 2dB, it indicates that the quality of the image has changed significantly. The greater the change in PSNR, the greater the decline in image quality, which also means that the noise reduction model needs to be optimized.

[0130] It should be noted that after the retraining is triggered, the new model will be pushed to the edge side through differential update technology. Differential update technology usually means only transmitting the incremental part of the model instead of the whole model. This way can greatly reduce the transmission bandwidth and storage consumption. The above model increment is about 3MB, which means that only the updated part of the model needs to be transmitted each time instead of the complete model, thus improving efficiency.

[0131] The system operation process is as follows:

[0132] Data Acquisition and Transmission: The multispectral sensor integration module first acquires image data in low-light environments. After preprocessing, this data is transmitted to the image feature extraction module;

[0133] Feature Extraction and Optimization: The image feature extraction module processes the transmitted data, extracts key features, and dynamically adjusts the fusion ratio of infrared and visible light data through a cross-modal attention unit to ensure the retention of details in low-light environments;

[0134] Edge Computing: In the edge computing module, the image data is adaptively compressed and inference-optimized according to the real-time computing resource situation, reducing the computing burden and dynamically switching the model structure when needed;

[0135] Cloud Guidance: The cloud computing guidance module generates a noise reduction kernel function based on the noise type in the image and optimizes it through a noise classification network and a dynamic autoencoder, pushing the generated noise reduction parameters to the edge side;

[0136] Task Scheduling and Mode Switching: The dynamic task scheduling module automatically adjusts the operating mode of the system according to environmental changes (such as light intensity, battery power, network latency, etc.) to ensure optimal performance in various scenarios;

[0137] Feedback Closed-loop and Continuous Optimization: In the edge-cloud feedback closed-loop module, the system uploads the quality index of the noise-reduced image every 10 seconds. When the index is lower than the set threshold, the cloud triggers a retraining operation to further optimize the system performance.

[0138] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image processing system based on the Internet of Things, characterized in that: It includes a multispectral sensor integration module, an image feature extraction module, an edge computing module, a cloud computing guidance module, a dynamic task scheduling module, and an edge-cloud feedback closed-loop module; The multi-spectral sensor integrated module is used to collect image data in a low-light environment through the collaborative work of a visible light sensor and an infrared light sensor; The image feature extraction module is used to extract features from image data collected by the multispectral sensor; The edge computing module is used to achieve adaptive compression and inference optimization, and optimize the utilization of computing resources by switching the model structure when the network load is high; The cloud computing guidance module includes a noise classification network and a dynamic autoencoder; The cloud computing guidance module classifies the noise types in the image through a noise classification network, and generates an optimized denoising kernel function according to different types of noise; The noise classification network identifies the noise type in the image through a pre-trained model and generates corresponding noise reduction parameters; These parameters are pushed to the edge to dynamically optimize the noise reduction process during system operation. The dynamic task scheduling module includes a high-performance mode, a basic mode and an emergency power-saving mode. The dynamic task scheduling module realizes switching between the high-performance mode, the basic mode or the emergency power-saving mode according to the ambient light intensity, the device power and the network delay factors; The edge-cloud feedback closed-loop module regularly uploads the denoised image quality indicators to the cloud, and pushes the new model to the edge through differential updates, thereby continuously optimizing the performance of the system.

2. The image processing system based on the Internet of Things according to claim 1, characterized in that: The image feature extraction module includes a high-frequency detail enhancement network, a semantic segmentation network, and a cross-modal attention unit; The high-frequency detail enhancement network is used to enhance the thermal radiation features captured by the infrared sensor to highlight the details in a low-light environment; The semantic segmentation network is used to retain the color and texture information captured by the visible light sensor and enhance the semantic representation of the scene; The cross-modal attention unit is used to dynamically allocate weights between sensor channels and automatically adjust the fusion ratio of infrared and visible light data according to the ambient light intensity to ensure that details are not lost in low-light conditions while retaining the color and texture information of bright areas.

3. The image processing system based on the Internet of Things according to claim 2, characterized in that: The fusion formula of the cross-modal attention unit is: ; Among them, F fused represents the fused image features, F IR is the image feature of the infrared channel, F VIS is the image feature of the visible light channel, and α is the fusion weight, which is dynamically adjusted according to the ambient light intensity.

4. The image processing system based on the Internet of Things according to claim 3, characterized in that: The calculation formula of the fusion weight α in the cross-modal attention unit is: ; Among them, I avg is the average light intensity of the scene.

5. The image processing system based on the Internet of Things according to claim 1, characterized in that: The edge computing module reduces unnecessary computation through a layer pruning strategy and only retains channels that are sensitive to infrared features, with a retention rate greater than 85%.

6. The image processing system based on the Internet of Things according to claim 4, characterized in that: The dynamic autoencoder is used to generate an optimized denoising kernel function according to the noise type.

7. The image processing system based on the Internet of Things according to claim 6, characterized in that: The formula of the denoising kernel function is as follows: ; Among them, N filter Represented as the filter kernel function after denoising, used for noise suppression in images, N base It is represented as the basic denoising kernel function, which is used to process Gaussian noise, salt and pepper noise, and image compression noise. β is represented as the dynamic adjustment weight according to the noise type and ambient light intensity. β is determined by the noise classification network and adjusted in the dynamic autoencoder using the following formula: 。 8. The image processing system based on the Internet of Things according to claim 7, characterized in that: The noise classification network is based on the ResNet-18 architecture, with the input being the frequency domain features of low-light images and the output being the noise type label; The denoising kernel function adjusts the size of the kernel function according to different noise types.

9. The image processing system based on the Internet of Things according to claim 1, characterized in that: When the light intensity is low, dual-channel full-resolution fusion is enabled to enter high-performance mode; When the battery is low or the light intensity is high, only the key areas of the infrared channel are processed to reduce the computing burden and enter the basic mode; When the network delay is too high, the local cache noise reduction parameters are used and the cloud backhaul link is closed to save bandwidth and battery consumption, thus entering the emergency power reduction mode.

10. The image processing system based on the Internet of Things according to claim 1, characterized in that: The edge-cloud feedback closed-loop module uploads the denoised image quality index to the cloud every 10 seconds and triggers a cloud retraining operation when the index is lower than the set threshold.

Citation Information

Patent Citations

  • Image enhancement processing method and device, computer equipment and medium

    CN114463223A

  • Vehicle-mounted reversing image system of heavy truck

    CN118521797A