An image processing method and system for industrial environments

By using an image enhancement network framework and lightweight algorithms to process low-light images in industrial environments, the problem of low image quality caused by poor lighting conditions is solved, achieving efficient safety behavior detection and early warning, and reducing system costs.

CN117710267BActive Publication Date: 2026-05-19CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2023-12-15
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing real-time monitoring technologies suffer from low image quality in industrial environments due to poor lighting conditions and uneven light distribution, affecting recognition accuracy and hindering the rapid and efficient detection and early warning of dangerous behaviors.

Method used

An image enhancement network framework is adopted, including an initialization module, an optimization module, and an image enhancement module. The initial reflectivity and illumination are obtained by training on low-light images. The reflection layer and illumination layer are iteratively refined to generate target illumination, which is then input into the image enhancement module for image enhancement. Combined with a lightweight algorithm and a self-correction module, the video quality is rapidly improved.

Benefits of technology

It improves the accuracy of image recognition in industrial environments, reduces false alarm rates, enables rapid and efficient detection and early warning of security behaviors, reduces system costs, and is suitable for scenarios with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117710267B_ABST
    Figure CN117710267B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of safety monitoring in an industrial environment, and discloses an image processing method and system for an industrial environment, which trains an input image enhancement network framework by taking a weak-light image in a video frame as the input image; in the training, the weak-light image is input into an initialization module to obtain initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, the initial illumination and normal illumination are input into an optimization module to iteratively refine a reflection layer and an illumination layer, so that target illumination is obtained; the target illumination is input into an image enhancement module to obtain an output enhanced image; and an image frame of the industrial environment collected in real time is input into the image enhancement model to obtain an enhanced image. The method can solve the problem that, due to poor lighting conditions and excessively high dust and smoke concentration in some industrial environments, the real-time monitoring technology causes low video image quality of a monitoring video, thereby affecting the final recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety monitoring technology in industrial environments, and in particular to an image processing method and system for industrial environments. Background Technology

[0002] In recent years, with the development of information technology in "high-tech mining," the intelligence level of underground digital video monitoring systems has been continuously improving, placing higher demands on safety management in poorly lit and harsh industrial environments. For areas where smoking is prohibited, flammable and explosive materials are present, areas requiring safety helmets, and areas prone to falls, precise monitoring methods are essential to promptly detect violations and maintain safety. In the development of modern science and technology, computer vision technology has become a popular field, widely applied in image recognition, autonomous driving, and security monitoring. Real-time monitoring technology, a branch of computer vision, can process, analyze, and identify images in a scene to achieve real-time monitoring, diagnosis, and early warning functions. In security monitoring, real-time monitoring technology, with its high efficiency, accuracy, and automation, has been successfully applied in various situations, becoming an important means of protecting people's lives and property.

[0003] However, real-time monitoring technology still has some limitations in practical applications. For example, in industrial environments, poor lighting conditions and uneven light distribution can lead to low image quality and thus low recognition accuracy. These problems prevent existing real-time monitoring technologies from quickly and efficiently detecting and warning of dangerous behaviors, dangerous situations, and distress calls in industrial environments, and from reliably preventing the occurrence of safety hazards. Summary of the Invention

[0004] This invention provides an image processing method and system for industrial environments to solve the problems in existing technologies, such as poor lighting conditions and uneven light distribution in industrial environments leading to low image quality and consequently low recognition accuracy.

[0005] To achieve the above objectives, the present invention employs the following technical solution:

[0006] In a first aspect, the present invention provides an image processing method for industrial environments, comprising:

[0007] S1: Capture video frames in an industrial environment;

[0008] S2: The low-light images in the video frames are used as input to train the image enhancement network framework to obtain the final image enhancement model. The image enhancement network framework includes an initialization module, an optimization module, and an image enhancement module. During training:

[0009] The low-light image is input into the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input into the optimization module to iteratively refine the reflection layer and illumination layer to obtain the target illumination; the target illumination is input into the image enhancement module to obtain the output enhanced image.

[0010] S3: Input the image frames of the industrial environment collected in real time into the image enhancement model to obtain the enhanced image.

[0011] Secondly, this application also provides an image processing system for industrial environments, comprising:

[0012] Monitoring devices are used to collect video frames in industrial environments;

[0013] An edge computing server is used to read video frames from industrial environments collected by monitoring devices. Low-light images from these video frames are used as input to train an image enhancement network framework to obtain the final image enhancement model. This network framework includes an initialization module, an optimization module, and an image enhancement module. During training: the low-light image is input to the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input to the optimization module to iteratively refine the reflection and illumination layers to obtain the target illumination; the target illumination is input to the image enhancement module to obtain the output enhanced image; real-time collected image frames from the industrial environment are input to the image enhancement model to obtain the enhanced image; the enhanced image is also used to resize and input into a target detection model to identify target objects in the enhanced image, determine relevant information about the target objects, including location coordinates, category information, and confidence values; and based on the relevant information about the target objects, behavioral warnings are generated to obtain alert information.

[0014] A visual interface device, which is used to receive and display alert information from an edge computing server;

[0015] Both the monitoring device and the visualization interface device are connected to the edge computing server.

[0016] Thirdly, this application provides an image processing system for industrial environments, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect above.

[0017] Beneficial effects:

[0018] The image processing method for industrial environments provided by this invention uses low-light images from video frames as input to train an image enhancement network framework to obtain the final image enhancement model. The image enhancement network framework includes an initialization module, an optimization module, and an image enhancement module. During training: the low-light image is input to the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input to the optimization module to iteratively refine the reflection layer and illumination layer to obtain the target illumination; the target illumination is input to the image enhancement module to obtain the output enhanced image; and real-time acquired image frames from the industrial environment are input to the image enhancement model to obtain the enhanced image. In this way, by enhancing the image, the problem of low-quality video images acquired by existing real-time monitoring technologies due to poor lighting conditions and high dust and smoke concentrations in some industrial environments can be solved, thus affecting the final recognition results.

[0019] In a further solution, a self-calibration module was used to improve the model, making it lighter and enabling faster and more efficient improvement of video quality. Attached Figure Description

[0020] Figure 1 This is a flowchart of an image processing method for industrial environments, which is a preferred embodiment of the present invention. Detailed Implementation

[0021] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "an" or "a" and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "connected" or "linked" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship also changes accordingly.

[0023] It should be understood that, based on Retinex theory, the observed image can be divided into an illumination layer and a reflection layer. In this application, a network is designed to decompose the reflection layer and the illumination layer, and adaptively enhance the illumination layer. At the same time, a self-correction module is introduced into it to achieve the final image enhancement.

[0024] It is worth noting that this application is applicable to the scenario of mines, where there are many low-light images due to poor lighting.

[0025] Please see Figure 1 This application provides an image processing method for industrial environments, comprising:

[0026] S1: Capture video frames in an industrial environment;

[0027] S2: The low-light images from the video frames are used as input to train the image enhancement network framework, resulting in the final image enhancement model. The image enhancement network framework includes an initialization module, an optimization module, and an image enhancement module. During training:

[0028] The low-light image is input into the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input into the optimization module to iteratively refine the reflection layer and illumination layer to obtain the target illumination; the target illumination is input into the image enhancement module to obtain the output enhanced image.

[0029] S3: Input the real-time acquired image frames of the industrial environment into the image enhancement model to obtain the enhanced image.

[0030] In this embodiment, acquiring video frames in an industrial environment can be achieved by acquiring video streams from an industrial environment using a camera or similar device, and then obtaining video frames to determine low-light images.

[0031] In this embodiment, taking a mine as an example, the method for collecting video frames in an industrial environment can be as follows: a monitoring device is set up within a preset area, and video stream data within the preset area is collected by the monitoring device. Video frames are then obtained based on the video stream data. Specifically, the monitoring device can be a terminal device and a camera used to collect video stream data.

[0032] Low-light images, also known as low-quality images, refer to images whose quality does not meet the set requirements due to uneven lighting or low illumination.

[0033] The image processing method described above for industrial environments can enhance images to solve the problem that existing real-time monitoring technologies suffer from low-quality video images due to poor lighting conditions and high dust and smoke concentrations in some industrial environments, which in turn affects the final recognition results.

[0034] The following is a detailed step-by-step description of the above-mentioned image processing method for industrial environments, using a complete example:

[0035] 1. Deploy edge server data: Pre-define information areas with uneven lighting distribution and low illumination in the mine within the edge server;

[0036] In a preferred embodiment of the present invention, specific warning signs are set up in the defined area, which can better enhance the image quality in the images extracted by the camera later. The appearance of the warning signs can also help the target detection model to more accurately classify abnormal behavior categories.

[0037] 2. Collect video stream data: Set up monitoring devices in a preset area, collect video stream data in the preset area through the monitoring devices, and send the video stream data to the edge computing server;

[0038] 3. Enhancement of images in mines with uneven lighting distribution and low illumination: The edge computing server decomposes the received video stream data frame by frame to obtain image information; after resizing the image information, it is input into a lightweight recognition-oriented low illumination image enhancement algorithm model to obtain the enhanced image;

[0039] The received video stream data is decomposed frame by frame to obtain image information; the lightweight, recognition-oriented industrial environment image enhancement algorithm is the SCIU algorithm, which consists of the following four steps:

[0040] Obtain the raw video frames;

[0041] The initial reflectance and illumination are generated by passing the low-light image from the video frame to the initialization module in the algorithm.

[0042] The optimization module iteratively refines the reflection layer and the illumination layer;

[0043] Optimize the lighting stage separately and adjust and enhance the lighting according to the user-defined enhancement ratio w.

[0044] The initial reflectivity and illumination are generated. The initialization module consists of three Conv+LeakyReLU layers, followed by convolutional layers and ReLU layers. The kernel size of the entire convolutional layer is set to 3. 3. The loss function expression is as follows:

[0045] ;

[0046] Among them, the first item It is the reconstruction loss, the second item. To maintain the overall structure of I, I represents the low-light image. Let μ be the initial reflectivity, and µ be a hyperparameter. It initializes the lighting and C∈{R,G,B} represents the RGB channels.

[0047] The reflection and illumination layers are iteratively refined. First, to avoid explicit regularization, illumination and reflectivity are adaptively recovered using deep learning, with the reflectivity of the normal light image as a reference. A structure-aware smoothing constraint is applied to the illumination of the normal light image, and then the loss function of the normal light image is decomposed as follows:

[0048] ;

[0049] in, , and These represent the image under normal light, the reflection of the image under normal light, and the illumination of the image under normal light, respectively. e and It's a hyperparameter. This is normal light image gradient calculation. This is the gradient calculation for normal light image illumination. F Represented as matrix norm;

[0050] Two networks, denoted as GL and GR, are introduced to update L and R respectively. GL is a simple fully convolutional network with five CONV layers, which learns the implicit prior on L through ReLU activation. Its loss function is designed as the sum of the loss functions of reflectance and illuminance, as expressed below:

[0051] ;

[0052] In the formula, T represents the total number of stages. This represents the auxiliary approximation of normal light reflection variables. This represents the normal light reflection variable. , , , All represent equilibrium parameters. Represents the illumination layer variable that assists in approximating normal light. This represents the normal illumination layer variable. Indicates the illumination at each stage, Represents the reflectivity at each stage. Represents the reflectance of a normal light image. ( ) represents feature extraction, and SSIM represents structural similarity loss.

[0053] This includes each stage and Mean square error (MSE) loss, MSE loss, structural similarity loss, and perceived loss between the final restored reflectance (RT) and each stage and MSE loss between stages, and each stage Total change loss.

[0054] In the illumination-stage optimization process, which adjusts and enhances the lighting, a relationship exists between the low-light image y and the desired sharp image z, according to Retinex theory. Illumination is typically considered a key component, and the primary focus of optimization is low-light image enhancement. Based on this theory, the estimated illuminance can be further used to obtain the enhanced output. This is achieved by introducing a mapping with parameter θ. To learn about lighting, we can model this task from a progressive perspective, with the basic unit written as:

[0055] ;

[0056] in, The residuals at stage t can be calculated in a way that greatly reduces computation and maintains stability, especially for exposure control. For illumination estimation network, and Regardless of the number of stages, the illumination estimation network maintains a shared structure and parameters at each stage. Let y represent the illumination at stage t, and y represent the low-light image.

[0057] In this application, the image enhancement process is divided into N stages, where N is a positive integer.

[0058] Generate basic units for image enhancement at each stage;

[0059] In each stage, the process of image enhancement is transformed into learning the residual between illumination and low-light observations for the parameter θ.

[0060] It's worth noting that the basic unit is a stage used to optimize the image. The basic unit enhances illumination in a staged manner. In this application, the progressive approach refers to enhancing the image using residuals. Compared to using a direct mapping between low-light observation and illumination, learning residual representations significantly reduces computational complexity, ensuring both performance and improved stability, especially for exposure control. This method greatly reduces computational load compared to directly enhancing the entire image, making it suitable for scenarios where computational resources and capabilities are significantly reduced in industrial environments.

[0061] Furthermore, this application introduces a self-calibration framework S to represent the difference between each level of input and the first level of input. The self-calibration module ensures that the outputs at different stages of the training process converge to the same state. The basic unit of the illumination optimization process is reformulated as:

[0062] ;

[0063] Where y is the low-light image, Z is the target image, S is the self-correcting map, and K is the target image. For parameterized operators, The parameters are learnable; Vt is the calibrated input used for the next stage; it is the basic unit of the illumination optimization process.

[0064] ;

[0065] In this way, the self-calibration module can reduce the computational load throughout the image enhancement process, making it suitable for scenarios where computing resources and capabilities are significantly reduced in industrial environments.

[0066] The illumination is adjusted according to a user-defined enhancement ratio w; the expression for w is:

[0067] ;

[0068] in, The lighting is for a realistic image. The lighting conditions are shown in the low-light diagram.

[0069] In a preferred embodiment of this invention, the UREtinex-Net neural network, comprising three learnable modules—an initialization module, an expansion and optimization module, and an illumination adjustment module—is used to enhance low-quality images transmitted from monitoring areas in industrial environments with poor lighting and high concentrations of smoke and dust. This method avoids the problems of relying on manually created priors and the time-consuming optimization process. It also introduces a self-calibration module to accelerate inference speed and solves the problem of requiring high-precision, fast, and lightweight models in industrial environments with numerous network cameras simultaneously transmitting multiple videos awaiting security checks.

[0070] 4. Target and behavior prediction: The image information obtained by the lightweight low-light image enhancement algorithm model for recognition is resized and then input into the target detection algorithm-based model to predict the target and behavior, and obtain the position coordinates, category information and confidence value of each prediction box;

[0071] In step 4, the image information size is adjusted to p p, where p is an integer multiple of 32; the target detection algorithm is the YOLOv5 algorithm;

[0072] The position coordinates of the predicted bounding box are (x, y, w, h), where x and y represent the coordinates of the center point of the predicted bounding box, and w and h represent the length and width of the predicted bounding box; the confidence score is defined and calculated according to the following relationship:

[0073] ;

[0074] in, This represents the confidence level of the j-th bounding box in the i-th cell. This indicates the probability that there is an object in the current box. This represents the IOU ratio between the actual detection bounding boxes and the predicted detection bounding boxes;

[0075] For each grid cell in the image, predict C conditional class probabilities simultaneously: ;

[0076] The probability of a certain type appearing in the bounding box and the degree of fit of the predicted bounding box to the target satisfy the following relationship:

[0077] ;

[0078] in, This represents the i-th category.

[0079] The YOLOv5 algorithm uses the CIOU_Loss loss function to calculate the overlap area between the detection box and the target box. The CIOU_Loss loss function is calculated according to the following formula:

[0080] ;

[0081] Where IOU is the ratio of the intersection to the union of the detection box and the target box; Distance_C is the diagonal distance of the minimum bounding rectangle; Distance_2 is the Euclidean distance between the center point of the detection box and the center point of the ground truth box; v is a parameter that measures the consistency of aspect ratio, and v satisfies the following relationship:

[0082] ;

[0083] in, and These are the width and height of the actual frame; and These are the width and height of the detection frame.

[0084] In a preferred embodiment of the present invention, the optimal anchor box values ​​are adaptively calculated for different training sets during each training iteration. In the YOLOv5 algorithm, initial anchor boxes with set widths and heights are provided for different datasets. During network training, the network outputs predicted boxes based on the initial anchor boxes, compares them with the ground truth boxes, calculates the difference, and then updates the network parameters iteratively. Ultimately, the algorithm calculates the most suitable anchor box size for the current dataset, thereby ensuring optimal training quality.

[0085] 5. Output safety behavior and personnel behavior monitoring results: After the processing in step 4, the model based on the target detection algorithm outputs the predicted target and target behavior results, provides timely early warning, and records and saves the results for statistical analysis. In a preferred embodiment of the present invention, all detection results are recorded to facilitate supervision and future early warning.

[0086] This invention presents a lightweight, recognition-oriented low-light image enhancement method for industrial environments. Compared to traditional monitoring systems, this method employs edge computing devices and lightweight algorithms for image enhancement and computational recognition. This enables rapid and efficient detection and early warning of safety and personnel behaviors in industrial environments, improving warning accuracy, reducing false alarm rates, effectively preventing safety hazards, and is applicable to complex scenarios, thus enhancing system accuracy and reliability. Furthermore, the edge computing device design reduces the performance and storage requirements of terminal devices, lowering system costs and facilitating system use and management. In industrial environments, where numerous network cameras typically transmit multiple videos simultaneously for security checks, a lighter model is essential to avoid decreased recognition accuracy and resource waste. This method utilizes a self-calibration module to improve the model's lightweight nature, enabling rapid and efficient improvement of video quality.

[0087] This application also provides an image processing system for industrial environments, comprising:

[0088] Monitoring devices are used to collect video frames in industrial environments;

[0089] An edge computing server is used to read video frames from industrial environments collected by monitoring devices. Low-light images from these video frames are used as input to train an image enhancement network framework to obtain the final image enhancement model. This network framework includes an initialization module, an optimization module, and an image enhancement module. During training: the low-light image is input to the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input to the optimization module to iteratively refine the reflection and illumination layers to obtain the target illumination; the target illumination is input to the image enhancement module to obtain the output enhanced image; real-time collected image frames from the industrial environment are input to the image enhancement model to obtain the enhanced image; the enhanced image is also used to resize and input into a target detection model to identify target objects in the enhanced image, determine relevant information about the target objects, including location coordinates, category information, and confidence values; and based on the relevant information about the target objects, behavioral warnings are generated to obtain alert information.

[0090] A visual interface device, which is used to receive and display alert information from an edge computing server;

[0091] Both the monitoring device and the visualization interface device are connected to the edge computing server.

[0092] This image processing system for industrial environments can implement the various embodiments of the above-described image processing methods for industrial environments and achieve the same beneficial effects, which will not be elaborated here.

[0093] This application also provides an image processing system for industrial environments, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the above-described method. This image processing system for industrial environments can implement various embodiments of the above-described image processing method for industrial environments and achieve the same beneficial effects; further details are omitted here.

[0094] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. An image processing method for industrial environments, characterized in that, include: S1: Capture video frames in an industrial environment; S2: The low-light images in the video frames are used as input to train the image enhancement network framework to obtain the final image enhancement model. The image enhancement network framework includes an initialization module, an optimization module, and an image enhancement module. During training: The low-light image is input into the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input into the optimization module to iteratively refine the reflection layer and illumination layer to obtain the target illumination; the target illumination is input into the image enhancement module to obtain the output enhanced image. S3: Input the real-time acquired image frames of the industrial environment into the image enhancement model to obtain the enhanced image; The step of inputting the initial reflectivity, initial illumination, and normal illumination into the optimization module to iteratively refine the reflection layer and illumination layer to obtain the target illumination includes: A structure-aware smoothing constraint is applied to the illumination of the normal lighting image, and the loss function for decomposing the normal lighting image is set as follows: ; in, , and These represent the image under normal light, the reflection of the image under normal light, and the illumination of the image under normal light, respectively. e and It's a hyperparameter. This is normal light image gradient calculation. This is the gradient calculation for normal light image illumination. F Represents the matrix norm; The total change in the lighting map is weighted by the gradient of the image to perform spatial smoothing of the lighting in a structure-aware manner. Networks GL and GR are introduced to update L and R respectively. GL is a fully convolutional network with five CONV layers, and then learns the implicit prior on L through ReLU activation. Its loss function is designed to be the sum of the loss functions of reflectance and illuminance, as shown in the following expression: ; In the formula, T represents the total number of stages. This represents the auxiliary approximation of normal light reflection variables. Indicates the normal light reflection variable. , , , All represent equilibrium parameters. Represents the illumination layer variable that assists in approximating normal light. Indicates the normal illumination layer variable. Indicates the illumination at each stage, Represents the reflectivity at each stage. Represents the reflectance of a normal light image. ( ) represents feature extraction, and SSIM represents structural similarity loss.

2. The image processing method for industrial environments according to claim 1, characterized in that, The kernel size of the convolutional layer in the initialization module is 3.

3. The loss function of the initialization module satisfies the following relationship: ; In the formula, I represents the low-light image. Let μ be the initial reflectivity, and µ be a hyperparameter. It initializes the lighting and C∈{R,G,B} represents the RGB channels.

3. The image processing method for industrial environments according to claim 1, characterized in that, The step of inputting the target illumination into the image enhancement module and obtaining the output enhanced image includes: Introduce a mapping with parameter θ Learning about lighting involves modeling basic units according to a set method, satisfying the following relationship, where the basic unit is a stage used to optimize the image: ; in, The residual at stage t, For illumination estimation network, and Regardless of the number of stages, the illumination estimation network maintains a shared structure and parameters at each stage. Let y represent the illumination at stage t, and y represent the low-light image.

4. The image processing method for industrial environments according to claim 3, characterized in that, The setting method is a progressive angle, and the modeling to obtain basic units according to the setting method includes: The image enhancement process is divided into N stages, where N is a positive integer. Generate basic units for image enhancement at each stage; In each stage, the process of image enhancement is transformed into learning the residual between illumination and low-light observations for the parameter θ.

5. The image processing method for industrial environments according to claim 3, characterized in that, During training, the image enhancement module includes a self-calibration module, and the basic unit of the illumination optimization process during image enhancement satisfies the following relationship: ; in: ; in, For low-light images, For the target image, For self-correcting mapping, For parameterized operators, This serves as the calibrated input for the next stage. Let t be the illumination during stage t.

6. The image processing method for industrial environments according to claim 1, characterized in that, The method further includes: S4: After resizing the enhanced image, input it into the target detection model to identify the target object in the enhanced image and determine the relevant information of the target object, including location coordinates, category information and confidence value; S5: Provide behavioral warnings based on the relevant information of the target object.

7. An image processing system for industrial environments, characterized in that, include: Monitoring devices are used to collect video frames in industrial environments; An edge computing server is used to read video frames from industrial environments collected by monitoring devices. Low-light images from these video frames are used as input to train an image enhancement network framework to obtain the final image enhancement model. This framework includes an initialization module, an optimization module, and an image enhancement module. During training: the low-light image is input to the initialization module to obtain the initial reflectivity and initial illumination output by the initialization module; the initial reflectivity, initial illumination, and normal illumination are input to the optimization module to iteratively refine the reflection and illumination layers to obtain the target illumination; the target illumination is input to the image enhancement module to obtain the output enhanced image; real-time collected image frames from the industrial environment are input to the image enhancement model to obtain the enhanced image; the enhanced image is also used to resize the enhanced image and input it into a target detection model to identify target objects in the enhanced image and determine relevant information about the target objects, including location coordinates, category information, and confidence values. Based on the relevant information of the target object, a warning message is generated by behavioral alert; A visual interface device, which is used to receive and display alert information from an edge computing server; Both the monitoring device and the visualization interface device are connected to the edge computing server; The step of inputting the initial reflectivity, initial illumination, and normal illumination into the optimization module to iteratively refine the reflection layer and illumination layer to obtain the target illumination includes: A structure-aware smoothing constraint is applied to the illumination of the normal lighting image, and the loss function for decomposing the normal lighting image is set as follows: ; in, , and These represent the image under normal light, the reflection of the image under normal light, and the illumination of the image under normal light, respectively. e and It's a hyperparameter. This is normal light image gradient calculation. This is the gradient calculation for normal light image illumination. F Represents the matrix norm; The total change in the lighting map is weighted by the gradient of the image to perform spatial smoothing of the lighting in a structure-aware manner. Networks GL and GR are introduced to update L and R respectively. GL is a fully convolutional network with five CONV layers, and then learns the implicit prior on L through ReLU activation. Its loss function is designed to be the sum of the loss functions of reflectance and illuminance, as shown in the following expression: ; In the formula, T represents the total number of stages. This represents the auxiliary approximation of normal light reflection variables. Indicates the normal light reflection variable. , , , All represent equilibrium parameters. Represents the illumination layer variable that assists in approximating normal light. Indicates the normal illumination layer variable. Indicates the illumination at each stage, Represents the reflectivity at each stage. Represents the reflectance of a normal light image. ( ) represents feature extraction, and SSIM represents structural similarity loss.

8. An image processing system for industrial environments, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.