Method and system for image fusion based on low-light image enhancement and visible-infrared image
By employing low-light image enhancement and visible-infrared image fusion in port clearance security monitoring, dynamic synergy between enhancement and fusion is achieved, generating high-quality multimodal image data. This solves the problems of inaccurate enhancement of details in low-light areas and insufficient complementarity of features in occluded areas in existing technologies, thereby improving the accuracy and real-time performance of abnormal behavior recognition.
Patent Information
- Application Number
- CN202610527694.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-21
AI Technical Summary
Existing infrared and visible light image fusion technology in port clearance security monitoring suffers from a lack of coordinated optimization between the enhancement and fusion stages. This results in inaccurate enhancement of details in low-light areas and insufficient complementarity of features in occluded areas, affecting the accuracy of abnormal behavior identification. Furthermore, the processing time is too long, making it difficult to meet real-time requirements.
A method based on low-light image enhancement and visible-infrared image fusion is adopted. Through preprocessing, parameter initialization and adaptation, bidirectional guided adjustment, adaptive weight acquisition and joint loss function training, high-quality multimodal image data is generated, realizing dynamic synergy between enhancement and fusion, and improving the robustness and accuracy of abnormal behavior recognition.
The system generates high-quality multimodal image data in low-light and occluded scenarios, improving the robustness and accuracy of identifying abnormal behavior of customs clearance personnel at ports, while also ensuring real-time processing, thus providing efficient and reliable technical support for port security management.
Smart Images

Figure CN122434749A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and port security monitoring technology, and in particular to a method and system based on low-light image enhancement and visible light-infrared image fusion. Background Technology
[0002] With the increasing frequency of international trade and cross-border personnel flows, port security monitoring has become a core support for ensuring public safety and improving customs clearance efficiency. Real-time and accurate detection of abnormal personnel behavior relies on high-quality multimodal image data. Port clearance monitoring scenarios generally face complex problems such as uneven lighting, localized low light, and obstructions from personnel or luggage. The fusion technology of infrared and visible light images has become a key means to solve the imaging quality problems in these scenarios: infrared images can penetrate partial obstructions to identify heat sources, while visible light images can provide rich color and texture information; the effective fusion of the two can achieve complementary advantages.
[0003] Existing infrared and visible light image fusion technologies suffer from the following technical shortcomings: The enhancement and fusion stages lack a collaborative optimization mechanism, often employing a unidirectional, independent processing mode of enhancement followed by fusion or fusion followed by enhancement, without forming a dynamic interaction or guidance relationship. This results in enhancement parameter adjustments failing to adapt to fusion requirements, and fusion strategy optimization not considering enhancement effects. Consequently, detail enhancement in low-light areas is inaccurate, and dual-modal feature complementarity in occluded areas is insufficient, leading to blurred key target features in the fused image and affecting the accuracy of subsequent abnormal behavior recognition. Furthermore, in the unidirectional, independent processing mode, to balance enhancement effects and fusion quality, models often require complex structures, resulting in excessively long single-frame processing times, making it difficult to meet the real-time requirements of port monitoring. Meanwhile, lightweight models, lacking collaborative design, suffer from insufficient feature learning, failing to adapt to complex scenarios at ports characterized by uneven lighting, localized low light, and frequent occlusion. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a method and system based on low-light image enhancement and visible light-infrared image fusion, which is applicable to low-light image enhancement and visible light-infrared image fusion processing in port clearance security monitoring environments. It can generate high-quality multimodal image data in harsh environments such as low light, low light, and complex lighting, providing accurate data support for the identification of abnormal behavior of port clearance personnel and improving the robustness of abnormal behavior identification.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: A first aspect is a method based on low-light image enhancement and visible-infrared image fusion, the method comprising: Infrared and visible light images of the same scene are preprocessed to obtain preprocessed standardized dual-modal image data; Based on the preprocessed standardized bimodal image data, a port scene-specific configuration file is loaded, and the parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module are initialized and adapted to obtain the adapted network parameters. The preprocessed standardized bimodal image data is input into the fusion and enhancement network. The network parameters are adapted and bidirectional guided adjustment of enhancement and fusion is performed. The enhancement gain parameters are dynamically optimized through the feedback signal of the fusion module, and the fusion strategy is dynamically adjusted according to the effect evaluation parameters of the enhancement module to generate bimodal features after passing through the brightness adjustment network. The dual-modal features after passing through the brightness adjustment network are input into the brightness feedback network, and feature enhancement, adaptive filtering, feedback signal generation and preliminary fusion operations are performed in sequence to obtain preliminary fused features. The initial fused features are input into the adaptive weight acquisition module to detect occlusion regions in the visible light image and generate occlusion masks; based on the occlusion masks, the fusion weights of infrared and visible light features are dynamically allocated, and feature enhancement and weighted fusion are performed to output the final fused features; Based on the final fusion features, the network is trained and optimized according to the joint loss function, and the network parameters are updated through backpropagation to obtain the optimized fusion model. The optimized fusion model is used to input the final fusion features into the post-processing module, where denormalization and edge smoothing operations are performed sequentially to generate the final fused image.
[0006] Furthermore, infrared and visible light images of the same scene are preprocessed to obtain preprocessed standardized dual-modal image data, including: The acquired infrared and visible light images are time-stamp aligned and verified to remove image pairs with time deviations exceeding a preset threshold, thus obtaining time-synchronized dual-modal image pairs. Pixel normalization is performed on the time-synchronized dual-modal image pairs to uniformly map the pixel value range of the infrared image and the visible light image to a preset range, thereby obtaining the normalized dual-modal image. The normalized bimodal image is subjected to data augmentation processing, including at least one of random cropping, horizontal flipping, and brightness perturbation, to obtain preprocessed standardized bimodal image data, which includes preprocessed standardized infrared image and preprocessed standardized visible light image.
[0007] Furthermore, based on the preprocessed standardized bimodal image data, a port-specific configuration file is loaded, and parameter initialization and adaptation are performed on the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module to obtain the adapted network parameters, including: Load the port scenario-specific configuration file, parse and extract the scenario prior configuration information from it; Based on the extracted scenario prior configuration information, the enhancement gain parameters and evaluation thresholds of the fusion and enhancement networks are initialized and assigned values. Based on the extracted scene prior configuration information, the filter coefficients and convolution kernel weights of the brightness feedback network are initialized and assigned values. Based on the prior configuration information of the scenario, the preset fusion weights of the occluded and unoccluded regions in the adaptive weight acquisition module are initialized and assigned values. By integrating all parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module after initialization, the adapted network parameters are obtained.
[0008] Furthermore, the preprocessed standardized bimodal image data is input into the fusion and enhancement network. Adapted network parameters are used to perform bidirectional guided adjustment of enhancement and fusion. The enhancement gain parameters are dynamically optimized through feedback signals from the fusion module, and the fusion strategy is dynamically adjusted based on the enhancement module's performance evaluation parameters. This generates bimodal features after passing through the brightness adjustment network, including: The preprocessed standardized bimodal image data is input into the fusion and enhancement network; Using the adapted network parameters, micro-light enhancement processing is performed on the visible light image to generate the enhanced visible light image and the corresponding enhancement effect evaluation parameters. The enhanced visible light image and the preprocessed standardized infrared image are input into the fusion module in the enhancement network. The fusion module performs feature matching and preliminary fusion to generate feature matching error, which is used as a feedback signal. Based on feature matching error and enhancement effect evaluation parameters, the enhancement gain parameters are dynamically updated by fusing and enhancing the parameter adjustment mechanism in the network, thus generating the updated enhancement gain parameters. By employing updated enhancement gain parameters, a new round of enhancement and fusion-guided adjustment is performed on the preprocessed standardized bimodal image data to generate bimodal features after passing through a brightness adjustment network.
[0009] Furthermore, the dual-modal features after passing through the brightness adjustment network are input into the brightness feedback network, and feature enhancement, adaptive filtering, feedback signal generation, and preliminary fusion operations are performed sequentially to obtain preliminary fused features, including: The dual-modal features after passing through the brightness adjustment network are input into the feature enhancement submodule of the brightness feedback network to enhance the infrared and visible light features respectively, thus obtaining the enhanced dual-modal features. The enhanced bimodal features are input into the adaptive filtering submodule of the luminance feedback network. The filtering intensity is dynamically adjusted based on the signal-to-noise ratio of the features, and the filtering process is performed to obtain the filtered bimodal features. The filtered dual-modal features are input into the preliminary fusion submodule of the brightness feedback network to perform weighted fusion of infrared and visible light features, thereby obtaining preliminary fused features.
[0010] Furthermore, the preliminary fused features are input into the adaptive weight acquisition module to detect occlusion regions in the visible light image and generate an occlusion mask; based on the occlusion mask, fusion weights for infrared and visible light features are dynamically allocated, and feature enhancement and weighted fusion are performed to output the final fused features, including: The preliminary fused features and the corresponding preprocessed standardized visible light image are input into the occlusion detection unit of the adaptive weight acquisition module to perform occlusion region detection and generate an occlusion mask. The initial fused features are input into the dual attention branch of the adaptive weight acquisition module, and attention enhancement processing is performed on the infrared features and visible light features respectively to obtain the attention-weighted dual-modal features. Based on the occlusion mask, the dynamic fusion weight of infrared features and visible light features is calculated in the dynamic weight allocation unit of the adaptive weight acquisition module. Based on the dynamic fusion weights, the attention-weighted bimodal features are weighted and fused to output the final fused features.
[0011] Furthermore, based on the final fusion features, training and optimization are performed according to the joint loss function, and the network parameters are updated through backpropagation to obtain the optimized fusion model, including: Based on the final fused features and the preprocessed standardized bimodal image data, a joint loss function value is calculated, which is composed of no-reference color loss, feature consistency loss and structural similarity loss. Based on the joint loss function value, the gradient is calculated through backpropagation, and the network parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module are updated to obtain the model with updated parameters. The updated model is used to process the validation set data to obtain the corresponding validation performance metrics. The training termination condition is determined based on the verification performance metrics. If the condition is met, the current model is used as the optimized fusion model. If the condition is not met, the joint loss function value is recalculated to continue the next round of iterative training.
[0012] Furthermore, the optimized fusion model is used to input the final fusion features into the post-processing module, where denormalization and edge smoothing operations are performed sequentially to generate the final fused image, including: The final fusion features are input into the optimized fusion model for processing to obtain an image matrix corresponding to the final fusion features; The image matrix is subjected to denormalization processing. Combined with the pixel extreme value parameters recorded in the preprocessing stage, the normalized pixel values are restored to the original numerical range to obtain the denormalized image. Edge smoothing filtering is performed on the denormalized image to eliminate image edge artifacts and generate the final fused image.
[0013] Secondly, systems based on low-light image enhancement and visible-infrared image fusion include: The data input and preprocessing module is used to preprocess the infrared and visible light images obtained in the same scene to obtain preprocessed standardized dual-modal image data. The port scene adaptation module is used to load a port scene-specific configuration file based on the preprocessed standardized bimodal image data, and to perform parameter initialization and adaptation of the fusion and enhancement network, brightness feedback network and adaptive weight acquisition module to obtain the adapted network parameters. The fusion and enhancement network module is used to take the input preprocessed standardized bimodal image data, use the adapted network parameters, and perform bidirectional guided adjustment of enhancement and fusion. It dynamically optimizes the enhancement gain parameters through the feedback signal of the fusion module, and dynamically adjusts the fusion strategy according to the effect evaluation parameters of the enhancement module to generate bimodal features after brightness adjustment network. The brightness feedback network module is used to sequentially perform feature enhancement, adaptive filtering, feedback signal generation and preliminary fusion operations on the obtained dual-modal features after brightness adjustment network to obtain preliminary fused features. The adaptive weight acquisition module is used to take the input preliminary fusion features, perform detection of occlusion regions in the visible light image and generate an occlusion mask; dynamically allocate fusion weights for infrared and visible light features based on the occlusion mask, and perform feature enhancement and weighted fusion to output the final fusion features; The training optimization module is used to train and optimize the model based on the final fusion features using a joint loss function, and to update the network parameters through backpropagation to obtain the optimized fusion model. The output post-processing module is used to obtain the final fusion features through the optimized fusion model, and then perform denormalization and edge smoothing operations in sequence to generate the final fused image.
[0014] The above-described solution of the present invention has at least the following beneficial effects: Because it adopts a fusion-enhancement bidirectional guided feedback loop framework, integrates feature enhancement, filtering, and fusion into a unified network, has a joint training module with no reference color loss, and has a port scene-specific adaptation and dynamic weight allocation mechanism, it overcomes the core problems of existing technologies such as lack of synergy between enhancement and fusion, insufficient feature complementarity in occluded scenes, easy color distortion, and poor scene adaptability. As a result, it realizes the generation of high-quality multimodal image data in low-light and occluded scenes, improves the robustness and accuracy of abnormal behavior recognition of port clearance personnel, and takes into account real-time processing, providing efficient and reliable technical support for port clearance security management. Attached Figure Description
[0015] Figure 1 This is a schematic flowchart of a method based on low-light image enhancement and visible-infrared image fusion provided by an embodiment of the present invention.
[0016] Figure 2 This is an overall framework diagram provided by an embodiment of the present invention.
[0017] Figure 3 This is a framework diagram of the adaptive weight acquisition module provided in an embodiment of the present invention.
[0018] Figure 4 This is a fusion and enhancement network framework diagram provided by an embodiment of the present invention.
[0019] Figure 5 This is a diagram of the brightness feedback network framework provided in an embodiment of the present invention. Detailed Implementation
[0020] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0021] like Figures 1-5 As shown, embodiments of the present invention propose a method based on low-light image enhancement and visible-infrared image fusion, the method comprising the following steps: Step 1: Preprocess the infrared and visible light images of the same scene to obtain preprocessed standardized dual-modal image data; Step 2: Based on the preprocessed standardized bimodal image data, load the port scene-specific configuration file, and perform parameter initialization and adaptation on the fusion and enhancement network, brightness feedback network and adaptive weight acquisition module to obtain the adapted network parameters. Step 3: Input the preprocessed standardized bimodal image data into the fusion and enhancement network, use the adapted network parameters, perform bidirectional guided adjustment of enhancement and fusion, dynamically optimize the enhancement gain parameters through the feedback signal of the fusion module, and dynamically adjust the fusion strategy according to the effect evaluation parameters of the enhancement module to generate bimodal features after brightness adjustment network. Step 4: Input the dual-modal features after the brightness adjustment network into the brightness feedback network, and perform feature enhancement, adaptive filtering, feedback signal generation and preliminary fusion operations in sequence to obtain preliminary fused features; Step 5: Input the preliminary fused features into the adaptive weight acquisition module to detect occlusion areas in the visible light image and generate an occlusion mask; dynamically allocate fusion weights for infrared and visible light features based on the occlusion mask, and perform feature enhancement and weighted fusion to output the final fused features; Step 6: Based on the final fusion features, train and optimize according to the joint loss function, update the network parameters through backpropagation, and obtain the optimized fusion model; Step 7: Using the optimized fusion model, the final fusion features are input into the post-processing module, and denormalization and edge smoothing operations are performed sequentially to generate the final fused image.
[0022] In this embodiment of the invention, preprocessing is completed through standardization and data augmentation to ensure the quality of input data and generalization ability; network parameters are accurately initialized based on port scenario-specific configuration to adapt to scene characteristics; bidirectional dynamic collaboration between enhancement and fusion is achieved through fusion and enhancement networks to avoid the disconnect between their needs; the brightness feedback network simultaneously completes feature enhancement, interference filtering, and preliminary fusion to reduce interference accumulation; the adaptive weight acquisition module accurately detects occluded areas and dynamically adjusts fusion weights to achieve complementary differential features; joint loss training ensures color reproduction and feature integrity, and post-processing eliminates edge artifacts and restores the original image range; ultimately, the quality of fused images in low-light and frequently occluded port scenarios is effectively improved, enhancing the robustness and accuracy of abnormal behavior recognition, while also taking into account real-time processing, providing efficient and reliable technical support for port clearance security monitoring.
[0023] In a preferred embodiment of the present invention, step 1 above may include: Step 1.1 involves performing timestamp alignment verification on the acquired infrared and visible light images to eliminate image pairs with time deviations exceeding a preset threshold, thus obtaining time-synchronized dual-modal image pairs. Specifically, this includes: starting the infrared and visible light cameras, setting the infrared camera exposure time to 10ms and the visible light camera exposure time to 5ms according to predefined parameters to ensure image quality meets the requirements of the port monitoring scenario; generating a high-precision synchronization trigger pulse to control both cameras to simultaneously capture images of the core area of the port clearance, including inspection channels and personnel passage areas, acquiring infrared and visible light images of the same scene; after capturing images, reading the capture time recorded in the DateTimeOriginal field of the EXIF metadata of each image, and extracting the time of capture from the infrared and visible light images. The timestamps are converted to millisecond-level format, and the difference between the two sets of timestamps is calculated to verify the time synchronization of the dual-light images. The preset threshold is set to 10ms, which is determined based on the speed of personnel passage at the port and the monitoring frame rate requirements. This ensures that the dual-light images accurately capture the scene at the same moment and avoids misalignment of target features during fusion due to time differences. When the time difference exceeds the threshold, the set of images is judged as invalid and discarded. At the same time, the reason for discarding (timestamp inconsistency) and the image number are recorded in the system log. When the time difference does not exceed 10ms, it is determined to be a time-synchronized dual-modal image pair. Finally, the image pair is stored in the specified folder according to the directory structure of shooting date, channel number, and image type. The storage format is PNG to ensure that the image is not compressed or distorted. The storage path and associated image ID are also recorded.
[0024] Step 1.2: Perform pixel normalization processing on the time-synchronized dual-modal image pairs to uniformly map the pixel value ranges of the infrared and visible light images to a preset interval, obtaining the normalized dual-modal image. Specifically, this includes: to eliminate the difference in pixel value ranges between the infrared and visible light images and avoid interference from different magnitudes of pixel values on the gradient of subsequent convolution operations, the pixel values of the two types of images need to be uniformly mapped to the preset interval [0, 1]. First, traverse all pixels of a single infrared image and a single visible light image, calculate the standard deviation σ of the image pixel values, exclude outliers exceeding the 2σ range, and then determine the minimum pixel value X of each single image. min With the maximum pixel value X max The raw infrared image data is 16-bit depth, which needs to be converted to 32-bit floating-point numbers first to avoid data overflow or precision loss in subsequent calculations; then, it is processed according to the formula. Perform pixel-by-pixel normalization calculation, where X represents the value of a single pixel in the original image. min X represents the minimum pixel value of a single image after removing outliers. max X represents the maximum pixel value of a single image after outlier removal. normThis represents the normalized pixel value. During the calculation, the image rows and columns are traversed in parallel to efficiently complete the pixel-by-pixel calculation. An outlier handling mechanism is also implemented; if the normalized value of a pixel exceeds the [0, 1] range, it is directly limited to 0 or 1. The final output is the normalized bimodal image, and the corresponding image ID and X are recorded in the log. min and X max The parameters provide data support for subsequent denormalization processing.
[0025] Step 1.3 involves performing data augmentation processing on the normalized bimodal image, including at least one of random cropping, horizontal flipping, and brightness perturbation, to obtain preprocessed standardized bimodal image data. This standardized bimodal image data includes preprocessed standardized infrared and preprocessed standardized visible light images. Specifically, to improve the generalization ability for changes in illumination, target posture, and local low-light environments in port clearance scenarios, data augmentation processing is performed on the normalized bimodal image. Random numbers in the range of 0 to 1 are generated to control the triggering of the augmentation strategy. The three augmentation strategies can be executed individually or in combination. When performing random cropping, the cropping area is limited to the effective range of the channel based on the coordinate range of the inspection channel area in the port scenario-specific configuration file. Within the effective range, the starting coordinates (x0, y0) of the cropping are determined, and the image is cropped to a size of 480×480. After cropping, the image is restored to a fixed size of 640×512 through pixel interpolation to ensure uniform input size for subsequent modules.
[0026] When performing horizontal flipping, if the generated random number is less than 0.5, the flip is triggered. By sharing a random number seed, the infrared image and the visible light image are flipped synchronously to avoid misalignment of dual-modal image features. When performing brightness perturbation, the infrared image is distinguished as a single channel and the visible light image as a three-channel image based on the number of channels. Only the visible light image is processed: it is converted from the RGB color space to the HSV color space, a brightness perturbation coefficient is generated within ±0.1, the brightness value of the V channel is adjusted with an adjustment step of 0.01, and then converted back to the RGB space. After adjustment, the average brightness of the visible light image is calculated. If the average value exceeds the range of [0.3, 0.7], it is readjusted to avoid overexposure or underexposure that leads to color distortion. After all enhancement operations are completed, the enhanced dual-modal image is quality checked by calculating the image sharpness (gradient mean) and brightness standard deviation to ensure no obvious distortion. Finally, the preprocessed standardized dual-modal image data is obtained, which includes the preprocessed standardized infrared image and the standardized visible light image.
[0027] In a preferred embodiment of the present invention, step 2 above may include: Step 2.1: Load the port scenario-specific configuration file, parse and extract the scenario prior configuration information, specifically including: loading a dedicated configuration file customized for the port clearance scenario, parsing the file content segment by segment according to the classification logic of area information, equipment parameters, and scenario characteristics, ensuring that all key prior information is extracted without omission; when parsing area information, accurately extract the coordinates of the upper left and lower right corners of the inspection channel, such as x1=100, y1=50, x2=540, y2=462, to determine the effective boundary of feature extraction and avoid invalid background interference; when parsing occlusion-related information, read the 128×128 size... High-frequency areas obstructed by luggage are marked with a mask. The area with a mask value of 1 is identified pixel by pixel, corresponding to the middle section of the passageway where luggage is densely packed and obstruction is frequent during customs clearance. The matrix data of this mask is stored. When analyzing equipment parameters, the standard exposure parameters of 10ms for infrared cameras and 5ms for visible light cameras are extracted as the basis for subsequent image acquisition quality calibration. When analyzing scene characteristics, the noise type of the port's nighttime surveillance images is confirmed to be Gaussian noise, and the noise variance is statistically determined to be 0.02, providing a precise reference for subsequent filter coefficient setting. All extracted prior information is stored as structured data according to categories for easy retrieval by subsequent modules.
[0028] Step 2.2: Based on the extracted scene prior configuration information, initialize the enhancement gain parameters and evaluation threshold of the fusion and enhancement network. Specifically, based on the low-light scene characteristics in the extracted scene prior configuration information, combined with the actual illumination statistics of the port clearance scene, the overall average illumination intensity of the port's nighttime inspection channel is less than 50 lux, while the illumination intensity in local areas such as channel corners and baggage placement areas is even less than 10 lux. Images in these areas generally suffer from insufficient brightness and loss of detail. To adapt to these scene characteristics and avoid excessive initial enhancement leading to highlight clipping or loss of detail, the initial value of the enhancement gain parameter of the fusion and enhancement network is set to 1.0. This value was determined through comparative testing of multiple sets of low-light images of the port. It can provide basic brightness compensation for low-light areas and reserve sufficient space for subsequent dynamic adjustment based on the fusion effect, ensuring... The initial enhancement effect will not conflict with the subsequent fusion requirements. Simultaneously, considering the core requirement of identifying abnormal behavior of personnel at the port, the key is to clearly capture facial features, body movements, and details of carried items. Therefore, the evaluation threshold is set to a structural similarity (SSIM) threshold of 0.75. This threshold was derived by analyzing over 1000 sets of port low-light image samples. When the SSIM value of the enhanced image and the original image is below 0.75, the blurring rate of key personnel features will exceed 30%, such as unclear facial contours and indistinguishable body movements, failing to meet the feature clarity requirements of subsequent fusion steps. At this point, a parameter adjustment mechanism will be triggered. This initial threshold is closely linked to the dynamic parameter optimization logic of subsequent steps, providing a clear and tailored judgment standard for the precise adjustment of enhancement gain parameters, ensuring that the enhancement effect always revolves around the fusion requirements.
[0029] Step 2.3: Based on the extracted scene prior configuration information, initialize the filter coefficients and convolution kernel weights of the brightness feedback network. Specifically, this includes: based on the Gaussian noise characteristics of the port nighttime monitoring images confirmed in the scene prior configuration information, actual testing shows that the noise variance is stable at around 0.02. This type of noise can cause fine graininess in the image, interfering with the extraction and matching of dual-modal features. Therefore, it is necessary to perform refined initialization of the filter coefficients and convolution kernel weights of the brightness feedback network; set the base coefficient of the adaptive filter to 1.0, which is determined by combining the noise variance. The correlation test with the filtering effect confirmed that it can play a basic filtering role for Gaussian noise with a current variance of 0.02, and can also ensure that the adjustment range of the filtering intensity is reasonable when dynamically adjusted according to the feature signal-to-noise ratio in the future, avoiding over-filtering that leads to blurred target features or under-filtering that leaves residual noise. As for the convolution kernels in the network, considering that 64 channels of high-dimensional features need to be output after subsequent feature enhancement, 64 3×3 convolution kernels are configured. The size of the convolution kernels can ensure sufficient receptive field to capture key target features, while also taking into account computational efficiency, which meets the design requirements of lightweight networks.
[0030] The convolutional kernel weights are generated using a He normal distribution. This distribution characteristic ensures that the weight mean is close to 0 and the variance matches the gradient propagation law of the ReLU activation function, effectively alleviating the gradient vanishing problem caused by unreasonable weight distribution in the early stages of training. At the same time, the bias terms of all convolutional layers are uniformly set to 0.1. This value has been verified through multiple experiments and ensures that neurons can still produce non-zero outputs when the weight values are small in the early stages of training. This avoids the ReLU activation function suppressing some useful features and ensures that the thermal features of infrared images and the color and texture features of visible light images can be effectively enhanced in the initialization stage, laying a high-quality feature foundation for subsequent filtering and fusion stages.
[0031] Step 2.4: Based on the scenario prior configuration information, initialize the preset fusion weights for occluded and unoccluded areas in the adaptive weight acquisition module. Specifically, this includes: combining the 128×128 baggage occlusion high-frequency area marker mask in the scenario prior configuration information, and the actual occlusion statistics for the port clearance scenario: In the middle section of the channel, due to dense baggage placement and crowded queues, the occlusion rate exceeds 60%, and the occluding objects are mostly opaque items such as baggage and backpacks. This type of occlusion severely obscures the detailed features of people in visible light images, while infrared images can penetrate this type of occlusion to capture the heat source signal of people. Based on this, refine the initialization of the preset fusion weights for occluded and unoccluded areas in the adaptive weight acquisition module; for occluded areas, adjust the preset fusion weights... The value is set to 0.7. This value was determined through multiple sets of port occlusion image fusion tests. It ensures that the thermal features of infrared images are used preferentially in occluded scenarios to accurately locate the position and outline of occluded personnel and avoid target loss due to the obstruction of visible light features. For unoccluded areas, the preset fusion weight is applied. When set to 0.3, the visible light image can clearly present key features such as clothing color, facial details, and body movements of people. This weight can maximize the retention of this information that is crucial for the identification of abnormal behavior. The setting of the two types of weights fully considers the type and frequency of port occlusion and the complementary advantages of dual-modal images. It is also deeply linked with the subsequent dynamic weight adjustment logic based on occlusion mask, providing a reliable initial benchmark for the accurate fusion of features in different scenarios.
[0032] Step 2.5 integrates all parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module after initialization to obtain the adapted network parameters. Specifically, this includes: classifying and organizing the parameters of each module after initialization into key-value pairs based on module name, parameter type, parameter value, and parameter purpose to construct a complete parameter dictionary; among them, the parameter entries of the fusion and enhancement network include the initial enhancement gain value: 1.0 (adapting to basic enhancement in low light) and the SSIM evaluation threshold: 0.75 (to determine whether the enhancement effect meets the standard); the parameter entries of the brightness feedback network include the basic filtering coefficient σ0: 1.0 (basic filtering of Gaussian noise), the convolution kernel weights: 64 3×3 He normal distribution weights (extracting bimodal features), and the bias term: 0.1, used to avoid the initial gradient vanishing; the parameter entries of the adaptive weight acquisition module include the preset weights for the occlusion region. Prioritize infrared features and pre-set weights for unobstructed areas Preferred visible light characteristics.
[0033] After the parameter dictionary is constructed, a multi-level parameter verification mechanism is initiated: First, the consistency of parameter data types is verified to ensure that all floating-point parameters have a uniform 64-bit precision; second, the rationality of parameter value ranges is verified, such as weight parameters needing to be within the range of 0 to 1, and convolution kernel dimensions needing to match the number of feature channels; finally, the adaptability of parameter uses to module functions is verified to avoid parameter mismatches; after verification, the parameter dictionary is stored as a binary file, and a parameter index table is generated to record the storage address and calling method of each parameter, ensuring that subsequent fusion and enhancement networks, brightness feedback networks, and adaptive weight acquisition modules can quickly and accurately call the corresponding parameters through the index table, ultimately forming a network parameter set that is fully adapted to the scenario of uneven lighting + local dim light + frequent occlusion at the port, providing stable and accurate parameter support for subsequent core collaborative processing links.
[0034] In a preferred embodiment of the present invention, step 3 above may include: Step 3.1 involves inputting the preprocessed standardized bimodal image data into the fusion and enhancement network. Specifically, this includes inputting the preprocessed standardized bimodal image data, comprising standardized infrared and visible light images, both with dimensions [B, C, H, W]. The infrared image has 1 channel (C=1) and uses 32-bit floating-point data, while the visible light image has 3 channels (C=3) and uses 32-bit floating-point data. The dimensions are fixed at H=512 and W=640, and pixel values are normalized to the [0, 1] range. This fusion and enhancement network is a collaborative core architecture specifically designed for scenarios involving uneven lighting, localized low light, and frequent occlusion at ports of entry, inheriting from deep learning. The network foundation class, by rewriting the initialization unit and forward propagation unit, supports forward feature processing and backward gradient propagation. Its core objective is to break down the barriers of traditional unidirectional processing and construct a dynamic collaborative closed loop between the brightness adjustment network and the fusion module. Internally, it integrates three core sub-modules: the signal interaction sub-module adopts a tensor high-speed communication mechanism to achieve data sharing through class instance variables, ensuring real-time bidirectional data transmission between the enhancement module and the fusion module; the enhancement effect evaluation sub-module uses structural similarity (SSIM) as the core indicator to quantitatively judge whether the enhancement quality is suitable for the fusion requirements; and the parameter adjustment sub-module dynamically optimizes the enhancement gain parameters based on the evaluation results and fusion feedback errors. The three modules form a close linkage through tensor data interaction.
[0035] After the network starts, the adapted network parameters are loaded first, including the initial enhancement gain value of 1.0 and the SSIM evaluation threshold of 0.75. These parameters are then assigned to the corresponding functional units in key-value pairs according to submodule name and parameter type to ensure accurate parameter calls in each submodule. Subsequently, a multi-level data verification mechanism is initiated: First, the data type and precision are verified to confirm that both bimodal images are 32-bit floating-point numbers, avoiding calculation errors caused by data type mismatches. Second, dimensional consistency is verified, checking whether the infrared image meets the dimensional requirements of [B, 1, 512, 640] and the visible light image meets the dimensional requirements of [B, 3, 512, 640], ensuring adaptation to the network input layer. Third, the pixel value range is verified, traversing the image pixels to confirm that all pixel values are in the [0, 1] range, removing pixels with abnormal overflow, and clipping them to 0 or 1 if any are found. Fourth, data integrity is verified, checking for missing pixels or damaged areas in the image to ensure the input data is free of quality defects.
[0036] After successful verification, the network initializes its internal signal interaction links and parameter storage units: the signal interaction links create dedicated data transmission buffers to adapt to different data formats for forward transmission (enhanced feature maps, effect evaluation parameters) and reverse transmission (feature matching error signals), while enabling interface adaptation mechanisms to ensure that data is transmitted between the enhancement module and the fusion module without delay or distortion; the parameter storage unit uses a high-speed cache to store key data such as current enhancement gain parameters and SSIM thresholds, supporting real-time reading and updating, providing rapid data retrieval support for subsequent collaborative processing; after completing the above initialization and preparation work, the network officially enters the enhancement and fusion collaborative processing flow, initiating a cyclical adjustment of enhancement, fusion, feedback, and optimization. The "bidirectional guided adjustment of enhancement and fusion" mechanism refers to the dynamic collaborative process between the enhancement module and the fusion module: the fusion module feeds back the feature matching error to the parameter adjustment submodule to optimize the enhancement gain parameters; simultaneously, the effect evaluation parameters of the enhancement module are fed back to the fusion module to dynamically adjust the fusion strategy, thereby forming a closed-loop optimization.
[0037] Step 3.2: Using the adapted network parameters, perform low-light enhancement processing on the visible light image to generate an enhanced visible light image and corresponding enhancement effect evaluation parameters. Specifically, this includes: based on the initial enhancement gain value of 1.0 in the adapted network parameters, perform adaptive low-light enhancement processing on the standardized visible light image: for areas with illumination below 50 lux, such as corners of the port inspection channel and baggage handling areas, focus on improving brightness and contrast; for the central area of the channel with relatively sufficient illumination of 50 to 100 lux, moderately enhance detail sharpness to avoid over-processing that leads to highlight clipping; after enhancement, generate the enhanced visible light image, and then calculate the enhancement effect evaluation parameters, using structural similarity (SSIM) as the core evaluation index, according to the following formula: ; The calculation is performed, where x is the enhanced visible light image and y is the original normalized visible light image. , These are the mean values within a 3×3 local window of each of the two images. , These represent the standard deviations within that local window. The covariance within this local window. =6.5025、 =58.5225 is a fixed constant used to avoid the denominator being zero; during the calculation, local statistics are collected by sliding a 3×3 local convolution kernel pixel by pixel, and the SSIM value is obtained by substituting it into the formula. The global mean is then taken as the final enhancement effect evaluation parameter to determine whether the enhancement quality meets the standard.
[0038] Step 3.3 involves inputting the enhanced visible light image and the preprocessed standardized infrared image into the fusion module of the fusion and enhancement network. The fusion module performs feature matching and preliminary fusion, generating a feature matching error, which is then used as a feedback signal. Specifically, this includes: inputting the enhanced visible light image and the preprocessed standardized infrared image into the fusion module of the fusion and enhancement network. This fusion module is the core sub-module of the fusion and enhancement network, responsible for matching and initially fusing the enhanced visible light features and infrared features, and generating a feature matching error as a feedback signal. The fusion module first extracts features through two 3×3 convolutional layers: from the enhanced visible light... Color and texture features such as color distribution, edge contours, and texture gradients are extracted from the image, while heat source features such as heat source intensity and target contours are extracted from the infrared image. Subsequently, feature complementarity analysis logic is used to match the two types of features in both channel and spatial dimensions: comparing the response intensity of the bimodal features in the channel dimension and aligning the target's position coordinates in the spatial dimension to determine the fit and complementarity of the two types of features. After matching, an equal-weighted (0.5 each) method is used to complete the initial fusion, generating an initial fused feature. The specific calculation method is: Initial fused feature = Enhanced visible light color and texture features × 0.5 + Infrared heat source features × 0.5; Simultaneously, the feature matching error is calculated. Specifically, at the same spatial location and channel dimension, the Euclidean distance between the enhanced visible light feature and the infrared feature is calculated, and then the average of the Euclidean distances at all locations is taken. The average value obtained is the feature matching error. This error is used to reflect the degree of adaptation between the current enhanced feature and the infrared feature, and finally it is transmitted as a feedback signal to the parameter adjustment mechanism.
[0039] Step 3.4: Based on the feature matching error and enhancement effect evaluation parameters, the enhancement gain parameters are dynamically updated through the parameter adjustment mechanism in the fusion and enhancement network to generate updated enhancement gain parameters. Specifically, this includes: based on the generated feature matching error... The enhancement gain parameters are dynamically updated using the obtained global SSIM value and the parameter adjustment mechanism in the fusion and enhancement network; the update process strictly follows the following formula: ; in This is the current enhancement gain parameter, with an initial value of 1.0. The updated enhancement gain parameters; The feature matching error is represented by 1 - SSIM. A larger error indicates a poorer fit between the enhanced and infrared features, requiring a more significant adjustment to the gain. The SSIM value represents the enhancement effect deviation; a lower SSIM value indicates a larger deviation and a worse enhancement effect, requiring targeted optimization. 0.1 and 0.05 are weighting coefficients, determined through experiments in multiple port scenarios, used to balance the impact of feature matching error and enhancement effect deviation on the gain parameters. After calculation, a coefficient boundary constraint mechanism is activated: if... If the value is below 0.5, it will be forced to be set to 0.5 to avoid insufficient enhancement causing blurring in low-light areas; if the value is above 2.0, it will be forced to be set to 2.0 to avoid overexposure distortion; finally, the updated enhancement gain parameters are generated and stored in real time through the parameter storage unit for use in the next round of enhancement processing.
[0040] Step 3.5: Using the updated enhancement gain parameters, perform a new round of enhancement and fusion guidance adjustment on the preprocessed standardized bimodal image data to generate bimodal features after passing through the brightness adjustment network. Specifically, this includes: using the updated enhancement gain parameters, performing a new round of low-light enhancement processing on the visible light image in the preprocessed standardized bimodal image data to ensure that the enhanced visible light features better match the fusion requirements of the infrared features; then, input the newly enhanced visible light image and the standardized infrared image back into the fusion module for re-feature matching and preliminary fusion, and dynamically adjust the fusion strategy: if the feature matching error calculated in the new round... The decrease compared to the previous round indicates improved adaptability of enhancement and fusion, maintaining the current equal-weight fusion logic. If the error does not decrease, the fusion weight of visible light features will be fine-tuned by ±0.1 to optimize feature matching. Considering the real-time requirements of port monitoring, single-frame processing is ≤50ms, and 1 to 2 iterations are sufficient: 1 iteration can quickly adapt to basic fusion requirements, and 2 iterations can further improve adaptation accuracy without significantly increasing processing time. After the above cyclic adjustment, a dual-modal feature after brightness adjustment network is generated. This feature retains the clear heat source target outline of the infrared image and contains rich color and texture details of the visible light image, maintaining the dimension [B,64,640,512], providing high-quality input for further processing by the subsequent brightness feedback network.
[0041] In a preferred embodiment of the present invention, step 4 above may include: Step 4.1: Input the bimodal features after the brightness adjustment network into the feature enhancement submodule of the brightness feedback network to enhance the infrared and visible light features respectively, obtaining the enhanced bimodal features. Specifically, this includes: taking the bimodal features after the brightness adjustment network, with dimensions [B, 64, 640, 512], and including the initially enhanced infrared heat source feature F... ir_opt With visible light color texture features F vis_optThe two types of features are arranged alternately by channels, with the first 32 channels being infrared features and the last 32 channels being visible light features. The feature enhancement submodule input to the brightness feedback network is specifically designed for low-light and occluded scenes at ports. Its core function is to further enhance the core information of the two types of features and weaken invalid background interference, providing more accurate feature input for subsequent filtering and fusion. First, through a channel separation mechanism, the dual-modal features are split into independent infrared feature branches and visible light feature branches according to the channel index, ensuring that the two types of features can be individually enhanced and avoiding cross-interference.
[0042] For the infrared feature branch, the focus is on enhancing the feature response of the heat source target area (port personnel, heat-generating items). First, the global response mean M of each infrared feature channel is calculated. ir_channel Then iterate through all pixels in each channel, and for those with response values higher than M... ir_channel The heat source region of ×1.2 is weighted and enhanced, and the specific mathematical expression is as follows: ; Among them, F ir_enhance To enhance the post-infrared signature, F ir_opt To initially optimize infrared features, 1.1 is used as the base enhancement factor, and M... ir_global The global gradient mean G is calculated for all infrared feature channels to highlight the outline of the heat source target and weaken background clutter. For the visible light feature branch, the focus is on enhancing color and texture details, especially the edge and outline details of the low-light areas at the port (channel corners, night passage areas). The global gradient mean G for each visible light feature map is calculated first. vis_global Then, the local gradient value G of each pixel is calculated using the 3×3 Sobel gradient operator. vis_local For gradient values higher than G vis_global The detailed areas are enhanced, and the specific mathematical expression is as follows: ; Among them, F vis_enhance To enhance the visible light characteristics, F vis_opt To initially optimize visible light features and prevent details in low-light areas from being obscured by the background, after enhancement processing, the two types of enhanced features are re-stitched in the original channel order. At the same time, the stitched features are normalized to ensure that the pixel values are in the range of [0,1], resulting in enhanced dual-modal features with dimensions maintained at [B,64,640,512]. This preserves the heat source identification of infrared features while improving the detail clarity of visible light features, fully meeting the feature extraction needs of port scenes.
[0043] It should be noted that the luminance feedback network (LFN) in this step is a dual-modal feature fusion core architecture specifically designed for monitoring scenarios with uneven lighting and frequent occlusion at ports. Its core design goal is to balance the fusion effect with the real-time requirements of port monitoring. While simplifying the network structure and reducing the number of parameters, it accurately retains the core information of infrared and visible light dual-modal features, avoiding the high time consumption problem caused by complex networks, and ensuring that the processing time of a single frame does not exceed 50ms, thus adapting to the actual deployment requirements of port real-time monitoring. The network adopts a modular design and mainly includes three core sub-modules: a feature enhancement sub-module, an adaptive filtering sub-module, and a preliminary fusion sub-module, which work together in sequence. The process is progressive: the feature enhancement submodule is responsible for strengthening the core details of the dual-modal features and weakening background interference, corresponding to step 4.1; the adaptive filtering submodule is responsible for filtering the noise introduced during the feature enhancement process to avoid the influence of artifacts, corresponding to step 4.2; the preliminary fusion submodule is responsible for realizing the basic weighted fusion of dual-modal features, providing support for subsequent adaptive weight allocation, corresponding to step 4.3; compared with traditional fusion networks, this brightness feedback network significantly reduces computational complexity by simplifying the convolutional kernel structure, reducing redundant channels, and reusing the feature parameters mentioned above, while also specifically optimizing the adaptability to port scenes, enabling efficient processing of dual-modal feature fusion tasks in low-light and occluded scenes.
[0044] Step 4.2 involves inputting the enhanced dual-modal features into the adaptive filtering submodule of the luminance feedback network. The filtering intensity is dynamically adjusted based on the signal-to-noise ratio of the features, and filtering is performed to obtain the filtered dual-modal features. Specifically, this includes inputting the enhanced dual-modal features, including the enhanced infrared feature F... ir_enhance Compared with enhanced visible light characteristics F vis_enhance The dimension [B, 64, 640, 512] is the adaptive filtering submodule of the input brightness feedback network. The core function of this submodule is to filter the noise introduced during the enhancement process, mainly Gaussian noise in the port nighttime monitoring scene. The measured variance is stable at around 0.02. Simultaneously, it avoids over-filtering that could blur the target features, balancing the filtering effect with the real-time requirements of port monitoring. First, the signal-to-noise ratio (SNR) is calculated for each branch of the enhanced dual-modal features. The SNR is used to quantify the ratio of effective signal to noise in the features. The specific mathematical expression is as follows: Where μ is the global mean of all pixel values in a single feature map, reflecting the overall strength of the feature signal, and is calculated as follows: ; In the formula, H=512 and W=640 are the feature map dimensions; σ is the global standard deviation of all pixel values in a single feature map, reflecting the dispersion of noise in the feature, and is calculated as follows: A lower SNR indicates more noise in the features, requiring stronger filtering. The filtering strength and coefficients are then dynamically adjusted based on the calculated signal-to-noise ratio. The adjustment follows the following mathematical expression: Where σ0 is the initial basic filter coefficient of 1.0, this formula ensures that the lower the SNR, the more noise there is, and the filter coefficient... The larger the SNR, the stronger the filtering strength, and the more effectively it can filter Gaussian noise in low-light areas; the higher the SNR (the clearer the features), the stronger the filtering coefficient. The smaller the value, the weaker the filtering strength, reducing the loss of target features.
[0045] The filtering process uses a 3×3 local window Gaussian filtering method. The specific mathematical expression for the filtered value of each pixel is as follows: ; in, After filtering The pixel value of the location, To enhance the 3×3 local window in the feature map The pixel value of the location, The filter coefficients are dynamically adjusted. This calculation achieves noise filtering through a weighted average, with the weights determined by the filter coefficients. Decide, The larger the value, the higher the weight ratio of neighboring pixels, and the more obvious the filtering effect. During the filtering process, edge pixels of the feature map are filled by copying edge pixels to avoid filtering distortion in edge areas. After adaptive filtering is performed on the infrared and visible light feature branches respectively, the filtered dual-modal feature (F) is obtained. ir_filter F vis_filter The dimensions remain [B, 64, 640, 512], and Gaussian noise in the features is effectively suppressed. The core features of port personnel, luggage and other targets are fully preserved, laying a high-quality foundation for subsequent fusion processing.
[0046] Step 4.3: Input the filtered dual-modal features into the preliminary fusion submodule of the luminance feedback network to perform weighted fusion of infrared and visible light features, obtaining preliminary fused features. Specifically, this includes: inputting the filtered dual-modal features and the filtered infrared feature F... ir_filter Compared with the filtered visible light feature F vis_filterThe dimension [B, 64, 640, 512] is the initial fusion submodule of the input brightness feedback network. The core function of this submodule is to perform preliminary weighted fusion of the two types of denoised features, achieving complementary advantages between the two modal features. Simultaneously, it provides stable feature support for the subsequent adaptive weight acquisition module to accurately fuse occluded areas. First, a sharpness score is calculated for each of the two types of filtered features. The sharpness score quantifies the quality of the features and serves as the core basis for fusion weight allocation, ensuring that features with higher sharpness have a higher proportion in the fusion process, fully leveraging the advantages of both types of features. Among them, the sharpness score of infrared features focuses on the identification of heat source targets, and the specific mathematical expression is: Among them, S ir For infrared feature sharpness scoring, P ir_peak M represents the peak response of the heat source target in the filtered infrared feature map, i.e., the maximum pixel value obtained after traversing all pixels of the infrared feature map. ir_filter The mean global response of the filtered infrared feature map is calculated as follows: ; A higher peak value and a lower mean value indicate a clearer infrared heat source target, resulting in a higher score. The clarity score for visible light features focuses on the richness of texture details, and the specific mathematical expression is as follows: ; Svis represents the visible light feature sharpness score. This is the sum of the local gradient values of all pixels in the filtered visible light feature map. The local gradients are calculated using a 3×3 Sobel operator, taking the sum of the absolute values of the gradients in the x and y directions, N. pixel The total number of pixels in the feature map is 640 × 512 = 327680. The higher the sum of the gradients, the richer the visible light texture details, and the higher the score. Then, the fusion weights of the two types of features are calculated based on the sharpness score, ensuring that the sum of the weights is 1. The specific mathematical expression is as follows: , ;where w ir For infrared feature fusion weights, w vis For visible light feature fusion weights, after calculation, the weights are constrained to ensure that 0 ≤ w ir ≤1、0≤w vis ≤1, to avoid fusion distortion caused by abnormal weights; this weight allocation logic is adapted to port scenarios: when visible light is blurred in low-light areas, S vis Lower, w vis Decrease, w ir Automatic lifting, prioritizing the preservation of infrared heat source characteristics; when there is sufficient light, S vis rise, w visIncrease the size while prioritizing the preservation of visible light texture details; finally, perform a weighted fusion calculation, the specific mathematical expression of which is: ; in, To initially fuse features, this calculation precisely integrates the core advantages of the two types of features. After fusion, the features are smoothed to reduce abrupt changes at feature edges, generating an initial fused feature with dimensions maintained at [B, 64, 640, 512]. This feature contains both the clear outline of the heat source target from infrared features and the rich texture details from visible light features, effectively compensating for the shortcomings of single-modal features and providing high-quality feature input for the subsequent adaptive weight acquisition module to further optimize the occluded area.
[0047] In a preferred embodiment of the present invention, step 5 above may include: Step 5.1: Input the preliminary fusion features and the corresponding preprocessed standardized visible light image into the occlusion detection unit of the adaptive weight acquisition module to perform occlusion region detection and generate an occlusion mask. Specifically, this includes: the adaptive weight acquisition module is a core module used to detect occlusion regions in the visible light image and dynamically allocate the fusion weights of infrared and visible light features based on the occlusion situation, including an occlusion detection unit, a dual attention branch, and a dynamic weight allocation unit; input the generated preliminary fusion features F fusion_init and the corresponding preprocessed normalized visible light image F vis_std The occlusion detection unit, which is jointly input into the adaptive weight acquisition module, aims to accurately identify occlusion areas (mainly luggage, backpacks, and crowding) in port clearance scenarios and generate occlusion masks for subsequent weight allocation. The occlusion detection unit first extracts texture features from the standardized visible light image and calculates the texture contrast of each pixel using a 3×3 local window. The specific mathematical expression is as follows: ; in, for Texture contrast at location, These are the pixel values within a 3×3 local window of the visible light image. This represents the average pixel value of the local window. The lower the texture contrast, the more likely the area is to be occluded, such as when the texture is blurred in an area obscured by luggage.
[0048] At the same time, the preliminary fusion feature F is extracted. fusion_init The target response value R(x,y) is calculated as follows: (Take the maximum response value of all channels at the current position). The response value will significantly decrease in occluded areas because the target features are masked. Then, calculate the occlusion probability for each pixel. The mathematical expression is: Where k=5 is the probability scaling coefficient, calibrated using port occlusion samples; λ=0.6 is the response value weight; and θ=0.3 is the occlusion judgment threshold, determined based on statistics from over 1000 sets of port occlusion images. The value range is [0,1], and the closer the value is to 1, the higher the occlusion probability; finally, a mask generation threshold of 0.5 is used, when P occ When (x,y)≥0.5, it is determined to be an occluded region, and the mask value is set to 1; when P occ When (x,y)<0.5, it is determined to be a non-occluded region, the mask value is set to 0, and a binary occlusion mask M with the same size as the input image ([B,1,640,512]) is generated to complete the occlusion region detection.
[0049] Step 5.2 involves inputting the preliminary fused features into the dual attention branch of the adaptive weight acquisition module, and performing attention enhancement processing on the infrared and visible light features respectively to obtain the attention-weighted dual-modal features. Specifically, this includes: inputting the preliminary fused features F... fusion_init The dual attention branch of the adaptive weight acquisition module is input, which includes an infrared feature attention sub-branch and a visible light feature attention sub-branch. Its core function is to enhance the key regions of infrared heat source features and visible light texture features respectively, weaken invalid background interference, and adapt to the complex distribution of targets and backgrounds in port scenes. First, the infrared feature branch F is extracted from the preliminary fused features through a channel separation mechanism. ir_branch The first 32 channels and the visible light characteristic branch F vis_branch For the last 32 channels, attention enhancement is performed separately for each of the two types of features. For the infrared feature attention sub-branch, a channel attention mechanism is used to focus on enhancing the channels corresponding to the heat source target, and the channel attention weight W is calculated. ir_channel The mathematical expression is: ; Where σ is the Sigmoid activation function, which normalizes the weights to [0,1], AvgPool is global average pooling, which extracts global features for each channel, MLP is a two-layer fully connected network (which restores channel dimensions after compression, enhancing channel correlation), and W ir_channel The dimension is [B, 32, 1, 1], used to characterize the importance of each infrared channel; then, the channel attention weights are multiplied by the infrared feature branches channel by channel to obtain the infrared feature F after channel attention enhancement. ir _ channel _ att .
[0050] Based on this, a spatial attention mechanism is added to calculate the spatial attention weight graph W. ir_spatial The mathematical expression is: ; Wherein, Conv3×3 is a 3×3 convolution, which integrates global max pooling and average pooling features, W ir_spatial The dimension is [B, 1, 640, 512], used to characterize the importance of infrared features at each spatial location, and is compared with F. ir _ channel _ att Pixel-by-pixel multiplication yields the attention-weighted infrared feature F. ir_att For the visible light feature attention sub-branch, the same channel + spatial attention structure is used, focusing on enhancing the channels and spatial positions corresponding to texture details, and calculating the channel attention weight W. vis_channel Spatial attention weight map W vis_spatial Finally, the attention-weighted visible light feature F is obtained. vis_att F ir_att With F vis_att All dimensions are maintained at [B, 32, 640, 512] to ensure dimensional matching during subsequent fusion.
[0051] Step 5.3: Based on the occlusion mask, the dynamic fusion weights of infrared and visible light features are calculated in the dynamic weight allocation unit of the adaptive weight acquisition module. Specifically, based on the generated occlusion mask M, the dynamic weight allocation unit of the adaptive weight acquisition module, combined with the initialized preset fusion weights, calculates the dynamic fusion weights of infrared and visible light features adapted to the port occlusion scenario, ensuring that both the occluded and unoccluded areas can achieve optimal complementarity of dual-modal features. First, the sharpness score of the attention-weighted dual-modal features is extracted, and the sharpness score S of the infrared features is calculated. ir_att (Focusing on heat source target identification) Calculation method is as follows Visible light feature sharpness score (Focusing on texture detail richness) Calculation method is as follows Where grad is the gradient calculation and num is the total number of pixels; then, the occlusion mask is combined. For the area to be covered, For the unoccluded region, the dynamic fusion weight is calculated, and the specific mathematical expression is as follows: ; in, For infrared feature dynamic weights, For visible light features, dynamic weights Preset weights for the initialized occlusion area. Preset weights for the initial unoccluded areas. and The formula scores the clarity of attentional features at the current location; the core logic of this formula is: occluded area. When infrared features are prioritized, their weights are further adjusted by combining them with infrared attention scores to ensure that heat source targets are captured even when occlusions are blocked. When the unobstructed region M(x,y)=0, visible light features are prioritized, and texture details are preserved by combining them with visible light attention scores. After calculation, the dynamic weights are subject to range constraints to ensure that... , To avoid fusion distortion caused by abnormal weights, a dynamic weight map with the same size as the occlusion mask [B,1,640,512] is finally obtained. and .
[0052] Step 5.4: Based on the dynamic fusion weights, perform weighted fusion on the attention-weighted bimodal features to output the final fused features. Specifically, this includes: calculating the dynamic fusion weights... and The attention-weighted bimodal features F ir_att With F vis_att The core purpose of pixel-by-pixel weighted fusion is to integrate the advantages of the two types of features to generate a final fused feature that is suitable for complex port scenarios and can be directly used for subsequent abnormal behavior identification. First, the dynamic weight map w... ir_dyn with w vis_dyn Each feature is multiplied pixel-by-pixel with its corresponding attention feature to ensure that the feature weights at each spatial location are accurately matched with the occlusion status and attention score. The specific calculation is as follows: ; in Infrared attention-weighted features, These are visible light attention-weighted features, all with dimensions [B, 32, 640, 512].
[0053] Subsequently, a weighted fusion operation is performed to generate the final fused feature, the specific mathematical expression of which is: ; Here, ⊕ represents the channel concatenation operation, which concatenates the infrared and visible light weighted features by channel, resulting in the dimension [B, 64, 640, 512]. Conv1×1 is a 1×1 convolution used to fuse channel features, reduce feature redundancy, and ensure that the feature dimension is compatible with subsequent recognition modules. For the final fused features, the dimensions are maintained at [B, 64, 640, 512]. After fusion, the final fused features are smoothed using a 5×5 Gaussian filter with a standard deviation σ=1.5 to eliminate abrupt changes in feature edges and ensure coherent feature contours and clear details. This fusion logic is well-suited for port scenarios: occluded areas (such as luggage obscuring people) are dominated by infrared attention features, clearly presenting the heat source contours of people; unoccluded areas are dominated by visible light attention features, fully preserving texture details such as clothing and movements, effectively compensating for the shortcomings of single-modal features. Simultaneously, 1×1 convolution optimizes the feature dimensions, providing high-quality, highly adaptable feature input for subsequent abnormal behavior recognition modules, and finally outputting the final fused features for use in subsequent steps.
[0054] In a preferred embodiment of the present invention, step 6 above may include: Step 6.1: Based on the final fused features and the preprocessed standardized bimodal image data, calculate the joint loss function value. The joint loss function value consists of no-reference color loss, feature consistency loss, and structural similarity loss, specifically including: based on the generated final fused features... Dimensions [B, 64, 640, 512] and preprocessed standardized bimodal image data, standardized infrared image Standardized visible light images The joint loss function is calculated, which is a weighted average of the no-reference color loss, feature consistency loss, and structural similarity loss. The specific mathematical expression is as follows: ; in, Loss weights calibrated for port scenarios; no-reference color loss. The method used to constrain the color deviation between the fused image and the visible light image is as follows: ; in, To fuse the YUV color channels of an image, To standardize the YUV color channels of visible light images, For the desired operation; feature consistency loss The information deviation between the fused features and the original bimodal features is constrained and calculated as follows: ; in the formula L2 norm; structural similarity loss The method used to constrain the structural deviation between the fused image and the visible light image is as follows: SSIM is the structural similarity index, and its calculation logic is completely consistent with that in step 3.2. The final output is the weighted joint loss function value. .
[0055] Step 6.2: Based on the joint loss function value, calculate the gradient using backpropagation, and update the network parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module to obtain the parameter-updated model. Specifically, this includes: calculating the gradient based on the obtained joint loss function value... The backpropagation algorithm is initiated; gradient calculation is strictly passed along the forward inference reverse path of the final fused feature, adaptive weight acquisition module, brightness feedback network, and fusion and enhancement network to accurately locate the contribution of each trainable parameter to the loss; among them, the trainable parameters of the fusion and enhancement network include the enhancement gain parameter γ and the weights of each convolutional layer; the trainable parameters of the brightness feedback network include the filter coefficients. The convolutional weights of the feature enhancement and filtering sub-modules; the trainable parameters of the adaptive weight acquisition module include the convolutional weights of the dual attention branch and the dynamic weight calibration parameters; subsequently, the Adam optimizer is used to perform parameter updates, with the optimizer's initial learning rate set to... A learning rate decay mechanism is introduced: every 20 iterations, the learning rate is decayed to 0.8 times the current value, thus balancing the rapid convergence in the early stages of training with the stability in the later stages; the mathematical expression for parameter update is: ; in the formula It is the set of trainable parameters of the model in the current iteration, including all convolution weights, enhancement gain parameters, filter coefficients, etc. This is the new set of parameters obtained after this update; This is the learning rate of the Adam optimizer, which controls the step size of parameter updates. The smaller the value, the smoother the parameter changes; the larger the value, the faster the convergence but the more prone to oscillation. It is the joint loss function with respect to the current parameters The gradient reflects the degree of influence of small parameter changes on the loss value. The gradient direction is the direction in which the loss increases the fastest, so the negative direction is used during the update to reduce the loss. During the update process, range constraints are imposed on the core parameters: the enhancement gain parameter γ is constrained to the range [0.5, 2.0] to avoid excessively high gain leading to overexposure of the image or excessively low gain leading to insufficient enhancement of low-light areas; the filter coefficients... The constraints are set in the range [0.1, 2.0] to prevent over-filtering from losing target features or under-filtering from leaving residual noise, and finally obtain the fusion model with updated parameters.
[0056] Step 6.3 involves processing the validation set data using the updated model to obtain the corresponding validation performance metrics. Specifically, this includes: processing the validation set data using the updated model; the validation set consists of 1000 sets of bimodal images of port scenes, covering typical scenarios such as low-light conditions at night, luggage obstruction, and crowding, ensuring that the validation results accurately reflect the model's adaptability in the actual port environment; the input data dimensions are completely consistent with the training set: standardized infrared images are [B, 1, 640, 512], and standardized visible light images are [B, 3, 640, 512]; the model will fully execute the inference process of preprocessing, feedback enhancement, lightweight fusion, and adaptive weight allocation to generate the final fusion features of the validation set; subsequently, validation performance metrics are calculated from three dimensions to comprehensively evaluate the model's performance: the first is the peak signal-to-noise ratio (PSNR), calculated as follows: ; PSNR is used to quantify the pixel-level error between the fused image and the reference image; a higher value indicates a smaller pixel deviation. MAX is the normalized maximum value of the image pixels. Since the input image has been normalized to the [0, 1] interval, it is set to 1. MSE is the mean square error between the fused image and the reference image. It is calculated as the average of the squares of the differences between the corresponding pixel values of the two images, reflecting the overall deviation of the pixel values. Then, the structural similarity SSIM is completely consistent with the calculation logic in step 3.2. It is used to evaluate the structural consistency between the fused image and the visible light image. The closer the value is to 1, the more complete the structure is preserved. The port personnel detection accuracy mAP is calculated based on the personnel annotation data of the validation set. It reflects the support capability of the fused features for abnormal behavior recognition through the average accuracy of the downstream detection tasks. The higher the value, the stronger the practicality of the features. Finally, the performance index set of the validation set is obtained, which provides a quantitative basis for the training termination judgment.
[0057] Step 6.4: Determine whether the training termination condition is met based on the validation performance index. If met, the current model is used as the optimized fusion model. If not met, return to recalculate the joint loss function value for the next round of iteration training. Specifically, this includes: determining whether the training termination condition is met based on the obtained validation performance index; the termination condition is designed to balance model convergence and generalization ability: Condition 1: The validation set SSIM does not improve for 3 consecutive rounds. This condition effectively prevents model overfitting and avoids the model from being over-optimized on the training set and losing its generalization ability to unknown scenes. Condition 2: The number of iterations reaches 100 rounds. This condition ensures training efficiency and avoids the waste of computational resources caused by unlimited iteration. If any termination condition is met, the model with updated parameters is saved as the optimized fusion model. The saved content includes: all trainable parameters, network structure configuration file, and initialized scene adaptation parameters, such as the preset weight α of the occlusion region. occ Preset weight α for unobstructed areasnonocc This facilitates subsequent deployment and inference; if the termination condition is not met, return to step 6.1, recalculate the joint loss function value using the current model, and continue to execute the next round of iteration training until the model converges to optimal performance.
[0058] In a preferred embodiment of the present invention, step 7 above may include: Step 7.1 involves inputting the final fusion features into the optimized fusion model for processing to obtain an image matrix corresponding to the final fusion features. Specifically, this includes: processing the generated final fusion features F... final The feature has dimensions [B, 64, 640, 512] and includes infrared thermal features and visible light texture features after attention enhancement and dynamic weight fusion. Pixel values are stably in the [0, 1] range, accurately carrying the core information of targets such as personnel, luggage, and channel backgrounds in the port scene. It is input into the optimized fusion model for specialized processing. The core purpose is to transform the high-dimensional fusion features into a two-dimensional image matrix that can be directly displayed and analyzed, ensuring that the fused features can be effectively transformed into a visualized image. The optimized fusion model has undergone parameter calibration for the port scene. During processing, a built-in 1×1 convolutional layer is first called to channel the final fused features. Dimensionality reduction precisely compresses the 64-channel high-dimensional features to 3 channels. This process fully preserves the core advantages of bimodal features: it does not lose the heat source contour information of infrared features, nor weaken the color and texture details of visible light features. At the same time, the Sigmoid activation function further constrains all feature values, ensuring that the feature values after dimensionality reduction remain stable in the [0,1] range, avoiding image distortion caused by numerical anomalies. Subsequently, a feature reshaping operation is performed, reshaping the 3-channel features after dimensionality reduction according to the standard format of an image matrix, transforming the original feature dimension [B,64,640,512] into an image matrix I with the same size as the input feature space. matrix The dimensions are [B, 3, 640, 512], where the three channels correspond to the R (red), G (green), and B (blue) components of a conventional visualization image, respectively. The value of each pixel precisely corresponds to the color intensity at that location, which can clearly present the heat source outline and clothing texture of people in the port scene, the edge details and placement of luggage, and the overall environment of the channel background. Finally, an image matrix corresponding to the final fusion features can be directly used for subsequent denormalization processing. Step 7.2: Perform denormalization on the image matrix. Combining the pixel extreme value parameters recorded in the preprocessing stage, restore the normalized pixel values to their original numerical range to obtain the denormalized image. Specifically, this includes: processing the obtained image matrix I... matrixPerforming denormalization is the complete inverse of the normalization process in step 1 (preprocessing). Its core purpose is to combine the pixel extreme value parameters recorded and uniformly saved in the preprocessing stage to accurately restore the normalized pixel values to the numerical range of the original image. This ensures the image can reproduce the brightness levels and color details of the original scene, meeting the display standards of conventional images and the requirements of subsequent analysis. The pixel extreme value parameters recorded in the preprocessing stage are the globally maximum pixel values obtained by statistically analyzing all original dual-modal images (original infrared images and original visible light images) in the training and test sets. and global minimum pixel value ,in , It fully conforms to the standard pixel value range of conventional RGB images, and this parameter is saved together with the model parameters to ensure the uniformity and accuracy of the denormalization process.
[0059] The specific mathematical expression for destandardization is: ; In the formula, For the final pixel value, For the normalized pixel values of the fused features, , These are the pixel extreme values of a single image recorded during the preprocessing stage. The implementation logic is as follows: First, the fused feature tensor is converted into an image matrix using a tensor-image matrix format conversion algorithm, and then the recovered value is calculated pixel by pixel using the formula. Next, a pixel value range cropping and bit-width conversion algorithm is used to constrain the pixel values to the range of 0~255, while simultaneously converting them to 8-bit integers to avoid image distortion caused by pixel value overflow.
[0060] Step 7.3 involves performing edge smoothing filtering on the denormalized image to eliminate image edge artifacts and generate the final fused image. Specifically, this includes: processing the obtained denormalized image I... denormThe core purpose of edge smoothing filtering is to eliminate edge artifacts generated during multi-round image fusion. These artifacts mainly originate from abrupt changes in edge pixels caused by operations such as dual-modal feature stitching, dynamic weight allocation, and channel dimensionality reduction. Specifically, they manifest as jagged edges, uneven light-dark boundaries, and subtle discontinuities in the fusion area. Without processing, these artifacts will severely affect the accuracy of subsequent identification of targets such as people and luggage at ports of entry, failing to meet the actual needs of identifying abnormal behavior at ports. The filtering process uses a 3×3 mean filtering method, identical to steps 4.3 and 5.4 above. This ensures the coherence of the entire invention's technical logic and the consistency of terminology, while also accurately adapting to the processing needs of fused images and ensuring the stability of the filtering effect. During the filtering process, a 3×3 filtering window slides pixel by pixel along the image from left to right and from top to bottom. The filtered value of each pixel is the average of its own pixel value and the pixel values of the eight pixels in its 3×3 neighborhood. Through this averaging calculation with equal weights, it can effectively smooth abrupt changes in edge pixels, weaken the impact of artifacts, and preserve the core details of the image to the maximum extent, avoiding image blurring caused by over-filtering.
[0061] To address the issue of edge pixels not being fully covered by the 3×3 filtering window, a method of copying edge pixels is used. Specifically, a 1-pixel-wide copy of the edge pixels is added to the image edge to ensure that the filtering window can fully cover 9 pixels when sliding to the edge. This avoids distortion and blurring at the edge due to incomplete filtering and ensures that the filtering effect is consistent between the image edge and the center area. After edge smoothing filtering, various edge artifacts in the image are effectively eliminated, and the fusion of infrared heat source contours and visible light texture details is more natural and harmonious. The contours of port personnel are smoothly transitioned, clothing textures are clearly discernible, the edge details of luggage are complete and without breaks, and the background of the passage is uniform and free of noise. The final fused image with dimensions [B,3,640,512] is generated, which is fully adapted to the requirements of port scene image display and subsequent abnormal behavior recognition.
[0062] The proposed LENFSion technology, a joint low-light enhancement and fusion technique for infrared and visible light image fusion, aims to address the issues of synergy and adaptability in dual-light image fusion under obstructed conditions during port clearance. The entire technical process follows a closed-loop logic of "data input - preprocessing - scene adaptation - core collaborative processing - training optimization - output post-processing," encompassing seven core functional parts. Each part and its corresponding modules form a tightly linked technical chain according to the code execution sequence. The overall workflow can be summarized as follows: (1) Data preparation stage: The infrared and visible light image pairs of the port clearance scene are acquired through the data acquisition module, and after standardization and enhancement by the preprocessing module, they are input into the system.
[0063] (2) Initialization and adaptation stage: The port scenario adaptation module loads exclusive configuration parameters and completes the initialization and adaptation of subsequent core modules.
[0064] (3) Core processing stage: The brightness calibration features output by the brightness adjustment network are input into the brightness feedback network. After brightness evaluation and interference filtering, they are input into the intermediate fusion module to obtain the preliminary fusion features. The preliminary fusion features and the output calibration features are input into the fusion and enhancement network. The enhancement and fusion parameters are adjusted through further enhancement optimization and feedback closed loop. The optimized features are input into the adaptive weight acquisition module to complete the accurate fusion of dual-modal features in the occlusion scene.
[0065] (4) Optimization and output stage: The network parameters are iteratively optimized based on the training optimization module with no reference color loss to ensure the fusion effect and color restoration accuracy; finally, the post-processing module outputs a high-quality fused image to support the detection of abnormal behavior at the port.
[0066] Detailed implementation of key modules: Data Input and Preprocessing Module This module serves as the starting point of the entire technical process. Its core function is to complete the entire data preparation chain, from "raw data acquisition" to "data quality verification," "standardization adaptation," and "enhanced generalization." It provides subsequent core modules with input data that is formatted uniformly, of acceptable quality, and possesses strong generalization capabilities. All functions are implemented through a closed-loop code structure. Classes and functions interact via parameter passing. Specifically, it comprises three sub-modules, each with clearly defined code mappings, as detailed below: (1) Data Acquisition Submodule: The core function is to realize the synchronous acquisition and storage of infrared and visible light images, ensure the timestamp alignment of the dual-light images, and avoid the misalignment of subsequent fusion features due to asynchronous acquisition. At the hardware control level, the infrared camera and the visible light camera are activated by the camera calling method, and the synchronous trigger signal is generated by the hardware synchronous trigger signal generation algorithm to realize synchronous shooting; at the data processing level, the module embeds an image synchronous reading unit, which integrates timestamp verification processing logic. The shooting time in the image EXIF information is parsed by EXIF metadata parsing and timestamp synchronization algorithm, and the time difference of the dual-light images is calculated. If the difference exceeds 10ms, the data is discarded by the invalid data filtering and discarding algorithm; at the storage level, after reading, the data is stored in the specified folder by the image data persistence storage algorithm. The storage format is PNG to ensure that the image is not compressed and distorted. The input of this submodule is the raw image data stream transmitted by the camera hardware, and the output is the timestamp-aligned infrared-visible light image pair. All input and output forms are constrained by data type.
[0067] (2) Standardization Submodule: The core function is to eliminate the difference in pixel value range between infrared and visible light images, and to normalize the two types of images to the [0,1] interval, so as to avoid the influence of the difference in pixel value magnitude on the gradient of subsequent convolution operations. The specific implementation logic is to first obtain the minimum pixel value of a single infrared image and a visible light image respectively. With the maximum value Then, normalization is achieved through pixel-by-pixel calculation; the function integrates a pixel extreme value extraction unit and a pixel-by-pixel normalization calculation unit, receives the image matrix as input, outputs the normalized image matrix, and simultaneously records the image's... and The core formula is: ; Where X is a single pixel value of the original image. and These are the minimum and maximum pixel values for a single image, respectively. For infrared images, the 16-bit data must first be converted to 32-bit floating-point numbers using a data type bit-width conversion algorithm before calculation to avoid overflow. To obtain the standardized pixel values, the code improves efficiency through a double loop or vectorized parallel operation algorithm and a single instruction multiple data vectorized parallel operation optimization algorithm. It traverses the rows and columns of the image to achieve pixel-by-pixel normalization calculation, and integrates an outlier handling unit.
[0068] (3) Data augmentation submodule: The core function is to improve the model’s generalization ability to changes in lighting and target posture in port clearance scenarios through diversified data transformations. The function integrates an enhancement strategy triggering unit and three types of data transformation processing units. It receives standardized image pairs as input and outputs enhanced image pairs. The triggering of the enhancement strategy is controlled by a uniformly distributed random number generation algorithm. The three enhancement strategies are: ① Random cropping: The starting coordinates of the cropping are determined by a random coordinate generation algorithm. The 480×480 size cropping is achieved through array slicing. After cropping, the image is restored to 640×512 by a bilinear interpolation image scaling algorithm to ensure uniform input size; ② Horizontal flipping: A random number between 0 and 1 is generated. If the random number is < 0.5, the flipping is triggered. The horizontal mirroring transformation is achieved by an image horizontal mirroring transformation algorithm. The infrared and visible light images are synchronized through a parameter synchronization transmission unit to ensure synchronous flipping; ③ Brightness perturbation: The infrared or visible light is distinguished by a channel number judgment unit. The image is converted to the HSV color space by an RGB-HSV color space conversion algorithm. After adjusting the V channel, it is converted back to the RGB space. A perturbation coefficient within the range of ±0.1 is generated by a uniformly distributed random number generation algorithm to avoid color distortion.
[0069] The brightness adjustment module is the core pre-processor for initial enhancement in low-light scenes. Located after data preprocessing and before the brightness module, its core function is to receive standardized infrared and visible light dual-modal input images, perform initial enhancement through an adaptive brightness mapping function and pixel-level nonlinear adjustment, and output a standardized dual-modal image with calibrated brightness. This provides high-quality input data for the subsequent brightness evaluation and fusion modules of the brightness feedback module. This module employs a reference-free color loss constraint and a training optimization module in conjunction, enabling adaptive brightness adjustment without relying on labeled data. This adapts to the dynamic changes in complex lighting scenes at ports of entry. Details are as follows: Adaptive Brightness Mapping Subunit: Its core function is to achieve differentiated brightness adjustment based on the illumination characteristics of dual-modal images. It receives standardized infrared and visible light images as input, splits the visible light RGB channels using channel separation and merging algorithms, and performs adaptive mapping between the V channel and the grayscale values of the infrared image.
[0070] The dual-modal brightness equalization subunit's core function is to balance the brightness differences between infrared and visible light images, preventing single-modal brightness imbalances from causing subsequent feature fusion deviations. It receives the brightness-adjusted dual-modal image as input, calculates the average brightness of the two images using a brightness mean alignment algorithm, fine-tunes the infrared image brightness based on the average brightness of the visible light image, ensuring the dual-modal brightness mean deviation is ≤0.05, and outputs a brightness-equalized dual-modal feature map.
[0071] Output verification subunit: Its core function is to verify the quality of the feature map after brightness adjustment, avoiding over-enhancement or under-enhancement. By calculating image contrast and detail retention, if the contrast is <0.1 or the gradient magnitude is <0.02, the coefficient readjustment logic is triggered. After the verification is passed, the feature map is transmitted to the brightness feedback network module, and the brightness adjustment parameters are stored synchronously for feedback and use by the fusion and enhancement network modules.
[0072] The port scenario adaptation module is the initialization and adaptation phase of the code execution. Its core function is to establish a mapping relationship between "port scenario prior information" and "core module parameters." By loading scenario-specific configurations and initialization parameters, the network adapts to the unique characteristics of a port scenario—high population density and occlusion interference—from the initial training stage, avoiding initial training biases caused by generalized parameters. It includes three core functional units: The scenario prior configuration submodule's core function is to read and parse the port scenario-specific configuration file, providing a scenario-based basis for parameter initialization. This is implemented in the code through a configuration loading unit, which receives the configuration file path as input, reads the configuration file through file reading and parsing logic, and outputs a dictionary containing all prior information. The configuration file stores key prior information categorized as "Regional Information - Device Parameters - Scene Characteristics," each with a specific purpose: ① Inspection channel area coordinate range: Subsequent core modules mark key areas using image region marking algorithms and extract features from these areas using image region ROI extraction algorithms, ignoring invalid background outside the channel; ② Baggage obstruction high-frequency area marking mask: This is a 128×128 binary matrix. Areas with a mask value of 1 are high-frequency baggage obstruction areas. The code loads the mask file using a binary mask file loading and parsing algorithm and passes it to the obstruction detection unit of the adaptive weight acquisition module to guide the detection of key areas; ③ Dual-light camera exposure parameters: Infrared exposure time is 10ms, and visible light exposure time is 5ms. Camera parameters are set using a camera parameter standardization configuration algorithm to calibrate the imaging quality of the input image and assist in standardization processing; ④ Scene noise characteristics: The noise type of the port night image is Gaussian noise, which is passed to the adaptive filtering unit of the brightness feedback network module as the initial setting basis for the filtering coefficients.
[0073] The parameter initialization submodule's core function is to initialize and assign values to the learnable parameters and hyperparameters of four core modules—the brightness adjustment network, the LFN brightness feedback network, the fusion and enhancement network, and the adaptive weight acquisition module—based on prior scene information, ensuring that the initial parameter values align with the characteristics of the port scene. This unit receives the configuration dictionary output by the scene configuration loading unit as input and outputs a parameter dictionary containing parameters for all core modules, specifically as follows: ① Adaptive weight acquisition module parameters: Based on the "luggage occlusion" characteristic, these parameters are initialized and assigned using a parameter initialization algorithm, serving as the initial values for infrared feature weights in occluded or unoccluded areas, respectively; ② Fusion and enhancement network module parameters: Based on the "low light" characteristic, these parameters are initialized and assigned using a parameter initialization algorithm, serving as the initial values for the enhancement gain coefficient and the SSIM evaluation threshold, respectively; ③ Brightness feedback network module parameters: Based on the Gaussian noise variance of 0.02, the adaptive filter base coefficients are assigned using a parameter initialization algorithm, the convolution kernel weights are initialized using a He normal distribution parameter initialization algorithm, and the bias term is set to 0.1 using a constant term initialization algorithm to avoid initial gradient vanishing. ④ Brightness Adjustment Network Module Parameters: Based on illumination characteristic parameters, the initial brightness adjustment coefficients are a=1.0, b=1.0, and c=0.05 to adapt to the initial enhancement requirements of the low-light scene at the port. The parameter dictionary is stored in key-value pair format, and other modules call the corresponding parameters through key_name to ensure the standardization of parameter transmission.
[0074] The Fusion and Enhancement Network module is the core component for achieving "fusion-enhancement bidirectional guidance." It breaks down the barrier of "independent processing of enhancement and fusion" in traditional technologies by constructing a closed-loop mechanism of "fusion feature output → further enhancement optimization → dynamic parameter adjustment." This makes the brightness adjustment parameters of the brightness adjustment network module dependent on the fusion effect, and the strategy optimization of the fusion module references the brightness enhancement quality and brightness feedback results, forming a collaborative optimization closed loop. All functions are implemented through a class that inherits from the network base class. This class requires overriding the initialization unit and forward propagation unit, supporting forward propagation and backward gradient transfer. Its execution order is after the brightness adjustment network and the LFN (brightness feedback network, intermediate fusion module, and adaptive weight acquisition module). Its core function is to implement secondary feature enhancement and fusion optimization, dynamically adjusting the brightness adjustment parameters of the brightness adjustment module and the weight allocation coefficients of the fusion module to ensure precise adaptation between the enhancement effect and fusion requirements. This module contains three sub-modules, which interact through tensor data. All functions have explicit code implementations, as detailed below: The signal interaction submodule's core function is to construct a bidirectional data transmission link between the brightness adjustment network module, the brightness feedback network module, and the intermediate fusion module. This ensures the real-time performance and integrity of data transmission, preventing feedback adjustment failures caused by data delays. This functionality is implemented through a class-based signal interaction unit. This unit achieves data sharing through instance variables of the class, employing a tensor high-speed communication mechanism. The specific interaction logic is as follows: ① Forward transmission: The brightness calibration feature map and brightness adjustment parameters output by the brightness adjustment module, the brightness feedback signal output by the brightness feedback module, and the preliminary fusion feature map output by the intermediate fusion module are stored in instance variables of the class using a tensor data format adaptation algorithm in the data storage unit. These are then available for use by the enhancement effect evaluation submodule and the parameter adjustment submodule. ② Reverse transmission: The optimized parameters output by the parameter adjustment submodule are stored in the data storage unit and fed back to the brightness enhancement module, brightness feedback module, and intermediate fusion module, respectively. To ensure data synchronization, a data verification unit is added to the submodule. This unit verifies dimensions using assertions. If a mismatch is detected, a warning is triggered, and the current batch of training is terminated via a training process abnormal termination mechanism.
[0075] Enhancement effect evaluation submodule: Its core function is to quantitatively evaluate the initial enhancement quality of the brightness adjustment network module, the feedback control effect of the brightness feedback network module, and the preliminary fusion effect, providing an objective basis for parameter adjustment and avoiding parameter adjustment deviations caused by subjective judgment. It uses structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) as dual evaluation indicators. This unit receives the enhanced image, the image after feedback control by the brightness feedback network, the preliminary fused image, and the original visible light image as input, and outputs SSIM and PSNR values. The core formula is: ; Where x represents the enhanced image and y represents the original visible light image. These are the local means of the two images, respectively. , Let be the local standard deviation of the two images. Let the local covariance of the two images be denoted as . =6.5025、 =58.5225 is a fixed constant. In the code, local statistics are calculated using a pixel-by-pixel local convolution statistical algorithm, and then substituted into the SSIM formula to complete the calculation. When the calculated SSIM < 0.75, the optimization logic of the parameter adjustment submodule is triggered by the threshold comparison algorithm.
[0076] The parameter adjustment submodule's core function is to dynamically optimize the enhancement gain coefficient based on the enhancement effect evaluation results and the feedback error from the fusion module, achieving coordinated adaptation between enhancement and fusion. The corresponding parameter adjustment unit in the code receives the current enhancement gain coefficient, SSIM value, and fusion error as input, and outputs the updated enhancement gain coefficient using the following formula: ; in, For the updated enhanced gain coefficient, This represents the current enhancement gain coefficient. The feature matching error of the fusion module is represented by 1-SSIM, and the enhancement effect deviation value is 0.1 and 0.05 in the formula. The formula calculation is implemented through a dynamic update algorithm for the gain coefficient. After calculation, the coefficient boundary constraint algorithm is applied, and the updated coefficients are stored in the data storage unit and transmitted in real time to the enhancement module for the next round of image enhancement.
[0077] The brightness feedback network module, at its core, implements the functions of "brightness assessment - interference filtering - feedback adjustment." It overcomes the interference accumulation problem caused by the independent separation of enhancement, feedback, and fusion in traditional technologies. Through a serial link structure, the output brightness calibration features complete "brightness verification - interference filtering - feedback generation" within the same module, providing high-quality features for brightness adaptation and interference removal for the subsequent fusion module, while also providing accurate feedback for the brightness feedback module. This module adopts an efficient and lightweight design, controlling computation time while ensuring the accuracy of brightness assessment and feedback. Its core function is to evaluate the brightness rationality of the enhanced image, filter abnormal brightness interference and noise, and generate a brightness feedback signal, ensuring the brightness adaptability and purity of the dual-modal features. The module uses a serial link structure, with each submodule executing sequentially in the code order of "feature alignment → brightness assessment → interference filtering → feedback generation." The output of the previous submodule serves as the input of the next submodule. Details are as follows: Feature Alignment Subunit: This unit adjusts the dimensions of the infrared and visible light features output by the network to unify brightness, avoiding deviations in subsequent brightness evaluation and feedback control caused by dimensionality differences. This unit receives the infrared and visible light features output from the LAN as input. Through a 1×1 convolutional dimensionality reduction and expansion algorithm, it maps both types of features to 64 channels, outputting aligned features with dimensions [B, 64, 640, 512]. The convolutional kernel weights are initialized using a He normal distribution parameter initialization algorithm, with the bias term set to 0.
[0078] Brightness Assessment and Interference Filtering Subunit: To assess the brightness rationality of aligned features, this subunit filters out abnormal brightness interference and Gaussian noise, preventing interfering features from entering subsequent fusion stages. This subunit receives aligned infrared or visible light features as input and outputs features with brightness calibration and interference filtering. Integrated Brightness Threshold Assessment and Dynamic Gaussian Filtering Logic: First, based on the illumination threshold of the scene adaptation module, excessively bright and dark areas are marked and appropriately corrected; then, the filtering intensity σ is dynamically adjusted according to the feature signal-to-noise ratio (S / N), with a higher filtering intensity for lower S / N. Core Formula: ; in, To enhance the post-features, GaussianBlur is used for Gaussian filtering, with ksize=3×3 as the fixed kernel size and σ as the dynamically adjusted filter coefficients. =1.0 is the basic filter coefficient, and S / N is the signal-to-noise ratio of the feature. The code calculates the mean and variance of the feature using a feature statistics calculation unit, obtains S / N through a signal-to-noise ratio quantization algorithm, substitutes it into the formula to calculate σ, and finally executes an adaptive Gaussian filtering algorithm. This submodule processes infrared and visible light features separately using a multimodal feature parallel processing algorithm, calling the filtering unit to ensure that noise in both types of features is effectively filtered.
[0079] Feedback Generation and Feature Output Subunit: This subunit generates a feedback signal based on the brightness assessment results and outputs purified features for subsequent fusion. It receives filtered infrared and visible light features and the brightness assessment results as input, generates a brightness feedback signal through threshold comparison and signal encoding, and transmits it to the LAN and RFN modules. Simultaneously, it performs a slight enhancement on the filtered features. The core formula is: ; The formula for visible light image feature enhancement is: ; in, These are the original infrared features. The original characteristics of visible light, , The weights of the convolution kernels for infrared and visible light feature enhancement are respectively. Conv3×3 represents a 3×3 convolution operation, and ReLU is the activation function. The infrared feature enhancement chain and the visible light feature enhancement chain run in parallel in the forward propagation unit to ensure processing efficiency. The enhanced feature dimensions are all [B, 64, 640, 512], and the number of channels is increased from 1 / 3 to 64, achieving high-dimensional representation of features.
[0080] The intermediate fusion module's core function is to receive the output luminance-adapted and interference-reduced dual-modal features, perform preliminary weighted fusion, and provide the foundational fusion features for further enhancement and fusion optimization in the fusion and enhancement network modules. This avoids the excessive computational burden caused by directly inputting the original purified features. The fusion formula is: ; in, , These are the infrared and visible light filtered features, respectively, with a fixed initial weight of 0.5. The code uses a dual-modal feature weighted fusion algorithm to calculate the formula, which is performed on the GPU to ensure processing efficiency. The dimensionality of the initially fused features remains unchanged, and they are subsequently transferred to the adaptive weight acquisition module via a feature data transmission scheduling algorithm.
[0081] The adaptive weight acquisition module is a crucial step in achieving accurate feature matching in occluded scenes. Its core function is to complete the entire processing chain: "occluded region recognition - dual-modal feature attention enhancement - adaptive weight acquisition - accurate fusion." Targeting the frequent occlusion issues at ports of entry, it implements differentiated fusion, prioritizing infrared penetration features in occluded areas and preserving visible light texture features in unoccluded areas. All functions are implemented through the AWAM class, running after the fusion and enhancement network module. The inputs are the optimized preliminary fusion features and the dual-modal features refined by the brightness feedback network; the output is the final fusion feature. This module contains four sub-modules, which are called in a linked manner within the forward propagation unit according to the code execution order of "occlusion detection → attention enhancement → weight acquisition → feature fusion," as detailed below: (1) Occlusion Region Detection Submodule: The core function is to accurately identify occlusion regions in visible light images and output an occlusion mask M for subsequent weight allocation. This function is implemented through an occlusion detection unit, which receives the enhanced visible light image as input and outputs an occlusion mask M. Specifically, the implementation integrates edge gradient detection and mask generation algorithms: the Sobel operator is used to calculate the image edge gradient, and the occlusion region is identified by utilizing the abrupt change in edge gradient in the occlusion region. First, the 3-channel visible light image is converted to grayscale using an RGB-to-grayscale algorithm, and then the gradients in the x and y directions are calculated using the Sobel edge gradient detection algorithm. The gradient formula is: ; ; in, , These are the edge gradients in the x and y directions, respectively. This is a visible light grayscale image. For convolution operations, G represents the total gradient. The total gradient is calculated using a total gradient magnitude calculation algorithm, with a gradient threshold of 0.3. An initial occlusion mask is generated using a binarization mask generation algorithm. To avoid false detections caused by isolated noise points, a morphological optimization processing unit is added to the code to optimize the occlusion mask, removing small noise areas. Finally, it is converted to tensor format output using a tensor-to-image matrix format conversion algorithm.
[0082] (2) Dual Attention Branch Submodule: The core function is to perform targeted attention enhancement on the filtered infrared and visible light features, increasing the weight of the target features and suppressing background interference features. This function is implemented through an infrared feature attention branch and a visible light feature attention branch, which run in parallel. Both branches use a convolutional attention module structure and are constructed through code. The inputs are the filtered infrared and visible light features, respectively, and the outputs are the attention-weighted infrared and visible light features, respectively. The formula for calculating the attention weights is as follows: ; ; In this context, GlobalAvgPool represents the global average pooling operation, Conv1×1 represents the 1×1 convolution operation, and Sigmoid is the activation function. , These are the attention weights for infrared and visible light features, respectively. The code defines the attention module using a Convolutional Attention Module (CAM) algorithm. In the forward propagation unit, it executes a process of "global average pooling → 1×1 convolution → Sigmoid" to obtain the attention weights. Then, an element-wise weighted enhancement algorithm is used to perform element-wise multiplication, strengthening high-weight target features. The convolution kernel parameters of the two attention branches are independently initialized using a branch-independent parameter initialization algorithm to adapt to the characteristics of the two types of features.
[0083] (3) Dynamic Weight Allocation and Feature Fusion Submodule: The core function is to dynamically adjust the fusion weights of infrared and visible light features based on the occlusion mask M to achieve differentiated fusion. This function is implemented through a dynamic weight allocation unit, which receives the occlusion mask M and attention-weighted infrared / visible light features as input and outputs the final fused features. The core logic consists of two steps, both implemented through code: the first step is to dynamically calculate the fusion weights, and the second step is to complete feature fusion based on the weights. The dynamic weight calculation formula is: ; Where M is the occlusion mask. , The preset weights are obtained from the scene adaptation module. α is the dynamically fused weight. Dynamic weight calculation is implemented in the code through a weight calculation unit, utilizing the broadcast mechanism of Tensors to ensure pixel-by-pixel weight matching. The feature fusion formula is: ; in, The features are ultimately fused and transmitted to the training optimization module through a feature data transmission scheduling algorithm, and are also used for subsequent output post-processing.
[0084] The training and optimization module is the core of model performance optimization. Its core function is to iteratively optimize all learnable parameters of the network by constructing a multi-loss joint optimization strategy of "no-reference color loss - feature consistency loss - structural similarity loss," ensuring high-quality fusion and color restoration accuracy in port scenarios. It contains three core functional units: a loss function construction unit, a parameter update unit, and a performance evaluation unit, running after the core processing module, with the final fused features as input. The original bimodal features and images are used as the output, which is the optimized network parameters. Specifically, it includes two sub-modules: (1) Loss Function Construction Submodule: The core function is to design a multi-loss joint optimization strategy, comprehensively constraining and fusing the color reproduction, feature consistency, and structural integrity of the image to avoid performance deviations caused by a single loss. The total loss function is: ; in , The specific calculation logic and code implementation for the three sub-losses, based on the loss weights, are as follows: ①No reference color loss: This is used to constrain the color distribution consistency between the fused image and the original visible light image, solving the color distortion problem of the fused image in low-light scenes. It is achieved by calculating the L1 distance between the color histograms of the two images, as shown in the following formula: ; In the formula, H=256 is the number of histogram bins. , The histogram counts are for the fused image and the original visible light image, respectively.
[0085] ② Feature consistency loss To ensure consistency between the final fused features and the mean values of the original bimodal features, and to prevent the loss of feature information during the fusion process, the mean squared error is used for calculation, as shown in the following formula: ; In the formula, N is the total number of characteristic elements. This is the i-th element of the final fused feature.
[0086] ③ Structural similarity loss Based on the Structural Similarity in Memory (SSIM) metric, the structural integrity of the fused image and the original visible light image is constrained to alleviate image blurring in low-light scenes. The formula is as follows: ; In the formula, To achieve reverse constraints on structural integrity by fusing the structural similarity between the image and the original visible light image.
[0087] (2) Parameter Update Submodule: Based on the constructed joint loss function, the network parameters are iteratively updated through an adaptive optimizer. Performance monitoring and early stopping strategies are introduced to balance model fitting performance and generalization ability, avoiding overfitting. The core process is as follows: ① Optimizer initialization: The Adam optimizer is adopted, which combines the advantages of momentum gradient descent and adaptive learning rate to improve convergence speed and stability.
[0088] ② Batch Iterative Training: The dataset is divided into multiple batches for sequential training using a batch iterative training algorithm. In each batch, forward propagation is used to obtain the loss terms and the total loss, followed by backpropagation to calculate the gradient. The Adam optimizer is then called to update all learnable parameters of the network, achieving a gradual decrease in loss. ③ Performance Evaluation and Early Stopping Strategy: Every 10 epochs, a performance evaluation unit is called to calculate the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) on the validation set to evaluate the model's generalization ability. An early stopping regularization strategy is employed. If there is no improvement in validation set performance for 15 consecutive epochs, training is terminated using a training process termination scheduling algorithm, preserving the current optimal parameters and preventing model overfitting.
[0089] The output post-processing module is the final stage of the technical process. Its core function is to convert the fused features into an image format suitable for practical applications. It receives the fused features after training and optimization, and the images stored during the standardization stage. and As input, the output is an image that can be directly used for port abnormal behavior detection, and it contains two sub-modules: (1) Denormalization Submodule: The core function is to convert the fusion features normalized to the [0,1] interval into 8-bit integer pixel values that meet the requirements of visual presentation and detection, thus restoring the original brightness range of the image. The core formula is the inverse operation of normalization: ; In the formula, For the final pixel value, For the normalized pixel values of the fused features, , These are the pixel extreme values of a single image recorded during the preprocessing stage. The implementation logic is as follows: First, the fused feature tensor is converted into an image matrix using a tensor-image matrix format conversion algorithm, and then the recovered value is calculated pixel by pixel using the formula. Next, a pixel value range cropping and bit-width conversion algorithm is used to constrain the pixel values to the range of 0~255, while simultaneously converting them to 8-bit integers to avoid image distortion caused by pixel value overflow.
[0090] (2) Edge Smoothing Submodule: The core function is to eliminate edge artifacts that may occur during the fusion process, improve the visual coherence of the image, and provide clearer edge features for subsequent anomaly detection. It is implemented using a fixed-parameter Gaussian filter. The filter kernel size is 5×5, and the standard deviation σ=1.5. Edge smoothing is achieved through a fixed-kernel Gaussian filter algorithm. The final output is a 3-channel RGB fused image with a resolution of 640×512, which is then directly transmitted to the port abnormal behavior detection system through the fused image detection system input interface adaptation algorithm.
[0091] To comprehensively verify the low-light image enhancement and visible-infrared image fusion method and system proposed in this invention for detecting abnormal behavior of personnel at ports of entry, a comparative experiment was specifically designed to meet the application requirements of actual security monitoring scenarios at ports. The experiment focused on three core objectives: fusion image quality assessment, model scenario adaptability verification, and real-time performance verification. It clarified the dataset composition, experimental hardware and software environment, implementation process, and evaluation indicators. Through a combination of quantitative analysis and qualitative evaluation, comparative tests were conducted with existing mainstream technical solutions to ensure the experimental results are authentic, reliable, and reproducible.
[0092] The experimental dataset was constructed using a "public dataset supplemented by a private port dataset" approach. This not only provides a foundation for comparing the general performance of different technical solutions but also ensures that the data closely reflects the actual customs clearance scenario at the port. The specific composition is as follows: (1) Data source ① Public datasets: The LLVIP dataset and the KAIST multispectral pedestrian dataset, two mainstream bimodal image fusion public datasets, were selected. Samples similar to the port scene were selected for benchmark comparison of the model's general performance. ② Port private datasets: The data were collected by infrared-visible light synchronous acquisition equipment deployed on-site at the port. After being authorized in compliance with regulations, the data were used for the model's special adaptation training and testing, covering the core scenarios of port clearance.
[0093] (2) Sample size and division The total experimental dataset contains 15,000 bimodal image pairs, each including simultaneously acquired infrared and visible light images, divided into training, validation, and test sets in a 7:2:1 ratio. The training and validation sets are comprised of 70% private port datasets and 30% public datasets, used for model parameter training and hyperparameter tuning. The test set includes 500 public dataset samples and 1,000 private port dataset samples, ensuring a fair comparison of the general performance and port-specific performance of different solutions. All sample images were uniformly adjusted to 640×512 resolution; infrared images were 16-bit single-channel grayscale images, and visible light images were 8-bit 3-channel RGB images.
[0094] (3) Data preprocessing The preprocessing process of the experimental dataset follows the implementation logic of the data preprocessing module of this invention. It mainly includes synchronization verification, standardization processing, and training set data augmentation to ensure that the data format of the input model is uniform and the quality meets the standards, forming a complete closed loop with the technical solution mentioned above.
[0095] The experiment was conducted strictly in accordance with standardized procedures, clearly defining the hardware and software environment configuration, training parameter settings, testing process design, and evaluation index system. This ensured that the experimental process was reproducible and the results had comparative value. Specific implementation details are as follows: (1) Hardware and software environment configuration ① Hardware Environment: The experimental hardware platform is tailored to the actual deployment needs of the port, with the following configuration: Intel Core i9-12900H CPU, NVIDIA RTX 3090 GPU, 32GB DDR5 4800MHz memory, and 1TB NVMe solid-state drive storage to ensure high efficiency in data reading and model computation; ② Software Environment: The operating system is Ubuntu 20.04 LTS Server, the programming language is Python 3.8, the deep learning framework is PyTorch 1.12.1, the image processing libraries are OpenCV 4.6.0 and PIL 9.2.0, and other core dependent libraries include NumPy 1.24.3, YAML 6.0, Matplotlib 3.7.1, and Scikit-learn 1.2.2.
[0096] (2) Model training parameter configuration The model training process is implemented based on the training optimization module of this invention. The core parameters are configured as follows: ① Optimizer: Adam adaptive optimizer is used, with β=0.9 and β=0.999, and the weight decay coefficient is 1e-5; ② Basic training parameters: batch size is set to 8, initial learning rate is 1e-4, and cosine annealing decay strategy is adopted; ③ Training epochs and early stopping strategy: the maximum number of training epochs is 200, and an early stopping mechanism is introduced. Training is terminated when the validation set evaluation index does not improve for 15 consecutive epochs, and the current optimal model parameters are saved; ④ Loss function: the joint loss function defined above is used, with loss weights λ=0.5 and λ=0.3; ⑤ Other configurations: CUDA is used for parallel acceleration during training, and performance evaluation is carried out on the validation set every 10 epochs to record the changes in core indicators.
[0097] (3) Test process design The experimental testing process is divided into three stages, comprehensively covering the key aspects of model performance evaluation: ① Model loading: Load the optimal model parameters saved through the early stopping strategy after training and set the model to inference mode; ② Data input: Test set samples are input into the model in batches, and sequentially undergo data preprocessing, scene adaptation, core module (RFN, LFN, AWAM) processing, and output post-processing to finally output the fused image; ③ Performance evaluation: Quantitative index calculation and qualitative effect analysis are performed on the output fused image, and the processing time of a single frame image throughout the entire process is statistically analyzed.
[0098] (4) Selection of comparison schemes To fully verify the comprehensive superiority of this invention, four existing mainstream bimodal image fusion and enhancement techniques were selected as comparison schemes, covering two major categories: traditional methods and deep learning methods. Specifically, the traditional methods selected were the wavelet transform-based fusion method WT and the multi-scale decomposition-based fusion method MSD; the deep learning methods selected were the convolutional neural network-based fusion method CNN-Fuse and the attention-based fusion method Attention-Fuse. All comparison schemes were configured with optimal parameters according to their original technical documents and tested on the same hardware and software environment and test set to ensure the fairness of the comparison results.
[0099] (5) Evaluation index system The experiment employs a dual evaluation system combining quantitative indicators and qualitative assessments to comprehensively measure the quality of fused images, the model's scene adaptability, and real-time performance. The quantitative indicators include 14 core metrics covering multiple dimensions such as information preservation, structural integrity, color reproduction, and sharpness. These specifically include information entropy (IE), standard deviation (SD), average gradient (AG), spatial frequency (SF), peak signal-to-noise ratio (PSNR), structural similarity (SSIM), color similarity (CS), mutual information (MI), normalized mutual information (NMI), correlation coefficient (CC), edge preservation (EP), visual information fidelity (VIF), noise suppression ratio (NSR), and runtime (RT).
[0100] To verify the superior performance and scene adaptability of this invention in the infrared and visible light dual-modal image fusion task, a comparative experiment was conducted using the LLVIP public dataset and the port clearance dataset. The test objects included this invention and mainstream fusion methods on the market such as wavelet transform (WT) and multi-scale decomposition (MSD); deep learning methods: convolutional neural network fusion CNN-Fuse and attention mechanism fusion Attention-Fuse.
[0101] The evaluation metrics selected are 14 core indicators, including entropy (EN), mutual information (MI), spatial frequency (SF), average gradient (AG), standard deviation (SD), visual information fidelity (VIF), correlation coefficient (CC), structural difference (SCD), mean square error (MSE), peak signal-to-noise ratio (PSNR), edge quality assessment (Qabf), normalized edge quality (Nabf), structural similarity (SSIM), and multi-scale structural similarity (MS_SSIM), covering key evaluation dimensions such as image information preservation, detail sharpness, structural consistency, and visual fidelity.
[0102] (1) Test results of the LLVIP public dataset The mean EN value of this invention is 7.503 (Std=0.241), and the mean MI value is 2.41 (Std=0.397), both higher than the WT method, MSD method, CNN-Fuse, and Attention-Fuse. Specifically, it improves EN by 2.08% and MI by 3.88% compared to Attention-Fuse, indicating that the fused image of this invention can more fully retain the original information of the two modalities, resulting in higher information richness. The mean AG value of this invention is 6.624 (Std=2.976), and the mean SF value is 20.11 (Std=7.871), significantly better than traditional methods. It improves by 13.62% and 11.41% respectively compared to CNN-Fuse (AG=5.83, SF=18.05), and by 6.67% and 5.73% respectively compared to Attention-Fuse, indicating that the edge details of the fused image of this invention are clearer and the spatial texture features are better presented. The average PSNR of this invention is 58.27 (Std=0.785), the average SSIM is 0.236 (Std=0.048), and the average MS_SSIM is 0.358 (Std=0.042). Specifically, the PSNR is 37.56% higher than WT, 31.21% higher than MSD, 12.62% higher than CNN-Fuse, and 8.67% higher than Attention-Fuse, with a Std of only 0.785, significantly lower than the comparison methods. This indicates that the fused image of this invention has higher visual fidelity, better structural consistency with the original image, and superior performance stability. The average VIF of this invention is 1.188 (Std=0.396), the average CC is 0.616 (Std=0.065), and the average Qabf is 0.368 (Std=0.099), all superior to various comparison methods, further confirming the superiority of this invention in overall image fusion performance.
[0103] (2) Test results of private dataset for port clearance scenarios The present invention achieves a mean EN of 7.686 (Std=0.041) and a mean MI of 4.783 (Std=0.163), both significantly improved compared to the LLVIP dataset, with extremely low Std values, far lower than the comparison methods. Specifically, it improves EN by 4.54% and MI by 10.52% compared to Attention-Fuse, and EN by 6.11% and MI by 12.34% compared to CNN-Fuse. This demonstrates that the present invention can still efficiently preserve bimodal information in complex real-world port scenarios, and its detail rendering capabilities are better suited to practical application needs. The mean AG value of this invention is 8.844 (Std=0.679), and the mean SF value is 35.365 (Std=1.833), representing improvements of 33.51% and 75.86% respectively compared to the LLVIP dataset. These improvements are more significant than those of the comparison methods, exceeding Attention-Fuse by 40.83% and 25.63%, CNN-Fuse by 46.18% and 34.36%, and the traditional WT method by 83.49% and 83.90%. Simultaneously, the mean Qabf value is 0.676 (Std=0.014), a 29.25% improvement over Attention-Fuse, indicating that this invention can accurately enhance target edge details in complex port scenarios, meeting the core requirement of port monitoring for image clarity. The PSNR of this invention reaches an average of 66.185 (Std=0.198), which is 13.58% higher than the LLVIP dataset, 17.29% higher than Attention-Fuse, 21.93% higher than CNN-Fuse, and 38.20% higher than the traditional MSD method. Furthermore, with a Std of only 0.198, it is the smallest among all comparison objects, indicating that this invention achieves extremely high visual fidelity and strong performance stability in real-world port scenarios. In addition, the SSIM average is 0.666 (Std=0.013) and the MS_SSIM average is 0.893 (Std=0.001), representing improvements of 27.83% and 17.19% respectively compared to Attention-Fuse, further demonstrating that the fused image and the original image have better structural consistency and are better suited for port abnormal behavior recognition tasks. The present invention achieves low performance across all metrics (Std) on the port clearance dataset, such as a mean CC of 0.971 (Std=0.005), a mean MSE of 0.016 (Std=0.001), and a mean SD of 61.839 (Std=1.865), all of which are lower than those of the LLVIP dataset and all comparison methods. This indicates that the present invention exhibits low performance fluctuations and strong robustness in complex and ever-changing real-world scenarios, making it more suitable for engineering deployment requirements.
[0104] Embodiments of the present invention also provide a system based on low-light image enhancement and visible-infrared image fusion, comprising: The data input and preprocessing module is used to preprocess the infrared and visible light images obtained in the same scene to obtain preprocessed standardized dual-modal image data. The port scene adaptation module is used to load a port scene-specific configuration file based on the preprocessed standardized bimodal image data, and to perform parameter initialization and adaptation of the fusion and enhancement network, brightness feedback network and adaptive weight acquisition module to obtain the adapted network parameters. The fusion and enhancement network module is used to take the input preprocessed standardized bimodal image data, use the adapted network parameters, and perform bidirectional guided adjustment of enhancement and fusion. It dynamically optimizes the enhancement gain parameters through the feedback signal of the fusion module, and dynamically adjusts the fusion strategy according to the effect evaluation parameters of the enhancement module to generate bimodal features after brightness adjustment network. The brightness feedback network module is used to sequentially perform feature enhancement, adaptive filtering, feedback signal generation and preliminary fusion operations on the obtained dual-modal features after brightness adjustment network to obtain preliminary fused features. The adaptive weight acquisition module is used to take the input preliminary fusion features, perform detection of occlusion regions in the visible light image and generate an occlusion mask; dynamically allocate fusion weights for infrared and visible light features based on the occlusion mask, and perform feature enhancement and weighted fusion to output the final fusion features; The training optimization module is used to train and optimize the model based on the final fusion features using a joint loss function, and to update the network parameters through backpropagation to obtain the optimized fusion model. The output post-processing module is used to obtain the final fusion features through the optimized fusion model, and then perform denormalization and edge smoothing operations in sequence to generate the final fused image.
[0105] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. Preprocess infrared and visible light images of the same scene to obtain preprocessed standardized dual-modal image data; Based on the preprocessed standardized bimodal image data, a port scene-specific configuration file is loaded, and the parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module are initialized and adapted to obtain the adapted network parameters. The preprocessed standardized bimodal image data is input into the fusion and enhancement network. The network parameters are adapted and bidirectional guided adjustment of enhancement and fusion is performed. The enhancement gain parameters are dynamically optimized through the feedback signal of the fusion module, and the fusion strategy is dynamically adjusted according to the effect evaluation parameters of the enhancement module to generate bimodal features after passing through the brightness adjustment network. The dual-modal features after passing through the brightness adjustment network are input into the brightness feedback network, and feature enhancement, adaptive filtering, feedback signal generation and preliminary fusion operations are performed in sequence to obtain preliminary fused features. The initial fused features are input into the adaptive weight acquisition module to detect occlusion regions in the visible light image and generate an occlusion mask; Based on the occlusion mask, the fusion weights of infrared and visible light features are dynamically assigned, and feature enhancement and weighted fusion are performed to output the final fused features; Based on the final fusion features, the network is trained and optimized according to the joint loss function, and the network parameters are updated through backpropagation to obtain the optimized fusion model. The optimized fusion model is used to input the final fusion features into the post-processing module, where denormalization and edge smoothing operations are performed sequentially to generate the final fused image.
2. The method based on low-light image enhancement and visible-infrared image fusion according to claim 1, characterized in that, Infrared and visible light images of the same scene are preprocessed to obtain preprocessed standardized dual-modal image data, including: The acquired infrared and visible light images are time-stamp aligned and verified to remove image pairs with time deviations exceeding a preset threshold, thus obtaining time-synchronized dual-modal image pairs. Pixel normalization is performed on the time-synchronized dual-modal image pairs to uniformly map the pixel value range of the infrared image and the visible light image to a preset range, thereby obtaining the normalized dual-modal image. The normalized bimodal image is subjected to data augmentation processing, including at least one of random cropping, horizontal flipping, and brightness perturbation, to obtain preprocessed standardized bimodal image data, which includes preprocessed standardized infrared image and preprocessed standardized visible light image.
3. The method based on low-light image enhancement and visible-infrared image fusion according to claim 2, characterized in that, Based on the preprocessed standardized bimodal image data, a port scene-specific configuration file is loaded, and the parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module are initialized and adapted to obtain the adapted network parameters, including: Load the port scenario-specific configuration file, parse and extract the scenario prior configuration information from it; Based on the extracted scenario prior configuration information, the enhancement gain parameters and evaluation thresholds of the fusion and enhancement networks are initialized and assigned values. Based on the extracted scene prior configuration information, the filter coefficients and convolution kernel weights of the brightness feedback network are initialized and assigned values. Based on the prior configuration information of the scenario, the preset fusion weights of the occluded and unoccluded regions in the adaptive weight acquisition module are initialized and assigned values. By integrating all parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module after initialization, the adapted network parameters are obtained.
4. The method based on low-light image enhancement and visible-infrared image fusion according to claim 3, characterized in that, The preprocessed standardized bimodal image data is input into the fusion and enhancement network. Adapted network parameters are used to perform bidirectional guided adjustment of enhancement and fusion. The enhancement gain parameters are dynamically optimized based on the feedback signal from the fusion module, and the fusion strategy is dynamically adjusted according to the effect evaluation parameters of the enhancement module. This generates bimodal features after passing through the brightness adjustment network, including: The preprocessed standardized bimodal image data is input into the fusion and enhancement network; Using the adapted network parameters, micro-light enhancement processing is performed on the visible light image to generate the enhanced visible light image and the corresponding enhancement effect evaluation parameters. The enhanced visible light image and the preprocessed standardized infrared image are input into the fusion module in the enhancement network. The fusion module performs feature matching and preliminary fusion to generate feature matching error, which is used as a feedback signal. Based on feature matching error and enhancement effect evaluation parameters, the enhancement gain parameters are dynamically updated by fusing and enhancing the parameter adjustment mechanism in the network, thus generating the updated enhancement gain parameters. By employing updated enhancement gain parameters, a new round of enhancement and fusion-guided adjustment is performed on the preprocessed standardized bimodal image data to generate bimodal features after passing through a brightness adjustment network.
5. The method based on low-light image enhancement and visible-infrared image fusion according to claim 4, characterized in that, The dual-modal features after brightness adjustment are input into the brightness feedback network, and feature enhancement, adaptive filtering, feedback signal generation, and preliminary fusion operations are performed sequentially to obtain preliminary fused features, including: The dual-modal features after passing through the brightness adjustment network are input into the feature enhancement submodule of the brightness feedback network to enhance the infrared and visible light features respectively, thus obtaining the enhanced dual-modal features. The enhanced bimodal features are input into the adaptive filtering submodule of the luminance feedback network. The filtering intensity is dynamically adjusted based on the signal-to-noise ratio of the features, and the filtering process is performed to obtain the filtered bimodal features. The filtered dual-modal features are input into the preliminary fusion submodule of the brightness feedback network to perform weighted fusion of infrared and visible light features, thereby obtaining preliminary fused features.
6. The method based on low-light image enhancement and visible-infrared image fusion according to claim 5, characterized in that, The initial fused features are input into the adaptive weight acquisition module to detect occlusion regions in the visible light image and generate an occlusion mask; Based on the dynamic allocation of fusion weights for infrared and visible light features using occlusion masks, feature enhancement and weighted fusion are performed to output the final fused features, including: The preliminary fused features and the corresponding preprocessed standardized visible light image are input into the occlusion detection unit of the adaptive weight acquisition module to perform occlusion region detection and generate an occlusion mask. The initial fused features are input into the dual attention branch of the adaptive weight acquisition module, and attention enhancement processing is performed on the infrared features and visible light features respectively to obtain the attention-weighted dual-modal features. Based on the occlusion mask, the dynamic fusion weight of infrared features and visible light features is calculated in the dynamic weight allocation unit of the adaptive weight acquisition module. Based on the dynamic fusion weights, the attention-weighted bimodal features are weighted and fused to output the final fused features.
7. The method based on low-light image enhancement and visible-infrared image fusion according to claim 6, characterized in that, Based on the final fusion features, training and optimization are performed according to the joint loss function. The network parameters are then updated via backpropagation to obtain the optimized fusion model, which includes: Based on the final fused features and the preprocessed standardized bimodal image data, a joint loss function value is calculated, which is composed of no-reference color loss, feature consistency loss and structural similarity loss. Based on the joint loss function value, the gradient is calculated through backpropagation, and the network parameters of the fusion and enhancement network, the brightness feedback network, and the adaptive weight acquisition module are updated to obtain the model with updated parameters. The updated model is used to process the validation set data to obtain the corresponding validation performance metrics. The training termination condition is determined based on the verification performance metrics. If the condition is met, the current model is used as the optimized fusion model. If the condition is not met, the joint loss function value is recalculated to continue the next round of iterative training.
8. The method based on low-light image enhancement and visible-infrared image fusion according to claim 7, characterized in that, The optimized fusion model is used to input the final fusion features into the post-processing module, where inverse normalization and edge smoothing operations are performed sequentially to generate the final fused image, including: The final fusion features are input into the optimized fusion model for processing to obtain an image matrix corresponding to the final fusion features; The image matrix is subjected to denormalization processing. Combined with the pixel extreme value parameters recorded in the preprocessing stage, the normalized pixel values are restored to the original numerical range to obtain the denormalized image. Edge smoothing filtering is performed on the denormalized image to eliminate image edge artifacts and generate the final fused image.
9. A system based on low-light image enhancement and visible-infrared image fusion, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The data input and preprocessing module is used to preprocess the infrared and visible light images obtained in the same scene to obtain preprocessed standardized dual-modal image data. The port scene adaptation module is used to load a port scene-specific configuration file based on the preprocessed standardized bimodal image data, and to perform parameter initialization and adaptation of the fusion and enhancement network, brightness feedback network and adaptive weight acquisition module to obtain the adapted network parameters. The fusion and enhancement network module is used to take the input preprocessed standardized bimodal image data, use the adapted network parameters, and perform bidirectional guided adjustment of enhancement and fusion. It dynamically optimizes the enhancement gain parameters through the feedback signal of the fusion module, and dynamically adjusts the fusion strategy according to the effect evaluation parameters of the enhancement module to generate bimodal features after brightness adjustment network. The brightness feedback network module is used to sequentially perform feature enhancement, adaptive filtering, feedback signal generation and preliminary fusion operations on the obtained dual-modal features after brightness adjustment network to obtain preliminary fused features. The adaptive weight acquisition module is used to take the initial fusion features of the input, perform detection of occlusion regions in the visible light image, and generate an occlusion mask. Based on the occlusion mask, the fusion weights of infrared and visible light features are dynamically assigned, and feature enhancement and weighted fusion are performed to output the final fused features; The training optimization module is used to train and optimize the model based on the final fusion features using a joint loss function, and to update the network parameters through backpropagation to obtain the optimized fusion model. The output post-processing module is used to obtain the final fusion features through the optimized fusion model, and then perform denormalization and edge smoothing operations in sequence to generate the final fused image.