Dual vision fusion image optimization processing method and system for night vision scope

By constructing a dual-light fusion virtual imaging model and edge processing nodes, the image adaptation problem of night vision sights in complex environments was solved, achieving accuracy in target recognition and stability in imaging, thus improving the target recognition and aiming effect of night vision sights.

CN122175819APending Publication Date: 2026-06-09SHENZHEN PARD TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN PARD TECH CO LTD
Filing Date
2026-02-11
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing dual-light fusion image processing technology in night vision sights cannot dynamically adapt to complex outdoor environmental changes, resulting in low image signal-to-noise ratio, blurred details, and difficulty in shielding interference areas, thus failing to meet the needs of efficient target recognition and accurate aiming.

Method used

By constructing a dual-light fusion virtual imaging model based on dual-light sensing data and ambient lighting parameters, and combining multi-dimensional simulation to determine the target recognition confidence threshold and image fusion adjustment strategy, edge processing nodes are deployed to collect noise signals, detail-sensitive features under the attention mechanism are set, and multi-level confidence separation masks are constructed for adaptive segmentation and interference area shielding.

Benefits of technology

It achieves accurate target recognition and stable imaging in complex environments, meets the needs of efficient aiming and target resolution, and improves the target recognition stability and aiming decision reliability of night vision sights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122175819A_ABST
    Figure CN122175819A_ABST
Patent Text Reader

Abstract

The application discloses a dual-light fusion image optimization processing method and system for a night-vision sighting telescope, relates to the technical field of intelligent image optimization processing, and comprises the following steps: collecting thermal imaging signal-to-noise ratio, visible light contrast and other imaging state data based on dual-light sensing and environmental illumination parameters; combining the data with an image degradation mode knowledge base, constructing a dual-light fusion virtual imaging model and simulating the model to determine a confidence threshold and a fusion adjustment strategy; collecting noise signals, setting sensitive features in combination with illumination attenuation trends, completing alarm matching and enhancement compensation, and simultaneously constructing a multi-level confidence separation mask to perform segmentation and shielding. The application solves the technical problems that the existing dual-light fusion image processing cannot dynamically adapt to environmental changes, and the optimization strategy lacks effective cooperation with image enhancement and interference exclusion, and achieves the technical effects of adapting to environmental changes, cooperatively optimizing image effects and accurately shielding interference, and improving the target recognition stability and sighting decision reliability of the night-vision sighting telescope.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent image optimization processing technology, and in particular to a dual-light fusion image optimization processing method and system for night vision sights. Background Technology

[0002] The image quality of night vision sights directly impacts target identification and accurate aiming in scenarios such as outdoor hunting and emergency search and rescue, making the optimized processing of dual-light fusion images crucial. Existing technologies mostly employ traditional dual-light overlay or single image enhancement methods, which have proven effective in stable environments. However, as applications extend to complex outdoor scenarios, traditional image data processing techniques reveal significant limitations. Due to the large fluctuations in outdoor lighting, noise interference, and target occlusion, traditional processing methods cannot accurately adapt to environmental changes, resulting in low signal-to-noise ratios, blurred details, difficulty in masking interference areas, and inaccurate image data, failing to meet the demands of efficient target identification and accurate aiming in night vision sights. Summary of the Invention

[0003] This application provides a dual-light fusion image optimization processing method and system for night vision sights, which solves the technical problems that existing dual-light fusion image processing cannot dynamically adapt to environmental changes, and that the optimization strategy lacks effective coordination with image enhancement and interference elimination.

[0004] The first aspect of this application provides a dual-light fusion image optimization processing method for night vision sights. The method includes: collecting imaging state data, including thermal imaging signal-to-noise ratio and visible light contrast, under the same time stamp based on dual-light sensing data and ambient light parameters; constructing a dual-light fusion virtual imaging model according to the dual-light sensing data and an image degradation mode knowledge base, and performing multi-dimensional simulation and deduction in combination with the imaging state data to determine the target recognition confidence threshold and image fusion adjustment strategy; deploying edge processing nodes to collect image noise intensity signals, and setting detail-sensitive features under an attention mechanism based on the light intensity attenuation trend within the time sliding window of the ambient light parameters, and performing adaptive alarm level matching and image enhancement command compensation; simultaneously, constructing a multi-level confidence separation mask under dual-light fusion for the night vision sight based on the target recognition confidence threshold and image fusion adjustment strategy, and using the multi-level confidence separation mask to perform adaptive segmentation and interference region shielding on the fused image.

[0005] A second aspect of this application provides a dual-light fusion image optimization processing system for night vision sights. The system includes: an imaging state data acquisition module, which collects imaging state data, including thermal imaging signal-to-noise ratio and visible light contrast, under the same time stamp, based on dual-light sensing data and ambient light parameters; an image fusion adjustment strategy acquisition module, used to construct a dual-light fusion virtual imaging model based on the dual-light sensing data and an image degradation mode knowledge base, and perform multi-dimensional simulation deduction in conjunction with the imaging state data to determine the target recognition confidence threshold and image fusion adjustment strategy; an image enhancement command compensation execution module, used to deploy edge processing nodes, collect image noise intensity signals, combine the light intensity attenuation trend within the time sliding window of the ambient light parameters, set detail-sensitive features under an attention mechanism, and perform adaptive alarm level matching and image enhancement command compensation; and a fusion image processing execution module, which simultaneously constructs a multi-level confidence separation mask under dual-light fusion for the night vision sight based on the target recognition confidence threshold and image fusion adjustment strategy, and uses the multi-level confidence separation mask to perform adaptive segmentation and interference region shielding on the fused image.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application collects dual-light sensing and ambient lighting-related data, constructs a dual-light fusion virtual model for multi-dimensional simulation, deploys edge processing nodes to collect noise signals and extract sensitive details, and combines confidence thresholds to construct multi-level separation masks to segment and shield interference, thereby accurately optimizing the dual-light fusion image effect. This makes the target recognition of the night vision sight more accurate and the imaging more stable in complex environments, meeting the needs of efficient aiming and target resolution. It achieves the technical effect of improving the target recognition stability and aiming decision reliability of the night vision sight by adapting to environmental changes, collaboratively optimizing image effects and accurately shielding interference. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a flowchart illustrating the dual-light fusion image optimization processing method for night vision sights provided in this application embodiment.

[0009] Figure 2 This is a schematic diagram of the structure of the dual-light fusion image optimization processing system for night vision sights provided in the embodiments of this application.

[0010] Figure labeling: 1. Imaging status data acquisition module; 2. Image fusion adjustment strategy acquisition module; 3. Image enhancement instruction compensation execution module; 4. Fusion image processing execution module. Detailed Implementation

[0011] This application provides a dual-light fusion image optimization processing method and system for night vision sights, which solves the technical problems that existing dual-light fusion image processing cannot dynamically adapt to environmental changes, and that the optimization strategy lacks effective coordination with image enhancement and interference elimination.

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0013] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0014] Example 1, as Figure 1 As shown, a dual-light fusion image optimization processing method for night vision sights is provided, wherein the method includes: Based on dual-light sensing data and ambient lighting parameters, imaging status data, including thermal imaging signal-to-noise ratio and visible light contrast, are collected under the same time stamp.

[0015] Specifically, firstly, the thermal imaging detector and night vision detector of the night vision sight collect raw thermal imaging data and raw night vision imaging data in real time, forming dual-light sensing data. Simultaneously, the light sensor on the night vision sight collects ambient light parameters such as light intensity and light incidence angle in real time. Then, a unified time base is used to timestamp the collected dual-light sensing data and ambient light parameters, ensuring that each set of raw thermal imaging data and raw night vision imaging data accurately corresponds to the ambient light parameters at the same moment, achieving strict synchronization of data timing and avoiding subsequent processing deviations caused by timing misalignment.

[0016] On the one hand, for the synchronized raw thermal imaging data, the effective signal area and noise area in the image are divided. The average gray value of the effective signal area and the standard deviation of the gray value of the noise area are calculated by statistical analysis. Based on the signal-to-noise ratio calculation formula, that is, the ratio of the average gray value of the signal to the standard deviation of the noise gray value, the thermal imaging signal-to-noise ratio is obtained, and the purity of the thermal imaging signal is quantified.

[0017] On the other hand, for the original night vision imaging data after synchronization, the target region of interest and the surrounding background region in the image are selected, and the average gray value of the two regions is calculated respectively. According to the contrast calculation formula, that is, the difference between the average gray value of the target region of interest and the background region and the ratio of the average gray value of the two regions, the visible light contrast is obtained, which represents the degree of distinction between the target and the background.

[0018] Finally, the original dual-light sensing data calibrated with the same timestamp, ambient lighting parameters, and the calculated thermal imaging signal-to-noise ratio and visible light contrast are correlated and integrated to form structured imaging state data, providing complete and synchronous basic data support for the subsequent construction and simulation of dual-light fusion virtual imaging models.

[0019] Based on the dual-light sensing data and the image degradation pattern knowledge base, a dual-light fusion virtual imaging model is constructed, and multi-dimensional simulation is performed in conjunction with the imaging state data to determine the target recognition confidence threshold and image fusion adjustment strategy.

[0020] Optionally, firstly, having acquired dual-light sensing data in the aforementioned steps and calculated the thermal imaging signal-to-noise ratio and visible light contrast ratio to obtain imaging state data, the acquisition of dual-light sensing data and extraction of core parameters are completed. Next, an image degradation mode knowledge base is acquired, specifically as follows: historical imaging data of night vision sights under different usage scenarios is collected, including imaging degradation samples under different environmental conditions such as temperature, humidity, and fog concentration. Simultaneously, imaging performance degradation data corresponding to different combinations of environmental parameters is obtained through controlled environment experiments. The above data is then organized and analyzed to extract the correspondence between environmental parameters such as temperature, humidity, and fog concentration and imaging degradation parameters such as signal-to-noise ratio attenuation rate, contrast reduction rate, and hot spot drift. These correspondences are stored in a structured format to construct an image degradation mode knowledge base, providing data support for subsequent model construction.

[0021] Next, based on the acquired dual-light sensing data and image degradation pattern knowledge base, a dual-light fusion virtual imaging model is constructed. This model integrates the fusion logic of dual-light images with the imaging degradation rules in the knowledge base, enabling it to simulate the dual-light fusion imaging process under different conditions. Combining the collected imaging state data under the same timestamp, simulations are conducted from multiple dimensions, including target distance, environmental conditions, and usage scenarios, to simulate the changes in imaging effects under different combinations of conditions. By analyzing the simulation results, the confidence threshold for accurately identifying the target is determined, i.e., the target recognition confidence threshold. Simultaneously, dual-light fusion parameter adjustment schemes adapted to different imaging scenarios, i.e., image fusion adjustment strategies, are formulated, providing a core basis for subsequent image optimization processing. This step will be explained in detail later.

[0022] Edge processing nodes are deployed to collect image noise intensity signals. Combined with the light intensity attenuation trend within the time sliding window of the ambient light parameters, detail-sensitive features under the attention mechanism are set, and adaptive alarm level matching and image enhancement command compensation are performed.

[0023] In this embodiment, the edge processing node is a lightweight computing module that integrates wavelet transform and other functional units and is deployed locally on the night vision sight. It can collect image noise signals nearby and perform operations such as feature processing, trend prediction and command triggering.

[0024] In one embodiment of this application, an edge processing node is deployed and an image noise intensity signal is acquired. The image noise intensity signal is first decomposed into multiple frequency band feature components by the wavelet transform unit integrated in the node. Energy entropy and kurtosis features are extracted from these feature components to generate a time-frequency domain sensitive feature vector. A subset of target edge sensitive features that are strongly correlated with target contour degradation is selected. This feature subset is then fused with the illumination intensity and incident angle in the ambient illumination parameters and used as the input of the gated recurrent unit network to predict the image sharpness change trend within the time sliding window. Combined with the normal sharpness range, it is determined whether to trigger an image enhancement or refocusing command. This step will be described in detail later.

[0025] Meanwhile, based on the target recognition confidence threshold and image fusion adjustment strategy, a multi-level confidence separation mask under dual-light fusion of the night vision aiming scope is constructed, and the multi-level confidence separation mask is used to perform adaptive segmentation and interference region shielding on the fused image.

[0026] Specifically, firstly, based on the target recognition confidence threshold and image fusion adjustment strategy determined in the aforementioned steps, a multi-level confidence separation mask is constructed. Specifically: according to the distribution range of the target recognition confidence threshold, the confidence level is divided into four levels: high confidence level (confidence ≥ 0.8), medium confidence level (0.6 ≤ confidence < 0.8), low confidence level (0.3 ≤ confidence < 0.6), and extremely low confidence level (confidence < 0.3), thus clarifying the level division rules. Next, a pixel value mapping method is used to establish a one-to-one correspondence between confidence levels and mask pixel values. High confidence level corresponds to pixel value 255, medium confidence level corresponds to pixel value 192, low confidence level corresponds to pixel value 128, and very low confidence level corresponds to pixel value 0. At the same time, combined with the fusion weight parameters in the image fusion adjustment strategy, the pixels at the edge of the mask are smoothed to avoid jagged distortion at the level boundaries, thereby completing the construction of a multi-level confidence separation mask.

[0027] Next, based on the constructed multi-level confidence separation mask, adaptive segmentation of the fused image is performed. A pixel-level threshold determination method is adopted, setting the segmentation criteria: for each pixel in the fused image, its corresponding target recognition confidence value is extracted and compared with the confidence level threshold associated with the pixel value at the corresponding position in the mask. Then, precise matching between the mask and image pixels is performed. All pixels in the fused image are traversed, and the confidence value of each pixel is matched with the level threshold at the corresponding position in the mask. If the pixel confidence value belongs to a high or medium confidence level, it is determined to be a pixel in the target region; if it belongs to a low confidence level, it is determined to be a pixel in the transition region; if it belongs to an extremely low confidence level, it is determined to be a pixel in the potential interference region. Through this matching mechanism, adaptive segmentation of the fused image is completed, achieving the initial division of the target region, transition region, and potential interference region.

[0028] Then, the interference area shielding operation is performed, clarifying the supplementary rules for defining the interference area and the pixel processing method. Based on the initial division, the interference area is further defined using gray-level variance analysis and edge density detection: the gray-level variance within the potential interference area is calculated. If the gray-level variance is lower than a preset threshold (determined based on historical interference area data of the night vision scope), and the edge density is lower than the set threshold, then the area is confirmed as an interference area. For transition areas, if the proportion of adjacent pixels to the interference area exceeds 50%, the transition area is classified as an interference area. The gray-level values ​​of all pixels within the confirmed interference area are set to 0, achieving direct shielding of the interference signal. Simultaneously, the image edges after shielding are processed using a mean filtering algorithm to avoid obvious boundary discontinuities between the shielded and unshielded areas, ensuring the overall image continuity.

[0029] By employing methods such as statistical interval division, pixel mapping, threshold determination, pixel-by-pixel matching, and gray-scale variance analysis, the construction of multi-level confidence separation masks, adaptive segmentation of fused images, and interference region shielding were fully realized. This effectively improved the target region recognition and imaging purity of the fused images, providing high-quality image data support for subsequent target recognition.

[0030] Furthermore, the method provided in this application embodiment includes: The thermal imaging signal-to-noise ratio and visible light contrast are bound to the usage scenario type and target distance information to establish an imaging-scene correlation matrix. Based on the imaging-scene correlation matrix, the imaging differences of the same night vision sight in different environmental stages are identified, and benchmark imaging feature values ​​are extracted. The benchmark imaging feature values ​​are used to automatically identify imaging abnormal scenes and push them to the user display interface.

[0031] Specifically, firstly, the thermal imaging signal-to-noise ratio (SNR) and visible light contrast are acquired through the thermal imaging detector and night vision detector integrated into the night vision sight. Simultaneously, the scene classification module integrated into the night vision sight pre-labels usage scenarios such as hunting, outdoor search and rescue, and night patrol. Target distance information is collected in real time through the laser ranging module. Using a timestamp synchronization method, the thermal imaging SNR and visible light contrast acquired at the same time are correlated one-to-one with the corresponding scene type labels and target distance data. A two-dimensional table is used to construct an imaging-scene correlation matrix. The row dimension of the matrix represents the combination of different scene types and target distances, and the column dimension represents the thermal imaging SNR and visible light contrast parameters. The correlated data is then filled into the corresponding positions in the matrix, completing the construction of the imaging-scene correlation matrix.

[0032] Next, according to preset usage scenario types, such as hunting, outdoor search and rescue, and night patrol, and target distance ranges, such as close range 0-100 meters, medium range 100-300 meters, and long range over 300 meters, the data in the imaging-scene correlation matrix are cross-grouped to ensure that each group of data corresponds to a unique combination of scene type and target distance. For each group of data, the mean and variance of thermal imaging signal-to-noise ratio and the mean and variance of visible light contrast are calculated to obtain the core statistical characteristics of each group of data. Subsequently, a two-way comparative analysis is performed: horizontally, the statistics of different target distance groups under the same scene type are compared to observe the changing trends of thermal imaging signal-to-noise ratio and visible light contrast with increasing distance, such as whether they show a decay trend and whether the decay rate is consistent; vertically, the statistics of different scene type groups under the same target distance are compared to analyze the degree of influence of environments such as low light, fog, and polar night on the two parameters, such as whether fog causes a significant increase in the variance of signal-to-noise ratio and whether polar night causes a significant decrease in the mean of contrast. By analyzing the results of bidirectional comparison, we can summarize the fluctuation range and variation of imaging parameters under different environmental stages. Based on these patterns, we can further clarify the deviation of imaging parameters from normal levels under different combinations of environments and distances, thereby accurately determining imaging differences.

[0033] Subsequently, the K-means clustering algorithm was used to perform cluster analysis on all parameter data under normal imaging conditions in the imaging-scene association matrix. Specifically: First, all data labeled as normal imaging conditions in the imaging-scene association matrix were filtered out, i.e., imaging data that had been judged as abnormal in the historical records, and this data was used as the input dataset for the K-means clustering algorithm. Referring to the number of scene types and data distribution characteristics involved in normal imaging, a preset value of K for the number of clusters was determined, typically between 3 and 5, which could be adjusted according to the actual amount of data. The K-means++ algorithm was used to initialize K cluster centers to avoid clustering bias caused by random selection of initial centers. The Euclidean distance from each sample in the input dataset to each cluster center was calculated, where each sample contains two parameters: thermal imaging signal-to-noise ratio and visible light contrast. Each sample was assigned to the category of the nearest cluster center, completing the first round of clustering assignment. Based on the results of the first round of assignment, the mean value of all samples in each cluster was recalculated, and the cluster center position was updated with this mean value. Then, the process of calculating the distance from the sample to the new cluster center, reassigning the sample category, and updating the cluster center was repeated. The iteration termination condition is set as follows: when the change in the position of the cluster center is less than a preset threshold of 0.001 in two consecutive iterations, or when the number of iterations reaches a preset upper limit of 100, the iteration stops, and the final clustering result is obtained. The number of samples in each cluster is counted, and the cluster with the highest percentage of samples is the dataset with the highest cluster density. The mean values ​​of thermal imaging signal-to-noise ratio and visible light contrast in this dataset are calculated, and these two means are used together as the baseline imaging feature value to ensure that the baseline value can reflect the imaging level when the equipment is working normally.

[0034] Finally, upper and lower floating thresholds are set based on the baseline imaging feature values ​​and the normal fluctuation range in historical operating data. Specifically, thermal imaging signal-to-noise ratio and visible light contrast data under normal imaging conditions are extracted from historical operating data. The standard deviation of this dataset is calculated. Considering the sensitivity requirements of the night vision sight for imaging anomalies in different scenarios such as hunting and outdoor search and rescue, a floating coefficient of 1.5-2.0 is set. The baseline imaging feature values ​​are then added to the product of the floating coefficient and the standard deviation, and subtracted from the product of the floating coefficient and the standard deviation to obtain the corresponding upper and lower floating thresholds. This ensures that the thresholds can capture real imaging anomalies in a timely manner while avoiding false alarms caused by minor fluctuations. The thermal imaging signal-to-noise ratio and visible light contrast under the current operating conditions are collected in real time and compared with the baseline imaging feature values. If the real-time parameters exceed the set threshold range, it is determined to be an imaging anomaly scene. The abnormal scene information is displayed on the user's display screen in the form of text prompts through the night vision sight's display screen driver module. At the same time, the device's built-in slight vibration module is triggered to push an anomaly alarm to the user, ensuring that the user is aware of it in a timely manner.

[0035] By binding imaging status data with scene, distance, and depth, and combining statistical analysis, clustering algorithms, and threshold judgment methods, accurate identification of imaging differences and automatic alarm for abnormal scenes are achieved, effectively improving the adaptability and imaging reliability of dual-light fusion images of night vision sights to different usage scenarios.

[0036] Furthermore, the method provided in this application embodiment includes: Digital gain, fusion weight, and AI inference frequency are used as variable parameters to set constraints in the dual-light fusion virtual imaging model; based on the constraints, an effective sample set is determined, and non-dominated sorting is performed to obtain the non-dominated solution set.

[0037] Specifically, firstly, three variable parameters are obtained: digital gain, fusion weight, and AI inference frequency. Digital gain is obtained through parameter calibration. Typical operating scenarios of the thermal imaging detector and night vision detector mounted on the night vision sight are selected, covering different environments such as low light, fog, and polar night. In each scenario, the gain value is gradually adjusted and imaging data is collected. Image noise analysis tools are used to statistically analyze the noise intensity under different gains, and the gain range with noise intensity below a preset threshold and signal clarity meeting requirements is selected. The median value of this range is taken as the initial digital gain parameter. Fusion weight is obtained based on the weighted fusion experimental method in dual-light fusion technology. A standard test platform is built in a laboratory environment, and different fusion ratios of thermal imaging and night vision imaging are set. The range is divided into 0.1 intervals within the 0 to 1 range. Imaging tests are performed on standard targets, and image quality evaluation indicators such as structural similarity index are used to score the fused image. The ratio with the highest score is selected as the initial fusion weight parameter. The AI ​​inference frequency was obtained through hardware adaptation testing methods. Combined with the CPU or GPU computing power parameters of the edge processing node of the night vision scope, computing power testing tools were used to simulate the algorithm running time at different inference frequencies. The frequency range with a time not exceeding the preset delay threshold and stable recognition accuracy was selected, and the commonly used value in this range was taken as the initial AI inference frequency parameter.

[0038] Next, specific constraints were set in the dual-light fusion virtual imaging model. For digital gain, based on the hardware performance limits of the detector, the maximum gain value with no significant noise was determined through actual measurement as the upper limit, while the minimum gain value required for target contour recognition was set as the lower limit, forming a constraint range for digital gain. For fusion weight, based on the basic logic of dual-light fusion, its value range was constrained to 0 to 1 to ensure a reasonable fusion ratio between thermal imaging data and night vision imaging data, avoiding excessive dominance of single-channel data in the imaging results. For AI inference frequency, combined with battery life test data from edge processing nodes, the frequency corresponding to the minimum power consumption required for continuous operation for a preset duration was set as the lower limit, while the maximum stable inference frequency corresponding to the upper limit of computing power was set as the upper limit, forming a constraint range for AI inference frequency. Simultaneously, imaging quality constraints were set. Imaging data under different parameter combinations were tested using target recognition algorithms such as YOLOv5s, with a target recognition accuracy of no less than 90% as a hard constraint, ensuring that parameter adjustments always meet the core imaging requirements.

[0039] Then, based on the aforementioned variable parameters and constraints, a valid sample set is determined. Sample data is collected, including measured imaging data of the night vision sight under different target distances and environmental conditions, as well as imaging effect data and power consumption data corresponding to different parameter combinations generated by simulation through a dual-light fusion virtual imaging model. Each collected raw sample data is verified, and invalid samples with digital gain, fusion weight, AI inference frequency exceeding the constraints, or target recognition accuracy below 90% are removed. The remaining samples constitute the valid sample set. Subsequently, a fast non-dominated sorting algorithm is used to evaluate each sample in the valid sample set for multi-target evaluation, with evaluation dimensions including target recognition accuracy and image processing power consumption. By comparing the dominance relationships between samples—that is, if the target recognition accuracy of one sample is not lower than that of another sample and the image processing power consumption is not higher than that of the first sample—the former is determined to dominate the latter. All samples not dominated by other samples are selected to form a non-dominated solution set, providing data support for the subsequent selection of the optimal balance point.

[0040] Core variable parameters were obtained through experimental calibration and statistical methods. Specific constraints were set in combination with hardware performance and imaging quality requirements. After data screening and processing with a fast non-dominated sorting algorithm, a reliable non-dominated solution set was efficiently obtained. This provided a solid data foundation for the simulation and deduction of the dual-light fusion virtual imaging model and the determination of the optimal fusion strategy, and improved the pertinence and feasibility of model optimization.

[0041] Furthermore, the method provided in this application embodiment includes: An environment-induced performance degradation coding sequence of the optical system is embedded in the dual-light fusion virtual imaging model. The environment-induced performance degradation coding sequence adopts discrete coding to characterize the performance degradation parameters of the dual-light imaging link under different temperatures, humidity, and fog concentrations. Based on the environment-induced performance degradation coding sequence, the signal-to-noise ratio attenuation process of the dual-light channel under different distances and weather combinations is simulated to generate the Pareto front for target recognition accuracy and the Pareto front for image processing power consumption. According to the current task priority, the optimal balance point in the non-dominated solution set is selected from the Pareto front for target recognition accuracy and the Pareto front for image processing power consumption.

[0042] Optionally, the dual-light fusion virtual imaging model adopts a four-level modular architecture: a data preprocessing layer, an environmental attenuation simulation layer, a dual-light fusion calculation layer, and a multi-objective optimization layer. Each layer is sequentially connected according to the data flow, forming a closed-loop operating logic. The data preprocessing layer, as the model's input interface, receives two types of core input data: one is dual-light sensing data, including raw thermal imaging image data, raw night vision image data, and derived thermal imaging signal-to-noise ratio and visible light contrast; the other is the mapping relationship data between different environmental parameters and imaging attenuation parameters stored in the image degradation mode knowledge base. The environmental attenuation simulation layer interacts bidirectionally with the data preprocessing layer, receiving preprocessed standardized data and feeding back attenuation correction parameters. The dual-light fusion calculation layer receives the attenuated dual-light data output from the environmental attenuation simulation layer and performs fusion calculations. The multi-objective optimization layer is directly connected to the dual-light fusion calculation layer, conducting optimization analysis based on the fusion results and ultimately outputting the target data.

[0043] Next, the embedding of the environmentally induced performance degradation coding sequence was carried out. An integer discrete coding method was used, first dividing temperature and humidity into several intervals: temperature was divided into 12 intervals ranging from -20℃ to 40℃ in 5℃ increments; humidity was divided into 7 intervals ranging from 30% to 90% in 10% increments; and fog concentration was divided into three levels based on visibility: no fog, moderate fog, and heavy fog. Then, a unique integer coding value was assigned to each temperature and humidity interval combination and fog concentration level. Performance degradation parameters of the dual-light imaging link under the corresponding coding values ​​were extracted from the image degradation pattern knowledge base, including the signal-to-noise ratio attenuation coefficient of the thermal imaging channel and the contrast reduction rate of the visible light channel. A one-to-one correspondence between the coding values ​​and the attenuation parameters was established. This correspondence was then embedded as the environmentally induced performance degradation coding sequence into the environmental degradation simulation layer, enabling the model to quantify the environmental impact.

[0044] Then, the parameters of each module of the model were configured and the functions were debugged. The data preprocessing layer adopted a data standardization method to normalize the input dual-light sensing data to the 0-1 range, eliminating dimensional differences. The dual-light fusion calculation layer adopted a weighted fusion algorithm, with preset initial fusion weights for thermal imaging data and night vision data. The weight values ​​were limited to the range of 0-1, and these weights would be dynamically adjusted by the multi-objective optimization layer. The multi-objective optimization layer integrated the open-source NSGA-II algorithm framework, configuring the algorithm iteration limit to 100 times and the convergence threshold to 0.001 to ensure optimization efficiency and accuracy. After the debugging of each module was completed, the model was built. Its core connection relationship is that the input data, after being standardized by the preprocessing layer, is passed to the environmental attenuation simulation layer for attenuation correction in combination with the encoding sequence. The corrected dual-light data is sent to the fusion calculation layer to complete the fusion. The fusion result and the corresponding processing power consumption data are input into the multi-objective optimization layer for analysis.

[0045] Subsequently, multi-dimensional simulations were conducted using the imaging state data from the imaging-scene correlation matrix obtained in the preceding steps. Target distance ranges were selected, including four dimensions: 0-100 meters, 100-300 meters, 300-500 meters, and above 500 meters. Weather types were selected, including four dimensions: sunny, cloudy, foggy, and rainy, forming 16 typical simulation scenarios. For each scenario, the environmental attenuation simulation layer called the attenuation parameters of the corresponding encoded sequence to perform signal-to-noise ratio attenuation processing on the dual-light sensing data. Then, the dual-light fusion calculation layer generated a fused image according to the current fusion weights. The lightweight target recognition algorithm YOLOv5s was used to perform target recognition on the fused image, and the recognition accuracy was calculated as the target recognition precision index. The computation time and energy consumption of the scenario data were processed using a power consumption monitoring tool and statistical model, serving as the image processing power consumption index. Each scenario was simulated 10 times, and the average value of the index was used to form the sample data corresponding to that scenario.

[0046] Subsequently, based on sample data from all scenarios, the NSGA-II algorithm in the multi-objective optimization layer is used to generate Pareto fronts. First, the algorithm performs a fast non-dominated sorting of the sample data, comparing the target recognition accuracy and image processing power consumption of each sample. If the recognition accuracy of one sample is not lower than that of another sample and its power consumption is not higher, then the former is determined to dominate the latter, and all non-dominated samples not dominated by other samples are selected. The non-dominated samples are sorted from high to low target recognition accuracy to form the target recognition accuracy Pareto front; and sorted from low to high image processing power consumption to form the image processing power consumption Pareto front. These two fronts together constitute the solution space of the multi-objective optimization.

[0047] Finally, the optimal balance point is selected based on the current task priority. If the task priority is target recognition, such as in a hunting scenario, samples that meet the preset accuracy requirements are selected from the Pareto front of target recognition accuracy. The sample with the lowest power consumption is then selected as the optimal balance point. The target recognition accuracy threshold corresponding to this sample is the target recognition confidence threshold. When the subsequent actual recognition accuracy is lower than this threshold, an image enhancement command is triggered. The combination of parameters such as digital gain, fusion weight, and AI inference frequency corresponding to this sample constitutes the image fusion adjustment strategy, clarifying the fusion ratio of dual-light data and the allocation rules of processing resources. If the task priority is low power consumption, such as in a long-term patrol scenario, samples that meet the preset power consumption constraints are selected from the Pareto front of image processing power consumption. The sample with the highest recognition accuracy is then selected as the optimal balance point, and the target recognition confidence threshold and image fusion adjustment strategy are determined using the same logic.

[0048] By constructing a four-level modular dual-light fusion virtual imaging model, embedding a discretely coded environmental attenuation sequence, and combining traversal simulation with the NSGA-II multi-objective optimization algorithm, the target recognition confidence threshold and image fusion adjustment strategy were accurately obtained, effectively improving the imaging adaptability and target recognition reliability of the night vision sight in complex environments.

[0049] Furthermore, the method provided in this application embodiment includes: A wavelet transform unit is integrated on the edge processing node to obtain a subset of target edge sensitive features. The subset of target edge sensitive features is fused with the illumination intensity and incident angle in the ambient illumination parameters and used as the input of the gated recurrent unit network to predict the image sharpness change trend within the time sliding window. Combined with the normal sharpness range, it is determined whether to trigger an image enhancement or refocusing command.

[0050] Specifically, firstly, in the wavelet transform unit integrated in the edge processing node, the image noise intensity signal is decomposed into feature components of multiple frequency bands. Then, energy entropy and kurtosis features are extracted from these feature components to generate time-frequency domain sensitive feature vectors. Finally, a subset of target edge sensitive features that are strongly correlated with target contour degradation is selected. This step will be explained in detail later.

[0051] Next, the acquired target edge sensitive feature subset and ambient lighting parameters are preprocessed and fused. Using a high-precision light sensor mounted on the night vision sight, real-time data on light intensity and incident angle are collected from the ambient lighting parameters. A moving average filtering method is used to smooth the collected lighting data, eliminating transient fluctuations. Simultaneously, the Min-Max normalization method is used to normalize the target edge sensitive feature subset and the preprocessed light intensity and incident angle data to the 0-1 range, eliminating dimensional differences between different parameters. Then, the normalized target edge sensitive feature subset (with N dimensions) is sequentially concatenated with the light intensity and incident angle data (each 1 dimension) to form an N+2 dimensional joint input feature vector, ensuring complete integration of feature information.

[0052] Secondly, a Gated Recurrent Unit (GRU) network adapted to edge processing nodes is configured. This network employs a lightweight architecture, consisting of an input layer, one hidden layer, and an output layer. The input layer dimension is consistent with the joint input feature vector dimension, i.e., N+2 dimensions. The number of hidden layer nodes is set to 64 to balance computational efficiency and prediction accuracy. The output layer dimension is 1 dimension, corresponding to the predicted image sharpness value. The tanh activation function is used for hidden layer neuron computation, and the sigmoid function is used for the logical judgment of updating and resetting the gates. The Adam optimizer is used, with a learning rate set to 0.001 and 500 iterations to avoid overfitting and low training efficiency. The training data uses historical data from a night vision scope under different environments (low light, fog, polar night) and target distances, including joint input features and corresponding actual image sharpness values. Supervised training allows the GRU network to learn the mapping relationship between features and sharpness changes.

[0053] Next, a time-sliding window is set up to capture continuous changing trends. A fixed-length sliding window method is adopted, with the window size set to 8 timestamps, each timetamp spaced 0.5 seconds apart, and the sliding step size being 1 timestamp. This ensures that the trend of light intensity decay and image sharpness changes over time are fully captured while maintaining real-time performance. The joint input feature vector within each sliding window is sequentially fed into a trained GRU network. The network memorizes historical feature information through the gating mechanism of the hidden layers and outputs the predicted image sharpness value at the end of the window, achieving accurate prediction of the trend of image sharpness changes within the time-sliding window.

[0054] Then, the normal image sharpness range is determined and an adaptive alarm level is matched. Based on historical imaging data of the night vision sight in various typical usage scenarios, the image sharpness baseline value for different scenarios is calculated. The normal sharpness range for that scenario is defined as the baseline value ± 2 standard deviations, and stored in the parameter configuration library of the edge processing node. The sharpness prediction value output by the GRU network is compared with the normal sharpness range of the corresponding scenario. If the prediction value is within the normal range, the "no alarm" level is matched, and no command is triggered. If the prediction value is below the lower limit of the normal range but the deviation is ≤30%, the "mild alarm" level is matched, and a mild image enhancement command is triggered. If the deviation of the prediction value is between 30% and 60%, the "moderate alarm" level is matched, and a moderate image enhancement command is triggered. If the deviation of the prediction value is >60%, the "severe alarm" level is matched, a refocus command is triggered, and high-intensity image enhancement is executed simultaneously, completing the adaptive alarm level matching and command triggering logic.

[0055] Finally, the specific compensation operations for image enhancement or refocusing commands are executed. Mild image enhancement commands use a Gamma correction algorithm to adjust the Gamma value to 1.2-1.4, improving image brightness and contrast. Medium image enhancement commands combine Gamma correction and bilateral filtering algorithms to improve image quality while preserving target edge details. High-intensity image enhancement commands use a multi-scale Retinex algorithm to enhance the image's dynamic range and detail. The refocusing command controls the night vision scope's motorized focusing module. Based on the current target distance data and the sharpness prediction deviation, it calls a preset focusing parameter mapping table to quickly adjust the lens focal length to the optimal state, achieving rapid compensation for image sharpness.

[0056] By employing feature fusion, lightweight GRU network prediction, sliding window analysis, and hierarchical alarm command triggering, adaptive alarm level matching and image enhancement command compensation are achieved, effectively improving the imaging stability and target recognition continuity of the night vision sight in dynamic environments, and adapting to the low computing power and low power consumption constraints of edge processing nodes.

[0057] Furthermore, the method provided in this application embodiment includes: The image noise intensity signal is decomposed into feature components of multiple frequency bands; energy entropy and kurtosis features are extracted from the feature components of the multiple frequency bands to generate time-frequency domain sensitive feature vectors, and a subset of target edge sensitive features that are strongly correlated with target contour degradation is selected.

[0058] Specifically, firstly, an edge processing node adapted to the night vision sight hardware is deployed. This node integrates an image noise acquisition module and a data preprocessing unit, and uses a low-power noise sensor to acquire the image noise intensity signal in real time during the dual-light fusion imaging process. For the acquired raw noise signal, a mean filtering algorithm is used for preliminary denoising to eliminate the influence of random pulse interference on the signal, resulting in a smoothed image noise intensity signal, which lays the foundation for subsequent feature extraction.

[0059] Next, based on the db4 wavelet transform method, the smoothed image noise intensity signal is decomposed into multiple scales in the wavelet transform unit of the edge processing node. The decomposition level is set to 3 levels, and the signal is decomposed into 1 low-frequency approximate feature component and 3 high-frequency detail feature components through wavelet transform. The low-frequency component reflects the overall trend of the signal, while the high-frequency components contain detailed information such as target edges and noise components, thus achieving signal separation in the time and frequency domains.

[0060] Then, the feature components of each frequency band obtained from the decomposition are processed. For each feature component of a frequency band, its energy entropy is calculated according to the information entropy calculation formula to quantify the uniformity of the energy distribution of the signal in that frequency band; at the same time, the kurtosis value of each component is solved using the kurtosis calculation formula to characterize the signal peak characteristics and the degree of deviation from the normal distribution. The energy entropy and kurtosis value of each frequency band are used as feature dimensions and combined in order of frequency band to form a time-frequency domain sensitive feature vector with an 8-dimensional dimension, i.e., 4 frequency bands × 2 features.

[0061] Subsequently, an attention mechanism is introduced to enhance and filter detail-sensitive features. A soft attention weight allocation method is employed, calculating the Pearson correlation coefficient between each feature dimension in the time-frequency domain sensitive feature vector and the target contour degradation. Attention weights are assigned based on the correlation coefficient, with higher correlation coefficients resulting in greater weight allocation, thus highlighting features strongly correlated with changes in the target contour. A weight threshold of 0.7 is set, filtering out feature dimensions with weights higher than this threshold and eliminating redundant features with low correlation, ultimately forming a subset of target edge-sensitive features enhanced by the attention mechanism.

[0062] By deploying edge processing nodes to collect and preprocess noise signals, and combining db4 wavelet transform, statistical feature extraction and attention mechanism weighted filtering, a subset of target edge sensitive features that are strongly correlated with target contour degradation was accurately obtained. This provides highly reliable feature support for subsequent prediction of image clarity change trends and adaptive command triggering, and improves the sensitivity of night vision sights to target details.

[0063] Furthermore, the method provided in this application embodiment includes: Based on the multi-level confidence separation mask under dual-light fusion of the night vision sight, key imaging indicators including hot spot drift rate and edge sharpness reduction rate are extracted; based on the key imaging indicators, combined with the current target distance and ambient lighting conditions, dynamic scheduling parameters and priority response rules for image enhancement resources are configured.

[0064] In one embodiment, firstly, based on the constructed dual-light fusion multi-level confidence separation mask for night vision sights, the effective target area and potential hotspot areas in the fused image are located. Using the Otsu adaptive threshold segmentation algorithm, a grayscale threshold is calculated for the fused image corresponding to the mask. Areas with grayscale values ​​higher than the threshold are identified as hotspot candidate areas. Interference areas corresponding to extremely low confidence levels in the mask are removed, resulting in a clean hotspot analysis area. The SIFT feature matching algorithm is used to extract the core feature points of the hotspot candidate areas from 10 consecutive frames of the fused image. The center coordinates of the hotspots in each frame are determined through feature point matching. The Euclidean distance between the center coordinates of the hotspots in adjacent frames is calculated and divided by the inter-frame time interval to obtain the hotspot drift rate, thus completing the extraction of this index.

[0065] Next, the edge sharpness reduction rate is extracted. Based on the high and medium confidence level regions in the multi-level confidence separation mask, the effective edge range of the target contour is determined. The Sobel operator edge detection method is used to calculate the gradient in the horizontal and vertical directions of the target edge range, obtaining the edge gradient magnitude matrix. The mean of all non-zero elements in this matrix is ​​calculated as the edge sharpness value of the current frame image. The baseline edge sharpness value in the baseline imaging feature value corresponding to the imaging-scene correlation matrix in the previous steps is retrieved. By calculating the percentage of (baseline edge sharpness value - current edge sharpness value) / baseline edge sharpness value, the edge sharpness reduction rate is obtained, thus achieving the complete extraction of key imaging indicators.

[0066] Then, the current target distance and ambient lighting conditions are acquired. Using the laser rangefinder module mounted on the night vision scope, the actual distance between the current target and the scope is measured and output, retaining the distance value to an integer. Real-time ambient light intensity and incident angle data are collected using a light sensor. A moving average filtering method is used to smooth the collected data, eliminating instantaneous fluctuations and obtaining stable ambient lighting parameters. The extracted hotspot drift rate and edge sharpness reduction rate are correlated with the current target distance and ambient lighting parameters to construct a four-dimensional feature dataset, providing data support for resource scheduling and configuration.

[0067] Subsequently, dynamic scheduling parameters for image enhancement resources are configured. Weights are assigned to hotspot drift rate, edge sharpness reduction rate, target distance, and ambient light intensity, set to 0.3, 0.3, 0.2, and 0.2 respectively. A resource scheduling requirement score is obtained through weighted summation. Dynamic scheduling parameters are configured based on the score: For example, when the score is below 30, the scheduling parameters are set to basic enhancement mode, with 2 threads allocated to the image enhancement algorithm and a processing frame rate of 30fps; when the score is between 30 and 60, the scheduling parameters are set to intermediate enhancement mode, with 4 threads allocated and a processing frame rate adjusted to 25fps; when the score is above 60, the scheduling parameters are set to advanced enhancement mode, with 6 threads allocated and a processing frame rate maintained at 20fps, ensuring precise matching between enhancement resources and imaging optimization requirements.

[0068] Finally, priority response rules were established. Specifically, based on different modes of dynamic scheduling parameters, response priorities were divided according to target distance and ambient lighting conditions: In advanced enhancement mode, if the target distance is ≤300 meters and the light intensity is <10 lux, the priority is set to level one, prioritizing the use of multi-scale Retinex enhancement algorithm and bilateral filtering algorithm to ensure clear target details; if the target distance is >300 meters and the light intensity is <10 lux, the priority is set to level two, prioritizing the use of Gamma correction algorithm and edge enhancement algorithm to balance sharpness and computational efficiency. In intermediate enhancement mode, regardless of the target distance, if the light intensity is ≥10 lux and the edge sharpness reduction rate is <20%, the priority is set to level three, calling the basic contrast enhancement algorithm; if the edge sharpness reduction rate is ≥20%, the priority is set to level two. In basic enhancement mode, the priority is uniformly set to level four, only enabling mild noise suppression and brightness adjustment algorithms to reduce resource consumption.

[0069] By extracting key imaging indicators and combining them with target distance and ambient lighting conditions, a weighted decision-making and hierarchical response mechanism was adopted to achieve precise configuration of dynamic scheduling parameters and priority response rules for image enhancement resources. This effectively improved the utilization efficiency of image enhancement resources for night vision sights and ensured the optimization targeting and computational economy under different imaging scenarios.

[0070] Furthermore, the method provided in this application embodiment includes: The image enhancement resources interact with the dual-light fusion virtual imaging model in real time; the processing load of the image enhancement resources is dynamically adapted to the battery life window of the edge processing node, and the image degradation mode knowledge base is iteratively updated using a deep reinforcement learning algorithm.

[0071] Optionally, firstly, a real-time interactive architecture is established between the image enhancement resource and the dual-light fusion virtual imaging model. A low-latency data transmission channel is constructed using the Socket TCP / IP communication protocol, and the standardized data format and communication timing of the interaction interface are clearly defined. The image enhancement resource, as the data sender, collects its own operating status data in real time, including the currently enabled enhancement algorithm type, the number of computing threads, the processing frame rate, the signal-to-noise ratio and contrast feedback value of the enhanced image, and pushes data to the dual-light fusion virtual imaging model at 0.05-second intervals after each frame is processed. The dual-light fusion virtual imaging model, as the data receiver and response end, receives the above status data in real time. At the same time, based on the latest results of multi-dimensional simulation and deduction, it feeds back dynamically updated fusion adjustment strategy parameters to the image enhancement resource, such as fusion weight correction values, target recognition confidence threshold update values, and real-time correction data of the environmental induced performance degradation coding sequence, ensuring the real-time and synchronous nature of the data interaction between the two.

[0072] Next, a real-time load monitoring module was deployed. Based on the hardware resource characteristics of the edge processing nodes, embedded system resource monitoring APIs were used, such as the perf tool for ARM architecture and the proc file system interface for embedded Linux, to collect core load indicators in real time, including CPU utilization, GPU computing load, memory usage, and data read / write bandwidth. The load monitoring sampling period was set to 0.1 seconds, and the raw load data was smoothed using a moving average filtering algorithm to eliminate instantaneous fluctuations and obtain stable processing load values, providing an accurate basis for subsequent adaptation decisions.

[0073] Simultaneously, a battery life monitoring module is deployed to acquire real-time battery data, including remaining battery percentage, current discharge current, battery temperature, and voltage, through a battery management system (BMS) integrated into the edge processing node. A battery life estimation algorithm is employed, based on the remaining battery percentage and current discharge rate, combined with power consumption models for different operating modes of the night vision sight, such as average power consumption data for basic imaging mode and enhanced imaging mode, to calculate the battery life window in real time—that is, the continuous working time supported by the current battery level. The battery life status data is updated at 0.5-second intervals and synchronized to the dynamic adaptation decision module.

[0074] Subsequently, a dynamic adaptation decision model was constructed, and a weighted scoring method was used to comprehensively evaluate processing load and battery life window. The weight of processing load was set to 0.6, and the weight of battery life window was set to 0.4. The processing load was divided into three levels according to CPU utilization: low load (≤50%), medium load (50%-80%), and high load (>80%), with corresponding scores of 80, 50, and 20 points, respectively. The battery life window was divided into three levels according to continuous working time: long battery life (≥3 hours), medium battery life (1-3 hours), and short battery life (<1 hour), with corresponding scores of 90, 60, and 30 points, respectively. A comprehensive adaptation score is obtained through weighted summation. Based on the score, a three-level adaptation strategy is formulated: When the comprehensive score is ≥70, a high-performance mode is executed, and complex enhancement algorithms such as multi-scale Retinex and bilateral filtering are used for image enhancement resources. The simulation parameters of the dual-light fusion virtual imaging model are configured with the highest precision. When the comprehensive score is between 50 and 70, a balanced mode is executed, and Gamma correction and basic edge enhancement algorithms are used for image enhancement resources. The simulation iteration number of the model is appropriately reduced to control the load. When the comprehensive score is <50, an energy-saving mode is executed, and mild noise suppression and brightness adjustment algorithms are used for image enhancement resources. The model calls lightweight simulation logic to reduce the amount of computation.

[0075] A dynamic feedback and adjustment mechanism for the adaptation strategy is then established. During the execution of the current adaptation strategy, changes in processing load and runtime window are continuously monitored. If the processing load exceeds the upper limit of the corresponding mode's threshold for three consecutive sampling periods, or the runtime window shortens at a rate exceeding 0.2 hours / minute, a strategy downgrade adjustment is triggered. If the processing load is below the lower limit of the threshold for five consecutive sampling periods, and the runtime window does not shorten significantly, a strategy upgrade adjustment is triggered. Simultaneously, the image enhancement resources provide real-time feedback of the adjusted operating status to the dual-light fusion virtual imaging model. The model synchronously optimizes simulation parameters and data push frequency, forming a closed-loop adaptation logic of "monitoring-decision-execution-feedback".

[0076] Finally, a deep reinforcement learning algorithm is used to iteratively update the image degradation pattern knowledge base. First, a two-dimensional reward function is set, which includes the first dimension of the target recognition accuracy and the matching degree of the actual target category, and the second dimension of the image processing delay. Then, based on this two-dimensional reward function, the optimal fusion strategy is learned in the discrete target type space through a deep Q network. This step will be explained in detail in the following content.

[0077] Through a standardized real-time interactive architecture, multi-dimensional status monitoring, and hierarchical dynamic adaptation strategy, efficient collaboration between image enhancement resources and dual-light fusion virtual imaging models, as well as precise adaptation between processing load and battery life window, are achieved. While ensuring imaging optimization effects, the resource utilization efficiency and battery life of edge processing nodes are maximized, adapting to the portable usage scenarios of night vision sights.

[0078] Furthermore, the method provided in this application embodiment includes: A two-dimensional reward function is set up, where the first dimension is the matching degree between the target recognition accuracy and the actual target category, and the second dimension is the image processing latency. Based on the two-dimensional reward function, a deep Q-network is used to learn the optimal fusion strategy in the discrete target type space.

[0079] In one embodiment, firstly, a discrete target type space is defined. Based on typical usage scenarios of night vision scopes, such as hunting, security, and outdoor reconnaissance, targets are divided into four discrete target types: human body, small and medium-sized animals, vehicles, and static obstacles. The core features of each type of target are clarified, such as the outline proportion of human body and the geometric shape of vehicle. A standardized target type feature library is constructed to provide a foundation for subsequent matching degree calculation and strategy learning.

[0080] Next, the specific calculation logic of the two-dimensional reward function is set. The first dimension is the matching degree between the target recognition accuracy and the actual target category. Target recognition algorithms such as YOLOv5s are used to perform target recognition on the processed image. The ratio of the number of correctly recognized targets to the total number of actual targets in the image is counted to obtain the target recognition accuracy. At the same time, by calculating the cosine similarity between the target category and the actual target category in the recognition result, and based on the core feature vectors of each category in the target type feature library, the category matching degree is obtained. The target recognition accuracy and the category matching degree are weighted and summed with a weight of 0.6:0.4, and this sum is used as the first dimension reward value, with a value range of 0-10. The higher the score, the better the recognition effect. The second dimension, image processing latency, employs embedded system timing tools, such as the `clock_gettime` function in the ARM architecture, to record the total time elapsed from the input of dual-light sensor data to the edge processing node to the output of the optimized image. A latency threshold of 100 milliseconds is set. When the latency is ≤100 milliseconds, the second-dimensional reward value is 10 points; for every 10 milliseconds exceeding the latency threshold, the reward value decreases by 1 point, down to a minimum of 0 points, forming a negative correlation between latency and reward. The final dual-dimensional reward function is the arithmetic sum of the reward values ​​from both dimensions, with the total reward value ranging from 0 to 20 points.

[0081] Then, a lightweight Deep Q-Network (DQN) architecture adapted to edge processing nodes is constructed. The network adopts a three-layer structure of "input layer-hidden layer-output layer". The input layer is set to 12-dimensional, including key state features such as ambient light intensity, target distance, thermal imaging signal-to-noise ratio, visible light contrast, hot spot drift rate, and edge sharpness reduction rate. The hidden layer has two layers, each with 128 nodes, and the ReLU function is used to avoid the gradient vanishing problem. The output layer dimension is consistent with the discrete action space dimension. The discrete action space contains six typical dual-light fusion strategies, such as high-weight thermal imaging fusion, balanced fusion, high-weight visible light fusion, enhanced edge fusion, low-latency fast fusion, and high-fidelity fusion. Each node in the output layer corresponds to the Q-value (action value function value) of one strategy. The network optimizer uses the Adam optimizer with a learning rate of 0.001, and the mean squared error loss function is used to minimize the deviation between the predicted Q-value and the target Q-value.

[0082] Next, the training dataset for the deep Q-network was prepared. Historical operational data for night vision sights was collected under various environments, including low light, fog, and polar night, at different target distances (0-100m, 100-300m, 300-500m, and over 500m), and for different target types. Each data sample included a state feature vector, the executed fusion strategy (action), the corresponding two-dimensional reward value, and the state feature vector for the next time step. A total of 100,000 valid samples were collected and divided into training and validation sets in an 8:2 ratio. An experience replay mechanism was employed, constructing an experience replay buffer with a capacity of 10,000 samples. During training, samples with a batch size of 32 were randomly selected from the buffer for batch training to avoid training instability caused by sample correlation.

[0083] Next, the iterative training and policy learning process of the deep Q-network is executed. The network parameters are initialized with random normal distribution values. The target network and evaluation network are set to update synchronously. The target network parameters are fixed every 100 iterations, while the evaluation network continuously updates its weights using training set samples. During training, an ε-greedy policy is used to select actions. The initial ε value is set to 0.9 (the probability of randomly selecting an action), and it decreases linearly with the number of iterations. Every 500 iterations, the ε value decreases by 0.1 until it reaches 0.1 and remains there, balancing exploration and exploitation. The state feature vectors from the training set are input into the evaluation network, which outputs the Q-values ​​for each action. A fusion policy is selected and executed according to the ε-greedy policy, obtaining the corresponding reward value and the next state. The data of state, action, reward, and next state are stored in the experience replay buffer. The network performance is monitored in real time using a validation set. When the average reward value on the validation set remains stable above 15 points for 1000 consecutive steps with a fluctuation range ≤5%, the deep Q-network training is considered converged, and training is stopped.

[0084] Finally, the image degradation mode knowledge base is updated based on the converged deep Q-network. The optimal fusion strategy learned by the network is associated with the corresponding state features, such as the optimal strategy under specific ambient lighting, target type, and imaging degradation indicators, to generate a policy-state mapping table, which is then added to the image degradation mode knowledge base. Simultaneously, during the actual operation of the night vision scope, new operational data is continuously collected, and the deep Q-network is incrementally trained periodically (e.g., every hour of cumulative operation) according to the above training process, updating the policy-state mapping table. This achieves dynamic iterative optimization of the image degradation mode knowledge base, ensuring that the knowledge base always adapts to the latest imaging scenarios and equipment states.

[0085] By constructing a dual-dimensional reward function and a lightweight deep Q-network, combined with experience replay and an ε-greedy strategy, the optimal fusion strategy is efficiently learned in the discrete target type space. This enables iterative updates of the image degradation mode knowledge base, improves the adaptability of the knowledge base to complex imaging scenarios, provides more accurate degradation mode support for the dual-light fusion virtual imaging model, and ensures continuous optimization of the imaging effect of the night vision sight.

[0086] In summary, the dual-light fusion image optimization processing method for night vision sights provided in this application has the following technical effects: This application collects data such as dual-light sensing and ambient light, and processes it through dual-light fusion virtual imaging model simulation, edge processing node feature extraction, multi-level confidence separation mask segmentation, etc. Combined with fusion adjustment strategy and knowledge base iterative optimization, it realizes dual-light fusion image optimization of night vision sights. It achieves the technical effect of improving the target recognition stability and aiming decision reliability of night vision sights by adapting to environmental changes, collaboratively optimizing image effects and accurately shielding interference.

[0087] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a dual-light fusion image optimization processing system for night vision sights, the system comprising: The imaging status data acquisition module 1 collects imaging status data, including thermal imaging signal-to-noise ratio and visible light contrast, under the same time stamp based on dual-light sensing data and ambient light parameters.

[0088] Image fusion adjustment strategy acquisition module 2 is used to construct a dual-light fusion virtual imaging model based on the dual-light sensing data and the image degradation mode knowledge base, and to perform multi-dimensional simulation and deduction in combination with the imaging state data to determine the target recognition confidence threshold and image fusion adjustment strategy.

[0089] Image enhancement instruction compensation execution module 3 is used to deploy edge processing nodes, collect image noise intensity signals, combine the light intensity attenuation trend within the time sliding window of the ambient light parameters, set detail-sensitive features under the attention mechanism, and perform adaptive alarm level matching and image enhancement instruction compensation.

[0090] The image processing execution module 4 integrates the target recognition confidence threshold and the image fusion adjustment strategy to construct a multi-level confidence separation mask under dual-light fusion of the night vision aiming scope. The multi-level confidence separation mask is used to perform adaptive segmentation and interference region shielding on the fused image.

[0091] Furthermore, the fused image processing execution module 4 is used to perform the following steps: Based on the multi-level confidence separation mask under dual-light fusion of the night vision sight, key imaging indicators including hot spot drift rate and edge sharpness reduction rate are extracted; based on the key imaging indicators, combined with the current target distance and ambient lighting conditions, dynamic scheduling parameters and priority response rules for image enhancement resources are configured.

[0092] Furthermore, the fused image processing execution module 4 is used to perform the following steps: The image enhancement resources interact with the dual-light fusion virtual imaging model in real time; the processing load of the image enhancement resources is dynamically adapted to the battery life window of the edge processing node, and the image degradation mode knowledge base is iteratively updated using a deep reinforcement learning algorithm.

[0093] Furthermore, the fused image processing execution module 4 is used to perform the following steps: A two-dimensional reward function is set up, where the first dimension is the matching degree between the target recognition accuracy and the actual target category, and the second dimension is the image processing latency. Based on the two-dimensional reward function, a deep Q-network is used to learn the optimal fusion strategy in the discrete target type space.

[0094] Furthermore, the imaging state data acquisition module 1 is used to perform the following steps: The thermal imaging signal-to-noise ratio and visible light contrast are bound to the usage scenario type and target distance information to establish an imaging-scene correlation matrix. Based on the imaging-scene correlation matrix, the imaging differences of the same night vision sight in different environmental stages are identified, and benchmark imaging feature values ​​are extracted. The benchmark imaging feature values ​​are used to automatically identify imaging abnormal scenes and push them to the user display interface.

[0095] Furthermore, the image fusion adjustment strategy acquisition module 2 is used to perform the following steps: An environment-induced performance degradation coding sequence of the optical system is embedded in the dual-light fusion virtual imaging model. The environment-induced performance degradation coding sequence adopts discrete coding to characterize the performance degradation parameters of the dual-light imaging link under different temperatures, humidity, and fog concentrations. Based on the environment-induced performance degradation coding sequence, the signal-to-noise ratio attenuation process of the dual-light channel under different distances and weather combinations is simulated to generate the Pareto front for target recognition accuracy and the Pareto front for image processing power consumption. According to the current task priority, the optimal balance point in the non-dominated solution set is selected from the Pareto front for target recognition accuracy and the Pareto front for image processing power consumption.

[0096] Furthermore, the image fusion adjustment strategy acquisition module 2 is used to perform the following steps: Digital gain, fusion weight, and AI inference frequency are used as variable parameters to set constraints in the dual-light fusion virtual imaging model; based on the constraints, an effective sample set is determined, and non-dominated sorting is performed to obtain the non-dominated solution set.

[0097] Furthermore, the image enhancement instruction compensation execution module 3 is used to perform the following steps: A wavelet transform unit is integrated on the edge processing node to obtain a subset of target edge sensitive features. The subset of target edge sensitive features is fused with the illumination intensity and incident angle in the ambient illumination parameters and used as the input of the gated recurrent unit network to predict the image sharpness change trend within the time sliding window. Combined with the normal sharpness range, it is determined whether to trigger an image enhancement or refocusing command.

[0098] Furthermore, the image enhancement instruction compensation execution module 3 is used to perform the following steps: The image noise intensity signal is decomposed into feature components of multiple frequency bands; energy entropy and kurtosis features are extracted from the feature components of the multiple frequency bands to generate time-frequency domain sensitive feature vectors, and a subset of target edge sensitive features that are strongly correlated with target contour degradation is selected.

[0099] The dual-light fusion image optimization processing system for night vision sights provided in this embodiment of the invention can execute the dual-light fusion image optimization processing method for night vision sights provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0100] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0101] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A dual-light fusion image optimization processing method for night vision sights, characterized in that, The method includes: Based on dual-light sensing data and ambient lighting parameters, imaging status data including thermal imaging signal-to-noise ratio and visible light contrast are collected under the same time stamp. Based on the dual-light sensing data and the image degradation mode knowledge base, a dual-light fusion virtual imaging model is constructed, and multi-dimensional simulation is performed in conjunction with the imaging state data to determine the target recognition confidence threshold and image fusion adjustment strategy. Deploy edge processing nodes to collect image noise intensity signals, combine the light intensity decay trend within the time sliding window of the ambient light parameters, set detail-sensitive features under the attention mechanism, and perform adaptive alarm level matching and image enhancement command compensation. Meanwhile, based on the target recognition confidence threshold and image fusion adjustment strategy, a multi-level confidence separation mask under dual-light fusion of the night vision aiming scope is constructed, and the multi-level confidence separation mask is used to perform adaptive segmentation and interference region shielding on the fused image.

2. The dual-light fusion image optimization processing method for night vision sights as described in claim 1, characterized in that, The method further includes performing adaptive segmentation and interference region masking on the fused image using the multi-level confidence separation mask: Based on the multi-level confidence separation mask under dual-light fusion of the night vision sight, key imaging indicators including hot spot drift rate and edge sharpness reduction rate are extracted. Based on the key imaging indicators, combined with the current target distance and ambient lighting conditions, configure dynamic scheduling parameters and priority response rules for image enhancement resources.

3. The dual-light fusion image optimization processing method for night vision sights as described in claim 2, characterized in that, The image enhancement resources interact with the dual-light fusion virtual imaging model in real time. The processing load of the image enhancement resources is dynamically adapted to the battery life window of the edge processing node, and the image degradation mode knowledge base is iteratively updated using a deep reinforcement learning algorithm.

4. The dual-light fusion image optimization processing method for night vision sights as described in claim 3, characterized in that, The method of iteratively updating the image degradation mode knowledge base using a deep reinforcement learning algorithm includes: A two-dimensional reward function is set up, where the first dimension is the matching degree between the target recognition accuracy and the actual target category, and the second dimension is the image processing latency. Based on the aforementioned dual-dimensional reward function, a deep Q-network is used to learn the optimal fusion strategy in the discrete target type space.

5. The dual-light fusion image optimization processing method for night vision sights as described in claim 1, characterized in that, The method for collecting imaging status data, including thermal imaging signal-to-noise ratio and visible light contrast, under the same timestamp, includes: By binding thermal imaging signal-to-noise ratio and visible light contrast with usage scenario type and target distance information, an imaging-scene correlation matrix is ​​established; Based on the imaging-scene correlation matrix, the imaging differences of the same night vision sight in different environmental stages are identified, and benchmark imaging feature values ​​are extracted. The reference imaging feature values ​​are used to automatically identify imaging anomalies and push them to the user display interface.

6. The dual-light fusion image optimization processing method for night vision sights as described in claim 5, characterized in that, A dual-light fusion virtual imaging model is constructed, and multi-dimensional simulations are performed using the imaging state data to determine the target recognition confidence threshold and image fusion adjustment strategy. The method includes: An environment-induced performance degradation coding sequence of the optical system is embedded in the dual-light fusion virtual imaging model. The environment-induced performance degradation coding sequence adopts discrete coding to characterize the performance degradation parameters of the dual-light imaging link under different temperatures, humidity, and fog concentrations. Based on the environmentally induced performance degradation coding sequence, the signal-to-noise ratio degradation process of dual optical channels under different distances and weather combinations is simulated to generate Pareto fronts for target recognition accuracy and image processing power consumption. Based on the current task priority, the optimal balance point is selected from the Pareto front of the target recognition accuracy and the Pareto front of the image processing power consumption in the non-dominated solution set.

7. The dual-light fusion image optimization processing method for night vision sights as described in claim 6, characterized in that, The method includes: Digital gain, fusion weight, and AI inference frequency are used as variable parameters to set constraints in the dual-light fusion virtual imaging model. Based on the constraints, a valid sample set is determined, and a non-dominated solution set is obtained by performing a non-dominated sort.

8. The dual-light fusion image optimization processing method for night vision sights as described in claim 1, characterized in that, Deploying edge processing nodes and acquiring image noise intensity signals, the method includes: By integrating wavelet transform units on the edge processing node, a subset of target edge-sensitive features is obtained; The target edge sensitive feature subset is fused with the illumination intensity and incident angle in the ambient illumination parameters and used as the input of the gated recurrent unit network to predict the image sharpness change trend within the time sliding window. Combined with the normal sharpness range, it is determined whether to trigger an image enhancement or refocusing command.

9. The dual-light fusion image optimization processing method for night vision sights as described in claim 8, characterized in that, Integrating wavelet transform units on the edge processing node to obtain a subset of target edge-sensitive features, the method includes: The image noise intensity signal is decomposed into feature components of multiple frequency bands; Energy entropy and kurtosis features are extracted from the feature components of the multiple frequency bands to generate time-frequency domain sensitive feature vectors, and a subset of target edge sensitive features that are strongly correlated with target contour degradation is selected.

10. A dual-light fusion image optimization processing system for night vision sights, characterized in that, The system is used to implement the dual-light fusion image optimization processing method for night vision sights according to any one of claims 1-9, the system comprising: The imaging status data acquisition module collects imaging status data, including thermal imaging signal-to-noise ratio and visible light contrast, under the same time stamp, based on dual-light sensor data and ambient light parameters. The image fusion adjustment strategy acquisition module is used to construct a dual-light fusion virtual imaging model based on the dual-light sensing data and the image degradation mode knowledge base, and to perform multi-dimensional simulation and deduction in combination with the imaging state data to determine the target recognition confidence threshold and the image fusion adjustment strategy. The image enhancement instruction compensation execution module is used to deploy edge processing nodes, collect image noise intensity signals, combine the light intensity attenuation trend within the time sliding window of the ambient light parameters, set detail-sensitive features under the attention mechanism, and perform adaptive alarm level matching and image enhancement instruction compensation. The image processing execution module integrates the target recognition confidence threshold and the image fusion adjustment strategy to construct a multi-level confidence separation mask for dual-light fusion of the night vision aiming scope. The multi-level confidence separation mask is then used to perform adaptive segmentation and interference region shielding on the fused image.