A vision system and method based on multi-scene and visible light image recognition

By introducing polarized light information and dynamically adjusting the visible light band, combined with cascaded neural networks and multi-level threshold discrimination mechanisms, the problem of unstable recognition in high-reflectivity environments of traditional vision systems is solved, achieving high accuracy and robust recognition in complex environments.

CN120563992BActive Publication Date: 2025-10-31HANGZHOU HUANYU VISION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511066166.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-31
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional vision systems struggle to suppress reflective interference in highly reflective environments, exhibit unstable recognition performance under varying lighting conditions, and lack sufficient accuracy and robustness in complex environments due to reliance solely on visible light feature extraction.

Method used

By introducing polarized light information and constructing a cascaded neural network architecture, the visible light band range and multi-level threshold discrimination mechanism are dynamically adjusted. The polarization and spectral features are fused together. Multi-dimensional data are collected synchronously using a CMOS array and a light intensity sensor. Reflection suppression weights and dynamic nonlinear compensation surfaces are constructed for image enhancement and feature extraction.

Benefits of technology

It effectively eliminates interference from high-reflectivity areas, improves image clarity and recognition accuracy, enhances the system's adaptability and stability in complex environments, and enables rapid response and precise positioning to abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563992B_ABST
    Figure CN120563992B_ABST
Patent Text Reader

Abstract

This invention discloses a vision system and method based on multi-scene and visible light image recognition, relating to the field of image recognition technology. The method includes: simultaneously acquiring polarization light information and light intensity information in the target scene to construct a vector parameter matrix; constructing reflection suppression weights based on the vector parameter matrix to generate a reflection-suppressed light intensity image, and constructing a dynamic nonlinear compensation surface by combining the light source color temperature and absolute illumination intensity to dynamically enhance the image; dynamically adjusting the visible light band range according to the light source color temperature, decomposing the enhanced image into three-band feature maps, and extracting polarization and spectral features respectively through a cascaded neural network architecture, fusing them to generate fused features; performing sliding window Fourier transforms on the fused features and the three-band feature maps respectively to extract frequency domain features, calculating the comprehensive difference degree, and triggering abnormal alarms or dynamically adjusting processing strategies based on a multi-level threshold discrimination mechanism to improve the adaptability and recognition capability of the vision system in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, specifically to a vision system and method based on multi-scene and visible light image recognition. Background Technology

[0002] In the field of image recognition technology, traditional vision systems mainly rely on single-dimensional visible light image data for target recognition and scene analysis. These systems generate two-dimensional images by capturing the light intensity distribution in the environment and use computer vision algorithms to detect, classify and track targets in the images.

[0003] However, in highly reflective environments such as tunnels, it is difficult to effectively suppress reflective interference, which leads to a decrease in image quality, difficulty in extracting target features, and thus affects the overall recognition accuracy. In scenarios with large changes in lighting conditions, such as day-night transitions or sudden weather changes, existing systems often cannot automatically adjust to adapt to the new lighting environment, resulting in unstable recognition performance.

[0004] Furthermore, feature extraction is often limited to the visible light band, ignoring information from other dimensions such as polarized light. This makes it difficult to fully capture target features in complex environments, reducing the accuracy and robustness of recognition. In particular, under low light or strong light interference conditions, single-band feature extraction methods are more susceptible to environmental factors. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a vision system and method based on multi-scene and visible light image recognition. By introducing polarized light information, dynamically adjusting the visible light band range, constructing a cascaded neural network architecture, and establishing a multi-level threshold discrimination mechanism, the adaptability and recognition capability of the vision system in complex environments are improved.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the present invention provides the following technical solution: a visual method based on multi-scene and visible light image recognition, comprising:

[0009] Simultaneously acquire polarization light information and light intensity information in the target scene to construct a vector parameter matrix;

[0010] Based on the vector parameter matrix, a reflection suppression weight is constructed to generate a light intensity image after reflection suppression. A dynamic nonlinear compensation surface is constructed by combining the light source color temperature and absolute light intensity to dynamically enhance the image.

[0011] The visible light band range is dynamically adjusted according to the color temperature of the light source. The enhanced image is decomposed into three-band feature maps, and polarization features and spectral features are extracted separately through a cascaded neural network architecture and fused to generate a fused feature.

[0012] Sliding window Fourier transforms are performed on the fused features and the three-band feature maps respectively to extract frequency domain features, calculate the comprehensive difference degree, and trigger anomaly alarms or dynamically adjust the processing strategy based on a multi-level threshold discrimination mechanism.

[0013] Furthermore, polarization information is acquired through a CMOS array deployed in the target scene. The CMOS array consists of multiple microlenses, and each node captures horizontal, vertical, 45° oblique, and 135° oblique polarization information through four-way polarization state separation. Light intensity information is acquired through a three-in-one light intensity sensor array deployed in the target scene. The light intensity sensor array simultaneously acquires absolute light intensity, light source color temperature, and light source azimuth angle.

[0014] Furthermore, a reflection suppression weight is constructed based on the vector parameter matrix. The reflection suppression weight includes calculating polarization feature weight and non-polarization feature weight. The polarization feature weight is used to suppress polarization interference in high-reflection areas, and the non-polarization feature weight is used to generate a reflection suppression image. Finally, the reflection suppression vector parameter matrix is ​​output.

[0015] Furthermore, a nonlinear compensation surface is constructed using absolute illumination intensity and light source color temperature to map pixel values ​​to the range of [0, 1]. Pixel-level transformation is performed based on the nonlinear compensation surface, and a compensation coefficient is applied to the shortwave band to suppress color shift of cold light source. The image is then dynamically enhanced, and the contrast of the enhanced image is calculated. When the contrast is less than the contrast threshold, the parameters of the nonlinear compensation surface are recalibrated.

[0016] Furthermore, when the color temperature of the light source is less than 4000K, the long-wave band is extended to 650nm, the short-wave band is compressed to 470~550nm, and the medium-wave band is fixed at 550~600nm; when the color temperature of the light source is greater than 5000K, the short-wave band is lowered to 430nm, the long-wave band is limited to 600~630nm, and the medium-wave band is adjusted to 500~580nm.

[0017] Based on the three characteristic bands of visible light, the enhanced image is decomposed to obtain characteristic maps of short-wave, medium-wave, and long-wave, i.e., three-band characteristic maps.

[0018] Furthermore, the cascaded neural network architecture includes a polarization branch and a spectral branch, with the polarization branch taking into account the linearly polarized components after reflection suppression. and Polarization features are extracted through deformable convolution and channel attention mechanism; spectral features are extracted from the three-band feature map input by spectral branch through residual network and attention mechanism.

[0019] Furthermore, the polarization features are upsampled to the same resolution as the spectral features using bilinear interpolation. The polarization and spectral features are then concatenated along the channel dimension to output a fused feature. ,in, For spectral characteristics, Polarization characteristics This indicates a splicing operation. R Let be the set of real numbers. H For height, W For width, C This represents the number of channels.

[0020] Furthermore, the difference in fused features is calculated: ,in, This represents the historical baseline value of the fusion feature under normal conditions. Frequency domain features for fusion features;

[0021] Calculate the difference in characteristics across the three bands: ,in, For the first i Frequency domain characteristics of the band For the first i Historical baseline values ​​for the band under normal conditions S It is a shortwave. M For medium wave, L For long waves, the difference between the fused features and the difference between the three-band features are weighted to generate a comprehensive difference.

[0022] Furthermore, the multi-level threshold discrimination mechanism includes: Level 1 discrimination: when the comprehensive difference is greater than or equal to D1, an audible and visual alarm is triggered, and the lighting intensity in the abnormal area is increased; Level 2 discrimination: when the comprehensive difference is greater than D2 and less than D1, the band range is dynamically adjusted according to real-time color temperature data; Level 3 discrimination: when the comprehensive difference is less than or equal to D2, the historical benchmark value is updated using a moving average strategy.

[0023] A vision system based on multi-scene and visible light image recognition, comprising:

[0024] The data acquisition module synchronously acquires polarization light information and light intensity information in the target scene and constructs a vector parameter matrix;

[0025] The image preprocessing module constructs reflection suppression weights based on the vector parameter matrix, generates a light intensity image after reflection suppression, and constructs a dynamic nonlinear compensation surface by combining the light source color temperature and absolute illumination intensity to dynamically enhance the image.

[0026] The feature extraction and fusion module dynamically adjusts the visible light band range according to the color temperature of the light source, decomposes the enhanced image into three-band feature maps, and extracts polarization features and spectral features respectively through a cascaded neural network architecture, and then fuses them to generate fused features.

[0027] The anomaly detection module performs a sliding window Fourier transform on the fused features to extract frequency domain features, calculates the comprehensive difference between the current fused features and the historical benchmark, and triggers anomaly alarms or dynamically adjusts the processing strategy based on a multi-level threshold discrimination mechanism.

[0028] (III) Beneficial Effects

[0029] This invention provides a vision system and method based on multi-scene and visible light image recognition, which has the following beneficial effects:

[0030] (1) By rationally deploying sensor arrays at the top and entrance of the tunnel, it is possible to accurately collect multi-dimensional data such as polarized light and light intensity, construct a vector parameter matrix, and ensure data accuracy through spatial and temporal alignment operations, providing a reliable foundation for subsequent processing, effectively coping with tunnel reflection interference, comprehensively perceiving environmental information, helping to accurately extract target features, and improving the stability and accuracy of the overall vision system.

[0031] (2) By polarization difference de-reflection, interference in high-reflection areas is effectively eliminated, improving image clarity. Illuminance and color temperature compensation construct a dynamic nonlinear compensation surface to ensure balanced image contrast, suppress color shift, and automatically recalibrate the surface parameters according to the image contrast after correction, making the pre-processed image quality better, providing a reliable foundation for subsequent feature processing and recognition, and enhancing the system's adaptability in complex tunnel environments.

[0032] (3) By dynamically adjusting the range of feature bands according to the color temperature of the light source to adapt to different lighting conditions, and by using a cascaded neural network architecture to extract polarization features and spectral features respectively and then fusing them, the comprehensiveness and accuracy of feature extraction are effectively improved, and the adaptability and recognition ability of the vision system in complex environments are enhanced.

[0033] (4) Frequency domain features are extracted by sliding window Fourier transform, and the difference between the current fused features and the historical normal state is calculated by combining cosine similarity. A multi-level threshold discrimination mechanism is established to realize rapid response and accurate positioning of abnormal situations. At the same time, the processing strategy is automatically adjusted according to the difference level, which effectively improves the alarm accuracy and environmental adaptability of the system and ensures that the vision system can still operate stably and reliably in complex and ever-changing scenarios. Attached Figure Description

[0034] Figure 1 This is a schematic diagram illustrating the steps of the visual method for multi-scene and visible light image recognition according to the present invention;

[0035] Figure 2 This is a schematic diagram of the visual system structure based on multi-scene and visible light image recognition of the present invention. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Please see Figure 1 This invention provides a visual method based on multi-scene and visible light image recognition, comprising the following steps:

[0038] Step 1: Synchronously acquire polarization light information and light intensity information in the target scene, and construct a vector parameter matrix;

[0039] Step one includes the following:

[0040] Step 101: Deploy a CMOS array (a sensor array composed of multiple microlenses) at fixed intervals (e.g., every 50 meters) in the target scene (e.g., the top of a tunnel). Each node achieves four-way polarization separation through the microlens array (i.e., captures polarized light information in the horizontal, vertical, 45° oblique and 135° oblique directions respectively).

[0041] Step 102: Synchronously acquire the Stokes vector of each node, which contains four components: , , , The Stokes vector data is transmitted to the central processing unit in real time to construct the vector parameter matrix: ;

[0042] It should be noted that in the vector parameter matrix x,y For spatial coordinates, t This is a time series, representing the timestamp of data collection. The total light intensity is the sum of the total radiation intensity received by the sensor, including the light intensity of all polarization states. ; and For linearly polarized components, The intensity difference between linearly polarized light at 0° and 90° represents the horizontal / vertical polarization dominance. , The intensity difference between linearly polarized light at 45° and 135° is used to characterize the oblique polarization properties. , The degree of circular polarization is represented by the intensity difference between right-handed and left-handed circularly polarized light. ,in, The intensity of right-handed circularly polarized light. To represent the intensity of left-handed circularly polarized light, the four components of the Stokes vector are acquired from the four microlens directions (0°, 45°, 90°, 135°) of the CMOS array. Among these, the circularly polarized component... right-handed With left-handed The intensity was measured after separation by a combination of a quarter-wave plate and a polarizer.

[0043] Step 103: Within a 200-meter radius of the tunnel entrance, deploy a three-in-one light intensity sensor array every 10 meters to simultaneously collect absolute light intensity, light source color temperature, and light source azimuth angle. Each node integrates: a light intensity sensor with a range of 0.01~88klux, a sampling rate of 100Hz, and outputs absolute light intensity; a spectrophotometer measuring the light source color temperature (2500~6500K) with a spectral resolution of 5nm; and a magnetoresistive azimuth sensor measuring the light source azimuth angle (0~360°) with an accuracy of ±0.5°.

[0044] Step 104: Spatially and temporally align the data collected by the CMOS array and the light intensity sensor array. Spatially, the SIFT algorithm is used to extract the common feature points of the CMOS and the light intensity sensor, and coordinate mapping is established through affine transformation. Temporally, cubic spline interpolation is used to synchronize the 30fps CMOS data and the 100Hz light intensity data to a unified timestamp.

[0045] When using this method, refer to steps 101 to 104:

[0046] By strategically deploying sensor arrays at the tunnel top and entrance, multi-dimensional data such as polarized light and illumination intensity can be accurately collected. A vector parameter matrix can be constructed, and spatial and temporal alignment operations ensure data accuracy, providing a reliable foundation for subsequent processing. This effectively addresses tunnel reflection interference, comprehensively perceives environmental information, helps accurately extract target features, and improves the stability and accuracy of the overall vision system.

[0047] Step 2: Construct reflection suppression weights based on the vector parameter matrix, generate a reflection-suppressed light intensity image, and construct a dynamic nonlinear compensation surface by combining the light source color temperature and absolute illumination intensity to dynamically enhance the image;

[0048] Step two includes the following:

[0049] Step 201: Construct reflection suppression weights based on the vector parameter matrix, including polarization feature weights and non-polarization feature weights: , ,in, For polarization feature weights, For non-polarization feature weights, The average light intensity at time t is given. When the angle between the azimuth of the light source and the normal of the CMOS array exceeds 60°, a gain coefficient of 1.2 is applied to the polarization feature weight to enhance the ability to suppress lateral reflection.

[0050] Step 202: Suppress polarization interference in high-reflection regions by using polarization feature weights: , , and The linearly polarized component after reflection suppression is used to generate a reflection-suppressed image through non-polarization feature weighting: , The output vector parameter matrix represents the light intensity after reflection suppression. ;

[0051] Step 203: By absolute light intensity and the color temperature of the light source Constructing a nonlinear compensation surface ,in, α Set a baseline offset (e.g., 0.8) to ensure the baseline is correct. c Values ​​are adapted to tunnel environments (too dark or too bright scenes). β The gain coefficient (e.g., 0.8) controls the nonlinear compensation intensity, avoiding over-enhancement or under-compensation. , , sigmoid ()for sigmoid function, For light intensity deviation: , This is the reference value for light intensity (generally taken as 20). Due to color temperature deviation, , The nonlinear compensation surface parameters, including illuminance and light source color temperature, are updated every 200ms using cubic spline interpolation, with a reference color temperature value (typically 4500). α (Baseline offset) and β (Gain coefficient), determined through calibration experiments: under reference illumination (20klux) and color temperature (4500K), , It can balance the compensation effect of overly dark and overly bright scenes;

[0052] Step 204: Dynamically enhance the image after reflection suppression, including:

[0053] Mapping pixel values ​​to the range [0, 1], for example, using 16-bit pixel values: , The normalized light intensity;

[0054] Pixel-level transformation based on nonlinear compensation surface: , The transformed light intensity;

[0055] Apply a compensation coefficient to the shortwave band (430~500nm) It suppresses color shift of cold light source. When the compensation coefficient exceeds the compensation threshold (e.g., 1.15), it automatically extends the lower limit of shortwave band to 410nm. The compensation threshold is based on historical data statistics. When the compensation coefficient exceeds 1.15, after extending the lower limit of shortwave band to 410nm, the accuracy of license plate recognition is significantly improved.

[0056] Step 205: Calculate the contrast of the enhanced image: , , These represent the maximum and minimum light intensities, respectively. , These are the maximum and minimum normalized illumination intensities of all pixels in the current frame image, obtained through global statistics. When the contrast is less than the contrast threshold (e.g., 0.2), a recalibration of the nonlinear compensation surface parameters is triggered, including updating the nonlinear compensation surface parameters using cubic spline interpolation in step 203 and readjusting the parameters. α (Baseline offset) and β (Gain coefficient); The output image is enhanced. The contrast threshold is based on historical data statistics. When the contrast is below 0.2, the distinction between the target (license plate, markings) and the background decreases significantly.

[0057] When using this method, refer to steps 201 to 205:

[0058] Polarization difference de-reflection effectively eliminates interference in high-reflection areas, improving image clarity. Illuminance and color temperature compensation construct a dynamic nonlinear compensation surface, ensuring balanced image contrast, suppressing color shift, and automatically recalibrating surface parameters based on the corrected image contrast, resulting in better pre-processed image quality. This provides a reliable foundation for subsequent feature processing and recognition, enhancing the system's adaptability in complex tunnel environments.

[0059] Step 3: Dynamically adjust the visible light band range according to the color temperature of the light source, decompose the enhanced image into three-band feature maps, and extract polarization features and spectral features respectively through a cascaded neural network architecture, and fuse them to generate a fused feature;

[0060] Step three includes the following:

[0061] Step 301: Obtain real-time light source color temperature data. When the light source color temperature is less than 4000K, the long-wave band is extended to 650nm to enhance the infrared radiation characteristics of the warm light source, the short-wave band is compressed to 470~550nm to suppress cold light source interference, and the medium-wave band is fixed at 550~600nm to maintain the reflective recognition capability of the markings.

[0062] When the color temperature of the light source is greater than 5000K, the short-wave band extends down to 430nm to enhance the extraction of blue and violet light features from cold light sources. The long-wave band is limited to 600~630nm to avoid overlapping with the red light band of vehicle lights. The mid-wave band is adjusted to 500~580nm to optimize the sensitivity of license plate recognition.

[0063] Step 302: Based on the three characteristic bands of visible light, decompose the enhanced image to obtain feature maps of three bands (short wave, medium wave, and long wave), i.e., three-band feature maps: short wave band: 430~500nm / 470~550nm; medium wave band: 500~580nm / 550~600nm; long wave band: 600~630nm / 630~650nm.

[0064] Step 303: Construct a cascaded neural network architecture, including polarization and spectral branches, and extract polarization and spectral features;

[0065] Specifically, the polarization branch: The linear polarization component is extracted from the input vector parameter matrix after reflection suppression, and then processed by a deformable convolutional network.

[0066] First layer: Input linear polarization component and splicing tensor Perform a 3×3 deformable convolution: ,in, This represents the first layer of features. GeLU () represents the GeLU activation function. This represents a 3×3 deformable convolution kernel, with settings for the number of channels (e.g., 64), offset learning rate (e.g., 0.1), and output resolution (e.g., 1024×1024); Second layer: Input Perform a 5×5 deformable convolution: ,in, This represents the second layer of features. LayerNorm () indicates the layer normalization operation. This represents a 5×5 deformable convolution kernel with a set number of channels (e.g., 128); after channel attention weighting, the output polarization features are shown. The SE module is used to perform global average pooling on the second-layer features to generate channel statistics. Channel weights are generated using two fully connected layers and the sigmoid function. ,in, , These are the parameters for the fully connected layer. , These are learnable parameters, and their final values ​​are automatically optimized during training using the backpropagation algorithm. The channel weights are multiplied by the second-layer features to output the polarization features. ;

[0067] Specifically, the spectral branch: Inputting a three-band feature map, the image data of the three bands are stacked along the channel dimension to generate a three-channel tensor. ,in, This represents a shortwave band characteristic map. This represents a mid-wave band characteristic map. To represent the long-wavelength band feature map, a residual network (such as ResNet34) and an attention mechanism module are used to extract spectral features. The residual network is used as the backbone network. Stage 1: 7×7 convolutions and activation functions (such as ReLU) are performed, followed by downsampling, for example, to 512×512. Stages 2 through 4: Features are extracted sequentially through residual blocks and an attention mechanism module (such as an SE module), with the number of channels set, for example, 128 / 256 / 512 respectively. The final output is the spectral feature. ;

[0068] Step 304: After upsampling the polarization features to the same resolution as the spectral features using bilinear interpolation, the polarization and spectral features are concatenated along the channel dimension to output the fused features. ,in, This indicates a splicing operation. R Represents the set of real numbers. H Indicates altitude, W Indicates width, C Indicates the number of channels;

[0069] When using this method, refer to steps 301 to 304:

[0070] By dynamically adjusting the feature band range according to the color temperature of the light source to adapt to different lighting conditions, and by using a cascaded neural network architecture to extract polarization features and spectral features separately and then fusing them, the comprehensiveness and accuracy of feature extraction are effectively improved, enhancing the adaptability and recognition ability of the vision system in complex environments.

[0071] Step 4: Perform a sliding window Fourier transform on the fused features to extract frequency domain features, calculate the comprehensive difference between the current fused features and the historical benchmark, and trigger anomaly alarms or dynamically adjust the processing strategy based on a multi-level threshold discrimination mechanism.

[0072] Step four includes the following:

[0073] Step 401: Perform a sliding window Fourier transform on the fused features, with a window size of 256×256 pixels and a stride of 64 pixels, to extract frequency domain features. ,in, F For Fast Fourier Transform, N This represents the number of windows generated along the height direction. , M This represents the number of windows generated along the width direction. , , They represent from the first x、y Starting with row 1, take 256 consecutive rows and perform sliding window Fourier transform on the three-band feature maps (shortwave, mediumwave, and longwave) to generate frequency domain features. , , ;

[0074] Step 402: Calculate the dissimilarity of the current fused features: ,in, To calculate the difference in the current three-band features based on the historical baseline value of the fused features under normal conditions: ,in, For the first i Band (shortwave) S , medium wave M Long waves L Based on historical baseline values ​​under normal conditions, a comprehensive difference score is generated by weighting the difference scores of the current fused features and the current three-band features: ,in, m The difference weight (default 0.8). Comparing different m Choose the value that optimizes alarm accuracy versus false alarm rate to achieve the best accuracy. m value;

[0075] Step 403: Establish a multi-level threshold discrimination mechanism:

[0076] Level 1 discrimination: When the overall difference is greater than or equal to D1 (e.g., 0.7), immediately activate the audible and visual alarm and red light warning, enforce speed limit, for example, limit the speed to 30km / h, and mark the coordinates of the abnormal area. Based on the light intensity data, automatically increase the lighting intensity of the abnormal area to 1.5 times the current light intensity value to suppress local reflection interference.

[0077] Secondary discrimination: When the overall difference is greater than D2 (e.g., 0.4) and less than D1 (e.g., 0.7), the band range is dynamically adjusted according to real-time color temperature data. When the light source color temperature is greater than 5000K, the lower limit of the short-wave band is extended to 410nm to enhance the extraction of blue-violet light features; when the light source color temperature is less than 4000K, the upper limit of the long-wave band is extended to 650nm to enhance the identification of infrared radiation features.

[0078] By comparing the alarm accuracy and false alarm rate under different thresholds through cross-validation, the thresholds that achieve the optimal accuracy are selected as D1 and D2.

[0079] Level 3 discrimination: When the overall difference is less than or equal to D2 (e.g., 0.4), the historical benchmark value is updated using a moving average strategy. , l The moving average coefficient, The default value is 0.95. Through calibration experiments, a higher value was selected. l (Close to 1) to ensure the smoothness of benchmark updates and avoid frequent benchmark changes due to minor environmental fluctuations, while also using (1− l New frequency domain features are gradually incorporated to achieve dynamic adaptation of the reference and verify the CMOS polarization state separation accuracy. If the deviation exceeds the deviation threshold (e.g., 5%), the reflection suppression weight is updated after calibration. The deviation threshold is set according to the requirements of CMOS polarization state separation accuracy and is set based on actual needs.

[0080] When using this method, please refer to the content of steps 401 to 403:

[0081] Frequency domain features are extracted by sliding window Fourier transform, and the difference between the current fused features and the historical normal state is calculated by combining cosine similarity. A multi-level threshold discrimination mechanism is established to achieve rapid response and accurate positioning of abnormal situations. At the same time, the processing strategy is automatically adjusted according to the difference level, which effectively improves the alarm accuracy and environmental adaptability of the system and ensures that the vision system can still operate stably and reliably in complex and ever-changing scenarios.

[0082] Please see Figure 2 The present invention also provides a vision system based on multi-scene and visible light image recognition, comprising:

[0083] The data acquisition module synchronously acquires polarization light information and light intensity information in the target scene and constructs a vector parameter matrix;

[0084] The image preprocessing module constructs reflection suppression weights based on the vector parameter matrix, generates a light intensity image after reflection suppression, and constructs a dynamic nonlinear compensation surface by combining the light source color temperature and absolute illumination intensity to dynamically enhance the image.

[0085] The feature extraction and fusion module dynamically adjusts the visible light band range according to the color temperature of the light source, decomposes the enhanced image into three-band feature maps, and extracts polarization features and spectral features respectively through a cascaded neural network architecture, and then fuses them to generate fused features.

[0086] The anomaly detection module performs a sliding window Fourier transform on the fused features to extract frequency domain features, calculates the comprehensive difference between the current fused features and the historical benchmark, and triggers anomaly alarms or dynamically adjusts the processing strategy based on a multi-level threshold discrimination mechanism.

[0087] In the application, the various formulas mentioned are all calculated by removing dimensions and taking their numerical values. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The coefficients in the formulas are set by those skilled in the art according to the actual situation.

[0088] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, and combinations thereof. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0089] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0090] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A visual method based on multi-scene and visible light image recognition, characterized in that: include: Simultaneously acquire polarization light information and light intensity information in the target scene to construct a vector parameter matrix; Based on the vector parameter matrix, a reflection suppression weight is constructed to generate a light intensity image after reflection suppression. A dynamic nonlinear compensation surface is constructed by combining the light source color temperature and absolute light intensity to dynamically enhance the image. The visible light band range is dynamically adjusted according to the color temperature of the light source. The enhanced image is decomposed into three-band feature maps, and polarization features and spectral features are extracted separately through a cascaded neural network architecture and fused to generate a fused feature. The cascaded neural network architecture includes a polarization branch and a spectral branch. The polarization branch takes into account the linearly polarized components after reflection suppression. and Polarization features are extracted through deformable convolution and channel attention mechanism; spectral features are extracted from the three-band feature map input by spectral branch through residual network and attention mechanism. The polarization features are upsampled to the same resolution as the spectral features using bilinear interpolation. The polarization and spectral features are then concatenated along the channel dimension to output the fused features. ,in, Spectral characteristics, Polarization characteristics R Let be the set of real numbers. H For height, W For width, C Number of channels; Sliding window Fourier transforms are performed on the fused features and the three-band feature maps respectively to extract frequency domain features, calculate the comprehensive difference degree, and trigger anomaly alarms or dynamically adjust the processing strategy based on a multi-level threshold discrimination mechanism.

2. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: The acquisition of polarized light information is achieved through a CMOS array deployed in the target scene. The CMOS array consists of multiple microlenses, and each node captures horizontal, vertical, 45° oblique, and 135° oblique polarized light information respectively through four-way polarization state separation. The acquisition of light intensity information is achieved through a three-in-one light intensity sensor array deployed in the target scene. The light intensity sensor array simultaneously acquires absolute light intensity, light source color temperature, and light source azimuth angle.

3. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: The reflection suppression weights are constructed based on the vector parameter matrix. The reflection suppression weights include calculating polarization feature weights and non-polarization feature weights. The polarization feature weights are used to suppress polarization interference in high-reflection areas, and the non-polarization feature weights are used to generate a reflection suppression image. Finally, the reflection suppression vector parameter matrix is ​​output.

4. The visual method based on multi-scene and visible light image recognition according to claim 3, characterized in that: A nonlinear compensation surface is constructed using absolute illumination intensity and light source color temperature to map pixel values ​​to the range [0, 1]. Pixel-level transformation is performed based on the nonlinear compensation surface, and a compensation coefficient is applied to the shortwave band to suppress color shift of cold light source. The image is then dynamically enhanced, and the contrast of the enhanced image is calculated. When the contrast is less than the contrast threshold, the parameters of the nonlinear compensation surface are recalibrated.

5. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: When the color temperature of the light source is less than 4000K, the long-wave band is extended to 650nm, the short-wave band is compressed to 470~550nm, and the medium-wave band is fixed at 550~600nm; when the color temperature of the light source is greater than 5000K, the short-wave band is lowered to 430nm, the long-wave band is limited to 600~630nm, and the medium-wave band is adjusted to 500~580nm. Based on the three characteristic bands of visible light, the enhanced image is decomposed to obtain characteristic maps of short-wave, medium-wave, and long-wave, i.e., three-band characteristic maps.

6. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: Calculate the difference of fused features: ,in, This represents the historical baseline value of the fusion feature under normal conditions. Frequency domain features for fusion features; Calculate the difference in characteristics across the three bands: ,in, For the first i Frequency domain characteristics of the band, For the first i Historical baseline values ​​for the band under normal conditions S It is a shortwave. M For medium wave, L For long waves, the difference between the fused features and the difference between the three-band features are weighted to generate a comprehensive difference.

7. A visual method based on multi-scene and visible light image recognition according to claim 6, characterized in that: The multi-level threshold discrimination mechanism includes: Level 1 discrimination: when the comprehensive difference is greater than or equal to D1, an audible and visual alarm is triggered, and the lighting intensity of the abnormal area is increased; Level 2 discrimination: when the comprehensive difference is greater than D2 and less than D1, the band range is dynamically adjusted according to real-time color temperature data; Level 3 discrimination: when the comprehensive difference is less than or equal to D2, the historical benchmark value is updated using a moving average strategy.

8. A vision system based on multi-scene and visible light image recognition, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The data acquisition module synchronously acquires polarization light information and light intensity information in the target scene and constructs a vector parameter matrix; The image preprocessing module constructs reflection suppression weights based on the vector parameter matrix, generates a light intensity image after reflection suppression, and constructs a dynamic nonlinear compensation surface by combining the light source color temperature and absolute illumination intensity to dynamically enhance the image. The feature extraction and fusion module dynamically adjusts the visible light band range according to the color temperature of the light source, decomposes the enhanced image into three-band feature maps, and extracts polarization features and spectral features respectively through a cascaded neural network architecture, and then fuses them to generate fused features. The cascaded neural network architecture includes a polarization branch and a spectral branch. The polarization branch takes into account the linearly polarized components after reflection suppression. and Polarization features are extracted through deformable convolution and channel attention mechanism; spectral features are extracted from the three-band feature map input by spectral branch through residual network and attention mechanism. The polarization features are upsampled to the same resolution as the spectral features using bilinear interpolation. The polarization and spectral features are then concatenated along the channel dimension to output the fused features. ,in, Spectral characteristics, Polarization characteristics R Let be the set of real numbers. H For height, W For width, C Number of channels; The anomaly detection module performs a sliding window Fourier transform on the fused features to extract frequency domain features, calculates the comprehensive difference between the current fused features and the historical benchmark, and triggers anomaly alarms or dynamically adjusts the processing strategy based on a multi-level threshold discrimination mechanism.

Citation Information

Patent Citations

  • Image recognition system and method based on deep learning

    CN120375075A