Visual system and method based on multiple scenes and visible light image recognition
By introducing polarized light information and dynamically adjusting the visible light band, combined with cascade neural network, the problem of reflective interference in the tunnel environment is solved, the recognition accuracy and stability of the visual system are improved, and adaptability and rapid response to complex environments are achieved.
Patent Information
- Application Number
- CN202511066166.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-31
AI Technical Summary
Traditional vision systems are difficult to suppress reflection interference in high-reflection environments such as tunnels, resulting in a decline in image quality, difficulty in extracting target features, and cannot be automatically adjusted in scenarios where lighting conditions change greatly, affecting the recognition accuracy and robustness.
By introducing polarized light information, a cascading neural network architecture is built, the visible band range is dynamically adjusted, and the multi-level threshold discrimination mechanism is combined to improve the adaptability and recognition capabilities of the visual system.
Effectively eliminate interference in high-reflection areas, improve image clarity and recognition accuracy, ensure the system operates stably in complex environments, and achieve rapid response and precise positioning of abnormal situations.
Smart Images

Figure CN120563992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a visual system and method based on multi-scene and visible light image recognition. Background Art
[0002] In the field of image recognition technology, traditional vision systems mainly rely on single-dimensional visible light image data for target recognition and scene analysis. These systems generate two-dimensional images by capturing the light intensity distribution in the environment, and use computer vision algorithms to detect, classify and track targets in the image.
[0003] However, in highly reflective environments such as tunnels, it is difficult to effectively suppress reflective interference, resulting in a decline in image quality and difficulty in extracting target features, which in turn affects the overall recognition accuracy. In scenarios with large changes in lighting conditions, such as day and night or sudden changes in weather, existing systems are often unable to automatically adjust to the new lighting environment, resulting in unstable recognition performance.
[0004] In addition, feature extraction is often limited to the visible light band, ignoring information in other dimensions such as polarized light. This makes it difficult to fully capture target features in complex environments, reducing the accuracy and robustness of recognition. Especially in low light or strong light interference conditions, the feature extraction method of a single band is more easily affected by environmental factors. Summary of the Invention
[0005] (1) Technical problems solved In response to the shortcomings of the existing technology, the present invention provides a visual system and method based on multi-scene and visible light image recognition. By introducing polarized light information, dynamically adjusting the visible light band range, constructing a cascaded neural network architecture, and establishing a multi-level threshold discrimination mechanism, the adaptability and recognition ability of the visual system in complex environments are improved.
[0006] (2) Technical solution To achieve the above objectives, the present invention is implemented through the following technical solutions: a visual method based on multi-scene and visible light image recognition, comprising: Synchronously collect polarization information and light intensity information in the target scene to construct a vector parameter matrix; Based on the vector parameter matrix, the reflection suppression weight is constructed to generate the light intensity image after reflection suppression. The dynamic nonlinear compensation surface is constructed by combining the light source color temperature and absolute light intensity to dynamically enhance the image. Dynamically adjust the visible light band range according to the color temperature of the light source, decompose the enhanced image into a three-band feature map, and use a cascaded neural network architecture to extract polarization features and spectral features respectively, and fuse them to generate a fusion feature; The fusion features and three-band feature maps are subjected to sliding window Fourier transform respectively to extract frequency domain features and calculate the comprehensive difference. Based on the multi-level threshold discrimination mechanism, abnormal alarms are triggered or the processing strategy is dynamically adjusted.
[0007] Furthermore, the collection of polarization light information is achieved through a CMOS array deployed in the target scene. The CMOS array is composed of multiple microlenses, and each node captures horizontal, vertical, 45° oblique and 135° oblique polarization light information through four-way polarization state separation; the collection of light intensity information is achieved through a three-in-one light intensity sensor array deployed in the target scene. The light intensity sensor array synchronously collects absolute light intensity, light source color temperature and light source azimuth.
[0008] Furthermore, a reflection suppression weight is constructed based on the vector parameter matrix. The reflection suppression weight includes calculating the polarization feature weight and the non-polarization feature weight. The polarization interference in the high-reflection area is suppressed by the polarization feature weight, and the reflection suppression image is generated by the non-polarization feature weight, and the vector parameter matrix after reflection suppression is output.
[0009] Furthermore, a nonlinear compensation surface is constructed based on the absolute light intensity and light source color temperature. The pixel values are mapped to the range of [0, 1]. Pixel-level transformation is performed according to the nonlinear compensation surface. A compensation coefficient is applied to the shortwave band to suppress the color cast of the cold light source. The image is dynamically enhanced and the contrast of the enhanced image is calculated. When the contrast is less than the contrast threshold, the nonlinear compensation surface parameter recalibration is triggered.
[0010] Furthermore, when the color temperature of the light source is less than 4000K, the longwave band is extended to 650nm, the shortwave band is compressed to 470~550nm, and the mediumwave band is fixed at 550~600nm; when the color temperature of the light source is greater than 5000K, the shortwave band is lowered to 430nm, the longwave band is limited to 600~630nm, and the mediumwave band is adjusted to 500~580nm; According to the three characteristic bands of visible light, the enhanced image is decomposed to obtain the characteristic maps of short wave, medium wave and long wave, that is, the three-band characteristic map.
[0011] Furthermore, the cascaded neural network architecture includes a polarization branch and a spectral branch, and the polarization branch inputs the linear polarization component after reflection suppression. and , polarization features are extracted through deformable convolution and channel attention mechanism; the spectral branch inputs the three-band feature map, and the spectral features are extracted through the residual network and attention mechanism.
[0012] Furthermore, after the polarization features are upsampled to the same resolution as the spectral features through bilinear interpolation, the polarization features and spectral features are spliced along the channel dimension to output the fused features: ,in, is the spectral characteristic, is the polarization characteristic, Represents a splicing operation, R is the set of real numbers, H is the height, W is the width, C is the number of channels.
[0013] Furthermore, the difference of the fusion features is calculated: ,in, is the historical benchmark value of the fusion feature under normal conditions, is the frequency domain feature of the fusion feature; Calculate the difference between the three-band features: ,in, For the i The frequency domain characteristics of the band, For the i The historical benchmark value of the band under normal conditions, S For shortwave, M For medium wave, L is the long wave; the difference of the fusion feature and the difference of the three-band feature are weighted to generate a comprehensive difference.
[0014] Furthermore, the multi-level threshold discrimination mechanism includes: first-level discrimination: when the comprehensive difference is greater than or equal to D1, an audible and visual alarm is triggered, and the lighting intensity in the abnormal area is enhanced; second-level discrimination: when the comprehensive difference is greater than D2 and less than D1, the band range is dynamically adjusted according to the real-time color temperature data; third-level discrimination: when the comprehensive difference is less than or equal to D2, a sliding average strategy is used to update the historical benchmark value.
[0015] A visual system based on multi-scene and visible light image recognition, comprising: The data acquisition module synchronously collects polarization information and light intensity information in the target scene and constructs a vector parameter matrix; The image preprocessing module constructs reflection suppression weights based on the vector parameter matrix, generates a light intensity image after reflection suppression, and constructs a dynamic nonlinear compensation surface based on the light source color temperature and absolute light intensity to dynamically enhance the image; The feature extraction and fusion module dynamically adjusts the visible light band range according to the color temperature of the light source, decomposes the enhanced image into a three-band feature map, and extracts polarization features and spectral features separately through a cascaded neural network architecture, and fuses them to generate a fusion feature; The anomaly detection module performs a sliding window Fourier transform on the fused features, extracts frequency domain features, calculates the comprehensive difference between the current fused features and the historical benchmark, and triggers an anomaly alarm or dynamically adjusts the processing strategy based on a multi-level threshold discrimination mechanism.
[0016] (3) Beneficial effects The present invention provides a visual system and method based on multi-scene and visible light image recognition, which has the following beneficial effects: (1) By rationally deploying sensor arrays at the top and entrance of the tunnel, multi-dimensional data such as polarized light and light intensity can be accurately collected, a vector parameter matrix can be constructed, and spatial and temporal alignment operations can be performed to ensure data accuracy, providing a reliable basis for subsequent processing, effectively dealing with tunnel reflection interference, and comprehensively perceiving environmental information, which helps to accurately extract target features and improve the stability and accuracy of the overall visual system.
[0017] (2) Polarization differential de-reflection is used to effectively eliminate interference from high-reflection areas and improve image clarity. Illumination and color temperature compensation are used to construct a dynamic nonlinear compensation surface to ensure balanced image contrast and suppress color deviation. The surface parameters can be automatically recalibrated according to the contrast of the corrected image, making the quality of the pre-processed image better, providing a reliable basis for subsequent feature processing and recognition, and enhancing the adaptability of the system in complex tunnel environments.
[0018] (3) By dynamically adjusting the characteristic band range according to the color temperature of the light source to adapt to different lighting conditions, and using a cascaded neural network architecture to extract polarization features and spectral features respectively and then fuse them, the comprehensiveness and accuracy of feature extraction are effectively improved, and the adaptability and recognition ability of the visual system in complex environments are enhanced.
[0019] (4) Frequency domain features are extracted through sliding window Fourier transform, and the difference between the current fusion features and the historical normal state is calculated by combining cosine similarity. A multi-level threshold discrimination mechanism is established to achieve rapid response and accurate positioning of abnormal situations. At the same time, the processing strategy is automatically adjusted according to the difference level, which effectively improves the alarm accuracy and environmental adaptability of the system, ensuring that the visual system can still operate stably and reliably in complex and changing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Schematic diagram of the steps of the visual method based on multi-scene and visible light image recognition of the present invention; Figure 2 This is a schematic diagram of the visual system structure based on multi-scene and visible light image recognition of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0022] See also Figure 1 The present invention provides a visual method based on multi-scene and visible light image recognition, comprising the following steps: Step 1: Synchronously collect polarization information and light intensity information in the target scene and construct a vector parameter matrix; The step 1 includes the following contents: Step 101: Deploy a CMOS array (a sensor array consisting of multiple microlenses) at a fixed interval (e.g., every 50 meters) in the target scene (e.g., the top of a tunnel). Each node uses the microlens array to achieve four-way polarization state separation (i.e., capturing polarization information of horizontal, vertical, 45° oblique, and 135° oblique light). Step 102: Synchronously collect the Stokes vector of each node, which includes four components: 、 、 、 , the Stokes vector data is transmitted to the central processing unit in real time to construct the vector parameter matrix: ; It should be noted that the vector parameter matrix x,y is the spatial coordinate, t is a time series, indicating the timestamp of data collection, is the total light intensity, the total radiation intensity received by the sensor, including the sum of the light intensities of all polarization states: ; and is the linear polarization component, The intensity difference between linearly polarized light at 0° and 90°, representing the horizontal / vertical polarization dominance: , The intensity difference between linearly polarized light at 45° and 135° characterizes the oblique polarization characteristics: , Represents the degree of circular polarization, which is the intensity difference between right-handed and left-handed circularly polarized light: ,in, is the intensity of right-handed circularly polarized light, is the intensity of left-handed circularly polarized light. The four components of the Stokes vector are collected by the four microlens directions (0°, 45°, 90°, and 135°) of the CMOS array, respectively. right-handed With left The intensity is measured after separation by a quarter-wave plate and polarizer combination; Step 103: Within 200 meters of the tunnel entrance, a three-in-one light intensity sensor array is deployed every 10 meters to simultaneously collect absolute light intensity, light source color temperature, and light source azimuth. Each node integrates: a light intensity sensor with a range of 0.01 to 88 klux, a sampling rate of 100 Hz, and outputs absolute light intensity; a spectrocolorimeter for measuring light source color temperature (2500 to 6500 K) with a spectral resolution of 5 nm; and a magnetoresistive azimuth sensor for measuring light source azimuth (0 to 360°) with an accuracy of ±0.5°. Step 104: The data collected by the CMOS array and the light intensity sensor array are spatially and temporally aligned. The SIFT algorithm is used for spatial alignment to extract the common feature points of the CMOS and light intensity sensors, and a coordinate mapping is established through affine transformation. The cubic spline interpolation method is used for temporal alignment to synchronize the 30fps CMOS data and the 100Hz light intensity data to a unified timestamp.
[0023] When using, combine the contents of steps 101 to 104: By rationally deploying sensor arrays at the top and entrance of the tunnel, it is possible to accurately collect multi-dimensional data such as polarized light and light intensity, construct a vector parameter matrix, and perform spatial and temporal alignment operations to ensure data accuracy, providing a reliable foundation for subsequent processing. This effectively responds to tunnel reflection interference, comprehensively perceives environmental information, helps to accurately extract target features, and improves the stability and accuracy of the overall visual system.
[0024] Step 2: Construct reflection suppression weights based on the vector parameter matrix to generate a light intensity image after reflection suppression. Combined with the light source color temperature and absolute light intensity, a dynamic nonlinear compensation surface is constructed to dynamically enhance the image. The second step includes the following contents: Step 201: Construct reflection suppression weights based on the vector parameter matrix, including polarization feature weights and non-polarization feature weights: , ,in, is the polarization feature weight, is the non-polarization feature weight, is the average light intensity at time t. When the angle between the light source azimuth and the CMOS array normal exceeds 60°, a 1.2x gain coefficient is applied to the polarization feature weight to enhance the lateral reflection suppression capability. Step 202: Suppress polarization interference in high-reflection areas using polarization feature weights: , , and is the linear polarization component after reflection suppression, and the reflection suppression image is generated by the non-polarization feature weight: , is the light intensity after reflection suppression, and outputs the vector parameter matrix after reflection suppression: ; Step 203: Absolute light intensity and light source color temperature , construct nonlinear compensation surface ,in, α is the baseline offset (such as 0.8), ensuring that the baseline c The value is adapted to the tunnel environment (too dark or too bright scene), β is the gain coefficient (such as 0.8), which controls the intensity of nonlinear compensation to avoid over-enhancement or under-compensation. , , sigmoid ()for sigmoid function, is the light intensity deviation: , is the light intensity reference value (usually 20), is the color temperature deviation, , The color temperature reference value (usually 4500) is used. The cubic spline interpolation is used every 200ms to update the nonlinear compensation surface parameters, including light intensity and light source color temperature. α (baseline offset) and β (Gain coefficient), determined by calibration experiments: Under reference light (20klux) and color temperature (4500K), 、 Can balance the compensation effect of too dark and too bright scenes; Step 204: Dynamically enhance the image after reflection suppression, including: Map pixel values to the range [0, 1], taking 16-bit pixel values as an example: , is the normalized light intensity; Perform pixel-level transformation based on nonlinear compensation surface: , is the light intensity after transformation; Apply compensation coefficient to shortwave band (430~500nm) , suppress the color cast of cold light sources, and automatically expand the shortwave band lower limit to 410nm when the compensation coefficient exceeds the compensation threshold (such as 1.15). The compensation threshold is based on historical data statistics. When the compensation coefficient exceeds 1.15, the shortwave band lower limit is expanded to 410nm, and the license plate recognition accuracy is significantly improved; Step 205: Calculate the contrast of the enhanced image: , 、 are the maximum and minimum light intensities, 、 are the maximum and minimum normalized illumination intensities of all pixels in the current frame image, respectively, obtained through global statistics; when the contrast is less than the contrast threshold (such as 0.2), the nonlinear compensation surface parameter recalibration is triggered, including the use of cubic spline interpolation to update the nonlinear compensation surface parameters in step 203 and readjustment of the parameters α (baseline offset) and β (Gain coefficient); output enhanced image, contrast threshold is based on historical data statistics, when the contrast is lower than 0.2, the distinction between the target (license plate, marking) and the background is significantly reduced; When using, combine the contents of steps 201 to 205: Polarization differential de-reflection effectively eliminates interference from high-reflection areas and improves image clarity. Illumination and color temperature compensation constructs a dynamic nonlinear compensation surface to ensure balanced image contrast and suppress color cast. The surface parameters can be automatically recalibrated based on the contrast of the corrected image, improving the quality of the pre-processed image, providing a reliable foundation for subsequent feature processing and recognition, and enhancing the system's adaptability in complex tunnel environments.
[0025] Step 3: Dynamically adjust the visible light band range according to the color temperature of the light source, decompose the enhanced image into a three-band feature map, and use a cascaded neural network architecture to extract polarization features and spectral features respectively, and fuse them to generate a fusion feature; The step three includes the following contents: Step 301: Acquire real-time light source color temperature data. When the light source color temperature is less than 4000K, the long-wave band is extended to 650nm to enhance the infrared radiation characteristics of the warm light source, the short-wave band is compressed to 470-550nm to suppress interference from cold light sources, and the medium-wave band is fixed at 550-600nm to maintain the ability to identify road marking reflective surfaces. When the color temperature of the light source is greater than 5000K, the shortwave band is lowered to 430nm to enhance the extraction of blue-violet light features of the cold light source. The longwave band is limited to 600-630nm to avoid overlap with the red light band of car lights. The mediumwave band is adjusted to 500-580nm to optimize the sensitivity of license plate recognition. Step 302: Decompose the enhanced image according to the three characteristic bands of visible light to obtain characteristic maps of the three bands (shortwave, medium wave, and long wave), i.e., three-band characteristic maps: shortwave band: 430-500nm / 470-550nm; medium wave band: 500-580nm / 550-600nm; longwave band: 600-630nm / 630-650nm; Step 303: constructing a cascade neural network architecture, including a polarization branch and a spectral branch, to extract polarization features and spectral features; Specifically, the polarization branch: inputs the vector parameter matrix after reflection suppression, extracts the linear polarization component, and processes it using a deformable convolutional network: First layer: input linear polarization component and The concatenated tensor , performing a 3×3 deformable convolution: ,in, represents the first layer features, GeLU () represents the GeLU activation function, Represents a 3×3 deformable convolution kernel, sets the number of channels (such as 64), offset learning rate (such as 0.1) and output resolution (such as 1024×1024); the second layer: input , performing a 5×5 deformable convolution: ,in, represents the second layer features, LayerNorm () represents the layer normalization operation, Represents a 5×5 deformable convolution kernel, sets the number of channels (such as 128); after channel attention weighting, outputs polarization features : Use the SE module to perform global average pooling on the second layer features to generate channel statistics: , generate channel weights through two layers of full connection and sigmoid function ,in, 、 is the fully connected layer parameter, 、 It is a learnable parameter whose final value is automatically optimized during the training process through the back propagation algorithm. The channel weight is multiplied by the second layer feature to output the polarization feature. ; Specifically, the spectral branch: inputs a three-band feature map, stacks the image data of the three bands along the channel dimension, and generates a three-channel tensor: ,in, Represents the shortwave band characteristic diagram, Represents the medium wave band characteristic diagram, Represent the long-wave band feature map, and use the residual network (such as ResNet34) and attention mechanism module to extract spectral features: use the residual network as the backbone network, stage one: perform 7×7 convolution and activation function (such as ReLU function) activation, and downsample, for example, downsample to 512×512, stage two to stage four: extract features through the residual block and attention mechanism module (such as SE module) in turn, set the number of channels, for example, the number of channels is set to 128 / 256 / 512 respectively, and finally output the spectral features ; Step 304: After the polarization features are upsampled to the same resolution as the spectral features through bilinear interpolation, the polarization features and spectral features are concatenated along the channel dimension to output the fused features: ,in, Represents a splicing operation, R represents the set of real numbers, H Indicates height, W Indicates width, C Indicates the number of channels; When using, combine the contents of steps 301 to 304: By dynamically adjusting the characteristic band range according to the color temperature of the light source to adapt to different lighting conditions, and using a cascaded neural network architecture to extract polarization features and spectral features separately and then fuse them, the comprehensiveness and accuracy of feature extraction are effectively improved, and the adaptability and recognition ability of the visual system in complex environments are enhanced.
[0026] Step 4: Perform a sliding window Fourier transform on the fusion features, extract the frequency domain features, calculate the comprehensive difference between the current fusion features and the historical benchmark, and trigger an abnormal alarm or dynamically adjust the processing strategy based on the multi-level threshold discrimination mechanism.
[0027] The fourth step includes the following contents: Step 401: Perform sliding window Fourier transform on the fusion features with a window size of 256×256 pixels and a step size of 64 pixels to extract frequency domain features. ,in, F is the fast Fourier transform, N is the number of windows generated along the height direction, , M is the number of windows generated along the width direction, , 、 Respectively represent x、y Starting from the first line, take 256 lines in a row and perform sliding window Fourier transform on the three-band feature maps (shortwave, medium wave, long wave) to generate frequency domain features. , , ; Step 402: Calculate the difference of the current fusion feature: ,in, To integrate the historical baseline values of the features under normal conditions, calculate the difference between the current three-band features: ,in, For the i Band (shortwave S , medium wave M , long wave L) Under normal conditions, the historical benchmark value is used to weight the difference between the current fusion feature and the current three-band feature to generate a comprehensive difference: ,in, m is the difference weight (default 0.8), , by cross-validation comparing different m The alarm accuracy and false alarm rate under the value, select the one that makes the accuracy the best m value; Step 403: Establish a multi-level threshold discrimination mechanism: Level 1 discrimination: When the comprehensive difference is greater than or equal to D1 (e.g. 0.7), the system immediately activates an audible and visual alarm and a red light warning, imposes a speed limit (e.g., to 30 km / h), and marks the coordinates of the abnormal area. Based on the light intensity data, the system automatically increases the lighting intensity in the abnormal area to 1.5 times the current value to suppress local reflective interference. Secondary discrimination: When the comprehensive difference is greater than D2 (such as 0.4) and less than D1 (such as 0.7), the band range is dynamically adjusted according to the real-time color temperature data. When the light source color temperature is greater than 5000K, the lower limit of the shortwave band is extended to 410nm to enhance the extraction of blue-violet light features; when the light source color temperature is less than 4000K, the upper limit of the longwave band is extended to 650nm to enhance the recognition of infrared radiation features. By comparing the alarm accuracy and false alarm rate under different thresholds through cross-validation, the thresholds with the best accuracy are selected as D1 and D2; Level 3 discrimination: When the comprehensive difference is less than or equal to D2 (such as 0.4), the sliding average strategy is used to update the historical benchmark value: , l is the sliding average coefficient, The default value is 0.95, which is verified by calibration experiments. l (close to 1) to ensure the smoothness of the benchmark value update and avoid frequent changes in the benchmark due to minor environmental fluctuations. At the same time, (1− l ) Gradually incorporate new frequency domain features to achieve dynamic adaptation of the benchmark and verify the CMOS polarization state separation accuracy. If the deviation exceeds a deviation threshold (e.g., 5%), the reflection suppression weight update is triggered after calibration. The deviation threshold is set based on the requirements of the CMOS polarization state separation accuracy and actual needs. When using, combine the contents of step 401 to step 403: Frequency domain features are extracted through sliding window Fourier transform, and the difference between the current fusion features and the historical normal state is calculated by combining cosine similarity. A multi-level threshold discrimination mechanism is established to achieve rapid response and precise positioning of abnormal situations. At the same time, the processing strategy is automatically adjusted according to the level of difference, which effectively improves the system's alarm accuracy and environmental adaptability, ensuring that the visual system can still operate stably and reliably in complex and changing scenarios.
[0028] See also Figure 2 The present invention also provides a visual system based on multi-scene and visible light image recognition, comprising: The data acquisition module synchronously collects polarization information and light intensity information in the target scene and constructs a vector parameter matrix; The image preprocessing module constructs reflection suppression weights based on the vector parameter matrix, generates a light intensity image after reflection suppression, and constructs a dynamic nonlinear compensation surface based on the light source color temperature and absolute light intensity to dynamically enhance the image; The feature extraction and fusion module dynamically adjusts the visible light band range according to the color temperature of the light source, decomposes the enhanced image into a three-band feature map, and extracts polarization features and spectral features separately through a cascaded neural network architecture, and fuses them to generate a fusion feature; The anomaly detection module performs a sliding window Fourier transform on the fused features, extracts frequency domain features, calculates the comprehensive difference between the current fused features and the historical benchmark, and triggers an anomaly alarm or dynamically adjusts the processing strategy based on a multi-level threshold discrimination mechanism.
[0029] In the application, the several formulas involved are all calculated by taking their numerical values after removing the dimensions, and the formula is a formula obtained by collecting a large amount of data and performing software simulation to obtain the latest real situation. The coefficients in the formula are set by technical personnel in this field according to actual conditions.
[0030] The above embodiments can be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0031] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0032] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A visual method based on multi-scene and visible light image recognition, characterized in that: include: Synchronously collect polarization information and light intensity information in the target scene to construct a vector parameter matrix; Based on the vector parameter matrix, the reflection suppression weight is constructed to generate the light intensity image after reflection suppression. The dynamic nonlinear compensation surface is constructed by combining the light source color temperature and absolute light intensity to dynamically enhance the image. Dynamically adjust the visible light band range according to the color temperature of the light source, decompose the enhanced image into a three-band feature map, and use a cascaded neural network architecture to extract polarization features and spectral features respectively, and fuse them to generate a fusion feature; The fusion features and three-band feature maps are subjected to sliding window Fourier transform respectively to extract frequency domain features and calculate the comprehensive difference. Based on the multi-level threshold discrimination mechanism, abnormal alarms are triggered or the processing strategy is dynamically adjusted.
2. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: The collection of polarization light information is achieved through a CMOS array deployed in the target scene. The CMOS array is composed of multiple microlenses. Each node captures horizontal, vertical, 45° oblique and 135° oblique polarization light information through four-way polarization state separation. The collection of light intensity information is achieved through a three-in-one light intensity sensor array deployed in the target scene. The light intensity sensor array simultaneously collects absolute light intensity, light source color temperature and light source azimuth.
3. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: The reflection suppression weight is constructed according to the vector parameter matrix. The reflection suppression weight includes calculating the polarization feature weight and the non-polarization feature weight. The polarization feature weight is used to suppress the polarization interference in the high-reflection area, and the non-polarization feature weight is used to generate a reflection suppression image. The vector parameter matrix after reflection suppression is output.
4. The visual method based on multi-scene and visible light image recognition according to claim 3, characterized in that: A nonlinear compensation surface is constructed based on the absolute light intensity and light source color temperature. The pixel values are mapped to the range of [0, 1]. Pixel-level transformation is performed based on the nonlinear compensation surface. A compensation coefficient is applied to the shortwave band to suppress the color cast of the cold light source. The image is dynamically enhanced and the contrast of the enhanced image is calculated. When the contrast is less than the contrast threshold, the nonlinear compensation surface parameters are recalibrated.
5. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: When the color temperature of the light source is less than 4000K, the long-wave band is extended to 650nm, the short-wave band is compressed to 470~550nm, and the medium-wave band is fixed at 550~600nm; when the color temperature of the light source is greater than 5000K, the short-wave band is lowered to 430nm, the long-wave band is limited to 600~630nm, and the medium-wave band is adjusted to 500~580nm; According to the three characteristic bands of visible light, the enhanced image is decomposed to obtain the characteristic maps of short wave, medium wave and long wave, that is, the three-band characteristic map.
6. The visual method based on multi-scene and visible light image recognition according to claim 5, characterized in that: The cascaded neural network architecture includes a polarization branch and a spectral branch. The polarization branch inputs the linear polarization component after reflection suppression. and , polarization features are extracted through deformable convolution and channel attention mechanism; the spectral branch inputs the three-band feature map, and the spectral features are extracted through the residual network and attention mechanism.
7. The visual method based on multi-scene and visible light image recognition according to claim 6, characterized in that: After the polarization features are upsampled to the same resolution as the spectral features through bilinear interpolation, the polarization features and spectral features are concatenated along the channel dimension to output the fused features: ,in, is the spectral characteristic, is the polarization characteristic, R is the set of real numbers, H is the height, W is the width, C is the number of channels.
8. The visual method based on multi-scene and visible light image recognition according to claim 1, characterized in that: Calculate the difference of fused features: ,in, is the historical benchmark value of the fusion feature under normal conditions, is the frequency domain feature of the fusion feature; Calculate the difference between the three-band features: ,in, For the i The frequency domain characteristics of the band, For the i The historical benchmark value of the band under normal conditions, S For shortwave, M For medium wave, L is the long wave; the difference of the fusion feature and the difference of the three-band feature are weighted to generate a comprehensive difference.
9. The visual method based on multi-scene and visible light image recognition according to claim 8, characterized in that: The multi-level threshold discrimination mechanism includes: first-level discrimination: when the comprehensive difference is greater than or equal to D1, an audible and visual alarm is triggered, and the lighting intensity in the abnormal area is enhanced; second-level discrimination: when the comprehensive difference is greater than D2 and less than D1, the band range is dynamically adjusted according to real-time color temperature data; third-level discrimination: when the comprehensive difference is less than or equal to D2, a sliding average strategy is used to update the historical benchmark value.
10. A visual system based on multi-scene and visible light image recognition, used to implement the method according to any one of claims 1 to 9, characterized in that: include: The data acquisition module synchronously collects polarization information and light intensity information in the target scene and constructs a vector parameter matrix; The image preprocessing module constructs reflection suppression weights based on the vector parameter matrix, generates a light intensity image after reflection suppression, and constructs a dynamic nonlinear compensation surface based on the light source color temperature and absolute light intensity to dynamically enhance the image; The feature extraction and fusion module dynamically adjusts the visible light band range according to the color temperature of the light source, decomposes the enhanced image into a three-band feature map, and extracts polarization features and spectral features separately through a cascaded neural network architecture, and fuses them to generate a fusion feature; The anomaly detection module performs a sliding window Fourier transform on the fused features, extracts frequency domain features, calculates the comprehensive difference between the current fused features and the historical benchmark, and triggers an anomaly alarm or dynamically adjusts the processing strategy based on a multi-level threshold discrimination mechanism.
Citation Information
Patent Citations
Camouflage object identification method and system based on polarized light acquisition and wave band operation
CN115359357A
Target polarization detection system under strong background and detection method thereof
CN116503704A
Method for suppressing glow effect of low-illumination image based on polarization characteristics
CN118691485A
Construction method and system of RGB intelligent night vision light source
CN119922794A
Visible light brightness detection method and system coping with camera
CN120343407A
Cited By
Acoustic module defect classification method and system based on multi-feature fusion
CN121280813A