Platform End Intrusion Detection System Based on Multimodal Front Fusion

By using multimodal pre-fusion technology, the system identifies feature failure areas using local information entropy maps and performs feature modulation and alignment, thus solving the problem of recognition stability and accuracy of the station monitoring system under extreme light field environments and achieving high-precision intrusion recognition in complex backgrounds.

CN121600476BActive Publication Date: 2026-04-17HUNAN YOULIANG ELECTRONIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN YOULIANG ELECTRONIC TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing platform monitoring systems have difficulty effectively identifying intrusion targets in extreme light field environments, and there are problems such as identification logic oscillation and false alarms. This is mainly due to insufficient feedback of the physical state of sensor imaging, which leads to the noise component of the failed mode contaminating the effective mode features.

Method used

The system employs multimodal pre-fusion technology, acquiring visible light and infrared image data through an image acquisition module, using an effectiveness evaluation module to calculate local information entropy maps to identify feature failure areas, a feature modulation module to generate spatial confidence masks for local suppression and infrared feature gain compensation, a spatial alignment module to perform coordinate alignment of heterogeneous feature maps, and a target detection module to perform intrusion identification by combining platform topology logic.

Benefits of technology

Maintaining the stability and accuracy of feature extraction under extreme light field environments, adaptive modulation is achieved through local information entropy field gradient constraints to eliminate the influence of noise components, thereby improving the accuracy and reliability of the recognition system in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600476B_ABST
    Figure CN121600476B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image recognition technology and discloses a platform end intrusion detection system based on multimodal pre-fusion, comprising: an image acquisition module for acquiring visible light and infrared images; an effectiveness evaluation module for identifying feature failure areas using the changing gradient of the local information entropy map; a feature modulation module for suppressing visible light failure areas and achieving infrared feature gain compensation based on mask weights to generate a fused feature map; a spatial alignment module for calculating offset vectors using feature anchor points to correct coordinate alignment deviations; and a target detection module for identifying intrusion targets using multimodal features. This invention utilizes information entropy to guide feature scheduling logic, blocking the flow of invalid features caused by high-contrast light fields during the extraction stage, eliminating physical drift of the sensor optical axis, and ensuring the robustness of the system in maintaining target recognition in the platform end environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a platform end intrusion detection system based on multimodal front fusion, belonging to the field of image recognition technology. Background Technology

[0002] Current platform monitoring tasks utilize visible light imaging sensors to acquire target images and employ infrared sensing technology to provide identification support in low-light environments. The physical topology of rail transit platforms exhibits a corridor effect, where the direct light from trains entering the station and the deep shadows of the platform structure create a high-contrast light field. Silicon-based sensors, limited by linear exposure characteristics, suffer from a limited dynamic range that leads to the physical loss of key features, resulting in oversaturated or deep shadow areas in the image. This causes semantic feature collapse in the visible light mode. For example, Chinese invention patent CN118298377B discloses a perimeter intrusion identification method and system based on joint video acquisition. It acquires the overall ambient brightness through a photosensor and performs selective activation or joint tracking between an infrared pan-tilt unit and a binocular camera. Although the global threshold switching mechanism can handle alternating light and dark conditions throughout the day, it belongs to the post-fusion processing logic and lacks refined feedback on the physical state of sensor imaging. In the face of strong backlighting and local shadows at the end of the platform, it cannot identify and shield the noise components generated by the failed modes, causing erroneous features to flow to the fusion layer, leading to oscillations or false alarms in the identification logic.

[0003] Existing fusion schemes mostly adopt fixed feature stitching or global attention weight allocation methods. Such methods lack feedback on the physical state of sensor imaging, which causes noise components generated by the failed modes to contaminate the normal features of the effective modes during the fusion stage, causing logical oscillations of the recognition system under extreme conditions. In addition, the feature extraction process based on global computation generates processing delays, making it difficult to adapt to the real-time recognition requirements of high-speed moving targets.

[0004] Therefore, how to shield and adaptively modulate failure information at the feature extraction front end based on the sensor's photosensitive characterization, and ensure the stability of feature extraction under extreme light field environments, has become the technical problem to be solved by this invention. Summary of the Invention

[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A platform end intrusion identification system based on multimodal front fusion, the system comprising:

[0006] The image acquisition module is used to acquire visible light image data and infrared image data of the monitoring area at the end of the platform;

[0007] The effectiveness evaluation module is used to traverse visible light image data using a sliding window to calculate the local information entropy map, and to identify feature failure areas based on the gradient of entropy value changes in local areas of the local information entropy map.

[0008] The feature modulation module is used to generate a spatial confidence mask based on the distribution of feature failure regions, perform local suppression on features in visible light image data using pixel-level weights determined by the spatial confidence mask, and perform feature gain compensation on the suppressed regions using features at corresponding locations in infrared image data to generate heterogeneous fused feature maps.

[0009] The spatial alignment module is used to extract gradient feature points in the high-confidence region indicated by the spatial confidence mask, calculate the cross-correlation offset vector of the gradient feature points in the corresponding region of the infrared image data, and perform coordinate alignment correction on the heterogeneous fused feature map according to the cross-correlation offset vector to generate the aligned heterogeneous fused feature map.

[0010] The target detection module is used to extract multimodal contour features based on the aligned and corrected heterogeneous fusion feature map, and perform classification and identification of intrusion targets in combination with the preset platform topology constraint logic.

[0011] Preferably, when generating the spatial confidence mask, the feature modulation module extracts the local gray-level gradient of the visible light image data, performs pixel-level dot product operation on the local gray-level gradient and the local information entropy map to generate a confidence matrix, and performs normalization processing on the confidence matrix to determine the weight value of the spatial confidence mask at each spatial coordinate point.

[0012] Preferably, when identifying feature failure areas, the effectiveness evaluation module identifies oversaturated areas caused by direct external light sources or low-light shadow areas formed by the platform structure by calculating the difference in local information entropy at the same coordinate positions between adjacent frames, and marks the oversaturated areas or low-light shadow areas as feature failure areas.

[0013] Preferably, the logic for the feature modulation module to perform local suppression and feature gain compensation follows the following formula: ,in, These are the modulated fusion feature values. These are the raw feature values ​​of the visible light image data. These are the raw feature values ​​of the infrared image data. These are the weighting coefficients corresponding to the spatial confidence mask. The preset infrared characteristic gain factor, and and All are dimensionless coefficients.

[0014] Preferably, when calculating the cross-correlation offset vector, the spatial alignment module performs sliding matching within the candidate region of the infrared image data using a search window, taking the position corresponding to the maximum cross-correlation coefficient as the matching point, and calculates the pixel displacement between the matching point and the gradient feature point as the basis for compensating the physical offset of the optical axis of the image acquisition module.

[0015] Preferably, after extracting multimodal contour features, the target detection module uses image connected component analysis logic to perform clustering processing on the pixels in the aligned and corrected heterogeneous fused feature map in order to extract the connected regions of candidate targets.

[0016] Preferably, the target detection module further includes a geometric constraint verification unit, which uses the parallax constraint generated by the installation baseline of the image acquisition module to perform spatial position verification on the connected regions of candidate targets and eliminate falsely reported targets that exceed the physical space of the platform.

[0017] Preferably, the spatial alignment module is also used to record historical data of cross-correlation offset vectors, calculate the fluctuation variance of historical data, and output maintenance warning instructions for the installation status of image acquisition equipment when the fluctuation variance exceeds a preset stability threshold.

[0018] Preferably, the system also includes a dynamic parameter update module, which performs a global average calculation on the local information entropy map to generate a global information entropy mean, and adjusts the confidence threshold of target determination in the target detection module in real time according to the global information entropy mean.

[0019] Preferably, when performing classification and recognition, the target detection module compares the multimodal contour features with the preset human skeleton structure features to determine whether there is a person intruding into the end area of ​​the platform.

[0020] Compared with the prior art, the beneficial effects of the present invention are:

[0021] 1. In the platform end intrusion identification system, a spatial confidence mask is constructed by utilizing the local gray-level distribution gradient of visible light images. During the feature extraction stage, local nonlinear suppression is performed on low-confidence areas, and the infrared feature gain is compensated according to the suppression intensity. This blocks the flow of invalid features caused by oversaturation or deep shadows at the signal acquisition front end, avoids the noise component generated by visible light mode failure from contaminating the fused feature layer, and ensures that the system maintains the physical closed loop and stability of semantic feature extraction in an extremely high dynamic range light field environment.

[0022] 2. By introducing local information entropy field gradient constraints, the feature scheduling logic is elevated from the dimension of light intensity perception to the dimension of information validity perception. The semantic collapse region is identified by using the local entropy value drop rate, enabling the feature extraction process to adjust the modal weights in real time according to the information abundance of imaging features. Under the premise of maintaining low computational load, high-precision adaptive modulation of heterogeneous features under complex background interference is achieved. The system uses the state representation generated by the spatial confidence mask as the detection carrier, extracts gradient feature anchor points using the high confidence region indicated by the mask, and calculates the cross-correlation offset vector of the anchor point in the corresponding neighborhood of the infrared feature map. This achieves sub-pixel level dynamic feature alignment. The algorithm logic not only eliminates the physical drift of the optical axis caused by the deformation of the sensor support surface or environmental vibration, but also ensures the pixel-level accuracy consistency of multimodal semantic edges during long-term operation through a multi-feature linkage mechanism.

[0023] 3. Combining thermodynamic connectivity analysis and parallax topology consistency verification, and utilizing the geometric constraints generated by the physical installation baseline of the sensors, the spatial displacement of the target centroid in the fused feature map is verified. This effectively removes virtual artifacts caused by water reflection and glass refraction. In conjunction with the nonlinear constraint rules of the human motion physical model, interference signals generated by non-biological heat sources are eliminated. Without increasing the active ranging hardware, the system's recognition reliability is improved in dark, high-temperature, and complex optical interface conditions. The exposure evaluation threshold is dynamically corrected using the photosensitive residual energy of the imaging background, achieving deep coupling between the fusion logic and the ambient light field quality. This ensures that the feature extraction engine always maintains the best gain balance in low-light or low-contrast conditions, enabling the system to have self-correction capabilities against hardware performance degradation and gradual environmental interference, thus enhancing the operational stability of the recognition task throughout its all-weather service life. Attached Figure Description

[0024] Figure 1 This is a flowchart of the multimodal pre-fusion recognition process guided by local information entropy in this invention;

[0025] Figure 2 This is a schematic diagram of the station end perception architecture and edge fusion computing engine of the present invention;

[0026] Figure 3 This is a diagram illustrating the business interaction logic and functional modules based on the multi-dimensional verification mechanism of this invention. Detailed Implementation

[0027] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0028] This invention provides a platform end intrusion detection system based on multimodal pre-fusion, comprising an image acquisition module, an effectiveness evaluation module, a feature modulation module, a spatial alignment module, and a target detection module. The image acquisition module acquires visible light and infrared image data of the platform end monitoring area. The effectiveness evaluation module determines the physical effectiveness of features based on the local information entropy field gradient of the visible light image. The feature modulation module uses a generated spatial confidence mask to perform dynamic suppression and compensation on heterogeneous feature streams. The spatial alignment module performs coordinate correction based on feature anchor points to eliminate optical axis drift. The target detection module combines topological constraint logic to classify and determine intrusion targets. Tensor-level information is transferred between the modules via a high-speed data bus. The image acquisition module acquires visible light and infrared image data of the platform end monitoring area. In the engineering environment of the platform end, due to the vibration generated by train entering the station and the intrinsic difference in the photosensitive frequencies of the visible light and infrared sensors, the image acquisition module ensures shutter trigger synchronization of the dual sensors within nanosecond precision through a precise clock synchronization mechanism. The visible light image data output by the image acquisition module is defined as a three-dimensional tensor. Infrared image data is defined as a three-dimensional tensor. ,in It has high spatial resolution to extract contour details. This dual-path synchronous sensing method provides a physical layer consistency basis for feature-level pre-fusion, enabling stable thermal source characteristics under extreme lighting conditions.

[0029] The effectiveness evaluation module uses a sliding window to traverse visible light image data to calculate a local information entropy map. Based on the gradient of entropy changes in local regions within the local information entropy map, it identifies feature failure areas. Due to the corridor effect in the physical topology of the platform ends, the direct light from trains entering the station and the deep shadows of the platform structure create a high-contrast light field, causing the silicon-based sensor to physically lose key semantic features within a limited dynamic range, resulting in oversaturated or deep shadow areas. To quantify the degree of information loss, the effectiveness evaluation module uses a size of... Sliding window traversal of visible light image tensors In this embodiment Values The effectiveness evaluation module calculates a histogram of pixel grayscale distribution within each window and calculates the local information entropy of that coordinate point based on the histogram. The calculation formula is as follows: ,in, For local information entropy, The grayscale value within a local window The frequency of pixel occurrence; the local information entropy of each coordinate point constitutes an information entropy distribution map aligned with the tensor space of the visible light image. The effectiveness evaluation module calculates the second-order spatial gradient of this information entropy distribution map. When the entropy value drop rate of a local region exceeds a preset stability threshold, it is determined that the region has undergone semantic collapse and is marked as a feature failure region, thereby elevating the feature scheduling logic from the dimension of light intensity perception to information effectiveness perception. The effectiveness evaluation module identifies the stability threshold calibration required for feature failure regions. During periods when no trains are entering the station and the platform lighting is constant and static, the image acquisition module acquires no less than 500 frames of visible light image sequences, calculates the local information entropy map corresponding to each frame, extracts the standard deviation of entropy value fluctuation within the time series of each spatial coordinate point, and takes the maximum value of the overall fluctuation standard deviation as the environmental background noise benchmark. The stability threshold is set to 3.5 times, As a dimensionless quantity, it characterizes the inherent random disturbance of the imaging system under a specific installation environment, and establishes a quantitative benchmark for characteristic failure judgment.

[0030] The feature modulation module generates a spatial confidence mask based on the distribution of feature failure regions. It then uses the pixel-level weights determined by the spatial confidence mask to perform local suppression of features in the visible light image data and uses features at corresponding locations in the infrared image data to perform feature gain compensation on the suppressed regions, generating a heterogeneous fusion feature map. When strong backlight interference exists at the platform end, large-scale noise components generated by the visible light mode can contaminate the fusion features, causing recognition logic oscillations. The feature modulation module extracts the local grayscale gradient of the visible light image data and performs a pixel-level dot product operation with it and the local information entropy map to generate a confidence matrix. The weight values ​​of the spatial confidence mask at each spatial coordinate point are determined by normalizing the confidence matrix. Weight value The range of values ​​is to The logic for local suppression and feature gain compensation performed by the feature modulation module follows the formula below: ,in, These are the modulated fusion feature values. These are the raw feature values ​​of the visible light image data. These are the raw feature values ​​of the infrared image data. These are the weighting coefficients corresponding to the spatial confidence mask. The preset infrared characteristic gain factor, and and All are dimensionless coefficients; when at a certain position Approaching At that time, the system forcibly shuts off the visible light feature flow at that location and simultaneously increases the transfer function slope of the infrared feature at the corresponding location, ensuring the dominance of the infrared mode at critical moments, thereby avoiding noise pollution. for , for , for , for For example, the fusion feature value is calculated. for The infrared characteristic gain factor required for the characteristic modulation module to perform gain compensation. Based on the online calibration of the infrared sensor's background noise level, at least 100 frames of raw infrared images were acquired in a zero-illuminance environment with the infrared lens shielded, and the standard deviation of the average dark current noise for all pixels was calculated. According to the proportional relationship Determine the gain coefficient. The infrared characteristic gain factor is a dimensionless constant. The grayscale unit is the standard deviation of dark current noise of the infrared sensor. By quantifying the fluctuation of the sensor's thermal sensitivity, the upper limit of energy gain for infrared modes during feature recovery in the visible light failure region is established.

[0031] The spatial alignment module extracts gradient feature points within the high-confidence region indicated by the spatial confidence mask, calculates the cross-correlation offset vector of the gradient feature points within the corresponding region of the infrared image data, and performs coordinate alignment correction on the heterogeneous fused feature map based on the cross-correlation offset vector to generate an aligned heterogeneous fused feature map. During long-term service, high-frequency vibrations caused by trains entering stations lead to surface deformation of the sensor bracket, causing sub-pixel-level dynamic drift of the optical axes of the visible light sensor and the infrared sensor. This results in semantic ghosting when multimodal features are cascaded. The spatial alignment module reuses the mask information generated by the validity evaluation module. Greater than Within a high-confidence region, pixels with gradient strength exceeding a preset threshold are identified as reference anchor points, using a region of size [missing information]. The search window performs sliding matching within the candidate region of the infrared image data, using the position corresponding to the maximum cross-correlation coefficient as the matching point. The pixel displacement between the matching point and the reference anchor point is calculated as the cross-correlation offset vector. The spatial alignment module uses this vector to apply to the coordinate mapping function of the spatial transformation module, performing local tensor displacement compensation to achieve integrated recognition and calibration. In the initial parameter settings of the spatial alignment module, the side length of the search window... The procedure determines that the image acquisition module is used to obtain continuous image data of the train at the preset maximum station entry speed, and the maximum pixel offset between sensors is identified by feature tracking. ,follow The calculation logic constructs a matching window. The unit for the search window side length is pixels. The maximum pixel offset is measured in pixels to ensure that the cross-correlation matching process covers the physical displacement envelope of the sensor's optical axis, thus eliminating the computational load of physical drift control feature retrieval.

[0032] The target detection module extracts multimodal contour features based on the aligned and corrected heterogeneous fusion feature map, and performs intrusion target classification and recognition in conjunction with preset platform topology constraint logic. In complex optical environments, ground water reflection or glass refraction can easily produce high-brightness virtual artifacts. The target detection module includes a geometric constraint verification unit, which uses the disparity constraints generated by the physical installation baseline between the sensors in the image acquisition module to perform spatial position verification on the connected regions of candidate targets. The target detection module calculates the feature centroid coordinates of the same semantic target in two tensors and compares its spatial displacement vector with the theoretical disparity threshold calculated based on the physical baseline. If the displacement vector does not conform to the preset human dynamics topology constraint rules, the target is determined to be a virtual artifact and the alarm command is canceled. At the same time, the target detection module compares the extracted multimodal contour features with the preset human skeleton structure features, and eliminates interference signals generated by non-biological heat sources by analyzing the evolution trajectory of the centroid in continuous frames. The system also includes a dynamic parameter update module, which performs global averaging calculation on the local information entropy map. The system generates a global information entropy mean and adjusts the confidence threshold for target determination in the target detection module in real time based on the global information entropy mean. When the environment is in a state of low global contrast caused by dense fog or sandstorm, the global information entropy mean decreases. The dynamic parameter update module uses a monotonic nonlinear function to correct the overexposure and underexposure discrimination thresholds in real time based on the change in the mean, ensuring the coupling between the fusion logic and the quality of the ambient light field, so that the system maintains gain balance throughout the day and night. The spatial alignment module also records the historical data of the cross-correlation offset vector, calculates the fluctuation variance of the historical data, and outputs maintenance warning instructions for the installation status of the image acquisition equipment when the fluctuation variance exceeds the preset stability threshold. By analyzing the statistical characteristics of the physical displacement of the sensor, the mechanical fatigue degree of the hardware support is characterized, realizing the self-diagnosis of equipment integrity and providing data support for the preventive maintenance of the rail transit monitoring system. In this embodiment, the geometric constraint verification unit establishes a parallax geometric model based on the horizontal installation baseline distance of the dual sensors in the image acquisition module (set to 150mm±0.5mm). When performing spatial position verification, the system reads the 3D laser point cloud topology map of the installation location and uses it as a physical fence for the movement range of candidate targets. When a target is in the candidate state, the system calculates its human proportion parameter R. This parameter is obtained by measuring the pixel height distance from the center of mass of the head and neck to the center of the torso and the pixel height distance from the torso to the hip joint. The system retrieves the imaging projection scale at the corresponding focal length in real time based on the current optical focal length range of the visible light lens (such as 6mm-12mm zoom travel). If the measured scale R deviates significantly from the normal physiological fluctuation range of 1.2 to 1.8, and the thermal energy center of mass of the target in consecutive frames exhibits a nonlinear scattering path consistent with the reflection law of surface water, the system determines that the signal source is an optical virtual image. Through this logic of deeply coupling human kinematic topology with optical propagation physical characteristics, physical shielding of non-biological heat source interference signals is achieved.

[0033] Example 1: In the monitoring area at the end of a rail transit platform, the system faces the situation where trains turn on their high-brightness headlights to enter the station at dusk. At this time, the directional strong backlight forms a high-contrast light field in the physical topology of the platform that exceeds the linear photosensitive dynamic range of the sensor, causing the visible light image tensor output by the image acquisition module to be affected. The system generates oversaturated regions where pixel values ​​reach their upper limit, resulting in physical information loss of semantic features originally used to extract target contours in the visible light mode. The effectiveness evaluation module executes an evaluation process based on information abundance perception, using a size of... Sliding window traversal of visible light image tensors Local information entropy is calculated by statistically analyzing the histogram of pixel grayscale distribution within a local window. The calculation process follows the formula: ,in, For local information entropy, The grayscale value within a local window The frequency of pixel occurrence; since the gray-level distribution in the oversaturated region tends to be uniform, the calculated value in this region... A decrease in numerical value occurs; the effectiveness evaluation module captures the second-order gradient change of the information entropy field in the spatial domain, and converts the local information entropy... Regions where the drop rate exceeds a preset stability threshold are identified as feature failure regions. This evaluation method anchors the feature scheduling logic to the physical representation of the sensor's imaging effectiveness, rather than simply the light intensity.

[0034] The feature modulation module generates a spatial confidence mask based on the distribution of feature failure regions and calculates the weight value of each coordinate point. In the region where the oversaturated feature fails, the weight value Downgraded to to Within the specified interval, the feature modulation module uses this spatial confidence mask to locally suppress high-noise features in the visible light image data, and simultaneously introduces features at the corresponding locations in the infrared image data to perform feature gain compensation. The calculation logic follows the formula: ,in, These are the modulated fusion feature values. These are the raw feature values ​​of the visible light image data. These are the raw feature values ​​of the infrared image data. These are the weighting coefficients corresponding to the spatial confidence mask. The system uses a preset infrared feature gain factor. By switching the weights of this feature channel, it blocks the transmission of visible light noise during the feature extraction stage and uses the stable heat source features provided by the infrared mode to compensate for missing semantic information. for , for , for , for For example, the fusion feature value is calculated. for The spatial alignment module simultaneously monitors the vibration of the sensor brackets caused by the train entering the station, within the spatial confidence mask. Greater than Gradient feature points are extracted from stable background regions by performing a process based on infrared image data. The cross-correlation matching calculation of the search window is affected by the cross-correlation offset vector generated by sensor displacement. The spatial alignment module feeds this vector back to the coordinate mapping function to perform sub-pixel-level displacement correction on the heterogeneous fused feature map. The target detection module extracts multimodal contour features on the aligned heterogeneous fused feature map. The geometric constraint verification unit compares the feature centroid coordinates of the same target in the dual-path tensor to eliminate virtual artifacts formed by ground water reflection. The system completes the identification of real intrusion targets in an environment where the visible light mode physically collapses. This feature modulation method based on information entropy guidance achieves dynamic weighting of visible light and infrared modes in the feature space without changing the physical characteristics of the sensor. Through the synergistic operation of local feature suppression and infrared gain compensation, the system transforms from a single-mode quality dependency to an adaptive management of heterogeneous feature streams, eliminating the interference of high-contrast light fields on the intrusion identification logic.

[0035] Example 2: In a simulation test at the end of a rail transit platform, the system faces an optical field contrast ratio of... And accompanied by The test platform for the power frequency electromagnetic interference was constructed in a test anechoic chamber simulating the physical topology of a platform. The data source was a sensing array consisting of visible light sensors and long-wave infrared sensors. The visible light sensors possess... Dynamic range, sampling frequency is The resolution is Infrared sensors have Thermal sensitivity, resolution is The signal-to-noise ratio of the actively superimposed signal source in the test signal source is Gaussian white noise was used to simulate signal degradation in a real industrial environment. This experiment was used to verify the accuracy of the local information entropy gradient in identifying failure features and the ability of the feature modulation module to suppress false targets.

[0036] Determine the sliding window size for the validity assessment module. At that time, the primary considerations for identification are the overflow current width generated by the image sensor at the oversaturation edge and the local connectivity of semantic features. The technical trade-off lies in the ability to smooth pixel-level isolated noise versus the degree of preservation of edge feature gradients. If the value is too small, then... The local entropy fluctuations caused by Gaussian noise are sensitive, leading to fragmentation of the spatial confidence mask. If the value is too large, the physical boundary of the oversaturated region will be blurred, suppressing the features of the normal region. The decision rule is based on the functional relationship between the signal-to-noise ratio of local features and the edge contrast, that is, when the background noise power spectral density increases... To enhance statistical stability, the value will be increased, taking into account the signal-to-noise ratio conditions of this experiment. Set as As an engineering example of applying this decision-making logic, the effectiveness of the technical solution was analyzed by comparing the detection performance of the control group and the sample group of the present invention. The control group adopted a feature splicing method based on fixed weights, while the sample group of the present invention executed a feature modulation process guided by local information entropy, with the light intensity increasing from... Increase to During the process, the sample group of this invention captures local information entropy. The second-order spatial gradient, in rate of fall exceeds The region generates a spatial confidence mask, making the weight values... Converging in the oversaturated core region nearby.

[0037] Table 1: Comparison of key data of different groups under strong backlight conditions in this embodiment

[0038]

[0039] Analysis of the data in Table 1 shows that the local information entropy Due to supersaturation decreasing At that time, the weight values ​​of the sample group in this invention are dynamically adjusted. to To suppress high noise characteristics in visible light and to compensate for infrared characteristic gain factors Set as Ultimately, the target detection accuracy remained at Compared to control group 1 It has an improvement, in which the out-of-range control group showed that when Below and Exceed The degradation effect at this time means that the background thermal noise of the infrared mode is amplified, causing the accuracy to drop back to [a certain value]. This proves that the limited parameter range is the optimized working window; regarding the infrared characteristic gain factor The setting follows the principle of nonlinear compensation. When the weight value of the spatial confidence mask is determined... In to In between, the system uses formulas Perform fused computing. These are the modulated fusion feature values. These are the raw feature values ​​of the visible light image data. These are the raw feature values ​​of the infrared image data. These are the weighting coefficients. This refers to the infrared characteristic gain factor; data shows that as the degree of visible light mode failure increases, i.e. Reduce, system upgrade The value is used to maintain the total energy balance of the fusion tensor, but when Exceed After this performance inflection point, the shot noise of the infrared sensor increases in the fused feature map, verifying that the gain compensation mechanism has a physical upper limit. The gain range determined in the experiment allows the heterogeneous feature flow to maintain smooth gain under various optical field conditions. The spatial alignment module demonstrates adaptive correction capability to mechanical vibration in the experiment. By monitoring the fluctuation variance of the historical data of the cross-correlation offset vector, the system detects the frequency generated by the sensor bracket. Amplitude is When shifting pixels, adjust the mapping relationship of the spatial transformation module to control the alignment deviation within a certain range. Within pixels.

[0040] Example 3: This example combines Figures 1 to 3 The following describes a platform end intrusion detection system based on multimodal front fusion, such as... Figure 1 As shown, for the monitoring area at the end of the platform containing high-contrast light fields or complex backgrounds, the image acquisition module acquires visible light image data and infrared image data to form a raw image stream. The stream then enters the effectiveness evaluation module to calculate the local information entropy map and uses entropy value change gradient analysis to identify feature failure areas and output failure area markers. After the data stream is transferred to the feature modulation module, a spatial confidence mask is generated and local suppression and feature gain compensation are performed to generate a heterogeneous fusion feature map. Then, the spatial alignment module calculates the cross-correlation offset vector based on gradient feature points and anchor points, corrects the coordinate deviation, and outputs the aligned and corrected features. Finally, the target detection module extracts multimodal contour features and identifies and classifies them by combining the platform topology constraint logic, and outputs the intrusion target identification result.

[0041] like Figure 2As shown, in a high-contrast and vibration environment at the end of the platform, the infrared thermal imaging stream and the visible light image stream are connected to the system through a shared anti-vibration bracket. After the data is transmitted to the edge fusion computing engine, local entropy evaluation, spatial mask generation, feature dynamic modulation, and coordinate alignment correction are performed sequentially. The fusion feature results output by this engine are used to trigger intrusion detection alarms on the one hand, and to generate maintenance warnings for hardware status by analyzing the feedback path on the other hand; for example... Figure 3 As shown, the system's interaction logic involves station monitoring personnel and system maintenance personnel. The core business flow is to identify intrusions at the station end. This process includes acquiring visible light and infrared image data, evaluating the effectiveness of features and analyzing the local entropy gradient, generating heterogeneous fusion feature maps, and performing spatial alignment correction. When performing intrusion target classification, it includes geometric constraint verification based on the parallax principle to eliminate virtual artifacts, and dynamically adjusting the judgment threshold based on the global entropy mean. In addition, the system is also configured with a function path to output maintenance warnings to system maintenance personnel when the support displacement is too large.

[0042] Example 4: In the process of algorithm deployment and initialization parameter calibration of the intrusion detection system at the end of a rail transit platform, the system needs to perform discretization configuration for the second-order spatial gradient operation in the effectiveness evaluation module to solve the problem of judgment benchmark drift caused by the drastic change in local information entropy field due to strong light caused by the train entering the station. The initial input is... Visible light image tensor with gray levels Its imaging noise level, as measured, has a standard deviation of not less than [a certain value]. After acquiring the local information entropy map, the grayscale and effectiveness evaluation module uses a discretized second-order Laplacian operator to extract the gradient of the information entropy field. This is achieved by summing the local information entropy of a local coordinate point and its four adjacent pixels (up, down, left, and right), and then subtracting the local information entropy of that coordinate point. To calculate the gradient of entropy change in a local region. Its operational logic satisfies the following formula: ,in, The gradient represents the entropy change in a local region. Coordinates of points in the local information entropy graph Local information entropy at the location, and For the spatial coordinate index of the pixel; when calculated When the preset stability threshold is exceeded, the system determines that there is a change in the light field distribution in the area and marks it as a characteristic failure area. This stability threshold is determined by continuously collecting data during static periods when no trains are entering the platform. Frame image sequence obtained, specific statistical analysis of background area The maximum fluctuation amplitude is used as the background benchmark, and this benchmark value is taken. A factor of 1.5 is used as the stability threshold for determining feature failure, with the measured background fluctuation amplitude as the threshold value. For example, the stability threshold is calculated as follows: .

[0043] Infrared characteristic gain factor in the characteristic modulation module The dark current noise tensor of the infrared sensor was determined through an online calibration process based on the background noise level of the thermal imaging sensor. This was achieved by acquiring the tensor in a zero-illuminance environment after the lens was blocked, and calculating its pixel-level standard deviation. Preset infrared characteristic gain factor Based on linear transformation logic The settings were configured to maintain the dynamic range of multimodal characteristics as the sensor's thermal sensitivity decays with fluctuations in ambient temperature, based on the measured dark current noise standard deviation. for For example, the calculation yields for , This is the infrared characteristic gain factor, which is a dimensionless constant. The standard deviation of the dark current noise of the infrared sensor; the weight values ​​determined by the feature modulation module based on the spatial confidence mask. The heterogeneous feature stream is dynamically allocated at the pixel level. After receiving the aligned heterogeneous fused feature map, the target detection module constructs a connected skeleton graph by extracting human joint nodes and verifies the structural features of the target using preset human dynamics topology constraint rules. The specific quantization logic of these rules is as follows: extract the longitudinal distance from the center point of the target's head and neck to the center point of the torso. and the longitudinal distance from the center point of the trunk to the center point of the hip joint of the lower limb. Calculate the ratio between the two. , For human body proportion parameters, This represents the vertical distance from the center point of the head and neck to the center point of the torso, in pixels. The distance from the center point of the torso to the center point of the hip joint, in pixels; system verification. Is it in to If the measured proportion deviates from the physiological range by more than [a certain amount], it indicates that the actual proportion is within the range. If the corresponding feature centroid displacement vector does not meet the preset uniform motion judgment, the system determines that the current target is a virtual artifact formed by mirror reflection and suppresses alarm triggering. This geometric constraint logic solves the recognition error caused by non-biological heat sources in complex optical environments, enabling the system to maintain safe recognition performance in high-contrast light field environments at the end of the platform.

[0044] Example 5: During the deployment and commissioning of the sensing system at a newly built rail transit platform, the system executes the environmental benchmark calibration procedure, continuously acquiring data under normal operating conditions with no trains entering the station. The effectiveness evaluation module utilizes a visible light image sequence of size [size missing]. A sliding window traverses the entire image field to calculate the local information entropy map, and the dynamic parameter update module calculates the local information entropy of each spatial coordinate point within the time series. The standard deviation of the fluctuation is calculated, and the maximum value of this standard deviation is extracted as the environmental disturbance benchmark. The stability threshold is set as the environmental disturbance benchmark. of Times, based on measured environmental disturbance benchmark for For example, the stability threshold calculated for this deployment site is: , For local information entropy, This serves as a benchmark for environmental disturbances.

[0045] When the system encounters interference from false human body images reflected by the glass of the shielding door, the target detection module runs a topology-consistent verification process. It extracts the centroids of connected regions of candidate targets from the heterogeneous fused feature map, uses preset joint detection operators to locate the center points of the head and neck, torso, and hip joints, and calculates the vertical pixel distance from the center point of the head and neck to the center point of the torso. and the vertical pixel distance from the center point of the torso to the center point of the hip joint According to the formula Obtain human body proportion parameters Verify human body proportion parameters Is it in to Within the physiological range, if the human body proportion parameters The measured value deviates from this physiological range, or in a continuous period of time. The frame's characteristic centroid displacement vector exhibits a nonlinear abrupt change, and the system determines that the target is a non-biological artifact and performs alarm suppression.

[0046] Example 6: During the deployment phase of the rail transit platform system, the system executes a search window calibration procedure based on physical vibration limits for the spatial alignment module, and uses the image acquisition module to acquire the displacement data of the train under the station entry condition to identify the maximum pixel offset. And according to the linear mapping formula Calculate the side length of the search window , The side length of the search window, in pixels. This is the maximum pixel offset, in pixels; the maximum pixel offset obtained through actual measurement. for For example, the side length of the search window is calculated. for This calibration value enables the cross-correlation matching process to cover the envelope of the sensor's physical displacement, achieving a balance between coordinate alignment correction accuracy and feature retrieval load.

[0047] When the system encounters low-contrast conditions caused by dense fog, the dynamic parameter update module runs a threshold correction program based on the global information entropy mean and extracts a preset benchmark detection threshold. and the reference entropy value under the standard light field The average global information entropy calculated using the current frame The decision logic of the target detection module was revised, and the decision threshold was calculated. Satisfying the relation ,in To determine the threshold, As the baseline detection threshold, The mean of global information entropy. The reference entropy value is used, and all the above variables are dimensionless parameters; the benchmark detection threshold is used. for And refer to the entropy value for For example, if the current global information entropy average is It dropped to due to decreased atmospheric visibility. The system will determine the threshold. Real-time adjustment This correction process compensates for the judgment error caused by weakened signal strength by adjusting the feature entry threshold, ensuring the system maintains the continuity of target recognition under all-weather environmental changes; during system operation, the spatial alignment module records continuous... The cross-correlation offset vector within each sampling period is calculated, and its variance is also calculated. When the measured variance of fluctuation Exceeding the preset stability threshold At that time, the system generates a maintenance warning signal for the installation status of the image acquisition equipment and sets a stability threshold. The measured background variance benchmark during the system initialization phase The ratio is times the measured background variance benchmark. For example, when the volatility variance The sensor mounting bracket was loosened and the height was increased. At that time, the system outputs a maintenance warning command.

[0048] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A platform end intrusion identification system based on multi-modal pre-fusion, characterized in that, The system includes: The image acquisition module is used to acquire visible light image data and infrared image data of the monitoring area at the end of the platform; The image acquisition module ensures that the dual sensors achieve shutter trigger synchronization with nanosecond-level accuracy through a precision clock synchronization mechanism; The effectiveness evaluation module is used to traverse visible light image data using a sliding window to calculate the local information entropy map, and to identify feature failure areas based on the gradient of entropy value changes in local areas of the local information entropy map. The feature modulation module is used to generate a spatial confidence mask based on the distribution of feature failure regions, perform local suppression on features in visible light image data using pixel-level weights determined by the spatial confidence mask, and perform feature gain compensation on the suppressed regions using features at corresponding locations in infrared image data to generate heterogeneous fused feature maps. The spatial alignment module is used to extract gradient feature points in the high-confidence region indicated by the spatial confidence mask, calculate the cross-correlation offset vector of the gradient feature points in the corresponding region of the infrared image data, and perform coordinate alignment correction on the heterogeneous fused feature map according to the cross-correlation offset vector to generate the aligned heterogeneous fused feature map. The target detection module is used to extract multimodal contour features based on the aligned and corrected heterogeneous fusion feature map, and perform classification and recognition of intrusion targets in combination with the preset platform topology constraint logic; In addition, when generating the spatial confidence mask, the feature modulation module extracts the local gray-level gradient of the visible light image data, performs pixel-level dot product operation on the local gray-level gradient and the local information entropy map to generate the confidence matrix, and performs normalization processing on the confidence matrix to determine the weight value of the spatial confidence mask at each spatial coordinate point. When identifying feature failure areas, the effectiveness evaluation module calculates the difference in local information entropy at the same coordinate positions between adjacent frames to identify oversaturated areas caused by direct external light sources or low-light shadow areas formed by the platform structure, and marks the oversaturated areas or low-light shadow areas as feature failure areas. When calculating the cross-correlation offset vector, the spatial alignment module performs sliding matching within the candidate region of the infrared image data using a search window. The position corresponding to the maximum cross-correlation coefficient is used as the matching point, and the pixel displacement between the matching point and the gradient feature point is calculated as the basis for compensating the physical offset of the optical axis of the image acquisition module.

2. The platform end intrusion identification system based on multi-modal pre-fusion according to claim 1, characterized in that, The feature modulation module performs the logic of local suppression and feature gain compensation following the formula: wherein, is the modulated fused feature value, is the original feature value of the visible light image data, is the original feature value of the infrared image data, is the weight coefficient corresponding to the spatial confidence mask, is a preset infrared feature gain factor, and and are both dimensionless coefficients.

3. The platform end intrusion identification system based on multi-modal pre-fusion according to claim 1, characterized in that, After extracting multimodal contour features, the target detection module uses image connected component analysis logic to perform clustering processing on the pixels in the aligned and corrected heterogeneous fused feature map in order to extract the connected regions of candidate targets.

4. The platform end intrusion identification system based on multi-modal pre-fusion according to claim 1, characterized in that, The target detection module also includes a geometric constraint verification unit, which uses the disparity constraints generated by the installation baseline of the image acquisition module to perform spatial position verification on the connected regions of candidate targets and eliminate false targets that exceed the physical space of the platform.

5. The platform end intrusion identification system based on multi-modal pre-fusion according to claim 1, characterized in that, The spatial alignment module is also used to record historical data of cross-correlation offset vectors, calculate the fluctuation variance of historical data, and output maintenance warning instructions for the installation status of image acquisition equipment when the fluctuation variance exceeds a preset stability threshold.

6. The multi-modal pre-fusion based platform end intrusion identification system of claim 1, wherein, The system also includes a dynamic parameter update module, which performs a global average calculation on the local information entropy map to generate a global information entropy mean, and adjusts the confidence threshold of target determination in the target detection module in real time based on the global information entropy mean.

7. The multi-modal pre-fusion based platform end intrusion identification system of claim 1, wherein, When performing classification and recognition, the target detection module compares multimodal contour features with preset human skeleton structure features to determine whether there are personnel intruding into the end area of ​​the platform.

Citation Information

Patent Citations

  • Perimeter intrusion identification method and system based on video joint acquisition

    CN118298377B

  • Track foreign body detection method and system

    CN119763073A

  • Vehicle-mounted laser radar point cloud real-time target detection method and system

    CN120630151A