Method and system for intelligent identification of irregular behavior based on multi-algorithm fusion
By using a multi-algorithm fusion approach, combining image sensors and laser sensors, and dynamically adjusting the laser curtain wall area, and employing an integrated learning network for deep analysis, the problem of identifying "passing bags" behavior in subway security checks has been solved. This has enabled highly accurate and adaptive identification of violations, improving the automation level and security of security checks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CONGWEN SOFTWARE TECHNOLOGICAL SHENZHEN CITY
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-10
AI Technical Summary
In subway security checks, existing technologies struggle to accurately identify passengers' "passing bags" behavior. The identification accuracy is insufficient, and the robustness is poor in complex and crowded environments. There is a lack of efficient solutions for identifying specific violations, resulting in blind spots in security control.
By using a multi-algorithm fusion approach, image sensors are used to acquire personnel density information, and the laser curtain wall area is dynamically adjusted. Intrusion features are captured by laser sensors, and deep analysis is performed through an integrated learning multi-algorithm network. The recognition results are then compensated and processed to achieve highly accurate identification of violations.
It achieves highly accurate and adaptive identification of violations in complex environments, reduces false alarm rates, and improves the automation level and security control efficiency of the security inspection process.
Smart Images

Figure CN121582880B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image analysis, in particular to a method and system for intelligent identification of irregular behavior based on multi-algorithm fusion. BACKGROUND
[0002] Metro security is a key link to ensure the public safety of rail transit, aiming to intercept all kinds of dangerous goods. With the development of the city, the passenger flow of the subway continues to grow, and the phenomenon of crowded queues in front of the security check point is becoming increasingly common. Some passengers try to avoid the normal security check procedure for the sake of convenience, among which, the behavior of "passing bags" is particularly prominent. This behavior refers to the fact that passengers do not check the bags through the X-ray security inspection machine, but directly pass the bags to the fellow passenger on the other side of the security area from the side or above, causing the bags to escape the device scanning, so that the prohibited items hidden in the bags may be taken into the subway, constituting a serious safety hazard.
[0003] At present, the monitoring of the metro security area mainly relies on the combination of manual duty and video monitoring backtracking. The security officer needs to consider both the security machine screen interpretation and the on-site order maintenance, and it is easy to miss due to fatigue or limited vision during the peak passenger flow period. The existing intelligent monitoring technology based on video analysis mostly focuses on detecting general scenes such as overall crowd density, abnormal gathering, abandoned items, or violent movements of personnel. However, for the "passing bags" behavior, which has a specific pattern, relatively concealed actions, and a similar appearance to normal passing of items, the recognition accuracy is often insufficient, with high false positive and false negative rates. Specifically, the existing technology cannot accurately distinguish the subtle differences between passengers' normal placement of bags, retrieval of bags from the security machine, and illegal delivery of bags, and cannot stably track the delivery path of the bags and determine whether they bypass the core scanning area of the security inspection device in a complex and crowded on-site environment. The existing technology lacks a technical solution that efficiently and robustly identifies this specific irregular behavior, leaving a significant blind spot in the security control of the metro security area. SUMMARY
[0004] The present application provides a method and system for intelligent identification of irregular behavior based on multi-algorithm fusion to address the technical problems of insufficient recognition accuracy of specific irregular behavior in the metro security area, poor robustness in complex crowded environments, and lack of effective dynamic adaptability adjustment capability in the existing technology.
[0005] The technical solution of the present application to solve the above technical problems is as follows:
[0006] In a first aspect, the present application provides a method for intelligent identification of irregular behavior based on multi-algorithm fusion, comprising:
[0007] Obtaining a sequence of regional images of a target area collected by an image sensor, performing personnel density identification, and obtaining personnel density information;
[0008] According to the personnel density information, laser curtain area scale decision is performed to obtain laser curtain area scale, and laser sensor is controlled to perform laser curtain adjustment to form a laser curtain area.
[0009] When the laser sensor determines that there is laser curtain area intrusion, the intrusion time and intrusion features are obtained, and an intrusion area image set is called.
[0010] Interactive violation behavior identification is performed on the intrusion area image set to obtain interactive violation behavior identification results, and the intrusion features are combined to obtain final violation behavior identification results, wherein compensation processing is performed according to the personnel density information and the intrusion features.
[0011] In the second aspect, the application provides a violation behavior intelligent identification system based on multi-algorithm fusion, comprising:
[0012] An image acquisition and processing module is configured to obtain a region image sequence of a target region collected by an image sensor, perform personnel density identification, and obtain personnel density information.
[0013] A laser curtain control module is configured to perform laser curtain area scale decision according to the personnel density information, obtain laser curtain area scale, control laser sensor to perform laser curtain adjustment, and form a laser curtain area.
[0014] An intrusion detection module is configured to obtain intrusion time and intrusion features when the laser sensor determines that there is laser curtain area intrusion, and call an intrusion area image set.
[0015] A violation behavior identification module is configured to perform interactive violation behavior identification on the intrusion area image set to obtain interactive violation behavior identification results, combine the intrusion features, and process to obtain final violation behavior identification results, wherein compensation processing is performed according to the personnel density information and the intrusion features.
[0016] The application has the following beneficial effects:
[0017] Compared with the prior art, the present application obtains the personnel density in real time through image analysis, and dynamically adjusts the monitoring scale of the laser curtain wall accordingly, so that the system can adapt to the change of passenger flow, effectively reduces the false alarm caused by environmental interference while ensuring the detection accuracy. The timing and spatial characteristics of physical intrusion are accurately captured by the laser sensor, providing objective and reliable event basis for subsequent analysis. After the event is triggered, the system uses a multi-algorithm fusion model to perform in-depth analysis on the related images, integrates the judgment results of multiple recognition networks, and improves the recognition accuracy of specific hidden illegal behaviors such as "handing over packages". Finally, by introducing personnel density and intrusion characteristics to double compensate and correct the recognition results, the system optimizes the decision threshold under different crowdedness and behavior patterns, realizes high-precision and high-adaptive illegal behavior intelligent identification in complex subway security scenes, and improves the automation level and safety control efficiency of the security link. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 A flowchart of the illegal behavior intelligent identification method based on multi-algorithm fusion provided by the present application is shown.
[0019] Figure 2 A structure diagram of the illegal behavior intelligent identification system based on multi-algorithm fusion provided by the present application is shown.
[0020] In the drawings, the components represented by each reference numeral are as follows:
[0021] Image acquisition and processing module 11, laser curtain wall control module 12, intrusion detection module 13, illegal behavior identification module 14. DETAILED DESCRIPTION
[0022] In example one, as shown, the present application provides an illegal behavior intelligent identification method based on multi-algorithm fusion, which comprises: Figure 1 S10: Obtain the regional image sequence of the target area collected by the image sensor, perform personnel density identification, and obtain personnel density information;
[0023] Specifically, obtaining the regional image sequence of the target area collected by the image sensor, performing personnel density identification, and obtaining personnel density information comprises:
[0024] Taking the regional image sequence of the target area collected by the image sensor, wherein the target area includes a first region and a second region, and the first region and the second region are divided by a spacing line;
[0025] Inputting the regional image sequence into the personnel density identification plug-in and outputting the obtained personnel density information.
[0026]
[0027] Firstly, a sequence of regional images of the target region is acquired by a continuously capturing image sensor. The image sensor is a device that converts optical images into electronic signals, such as a surveillance camera deployed in a subway station. The image sensor is fixedly installed at a specific position above the security check area to continuously capture a surveillance video stream of the target region at a preset constant frame rate, thereby forming a sequence of regional images that are continuous and ordered in the time dimension. Specifically, the target region is the complete monitoring field of view covering the subway security check machine and its adjacent passageway, which is explicitly divided into two adjacent sub-regions, i.e., the first region and the second region, by a physical or virtual separation line. In the actual scene, the separation line can be embodied as a physical isolation fence used to separate the flow of people at the subway security check point, which clearly defines the pre-check passageway before security check and the post-check passageway after security check.
[0028] Therefore, each frame of image in the acquired sequence of regional images completely contains real-time scene information of the pre-check passageway and the post-check passageway, and can synchronously reflect the dynamic distribution and behavior of the people in the passageway.
[0029] Further, personnel density analysis is performed based on the acquired sequence of regional images. Specifically, the sequence of regional images containing the first region and the second region collected above is input into a personnel density recognition plug-in that is pre-trained in real time. The personnel density recognition plug-in maps a complex functional relationship from image features to personnel density, can analyze each frame of input image, automatically recognize and count the number or distribution characteristics of individuals appearing in the first region and the second region in the image, and finally output a quantitative or hierarchical personnel density information that can comprehensively reflect the real-time congestion degree of the pre-check passageway and the post-check passageway respectively or as a whole.
[0030] Specifically, the personnel density recognition plug-in is trained and configured in the control system by the following steps:
[0031] According to the historical monitoring data of the target region, a sample regional image set is collected;
[0032] The personnel density in each sample regional image is labeled to obtain a sample personnel density information set;
[0033] Based on a convolutional neural network, a personnel density recognition plug-in is constructed;
[0034] The sample regional image set and the sample personnel density information set are used to supervise the training and testing of the personnel density recognition plug-in, and after the testing converges, the plug-in is embedded into the control system.
[0035] First, according to the target area, that is, the historical monitoring data of the subway security channel, a large number of scene pictures containing different time periods and different passenger flow conditions are systematically collected to form a diversified sample area image set. This sample area image set needs to cover various personnel density scenes from sparse to crowded to ensure that the subsequent trained model has wide adaptability.
[0036] Secondly, data labeling is performed. Specifically, for each image in the sample area image set, a professional or an auxiliary labeling tool that has been verified is used to accurately quantify or objectively grade the personnel density presented in the image according to a pre-defined unified standard.
[0037] For example, the quantitative labeling method can be to mark each visible human body target in the image one by one and count the total number to obtain the accurate personnel count value of the frame image. In addition, the optional grading labeling method can be to divide the personnel density into three levels of "low density", "medium density" and "high density" according to pre-set threshold ranges, for example, defining the case of less than 10 people in the image as low density, the case of 10-30 people as medium density, and the case of more than 30 people as high density. Through any one or a combination of the above labeling methods, a sample personnel density information set corresponding to the sample area image set and containing accurate and real labels can be finally formed.
[0038] Further, a convolutional neural network is selected as the basic architecture to build a personnel density recognition plug-in. The convolutional neural network is a deep learning model specially designed for processing image data, which automatically extracts spatial hierarchical features in images through multiple layers of convolution and pooling operations, and is suitable for recognizing and locating human body targets from complex scenes.
[0039] Further, the sample area image set and the sample personnel density information set are used to perform end-to-end supervised training on the constructed convolutional neural network model. In this process, the model continuously adjusts its internal parameters to learn the mapping relationship from the input image to the corresponding personnel density label. After training, the model performance is evaluated using an independent test data set until it converges. The convergence condition is pre-set according to the requirements of accuracy and stability in actual application scenarios, for example, it can be set that the average absolute error fluctuation range of the model on the test set is less than 2% in the last 20 training cycles, and the final average absolute error value is less than 2 people. Finally, the trained model is solidified as a personnel density recognition plug-in, which is formally embedded into the software framework of the control system in the form of an application programming interface or a dynamic link library, thereby realizing the online and automatic personnel density analysis function of real-time video streams.
[0040] Exemplarily, since there is a high degree of complexity and nonlinearity in the visual association between the distribution, posture, occlusion and density evaluation of the personnel in the image, and the convolutional neural network has an advantage in automatically extracting deep spatial features of the image and processing such visual pattern recognition problems, the convolutional neural network is selected as the basic architecture to build the personnel density recognition plug-in.
[0041] Specifically, the main model of the personnel density recognition plug-in mainly consists of a feature extraction backbone network, a multi-scale feature fusion module and a density regression output layer. The input layer receives a single-frame area image that has been preprocessed by size normalization and pixel standardization. The feature extraction backbone network adopts a residual network structure pre-trained on the ImageNet dataset, for example, ResNet34, to utilize its powerful general feature extraction capability. The network layer parameters after that will be fine-tuned in the training process to adapt to the specific task. The multi-scale feature fusion module fuses the feature maps output by the backbone network at different stages through a spatial pyramid pooling structure to enhance the model's perception ability of personnel targets at different scales. The output layer adopts a fully connected layer, and a linear activation function is used to finally map the fused high-dimensional features to a continuous real value representing the personnel density. This value can be an accurate count or a normalized density level.
[0042] During the training process, the key hyperparameters include an initial learning rate set to 0.0001, a total number of training cycles set to 100, and a number of image samples input to the model per batch set to 16. The learning rate setting aims to ensure the stability of the pre-trained weight fine-tuning, the training cycle number setting ensures sufficient learning of the model on the task dataset, and the batch size selection needs to balance the computational resource limitation and the stability of gradient update. The specific training data is the sample area image set obtained by the foregoing collection, and the corresponding sample personnel density information set, which is randomly divided into a training set, a validation set and a test set in a ratio of 6:2:2.
[0043] Further, the image samples in the training set are input as input, and their corresponding personnel density label values are used as supervision signals. Through the back propagation algorithm, combined with the stochastic gradient descent optimizer, all weight parameters of the model are iteratively optimized. A smooth L1 loss function is used to measure the deviation between the model's predicted density value and the true label value. This loss function is relatively insensitive to outliers, which helps to stabilize the training. The training process is monitored by the validation set. When the validation set loss function value does not show a downward trend for 10 consecutive training cycles, and the average absolute error of the model on the validation set reaches the preset convergence standard, for example, less than 2 people, the training is terminated, thereby obtaining the converged personnel density recognition model.
[0044] Finally, the collected area image sequence is input into the personnel density recognition plug-in with completed configuration, and the personnel density information is obtained.
[0045] S20: performing laser curtain area scale decision according to the personnel density information, obtaining a laser curtain area scale, and controlling the laser sensor to perform laser curtain adjustment to form a laser curtain area;
[0046] Specifically, performing laser curtain area scale decision according to the personnel density information, obtaining a laser curtain area scale, and controlling the laser sensor to perform laser curtain adjustment to form a laser curtain area, comprises:
[0047] obtaining an average personnel density and a preset laser curtain area scale;
[0048] performing decision calculation adjustment on the preset laser curtain area scale according to the personnel density information and the average personnel density, and obtaining a laser curtain area scale;
[0049] controlling the laser sensor to perform laser curtain adjustment according to the laser curtain area scale to form a laser curtain area of the laser curtain area scale, wherein the laser curtain area comprises a first sub-laser curtain area and a second sub-laser curtain area falling into a first area and a second area divided by a spacing line.
[0050] First, two key reference parameters are obtained. The first is the average personnel density value obtained from historical data analysis or experience configuration, which represents the crowdedness of the target area in the normal or typical state. The second is the preset laser curtain area scale. The laser curtain area refers to an invisible detection area constructed by the emission and receiving devices of the laser sensor in the physical space on both sides of the spacing line in the security channel, representing the sensitive space range of the system actively detecting intrusion behavior. The preset laser curtain area scale is the initial monitoring physical boundary width reference value of the invisible detection area.
[0051] Secondly, the personnel density information obtained in real time is compared and analyzed with the aforementioned average personnel density, and the preset laser curtain area scale is dynamically adjusted, so as to output a laser curtain area scale suitable for the current flow condition.
[0052] For example, when the real-time personnel density is significantly higher than the average density, it indicates that the passage is more crowded, the personnel body movement space is limited, and the probability of unintentional body approaching the interval line during normal passage is increased. To avoid generating too many invalid alarms, the decision calculation may tend to appropriately reduce the laser curtain area size, i.e., narrow the width of the monitoring sensitive band, so that the system triggers only when the object is very close to the interval line, thereby realizing a more relaxed and more humanized monitoring strategy in a crowded environment. Conversely, when the real-time personnel density is lower than the average density, the passage is relatively empty, and the personnel movement space is large. At this time, the decision calculation may tend to increase the laser curtain area size, i.e., widen the width of the monitoring sensitive band, so that any object moving close to the interval line is kept highly vigilant, and a more stringent detection of potential illegal package delivery behavior is realized.
[0053] Specifically, according to the personnel density information and the average personnel density, the preset laser curtain area size is adjusted by decision calculation to obtain a laser curtain area size, including:
[0054] The ratio of the personnel density information and the average personnel density is calculated as a density adjustment coefficient;
[0055] The preset laser curtain area size is adjusted by decision calculation using the density adjustment coefficient to obtain a laser curtain area size.
[0056] First, the real-time personnel density information is taken as the numerator, and the preset average personnel density is taken as the denominator to calculate the ratio between the two, which is defined as the density adjustment coefficient.
[0057] For example, if the real-time personnel density is 45 people and the average personnel density is 30 people, the calculated density adjustment coefficient is 1.5. The density adjustment coefficient directly and quantitatively reflects the deviation of the current scene congestion degree from the normal level. When the density adjustment coefficient is greater than 1, it indicates that the current personnel flow is higher than the average level, and the environment is more crowded. When the density adjustment coefficient is less than 1, it indicates that the current personnel flow is sparse, which is lower than the average level.
[0058] Second, the density adjustment coefficient calculated is used for size adjustment decision. Specifically, the preset laser curtain area size is adjusted by decision calculation using the density adjustment coefficient. Optionally, the preset laser curtain area size is directly multiplied by the reciprocal of the density adjustment coefficient, thereby outputting the final executed laser curtain area size.
[0059] For example, when the density adjustment coefficient is greater than 1, indicating that the current personnel density is higher than the average level, the reciprocal is less than 1, so that the laser curtain area size obtained after multiplication is smaller than the preset laser curtain area size, thereby narrowing the monitoring range in a crowded environment and reducing the probability of false triggering.
[0060] Finally, according to the calculated laser curtain area size, the laser sensor is controlled to perform physical adjustment of the laser curtain. Specifically, the laser sensor is a non-contact distance and position detection device based on optical principles, which is deployed above the target area, i.e., the interval line of the subway security channel, to generate an array of invisible straight detection beams within a set spatial range to form an optical detection curtain. The laser sensor usually includes a pair of laser transmitter modules and laser receiver modules that work cooperatively. The control instruction can accurately adjust the beam emission angle of the laser transmitter module or the effective sensing range and direction of the laser receiver module, so as to accurately construct a laser curtain area in the physical space, which matches the width, position and calculated laser curtain area size parameters.
[0061] Specifically, the finally obtained laser curtain area completely corresponds to the target area defined by the image in terms of spatial layout, and is also divided into two parts by the physical interval line: one part of the beam array falls into the first area, i.e., the pre-inspection channel of the security check, to form a first sub-laser curtain area; the other part of the beam array falls into the second area, i.e., the post-inspection channel of the security check, to form a second sub-laser curtain area.
[0062] Therefore, on both sides of the security channel interval line, an invisible laser detection curtain with variable width is formed. When an object, such as a package or a person's arm, intrudes and blocks any sub-laser curtain area, the laser receiver will detect the interruption of the light path, thereby determining that an intrusion event has occurred.
[0063] S30: When the laser sensor determines that there is an intrusion into the laser curtain area, the intrusion time and intrusion characteristics are obtained, and the intrusion area image set is retrieved;
[0064] Specifically, when the laser sensor determines that there is an intrusion into the laser curtain area, the intrusion time and intrusion characteristics are obtained, and the intrusion area image set is retrieved, including:
[0065] When the laser sensor determines that there is an intrusion into the laser curtain area, the intrusion time and the intrusion sub-laser curtain area are obtained, wherein the intrusion sub-laser curtain area is the first sub-laser curtain area and the second sub-laser curtain area;
[0066] Continue to monitor the interactive intrusion time when another sub-laser curtain area is intruded, and calculate the time interval from the intrusion time as the intrusion characteristics, wherein if no intrusion into another sub-laser curtain area is monitored within a preset time range, it is determined that no intrusion into the laser curtain area has occurred, and the monitoring continues;
[0067] According to the intrusion time and the intrusion characteristics, the intrusion area image set is retrieved.
[0068] Firstly, when the laser sensor first detects that the light beam of any sub-laser curtain area is blocked, it immediately determines that an intrusion event has occurred, and needs to accurately record the time of this event, which is recorded as the intrusion time, and record the specific sub-laser curtain area that is triggered, that is, to determine whether the intrusion occurs in the first sub-laser curtain area or the second sub-laser curtain area, which is recorded as the intrusion sub-laser curtain area.
[0069] Since the typical illegal package delivery behavior is physically manifested as continuous interaction on both sides of the security interval line, that is, the behavior of one side delivering goods and the other side taking goods closely connects in time and space, this behavior pattern is expected to cause the first sub-laser curtain area and the second sub-laser curtain area to be triggered in succession in a short time. Based on this understanding, after the laser sensor first detects that an intrusion occurs in a sub-laser curtain area on one side and records the intrusion time, it will automatically start continuous monitoring within a preset time range.
[0070] Specifically, the preset time range refers to a limited time interval starting from the first intrusion time, and the specific value of this interval is pre-set according to the statistical analysis results of the maximum reasonable interval time of the two actions in the typical package delivery behavior pattern, for example, it can be set to 3 seconds.
[0071] Within the preset time range, the sub-laser curtain area on the opposite side is continuously monitored for subsequent intrusion events. If an intrusion is detected in the opposite sub-laser curtain area within the preset time range, the time of this subsequent intrusion is accurately recorded and defined as the interactive intrusion time. The absolute time difference between the first intrusion time and the interactive intrusion time is calculated, and this difference is defined as the intrusion feature of the current composite intrusion event. This intrusion feature parameter quantifies the close relationship between the two independent physical triggering events from the time dimension, and can provide a basis for subsequent judgment of whether the behavior constitutes a two-way interaction with spatial span.
[0072] On the contrary, if no intrusion event is detected in the opposite sub-laser curtain area within the preset time range, it is determined that there is only one isolated one-sided triggering. This situation usually corresponds to non-interactive scenarios such as pedestrians unknowingly approaching the fence, accidentally swinging their belongings, or small-range dropping of belongings, and its behavior pattern does not conform to the typical characteristics of illegal package delivery. Therefore, this one-sided triggering is determined not to constitute an effective laser curtain area intrusion threat, and the system returns to the state of continuous and synchronous monitoring of the two sub-laser curtain areas.
[0073] Finally, when an effective intrusion event and the corresponding intrusion feature are confirmed, according to the recorded intrusion time and the calculated intrusion feature, the intrusion area image set is retrieved from the video stream continuously recorded by the image sensor.
[0074] Specifically, according to the intrusion time and the intrusion feature, an intrusion region image set is retrieved, including:
[0075] The region image corresponding to the intrusion time and the region image corresponding to the interaction intrusion time within the intrusion feature after the intrusion time are retrieved to obtain the intrusion region image set.
[0076] The system locates and extracts one or more frames of region images corresponding to the first intrusion time from the video stream data continuously stored by the image sensor according to the recorded first intrusion time. Further, since the intrusion feature represents the time interval between the first intrusion and the subsequent interaction intrusion, the system will continue to retrieve the image sequence from the video stream based on the time interval parameter from the first intrusion time to the interaction intrusion time within a time interval, which includes one or more key images recording the interaction intrusion time.
[0077] Finally, by retrieving the images of the above two time points and the process therebetween, the obtained intrusion region image set completely covers the entire suspicious behavior time period from the single-sided first trigger to the interactive trigger on the opposite side. The intrusion region image set can provide a data basis for subsequent image analysis-based violation behavior identification.
[0078] S40: interactive violation behavior identification is performed on the intrusion region image set to obtain an interactive violation behavior identification result, and a final violation behavior identification result is processed in combination with the intrusion feature, wherein compensation processing is performed according to the personnel density information and the intrusion feature.
[0079] Specifically, interactive violation behavior identification is performed on the intrusion region image set to obtain an interactive violation behavior identification result, and a final violation behavior identification result is processed in combination with the intrusion feature, including:
[0080] A violation behavior identification network group constructed based on ensemble learning is obtained, wherein the violation behavior identification network group includes multiple violation behavior identification networks, the training data of each violation behavior identification network is not completely the same, the input feature in supervised training is a sample intrusion region image set, and the supervised label is a sample violation behavior identification result;
[0081] The intrusion region image set is input into the violation behavior identification network group, and multiple violation behavior identification results are output, wherein the violation behavior identification result includes violation or no violation;
[0082] The proportion of violations in the multiple violation behavior identification results is calculated to obtain a violation rate as an interactive violation behavior identification result;
[0083] According to the interactive violation behavior identification result and the intrusion feature, a final violation behavior identification result is processed.
[0084] First, a pre-constructed group of rule violation recognition networks is called. The group of rule violation recognition networks is constructed based on ensemble learning, and contains multiple rule violation recognition networks with the same or similar structures. Each rule violation recognition network is independently trained using different training data sets, for example, different subsets are extracted from the total samples by the bootstrap sampling method for training, so as to introduce the diversity of the model. In the training stage, the input of each rule violation recognition network is a sample intrusion region image set, and the supervised label is the sample rule violation recognition result of the corresponding image sequence determined by artificial judgment, for example, marked as “violation” or “non-violation”.
[0085] For example, in the subway security scene, the “handing over a bag” behavior is shown as the rapid transfer of objects across the physical interval line and the interaction with the hands of the personnel in the image sequence. Its visual pattern is complex, easily blocked and similar to the normal luggage taking and placing action, and there is a challenge that it is difficult to be stably recognized by a single model. Ensemble learning can effectively improve the generalization ability and decision robustness of the model by combining the prediction results of multiple base learners, so the rule violation recognition network group is constructed based on ensemble learning.
[0086] Specifically, the rule violation recognition network group is composed of multiple parallel rule violation recognition networks, each of which uses a three-dimensional convolutional neural network as a basic architecture to process the intrusion region image set with time dimension. The input layer receives the preprocessed and time-sampled image sequence. The network body contains multiple three-dimensional convolutional layers and three-dimensional pooling layers for jointly extracting motion and appearance features in the video clip from the space-time dimension. The fully connected layer is located at the end of the network, which integrates the extracted space-time features. The output layer uses the Softmax activation function to map the features to the probability distribution of the “violation” and “non-violation” two categories. Each independent rule violation recognition network remains consistent in architecture, but the difference in model parameters is obtained through a differentiated training process.
[0087] In the training process, the key hyperparameters are set as follows: the initial learning rate is 0.0005, the number of training iterations is 120, and the batch size is set to 8 image sequences. The setting of the learning rate takes into account the stability of the three-dimensional convolutional network training, the number of iterations ensures that the model learns the space-time pattern sufficiently, and the selection of the batch size adapts to the high demand of video data for video memory. The specific training data comes from a large number of sample intrusion region image sets intercepted from historical monitoring events, and each image set is attached with the sample rule violation recognition result determined by artificial judgment as the supervised label. Through the bootstrap sampling method, multiple overlapping but not completely identical subsets are randomly extracted from the total training data set.
[0088] Further, each subset of samples is independently used to train a rule violation recognition network. In the training of a single network, the sample image sequence in the sample subset is taken as the input, and the corresponding sample rule violation recognition result is taken as the supervision signal. The network parameters are iteratively optimized using the back propagation algorithm and the Adam optimizer. The cross-entropy loss is selected as the loss function to measure the difference between the network prediction probability distribution and the true label. After all the networks are independently trained, they collectively constitute a rule violation recognition network group. The rule violation recognition network group can learn the spatiotemporal features of the package delivery behavior from multiple slightly different data perspectives. Its collective decision mechanism effectively reduces the risk of overfitting to a single noise pattern or a specific scene, thereby exhibiting stronger generalization performance and more stable recognition accuracy when facing new intrusion event images.
[0089] Further, the current set of intrusion region images to be determined is simultaneously input into each rule violation recognition network in the rule violation recognition network group. Each rule violation recognition network independently performs forward reasoning and outputs a binary preliminary recognition result, i.e., determining whether the behavior exhibited by the current image sequence is "rule violation" or "non-rule violation". Second, multiple rule violation recognition results are summarized, the number of results determined as "rule violation" is counted, and the proportion of this number to the total number of the network group is calculated to obtain a quantitative rule violation rate. This rule violation rate reflects the collective decision-making tendency based on visual information and is defined as the interactive rule violation recognition result.
[0090] Finally, according to the obtained interactive rule violation recognition result and intrusion feature, the final rule violation recognition result is processed.
[0091] Specifically, according to the interactive rule violation recognition result and the intrusion feature, the final rule violation recognition result is processed, including:
[0092] Obtaining a benchmark intrusion feature of a rule violation;
[0093] Calculating the similarity between the intrusion feature and the benchmark intrusion feature to obtain an intrusion rule violation rate;
[0094] According to the intrusion rule violation rate and the rule violation rate in the interactive rule violation recognition result, a fusion rule violation rate is calculated;
[0095] Based on the personnel density information and the average personnel density, the fusion rule violation rate is compensated to obtain a compensated rule violation rate as the final rule violation recognition result.
[0096] Firstly, the benchmark intrusion feature of the violation behavior is obtained. The benchmark intrusion feature is a quantitative time interval value representing the expected time span from the first laser curtain trigger on one side of the security channel to the interactive trigger on the other side corresponding to a typical violation behavior verified in historical data. The benchmark intrusion feature is obtained by statistically analyzing the laser trigger time interval data in a large number of confirmed violation event samples and calculating the central tendency, such as the arithmetic mean or median.
[0097] Secondly, the actual intrusion feature obtained by the current actual monitoring, i.e. the actual time interval between the first trigger and the interactive trigger, is compared with the above-mentioned benchmark intrusion feature. The similarity between the two is calculated to quantify the consistency of the time pattern of the current event with the typical violation time pattern. For example, a similarity calculation method based on a Gaussian function can be used, in which the benchmark intrusion feature is the center of the Gaussian distribution. The closer the actual time interval is to the center benchmark value, the higher the calculated similarity value will be. This similarity value is directly defined as the violation probability based on the time feature, called the intrusion violation rate. The intrusion violation rate reflects the degree of agreement between the spatio-temporal behavior pattern of the current event and the typical violation pattern.
[0098] Further, the obtained intrusion violation rate and the violation rate in the interactive violation behavior recognition result obtained from the violation behavior recognition network group are comprehensively calculated by a predetermined fusion algorithm. For example, the fusion algorithm can be a weighted average, i.e. the visual violation rate and the time violation rate are added after being assigned different weights. The calculation formula of the fusion violation rate can be represented as: fusion violation rate = a x visual violation rate + (1-a) x intrusion violation rate, where a is the weight coefficient of the visual violation rate, which is pre-set according to the historical accuracy and confidence of the visual recognition model on the independent test set, for example, a can be set to 0.7. Through this calculation, a fusion violation rate considering both visual evidence and spatio-temporal behavior evidence can be obtained.
[0099] Finally, considering the potential impact of the density of people on the reliability of behavior recognition, the above-mentioned fusion violation rate needs to be further compensated by using the real-time acquired people density information and the pre-set average people density.
[0100] For example, when the real-time people density is much higher than the average density, the on-site crowd may cause the image clarity to decrease and the behaviors to be mutually occluded, at this time the recognition confidence based on the image will decrease, and the system can appropriately down-regulate the fusion violation rate through a compensation function to avoid the false positive tendency that may be caused by the decrease of image quality. Conversely, in a sparse crowd, the confidence in the fusion result can be maintained or enhanced. The value obtained after this environmental adaptive compensation calculation is called the compensated violation rate. Specifically, compensated violation rate = fusion violation rate x (average people density / real-time people density).
[0101] When the real-time person density is equal to the average person density, the ratio is 1, the compensation violation rate is equal to the fusion violation rate, and the system does not perform environmental compensation adjustment. When the real-time person density is higher than the average person density, the ratio is less than 1, the compensation violation rate is lower than the fusion violation rate, and the system moderately reduces the final violation probability based on the judgment of environmental congestion to inhibit the false alarm tendency caused by the decline in image quality and the mutual occlusion of behaviors. When the real-time person density is lower than the average person density, the ratio is greater than 1, the compensation violation rate is higher than the fusion violation rate, and the system maintains or enhances the confidence in the fusion result based on the judgment of environmental laxity, thereby improving the detection strictness and early warning sensitivity to potential violation behaviors.
[0102] The compensation violation rate integrates three types of information, namely visual analysis, spatiotemporal behavior pattern and environmental congestion degree, and is output as the final, more robust violation behavior recognition result.
[0103] In summary, the embodiments of the present application have at least the following technical effects:
[0104] Compared with the prior art, the present application first perceives the person density in the security check area through the image sensor in real time, and dynamically decides and adjusts the monitoring range of the laser curtain according to the density information, so that the physical detection boundary can adapt to the change in the flow of people, avoids false triggering in congestion, and improves the sensitivity in laxity, thereby realizing the intelligentization and flexibility of the monitoring strategy. Secondly, the laser curtain is used as a high-precision and low-delay first-level triggering mechanism, which can accurately capture the intrusion event and its accurate timing characteristics when the package crosses the preset physical boundary, thereby providing reliable event starting point and key parameters for subsequent analysis.
[0105] Thirdly, after the laser triggering, the system retrieves the local image sequence associated with the event, and uses a multi-algorithm fusion network based on ensemble learning to perform deep behavior analysis on the image, thereby improving the robustness and accuracy of the identification of specific violation behaviors such as package delivery by integrating the judgment results of multiple models, and effectively overcoming the limitations of a single algorithm in complex scenes. Finally, the person density information and the intrusion feature are innovatively introduced into the final decision-making link to perform dual compensation correction on the preliminary identification result in terms of environment and behavior pattern, thereby further reducing the misjudgment and omission caused by scene congestion and high behavior similarity, and finally realizing efficient, accurate and environment-adaptive intelligent identification of violation package delivery behaviors in the security check area of the subway.
[0106] Embodiment two, as Figure 2 shown, based on the same inventive concept of the multi-algorithm fusion-based violation behavior intelligent identification method provided in embodiment one, the present application embodiment further provides a multi-algorithm fusion-based violation behavior intelligent identification system, which comprises:
[0107] The image acquisition and processing module 11 is configured to acquire a region image sequence of a target region collected by an image sensor, perform personnel density recognition, and obtain personnel density information.
[0108] The laser curtain control module 12 is configured to make a laser curtain region scale decision according to the personnel density information, obtain a laser curtain region scale, control a laser sensor to adjust the laser curtain, and form a laser curtain region.
[0109] The intrusion detection module 13 is configured to acquire an intrusion time and an intrusion feature when the laser sensor determines that there is an intrusion in the laser curtain region, and call an intrusion region image set.
[0110] The violation behavior recognition module 14 is configured to perform interactive violation behavior recognition on the intrusion region image set, obtain an interactive violation behavior recognition result, and process a final violation behavior recognition result in combination with the intrusion feature, wherein compensation processing is performed according to the personnel density information and the intrusion feature.
[0111] The image acquisition and processing module 11 is configured to:
[0112] The image acquisition and processing module 11 is configured to:
[0113] The image acquisition and processing module 11 is configured to:
[0114] The image acquisition and processing module 11 is configured to:
[0115] Specifically, the personnel density recognition plug-in is trained and configured in the control system by the following steps:
[0116] According to historical monitoring data of the target region, a sample region image set is collected.
[0117] The personnel density in each sample region image is labeled to obtain a sample personnel density information set.
[0118] Based on a convolutional neural network, a personnel density recognition plug-in is constructed.
[0119] The personnel density recognition plug-in is supervised and trained and tested by using the sample region image set and the sample personnel density information set, and after convergence in the test, the personnel density recognition plug-in is embedded into the control system.
[0120] The laser curtain control module 12 is configured to:
[0121] According to the personnel density information, a laser curtain area scale decision is made to obtain a laser curtain area scale, and a laser sensor is controlled to adjust the laser curtain to form a laser curtain area, including:
[0122] An average personnel density and a preset laser curtain area scale are obtained.
[0123] According to the personnel density information and the average personnel density, the preset laser curtain area scale is adjusted to obtain a laser curtain area scale.
[0124] According to the laser curtain area scale, the laser sensor is controlled to adjust the laser curtain to form a laser curtain area of the laser curtain area scale, wherein the laser curtain area includes a first sub-laser curtain area and a second sub-laser curtain area falling into a first area and a second area divided by a spacing line.
[0125] Specifically, according to the personnel density information and the average personnel density, the preset laser curtain area scale is adjusted to obtain a laser curtain area scale, including:
[0126] The ratio of the personnel density information and the average personnel density is calculated as a density adjustment coefficient.
[0127] The density adjustment coefficient is used to adjust the preset laser curtain area scale to obtain a laser curtain area scale.
[0128] The intrusion detection module 13 is specifically configured to:
[0129] When the laser sensor determines that there is an intrusion into the laser curtain area, the intrusion time and the intrusion feature are obtained, and an intrusion area image set is called, including:
[0130] When the laser sensor determines that there is an intrusion into the laser curtain area, the intrusion time and the intrusion sub-laser curtain area are obtained, wherein the intrusion sub-laser curtain area is the first sub-laser curtain area and the second sub-laser curtain area.
[0131] The interaction intrusion time of another sub-laser curtain area intrusion is continuously monitored and obtained, and the time interval from the intrusion time is calculated as the intrusion feature, wherein if another sub-laser curtain area intrusion is not monitored within a preset time range, it is determined that there is no laser curtain area intrusion, and the monitoring continues.
[0132] According to the intrusion time and the intrusion feature, the intrusion area image set is called.
[0133] Specifically, according to the intrusion time and the intrusion feature, the intrusion area image set is called, including:
[0134] The region image corresponding to the intrusion moment and the region image of the interaction intrusion moment after the intrusion moment in the intrusion feature are obtained to obtain an intrusion region image set.
[0135] The violation behavior recognition module 14 is specifically configured to:
[0136] The intrusion region image set is subjected to interactive violation behavior recognition to obtain an interactive violation behavior recognition result, and the intrusion feature is combined to process a final violation behavior recognition result, including:
[0137] An integrated learning-based violation behavior recognition network group is obtained, wherein the violation behavior recognition network group includes multiple violation behavior recognition networks, the training data of each violation behavior recognition network is not completely the same, the input feature in the supervised training is a sample intrusion region image set, and the supervised label is a sample violation behavior recognition result.
[0138] The intrusion region image set is input into the violation behavior recognition network group to output multiple violation behavior recognition results, wherein the violation behavior recognition result includes violation or no violation.
[0139] The proportion of violation in the multiple violation behavior recognition results is calculated to obtain a violation rate as an interactive violation behavior recognition result.
[0140] According to the interactive violation behavior recognition result and the intrusion feature, a final violation behavior recognition result is processed.
[0141] Specifically, according to the interactive violation behavior recognition result and the intrusion feature, a final violation behavior recognition result is processed, including:
[0142] A reference intrusion feature of a violation behavior is obtained.
[0143] The similarity of the intrusion feature and the reference intrusion feature is calculated to obtain an intrusion violation rate.
[0144] According to the intrusion violation rate and the violation rate in the interactive violation behavior recognition result, a fusion violation rate is calculated.
[0145] Based on the personnel density information and the average personnel density, the fusion violation rate is compensated to obtain a compensated violation rate as the final violation behavior recognition result.
[0146] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for intelligent identification of irregular behavior based on multi-algorithm fusion, characterized in that, The method comprises: acquiring a region image sequence of a target region collected by an image sensor, performing personnel density recognition to obtain personnel density information; based on the personnel density information, performing laser curtain region scale decision to obtain a laser curtain region scale, and controlling a laser sensor to perform laser curtain adjustment to form a laser curtain region; when the laser sensor determines that there is laser curtain region intrusion, acquiring an intrusion time and an intrusion feature, and calling an intrusion region image set, comprising: when the laser sensor determines that there is laser curtain region intrusion, acquiring an intrusion time and an intrusion sub-laser curtain region, wherein the intrusion sub-laser curtain region is a first sub-laser curtain region and a second sub-laser curtain region; continuously monitoring to acquire an interactive intrusion time of another sub-laser curtain region intrusion, calculating a time interval from the intrusion time as an intrusion feature, wherein if no intrusion of another sub-laser curtain region is monitored within a preset time range, it is determined that no laser curtain region intrusion has occurred, and monitoring is continued; based on the intrusion time and the intrusion feature, calling an intrusion region image set, comprising: calling a region image corresponding to the intrusion time, and a region image of an interactive intrusion time within the intrusion feature after the intrusion time, to obtain an intrusion region image set; performing interactive violation behavior recognition on the intrusion region image set to obtain an interactive violation behavior recognition result, and processing to obtain a final violation behavior recognition result in combination with the intrusion feature, wherein compensation processing is performed based on the personnel density information and the intrusion feature. 2.The method of claim 1, wherein, acquiring a region image sequence of a target region collected by an image sensor, performing personnel density recognition to obtain personnel density information, comprising: acquiring a region image sequence of a target region collected by an image sensor, wherein the target region comprises a first region and a second region, and the first region and the second region are divided by a spacing line; inputting the region image sequence into a personnel density recognition plug-in to output and obtain personnel density information. 3.The method of claim 2, wherein, The personnel density recognition plug-in is configured in a control system by the following steps: based on historical monitoring data of the target region, a sample region image set is collected; the personnel density in each sample region image is labeled to obtain a sample personnel density information set; based on a convolutional neural network, a personnel density recognition plug-in is constructed; the sample region image set and the sample personnel density information set are used to supervise training and testing of the personnel density recognition plug-in, and after convergence in testing, the personnel density recognition plug-in is embedded into the control system.
4. The method of claim 1, wherein the method is characterized by, based on the personnel density information, performing laser curtain region scale decision to obtain a laser curtain region scale, and controlling a laser sensor to perform laser curtain adjustment to form a laser curtain region, comprising: acquiring an average personnel density and a preset laser curtain region scale; based on the personnel density information and the average personnel density, performing decision calculation adjustment on the preset laser curtain region scale to obtain a laser curtain region scale; According to the laser curtain area scale, the laser sensor is controlled to adjust the laser curtain to form a laser curtain area of the laser curtain area scale, wherein the laser curtain area includes first and second sub-laser curtain areas divided by the interval line and falling into the first and second areas.
5. The method of claim 4, wherein the method is based on multi-algorithm fusion. According to the personnel density information and the average personnel density, the preset laser curtain area scale is adjusted by decision calculation to obtain a laser curtain area scale, including: The ratio of the personnel density information and the average personnel density is calculated as a density adjustment coefficient; The preset laser curtain area scale is adjusted by decision calculation using the density adjustment coefficient to obtain a laser curtain area scale. 6.The method of claim 1, wherein, Interactive violation behavior identification is performed on the intrusion area image set to obtain an interactive violation behavior identification result, and the final violation behavior identification result is processed in combination with the intrusion feature, including: An integrated learning-based violation behavior identification network group is obtained, wherein the violation behavior identification network group includes multiple violation behavior identification networks, the training data of each violation behavior identification network is not completely the same, the input feature in supervised training is a sample intrusion area image set, and the supervised label is a sample violation behavior identification result; The intrusion area image set is input into the violation behavior identification network group, and multiple violation behavior identification results are output, wherein the violation behavior identification result includes violation or no violation; The proportion of violations in the multiple violation behavior identification results is calculated to obtain a violation rate as an interactive violation behavior identification result; According to the interactive violation behavior identification result and the intrusion feature, the final violation behavior identification result is processed.
7. The method according to claim 6, wherein, According to the interactive violation behavior identification result and the intrusion feature, the final violation behavior identification result is processed, including: A reference intrusion feature of a violation behavior is obtained; The similarity between the intrusion feature and the reference intrusion feature is calculated to obtain an intrusion violation rate; According to the intrusion violation rate and the violation rate in the interactive violation behavior identification result, a fusion violation rate is calculated; Based on the personnel density information and the average personnel density, the fusion violation rate is compensated to obtain a compensated violation rate as the final violation behavior identification result.
8. The intelligent identification system of violation behavior based on multi-algorithm fusion, characterized in that, A multi-algorithm fusion-based violation behavior intelligent identification method according to any one of claims 1-7, including: An image acquisition and processing module for acquiring a region image sequence of a target region collected by an image sensor, performing personnel density identification, and obtaining personnel density information; A laser curtain control module for performing laser curtain area scale decision according to the personnel density information to obtain a laser curtain area scale, controlling a laser sensor to adjust the laser curtain, and forming a laser curtain area; An intrusion detection module for acquiring an intrusion time and an intrusion feature when the laser sensor determines that there is an intrusion into the laser curtain area, and calling an intrusion area image set, including: When the laser sensor determines that there is an intrusion into the laser curtain area, an intrusion time and an intrusion sub-laser curtain area are acquired, wherein the intrusion sub-laser curtain area is the first and second sub-laser curtain areas. Continue to monitor to obtain another sub-laser curtain area intrusion interaction intrusion time, calculate the time interval with the intrusion time as the intrusion feature, wherein if no other sub-laser curtain area intrusion is monitored within a preset time range, it is judged that no laser curtain area intrusion occurs, and the monitoring continues; According to the intrusion time and the intrusion feature, the intrusion area image set is called, including: calling the area image corresponding to the intrusion time, and the area image of the interactive intrusion time within the intrusion feature after the intrusion time, obtaining the intrusion area image set; The violation behavior recognition module is used for interactive violation behavior recognition on the intrusion area image set, obtaining interactive violation behavior recognition result, combining the intrusion feature, and processing to obtain the final violation behavior recognition result, wherein compensation processing is performed according to the personnel density information and the intrusion feature.
Citation Information
Patent Citations
Detection model training method, detection method and related device
CN111091098A
Target area monitoring method and device and storage medium
CN120802383A
Radar monitoring method and device, electronic equipment and storage medium
CN121299605A