A method and device for identifying abnormal wandering persons in a subway car
By installing multiple cameras in the subway car, images of abnormal wandering people are identified and filtered out, and combined with abnormal behavior recognition model and visual language model, the problem of difficult monitoring of abnormal wandering behavior in the subway car is solved, and efficient and accurate abnormal recognition and monitoring is achieved.
Patent Information
- Application Number
- CN202411377663.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-30
AI Technical Summary
Abnormal hesitation behaviors in the subway car are difficult to monitor and warn in real time, affecting operational safety and passenger experience.
By installing multiple cameras in each car, video images are obtained and recognition processing are performed, target personnel characteristics are extracted, and images of wandering personnel that meet the preset wandering time and space conditions are selected. Combined with an abnormal behavior recognition model and a visual language model, the abnormal hovering recognition results are determined.
Accurate identification and real-time monitoring of abnormal wandering people in subway cars is achieved, operational safety and passenger experience are improved, human interference is reduced, and the objectivity and accuracy of detection results are enhanced.
Smart Images

Figure CN118887707B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method and device for identifying abnormally wandering persons in a subway car. Background Art
[0002] As an important part of urban public transportation, the subway carries a large number of people, and its safety is very important. Sometimes people walk between multiple carriages, wander back and forth, and communicate with other passengers. These behaviors are considered abnormal wandering behaviors on the subway and may be accompanied by potential risks. In addition, subway carriages are closed public spaces, and passengers need a quiet and orderly environment to complete their commutes. The presence of abnormal wandering people will interfere with the normal riding experience of other passengers and disrupt the order in the carriage.
[0003] Therefore, it is necessary to use technical means to monitor and warn of these abnormal wandering and abnormal behaviors in real time, so as to help operators take timely measures and reduce operational interference caused by such incidents. Summary of the invention
[0004] In view of this, the purpose of the present application is to provide a method and device for identifying people wandering abnormally in a subway car, which can accurately identify people wandering abnormally on the subway, so as to monitor and warn of abnormal wandering and abnormal behavior in the subway car in real time.
[0005] The embodiment of the present application provides a method for identifying abnormally wandering persons in a subway car, wherein a camera is installed in each car of the subway; the identification method comprises:
[0006] Acquire video images captured by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the multiple cameras have different shooting ranges;
[0007] By performing recognition processing on the video images captured by the camera of each carriage, a moving person image is recognized from the video image, and the target person features corresponding to the camera are extracted based on the moving person image;
[0008] Processing the target person features corresponding to different cameras within the second preset time period, and screening out wandering person images that meet the preset wandering time condition and the preset wandering space condition from the moving person images of different cameras;
[0009] The screened wandering person images are processed to determine abnormal behavior recognition results of the wandering person in the wandering person images, and abnormal wandering recognition results are determined.
[0010] In some embodiments, in the method for identifying abnormal wandering persons in a subway car, the processing of the screened wandering person images, determining the abnormal behavior recognition result of the wandering person in the wandering person images, and determining the abnormal wandering recognition result include:
[0011] Inputting the screened wandering person images into a trained abnormal behavior recognition model, the abnormal behavior recognition model processes the wandering person images, and outputs the confidence level of the abnormal behavior of the wandering person in the wandering person images;
[0012] When the abnormal behavior confidence exceeds a first preset threshold, determining that the abnormal wandering identification result is that there is an abnormal wandering person in the subway car;
[0013] When the abnormal behavior confidence is less than a second preset threshold, determining that the abnormal wandering identification result is that there is no abnormal wandering person in the subway car;
[0014] When the abnormal behavior confidence is greater than or equal to the second preset threshold and less than or equal to the first preset threshold, the video image corresponding to the wandering person image and the pre-set reference question are input into the trained visual language model, so that the visual language model recognizes the image information of the wandering person image, and outputs the reference answer to the reference question as the abnormal wandering identification result based on the recognized image information.
[0015] In some embodiments, in the method for identifying abnormal wandering persons in a subway car, the trained visual language large model is obtained by training using the following method:
[0016] Identify the target weight matrix that needs to be fine-tuned from the existing large visual language model;
[0017] Introducing a first low-rank matrix and a second low-rank matrix, and training the first low-rank matrix and the second low-rank matrix based on a sample training set; the sample training set is obtained based on multiple scene images of a carriage;
[0018] The product of the trained first low-rank matrix and the second low-rank matrix is superimposed on the target weight matrix to obtain a large visual language model after the target weight matrix is adjusted.
[0019] In some embodiments, in the method for identifying abnormally wandering persons in a subway car, the step of performing recognition processing on the video images captured by the camera of each car to identify the images of the moving persons from the video images includes:
[0020] Obtaining the door opening time and door closing time of the carriage, and filtering the video images captured by the camera of each carriage based on the time difference between the shooting time of the video image and the door opening time or door closing time, to obtain a filtered video image;
[0021] Inputting the video image after the first screening into a mobile person recognition model to recognize images of mobile persons within the shooting range of the camera;
[0022] The images of mobile personnel within the shooting range of the camera are identified and matched with preset staff features, and the video images are secondary screened based on the matching results to obtain secondary screened images of mobile personnel.
[0023] In some embodiments, in the method for identifying abnormally wandering persons in a subway car, extracting the target person features corresponding to the camera based on the moving person image includes:
[0024] Based on multiple image quality dimensions, face capture images and / or body capture images meeting preset quality conditions are screened out from multiple images of moving persons captured by the camera;
[0025] Process the facial capture image based on the trained facial feature extraction model to obtain facial features;
[0026] Process the human body capture based on the trained human body feature extraction model to obtain human body features;
[0027] Based on the facial features and body features, the target person features corresponding to the camera are obtained.
[0028] In some embodiments, in the method for identifying abnormally wandering persons in a subway car, the step of processing a human body capture image based on a trained human body feature extraction model to obtain human body features includes:
[0029] Processing human body captures based on the trained human body feature extraction model to extract posture features of three granularities; the three granularities include: global granularity, upper and lower granularity, and upper, middle and lower granularity; different granularities correspond to different body regions; the global granularity represents the entire body region, the upper and lower granularity represents dividing the entire body region into an upper body region and a lower body region, and the upper, middle and lower granularity represents dividing the entire body region into three body regions from top to bottom;
[0030] Combining the posture features of three granularities corresponding to the target body area in the human body capture image to obtain the combined posture features corresponding to the target body area;
[0031] The combined posture features and the posture features of three granularities are used as human body features.
[0032] In some embodiments, in the method for identifying abnormal wandering persons in a subway car, the processing of target person features corresponding to different cameras within a second preset time period, and screening out wandering person images that meet preset wandering time conditions and preset wandering space conditions from moving person images of different cameras, includes:
[0033] Taking each target person feature corresponding to different cameras in the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result;
[0034] Based on the clustering results, determining the shooting time and camera of the mobile person image corresponding to the target person feature of each cluster of the clustering results;
[0035] Based on the shooting time and camera of the mobile person images corresponding to the target person features in each cluster, the wandering person images meeting the preset wandering time condition and the preset wandering space condition are screened out from the mobile person images of different cameras.
[0036] In some embodiments, in the method for identifying abnormally wandering persons in a subway car, the target person features include facial features and body features;
[0037] The method of taking each target person feature corresponding to different cameras within the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result includes:
[0038] Taking each target person feature of different cameras in the second preset time period as a node, determining whether both nodes have facial features;
[0039] If so, calculate the similarity of facial features as the similarity between the two nodes;
[0040] If not, the similarity of human features is calculated as the similarity between the two nodes;
[0041] Based on the similarity between any two nodes, the nodes are clustered to obtain a clustering result.
[0042] In some embodiments, in the method for identifying abnormal wandering persons in a subway car, clustering the nodes based on the similarity between any two nodes to obtain a clustering result includes:
[0043] Based on the human feature similarity between two nodes, when clustering the nodes, when the human feature similarity between a first node and a plurality of second nodes is greater than a preset human feature similarity threshold, the time interval, trajectory direction, and camera distance of the moving person image corresponding to the human features of the first node and each second node are determined; the camera distance is the difference in the serial numbers of the cameras of the first node and the second node;
[0044] Determine the trajectory direction of the first node and each second node, the camera distance and the priority matching result of the preset wandering rules of different priorities; wherein the preset wandering rules of different priorities correspond to different weights, and the preset wandering rules with higher priorities have higher weights;
[0045] Based on the priority matching result and the weights corresponding to the preset wandering rules of different priorities, the human feature similarities of the first node and the plurality of second nodes are updated, and the clustering result is determined based on the updated human feature similarities.
[0046] In some embodiments, a device for identifying abnormally wandering persons in a subway car is further provided, wherein a camera is installed in each car of the subway; the identification device comprises:
[0047] An acquisition module is used to acquire video images captured by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the shooting ranges of the multiple cameras are different;
[0048] An extraction module is used to identify and process the video images taken by the camera of each carriage, identify the images of moving persons from the video images, and extract the target person features corresponding to the camera based on the images of moving persons;
[0049] a processing module, used for processing the target person features corresponding to different cameras within a second preset time period, and screening out images of wandering persons that meet preset wandering time conditions and preset wandering space conditions from the images of moving persons captured by different cameras;
[0050] The determination module is used to process the screened wandering person images, determine the abnormal behavior recognition result of the wandering person in the wandering person images, and determine the abnormal wandering recognition result.
[0051] In an embodiment of the present application, a method and device for identifying abnormal wandering persons in a carriage are provided. The identification method obtains video images taken by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the shooting ranges of the multiple cameras are different; by performing recognition processing on the video images taken by the camera in each carriage, a moving person image is identified from the video image, and the target person features corresponding to the camera are extracted based on the moving person image; the target person features corresponding to different cameras within a second preset time period are processed, and wandering person images that meet preset wandering time conditions and preset wandering space conditions are screened out from the moving person images of different cameras; the screened wandering person images are processed to determine the abnormal behavior recognition result of the wandering person in the wandering person images, and the abnormal wandering recognition result is determined; thereby, by detecting the target image frame of a single camera, and The mobile personnel are identified in real time, and then the wandering personnel are determined based on the characteristics of abnormal wandering behavior in time and space, and finally it is determined whether the wandering personnel make specific abnormal behaviors, and finally the abnormal wandering personnel are identified; in this way, through the real-time detection of the target in the image frame, the analysis of the time characteristics and spatial characteristics, and the identification of specific abnormal behaviors, a multi-dimensional and multi-level analysis of abnormal wandering personnel is realized, and the characteristics of abnormal behaviors can be captured more comprehensively, thereby improving the accuracy of detection, and taking into account the amount of calculation; with the help of image processing technology and artificial intelligence algorithms, intelligent processing of surveillance videos is realized, the processing speed is improved, and the interference of human factors is reduced, making the detection results more objective and accurate; finally, real-time detection enables an alarm to be issued immediately or other measures to be taken once an abnormal wandering person is found, so that it is convenient for staff to stop the occurrence of abnormal behaviors in time and ensure the safety and order of subway cars. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0053] Figure 1 A flow chart showing a method for identifying an abnormally wandering person in a subway car according to an embodiment of the present application is shown;
[0054] Figure 2 A flow chart of a method for identifying an image of a moving person from a video image according to an embodiment of the present application is shown;
[0055] Figure 3A flow chart of a method for extracting target person features corresponding to a camera based on the mobile person image according to an embodiment of the present application is shown;
[0056] Figure 4 A flow chart of a method for processing target person features corresponding to different cameras within a second preset time period according to an embodiment of the present application is shown;
[0057] Figure 5 A flow chart of a method for processing the screened wandering person image according to an embodiment of the present application is shown;
[0058] Figure 6 A schematic diagram of the structure of a device for identifying abnormally wandering persons in a subway car according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0059] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.
[0060] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0061] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0062] As an important part of urban public transportation, the subway carries a large number of people, and its safety is very important. Sometimes people walk between multiple carriages, wander back and forth, and communicate with other passengers. These behaviors are considered abnormal wandering behaviors on the subway and may be accompanied by potential risks. In addition, subway carriages are closed public spaces, and passengers need a quiet and orderly environment to complete their commutes. The presence of abnormal wandering people will interfere with the normal riding experience of other passengers and disrupt the order in the carriage.
[0063] Based on this, in an embodiment of the present application, a method and device for identifying abnormal wandering persons in a carriage are provided, wherein the identification method obtains video images taken by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the shooting ranges of the multiple cameras are different; by performing recognition processing on the video images taken by the camera in each carriage, images of moving persons are identified from the video images, and target person features corresponding to the camera are extracted based on the moving person images; target person features corresponding to different cameras within a second preset time period are processed, and images of wandering persons that meet preset wandering time conditions and preset wandering space conditions are screened out from the moving person images of different cameras; the screened wandering person images are processed to determine the abnormal behavior recognition results of the wandering persons in the wandering person images, and the abnormal wandering recognition results are determined; thereby, by detecting the target image frames of a single camera, the target person features corresponding to different cameras within a second preset time period are screened out, and the wandering person images that meet the preset wandering time conditions and preset wandering space conditions are screened out; the screened wandering person images are processed to determine the abnormal behavior recognition results of the wandering persons in the wandering person images, and the abnormal wandering recognition results are determined. , timely identify mobile personnel, and then determine the wandering personnel based on the characteristics of abnormal wandering behavior in time and space, and finally determine whether the wandering personnel have made specific abnormal behaviors, and finally identify the abnormal wandering personnel; in this way, through real-time detection of targets in image frames, analysis of time characteristics and spatial characteristics, and identification of specific abnormal behaviors, multi-dimensional and multi-level analysis of abnormal wandering personnel is realized, which can capture the characteristics of abnormal behaviors more comprehensively, thereby improving the accuracy of detection and taking into account the amount of calculation; with the help of image processing technology and artificial intelligence algorithms, intelligent processing of surveillance videos is realized, the processing speed is improved, and the interference of human factors is reduced, making the detection results more objective and accurate; finally, real-time detection enables an alarm to be issued immediately or other measures to be taken once an abnormal wandering person is found, so that it is convenient for staff to stop the occurrence of abnormal behaviors in time and ensure the safety and order of subway cars.
[0064] Please refer to Figure 1 , Figure 1 The flowchart of the method for identifying abnormal wandering persons in a subway car according to an embodiment of the present application is shown, wherein a camera is installed in each car of the subway; Figure 1 As shown, the identification method includes the following steps S101-S104:
[0065] S101, obtaining video images captured by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the shooting ranges of the multiple cameras are different;
[0066] S102, by performing recognition processing on the video images captured by the camera of each carriage, identifying the images of moving persons from the video images, and extracting the target person features corresponding to the camera based on the images of moving persons;
[0067] S103, processing the target person features corresponding to different cameras within the second preset time period, and screening out wandering person images that meet preset wandering time conditions and preset wandering space conditions from the moving person images of different cameras;
[0068] S104: Process the screened wandering person images, determine abnormal behavior recognition results of the wandering persons in the wandering person images, and determine abnormal wandering recognition results.
[0069] In step S101, video images captured by a camera in each carriage within a first preset time period are obtained.
[0070] Each carriage of the subway is equipped with a camera, and the number of the camera can be one or more. Multiple cameras can monitor the situation in the subway carriage more comprehensively.
[0071] When multiple cameras are installed in each carriage, the shooting ranges of the multiple cameras are different.
[0072] Exemplarily, in an embodiment of the present application, two cameras are installed in each car, located at the front and rear positions of the car respectively; because people who wander abnormally on the subway usually walk from the first car to the last car of the subway and wander repeatedly in order to communicate with passengers, while normal passengers may move in a single car, such as looking for seats, but generally do not walk between different cars; therefore, two cameras are respectively set at the front and rear positions of the car, which can effectively detect people wandering in different cars, and exclude people who move normally in the car to a certain extent, thereby reducing the amount of calculation and improving detection accuracy.
[0073] The video image captured by the camera within the first preset time period is specifically a video image within the shooting range of the camera.
[0074] The first preset time period, for example, may be 2 minutes, 3 minutes, etc.
[0075] Exemplarily, the camera in the subway car is connected to the subway central monitoring system or data storage server, and the collected video images are uploaded to the subway central monitoring system or data storage server in real time. When the first preset time period is reached, the subway central monitoring system obtains the video images taken by the camera in each car during the first preset time period, and processes and analyzes them.
[0076] During the entire process of video image acquisition, processing, storage and utilization, it is necessary to ensure data security and compliance, and desensitize video data involving personal privacy or restrict access rights.
[0077] In some embodiments, after obtaining the original video images taken by the camera in each car within a first preset time period from the subway central monitoring system or data storage server, the video images are preprocessed. Through preprocessing, adverse factors such as noise, blur, deformation, etc. in the image can be removed or weakened, making the image clearer and richer in details.
[0078] Exemplarily, the preprocessing includes image denoising, image enhancement, etc.; image denoising can adopt methods such as mean filtering, median filtering, Gaussian low-pass filtering, Fourier transform, wavelet transform, etc.; image enhancement can adopt histogram equalization, contrast stretching, sharpening, etc.
[0079] It is particularly important to note that in the embodiment of the present application, the video image is pre-processed, including the deformation correction processing of the video image. Because the camera (i.e., the camera) in the car uses a fisheye camera with a wide viewing angle, a single lens can cover a large area, but the distortion of the picture is extremely serious, which makes it difficult to match the human body and the face, so image correction is required.
[0080] Specifically, in the embodiment of the present application, the longitude and latitude expansion method is used to correct the distorted image in the video image into a square image to eliminate the picture distortion; the principle of the longitude and latitude expansion method is to regard the image as a projection of a sphere based on the geometric characteristics of fisheye lens distortion, and each pixel point is regarded as the longitude and latitude coordinates corresponding to the sphere, and the longitude and latitude are mapped to the horizontal coordinate system of the plane to complete the correction process.
[0081] In step S102, by performing recognition processing on the video images captured by the camera of each carriage, the images of moving persons are recognized from the video images, and the features of the target persons corresponding to the camera are extracted based on the images of moving persons.
[0082] Here, the moving person image represents the moving trajectory of the moving person; when the same camera and the same moving person move in the forward direction and the reverse direction, they have different moving trajectories.
[0083] Here, the video images captured by the camera of each carriage are processed for recognition, and the images of moving persons are identified from the video images. Specifically, the video images captured by the camera of each carriage are processed for recognition through a pre-trained mobile person recognition model, and the images of moving persons are identified from the video images.
[0084] The pre-trained mobile personnel recognition model specifically uses a target detection and tracking algorithm to recognize the mobile personnel image from the video image.
[0085] Specifically, the mobile person recognition model identifies the person in the image by analyzing the features in the video image (such as color, texture, shape, etc.), and gives the person's boundary box to identify the person's position information; in the video image of continuous frames, it tracks the identified person, records the person's movement trajectory, and updates their position information, thereby identifying the mobile person image from the video image.
[0086] Specifically, in the embodiment of the present application, when a person in an image is identified and a bounding box of the person is given, a pedestrian re-identification algorithm is used to perform global feature extraction and local feature extraction to identify the person in the image. Global feature extraction performs feature extraction on the entire person image to obtain the overall appearance information of the pedestrian. However, global features may be affected by factors such as posture, lighting, and occlusion, resulting in reduced recognition performance. Local feature extraction is performed on a certain area or component of the image to capture more detailed information.
[0087] In this way, in the embodiment of the present application, when identifying images of moving pedestrians, target detection and tracking algorithms and pedestrian re-identification algorithms are combined to solve problems such as deformation and partial occlusion of the tracked target during the tracking process.
[0088] The moving person image is identified from the video image. Since the identified person is tracked and the person's movement trajectory is recorded, the moving person image is often a continuous multiple images. Since the person's boundary frame has been identified, the moving person image is specifically the moving person image obtained by image segmentation, which only includes the moving person to reduce the subsequent calculation amount.
[0089] In some embodiments, some special video images may be screened out according to some rules to reduce the subsequent calculation amount or interference factors.
[0090] For details, please refer to Figure 2 , Figure 2A flow chart of a method for identifying a moving person image from a video image according to an embodiment of the present application is shown; the method for identifying a moving person image from the video image by performing recognition processing on the video image taken by a camera of each carriage comprises the following steps S201-S203:
[0091] S201, obtaining the door opening time and door closing time of the carriage, and filtering the video images captured by the camera of each carriage based on the time difference between the shooting time of the video image and the door opening time or door closing time to obtain a filtered video image;
[0092] S202, inputting the video image after the first screening into a mobile person recognition model to recognize images of mobile persons within the shooting range of the camera;
[0093] S203, matching the image of the mobile person within the shooting range of the camera with the preset staff features, and secondary screening the video image based on the matching result to obtain the secondary screened image of the mobile person.
[0094] The opening and closing time of the carriage door is the key moment for passengers to enter and exit the carriage. A large number of passengers are moving, and a large number of moving person images will be identified, but these moving person images have not actually moved subsequently; and abnormal wandering behavior usually occurs in the carriage when the door is not opened or closed, rather than talking to passengers when they are in a hurry to get on and off the carriage. Therefore, a part of the moving person images can be screened out based on time, which can effectively reduce the amount of video image data to be processed subsequently and improve processing efficiency.
[0095] The video images captured by the camera of each carriage are screened based on the time of capturing the video images and the time difference between the door opening time or the door closing time. Specifically, the video images whose time difference between the door opening time or the door closing time is less than a third preset threshold are screened out.
[0096] The third preset threshold value can be obtained by testing the boarding and alighting scenarios of a subway car.
[0097] Although a large number of moving person images are removed from the video images after one screening, the images of moving persons identified may still contain images of staff members. By matching the preset staff features (such as work clothes color, badges, etc.), the images of moving persons of staff members can be further removed. Since staff members often wander around in subway cars due to work needs, if abnormal wandering behavior is identified and reminded to other staff members every time, it will increase the workload and is also a false detection compared to detecting abnormal behavior.
[0098] In the step S102, the target person features corresponding to the camera are extracted based on the moving person image.
[0099] For details, please refer to Figure 3 , the step of extracting target person features corresponding to the camera based on the moving person image includes the following steps S301-S304:
[0100] S301, filtering out face capture images and / or body capture images that meet preset quality conditions from multiple images of moving persons captured by a camera based on multiple image quality dimensions;
[0101] S302, processing the face capture image based on the trained face feature extraction model to obtain face features;
[0102] S303, processing the human body captured image based on the trained human body feature extraction model to obtain human body features;
[0103] S304: Based on the facial features and body features, obtain the target person features corresponding to the camera.
[0104] The multiple images of moving persons captured by the camera essentially represent the moving trajectories of the moving persons within the range of the camera; the multiple image quality dimensions include detection category, detection object size, image clarity, face angle, face alignment result, degree of occlusion, etc.; the detection category includes face detection and body detection.
[0105] The detection categories clearly distinguish between face detection and body detection. Face detection focuses on facial features, while body detection focuses on the overall body shape. Sometimes it is difficult to capture high-quality face images in images of moving people, so combining body detection can capture the information of moving people more comprehensively.
[0106] The size of the detected object is used to ensure that the target object (face or body) in the image occupies a certain proportion to avoid recognition difficulties caused by objects that are too small or too large.
[0107] Image clarity is a key factor in image quality. High-definition images can provide more details, allowing recognition algorithms to extract features more accurately, thereby improving recognition accuracy.
[0108] The angle of the face has a great influence on the recognition results. Frontal or near-frontal face images usually contain the richest feature information, which is conducive to the processing of the recognition algorithm. By screening face images with appropriate angles, the amount of calculation can be reduced and the recognition efficiency can be improved.
[0109] Face alignment is the process of adjusting a face image to a standard position (such as frontal, horizontal, etc.). The aligned image helps the recognition algorithm to extract features more accurately, thereby improving recognition speed and accuracy.
[0110] The degree of occlusion is one of the important factors affecting the recognition effect. Severely occluded face or body images may lead to recognition failure or false alarms. By screening images with lower occlusion levels, false alarms and missed alarms can be reduced.
[0111] Therefore, in actual applications, the content of the images inside the car taken by the camera is complex and changeable, including background interference, crowded people, occlusion and other factors. Comprehensive evaluation based on multiple image quality dimensions can more comprehensively judge whether the image quality meets the preset conditions, screen out face captures and / or body captures that meet the preset quality conditions, improve recognition accuracy and optimize recognition efficiency.
[0112] Based on the above analysis, in the embodiments of the present application, it is necessary to comprehensively consider factors such as the detection category, object size, image clarity, face angle, face alignment results, and degree of occlusion, and use specific rules to screen out face captures and body captures that can represent the movement trajectory of moving people within the camera range, and then use the trained model to extract face features and body features as information for subsequent matching of the same moving person with different cameras.
[0113] Here, the face snapshot and / or body snapshot selected may be one or multiple. Specifically, the face snapshot and the body snapshot may be one image or different images, because a single image may only have the face snapshot or only the body snapshot, or both the face snapshot and the body snapshot; the face snapshot may be one or multiple; the body snapshot may be multiple or one.
[0114] The target person feature corresponding to the camera is the target person feature of the moving person photographed by the camera. When the camera photographs multiple moving persons, each moving person corresponds to a target person feature.
[0115] In other words, if the camera corresponds to multiple target person features, the multiple target person features are features of different moving persons.
[0116] Preferably, in the embodiment of the present application, one face capture and one body capture with the best quality are selected.
[0117] The method of processing the human body captured image based on the trained human body feature extraction model to obtain human body features includes:
[0118] In the embodiment of the present application, a human body capture is processed based on a trained human body feature extraction model to extract posture features of three granularities; the three granularities include: global granularity, upper and lower granularity, and upper, middle and lower granularity; different granularities correspond to different body regions; the global granularity represents the entire body region, the upper and lower granularity represents the division of the entire body region into an upper body region and a lower body region, and the upper, middle and lower granularity represents the division of the entire body region into three body regions from top to bottom;
[0119] Combining the posture features of three granularities corresponding to the target body area in the human body capture image to obtain the combined posture features corresponding to the target body area;
[0120] The combined posture features and the posture features of three granularities are used as human body features.
[0121] The human feature extraction network uses a multiple granularity network, which takes into account both overall and local features and gives the model a certain degree of interpretability. This feature will also have excellent performance when combined with business. The network extracts features of three granular structures: global, top-bottom, and top-middle-bottom.
[0122] Posture features of three granularities These features can also be understood as features of corresponding positions on the body and can be split and used, so that when searching for people, you can set the tops and bottoms separately.
[0123] The posture features of three granularities corresponding to the target body area in the human body capture are combined to obtain the combined posture features corresponding to the target body area. For example, the features related to the upper body are combined, such as combining the features of the upper body area in the upper and lower granularities and the upper 1 / 3 area in the upper, middle and lower granularities to characterize the features of the top, thereby matching subsequent features based on the body features of the upper body area.
[0124] Specifically, the combined posture features corresponding to the target body area here are the combined posture features corresponding to the upper body area and the combined posture features corresponding to the lower body area.
[0125] The combined posture features and the posture features of three granularities are used as human body features. At this time, the human body features include the posture features of the entire body area at three granularities and the combined posture features corresponding to the upper body area and the lower body area respectively. Through the posture features of three granularities (different abstract levels from coarse granularity to fine granularity), the human body data in the image of the moving person can be abstractly represented at multiple levels, and rich information from the whole to the part can be captured, thereby providing a more comprehensive human body description; regional subdivision helps to more accurately describe the posture changes of different parts of the human body and improve the recognition accuracy; since the posture features of multiple granularities and regions are included, the human body features can better adapt to the diversity and variability of human body postures. When the posture of the moving person changes or part of the body area is blocked, the posture features of other regions and granularities can still provide useful information, thereby improving the recognition accuracy of subsequent abnormal wandering persons.
[0126] In step S103, the target person features corresponding to different cameras within the second preset time period are processed, and wandering person images that meet the preset wandering time condition and the preset wandering space condition are screened out from the moving person images of different cameras.
[0127] The preset wandering time condition and the preset wandering space condition are determined based on the time characteristics and space characteristics of the abnormal wandering person when he / she wanders abnormally in the subway car.
[0128] After analysis, it was found that abnormal wandering people usually travel through multiple carriages, that is, they will be photographed by cameras in different locations at different times; in special circumstances, they will be photographed by the same camera, such as moving to the end of the train and turning around.
[0129] Therefore, it is necessary to analyze whether the same moving person appears in the shooting range of different cameras to identify the wandering person.
[0130] Please refer to Figure 4 The processing of target personnel features corresponding to different cameras within the second preset time period described in step S103, and screening out wandering personnel images that meet the preset wandering time condition and the preset wandering space condition from the moving personnel images of different cameras, includes the following steps S401-S403:
[0131] S401, taking each target person feature corresponding to different cameras within the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result;
[0132] S402, based on the clustering result, determining the shooting time and camera of the moving person image corresponding to the target person feature of each cluster of the clustering result;
[0133] S403 , based on the shooting time and camera of the mobile person images corresponding to the target person features in each cluster, filter out the wandering person images that meet the preset wandering time condition and the preset wandering space condition from the mobile person images of different cameras.
[0134] Specifically, clustering refers to dividing a data set into different classes or clusters according to a specific standard (such as similarity), so that the similarity of data objects in the same cluster is as large as possible, and the difference of data objects in different clusters is as large as possible.
[0135] In the embodiment of the present application, a cluster includes target person features of different moving trajectories of the same moving person under a camera; the different moving trajectories are different moving trajectories of the same moving person under different cameras, and / or moving trajectories of the same moving person in different moving directions under a camera.
[0136] The general process of the clustering method based on information graph is:
[0137] Initialize and treat each node (which may contain facial features and body features) as an independent group; calculate the similarity between any two nodes as the transition probability; randomly sample a sequence of nodes in the graph, and try to assign each node to the group where the neighboring node belongs in order, and assign the group with the largest average bit drop to the node. If there is no drop, the group to which the node belongs remains unchanged; in order to avoid random walks entering isolated areas, the crossing probability is introduced; iteratively repeat the above steps until the optimal solution is reached.
[0138] In an embodiment of the present application, the nodes are clustered to obtain clustering results, in which each cluster is a cluster of the same mobile person captured by the camera; next, it is necessary to determine whether the mobile person conforms to the spatiotemporal law of abnormal wandering based on the shooting time and camera of the mobile person image corresponding to the target person characteristics.
[0139] In the method for identifying abnormally wandering persons in a subway car described in an embodiment of the present application, the target person features include facial features and body features.
[0140] In the embodiment of the present application, when calculating the similarity between two nodes, facial features are preferably used for matching. Facial features are highly unique and recognizable, and each person's facial features are unique (except for some extremely similar twins). Therefore, in most cases, facial features can be used to quickly and accurately match the target person's features under different cameras.
[0141] The method of taking each target person feature corresponding to different cameras within the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result includes:
[0142] Taking each target person feature of different cameras in the second preset time period as a node, determining whether both nodes have facial features;
[0143] If so, calculate the similarity of facial features as the similarity between the two nodes;
[0144] If not, the similarity of human features is calculated as the similarity between the two nodes;
[0145] Based on the similarity between any two nodes, the nodes are clustered to obtain a clustering result.
[0146] Wherein, based on the similarity between any two nodes, the nodes are clustered to obtain a clustering result, including:
[0147] Based on the human feature similarity between two nodes, when clustering the nodes, when the human feature similarity between a first node and a plurality of second nodes is greater than a preset human feature similarity threshold, the time interval, trajectory direction, and camera distance of the moving person image corresponding to the human features of the first node and each second node are determined; the camera distance is the difference in the serial numbers of the cameras of the first node and the second node;
[0148] Determine the trajectory direction of the first node and each second node, the camera distance and the priority matching result of the preset wandering rules of different priorities; wherein the preset wandering rules of different priorities correspond to different weights, and the preset wandering rules with higher priorities have higher weights;
[0149] Based on the priority matching result and the weights corresponding to the preset wandering rules of different priorities, the human feature similarities of the first node and the plurality of second nodes are updated, and the clustering result is determined based on the updated human feature similarities.
[0150] Specifically, in the embodiment of the present application, the feature similarity between two nodes adopts the cosine distance, which is represented by the function Represents the distance between the target person feature vector A and the target person feature vector B.
[0151]
[0152] Let the node set P be For Node facial features, among which For Node The human features of , the human feature similarity threshold is set to .
[0153] First, determine whether the camera nodes have high-quality facial features. If both nodes have high-quality facial features, use facial similarity as the judgment criterion and find all nodes that meet the criteria. Node pairs and aggregate nodes and , otherwise go to the human feature identification logic.
[0154] Since human body capture may be blocked and have angle changes in the car environment, only the lower limit threshold is set here. When the node and will not be considered as snapshots of the same moving person, but , then we can only consider the node and The snapshots that are not of the same mobile person are not excluded, that is, they may be snapshots of the same mobile person; this strategy will result in multiple matches for a node, so space and time constraints are introduced, and the two nodes with the highest matching scores that meet the constraints are selected for aggregation, thereby improving the accuracy of aggregation and the accuracy of wandering identification.
[0155] The spatial and temporal constraint strategies are as follows: First, set the camera numbers from the front to the rear of the vehicle in a continuous increment, starting from 0 and recorded as ,camera and The distance is ,Notice , Can be negative. If the node The camera belongs to ,node The camera belongs to , then the function Representation Node and The camera distance is the difference in the camera numbers, for example, the first camera and the second camera are 1.
[0156] Then set the trajectory direction for each capture (from the front to the rear of the vehicle is forward, and vice versa is reverse). During the matching process, set seven priority levels of preset wandering rules, with priority 1 being the highest. The higher the priority, the greater the possibility of occurrence and the greater the weight.
[0157] Priority 1: same direction, distance 1, that is, walking forward, and being captured by all adjacent cameras; Priority 2: opposite direction, distance 0, that is, turning around under the same camera, and being captured twice by the same camera; Priority 3: same direction, distance 2, that is, walking forward and passing through adjacent cameras ABC, but missed by camera B; Priority 4: opposite direction, distance -1, walking forward and then turning around in the opposite direction, and being captured by all adjacent cameras; Priority 5: opposite direction, distance 1, (walking forward and passing through adjacent cameras ABC, but missed by camera C; Priority 6: same direction, distance 3, that is, walking forward and passing through adjacent cameras ABCD, but missed by camera BC; Priority 7: opposite direction, distance 2, that is, walking forward and passing through adjacent cameras ABCD and then turning around, but missed by camera BC.
[0158] Therefore, the preset loitering rules of different priorities are determined based on the movement trajectory of the loitering person and the spatial relationship of the camera.
[0159] Assume that we are looking for a node Is there a matching node? , traverse all The node pair, set the priority , starting from u=1, determine the node pair Whether the priority u is met, if it is met, the sample pair Add to collection , when the collection When not empty, find the set middle The largest node pairs are aggregated. Represents the priority weight. The highest weight indicates that the possibility of being the same person is very high, so they are aggregated to obtain the aggregated result.
[0160] After the clustering result is determined, the preset wandering time condition needs to be considered. The wandering person images that meet the preset wandering time condition and the preset wandering space condition are screened out from the moving person images based on the shooting time and camera of the moving person images corresponding to the target person features in each cluster, including:
[0161] Determine the shooting time and camera of the mobile person image corresponding to the target person features of the two nodes, and determine the time interval, trajectory direction, and camera distance of the same mobile person passing through the camera shooting area twice; the camera distance is the difference in the serial numbers of the cameras photographed twice; different trajectory directions and camera distances correspond to different time constraint thresholds;
[0162] If the time interval between two times when the same mobile person passes through the shooting area of the camera exceeds the corresponding time constraint threshold, it is determined that the mobile persons in the mobile person images corresponding to the two nodes are not the same person, and the mobile person images corresponding to the two nodes are not wandering person images.
[0163] Specifically, considering the time constraint, the longest passing time through the shooting range of the two cameras is , when the node pair The time interval between snapshot timestamps exceeds the distance , then the sample pair They do not belong to the same person, that is, a wandering person must pass through the shooting area of the camera multiple times within a certain time constraint threshold, and the time constraint threshold is related to the distance between the cameras.
[0164] To sum up, the preset wandering time condition and the preset wandering space condition described in the embodiments of the present application are: the mobile person passes through the shooting area of the camera multiple times within a certain time constraint threshold, and the time constraint threshold is related to the distance between the cameras; it is particularly important to note that the camera distance (camera distance) can be 0, that is, the mobile person repeatedly passes through the shooting range of a camera.
[0165] In other words, the preset wandering time condition and the preset wandering space condition, that is, the preset wandering space condition is: the moving trajectory of the moving person is photographed by different cameras, or two moving trajectories in different moving directions are photographed by the same camera; the preset wandering time condition, that is, the time difference between the two moving trajectories of the moving person meets the preset time constraint condition, and the preset time constraint condition is related to the camera distance between the two moving trajectories.
[0166] In the step S104, the screened wandering person images are processed to determine abnormal behavior recognition results of the wandering person in the wandering person images, and abnormal wandering recognition results are determined.
[0167] Specifically, the screened wandering person images are processed to determine abnormal behavior recognition results of the wandering person in the wandering person images using a trained abnormal behavior recognition model.
[0168] Exemplarily, the abnormal behavior recognition model uses a target detection algorithm to detect whether there are specific items, such as mobile phones, musical instruments, etc., in the image of the wandering person.
[0169] Please refer to Figure 5 The processing of the screened wandering person image, determining the abnormal behavior recognition result of the wandering person in the wandering person image, and determining the abnormal wandering recognition result includes the following steps S501-S504:
[0170] S501, inputting the screened wandering person image into a trained abnormal behavior recognition model, wherein the abnormal behavior recognition model processes the wandering person image and outputs the abnormal behavior confidence of the wandering person in the wandering person image;
[0171] S502: when the abnormal behavior confidence exceeds a first preset threshold, determining that the abnormal wandering identification result is that there is an abnormal wandering person in the subway car;
[0172] S503: when the abnormal behavior confidence level is less than a second preset threshold, determining that the abnormal wandering identification result is that there is no abnormal wandering person in the subway car;
[0173] S504. When the abnormal behavior confidence is greater than or equal to the second preset threshold and less than or equal to the first preset threshold, the video image corresponding to the wandering person image and the pre-set reference question are input into the trained visual language large model, so that the visual language large model recognizes the image information of the wandering person image, and outputs the reference answer to the reference question as the abnormal wandering identification result based on the recognized image information.
[0174] Here, the visual language model can adopt the CogVLM2 model. CogVLM2 is a powerful open source visual language base model. Unlike the popular shallow alignment method that maps image features to the input space of the language model, CogVLM bridges the gap between the frozen pre-trained language model and the image encoder by introducing trainable visual expert modules in the attention and FFN (fully connected feedforward network) layers.
[0175] The Visual Language Big Model is an artificial intelligence technology that combines visual and textual information. It can learn from images and text simultaneously, and demonstrates excellent ability in understanding and generating content involving images and text. The Visual Language Big Model can generate image descriptions and automatically generate descriptive text for images. It can capture all key information in the video images taken by the camera, not just the information of moving people, and accurately describe them in natural language. The Visual Language Big Model can understand natural language questions and provide answers based on the image content.
[0176] Fine-tuning of the large visual language model can integrate the knowledge base of identifying abnormal wandering people in subway cars for search and build a question-answering system. The fine-tuning process described in the embodiment of the present application introduces a small, low-rank matrix in the decisive layer of the model to achieve fine-tuning of the model behavior without significantly modifying the entire model structure. Under the premise of not significantly increasing the additional computational burden, the model can be effectively fine-tuned while retaining the original performance level of the model.
[0177] In the embodiment of the present application, the trained visual language model is obtained by training using the following method:
[0178] Identify the target weight matrix that needs to be fine-tuned from the existing large visual language model;
[0179] Introducing a first low-rank matrix and a second low-rank matrix, and training the first low-rank matrix and the second low-rank matrix based on a sample training set; the sample training set is obtained based on multiple scene images of a carriage;
[0180] The product of the trained first low-rank matrix and the second low-rank matrix is superimposed on the target weight matrix to obtain a large visual language model after the target weight matrix is adjusted.
[0181] Here, the target weight matrix that needs to be fine-tuned is generally located in the multi-head self-attention and feedforward neural network parts of the model.
[0182] Introducing the first low-rank matrix A and the second low-rank matrix B, assuming that the size of the original weight matrix is dd, the sizes of A and B may be dr and r*d, where r is much smaller than d.
[0183] The product AB of these two low-rank matrices generates a new matrix whose rank (i.e., r) is much smaller than the rank of the original weight matrix. This product is actually a low-rank approximation adjustment of the original weight matrix; finally, the newly generated low-rank matrix AB is superimposed on the original weight matrix. Therefore, the original weights have been slightly adjusted, but most of the weights remain unchanged. This process can be described by a mathematical expression: new weight = original weight + AB.
[0184] The fine-tuning-based visual language model training method described in the embodiment of the present application does not need to directly modify the existing large number of weights of the visual language model. Instead, it only needs to introduce low-rank matrices in key parts of the model and perform effective weight adjustments through the product of these matrices. The visual language model can better adapt to the professional language and terminology in the field of abnormal wandering detection, while also avoiding the need for large-scale weight adjustments and retraining.
[0185] In the embodiment of the present application, the sample training set is obtained based on multiple scene images of the carriage; specifically, the sample training set consists of two parts: scenes during subway operation and crowd actor simulation; the scenes during train operation will include common scenes such as the starting station (empty car->someone), the terminal station (someone->empty car), getting on and off the train, and passengers walking through to find seats, and the crowd actor simulation will include a variety of abnormal behaviors. During the data collection process, the timing slice is considered to automatically identify the time segments of door opening and closing and walking through by using methods such as door opening and closing detection and trajectory tracking to improve efficiency. In order to further improve the utilization efficiency under limited data and allow the visual language model to further adapt to the carriage scene, the fine-tuning of the question words set will not be limited to abnormal behaviors. For example, you can set questions such as how many people are on the seats above the image, their respective clothes, and gender, so that the model can transfer the knowledge understanding of daily scenes to the carriage scene.
[0186] In the embodiment of the present application, if the confidence level of the abnormal behavior of the wandering person in the output wandering person image is very high, it can be directly considered that the person has abnormal behavior; if it is very low, it can be directly considered that the person has no abnormal behavior; for the recognition results whose confidence level of abnormal behavior is neither too high nor too low, the detailed information of the wandering person image is analyzed with the help of the visual language large model; the visual language large model can understand and process the complex scenes and relationships in the image, and at the same time, combined with natural language processing technology, analyze the reactions of other passengers in the car, environmental background and other factors, which can often provide important clues for judging whether the behavior of the wandering person is abnormal; by analyzing these detailed information, a more accurate and comprehensive reference answer as to whether the passenger has behaved abnormally is given.
[0187] Based on the same inventive concept, an embodiment of the present application also provides a device for identifying abnormal wandering persons in a subway car corresponding to the method for identifying abnormal wandering persons in a subway car. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0188] Please refer to Figure 6 , Figure 6 The schematic diagram of the structure of the device for identifying abnormal wandering persons in a subway car according to an embodiment of the present application is shown; the device for identifying abnormal wandering persons in a subway car comprises:
[0189] An acquisition module 601 is used to acquire video images captured by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the multiple cameras have different shooting ranges;
[0190] The extraction module 602 is used to identify the moving person image from the video image by performing recognition processing on the video image taken by the camera of each carriage, and extract the target person features corresponding to the camera based on the moving person image;
[0191] The processing module 603 is used to process the target person features corresponding to different cameras within the second preset time period, and select the wandering person images that meet the preset wandering time condition and the preset wandering space condition from the moving person images of different cameras;
[0192] The determination module 604 is used to process the screened wandering person images, determine the abnormal behavior recognition result of the wandering person in the wandering person images, and determine the abnormal wandering recognition result.
[0193] In some embodiments, the determination module in the device for identifying abnormal wandering persons in a subway car, when processing the screened wandering person images, determining the abnormal behavior recognition result of the wandering person in the wandering person images, and determining the abnormal wandering recognition result, is specifically used to:
[0194] Inputting the screened wandering person images into a trained abnormal behavior recognition model, the abnormal behavior recognition model processes the wandering person images, and outputs the confidence level of the abnormal behavior of the wandering person in the wandering person images;
[0195] When the abnormal behavior confidence exceeds a first preset threshold, determining that the abnormal wandering identification result is that there is an abnormal wandering person in the subway car;
[0196] When the abnormal behavior confidence is less than a second preset threshold, determining that the abnormal wandering identification result is that there is no abnormal wandering person in the subway car;
[0197] When the abnormal behavior confidence is greater than or equal to the second preset threshold and less than or equal to the first preset threshold, the video image corresponding to the wandering person image and the pre-set reference question are input into the trained visual language model, so that the visual language model recognizes the image information of the wandering person image, and outputs the reference answer to the reference question as the abnormal wandering identification result based on the recognized image information.
[0198] In some embodiments, the device for identifying abnormal wandering persons in a subway car further includes a training module for identifying a target weight matrix that needs to be fine-tuned from an existing large visual language model;
[0199] Introducing a first low-rank matrix and a second low-rank matrix, and training the first low-rank matrix and the second low-rank matrix based on a sample training set; the sample training set is obtained based on multiple scene images of a carriage;
[0200] The product of the trained first low-rank matrix and the second low-rank matrix is superimposed on the target weight matrix to obtain a large visual language model after the target weight matrix is adjusted.
[0201] In some embodiments, the extraction module in the device for identifying abnormally wandering persons in the subway car, when identifying the image of the moving person from the video image by performing recognition processing on the video image taken by the camera of each car, is specifically used to:
[0202] Obtaining the door opening time and door closing time of the carriage, and filtering the video images captured by the camera of each carriage based on the time difference between the shooting time of the video image and the door opening time or door closing time, to obtain a filtered video image;
[0203] Inputting the video image after the first screening into a mobile person recognition model to recognize images of mobile persons within the shooting range of the camera;
[0204] The images of mobile personnel within the shooting range of the camera are identified and matched with preset staff features, and the video images are secondary screened based on the matching results to obtain secondary screened images of mobile personnel.
[0205] In some embodiments, the extraction module in the device for identifying abnormally wandering persons in a subway car, when extracting the target person features corresponding to the camera based on the moving person image, is specifically used to:
[0206] Based on multiple image quality dimensions, face capture images and / or body capture images meeting preset quality conditions are screened out from multiple images of moving persons captured by the camera;
[0207] Process the facial capture image based on the trained facial feature extraction model to obtain facial features;
[0208] Process the human body capture based on the trained human body feature extraction model to obtain human body features;
[0209] Based on the facial features and body features, the target person features corresponding to the camera are obtained.
[0210] In some embodiments, the extraction module in the device for identifying abnormal wandering persons in a subway car, when processing a captured human body image based on a trained human body feature extraction model to obtain human body features, is specifically used to:
[0211] Processing human body captures based on the trained human body feature extraction model to extract posture features of three granularities; the three granularities include: global granularity, upper and lower granularity, and upper, middle and lower granularity; different granularities correspond to different body regions; the global granularity represents the entire body region, the upper and lower granularity represents dividing the entire body region into an upper body region and a lower body region, and the upper, middle and lower granularity represents dividing the entire body region into three body regions from top to bottom;
[0212] Combining the posture features of three granularities corresponding to the target body area in the human body capture image to obtain the combined posture features corresponding to the target body area;
[0213] The combined posture features and the posture features of three granularities are used as human body features.
[0214] In some embodiments, the processing module in the device for identifying abnormal wandering persons in a subway car, when processing the target person features corresponding to different cameras in the second preset time period and screening out wandering person images that meet the preset wandering time condition and the preset wandering space condition from the moving person images of different cameras, is specifically used to:
[0215] Taking each target person feature corresponding to different cameras in the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result;
[0216] Based on the clustering results, determining the shooting time and camera of the mobile person image corresponding to the target person feature of each cluster of the clustering results;
[0217] Based on the shooting time and camera of the mobile person images corresponding to the target person features in each cluster, the wandering person images meeting the preset wandering time condition and the preset wandering space condition are screened out from the mobile person images of different cameras.
[0218] In some embodiments, in the device for identifying abnormally wandering persons in a subway car, the target person features include facial features and body features;
[0219] The processing module, when taking each target person feature corresponding to different cameras within the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result, is specifically used to:
[0220] Taking each target person feature of different cameras in the second preset time period as a node, determining whether both nodes have facial features;
[0221] If so, calculate the similarity of facial features as the similarity between the two nodes;
[0222] If not, the similarity of human features is calculated as the similarity between the two nodes;
[0223] Based on the similarity between any two nodes, the nodes are clustered to obtain a clustering result.
[0224] In some embodiments, in the device for identifying abnormally wandering persons in a subway car, the processing module, when clustering the nodes based on the similarity between any two nodes to obtain the clustering result, is specifically used to:
[0225] Based on the human feature similarity between two nodes, when clustering the nodes, when the human feature similarity between a first node and a plurality of second nodes is greater than a preset human feature similarity threshold, the time interval, trajectory direction, and camera distance of the moving person image corresponding to the human features of the first node and each second node are determined; the camera distance is the difference in the serial numbers of the cameras of the first node and the second node;
[0226] Determine the trajectory direction of the first node and each second node, the camera distance and the priority matching result of the preset wandering rules of different priorities; wherein the preset wandering rules of different priorities correspond to different weights, and the preset wandering rules with higher priorities have higher weights;
[0227] Based on the priority matching result and the weights corresponding to the preset wandering rules of different priorities, the human feature similarities of the first node and the plurality of second nodes are updated, and the clustering result is determined based on the updated human feature similarities.
[0228] Based on the same inventive concept, an electronic device corresponding to the method for identifying abnormally wandering people in a subway car is also provided in the embodiment of the present application. Since the principle of solving the problem by the electronic device in the embodiment of the present application is similar to the above-mentioned method in the embodiment of the present application, the implementation of the electronic device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0229] An electronic device comprises: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the method for identifying abnormal wandering persons in a vehicle compartment are executed.
[0230] Based on the same inventive concept, the embodiment of the present application also provides a computer-readable storage medium corresponding to the method for identifying abnormal wandering persons in a subway car. Since the principle of solving the problem by the computer-readable storage medium in the embodiment of the present application is similar to the above-mentioned method in the embodiment of the present application, the implementation of the computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be repeated.
[0231] A computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for identifying an abnormally wandering person in a carriage.
[0232] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0233] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0234] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0235] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a platform server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.
[0236] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for identifying abnormal wandering persons in a subway car, characterized in that: A camera is installed in each carriage of the subway; the identification method includes: Acquire video images captured by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the multiple cameras have different shooting ranges; By performing recognition processing on the video images captured by the camera of each carriage, a moving person image is recognized from the video image, and the target person features corresponding to the camera are extracted based on the moving person image; Processing the target person features corresponding to different cameras within the second preset time period, and screening out images of wandering persons that meet preset wandering time conditions and preset wandering space conditions from the images of moving persons captured by different cameras; Processing the screened wandering person images, determining abnormal behavior recognition results of the wandering person in the wandering person images, and determining abnormal wandering recognition results; The processing of target person features corresponding to different cameras within the second preset time period, and screening out wandering person images that meet preset wandering time conditions and preset wandering space conditions from the moving person images of different cameras, includes: Taking each target person feature corresponding to different cameras in the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result; Based on the clustering results, determining the shooting time and camera of the mobile person image corresponding to the target person feature of each cluster of the clustering results; Based on the shooting time and camera of the mobile person images corresponding to the target person features in each cluster, images of wandering persons that meet the preset wandering time condition and the preset wandering space condition are screened out from the mobile person images of different cameras; The method of performing recognition processing on the video images captured by the camera of each carriage to recognize the images of moving persons from the video images includes: Obtaining the door opening time and door closing time of the carriage, and filtering the video images captured by the camera of each carriage based on the time difference between the shooting time of the video image and the door opening time or door closing time, to obtain a filtered video image; Inputting the video image after the first screening into a mobile person recognition model to recognize images of mobile persons within the shooting range of the camera; Matching the image of the mobile person within the shooting range of the camera with the preset staff features, and secondary screening the video image based on the matching result to obtain the secondary screened mobile person image; extracting the target person features corresponding to the camera based on the mobile person image, including: Based on multiple image quality dimensions, face capture images and / or body capture images meeting preset quality conditions are screened out from multiple images of moving persons captured by the camera; Process the facial capture image based on the trained facial feature extraction model to obtain facial features; Process the human body capture based on the trained human body feature extraction model to obtain human body features; Based on the facial features and body features, obtaining the target person features corresponding to the camera; The method of processing the human body captured image based on the trained human body feature extraction model to obtain human body features includes: Processing human body captures based on the trained human body feature extraction model to extract posture features of three granularities; the three granularities include: global granularity, upper and lower granularity, and upper, middle and lower granularity; different granularities correspond to different body regions; the global granularity represents the entire body region, the upper and lower granularity represents dividing the entire body region into an upper body region and a lower body region, and the upper, middle and lower granularity represents dividing the entire body region into three body regions from top to bottom; Combining the posture features of three granularities corresponding to the target body area in the human body capture image to obtain the combined posture features corresponding to the target body area; The combined posture feature and the posture features of three granularities are used as human body features; The target person features include facial features and body features; The method of taking each target person feature corresponding to different cameras within the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result includes: Taking each target person feature of different cameras in the second preset time period as a node, determining whether both nodes have facial features; If so, calculate the similarity of facial features as the similarity between the two nodes; If not, the similarity of human features is calculated as the similarity between the two nodes; Based on the similarity between any two nodes, clustering the nodes is performed to obtain a clustering result; Based on the similarity between any two nodes, clustering is performed on the nodes to obtain clustering results, including: Based on the human feature similarity between two nodes, when clustering the nodes, when the human feature similarity between a first node and a plurality of second nodes is greater than a preset human feature similarity threshold, the time interval, trajectory direction, and camera distance of the moving person image corresponding to the human features of the first node and each second node are determined; the camera distance is the difference in the serial numbers of the cameras of the first node and the second node; Determine the trajectory direction of the first node and each second node, the camera distance and the priority matching result of the preset wandering rules of different priorities; wherein the preset wandering rules of different priorities correspond to different weights, and the preset wandering rules with higher priorities have higher weights; Based on the priority matching result and the weights corresponding to the preset wandering rules of different priorities, updating the human feature similarities of the first node and the plurality of second nodes, and determining the clustering result based on the updated human feature similarities; After the clustering result is determined, the image of the wandering person that meets the preset wandering time condition and the preset wandering space condition is screened out from the moving person image based on the shooting time and camera of the moving person image corresponding to the target person feature in each cluster, including: Determine the shooting time and camera of the mobile person image corresponding to the target person features of the two nodes, and determine the time interval, trajectory direction, and camera distance of the same mobile person passing through the camera shooting area twice; the camera distance is the difference in the serial numbers of the cameras photographed twice; different trajectory directions and camera distances correspond to different time constraint thresholds; If the time interval between two times when the same mobile person passes through the shooting area of the camera exceeds the corresponding time constraint threshold, it is determined that the mobile persons in the mobile person images corresponding to the two nodes are not the same person, and the mobile person images corresponding to the two nodes are not wandering person images.
2. The method for identifying abnormal wandering persons in a subway car according to claim 1, characterized in that: The processing of the screened wandering person images, determining the abnormal behavior recognition result of the wandering person in the wandering person images, and determining the abnormal wandering recognition result include: Inputting the screened wandering person images into a trained abnormal behavior recognition model, the abnormal behavior recognition model processes the wandering person images, and outputs the confidence level of the abnormal behavior of the wandering person in the wandering person images; When the abnormal behavior confidence exceeds a first preset threshold, determining that the abnormal wandering identification result is that there is an abnormal wandering person in the subway car; When the abnormal behavior confidence is less than a second preset threshold, determining that the abnormal wandering identification result is that there is no abnormal wandering person in the subway car; When the abnormal behavior confidence is greater than or equal to the second preset threshold and less than or equal to the first preset threshold, the video image corresponding to the wandering person image and the pre-set reference question are input into the trained visual language model, so that the visual language model recognizes the image information of the wandering person image, and outputs the reference answer to the reference question as the abnormal wandering identification result based on the recognized image information.
3. The method for identifying abnormal wandering persons in a subway car according to claim 2, characterized in that: The trained visual language model is obtained by training using the following method: Identify the target weight matrix that needs to be fine-tuned from the existing large visual language model; Introducing a first low-rank matrix and a second low-rank matrix, and training the first low-rank matrix and the second low-rank matrix based on a sample training set; the sample training set is obtained based on multiple scene images of a carriage; The product of the trained first low-rank matrix and the second low-rank matrix is superimposed on the target weight matrix to obtain a large visual language model after the target weight matrix is adjusted.
4. A device for identifying abnormally wandering persons in a subway car, characterized in that: A camera is installed in each carriage of the subway; the recognition device includes: An acquisition module is used to acquire video images captured by a camera in each carriage within a first preset time period; each carriage is equipped with multiple cameras and the shooting ranges of the multiple cameras are different; An extraction module is used to identify and process the video images taken by the camera of each carriage, identify the images of moving persons from the video images, and extract the target person features corresponding to the camera based on the images of moving persons; a processing module, used for processing the target person features corresponding to different cameras within a second preset time period, and screening out images of wandering persons that meet preset wandering time conditions and preset wandering space conditions from the images of moving persons captured by different cameras; a determination module, used for processing the screened wandering person images, determining the abnormal behavior recognition result of the wandering person in the wandering person images, and determining the abnormal wandering recognition result; The processing module, when processing the target person features corresponding to different cameras within the second preset time period and screening out wandering person images that meet the preset wandering time condition and the preset wandering space condition from the moving person images of different cameras, is specifically used to: Taking each target person feature corresponding to different cameras in the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result; Based on the clustering results, determining the shooting time and camera of the mobile person image corresponding to the target person feature of each cluster of the clustering results; Based on the shooting time and camera of the mobile person images corresponding to the target person features in each cluster, images of wandering persons that meet the preset wandering time condition and the preset wandering space condition are screened out from the mobile person images of different cameras; The extraction module, when performing recognition processing on the video images captured by the camera of each carriage and identifying the images of the moving persons from the video images, is specifically used to: Obtaining the door opening time and door closing time of the carriage, and filtering the video images captured by the camera of each carriage based on the time difference between the shooting time of the video image and the door opening time or door closing time, to obtain a filtered video image; Inputting the video image after the first screening into a mobile person recognition model to recognize images of mobile persons within the shooting range of the camera; Matching the image of the mobile person within the shooting range of the camera with the preset staff features, and secondary screening the video image based on the matching result to obtain the secondary screened image of the mobile person; The extraction module, when extracting the target person features corresponding to the camera based on the moving person image, is specifically used to: Based on multiple image quality dimensions, face capture images and / or body capture images meeting preset quality conditions are screened out from multiple images of moving persons captured by the camera; Process the facial capture image based on the trained facial feature extraction model to obtain facial features; Process the human body capture based on the trained human body feature extraction model to obtain human body features; Based on the facial features and body features, obtaining the target person features corresponding to the camera; The extraction module, when processing a human body capture image based on a trained human body feature extraction model to obtain human body features, is specifically used to: Processing human body captures based on the trained human body feature extraction model to extract posture features of three granularities; the three granularities include: global granularity, upper and lower granularity, and upper, middle and lower granularity; different granularities correspond to different body regions; the global granularity represents the entire body region, the upper and lower granularity represents dividing the entire body region into an upper body region and a lower body region, and the upper, middle and lower granularity represents dividing the entire body region into three body regions from top to bottom; Combining the posture features of three granularities corresponding to the target body area in the human body capture image to obtain the combined posture features corresponding to the target body area; The combined posture feature and the posture features of three granularities are used as human body features; The target person features include facial features and body features; The processing module, when taking each target person feature corresponding to different cameras within the second preset time period as a node, calculating the similarity between any two nodes, and clustering the nodes based on the similarity to obtain a clustering result, is specifically used to: Taking each target person feature of different cameras in the second preset time period as a node, determining whether both nodes have facial features; If so, calculate the similarity of facial features as the similarity between the two nodes; If not, the similarity of human features is calculated as the similarity between the two nodes; Based on the similarity between any two nodes, clustering the nodes is performed to obtain a clustering result; The processing module, when clustering the nodes based on the similarity between any two nodes to obtain a clustering result, is specifically used to: Based on the human feature similarity between two nodes, when clustering the nodes, when the human feature similarity between a first node and a plurality of second nodes is greater than a preset human feature similarity threshold, the time interval, trajectory direction, and camera distance of the moving person image corresponding to the human features of the first node and each second node are determined; the camera distance is the difference in the serial numbers of the cameras of the first node and the second node; Determine the trajectory direction of the first node and each second node, the camera distance and the priority matching result of the preset wandering rules of different priorities; wherein the preset wandering rules of different priorities correspond to different weights, and the preset wandering rules with higher priorities have higher weights; Based on the priority matching result and the weights corresponding to the preset wandering rules of different priorities, updating the human feature similarities of the first node and the plurality of second nodes, and determining the clustering result based on the updated human feature similarities; After the clustering result is determined, the image of the wandering person that meets the preset wandering time condition and the preset wandering space condition is screened out from the moving person image based on the shooting time and camera of the moving person image corresponding to the target person feature in each cluster, including: Determine the shooting time and camera of the mobile person image corresponding to the target person features of the two nodes, and determine the time interval, trajectory direction, and camera distance of the same mobile person passing through the camera shooting area twice; the camera distance is the difference in the serial numbers of the cameras photographed twice; different trajectory directions and camera distances correspond to different time constraint thresholds; If the time interval between two times when the same mobile person passes through the shooting area of the camera exceeds the corresponding time constraint threshold, it is determined that the mobile persons in the mobile person images corresponding to the two nodes are not the same person, and the mobile person images corresponding to the two nodes are not wandering person images.
Citation Information
Patent Citations
Abnormal behavior early warning method and system for hovering event space-time big data analysis
CN105678247A
AI model-based loitering person detection method, edge device and storage medium
CN116189086A