A method for detecting violations of subway car passengers based on image segmentation technology
By segmenting the image of seats in the subway car and detecting the passenger's bone point, the passenger's violations are automatically identified, and the problem of detection of uncivilized behaviors in the subway car is solved, the detection accuracy is improved and the calculation volume is reduced, ensuring the normal operation of subway operations.
Patent Information
- Application Number
- CN202210823432.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The existing technology lacks effective detection methods for passengers in subway cars, which leads to uncivilized behaviors that cannot be stopped in time, increasing the workload of subway staff and affecting normal operations.
The seats in the subway car are accurately segmented based on image segmentation technology, and the inter-frame differential optical flow field algorithm and bone point detection are used to automatically identify violations based on passengers' behavior and key points of the feet.
Automatic detection of uncivilized behavior in subway cars is realized, the detection accuracy is improved, the calculation volume is reduced, the dependence on subway staff is reduced, and the normal operation of subway operations is ensured.
Smart Images

Figure CN115082965B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of subway management. Specifically, it relates to a method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology. Background Art
[0002] Subways have become the preferred means of transportation for more and more citizens around the world due to their speed and convenience. To improve the riding experience of citizens when taking the subway, it is necessary to stop the illegal behaviors of subway carriage passengers. If subway staff are used for patrol detection, it will inevitably increase the workload of the staff and affect the normal operation of busy subway sections.
[0003] Patent No. 201910011455.6 discloses a method, device, and electronic device for detecting vehicle illegal events, which includes collecting image data of all vehicles in a preset area in real time, determining the movement trajectory of each vehicle in the preset area according to the image data; and determining the current traffic event to which each vehicle belongs according to the movement trajectory of each vehicle and the traffic event determination model, and can accurately detect vehicle illegal events in various scenarios.
[0004] Patent No. 201910323569.4 discloses a on-site illegal behavior recognition system, which includes a signal triggering device, a passenger car identification device, and a compression and transmission device. When it detects a passenger car object, it compresses and wirelessly transmits the customized on-site image to a traffic management server for evidence preservation, thereby effectively identifying on-site illegal behaviors.
[0005] However, the above patents have the following problems when in use: Due to the increasing demand for civilized travel, uncivilized behaviors in the subway are resisted by more and more people. There will be some uncivilized and illegal behaviors in the subway, such as stepping on seats, lying on seats, etc. And it can be seen from the prior art that the prior art mainly detects vehicle illegal events for traffic violations, and lacks the detection of illegal behaviors of passengers inside the vehicle, that is, in the subway.
[0006] No effective solution has been proposed for the problems in the related art. Summary of the Invention
[0007] In view of the problems in the related art, the present invention proposes a method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology to overcome the above technical problems existing in the prior related art.
[0008] To this end, the specific technical solution adopted by the present invention is as follows:
[0009] A method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology, the method includes the following steps:
[0010] Read the current picture inside the carriage based on the surveillance camera;
[0011] Use real-time segmentation technology to accurately segment the seats inside the subway carriage and initialize the seat area;
[0012] Use the inter-frame difference optical flow field algorithm to determine whether to read the surveillance picture inside the carriage;
[0013] If the surveillance picture inside the carriage is read, detect the skeletal points of the passengers in the picture;
[0014] Classify their behavior actions based on the spatial structure information of the skeletal points;
[0015] If the passenger's behavior is abnormal, locate the key points of their feet and combine them with the effective area of the seat to determine whether there is any illegal behavior;
[0016] The step of using the inter-frame difference optical flow field algorithm to determine whether to read the surveillance picture inside the carriage further includes the following steps:
[0017] When an object moves in the surveillance scene, obtain the absolute value of the difference in grayscale values of two frames by subtracting two frames;
[0018] Calculate the frame difference images between the current frame and the previous frame and the next frame, and perform thresholding on the images after inter-frame difference to obtain the binary images after inter-frame difference and detect moving objects;
[0019] Perform optical flow analysis on the images of the surveillance picture, and set the pixels where the values in the difference map are not zero to correspond to the points with large grayscale gradients;
[0020] Calculate the speed, coordinates, direction, and grayscale value of the moving points of the moving object, determine whether there is an object in the surveillance picture, and if there is an object, determine to read the surveillance picture inside the carriage;
[0021] The step of classifying their behavior actions based on the spatial structure information of the skeletal points further includes the following steps:
[0022] Obtain the spatial structure information based on the skeletal points in the surveillance image, construct a node tree with a support vector machine, and mark the regions divided by the hyperplane in each step of the node tree as non-hard classes;
[0023] Use a non-linear support vector machine based on the radial basis kernel function to classify the spatial structure information based on the skeletal points, where it is determined that the passenger actions include standing, lying, sitting, and crossing legs.
[0024] Furthermore, the step of using real-time segmentation technology to accurately segment the seats inside the subway carriage and initialize the seat area further includes the following steps:
[0025] Pre-train a short-term dense connection module, integrate the short-term dense connection module into the U-net architecture to form a short-term dense connection module network, and configure a decoder at the same time;
[0026] Use the short-term dense connection module network as the backbone network of the encoder, and adopt the context path of BiSeNet to encode the context information of the seat image;
[0027] Guide the low layer to learn spatial information in a single-stream manner through a detail guidance module to obtain spatial details;
[0028] Fuse the spatial details with the context features of the encoder depth block, output the segmentation result, and initialize the seat area.
[0029] Furthermore, each layer in the short-term dense connection module encodes the input image or features at different scales and in their respective domains, and gradually reduces the convolution kernel size of the layer.
[0030] Furthermore, the short-term dense connection module is divided into several sub-modules, and ConvX i represents the operation of the i-th sub-module, and the output of the i-th sub-module is:
[0031] ;
[0032] Among them, F is the fusion operation, is the feature map of all blocks, and n is a non-zero natural number.
[0033] Furthermore, using the short-term dense connection module network as the backbone network of the encoder and adopting the context path of BiSeNet to encode the context information of the seat image further includes the following steps:
[0034] Set that in addition to the input layer and the prediction layer, the short-term dense connection module network also includes 6 stages;
[0035] In the 3rd to 5th stages, feature maps with downsampling rates of 1 / 8, 1 / 16, and 1 / 32 are generated respectively, and global average pooling is used to generate global context information with a large receptive field;
[0036] Adopt a U-shape structure to upsample the global features, combine them with the features of the last 2 stages in the encoding stage, and use an attention module to refine the combined features of every two stages after BiSeNet;
[0037] Among them, in the final semantic segmentation prediction, the feature fusion module in BiSeNet is used to fuse the 1 / 8 downsampled feature obtained from the 2nd stage with the 1 / 8 downsampled feature obtained from the decoder.
[0038] Furthermore, the step of guiding the lower layer to learn spatial information in a single-stream manner through the detail guidance module to obtain spatial details further includes the following steps:
[0039] Use the Laplacian operator to perform edge detection on the segmented real background image to generate an edge detail map, and insert a detail Head in the third stage to generate a detail feature map;
[0040] Use the details of the ground truth as the guidance of the detail feature map to guide the lower layer to learn;
[0041] Among them, the Laplacian operator generates detail feature maps with different step sizes to obtain multi-scale detail information, and uses the upsampling method to map the detail feature maps to the original size, and at the same time fuses a trainable 11-1 convolution for dynamic reweighting;
[0042] Use the boundary and corner information to convert the predicted details into the edge binary image of the final original image with a threshold of 0.1;
[0043] The detail loss formula of the detail feature map is as follows:
[0044] ;
[0045] In the formula, is the predicted detail, is the corresponding ground truth of the predicted detail, H is the height, W is the width, i is a non-zero natural number, is the smoothing factor.
[0046] Furthermore, the step of positioning the key points of its feet further includes the following steps:
[0047] Input the image to be detected obtained by photographing the passengers in the subway car into OpenPose, and output the key point confidence map and the vector field of the affinity between key points;
[0048] According to the heat map of the connection relationship between the key points and the keys, extract all key points, and group all key points so that all key points of the same person are assigned to the corresponding person, and at the same time obtain the foot points of any person;
[0049] Among them, OpenPose uses MobileNet as the backbone network and improves the receptive field by using dilated convolution.
[0050] Furthermore, the step of inputting the image to be detected obtained by photographing the passengers in the subway car into OpenPose and outputting the key point confidence map and the vector field of the affinity between key points further includes the following steps:
[0051] Input the image to be detected obtained by photographing passengers in the subway car into OpenPose;
[0052] Merge the two prediction branches in OpenPose into one branch;
[0053] At the output stage, two branches are separated through 1*1 convolution;
[0054] The two branches respectively output the confidence map of the key points and the vector field of the affinity between the key points.
[0055] Furthermore, 1*1, 3*3, 3*3 convolutional cascades are arranged in the two branches. In the last 3*3 convolution, dilated convolution with dilation = 2 is used, and residual connection structures are used in each convolution block.
[0056] The beneficial effects of the present invention are as follows: By segmenting the seats in the subway car image, the seat area is obtained, and the foot information of the passengers in the car is detected. If the feet of the passengers fall into the seat area, it is defined as a violation. Thus, uncivilized behaviors in the subway car can be automatically detected, contributing to civilized travel. Moreover, there is no need for subway staff to conduct patrol inspections, so it will not increase the workload of subway staff and will not have an adverse impact on the normal operation of the subway. The method for segmenting seats in the car image of the present invention has high speed and accuracy. At the same time, the method for detecting the bone points of car passengers of the present invention has a low computational complexity and can improve the accuracy by 6.5 times or more. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0058] Figure 1 It is a flowchart of a method for detecting violations of subway car passengers based on image segmentation technology according to an embodiment of the present invention. Detailed Embodiments
[0059] To further illustrate the embodiments, the present invention provides drawings. These drawings are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the drawings are not drawn to scale, and similar component symbols are usually used to represent similar components.
[0060] According to an embodiment of the present invention, a method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology is provided.
[0061] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments. As Figure 1 shown, the method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology according to an embodiment of the present invention includes the following steps:
[0062] S1. Read the image of the current carriage based on a surveillance camera;
[0063] S2. Accurately segment the seats in the subway carriage using real-time segmentation technology and initialize the seat area;
[0064] Among them, the step of accurately segmenting the seats in the subway carriage using real-time segmentation technology and initializing the seat area further includes the following steps:
[0065] Pre-train a short-term dense connection module and integrate the short-term dense connection module into the U-net (an algorithm for semantic segmentation using a fully convolutional network) architecture to form a short-term dense connection module network, and configure a decoder at the same time;
[0066] Use the short-term dense connection module network as the backbone network of the encoder, and adopt the context path of BiSeNet (a real-time semantic segmentation network structure) to encode the context information of the seat image;
[0067] Guide the lower layer to learn spatial information in a single-stream manner through a detail guidance module to obtain spatial details;
[0068] Fuse the spatial details with the context features of the encoder depth block, output the segmentation result, and initialize the seat area.
[0069] Among them, each layer in the short-term dense connection module encodes the input image or features at different scales and in their respective domains, and gradually reduces the convolution kernel size of the layer.
[0070] The short-term dense connection module is divided into several sub-modules, and ConvX i represents the operation of the i-th sub-module. The output of the i-th sub-module is:
[0071] ;
[0072] Among them, and are the input and output of the i-th sub-module respectively, includes a convolutional layer, a BN layer, and a ReLU, is the kernel size of the convolutional layer, and i is a non-zero natural number.
[0073] The final output of the short-term dense connection module is:
[0074] ;
[0075] where F is the fusion operation, is the feature map of all blocks, and n is a non-zero natural number.
[0076] Using the short-term dense connection module network as the backbone network of the encoder and adopting the context path of BiSeNet to encode the context information of the seat image further includes the following steps:
[0077] It is set that except for the input layer and the prediction layer, the short-term dense connection module network also includes 6 stages;
[0078] In the 3rd to 5th stages, feature maps with downsampling rates of 1 / 8, 1 / 16, and 1 / 32 are generated respectively, and global average pooling is used to generate global context information with a large receptive field;
[0079] The global features are upsampled using a U-shape structure, and combined with the features of the last 2 stages of the encoding stage. Moreover, an attention module is used after BiSeNet to refine the combined features of every two stages;
[0080] Among them, in the final semantic segmentation prediction, the feature fusion module in BiSeNet is used to fuse the 1 / 8 downsampled feature obtained from the 2nd stage with the 1 / 8 downsampled feature obtained from the decoder.
[0081] The guiding of the low layer to learn spatial information in a single-stream manner through the detail guidance module to obtain spatial details further includes the following steps:
[0082] The Laplacian operator (the Laplacian operator is a second-order differential operator in an n-dimensional Euclidean space, defined as the divergence div of the gradient grad, and this theorem can be calculated using an operation template) is used to perform edge detection on the segmented real background image to generate an edge detail map, and a detail Head is inserted in the 3rd stage to generate a detail feature map;
[0083] The details of the ground truth are used as the guidance of the detail feature map to guide the low layer to learn;
[0084] Among them, the Laplacian operator generates detail feature maps with different step sizes to obtain multi-scale detail information, and uses upsampling to map the detail feature maps to the original size. At the same time, a trainable 11-1 convolution is fused for dynamic reweighting;
[0085] Using the boundary and corner information, the predicted details are converted into the edge binary image of the final original image with a threshold of 0.1;
[0086] The detail loss formula of the detail feature map is as follows:
[0087] ;
[0088] is the predicted detail, is the corresponding true value of the predicted detail, is the binary cross-entropy loss, is the dice loss, the height of the detail map is H, and the width of the detail map is W;
[0089] ;
[0090] In the formula, is the predicted detail, is the corresponding true value of the predicted detail, H is the height, W is the width, i is a non-zero natural number, is the smoothing factor.
[0091] S3. Use the inter-frame difference optical flow field algorithm to determine whether to read (analyze) the monitoring video in the carriage; the purpose of using the frame difference method is to save the computing resources of the edge device. Only when there is target movement, the edge device will be triggered to read the video information for intelligent analysis.
[0092] Among them, the step of using the inter-frame difference optical flow field algorithm to determine whether to read the monitoring video in the carriage further includes the following steps:
[0093] When an object moves in the monitoring scene, the absolute value of the difference in gray values of two frames is obtained by subtracting two frames;
[0094] Calculate the frame difference images between the current frame and the previous frame and the next frame, and perform thresholding on the image after inter-frame difference to obtain the binary image after inter-frame difference, and detect moving objects;
[0095] Perform optical flow analysis on the image of the monitoring video, and set the pixels where the difference map is not zero to correspond to the points with large gray gradients;
[0096] Calculate the speed, coordinates, direction, and gray value of the moving points of the moving object, determine whether there is an object in the monitoring video, and if there is an object, determine to read the monitoring video in the carriage;
[0097] Inter-frame difference optical flow algorithm: When an object moves in the monitoring scene, there will be obvious differences between frames. That is, the absolute value of the brightness difference between the two frames is obtained by subtracting the two frames. By judging whether it is greater than the threshold, the motion characteristics of the video or image sequence are analyzed to determine whether there is an object in the image sequence.
[0098] S4. If the monitoring picture in the carriage is read (i.e., the current frame is determined to be analyzed), the skeleton points of the passengers in the picture are detected.
[0099] S5. Classify the behavior based on the spatial structure information of the skeleton points. Since it is too simple to judge whether the passenger has violated the rules by only relying on the coordinates of the key points of the feet and the seat area, it is easy to produce false alarms. Therefore, it is necessary to combine action recognition to improve the recognition rate and reduce the false alarm rate.
[0100] Wherein, the classification of the behavior action based on the spatial structure information of the skeleton points further comprises the following steps:
[0101] Obtain the spatial structure information based on the skeleton points in the monitoring image, and build a node tree with a support vector machine, and mark the area divided by the hyperplane in each step of the node tree as a non-hard class;
[0102] A nonlinear support vector machine based on radial basis kernel function is used to classify the spatial structure information based on skeleton points, wherein the passenger actions are judged to include standing, lying, sitting and crossing legs.
[0103] S6. If the passenger's behavior is abnormal, the key points of his feet are located and combined with the effective area of the seat to determine whether he has violated the rules; that is, if it is determined that the passenger is lying down or stepping on something, the key points of his feet are located to determine whether they are located in the effective area of the seat, and the violation is related to the feet.
[0104] Wherein, the positioning of the key points of the foot also includes the following steps:
[0105] The image to be detected obtained by photographing passengers in the subway car is input into OpenPose (OpenPose human posture recognition project is an open source library developed by Carnegie Mellon University in the United States based on convolutional neural networks and supervised learning and using Caffe as the framework. It can realize posture estimation of human movements, facial expressions, finger movements, etc. It is suitable for single and multi-person and has excellent robustness. It is the world's first real-time multi-person two-dimensional posture estimation application based on deep learning, and examples based on it have sprung up like mushrooms after rain. Human posture estimation technology has broad application prospects in sports fitness, motion collection, 3D fitting, public opinion monitoring and other fields), and outputs the key point confidence map and the vector field of affinity between key points;
[0106] According to the heatmap of the connection relationship between key points and key correspondences, extract all key points and group all key points so that all key points of the same person are assigned to the corresponding person, and at the same time obtain the foot points of any person;
[0107] Among them, the OpenPose uses MobileNet as the backbone network and improves the receptive field by using dilated convolution.
[0108] The step of inputting the to-be-detected image obtained by photographing the passengers in the subway car into OpenPose and outputting the confidence map of key points and the vector field of the affinity between key points further includes the following steps:
[0109] Input the to-be-detected image obtained by photographing the passengers in the subway car into OpenPose;
[0110] Merge the two prediction branches in OpenPose into one branch;
[0111] At the output stage, two branches are separated by a 1*1 convolution;
[0112] The two branches respectively output the confidence map of key points and the vector field of the affinity between key points.
[0113] In the two branches, there is a cascade of 1*1, 3*3, and 3*3 convolutions. In the last 3*3 convolution, dilated convolution with dilation = 2 (dilation is the dilation of the convolution kernel) is used, and a residual connection structure is used in each convolution block.
[0114] In summary, the present invention divides the seats in the subway car image to obtain the seat area, and detects the foot information of the passengers in the car. If the feet of the passengers fall into the seat area, it is defined as a violation. Furthermore, it can automatically detect uncivilized behaviors in the subway car, contribute to civilized travel, and does not require subway staff to conduct patrol inspections, thus not increasing the workload of subway staff and not having an adverse impact on the normal operation of the subway. The method for dividing seats in the car image of the present invention has high speed and accuracy. At the same time, the method for detecting the bone points of the car passengers of the present invention has a low computational amount and can improve the accuracy by 6.5 times or more.
[0115] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting illegal behaviors of passengers in subway carriages based on image segmentation technology, characterized in that, Including: Read the current picture inside the carriage based on a surveillance camera; Adopt real-time segmentation technology to accurately segment the seats inside the subway carriage and initialize the seat area; Use the inter-frame difference optical flow field algorithm to determine whether to read the surveillance picture inside the carriage: When an object moves in the surveillance scene, obtain the absolute value of the difference in grayscale values of two frames by subtracting two frames; Calculate the frame difference image between the current frame and the previous frame and the next frame, perform thresholding on the image after inter-frame difference, obtain the binary image after inter-frame difference, and detect moving objects; Perform optical flow analysis on the image of the surveillance picture, and set the pixels where the value in the difference map is not zero to correspond to points with large grayscale gradients; Calculate the speed, coordinates, direction, and grayscale value of the moving points of the moving object, determine whether there is an object in the surveillance picture, and if there is an object, determine to read the surveillance picture inside the carriage; If the surveillance picture inside the carriage is read, detect the skeleton points of the passengers in the picture; Classify their behavioral actions based on the spatial structure information of the skeleton points: Obtain the spatial structure information based on the skeleton points in the surveillance image, construct a node tree with a support vector machine, and mark the regions divided by the hyperplane in each step of the node tree as non-hard classes; Use a non-linear support vector machine based on the radial basis kernel function to classify the spatial structure information based on the skeleton points; If the passenger's behavior is abnormal, locate the key points of their feet, and input the image to be detected obtained by photographing the passengers inside the subway carriage into OpenPose, and output the confidence map of the key points and the vector field of the affinity between the key points; Combine with the effective area of the seat to judge whether there is any violation behavior.
2. The subway carriage personnel violation detection method based on image segmentation technology according to claim 1, wherein, The step of accurately segmenting the seats inside the subway carriage by using real-time segmentation technology and initializing the seat area further includes the following steps: Pre-train the short-term dense connection module, and integrate the short-term dense connection module into the U-net architecture to form a short-term dense connection module network, and configure the decoder at the same time; Use the short-term dense connection module network as the backbone network of the encoder, and adopt the context path of BiSeNet to encode the context information of the seat image; Guide the low layer to learn spatial information in a single-stream manner through the detail guidance module to obtain spatial details; Fuse the spatial details with the context features of the encoder depth block, output the segmentation result, and initialize the seat area.
3. The subway carriage personnel violation detection method based on image segmentation technology according to claim 2, wherein, Each layer in the short-term dense connection module encodes the input image or features at different scales and their respective domains, and gradually reduces the convolution kernel size of the layer.
4. The subway carriage personnel violation detection method based on image segmentation technology according to claim 2, wherein The short-term dense connection module is divided into several sub-modules, and ConvX i represents the operation of the i-th sub-module. The output of the i-th sub-module is as follows: ; Among them, and are the input and output of the i-th sub-module respectively, including a convolutional layer, a BN layer, and ReLU, is the kernel size of the convolutional layer, and i is a non-zero natural number.
5. The subway carriage personnel violation detection method based on image segmentation technology according to claim 4, characterized in that, The final output of the short-term dense connection module is: ; Among them, F is the fusion operation, is the feature map of all blocks, and n is a non-zero natural number.
6. The subway carriage personnel violation detection method based on image segmentation technology according to claim 2, characterized in that The step of using the short-term dense connection module network as the backbone network of the encoder and adopting the context path of BiSeNet to encode the context information of the seat image further includes the following steps: Set that in addition to the input layer and the prediction layer, the short-term dense connection module network also includes 6 stages; The 3rd stage to the 5th stage respectively generate feature maps with downsampling rates of 1 / 8, 1 / 16, and 1 / 32, and use global average pooling to generate global context information with a large receptive field; The global features are upsampled using a U-shape structure and combined with the features of the last two stages in the encoding phase, and an attention module is used after BiSeNet to refine the combined features of every two stages; Among them, in the final semantic segmentation prediction, the feature fusion module in BiSeNet is used to fuse the 1 / 8 downsampled features obtained from the second stage with the 1 / 8 downsampled features obtained from the decoder.
7. A method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology according to claim 2, characterized in that, The step of guiding the low layer to learn spatial information in a single-stream manner through the detail guidance module to obtain spatial details further includes the following steps: The Laplacian operator is used to detect the edges of the segmented real background image to generate an edge detail map, and a detail Head is inserted in the third stage to generate a detail feature map; The details of the ground truth are used as the guidance of the detail feature map to guide the low layer to learn; Among them, the Laplacian operator generates detail feature maps with different strides to obtain multi-scale detail information, and the detail feature maps are mapped to the original size by upsampling, and a trainable 11-1 convolution is fused for dynamic reweighting; The predicted details are converted into the edge binary image of the final original image by using the boundary and corner information with a threshold of 0.1; The detail loss formula of the detail feature map is as follows: ; For the details of the prediction, For the corresponding ground truth of the prediction details, For the binary cross-entropy loss, For the dice loss, the height of the detail map is H and the width of the detail map is W; ; Wherein, is the predicted detail, is the true value of the corresponding detail of the prediction, H is the height, W is the width, and i is a non-zero natural number, is the smoothing factor.
8. A method for detecting illegal behaviors of subway carriage passengers based on image segmentation technology according to claim 1, characterized in that The step of locating the key points of its feet further includes: According to the heat maps of the key points and the connection relationships corresponding to the keys, all key points are extracted and grouped so that all the key points of the same person are assigned to the corresponding person, and at the same time, the foot points of any person are obtained; Among them, OpenPose uses MobileNet as the backbone network and uses dilated convolution to increase the receptive field.
9. The subway carriage personnel violation detection method based on image segmentation technology according to claim 8, characterized in that The step of inputting the image to be detected obtained by photographing the passengers in the subway car into OpenPose and outputting the key point confidence map and the vector field of the affinity between the key points further includes the following steps: Input the image to be detected obtained by photographing the passengers in the subway car into OpenPose; Merge the two prediction branches in OpenPose into one branch; At the output stage, two branches are separated by a 1*1 convolution; The two branches respectively output the confidence map of the key points and the vector field of the affinity between the key points.
10. A method for detecting illegal behaviors of subway car passengers based on image segmentation technology according to claim 9, characterized in that, There are 1*1, 3*3, 3*3 convolution cascades in the two branches, and in the last 3*3 convolution, dilated convolution with dilation = 2 is used, and a residual connection structure is used in each convolution block.
Citation Information
Patent Citations
A method, apparatus, and electronic device for detecting vehicle violations.
CN109784254B
On-site violation identification system
CN111028515B
In-vehicle passenger detection apparatus and method of controlling same
CN111152744A
Subway illegal behavior early warning method based on improved HigherHRNet model and DNN network
CN113449609A
Semantic segmentation method and system based on low-illumination complex road scene
CN113902915A