A method, device, equipment and storage medium for detecting illegal passenger carrying by vehicles

By detecting target vehicles and pedestrians on multiple frame images in driving videos, and judging the positional relationship between vehicles and pedestrians with attention mechanisms, the problem of rapid and accurate illegal manned inspection of agricultural vehicles is solved, and effective supervision and management of illegal manned aircraft for agricultural vehicles is achieved.

CN114639038BActive Publication Date: 2025-06-20QINGDAO HISENSE TRANS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210217438.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-06-20
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

How to quickly and effectively detect whether agricultural vehicles are illegally carrying people to ensure travel safety.

Method used

By obtaining multiple frame images in the driving video, the target vehicle and pedestrian are detected on each frame of the image, combined with the spatial attention mechanism and the channel attention mechanism, the position relationship between the vehicle and the pedestrian and the overlap of the detection area are determined, and whether the vehicle is illegally carrying a manned.

Benefits of technology

It has achieved rapid and accurate detection of illegal manned personnel in agricultural vehicles, improved the efficiency of supervision and management, and ensured the safety of road travel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114639038B_ABST
    Figure CN114639038B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, equipment and storage medium for detecting illegal passenger carrying in vehicles, which relates to the field of intelligent transportation technology. The method includes: respectively performing target vehicle detection and pedestrian detection on multiple frames of driving images in a driving video to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; if the target vehicle detection result of any frame of driving image includes the detection information of the target vehicle and the pedestrian detection result includes the detection information of multiple pedestrians, then according to the detection information of the target vehicle and the detection information of each pedestrian in this frame of driving image, determine the positional relationship between the target vehicle and each pedestrian and the detection area overlap degree, and based on this, determine the passenger carrying determination result of the target vehicle in this frame of driving image; based on the passenger carrying determination results of the target vehicles of each obtained frame of driving image, determine whether the target vehicle is carrying passengers illegally. The above solution can quickly and effectively detect the illegal passenger carrying behavior of the target vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation technology, and particularly to a method, device, equipment and storage medium for detecting illegal passenger carrying of vehicles. Background Art

[0002] Some agricultural vehicles (such as tricycles, etc.) are the most common means of transportation in rural areas, which are convenient for farmers to transport agricultural materials and do business, etc., and have a very high usage frequency. However, due to the poor safety and stability of agricultural vehicles, there is a very high risk for passenger transportation. Once an accident occurs, it often causes extremely serious traffic accidents. Therefore, relevant laws clearly stipulate that it is prohibited for agricultural vehicles to carry passengers.

[0003] In order to assist in supervising the illegal passenger carrying behavior of agricultural vehicles and protect people's travel safety, how to quickly and effectively detect whether an agricultural vehicle is carrying passengers illegally is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a method, device, equipment and storage medium for detecting illegal passenger carrying of vehicles, so as to accurately detect whether an agricultural vehicle is carrying passengers illegally.

[0005] The specific technical solutions provided by the embodiments of this application are as follows:

[0006] In a first aspect, an embodiment of this application provides a method for detecting illegal passenger carrying of vehicles, including:

[0007] Obtain multiple frames of driving images in the driving video to be detected;

[0008] Perform target vehicle detection and pedestrian detection on the multiple frames of driving images respectively, to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target vehicle is a vehicle restricted from carrying passengers;

[0009] If the target vehicle detection result of any frame of driving image includes the detection information of the target vehicle, and the pedestrian detection result includes the detection information of multiple pedestrians, then according to the detection information of the target vehicle and the detection information of each pedestrian in the any frame of driving image, determine the position relationship and detection area overlap degree between the target vehicle and each pedestrian in the any frame of driving image;

[0010] According to the position relationship and detection area overlap degree between the target vehicle and the multiple pedestrians in the any frame of driving image, determine the passenger carrying determination result of the target vehicle in the any frame of driving image;

[0011] Based on the obtained passenger carrying determination results of the target vehicle corresponding to each frame of driving image, determine whether the target vehicle is carrying passengers illegally.

[0012] In the embodiments of the present application, target vehicle detection and pedestrian detection are performed on multiple frames of driving images in a driving video. For each frame of driving image, if a target vehicle and multiple pedestrians are detected, the manned determination result of the target vehicle in this frame of driving image is determined, for example: manned state or unmanned state; then, according to the manned determination results of the target vehicle in multiple frames of driving images, it is determined whether the target vehicle is illegally carrying passengers. In this way, it is possible to quickly and effectively detect whether the target vehicle is illegally carrying passengers.

[0013] In some optional embodiments, the performing target vehicle detection and pedestrian detection on the multiple frames of driving images respectively to obtain the target vehicle detection results and pedestrian detection results of each of the multiple frames of driving images includes:

[0014] Performing target vehicle detection and pedestrian detection on the multiple frames of driving images respectively through a target detection model to obtain the target vehicle detection results and pedestrian detection results of each of the multiple frames of driving images; wherein, the target detection model at least includes a spatial attention mechanism network and a channel attention mechanism network, and the spatial attention mechanism network and the channel attention mechanism network are used to focus on the visible areas of pedestrians in the driving images.

[0015] In the embodiments of the present application, considering the problem that pedestrians are easily missed when occluded, a spatial attention mechanism and a channel attention mechanism are introduced into the target detection model. When detecting pedestrians, the visible areas of pedestrians can be focused on, avoiding the problem of missing pedestrians, improving the recall rate of pedestrian detection, and thus more accurately detecting whether the target vehicle is illegally carrying passengers.

[0016] In some optional embodiments, the target detection model further includes a feature pyramid network, a first feature fusion network, and a second feature fusion network composed of multiple sequentially connected convolutional layers;

[0017] The first convolutional layer is connected to the spatial attention mechanism network, and some of the convolutional layers in the multiple convolutional layers and the spatial attention mechanism network are respectively connected to the first feature fusion network;

[0018] The second convolutional layer is connected to the channel attention mechanism network, and some of the convolutional layers in the multiple convolutional layers and the channel attention mechanism network are respectively connected to the second feature fusion network.

[0019] In the target detection model with the above structure, the feature pyramid network can achieve multi-scale fine-grained feature extraction. At the same time, in order to better achieve the accurate detection and separation of pedestrians and target vehicles, a spatial attention mechanism and a channel attention mechanism are introduced, which can ensure the detection accuracy of pedestrians and target vehicles, avoid the occurrence of missed detection situations, and thus more accurately detect whether the target vehicle is illegally carrying passengers.

[0020] In some alternative embodiments, the target detection network is obtained by training an initial target detection network with an image sample set. The loss function in the training process includes a classification loss function, a regression loss function, and a loss function of the spatial attention mechanism network;

[0021] Among them, the regression loss function uses the intersection over union of the predicted pedestrian region of each image sample and the actual visible pedestrian region as the weight of the loss value of the image sample.

[0022] In this embodiment, for an image sample, if its predicted pedestrian region overlaps more with the actual visible pedestrian region, then the loss generated by this image sample is more credible and a higher weight can be assigned; conversely, a lower weight is assigned. In this way, the trained target detection model can accurately detect the visible region of pedestrians, avoid the problem of missed detection caused by pedestrians being occluded, and thus more accurately detect whether the target vehicle illegally carries passengers.

[0023] In some alternative embodiments, the detection information of the target vehicle includes the target vehicle detection region and the position information of the target vehicle detection region, and the detection information of each pedestrian includes the pedestrian detection region and the position information of the pedestrian detection region;

[0024] Determining the position relationship and the detection region overlap degree between the target vehicle and each pedestrian in the arbitrary frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in the arbitrary frame of driving image includes:

[0025] According to the position information of the target vehicle detection region and the position information of each pedestrian detection region in the arbitrary frame of driving image, determining the position relationship of the center coordinates of the target vehicle detection region and each pedestrian detection region;

[0026] Taking the position relationship of the center coordinates of the target vehicle detection region and each pedestrian detection region as the position relationship between the target vehicle and each pedestrian;

[0027] Determining the ratio of the intersection and the union of the target vehicle detection region and each pedestrian detection region, and taking the ratio of the intersection and the union as the detection region overlap degree between the target vehicle and each pedestrian.

[0028] In the above embodiment, according to the position relationship between the center coordinates of the pedestrian detection region of each pedestrian and the corresponding target vehicle detection region, the position relationship between each pedestrian and the corresponding target vehicle can be determined; according to the ratio of the intersection and the union of the target vehicle detection region and each pedestrian detection region, the detection region overlap degree between the target vehicle and each pedestrian can be determined.

[0029] In some alternative embodiments, determining the passenger-carrying determination result of the target vehicle in any frame of driving image based on the positional relationship between the target vehicle and multiple pedestrians and the overlap degree of the detection regions in the any frame of driving image includes:

[0030] Determining target pedestrians among the multiple pedestrians whose overlap degree with the detection region of the target vehicle reaches a preset value;

[0031] If the number of the target pedestrians exceeds a preset number and the central coordinates of the pedestrian detection region of each target pedestrian are within the detection region of the target vehicle, determining that the target vehicle in the any frame of driving image is in a passenger-carrying state.

[0032] In the above embodiment, when determining the passenger-carrying determination result of the target vehicle, not only the positional relationship between the target vehicle and the pedestrians is considered, but also the overlap degree of the pedestrians and the detection region of the target vehicle is considered; when the overlap degree of a pedestrian and the detection region of the target vehicle meets the conditions and the central coordinates of the pedestrian detection region of the pedestrian are within the detection region of the target vehicle, it can be considered that the pedestrian is on the target vehicle, otherwise the pedestrian is considered to be a passing pedestrian, which can avoid false detection of illegal passenger-carrying caused by passing pedestrians. At the same time, considering the driver of the target vehicle, when the number of pedestrians on the target vehicle exceeds a preset number, the target vehicle is considered to be in a passenger-carrying state.

[0033] In some alternative embodiments, determining whether the target vehicle illegally carries passengers based on the passenger-carrying determination results of the target vehicle corresponding to each frame of driving image obtained includes:

[0034] If, in each frame of driving image, the proportion of the number of driving images in which the target vehicle is in a passenger-carrying state reaches a preset proportion and the target vehicle is in a non-stationary state in each frame of driving image, determining that the target vehicle illegally carries passengers.

[0035] In the above embodiment, in order to avoid false detection of passenger-carrying caused by spatio-temporal position overlap, when the target vehicle is in a non-stationary state, determining whether the target vehicle illegally carries passengers according to the passenger-carrying determination results of the target vehicle in multiple frames of driving images can improve the detection accuracy of illegal passenger-carrying of the target vehicle.

[0036] In a second aspect, an embodiment of the present application provides a vehicle illegal passenger-carrying detection device, including:

[0037] An acquisition module, configured to acquire multiple frames of driving images in a driving video to be detected;

[0038] A detection module, configured to perform target vehicle detection and pedestrian detection on the multiple frames of driving images respectively, to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target vehicle is a vehicle restricted from carrying passengers.

[0039] A position determination module, configured to, if the target vehicle detection result of any one frame of driving image includes the detection information of the target vehicle and the pedestrian detection result includes the detection information of multiple pedestrians, determine the position relationship and detection area overlap degree between the target vehicle and each pedestrian in the any one frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in the any one frame of driving image.

[0040] A status determination module, configured to determine the passenger-carrying determination result of the target vehicle in the any one frame of driving image according to the position relationship and detection area overlap degree between the target vehicle and the multiple pedestrians in the any one frame of driving image.

[0041] A passenger-carrying determination module, configured to determine whether the target vehicle illegally carries passengers based on the passenger-carrying determination results of the target vehicle corresponding to each frame of driving image obtained.

[0042] In some optional embodiments, the detection module is further configured to:

[0043] Perform target vehicle detection and pedestrian detection on the multiple frames of driving images respectively through a target detection model, to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target detection model at least includes a spatial attention mechanism network and a channel attention mechanism network, and the spatial attention mechanism network and the channel attention mechanism network are used to focus on the visible areas of pedestrians in the driving image.

[0044] In some optional embodiments, the target detection model further includes a feature pyramid network, a first feature fusion network, and a second feature fusion network composed of multiple sequentially connected convolutional layers;

[0045] The first convolutional layer is connected to the spatial attention mechanism network, and some of the convolutional layers in the multiple convolutional layers and the spatial attention mechanism network are respectively connected to the first feature fusion network;

[0046] The second convolutional layer is connected to the channel attention mechanism network, and some of the convolutional layers in the multiple convolutional layers and the channel attention mechanism network are respectively connected to the second feature fusion network.

[0047] In some alternative embodiments, the target detection network is obtained by training an initial target detection network with an image sample set, and the loss function in the training process includes a classification loss function, a regression loss function, and the loss function of the spatial attention mechanism network;

[0048] Wherein, the regression loss function uses the intersection over union of the predicted pedestrian region of each image sample and the actual pedestrian visible region as the weight of the loss value of the image sample.

[0049] In some alternative embodiments, the detection information of each target vehicle includes the target vehicle detection region and the position information of the target vehicle detection region, and the detection information of each pedestrian includes the pedestrian detection region and the position information of the pedestrian detection region;

[0050] The position determination module is further configured to:

[0051] Determine the positional relationship between the center coordinates of the target vehicle detection region and each pedestrian detection region according to the position information of the target vehicle detection region and the position information of each pedestrian detection region in any frame of driving image;

[0052] Use the positional relationship between the center coordinates of the target vehicle detection region and each pedestrian detection region as the positional relationship between the target vehicle and each pedestrian;

[0053] Determine the ratio of the intersection and union of the target vehicle detection region and each pedestrian detection region, and use the ratio of the intersection and union as the overlap degree of the detection regions between the target vehicle and each pedestrian.

[0054] In some alternative embodiments, the state determination module is further configured to:

[0055] Determine the target pedestrians among the multiple pedestrians whose overlap degree with the target vehicle detection region reaches a preset value;

[0056] If the number of target pedestrians exceeds a preset number and the center coordinates of the pedestrian detection region of each target pedestrian are within the target vehicle detection region, determine that the target vehicle in any frame of driving image is in a manned state.

[0057] In some alternative embodiments, the manned determination module is further configured to:

[0058] If the proportion of the number of driving images in which the target vehicle is in a manned state among all frames of driving images reaches a preset proportion and the target vehicle is in a non-stationary state in all frames of driving images, determine that the target vehicle is illegally carrying passengers.

[0059] In a third aspect, an embodiment of the present application provides a vehicle illegal passenger-carrying detection device, including a processor and a data receiving unit;

[0060] The data receiving unit is configured to: receive a driving video to be detected;

[0061] The processor is configured to: execute the method according to any one of the first aspect.

[0062] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method according to any one of the first aspect is implemented.

[0063] For the technical effects brought by any implementation manner of the second aspect to the fourth aspect, reference may be made to the technical effects brought by the corresponding implementation manner in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 It is a schematic diagram of an application scenario of a vehicle illegal passenger-carrying detection method provided in an embodiment of the present application;

[0066] Figure 2 It is a flowchart of a vehicle illegal passenger-carrying detection method provided in an embodiment of the present application;

[0067] Figure 3 It is a schematic diagram of the intersection-over-union ratio of a target vehicle detection area and a pedestrian detection area provided in an embodiment of the present application;

[0068] Figure 4 It is a schematic diagram of the detection of pedestrians and target vehicles in a driving image provided in an embodiment of the present application;

[0069] Figure 5 It is a flowchart of another vehicle illegal passenger-carrying detection method provided in an embodiment of the present application;

[0070] Figure 6 It is a schematic diagram of the structure of a target detection model provided in an embodiment of the present application;

[0071] Figure 7 It is a schematic diagram of the structure of a feature extraction network based on a spatial attention mechanism provided in an embodiment of the present application;

[0072] Figure 8 It is a schematic structural diagram of a feature extraction network based on a channel attention mechanism provided in an embodiment of the present application;

[0073] Figure 9 It is a schematic structural diagram of a feature fusion network provided in an embodiment of the present application;

[0074] Figure 10 It is a schematic structural diagram of a vehicle illegal passenger-carrying detection device provided in an embodiment of the present application;

[0075] Figure 11 It is a schematic structural diagram of a detection device provided in an embodiment of the present application. Detailed implementation manners

[0076] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0077] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0078] In an intelligent transportation system, a high-definition camera is used to capture a driving video, and the driving video is processed according to video image processing technology to detect illegal driving behaviors of vehicles. In order to quickly and effectively detect whether an agricultural vehicle illegally carries passengers, an embodiment of the present application provides a vehicle illegal passenger-carrying detection method, device, equipment, and storage medium, which perform target vehicle and pedestrian detection on multiple frames of driving images in the driving video. For each frame of driving image, if a target vehicle and a pedestrian are detected, the passenger-carrying determination result of the target vehicle in this frame of driving image is determined. Then, according to the passenger-carrying determination results of the target vehicles in multiple frames of driving images, it is determined whether the target vehicle illegally carries passengers. In this way, it is possible to quickly and effectively detect whether the target vehicle illegally carries passengers.

[0079] The following provides an exemplary introduction to the application scenarios of the vehicle illegal passenger-carrying detection method in the embodiments of the present application.

[0080] Refer to Figure 1, which is a schematic diagram of the application scenario of the vehicle illegal passenger-carrying detection method provided by the embodiments of the present application. The application scenario includes a camera 100 and a detection device 200, and the camera 100 and the detection device 200 can be connected through a wired or wireless communication network.

[0081] Among them, the detection device 200 can be any device capable of performing video image processing. The camera 100 is used to capture a driving video and send the driving video to the detection device 200 in real time. After receiving the driving video, the detection device 200 can perform target vehicle detection and pedestrian detection on multiple frames of driving images, and then determine whether there is a behavior of a target vehicle illegally carrying passengers.

[0082] The vehicle illegal passenger-carrying detection method of the present application will be introduced in detail below with reference to the accompanying drawings and specific embodiments.

[0083] As Figure 2 shown, the embodiments of the present application provide a vehicle illegal passenger-carrying detection method, which can be executed by a detection device and includes the following steps:

[0084] Step S201, obtain multiple frames of driving images in the driving video to be detected.

[0085] Among them, the driving video can be obtained by the camera. By decoding each frame of the driving video and encoding it into a driving image in a preset format, multiple frames of driving images are obtained.

[0086] It should be noted that in the embodiments of the present application, the acquisition and use of the driving video all comply with the requirements of relevant national laws and regulations.

[0087] Step S202, perform target vehicle detection and pedestrian detection on multiple frames of driving images respectively, and obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; among them, the target vehicle is a vehicle restricted from carrying passengers.

[0088] In the embodiments of the present application, the target vehicle can be a vehicle not allowed to carry passengers, such as a tricycle, a freight vehicle, etc. When detecting multiple frames of driving images respectively, it is possible to detect each frame of driving image after decoding one frame of driving image, or to detect multiple frames of driving images respectively after decoding multiple frames of driving images. Specifically, a target detection model can be used to perform target vehicle detection and pedestrian detection on each frame of driving image.

[0089] Step S203, if the target vehicle detection result of any frame of driving image includes the detection information of the target vehicle, and the pedestrian detection result includes the detection information of multiple pedestrians, then determine the positional relationship and detection area overlap degree between the target vehicle and each pedestrian in any frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in any frame of driving image.

[0090] Among them, when a target vehicle and pedestrians are included in any frame of driving image, by performing target vehicle detection and pedestrian detection on the driving image, detection information of the target vehicle and detection information of the pedestrians can be obtained. Considering that there is a driver in the target vehicle, generally, the number of drivers in the target vehicle can be one or two; assuming the number of drivers is 1, when at least two pedestrians are included in any frame of driving image, it can be further determined whether the target vehicle in the driving image is carrying people; in addition, the number of drivers may also be 2, and at this time, when at least three pedestrians are included in any frame of driving image, it can be further determined whether the target vehicle in the driving image is carrying people.

[0091] Furthermore, according to the detection information of the target vehicle and the detection information of each pedestrian, the positional relationship and the overlap degree of the detection regions between the target vehicle and each pedestrian can be determined, and based on this positional relationship and the overlap degree of the detection regions, it can be determined whether a pedestrian is on the target vehicle.

[0092] It should be noted that multiple target vehicles may also be included in any frame of driving image. At this time, the positional relationship between each target vehicle and each pedestrian can be determined, so as to determine the number of pedestrians on each target vehicle. In the following embodiments of the present application, one target vehicle is taken as an example for illustration.

[0093] Assume that there is one driver on the target vehicle. When there is a target vehicle in any frame of driving image, but there is only one pedestrian, steps S203 and the following step S204 may not be executed, and it can be directly determined that the target vehicle in this frame of driving image is in an unoccupied state.

[0094] In the above step S203, the detection information of each target vehicle may include the target vehicle detection region and the position information of the target vehicle detection region, and the detection information of each pedestrian may include the pedestrian detection region and the position information of the pedestrian detection region.

[0095] Based on this, when performing the above step S203, the following steps A1 - A3 can be specifically executed:

[0096] Step A1: According to the position information of the target vehicle detection region and the position information of each pedestrian detection region in any frame of driving image, determine the positional relationship between the center coordinates of the target vehicle detection region and each pedestrian detection region.

[0097] Among them, the positional relationship between the center coordinates of the target vehicle detection region and each pedestrian detection region includes: the center coordinates of each pedestrian detection region are within the target vehicle detection region, or the center coordinates of each pedestrian detection region are not within the target vehicle detection region.

[0098] Step A2: Use the positional relationship between the center coordinates of the target vehicle detection area and each pedestrian detection area as the positional relationship between the target vehicle and each pedestrian.

[0099] Step A3: Determine the ratio of the intersection to the union of the target vehicle detection area and each pedestrian detection area, and use the ratio of the intersection to the union as the overlap degree of the detection areas between the target vehicle and each pedestrian.

[0100] Exemplarily, as Figure 3 shown, A represents the pedestrian detection area, B represents the target vehicle detection area, and the ratio of the intersection to the union of the target vehicle detection area and the pedestrian detection area

[0101] Step S204: Based on the positional relationships and detection area overlap degrees between the target vehicle in any frame of driving image and multiple pedestrians respectively, determine the occupancy determination result of the target vehicle in any frame of driving image.

[0102] Among them, based on the positional relationship and detection area overlap degree between the target vehicle and each pedestrian, it can be determined whether each pedestrian is on the target vehicle, thereby determining the number of pedestrians on the target vehicle. When the number of pedestrians is greater than a preset number (for example, 1), it can be determined that the target vehicle is occupied.

[0103] Generally, when the center coordinate of a pedestrian detection area is within a target vehicle detection area, it can be considered that the pedestrian is on the target vehicle corresponding to the target vehicle detection area. Further, in order to more accurately determine the pedestrians on the target vehicle and avoid false occupancy detection caused by passing pedestrians, on the basis of considering the positional relationship between the center coordinate of the pedestrian detection area and the target vehicle detection area, the overlap degree between the pedestrian detection area and the target vehicle detection area can also be considered simultaneously, that is, the ratio of the intersection to the union of the pedestrian detection area and the target vehicle detection area.

[0104] Optionally, when performing the above Step S204, the following Steps B1 - B2 can be executed:

[0105] Step B1: Determine the target pedestrians among the multiple pedestrians whose overlap degree with the detection area of the target vehicle reaches a preset value.

[0106] Among them, the overlap degree of the detection areas between the target vehicle and the pedestrian, that is, the ratio of the intersection to the union of the target vehicle detection area and the pedestrian detection area (hereinafter referred to as the intersection - union ratio). The above - mentioned preset value can be set as needed, for example, 0.6, 0.7, etc., and is not limited here.

[0107] Step B2: If the number of target pedestrians exceeds the preset number, and the central coordinates of the pedestrian detection area of each target pedestrian are within the target vehicle detection area, then determine that the target vehicle in any frame of driving image is in a manned state.

[0108] Among them, the preset number can be the number of drivers of the target vehicle, usually one or two. That is to say, when the number of pedestrians on the target vehicle exceeds the number of drivers, it is considered that the target vehicle is manned. In this embodiment, when a pedestrian simultaneously meets the following two conditions, it can be considered that the pedestrian is on the target vehicle: Condition 1, the intersection over union (IoU) of the pedestrian detection area and the target vehicle detection area reaches a preset value; Condition 2, the central coordinates of the pedestrian detection area are within the target vehicle detection area; On the contrary, when a pedestrian does not meet one of the above two conditions, it is considered that the pedestrian is a passing pedestrian around the target vehicle.

[0109] Exemplarily, as Figure 4 shown, a frame of driving image includes a target vehicle detection area and two pedestrian detection areas. The central coordinates of these two pedestrian detection areas are both within the target vehicle detection area, and moreover, the overlap degree of each of these two pedestrian detection areas with the target vehicle detection area reaches the preset value. At this time, it can be determined that the target vehicle in this frame of driving image is in a manned state.

[0110] In the embodiments of the present application, when determining the manned determination result of the target vehicle, not only the positional relationship between the target vehicle and the pedestrian is considered, but also the intersection over union of the pedestrian detection area and the target vehicle area is considered, which can avoid false detection of illegal manned caused by passing pedestrians. At the same time, considering the driver of the target vehicle, when the number of pedestrians on the target vehicle exceeds the preset number, it is considered that the target vehicle is in a manned state.

[0111] Step S205: Based on the manned determination results of the target vehicle corresponding to each frame of driving image obtained, determine whether the target vehicle is illegally manned.

[0112] After obtaining each frame of driving image containing the target vehicle, when there are multiple driving images in which the target vehicle is in a manned state, it can be determined that the target vehicle is illegally manned. In this way, it can avoid misjudgment of illegal manned caused by false detection of manned in a single frame of driving image.

[0113] Optionally, when determining whether the target vehicle is illegally manned in step S205, it can be determined by the following method:

[0114] If, among the frames of driving images containing the target vehicle, the proportion of the number of driving images in which the target vehicle is in a manned state reaches a preset proportion, and the target vehicle is in a non-stationary state in each frame of driving image, then determine that the target vehicle is illegally manned.

[0115] Among them, the preset proportion can be set as needed. Suppose the preset proportion is 50%. If a certain target vehicle appears in 3 seconds of the driving video and the target vehicle is in a non-stationary state, by detecting each frame of the driving image in these 3 seconds of the driving video, the manned determination result of the target vehicle corresponding to each frame of the driving image is obtained. If the manned determination result of more than 50% of the driving images is the manned state, it can be determined that the target vehicle is carrying passengers illegally. At this time, an alarm message can be pushed. In this way, the misdetection of the manned state of a single-frame driving image caused by the overlap of space-time positions can be avoided.

[0116] In the embodiments of the present application, target vehicle and pedestrian detections are performed on a continuous multi-frame driving image. For each frame of the driving image, if a target vehicle and a pedestrian are detected, the manned determination result of the target vehicle in this frame of the driving image is determined, for example: the manned state or the unmanned state; then, according to the manned determination results of the target vehicles in the multi-frame driving images, it is determined whether the target vehicle is carrying passengers illegally. In this way, it can be quickly and effectively detected whether the target vehicle is carrying passengers illegally.

[0117] Next, an exemplary introduction is given to the method of performing target vehicle detection and pedestrian detection on the driving image in step S202 above.

[0118] In some embodiments, as Figure 5 shown, the above step S202 may include the following steps:

[0119] Step S2021, perform target vehicle detection and pedestrian detection on each of the multi-frame driving images through a target detection model, and obtain the target vehicle detection results and pedestrian detection results of each of the multi-frame driving images; among them, the target detection model at least includes a spatial attention mechanism network and a channel attention mechanism network, and the spatial attention mechanism network and the channel attention mechanism network are used to focus on the visible area of pedestrians in the driving image.

[0120] In the embodiments of the present application, the target detection model may introduce the above spatial attention mechanism network and channel attention mechanism network on the basis of the target detection network. For example, the target detection network may be an SSD (Single Shot MultiBox Detector) target detection network or other target detection networks, so that when the target detection model detects pedestrians, it pays more attention to the visible area of pedestrians, avoids missed detection due to pedestrians being blocked, improves the recall rate of pedestrian detection, and thus more accurately detects whether the target vehicle is carrying passengers illegally.

[0121] Next, an exemplary introduction is given to the structure of the target detection model in the embodiments of the present application.

[0122] In some alternative embodiments, the object detection model further includes a feature pyramid network, a first feature fusion network, and a second feature fusion network, which are composed of a plurality of sequentially connected convolutional layers;

[0123] The first convolutional layer is connected to the spatial attention mechanism network, and some of the convolutional layers in the plurality of convolutional layers and the spatial attention mechanism network are respectively connected to the first feature fusion network; the second convolutional layer is connected to the channel attention mechanism network, and some of the convolutional layers in the plurality of convolutional layers and the channel attention mechanism network are respectively connected to the second feature fusion network.

[0124] Exemplarily, as Figure 6 shown, the object detection model can use the SSD object detection model as the baseline network model, and respectively introduce the Residual Network (ResNet), the feature pyramid networks (FPN), the spatial attention mechanism network, the channel attention mechanism network, the first feature fusion network, and the second feature fusion network to reconstruct the object detection network.

[0125] Among them, the feature pyramid network includes convolutional layers such as Conv3 (Convolution3, the 3rd convolution), Conv4, Conv5, Conv6_2 (the 2nd convolutional layer in the 6th convolution), Conv7_2, and Conv8_2, to achieve multi-scale fine-grained feature extraction. The spatial attention mechanism network, Conv3, Conv4, and Conv5 form a residual network. Conv3 is connected to the spatial attention mechanism network, and the spatial attention mechanism network, Conv4, and Conv5 are respectively connected to the first feature fusion network; Conv4 is connected to the channel attention mechanism network, and the channel attention mechanism network, Conv5, and Conv6_2 are respectively connected to the second feature fusion network.

[0126] In the embodiments of the present application, considering the problem of pedestrian occlusion, a spatial attention mechanism and a channel attention mechanism are introduced into the above object detection model to focus on the unoccluded area (i.e., the visible area) of the pedestrian, increase the feature weight of the key parts of the pedestrian, so as to avoid the influence of interference information such as background occlusion. Different attention mechanisms are adopted for the classification and localization aspects in object detection: the spatial attention mechanism is adopted in the localization branch, and the channel attention mechanism is adopted in the classification branch.

[0127] Among them, the spatial attention mechanism focuses on the spatial position of the visible area of the pedestrian. As Figure 7As shown in the figure, the spatial attention mechanism network uses a Down-Up (up and down) stacked hourglass feature extraction network based on residual blocks to screen the key feature information of the detection target, and increase the context information association of the detection target. The entire network structure of the stacked hourglass feature extraction network is composed of fully convolutional layers. Among them, Max pool refers to max pooling, pad 1 means padding the edge by 1, and stride2 means the pooling stride is 2. The network structure first uses downsampling operations combined with residual blocks to compress the input feature map, and can effectively filter out noise and background information while retaining the multi-scale feature information of the target. Next, bilinear upsampling operations are used to increase the dimension of the feature map, ensuring that the feature map can output a feature vector of a specified dimension after passing through the residual attention network module to support the subsequent convolutional operations of the entire network.

[0128] The channel attention mechanism focuses on the feature channels of the visible area of the pedestrian, and different feature channels encode the features of different parts of the pedestrian. As Figure 8 shown, the channel attention mechanism network first performs global pooling and max pooling on the feature vectors of the classification branch respectively; the feature vector F avg after global pooling and the feature vector F max after max pooling are sent to the fully connected layers (FC), and "compression" and "stretching" operations are performed on these two weight vectors. Specifically, these two feature vectors are sequentially input into FC1 (scale of 1*1*16) and FC2 (scale of 1*1*256), and the ReLU (Rectified linear unit) is used for activation between FC1 and FC2. ReLU is the activation function of the neuron; then, the components of the feature vector are restricted between 0 and 1 through the sigmoid function of the neural network activation function, and the feature vectors obtained through global pooling and max pooling are weighted and fused to obtain the final feature representation. This method is different from only using average pooling operations in the original structure. We use a dual-channel operation method of global pooling and max pooling, which can highlight the main features while retaining the average features of each channel, making the network pay more attention to the visible parts of the target.

[0129] Furthermore, in order to enable the target feature map output by the spatial attention mechanism network to provide context information at the position where the target needs to be detected, the target feature map can be fused with higher-level context features. Among them, the higher-level context features can be Conv4 and Conv5 in the above Figure 6 , and the target feature map and the context features are fused through the first feature fusion network.

[0130] Similarly, in order to enable the target feature map output by the channel attention mechanism network to provide context information at the positions where the targets need to be detected, the target feature map can be fused with higher-level context features. Among them, the higher-level context features can be Conv5 and Conv6_2 in the above Figure 6 . The target feature map and the context features are fused through the second feature fusion network.

[0131] The structures of the above first feature fusion network and second feature fusion network are as Figure 9 shown. Since the feature maps of different layers have different spatial sizes and cannot be directly weighted and fused, before fusing through concatenated features, deconvolution is performed on the context features to make them have the same spatial size as the target features, achieving the alignment of feature sizes and the number of channels. At the same time, the normalization operation before concatenating features is very important because each feature value in different layers has a different scale. Therefore, batch normalization (e.g., normalization by L2 norm) and ReLU activation are performed after each layer. Finally, the target features and the context features are connected by stacking features.

[0132] The object detection model with the above structure in the embodiments of the present application can achieve multi-scale fine-grained feature extraction through the feature pyramid network. At the same time, in order to better achieve the accurate detection and separation of pedestrians and target vehicles, the spatial attention mechanism and the channel attention mechanism are introduced, which can ensure the detection accuracy of pedestrians and target vehicles and avoid the occurrence of missed detections, so as to more accurately detect whether the target vehicle is illegally carrying people.

[0133] The above object detection network is obtained by training the initial object detection network with an image sample set. The loss function in the training process includes a classification loss function, a regression loss function, and a loss function of the spatial attention mechanism network. Among them, the regression loss function uses the intersection over union of the predicted pedestrian region and the actual pedestrian visible region of each image sample as the weight of the loss value of the image sample.

[0134] Specifically, in the embodiments of the present application, a multi-task loss function is used to jointly optimize the parameters of each network. The loss function is composed of the above classification loss function, regression loss function, and loss function of the spatial attention mechanism network, as shown in the following formula:

[0135]

[0136] Among them, L c (p n ,p n *) is the classification loss function, and its basic form is a weighted cross-entropy loss function. Its main purpose is to improve the problem of extreme imbalance between positive and negative samples in regression-based object detection algorithms; Mc is the number of all target bounding boxes predicted in each training image; A represents the number of categories that can be detected, and p n , p n *, respectively, represent the category probability of the nth predicted target bounding box and the corresponding actual category; the regression loss function L r (t, t*) is the regression loss function, which can independently design the weight size according to different occlusion degrees. By using the intersection over ground truth (IOG) between the predicted target bounding box and the boundary box of the actual target visible area instead of the Intersection-over-Union (IOU) between the predicted target bounding box and the boundary box of the actual target area, the loss function weight generated by each positive sample is calculated, which can better handle the target occlusion problem; t and t* respectively represent the detection coordinate box and the true coordinate box of the target bounding box; M r is the number of target bounding boxes that are only considered as background among all the target bounding boxes; L a (m, m*) is the loss function of the spatial attention mechanism sub-network, which is actually a cross-entropy loss function based on masked pixels; are the mask generated by the spatial attention mechanism and its corresponding mask label respectively; λ1 and λ2 are parameters used to balance the sub-loss functions and can be set as needed, for example, both are set to 1.

[0137] During the training process, in order to further solve the problem of missed detection caused by pedestrians being occluded, when calculating the regression loss function, for positive image samples, the intersection over ground truth (IOG) between the predicted pedestrian bounding box and the boundary box of the actual pedestrian visible area is used as the weight of the loss function generated by this positive image sample. That is, if the predicted positive image sample bounding box overlaps more with the pedestrian visible area, then the loss it generates is more credible and a higher weight is assigned, otherwise a lower weight is assigned. In this way, the trained object detection model can accurately detect the visible area of pedestrians, avoid the problem of missed detection caused by pedestrians being occluded, and thus more accurately detect whether the target vehicle is illegally carrying passengers.

[0138] As Figure 10 shown, based on the same inventive concept, an embodiment of the present application provides a vehicle illegal passenger carrying detection device, including:

[0139] An acquisition module 101, configured to acquire multiple frames of driving images in a driving video to be detected;

[0140] A detection module 102, configured to perform target vehicle detection and pedestrian detection on multiple frames of driving images respectively, and obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target vehicle is a vehicle restricted from carrying passengers;

[0141] A position determination module 103, configured to, if the detection result of the target vehicle in any frame of driving image includes the detection information of the target vehicle and the pedestrian detection result includes the detection information of multiple pedestrians, determine the positional relationship and detection area overlap degree between the target vehicle and each pedestrian in any frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in any frame of driving image;

[0142] A status determination module 104, configured to determine the manned determination result of the target vehicle in any frame of driving image according to the positional relationship and detection area overlap degree between the target vehicle and multiple pedestrians in any frame of driving image;

[0143] A manned determination module 105, configured to determine whether the target vehicle is illegally carrying passengers based on the manned determination results of the target vehicle corresponding to each frame of driving image obtained.

[0144] In some alternative embodiments, the detection module 102 is further configured to:

[0145] Perform target vehicle detection and pedestrian detection on multiple frames of driving images respectively through a target detection model, and obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target detection model at least includes a spatial attention mechanism network and a channel attention mechanism network, and the spatial attention mechanism network and the channel attention mechanism network are used to focus on the visible area of pedestrians in the driving image.

[0146] In some alternative embodiments, the target detection model further includes a feature pyramid network, a first feature fusion network, and a second feature fusion network composed of multiple sequentially connected convolutional layers;

[0147] The first convolutional layer is connected to the spatial attention mechanism network, and some of the convolutional layers in the multiple convolutional layers and the spatial attention mechanism network are respectively connected to the first feature fusion network;

[0148] The second convolutional layer is connected to the channel attention mechanism network, and some of the convolutional layers in the multiple convolutional layers and the channel attention mechanism network are respectively connected to the second feature fusion network.

[0149] In some alternative embodiments, the target detection network is obtained by training an initial target detection network with an image sample set, and the loss function in the training process includes a classification loss function, a regression loss function, and a loss function of the spatial attention mechanism network;

[0150] Among them, the regression loss function uses the intersection over union of the predicted pedestrian area and the actual pedestrian visible area of each image sample as the weight of the loss value of the image sample.

[0151] In some alternative embodiments, the detection information of the target vehicle includes the target vehicle detection area and the position information of the target vehicle detection area, and the detection information of each pedestrian includes the pedestrian detection area and the position information of the pedestrian detection area;

[0152] The position determination module 103 is further configured to:

[0153] According to the position information of the target vehicle detection area and the position information of each pedestrian detection area in any frame of driving image, determine the positional relationship between the center coordinates of the target vehicle detection area and each pedestrian detection area;

[0154] Use the positional relationship between the center coordinates of the target vehicle detection area and each pedestrian detection area as the positional relationship between the target vehicle and each pedestrian;

[0155] Determine the ratio of the intersection and union of the target vehicle detection area and each pedestrian detection area, and use the ratio of the intersection and union as the overlap degree of the detection areas of the target vehicle and each pedestrian.

[0156] In some alternative embodiments, the status determination module 104 is further configured to:

[0157] Determine the target pedestrians among multiple pedestrians whose overlap degree with the target vehicle detection area reaches a preset value;

[0158] If the number of target pedestrians exceeds a preset number and the center coordinates of the pedestrian detection area of each target pedestrian are within the target vehicle detection area, determine that the target vehicle in any frame of driving image is in a manned state.

[0159] In some alternative embodiments, the manned determination module 105 is further configured to:

[0160] If, in each frame of driving image, the proportion of the number of driving images in which the target vehicle is in a manned state reaches a preset proportion and the target vehicle is in a non-stationary state in each frame of driving image, determine that the target vehicle is illegally carrying passengers.

[0161] Since the principle of the above device for solving problems is similar to the method for detecting illegal passenger-carrying of vehicles, the implementation of the above device can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0162] As Figure 11 shown, based on the same inventive concept, an embodiment of the present application provides a device for detecting illegal passenger-carrying of vehicles, and the detection device includes: a processor 1101 and a data receiving unit 1102.

[0163] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0164] The data receiving unit 1102 is configured to: receive the driving video to be detected;

[0165] The processor 1101 is configured to:

[0166] Obtain multiple frames of driving images in the driving video to be detected;

[0167] Perform target vehicle detection and pedestrian detection on multiple frames of driving images respectively to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target vehicle is a vehicle restricted from carrying people;

[0168] If the target vehicle detection result of any frame of driving image includes the detection information of the target vehicle, and the pedestrian detection result includes the detection information of multiple pedestrians, then according to the detection information of the target vehicle and the detection information of each pedestrian in any frame of driving image, determine the position relationship and detection area overlap degree between the target vehicle and each pedestrian in any frame of driving image;

[0169] According to the position relationship and detection area overlap degree between the target vehicle in any frame of driving image and multiple pedestrians respectively, determine the manned determination result of the target vehicle in any frame of driving image;

[0170] Based on the manned determination results of the target vehicles corresponding to each frame of driving image obtained, determine whether the target vehicle is illegally carrying people.

[0171] In some exemplary embodiments, when performing target vehicle detection and pedestrian detection on multiple frames of driving images respectively to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image, the processor 1101 is further configured to:

[0172] Perform target vehicle detection and pedestrian detection on multiple frames of driving images respectively through a target detection model to obtain the target vehicle detection results and pedestrian detection results of each frame of driving image; wherein, the target detection model at least includes a spatial attention mechanism network and a channel attention mechanism network, and the spatial attention mechanism network and the channel attention mechanism network are used to focus on the visible area of pedestrians in the driving image.

[0173] In some exemplary embodiments, the object detection model further includes a feature pyramid network, a first feature fusion network, and a second feature fusion network, which are composed of a plurality of sequentially connected convolutional layers;

[0174] The first convolutional layer is connected to the spatial attention mechanism network, and some of the convolutional layers in the plurality of convolutional layers and the spatial attention mechanism network are respectively connected to the first feature fusion network;

[0175] The second convolutional layer is connected to the channel attention mechanism network, and some of the convolutional layers in the plurality of convolutional layers and the channel attention mechanism network are respectively connected to the second feature fusion network.

[0176] In some exemplary embodiments, the object detection network is obtained by training an initial object detection network with an image sample set, and the loss function in the training process includes a classification loss function, a regression loss function, and a loss function of the spatial attention mechanism network;

[0177] Among them, the regression loss function uses the intersection over union of the predicted pedestrian region and the actual pedestrian visible region of each image sample as the weight of the loss value of the image sample.

[0178] In some exemplary embodiments, the detection information of each target vehicle includes the target vehicle detection region and the position information of the target vehicle detection region, and the detection information of each pedestrian includes the pedestrian detection region and the position information of the pedestrian detection region;

[0179] When determining the positional relationship and detection region overlap degree between the target vehicle and each pedestrian in any frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in any frame of driving image, the processor 1101 is specifically configured to:

[0180] According to the position information of the target vehicle detection region and the position information of each pedestrian detection region in any frame of driving image, determine the positional relationship between the center coordinates of the target vehicle detection region and each pedestrian detection region;

[0181] Use the positional relationship between the center coordinates of the target vehicle detection region and each pedestrian detection region as the positional relationship between the target vehicle and each pedestrian;

[0182] Determine the ratio of the intersection and union of the target vehicle detection region and each pedestrian detection region, and use the ratio of the intersection and union as the detection region overlap degree between the target vehicle and each pedestrian.

[0183] In some exemplary embodiments, when determining the passenger-carrying determination result of the target vehicle in any frame of driving image according to the positional relationship and detection region overlap degree between the target vehicle and multiple pedestrians in any frame of driving image, the processor 1101 is further configured to:

[0184] Determine target pedestrians among multiple pedestrians whose overlap degree with the detection area of the target vehicle reaches a preset value;

[0185] If the number of target pedestrians exceeds a preset number, and the central coordinates of the pedestrian detection area of each target pedestrian are within the target vehicle detection area, then determine that the target vehicle in any frame of driving image is in a manned state.

[0186] In some exemplary embodiments, when determining whether the target vehicle illegally carries passengers based on the manned determination results of the target vehicles corresponding to each frame of driving image obtained, the processor 1101 is specifically configured to:

[0187] If the proportion of the number of driving images in which the target vehicle is in a manned state among all frames of driving images reaches a preset proportion, and the target vehicle is in a non-stationary state in all frames of driving images, then determine that the target vehicle illegally carries passengers.

[0188] The embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to make a computer execute any one of the vehicle illegal passenger-carrying detection methods in the above embodiments.

[0189] The computer-readable storage medium in the above embodiments may be any available medium or data storage device that can be accessed by the processor in the device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSDs), etc.

[0190] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0191] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0194] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A method for detecting illegal passenger carrying by vehicles, characterized in that, Including: Obtaining multiple driving images in a driving video to be detected; Performing target vehicle detection and pedestrian detection on each of the multiple driving images respectively to obtain the target vehicle detection result and pedestrian detection result of each of the multiple driving images; wherein, the target vehicle is a vehicle restricted from carrying passengers; If the target vehicle detection result of any one of the driving images includes the detection information of the target vehicle and the pedestrian detection result includes the detection information of multiple pedestrians, then according to the detection information of the target vehicle and the detection information of each pedestrian in the any one of the driving images, determining the positional relationship and detection area overlap degree between the target vehicle and each pedestrian in the any one of the driving images; Determining the target pedestrians among the multiple pedestrians whose detection area overlap degree with the target vehicle reaches a preset value; if the number of the target pedestrians exceeds a preset number and the center coordinates of the pedestrian detection area of each target pedestrian are within the target vehicle detection area, then determining that the target vehicle in the any one of the driving images is in a state of carrying passengers; Based on the passenger-carrying determination results of the target vehicle corresponding to each of the obtained driving images, determining whether the target vehicle illegally carries passengers.

2. The method according to claim 1, characterized in that, The performing target vehicle detection and pedestrian detection on each of the multiple driving images respectively to obtain the target vehicle detection result and pedestrian detection result of each of the multiple driving images includes: Performing target vehicle detection and pedestrian detection on each of the multiple driving images respectively through a target detection model to obtain the target vehicle detection result and pedestrian detection result of each of the multiple driving images; wherein, the target detection model at least includes a spatial attention mechanism network and a channel attention mechanism network, and the spatial attention mechanism network and the channel attention mechanism network are used to focus on the visible area of pedestrians in the driving image.

3. The method according to claim 2, characterized in that, The target detection model further includes a feature pyramid network, a first feature fusion network, and a second feature fusion network composed of multiple sequentially connected convolutional layers; The first convolutional layer is connected to the spatial attention mechanism network, and some of the convolutional layers among the multiple convolutional layers and the spatial attention mechanism network are respectively connected to the first feature fusion network; The second convolutional layer is connected to the channel attention mechanism network, and some of the convolutional layers among the multiple convolutional layers and the channel attention mechanism network are respectively connected to the second feature fusion network.

4. The method according to claim 3, characterized in that, The target detection network is obtained by training an initial target detection network with an image sample set, and the loss function in the training process includes a classification loss function, a regression loss function, and a loss function of the spatial attention mechanism network; Wherein, the regression loss function uses the intersection over union of the predicted pedestrian area and the actual visible area of pedestrians of each image sample as the weight of the loss value of the image sample.

5. The method according to any one of claims 1 to 4, characterized in that, The detection information of the target vehicle includes the target vehicle detection area and the position information of the target vehicle detection area, and the detection information of each pedestrian includes the pedestrian detection area and the position information of the pedestrian detection area; Determining the positional relationship and detection area overlap degree between the target vehicle and each pedestrian in the arbitrary frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in the arbitrary frame of driving image includes: Determining the positional relationship of the central coordinates of the target vehicle detection area and the detection area of each pedestrian according to the position information of the target vehicle detection area and the position information of the detection area of each pedestrian in the arbitrary frame of driving image; Using the positional relationship of the central coordinates of the target vehicle detection area and the detection area of each pedestrian as the positional relationship between the target vehicle and each pedestrian; Determining the ratio of the intersection and union of the target vehicle detection area and the detection area of each pedestrian, and using the ratio of the intersection and union as the detection area overlap degree between the target vehicle and each pedestrian.

6. The method according to any one of claims 1 to 4, characterized in that, Determining whether the target vehicle illegally carries passengers based on the obtained results of judging whether the target vehicle carries passengers corresponding to each frame of driving image includes: If, among the frames of driving images, the proportion of the number of driving images in which the target vehicle is in a passenger-carrying state reaches a preset proportion, and the target vehicle is in a non-stationary state in the frames of driving images, it is determined that the target vehicle illegally carries passengers.

7. A device for detecting illegal passenger carrying by vehicles, characterized in that, Including: An acquisition module, configured to acquire multiple frames of driving images in a driving video to be detected; A detection module, configured to respectively perform target vehicle detection and pedestrian detection on the multiple frames of driving images to obtain the target vehicle detection results and pedestrian detection results of the multiple frames of driving images respectively; wherein, the target vehicle is a vehicle restricted from carrying passengers; A position determination module, configured to, if the target vehicle detection result of an arbitrary frame of driving image includes the detection information of the target vehicle and the pedestrian detection result includes the detection information of multiple pedestrians, determine the positional relationship and detection area overlap degree between the target vehicle and each pedestrian in the arbitrary frame of driving image according to the detection information of the target vehicle and the detection information of each pedestrian in the arbitrary frame of driving image; A state determination module, configured to determine the target pedestrians whose detection area overlap degree with the target vehicle reaches a preset value among the multiple pedestrians; if the number of target pedestrians exceeds a preset number and the central coordinates of the pedestrian detection area of each target pedestrian are within the target vehicle detection area, it is determined that the target vehicle in the arbitrary frame of driving image is in a passenger-carrying state; A passenger-carrying determination module, configured to determine whether the target vehicle illegally carries passengers based on the obtained results of judging whether the target vehicle carries passengers corresponding to each frame of driving image.

8. A device for detecting illegal passenger carrying by vehicles, characterized in that, Including a processor and a data receiving unit; The data receiving unit is configured to: receive the driving video to be detected; The processor is configured to: execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, in which a computer program is stored, characterized in that: When the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Vehicle illegal video processing method and device, computer equipment and storage medium

    CN110675637A

  • Video-based motor vehicle illegal manned detection method

    CN114120250A