Ghost probe detection system based on computer vision and deep learning

By using computer vision and deep learning technology in the ghost probe detection system, combined with YOLOv5 and Deepsort algorithms, accurate detection and early alarm of ghost probe accidents are achieved, and the problems of inaccurate detection and high cost in the existing technology are solved, and detection efficiency and reliability are improved.

CN119964041APending Publication Date: 2025-05-09NANJING HONGSHANYUAN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311474754.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When detecting ghost probe accidents, the existing technology is greatly affected by the ambient temperature and it is difficult to distinguish different types of items, and infrared detection technology is inaccurate; while the lidar detection technology is expensive and is disturbed in bad weather, which cannot effectively solve the detection problem of ghost probe accidents.

Method used

The ghost probe detection system based on computer vision and deep learning is adopted, combined with the YOLOv5 object detection model and the Deepsort multi-object tracking algorithm, and the processing and analysis of video stream data is carried out to accurately detect and track multiple targets, and speed, distance and direction are calculated based on the target's position and trajectory information, and alarm rules are set to alarm in advance.

Benefits of technology

Accurate detection and tracking of multiple targets in the video stream is achieved, and can work effectively under different weather conditions, reduce detection costs, and reduce the possibility of accidents through early alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a ghost probe detection system based on computer vision and deep learning, which is applied to the field of intelligent transportation. The method comprises the following steps: S1, acquiring video stream data; s2, training and loading the model; s3, initializing the tracking module; s4, performing target detection and recognition to obtain a detection result; s5, reasoning is carried out, a data structure is adjusted, and the data structure is transmitted into Deepsort; s6, calling a function to calculate information such as speed, distance and direction of the target; s7, alarm information is sent to a roadside information board and a vehicle-mounted terminal through a TCP protocol and a V2X protocol; and S8, the roadside information board and the vehicle-mounted terminal carry out sound-light alarm. According to the method, the YOLOv5 target detection algorithm and the Deepsort multi-target tracking algorithm are combined, accurate detection and tracking of the multiple targets in the video stream are achieved, the speed, the distance and the movement direction of the targets are calculated according to the position and track information of the targets, analysis and judgment of target behaviors are achieved, and the target tracking accuracy is improved. The judgment capability of the ghost probe phenomenon is further improved, and the possibility of accidents caused by the ghost probe is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent traffic detection technology, and in particular to a ghost detection system based on computer vision and deep learning. Background Art

[0002] "Ghost poking" accidents often occur at bus stops, zebra crossings, parking lots and other places where there are vehicles parked on both sides.

[0003] Traditional detection methods are based on infrared detection technology and laser radar detection technology. Infrared detection technology uses the different emission capabilities of infrared rays to objects of different temperatures. It obtains the temperature distribution image of the detected object by receiving infrared rays. However, it is greatly affected by the ambient temperature and it is difficult to distinguish different types of objects.

[0004] LiDAR detection technology emits a detection signal (laser beam) to the target and then obtains relevant information about the target, but it is expensive and will be greatly disturbed in rain, fog, wind and sand. With the development and application of deep learning technology, we propose a ghost detection method based on computer vision and deep learning, which can effectively solve the problems and shortcomings of existing technologies. Summary of the invention

[0005] The purpose of the present invention is to solve at least one of the technical problems existing in the prior art, and to provide a ghost detection system based on computer vision and deep learning, which combines the YOLOv5 target detection model and the Deepsort multi-target tracking algorithm to achieve accurate detection and tracking of multiple targets in a video stream.

[0006] The present invention also provides the above-mentioned ghost detection system based on computer vision and deep learning, comprising the following steps:

[0007] S1. Obtain video stream data: obtain the video stream data captured by the traffic camera by calling the improved method of the OpenCV library, and input the video stream data to the edge computing device MEC for processing;

[0008] S2. Model training and loading: making data sets, image preprocessing, YOLOv5 model training, and loading trained weight files;

[0009] S3. Tracking module initialization: import the YOLOv5 weight file and initialize the Deepsort multi-target tracking module;

[0010] S4. Target detection and recognition, and obtaining detection results: Use the trained weight file to detect the loaded video stream and identify the coordinates, category and other information of the target;

[0011] S5. Perform reasoning, adjust data structure and pass it into Deepsort: The detection results obtained in S4 are passed into the Deepsort module to perform reasoning, adjust data structure, NMS (non-maximum suppression) and other operations to track the trajectory of the detection target;

[0012] S6. Call the function to calculate the target's speed, distance, direction and other information, and determine whether to alarm:

[0013] 1) Calculate the speed and distance of the target based on its location information;

[0014] 2) According to the location information of the target, a judgment line is set in the video stream image, and the direction of the target is determined by the positional relationship between the coordinates of the target point and the judgment line, and recorded in the direction list;

[0015] 3) Set alarm rules according to needs and scenarios, and judge the speed and distance of the target according to the set alarm rules. If a target triggers the alarm rules, the system will send out an alarm signal;

[0016] S7. Send the alarm information to the roadside information board and the vehicle terminal respectively through TCP and V2X protocols: Send the alarm information to the roadside unit RSU and the roadside information board through the TCP protocol, and the roadside unit forwards the alarm information to the vehicle terminal through the V2X protocol.

[0017] S8. The roadside information board and the vehicle-mounted terminal make sound and light alarms: The roadside information board and the vehicle-mounted terminal make sound and light alarms (voice broadcast and text display) after receiving the alarm information.

[0018] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the improved method of S1 acquiring video stream includes: multi-threaded processing, reading the latest frame of the video, etc.

[0019] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the preparation of the data set in S2 includes: collecting target images, image target annotation, and constructing a training image set.

[0020] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the preprocessing of the training image set in S2 includes: graying, image enhancement, denoising, etc.

[0021] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the speed of the target in S6 is obtained by obtaining the center point position coordinates of the target through target data, and calculated using the Euclidean distance formula.

[0022] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the distance of the target in S6 is calculated based on the principle of similar triangles, using the ratio of the height of the target to the height of the camera being equal to the ratio of the distance between the target and the camera to the focal length.

[0023] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the direction of the target in S6 is determined by drawing a judgment line through OpenCV, and the position of the point relative to the straight line is determined by the positive and negative values ​​of the vector cross product, thereby obtaining the target's forward direction.

[0024] According to a ghost detection system based on computer vision and deep learning provided by the present invention, the alarm rule in S6 calculates and predicts whether there is a possibility of collision between the pedestrian and the vehicle's forward route through the target speed, direction and trajectory, thereby issuing an alarm.

[0025] The overall architecture of a ghost detection system provided by the present invention includes: a perception module, which is used to monitor the current area, perceive and obtain video stream data of the monitored area; an identification module, which is used to use the trained weight file to identify the above video stream data to obtain a detection result; a calculation module, which is used to use the above detection result to calculate the speed, distance, and direction of the target object, and predict and determine whether a ghost phenomenon occurs, and generate an alarm message; a communication module, which is used to forward the above alarm message to a roadside information board and a vehicle-mounted unit; and a terminal display module, which is used to perform an audible and visual alarm on the roadside information board and the vehicle-mounted unit.

[0026] According to the present invention, a software architecture of an edge computing device MEC of a ghost detection system is provided, wherein a computer program is stored, and when the computer program is executed, a ghost detection processing method is implemented.

[0027] Beneficial Effects

[0028] 1. The present invention combines the YOLOv5 target detection algorithm and the Deepsort multi-target tracking algorithm to achieve accurate detection and tracking of multiple targets in the video stream;

[0029] 2. The present invention uses the position and trajectory information of the target to calculate the speed and distance of the target, thereby realizing the analysis and judgment of the target behavior;

[0030] 3. The present invention designs different alarm rules for different usage scenarios, and realizes early alarm of the "ghost poking" phenomenon by judging the speed, distance, direction, etc. of the target. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The present invention is further described below in conjunction with the accompanying drawings and embodiments;

[0032] Figure 1 This is a schematic diagram of the overall process of a ghost detection system based on computer vision and deep learning in the present invention;

[0033] Figure 2 A schematic diagram of speed calculation of a ghost detection system based on computer vision and deep learning according to the present invention;

[0034] Figure 3 This is a schematic diagram of direction judgment of a ghost detection system based on computer vision and deep learning in the present invention.

[0035] Figure 4 A schematic diagram of the overall architecture of a ghost detection system of the present invention;

[0036] Figure 5 The figure is a software architecture diagram of MEC, an edge computing device in a ghost detection system. DETAILED DESCRIPTION

[0037] This section will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the accompanying drawings is to supplement the description of the text part of the specification with graphics, so that people can intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it cannot be understood as a limitation on the scope of protection of the present invention.

[0038] Reference Figure 1 , an embodiment of the present invention provides a ghost detection system based on computer vision and deep learning, which includes the following steps:

[0039] S1. Get video stream data:

[0040] First, you need to install the network camera in the area that needs to be monitored, adjust the angle and height of the camera to ensure that the scene that needs to be monitored can be captured; then connect the camera and the edge computing device to the same local area network through a network cable to ensure that the video stream data can be transmitted to the edge computing device; finally, obtain the video stream data captured by the camera through the edge computing device MEC;

[0041] S2. Model training and loading:

[0042] Collect images: First, you need to collect images with target objects. These images should be taken in actual application scenarios to ensure that the dataset covers a variety of backgrounds, lighting conditions, and object poses.

[0043] Label the object: For each image, you need to use a labeling tool to label the target object. Usually, a rectangular bounding box is used to label the location and size of the target object. In some cases, the category of the object also needs to be labeled.

[0044] Data preprocessing: Using data augmentation techniques, such as random cropping, scaling, translation, grayscale, and other operations, to transform images can increase the diversity and robustness of the data set.

[0045] Divide the data set: Divide the data set into a training set (train), a validation set (val), and a test set (test). The training set is used for model training and parameter optimization, the validation set is used to adjust the model's hyperparameters and evaluate the model's performance, and the test set is used to ultimately evaluate the performance of the trained model.

[0046] Generate label file: For each image, a corresponding label file needs to be generated, which contains the category of each target object and the location and size information of the bounding box.

[0047] Training model: Use the dataset to train the model. Use the implemented YOLOv5n model to train in the PyTorch framework to get a weight file.

[0048] Loading model: Load the trained weight file for detection. By calibrating the bounding box and determining the category of each detected target, the target's location, confidence, category and other information can be obtained.

[0049] S3. Tracking module initialization:

[0050] Initialize the Deepsort multi-target tracking module and use the target detection results to track multiple targets. By using the target's location, category and other information, the targets in adjacent frames are matched and tracked to form target trajectories;

[0051] S4. Target detection and recognition, and obtaining detection results:

[0052] The YOLOv5n model performs target detection on the preprocessed image, and obtains the target's location, confidence, category and other information by calibrating the bounding box and determining the category of each detected target;

[0053] S5. Perform inference, adjust data structure and pass it into Deepsort: load the video stream data obtained in S1 into the model of S2 and the module of S3;

[0054] S6. Call the function to calculate the target's speed, distance, direction and other information, and determine whether to alarm:

[0055] 1) Calculate the speed and distance of the target based on its location information;

[0056] The speed of the target

[0057] In the program, first set the initial position coordinates of each target (curr_x=0,curr_y=0), then obtain the target data x1,y1,x2,y2 of the first frame, then the center point position coordinates of the target are ((x1+x2) / 2,(y1+y2) / 2), that is (x,y). Finally, the displacement of the target can be calculated using the Euclidean distance formula:

[0058]

[0059] At the same time, the time of each frame can be obtained by the frame rate fps of the video stream:

[0060]

[0061] Finally, the velocity of the target can be found by dividing the displacement by the time:

[0062]

[0063] In the next frame, assign the values ​​of x and y to curr_x and curr_y, and get the new x and y, and repeat the calculation, such as Figure 2 As shown;

[0064] ●Distance to target:

[0065] In the program, we first need to obtain the target's location information and ID. We need to modify them according to the actual data and enter the initial values ​​as follows:

[0066] actual_height: the actual height of the object;

[0067] cutual_width: actual object width;

[0068] triangle_height: the height of the camera;

[0069] triangle_width: the width of the camera;

[0070] focal_length: The focal length of the camera.

[0071] According to the principle of similar triangles, the following proportional relationship can be established:

[0072]

[0073] Through deformation we can get:

[0074]

[0075] Returns the distance between the object and the camera, rounded to an integer.

[0076] 2) According to the location information of the target, a judgment line is set in the video stream image, and the direction of the target is determined by the positional relationship between the coordinates of the target point and the judgment line, and recorded in the direction list;

[0077] In the program, we first need to obtain the target's position information, and then use OpenCV to draw the judgment. Given two points pt1 and pt2 on the line, and the target coordinate point pt to be judged, we use the properties of the vector cross product to judge the relationship between pt and both sides of the line pt1pt2. The positive and negative values ​​of the vector cross product can determine whether the point is above or below the line.

[0078] If the judgment point pt is above the straight line, it returns 0.

[0079] pt1:(pt1.x,pt1.y,1)

[0080] pt2:(pt2.x,pt2.y,1)

[0081] pt:(pt.x,pt.y,1)

[0082]

[0083] Finally, two states are set. State 1 is when the judgment point is above the straight line, and state 2 is when the judgment point is below the straight line. From state 1 to state 2, the target is classified as southward; otherwise, it is northward. The target ID is stored in the direction list. The schematic diagram is as follows Figure 3 As shown, after the target direction is determined, the ID information of the target is stored in the corresponding direction storage list, and the counter is increased by 1 and displayed in the upper left corner of the detection screen.

[0084] 3) Set alarm rules and thresholds according to scene requirements, and judge the speed and distance of the target based on the set alarm rules and thresholds. If a target triggers the alarm rules, the system will issue an alarm signal.

[0085] S7. Send the alarm information to the roadside information board and the vehicle terminal respectively through TCP and V2X protocols: Send the alarm information to the roadside unit RSU and the roadside information board through the TCP protocol, and the roadside unit forwards the alarm information to the vehicle terminal through the V2X protocol.

[0086] S8. Roadside information board and vehicle terminal provide sound and light alarm: Output the detected target category, speed, distance and alarm information. The detection results can be output in the form of graphical interface, text file and voice notification, which is convenient for users to conduct further analysis and decision-making.

[0087] Reference Figure 4 , an overall architecture of a ghost detection system according to an embodiment of the present invention.

[0088] Reference Figure 5 , an embodiment of the present invention provides a software architecture of an edge computing device MEC of a ghost detection system.

[0089] Application of this system:

[0090] 1. This system obtains the video stream data of a certain area captured by the camera in real time, and performs target detection and tracking. At the same time, according to the set alarm rules and thresholds, it accurately judges the occurrence of "ghost peeking" phenomenon and issues an early alarm;

[0091] 2. The system will display the location, confidence, category and other information of the detected target and the alarm information in real time, and can output the data to related devices or storage media to facilitate users to conduct real-time monitoring and data analysis. The information is sent to the vehicle-mounted unit through the roadside unit RSU to remind the driver; it is sent to the roadside display screen through the edge computing device MEC for sound and light alarm to remind roadside pedestrians.

[0092] The embodiments of the present invention are described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in the technical field without departing from the purpose of the present invention.

Claims

1. A ghost detection system based on computer vision and deep learning, characterized in that: The following steps are involved: S1. Get video stream data; S2. Model training and loading; S3. Tracking module initialization; S4. Target detection and recognition, and obtaining detection results; S5. Perform reasoning, adjust the data structure and pass it into Deepsort; S6. Call the function to calculate the target's speed, distance, direction and other information, and determine whether to alarm; S7. Send the alarm information to the roadside information board and the vehicle terminal respectively through TCP and V2X protocols; S8. The roadside information board and the vehicle terminal will give sound and light alarms.

2. A ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The improved method for obtaining a video stream comprises: Acquire video stream data, wherein the video stream data is captured by a traffic camera and transmitted via an RTSP (real-time streaming) protocol; Dynamically determine whether the video stream has started, and stop reading frames if the video stream has not started or has ended; Use multi-threading to read video frames in parallel, and read the latest frame of the video through sub-threads, which has high real-time performance; It can detect the quality parameters of the video stream, such as frame rate, resolution, etc. in real time, and can set the sampling rate or frame rate of the video stream according to specific needs.

3. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The camera in S1 is installed in the area to be monitored, and the angle and height of the camera are adjusted.

4. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The camera in S1 and the edge computing device MEC are connected in the same local area network.

5. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The data set produced in S2 includes: According to the usage scenario, collect the target images to be detected (pedestrians, bicycles, electric vehicles, cars, trucks, buses); Annotating the target object on the target image to obtain a training image; A data set is constructed based on the labeled training images and divided into training set (train), validation set (val), and test set (test) according to a certain ratio.

6. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The image preprocessing in S2 includes graying, image enhancement, denoising and other operations on the images in the data set.

7. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The speed of the target in S6 is obtained by obtaining the center point position coordinates of the target through target detection data and calculated using the Euclidean distance formula.

8. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The distance of the target in S6 is calculated by the principle of similar triangles, that is, the ratio of the height of the target to the height of the camera is equal to the ratio of the distance between the target and the camera to the focal length, thereby obtaining the distance between the target and the camera.

9. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The direction of the target in S6 is determined by drawing a judgment line through OpenCV, and the position of the point relative to the straight line is determined by the positive and negative values ​​of the vector cross product, so as to obtain the moving direction of the target.

10. The ghost detection system based on computer vision and deep learning according to claim 1, characterized in that: The alarm rules in S6 include: Pedestrians and vehicles appear in the monitoring area at the same time; By calculating and predicting the target speed, direction and trajectory, there is a possibility of collision between pedestrians and vehicles.

11. An overall architecture of a ghost detection system, characterized in that: include: The perception module is used to monitor the current area, perceive and obtain the video stream data of the monitored area; An identification module is used to identify the video stream data using the trained weight file to obtain a detection result; A calculation module is used to calculate the speed, distance and direction of the target object using the above detection results, and to predict whether a ghosting phenomenon occurs and generate an alarm message; A communication module, used to forward the above alarm information to the roadside information board and the vehicle-mounted unit; Terminal display module, used for sound and light alarm on roadside information board and vehicle-mounted unit.

12. A MEC software architecture for edge computing devices in a ghost detection system, characterized in that: include: For storing computer programs; A ghost detection processing method for executing the S1-S7 process in claim 1 above.