Vehicle size recognition method, device, storage medium and product
By detecting the vertical coordinates and height of the center point of the small vehicle target box in the vehicle video stream and performing polynomial fitting, the problems of high cost and low efficiency of vehicle size classification labeling in the prior art are solved, and more accurate and fast vehicle size recognition is achieved.
Patent Information
- Application Number
- CN202410298702.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-03-15
AI Technical Summary
The existing vehicle size classification method has high labeling costs and low multi-objective classification efficiency, making it difficult to efficiently identify vehicle size in traffic road scenarios.
By obtaining video frames in the vehicle flow video stream for object detection, using deep learning models to identify the vertical coordinates and height of the center point of the small vehicle object detection frame, perform polynomial fitting, and calculate vehicle size categories based on polynomial expression and 3Sigma principle.
It realizes more accurate and flexible vehicle size recognition, reduces labeling and deployment costs, and improves recognition speed and accuracy in multi-target scenarios.
Smart Images

Figure CN118115809B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a vehicle size recognition method, device, storage medium and product. Background Art
[0002] With the acceleration of the urbanization process, the number of vehicles on the road has increased year by year, and the problem of road congestion is severe, bringing great challenges to traffic signal management and urban road construction. In recent years, intelligent transportation has used big data and artificial intelligence technologies to extract structured information of vehicles on the road, and provided guiding opinions for urban road construction and traffic signal optimization by obtaining the structured information of vehicles.
[0003] Obtaining the structured information of vehicles on the road is the basis for traffic road analysis. The current mainstream technical solution is to perform object detection and classification on vehicles on the road through deep learning methods. However, there are a large number of vehicle objects on the road, resulting in low efficiency of this classification method, and there is also diversity in the types of vehicles, and the classification accuracy is also low.
[0004] Currently, the methods based on video processing are roughly divided into two types: The first is to set empirical features by using traditional image processing methods in the video sequence, such as the patent documents with publication numbers CN107122758A and CN110751119A. This method is highly dependent on the image quality, and the robustness of the algorithm is low; the second is to classify vehicles once or multiple times through a deep learning classification model. The robustness of this method will be improved to some extent, but for the case of multiple targets, the processing efficiency of the algorithm is low.
[0005] The radar-vision device is a better solution in the intelligent transportation scenario in recent years. The roadside millimeter-wave radar has the advantages of accurate positioning, precise speed measurement, and all-weather operation, but it cannot provide rich target structured information; video processing can bring richer target structured information, but it cannot provide precise target positioning and accurate speed information; the radar-vision device is a more complete product that combines the advantages of both.
[0006] Existing vehicle size classification algorithms usually perform data annotation on different types of vehicles, and then construct a multi-classification model in combination with a deep learning algorithm, or construct multiple classification models in a cascaded manner, and finally classify the vehicle size according to the category. Taking trucks as an example, trucks can be divided into small trucks, medium trucks, and large trucks according to size, and trucks can be divided into cargo trucks, heavy trucks, etc. according to category. The current classification method has the following two problems:
[0007] (1) The annotation cost is high in the case of a large number of existing vehicle models, and the classification model needs to be re-optimized when transferred to a new usage scenario.
[0008] (2) In the traffic road scenario, it is usually a multi-target situation, and the efficiency of multi-target classification is low, which will take up a large amount of running time of the algorithm in actual use. Summary of the Invention
[0009] The purpose of the present invention is to provide a vehicle size recognition method, device, storage medium and product to solve the problems of high labeling cost of traditional methods and low efficiency of multi-target classification.
[0010] The present invention solves the above technical problems through the following technical solutions: A vehicle size recognition method, the recognition method includes the following steps:
[0011] Obtain a traffic flow video stream, perform object detection on the video frames in the traffic flow video stream to obtain vehicle object detection information; wherein, the vehicle object detection information includes the position, height, width and category of the object detection frame.
[0012] When the category of the object detection frame is a small vehicle, calculate the center point ordinate of the small vehicle object detection frame according to the position, height and width of the small vehicle object detection frame.
[0013] Collect data of the set corresponding to each sub-region according to the center point ordinate and height of the small vehicle object detection frame and the ordinate range of each sub-region; wherein, the sub-region is obtained by equally dividing a certain video frame in the image to be recognized or the traffic flow video stream along the image ordinate direction into N parts, and the sets corresponding to each sub-region are sorted in ascending order according to the ordinate range of each sub-region.
[0014] Calculate the average value of all data in each set, and use the average value as the average height value of the small vehicle detection frame corresponding to the set.
[0015] Perform polynomial fitting on N groups of set numbers and their average height values of small vehicle detection frames to obtain a polynomial expression of the set number and the average height value of the small vehicle detection frame.
[0016] Obtain the object detection frame to be recognized, and calculate the center point ordinate of the object detection frame to be recognized.
[0017] According to the center point ordinate of the object detection frame to be recognized, the height of the image or video frame to be recognized, the number of sub-regions, and the polynomial expression, obtain the average height value of the small vehicle detection frame corresponding to the center point ordinate of the object detection frame to be recognized.
[0018] Determine the vehicle size category of the object detection frame to be recognized according to the height of the object detection frame to be recognized and the average height value of the small vehicle detection frame corresponding to the center point ordinate of the object detection frame to be recognized.
[0019] Further, the specific implementation process of performing object detection on the video frames in the traffic flow video stream is as follows:
[0020] Obtain vehicle environment images, label vehicles with 5 seats or less in the vehicle environment images as small vehicles, and label vehicles with more than 5 seats as other categories, and construct a small vehicle detection data set;
[0021] Construct a deep learning object detection model, and use the small vehicle detection data set to train and validate the deep learning object detection model to obtain an object detection model;
[0022] Use the object detection model to perform object detection on the video frames in the traffic flow video stream to obtain vehicle object detection information.
[0023] Further, according to the vertical coordinate of the center point and the height of the small vehicle object detection frame and the vertical coordinate ranges of each sub-region, data collection for the set corresponding to each sub-region is performed. The specific implementation process is as follows:
[0024] Compare the vertical coordinate of the center point of the small vehicle object detection frame with the vertical coordinate ranges of each sub-region. When the vertical coordinate of the center point of the small vehicle object detection frame is within the vertical coordinate range of the sub-region, add the height of the small vehicle object detection frame to the set corresponding to this sub-region;
[0025] Calculate the size of the set corresponding to each sub-region, and determine whether the size of each set reaches the set threshold;
[0026] When the size of the set reaches the set threshold, complete the data collection for this set;
[0027] When the size of the set does not reach the set threshold, repeat the steps of object detection of video frames, calculation of the vertical coordinate of the center point of the small vehicle object detection frame, and data collection of the set until the data collection for all sets is completed.
[0028] Further, according to the vertical coordinate of the center point of the object detection frame to be recognized, the height of the image or video frame of the scene to be recognized, the number of sub-regions, and the multiple expression, obtain the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the object detection frame to be recognized. The specific implementation process is as follows:
[0029] Calculate the serial number of the set to which the object detection frame to be recognized belongs according to the vertical coordinate of the center point of the object detection frame to be recognized, the height of the image or video frame of the scene to be recognized, and the number of sub-regions. The specific calculation formula is:
[0030] Index = floor(vehicle_center_y * N / H);
[0031] Among them, Index represents the serial number of the set, floor() represents the floor function, vehicle_center_y represents the vertical coordinate of the center point of the target detection frame to be recognized, N represents the number of sub-regions or the number of sets, and H represents the height of the video frame in the scene image to be recognized or the traffic flow video stream;
[0032] Calculate the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized according to the serial number of the set to which the target detection frame to be recognized belongs and the multiple expressions.
[0033] Further, determine the vehicle size category of the target detection frame to be recognized according to the height of the target detection frame to be recognized and the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized. The specific implementation process is as follows:
[0034] Calculate the ratio of the height of the target detection frame to be recognized to the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized;
[0035] Determine the vehicle size category of the target detection frame to be recognized according to the ratio.
[0036] Further, determine the vehicle size category of the target detection frame to be recognized according to the ratio. The specific implementation process is as follows:
[0037] When the ratio < 1, the target detection frame to be recognized is a small vehicle;
[0038] When 1 ≤ ratio ≤ 5 / 3, the target detection frame to be recognized is a medium vehicle;
[0039] When 5 / 3 < ratio, the target detection frame to be recognized is a large vehicle.
[0040] Further, after collecting the data of the sets corresponding to each sub-region and before calculating the average value of all the data in each set, the recognition method further includes eliminating the noise data in each set by using the 3Sigma principle.
[0041] Based on the same concept, the present invention also provides a thunder and vision device, including a memory, a processor, and a computer program / instructions stored on the memory. The processor executes the computer program / instructions to implement the vehicle size recognition method as described above.
[0042] Based on the same concept, the present invention also provides a computer-readable storage medium, on which a computer program / instructions is stored. When the computer program / instructions is executed by a processor, the vehicle size recognition method as described above is implemented.
[0043] Based on the same inventive concept, the present invention further provides a computer program product, including computer programs / instructions, which when executed by a processor implement the vehicle size recognition method described above.
[0044] Advantageous effects
[0045] Compared with the prior art, the advantages of the present invention are as follows:
[0046] Based on data statistics and polynomial fitting, the present invention can more accurately estimate the average height value of the small vehicle detection frame at the ordinate pixel position. This estimation method has higher accuracy and better flexibility, can adaptively calculate for different usage scenarios, has less interference from human experience, and has better robustness in actual usage scenarios.
[0047] The present invention uses the ratio of the height of the detection frame to be recognized to the average height value of the small vehicle detection frame corresponding to the ordinate of the center point of the detection frame to be recognized as the discrimination criterion for large vehicles, medium vehicles, and small vehicles, that is, the height ratio relationship of the detection frame to be recognized approximates the actual vehicle length ratio relationship. This approximation method is more accurate, the discrimination conforms to the classification standard, and the classification accuracy in actual use is higher, meeting the requirements of the actual scenario.
[0048] The present invention uses the height information of the detection frame to be recognized as the classification feature, with lower calculation cost; compared with the fine classification method, the present invention does not require the collection and annotation of classification data, with lower annotation cost; compared with the cascade classification method, there is only one detection model, and the deployment cost of the present invention is lower and the recognition speed is faster in a multi-target scenario. Description of the drawings
[0049] In order to more clearly illustrate the technical solution of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0050] Figure 1 is the flowchart of the vehicle size recognition method in the embodiment of the present invention;
[0051] Figure 2 is the traffic flow distribution diagram of a certain video frame in the traffic flow video stream in the embodiment of the present invention;
[0052] Figure 3 is the height distribution diagram of the small vehicle target detection frames of a certain set in the embodiment of the present invention;
[0053] Figure 4It is the 6th-order polynomial fitting curve of the set serial numbers of 432 groups and the average height values of the small vehicle detection frames in the embodiments of the present invention;
[0054] Figure 5 It is the vehicle size classification result diagram in the embodiments of the present invention. Specific embodiments
[0055] Combined with the accompanying drawings in the embodiments of the present invention below, the technical solutions in the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0056] The technical solutions of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0057] Embodiment 1
[0058] As Figure 1 shown, a vehicle size recognition method provided by an embodiment of the present invention includes the following steps:
[0059] Step S1: Obtain any video frame in the to-be-recognized scene image or the vehicle flow video stream, equally divide a certain video frame in the to-be-recognized scene image or the vehicle flow video stream into N sub-regions along the vertical coordinate direction of the image, and construct sets corresponding to each sub-region, and sort each set in ascending order according to the vertical coordinate range of each sub-region.
[0060] The vehicle size recognition method of the present invention can be applied to a radar-vision device, mainly optimizing the video processing part in the radar-vision device. A vehicle refers to a motor vehicle with four or more wheels, excluding tricycles and two-wheeled motorcycles.
[0061] In this embodiment, the radar-vision device is installed forwardly on the road, and the radar-vision device is debugged so that the radar-vision device can collect the traffic flow images in the road, that is, the images of the scene to be recognized, to realize road traffic monitoring. The image of the scene to be recognized is equally divided into N sub-regions along the vertical coordinate direction of the image, and each sub-region is sorted from bottom to top (that is, sorted in the order of the ascending vertical coordinate range of each sub-region). Then, the vertical coordinate range of the i-th sub-region is (i−1)h to ih, where i = 1, 2, …, N, h represents the height of the sub-region, and h ≤ 7. That is, the vertical coordinate range of the first sub-region (the bottommost sub-region) is 0 to h, the vertical coordinate range of the second sub-region is h to 2h, and so on. The vertical coordinate range of the N-th sub-region (the topmost sub-region) is (N−1)h to Nh, and H = Nh, where H represents the pixel height of the image of the scene to be recognized. Therefore, the number of sets is N, and the serial numbers of the sets corresponding to each sub-region are 1 to N in sequence. That is, the set corresponding to the first sub-region is the first set (its serial number is 1), the set corresponding to the second sub-region is the second set (its serial number is 2), and so on. The set corresponding to the N-th sub-region is the N-th set (its serial number is N).
[0062] Figure 2 Fig. shows the traffic flow distribution of a certain video frame in the traffic flow video stream. The video frame is equally divided into 432 sub-regions, that is, N = 432. Each sub-region corresponds to a set. The height h of each sub-region is 5. Each set is used to record or store the height of the small vehicle target detection frame whose central point vertical coordinate is within the vertical coordinate range. The pixel height of the video frame is 2160.
[0063] In another embodiment of the present invention, a grid division method is used to divide the image of the scene to be recognized or a certain video frame in the traffic flow video stream. Whether it is equally divided along the vertical coordinate direction of the image (i.e., horizontally sliced) or based on the grid division method, the purpose is to obtain the target detection frame information at the pixel position of the central point vertical coordinate.
[0064] Step S2: Obtain the traffic flow video stream, perform target detection on the video frames in the traffic flow video stream, and obtain vehicle target detection information.
[0065] In the embodiment of the present invention, the target detection model is used to perform target detection on the video frames in the traffic flow video stream. The specific implementation process is as follows:
[0066] Step S2.1: Obtain vehicle environment images, label the vehicles with 5 seats or less in the vehicle environment images as small vehicles, and label the vehicles with more than 5 seats as other categories to construct a small vehicle detection data set;
[0067] Step S2.2: Construct a deep learning target detection model, and use the small vehicle detection data set to train and verify the deep learning target detection model to obtain the target detection model;
[0068] Step S2.3: Use the object detection model to perform object detection on the video frames in the traffic flow video stream to obtain vehicle object detection information.
[0069] Traffic management includes traffic flow statistics, queue detection, and congestion detection, all of which require the use of vehicle size information. Taking traffic flow statistics as an example, only the flow of large, medium, and small vehicles needs to be counted. Existing recognition methods identify categories such as trucks, buses, and freight vehicles, which are not directly used in traffic flow statistics. Instead, they need to be classified into large, medium, and small according to the actual category of the vehicle.
[0070] The object detection model of the present invention only identifies small vehicles, providing a reference standard for subsequent classification of small, medium, and large vehicles. In this embodiment, vehicle environment images are obtained from the publicly available dataset COCO, and vehicles with 5 seats or less are labeled as small vehicles. The deep learning object detection model can be the YOLOv5 model. The object detection model can be obtained by fine-tuning on the basis of the existing model using the small vehicle detection dataset. In another embodiment of the present invention, an existing object detection model capable of identifying small vehicles can also be used to perform object detection on the video frames in the traffic flow video stream.
[0071] In this embodiment, the vehicle object detection information includes the position of the object detection frame (such as the upper left corner coordinates), height, width, and category (i.e., small vehicle or other category). Only the object detection frames with the category of small vehicle are selected to construct a polynomial expression. If a video frame contains M vehicle objects, M object detection frames are obtained by performing object detection on this video frame. For each object detection frame with the category of small vehicle, the calculation of the vertical coordinate of the center point and the comparison of the vertical coordinate of the center point with the vertical coordinate ranges of each sub-region are performed to collect the data of the set in step S4.
[0072] Step S3: When the category of the object detection frame is a small vehicle, calculate the vertical coordinate of the center point of the small vehicle object detection frame according to the upper left corner coordinates, height, and width of the small vehicle object detection frame.
[0073] Step S4: Collect the data of the set corresponding to each sub-region according to the vertical coordinate of the center point of the small vehicle object detection frame, the height, and the vertical coordinate ranges of each sub-region.
[0074] In the embodiment of the present invention, collecting the data of the set corresponding to each sub-region according to the vertical coordinate of the center point of the small vehicle object detection frame, the height, and the vertical coordinate ranges of each sub-region, the specific implementation process is as follows:
[0075] Step S4.1: Compare the ordinate of the center point of the small vehicle target detection frame with the ordinate range of each sub-region. When the ordinate of the center point of the small vehicle target detection frame is within the ordinate range of the sub-region, add the height of the small vehicle target detection frame to the set corresponding to the sub-region;
[0076] Step S4.2: Calculate the size of the set corresponding to each sub-region, and determine whether the size of each set reaches a set threshold;
[0077] When the size of a set reaches the set threshold, the data collection of the set is completed, and the height of the small vehicle target detection box is no longer added to the set;
[0078] When the size of the set does not reach the set threshold, repeat steps S3 to S4 for the next small vehicle target detection frame until data collection for all sets is completed. If all video frames in the traffic video stream are subjected to target detection, center point ordinate calculation, center point ordinate comparison with the ordinate range of each sub-region, and height addition to the corresponding set, and there are still sets that have not completed data collection, then continue to obtain the traffic video stream of the next time period until data collection for all sets is completed.
[0079] For example, if the ordinate of the center point of a small vehicle target detection frame is in the range of h to 2h (i.e., the ordinate range of the second sub-region), the height of the small vehicle target detection frame is added to the set corresponding to the sub-region (i.e., the second set). The data in each set is the height of the small vehicle target detection frame, and the size of each set refers to the number of heights of the small vehicle target detection frame. For example, if the second set records or stores 2 heights (of the small vehicle target detection frame), the size of the second set is 2.
[0080] Step S5: Use the 3Sigma principle to remove noise data in each set.
[0081] After completing the data collection of each set, each set contains M heights of small vehicle target detection frames, where M is the set threshold. In order to avoid errors caused by inaccurate small vehicle target detection frames, the 3Sigma principle is used to remove noise data (i.e., singular points) in each set, that is, to remove the heights of some small vehicle target detection frames with large errors. Figure 3 The figure shows the height distribution of the small vehicle target detection frame contained in a certain set. The horizontal axis represents the vertical axis range of the sub-area corresponding to the set, the vertical axis represents the pixel height of the target detection frame, the red points represent the height points that meet the requirements, and the green points represent the noise data (or singular points) that are removed.
[0082] Step S6: Calculate the average value of all data within each set, and use this average value as the average height value of the small vehicle detection frame corresponding to the set.
[0083] Exemplarily, the second set contains the heights of m small vehicle target detection frames, where m ≤ M. Then, calculate the average value of the heights of the m small vehicle target detection frames, and use this average value as the average height value of the small vehicle detection frame of the second set, that is, use this average value as the average height value of the small vehicle detection frame within the ordinate range h to 2h.
[0084] Step S7: Perform polynomial fitting on the N sets of set numbers and their corresponding average height values of small vehicle detection frames to obtain a polynomial expression of the set number and the average height value of the small vehicle detection frame.
[0085] Each set corresponds to a number and an average height value of small vehicle detection. For example, the number of the second set is 2, and the average height value of the small vehicle detection frame of the second set is calculated according to Step S6. Since there are N sub-regions and N sets, there are N groups of set numbers and their corresponding average height values of small vehicle detection frames. In the polynomial expression of the set number and the average height value of the small vehicle detection frame obtained through polynomial fitting, the set number is the independent variable, and the dependent variable is the average height value of the small vehicle detection frame corresponding to this number.
[0086] Figure 4 Shows a 6th-order polynomial fitting curve of N = 432 points. In actual use, polynomial fitting of 2nd order or higher can be adopted. The abscissa represents the set number, and the ordinate represents the average height value of the small vehicle detection frame corresponding to the number. The red curve represents the true statistical value over a period of time, and the green curve represents the fitting curve. Figure 4 It can be seen that the polynomial fitting value of the present invention is basically consistent with the true statistical value, with high fitting accuracy, ensuring the classification accuracy.
[0087] Step S8: Obtain the target detection frame to be recognized, and calculate the ordinate of the center point of the target detection frame to be recognized.
[0088] When it is necessary to identify the vehicle size of a video frame in a vehicle flow video stream, use the target detection model to perform target detection on this video frame to obtain the target detection frame to be recognized and its position, height, and width. Calculate the ordinate of the center point of the target detection frame to be recognized according to the position, height, and width of the target detection frame to be recognized.
[0089] Step S9: Obtain the average height value of the small vehicle detection frame corresponding to the ordinate of the center point of the target detection frame to be recognized according to the ordinate of the center point of the target detection frame to be recognized, the height of the video frame, the number of sub-regions, and the polynomial expression.
[0090] In an embodiment of the present invention, according to the vertical coordinate of the center point of the target detection frame to be recognized, the height of the video frame, the number of sub-regions, and the multiple expressions, the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized is obtained. The specific implementation process is as follows:
[0091] Step S9.1: Calculate the serial number of the set to which the target detection frame to be recognized belongs according to the vertical coordinate of the center point of the target detection frame to be recognized, the height of the video frame, and the number of sub-regions. The specific calculation formula is:
[0092] Index = floor(vehicle_center_y * N / H) (1)
[0093] where Index represents the serial number of the set, floor() represents the floor function, vehicle_center_y represents the vertical coordinate of the center point of the target detection frame to be recognized, N represents the number of sub-regions or the number of sets, and H represents the height of the video frame in the image of the scene to be recognized or the traffic flow video stream;
[0094] Step S9.2: Calculate the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized according to the serial number of the set to which the target detection frame to be recognized belongs and the multiple expressions.
[0095] Substitute the serial number of the set to which the target detection frame to be recognized belongs into the multiple expressions obtained in step S7, and the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized can be obtained.
[0096] Step S10: Determine the vehicle size category of the target detection frame to be recognized according to the height of the target detection frame to be recognized and the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized.
[0097] Calculate the ratio of the height of the target detection frame to be recognized to the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized; determine the vehicle size category of the target detection frame to be recognized according to the ratio.
[0098] Refer to the standard of GB / T 918.1 - 1989 "Classification and Coding of Road Vehicles - Motor Vehicles" (referred to as the classification standard), and some classification rules are shown in Table 1.
[0099] Table 1 Some Classification Rules of "Classification and Coding of Road Vehicles - Motor Vehicles"
[0100] Code Vehicle category name Reference length range (m) 1 Extra-large vehicle >16m 2 Large vehicle 10m - 16m 3 Medium vehicle 6m - 10m 4 Small vehicle 3.5m - 6m 5 Miniature vehicle ≤3.5m
[0101] According to the classification standard and combined with the actual road requirements, vehicle sizes are divided into three categories, namely large vehicles, medium-sized vehicles, and small vehicles. Ultra-large vehicles and large vehicles are classified as large vehicles, medium-sized vehicles are classified as medium-sized vehicles, and small vehicles and micro vehicles are classified as small vehicles. The classification standard gives the actual vehicle length ranges. Based on the length range of small vehicles, the ratios of the length ranges of large vehicles to small vehicles and medium-sized vehicles to small vehicles are calculated. The vertical coordinate of the center point of the target detection box of the vehicle is used to estimate the position of the vehicle. For vehicles at the same position, the ratio of the vehicle length range to the length range of small vehicles can reflect the size relationship of the vehicles. Therefore, the ratio of the vehicle length range to the length range of small vehicles is used as the classification threshold, and combined with the classification relationship in step S10, the size classification of the vehicles is completed.
[0102] In Table 1, the length range of small vehicles is less than 6 meters, the length range of medium-sized vehicles is 6 meters to 10 meters, and the length range of large vehicles is more than 10 meters. Then the ratio range of the length range of medium-sized vehicles to small vehicles is 1 to 1.66, the ratio of the length range of large vehicles to small vehicles is greater than 1.66, and the ratio of the length range of small vehicles to small vehicles is less than 1. Therefore, when the ratio < 1, the target detection box to be recognized is a small vehicle; when 1 ≤ ratio ≤ 5 / 3, the target detection box to be recognized is a medium-sized vehicle; when 5 / 3 < ratio, the target detection box to be recognized is a large vehicle.
[0103] When the height h of each sub-region is equal to 1 pixel, the average height value of the small vehicle detection box corresponding to each vertical coordinate in the video frame can be obtained. Then, the ratio of the height of the target detection box of the vehicle to be recognized (i.e., the target detection box to be recognized) to the average height value of the small vehicle detection box corresponding to the vertical coordinate of the center point of this target detection box is calculated, and this ratio is used as the classification relationship.
[0104] Exemplarily, if the ratio is equal to 2, the vehicle category corresponding to the target detection box to be recognized is a large vehicle. Figure 5 The vehicle size classification result diagram is shown. Figure 5 In it, small represents small vehicles, medium represents medium-sized vehicles, and big represents large vehicles. From Figure 5 it can be seen that the present invention can accurately identify the sizes of various vehicles.
[0105] Using the present invention for vehicle size classification, there is less human experience intervention, higher classification accuracy, it can be adaptively counted according to different usage scenarios, has better generalization, and there is only a process of comparing one detection box height threshold, so the recognition speed is faster.
[0106] Embodiment 2
[0107] The embodiment of the present invention also provides a radar-vision device, which includes: a memory, a processor, and a computer program / instructions stored on the memory. The processor executes the computer program / instructions to implement the vehicle size recognition method described in Embodiment 1.
[0108] Although not shown, the radar-vision device includes a processor that can perform various appropriate operations and processes according to programs and / or data stored in a read-only memory (ROM) and / or programs and / or data loaded from a storage section into a random access memory (RAM). The processor can be a multi-core processor or can include multiple processors. In some embodiments, the processor can include a general-purpose main processor and one or more special co-processors, such as a central processing unit, a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processing unit (DSP), and so on. In the RAM, various programs and data required for the operation of the radar-vision device are also stored. The processor, the ROM, and the RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0109] The above-mentioned processor and memory are jointly used to execute the programs / instructions stored in the memory, and when the programs / instructions are executed by a computer, they can implement the methods, steps, or functions described in the above embodiments.
[0110] Although not shown, an embodiment of the present invention also provides a computer-readable storage medium on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the vehicle size recognition method described in Embodiment 1 is implemented.
[0111] In the embodiments of the present invention, the storage medium includes permanent and non-permanent, removable and non-removable articles that can implement information storage by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only
[0112] memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0113] A readable storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0114] Although not shown, embodiments of the present invention also provide a computer program product, including: computer programs / instructions, which, when executed by a processor, implement the vehicle size recognition method as described in Embodiment 1.
[0115] The above-disclosed are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or variations, which should all be covered within the protection scope of the present invention.
Claims
1. A vehicle size recognition method, characterized in that, The recognition method includes the following steps: Obtain a traffic flow video stream, perform object detection on video frames in the traffic flow video stream to obtain vehicle object detection information; wherein, the vehicle object detection information includes the position, height, width, and category of the object detection box; When the category of the object detection box is a small vehicle, calculate the center point ordinate of the small vehicle object detection box according to the position, height, and width of the small vehicle object detection box; Collect data of sets corresponding to each sub-region according to the center point ordinate and height of the small vehicle object detection box and the ordinate ranges of each sub-region; wherein, the sub-regions are obtained by equally dividing a certain video frame in the image to be recognized or the traffic flow video stream along the image ordinate direction into N parts, and the sets corresponding to each sub-region are sorted in ascending order according to the ordinate ranges of each sub-region; Calculate the average value of all data in each set, and use the average value as the average height value of the small vehicle detection box corresponding to the set; Perform polynomial fitting on N groups of set numbers and their average height values of the small vehicle detection box to obtain a polynomial expression of the set number and the average height value of the small vehicle detection box; Obtain the object detection box to be recognized, and calculate the center point ordinate of the object detection box to be recognized; Obtain the average height value of the small vehicle detection box corresponding to the center point ordinate of the object detection box to be recognized according to the center point ordinate of the object detection box to be recognized, the height of the image or video frame to be recognized, the number of sub-regions, and the polynomial expression; Determine the vehicle size category of the object detection box to be recognized according to the height of the object detection box to be recognized and the average height value of the small vehicle detection box corresponding to the center point ordinate of the object detection box to be recognized.
2. The vehicle size recognition method according to claim 1, wherein The specific implementation process of performing object detection on video frames in the traffic flow video stream is as follows: Obtain a vehicle environment image, label vehicles with 5 seats or less in the vehicle environment image as small vehicles, and label vehicles with more than 5 seats as other categories to construct a small vehicle detection data set; Construct a deep learning object detection model, and use the small vehicle detection data set to train and verify the deep learning object detection model to obtain an object detection model; Use the object detection model to perform object detection on video frames in the traffic flow video stream to obtain vehicle object detection information.
3. The vehicle size recognition method according to claim 1, wherein The specific implementation process of collecting data of sets corresponding to each sub-region according to the center point ordinate and height of the small vehicle object detection box and the ordinate ranges of each sub-region is as follows: Compare the center point ordinate of the small vehicle object detection box with the ordinate ranges of each sub-region. When the center point ordinate of the small vehicle object detection box is within the ordinate range of the sub-region, add the height of the small vehicle object detection box to the set corresponding to the sub-region; Calculate the size of the sets corresponding to each sub-region, and determine whether the size of each set reaches a set threshold; When the size of the set reaches the set threshold, complete the data collection of the set; When the size of the set does not reach the set threshold, repeat the steps of target detection of video frames, calculation of the vertical coordinate of the center point of the small vehicle target detection frame, and data collection of the set until the data collection of all sets is completed.
4. The vehicle size recognition method according to claim 1, characterized in that, According to the vertical coordinate of the center point of the target detection frame to be recognized, the height of the image or video frame of the scene to be recognized, the number of sub-regions, and the multiple expressions, obtain the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized. The specific implementation process is as follows: Calculate the serial number of the set to which the target detection frame to be recognized belongs according to the vertical coordinate of the center point of the target detection frame to be recognized, the height of the image or video frame of the scene to be recognized, and the number of sub-regions. The specific calculation formula is: Index = floor(vehicle_center_y * N / H); Where Index represents the serial number of the set, floor() represents the floor function, vehicle_center_y represents the vertical coordinate of the center point of the target detection frame to be recognized, N represents the number of sub-regions or the number of sets, and H represents the height of the image to be recognized or the video frame in the vehicle flow video stream; Calculate the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized according to the serial number of the set to which the target detection frame to be recognized belongs and the multiple expressions.
5. The vehicle size recognition method according to claim 1, characterized in that Determine the vehicle size category of the target detection frame to be recognized according to the height of the target detection frame to be recognized and the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized. The specific implementation process is as follows: Calculate the ratio of the height of the target detection frame to be recognized to the average height value of the small vehicle detection frame corresponding to the vertical coordinate of the center point of the target detection frame to be recognized; Determine the vehicle size category of the target detection frame to be recognized according to the ratio.
6. The vehicle size recognition method according to claim 5, wherein Determine the vehicle size category of the target detection frame to be recognized according to the ratio. The specific implementation process is as follows: When the ratio < 1, the target detection frame to be recognized is a small vehicle; When 1 ≤ ratio ≤ 5 / 3, the target detection frame to be recognized is a medium-sized vehicle; When 5 / 3 < ratio, the target detection frame to be recognized is a large vehicle.
7. The vehicle size recognition method according to any one of claims 1 to 6, characterized in that After collecting the data of the sets corresponding to each sub-region and before calculating the average value of all the data in each set, the recognition method further includes eliminating the noise data in each set by using the 3Sigma principle.
8. A radar-vision device, comprising a memory, a processor, and computer programs / instructions stored on the memory, characterized in that The processor executes the computer program / instructions to implement the vehicle size recognition method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the vehicle size recognition method according to any one of claims 1 to 7 is implemented.
10. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the vehicle size recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Vehicle type identification and traffic flow detection method
CN107122758A
Traffic flow statistics and vehicle type classification method and device
CN110751119A
Exception event detection method for expressway
CN114170580A
Distance measurement method and device based on target and terminal equipment
CN115035188A