Vehicle speed detection method, electronic device, readable storage medium and program product

By acquiring image data at reference locations for vehicle recognition and trajectory tracking, and combining it with path planning data from edge servers, the problem of insufficient accuracy in vehicle speed detection under complex environments is solved, achieving efficient and accurate vehicle speed detection.

CN120668955BActive Publication Date: 2025-12-16LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511188324.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-12-16
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing vehicle speed detection methods lack accuracy in complex environments and cannot meet high-precision requirements, especially in scenarios with complex road networks and frequent changes in vehicle speed within closed areas, where vehicle tracking fails and speed measurement is inaccurate.

Method used

Vehicle identification is performed by acquiring image data at reference locations, matching images of vehicles are determined, and vehicle tracking is performed by combining edge servers with vehicle driving trajectory and speed measurement area path planning data, thereby reducing computing resource requirements and improving vehicle tracking efficiency and speed detection accuracy.

Benefits of technology

It effectively improves the accuracy of vehicle speed detection, reduces computing resource consumption, is suitable for edge device deployment, adapts to complex road conditions, and reduces speed calculation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668955B_ABST
    Figure CN120668955B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle speed detection method, electronic equipment, readable storage medium and program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: collecting a vehicle matching image used for subsequent vehicle tracking according to a vehicle recognition result of an image collected from a target reference position of a speed measurement area. Vehicle tracking data matched with a vehicle to be measured is determined from a video stream image frame of a current field of view. After the vehicle drives out of the current field of view, a field of view for extracting a next image data is determined, and vehicle tracking data of a new field of view is determined again according to the latest vehicle matching image. The average driving speed of the vehicle in the entire speed measurement area and the field of view driving speed under each field of view are determined according to all the vehicle tracking data, and the driving speed of the vehicle to be measured is determined according to the speed data. The application can solve the problems of vehicle tracking failure and inaccurate speed measurement, and effectively improve the accuracy of vehicle speed detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a vehicle speed detection method, an electronic device, a computer readable storage medium and a computer program product. BACKGROUND

[0002] The vehicle speed detection accuracy increases with the increasing safety requirements of the vehicle driving scene. In the related art, there are problems of vehicle tracking failure and inaccurate speed measurement in the vehicle speed measurement process, which cannot meet the high-precision vehicle speed detection requirements of users. SUMMARY

[0003] The present application provides a vehicle speed detection method, an electronic device, a computer readable storage medium and a computer program product, which effectively improve the vehicle speed detection accuracy.

[0004] To solve the above technical problems, the present application provides the following technical solutions:

[0005] The present application provides a vehicle speed detection method, comprising:

[0006] According to the vehicle recognition result of the target reference area image corresponding to the target reference position of the speed measurement area, a vehicle matching image containing the vehicle to be measured is determined. Based on the vehicle matching image, vehicle tracking data matched with the vehicle to be measured is determined from the video stream image frame of the current field of view. When it is determined that the vehicle drives away from the current field of view according to the vehicle tracking data, the target field of view corresponding to the current driving position of the vehicle is determined according to the speed measurement area path planning data, and new vehicle tracking data is determined again in the target field of view according to the current vehicle matching image. According to the vehicle tracking data of each field of view passed through by the vehicle to be measured during driving, the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of each field of view are determined, and the driving speed of the vehicle to be measured is determined according to the average driving speed and the driving speed of each field of view.

[0007] The present application also provides an electronic device comprising a memory and a processor, wherein the processor is configured to execute the steps of any of the above vehicle speed detection methods when executing the computer program stored in the memory.

[0008] The present application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of any of the above vehicle speed detection methods.

[0009] Finally, the present application also provides a computer program product comprising a computer program / instruction, wherein the computer program / instruction is executed by a processor to implement the steps of any of the above vehicle speed detection methods.

[0010] The vehicle speed detection method provided by the present application has the advantages that the vehicle recognition result of the image data collected by the reference position is used as the vehicle matching image for tracking the vehicle to be tested, and the image collection devices passed by the vehicle to be tested in the speed detection area are determined according to the vehicle driving track and the path planning data of the speed detection area, so that vehicle detection is not required for all image collection devices in the speed detection area, the vehicle tracking efficiency is improved, the speed detection efficiency is improved, the calculation resources are effectively saved, the method is suitable for deployment on an edge device, the final driving speed of the vehicle is determined according to the driving speed of the vehicle in each single camera area and the entire speed detection area, the speed calculation error is effectively reduced, and the vehicle speed detection accuracy is improved.

[0011] In addition, the present application also provides a corresponding electronic device, a computer readable storage medium and a computer program product for the vehicle speed detection method, which further makes the method more practical, and the electronic device, the computer readable storage medium and the computer program product have corresponding advantages. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the present application or related technologies, the following will briefly introduce the drawings needed to be used in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0013] Figure 1 The hardware composition framework schematic diagram of the vehicle speed detection method provided by the present application;

[0014] Figure 2 The flowchart schematic diagram of the vehicle speed detection method provided by the present application;

[0015] Figure 3 The network model structure schematic diagram of the vehicle detection network model provided by the present application in an exemplary application scenario;

[0016] Figure 4 The network model structure schematic diagram of the initial feature extraction layer provided by the present application in an exemplary application scenario;

[0017] Figure 5 The network model structure schematic diagram of the multi-scale feature extraction layer provided by the present application in an exemplary application scenario;

[0018] Figure 6 The network model structure schematic diagram of the feature extraction block provided by the present application in an exemplary application scenario;

[0019] Figure 7A first output layer provided by the present application is a network model structure schematic diagram in an exemplary application scenario;

[0020] Figure 8 A vehicle matching network model provided by the present application is a network model structure schematic diagram in an exemplary application scenario;

[0021] Figure 9 A detection head provided by the present application is a network model structure schematic diagram in an exemplary application scenario;

[0022] Figure 10 A flowchart of another vehicle speed detection method provided by the present application;

[0023] Figure 11 A structure framework diagram of an exemplary embodiment of a vehicle speed detection device provided by the present application;

[0024] Figure 12 A structure schematic diagram of an exemplary embodiment of an electronic device provided by the present application. DETAILED DESCRIPTION

[0025] In order to make the person skilled in the art better understand the technical solutions of the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments. In the specification and the above-mentioned drawings, the terms "first", "second", "third", "fourth" and the like are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "as an example, embodiment or illustration". Any embodiment described as "exemplary" herein is not necessarily interpreted as superior or better than other embodiments.

[0026] The vehicle speed detection accuracy increases with the increase of the safety demand of the vehicle driving scene. For a closed area with an entrance and an exit, such as a logistics park, a campus, and a scenic area, in order to avoid traffic accidents caused by overspeed, it is necessary to strictly control the maximum speed of the vehicle driving speed. The current vehicle speed detection method includes non-vision speed detection method (such as radar speed measurement technology, ground inductor detection technology), camera-based speed measurement technology (such as traditional video analysis method, computer vision technology), and a combination of the two, such as visual + radar combination to realize speed measurement. Such method has high cost and poor universality.

[0027] Among them, the radar speed measurement technology will first emit electromagnetic waves, then receive the vehicle reflection signal, and finally calculate the vehicle speed based on the Doppler effect. This method will be linked with the camera (such as radar and vision integrated machine), realize the speed capture, the speed measurement accuracy is ±0.5km / h, and supports 100 meters range multi-lane coverage. However, it is easy to misjudge due to the interference of metal objects, and the accuracy decreases in rainy and foggy weather. In addition, this method only outputs speed data and cannot identify vehicle features (such as license plate and vehicle model). The inductive coil detection technology will bury the inductive coil in the road surface, and calculate the speed v according to the electromagnetic field change when the vehicle passes through. This method has low cost and high reliability, but it needs to damage the road surface for construction, and the road surface needs to be changed again, and the coil needs to be buried again, which has high maintenance cost and poor scalability, and cannot identify vehicle feature information. When multiple vehicles pass through, the speed measurement is prone to error. In addition, infrared speed measurement uses vehicle infrared radiation monitoring, but the effective distance is short (less than 50 meters), and it is greatly affected by weather. The ultrasonic speed measurement has strong anti-rain and anti-fog ability, but the beam divergence angle is large, the resolution is low, and the error is significant.

[0028] Among them, the traditional video analysis method refers to setting a virtual detection line in the video frame, triggering vehicle position analysis through gray scale change, and calculating the speed combined with frame interval time. This method analyzes the two-dimensional displacement of the vehicle in the continuous video frame, and converts the actual speed combined with the calibration parameters. Due to the change of light (such as night, shadow), the image noise increases, which leads to high false detection rate, in addition, manual calibration of reference objects is required, the algorithm robustness is poor, and the data processing depends on high-performance computing devices, which has poor real-time performance. Computer vision technology realizes vehicle dynamic tracking by introducing deep learning algorithm and neural network model, can realize license plate recognition and trajectory prediction, reduces the dependence on manual calibration, can improve the adaptability of complex scenes, and can output structured data such as vehicle model and color. However, this method has poor tracking effect, and when the vehicle overtakes or overlaps, it is easy to cause tracking failure and speed measurement error. For speed measurement methods that require camera participation, since it needs to rely on camera image acquisition, there are problems of poor environmental adaptability: narrow road, curved road and other complex road conditions lead to poor vehicle tracking effect, frequent target occlusion; day and night light difference, tree shadow and other factors reduce image quality, affect the overall speed measurement accuracy.

[0029] From the above, the related technology can realize the speed measurement function, but there are bottlenecks in the aspects of environmental robustness, deployment flexibility and data intelligentization, especially for the scene of complex road network and frequent vehicle speed change in a closed area. In view of this, the present application provides a low dependence, high compatibility and strong adaptation pure visual speed measurement method, which uses the vehicle recognition result of the image data collected by the reference position as the vehicle matching image for tracking the vehicle to be measured in the subsequent process, determines the image collection equipment passed by the vehicle to be measured in the driving process in the speed measurement area according to the vehicle driving track and the speed measurement area path planning data, and determines the final driving speed of the vehicle according to the driving speed of the vehicle in each single camera area and the whole speed measurement area, thereby efficiently and accurately detecting the vehicle speed in the speed measurement area on the basis of reducing the calculation resources used in the speed measurement process.

[0030] In combination with the specific application environment architecture or specific hardware architecture on which the vehicle speed detection method is dependent, the specific application environment architecture or specific hardware architecture is described herein. The following describes the vehicle speed detection method in combination with Figure 1 Some possible application scenarios related to the technical solutions of the present application are exemplarily introduced, which can include the following contents:

[0031] An entrance camera is deployed at the entrance of the speed measurement area, a plurality of cameras are deployed around the vehicle driving road according to the path planning of the speed measurement area, and each camera will transmit the image data collected in the field of view to the edge server in real time.

[0032] The edge server pre-deploys any kind of target detection network model capable of recognizing vehicles, when receiving the image data collected by the entrance camera at the entrance of the speed measurement area, transmits the image data to the target detection network model, identifies whether there is a vehicle, if there is, outputs a vehicle area image containing a vehicle tile, and uses it as a vehicle matching image for tracking the vehicle. Based on the vehicle matching image, the vehicle tracking data matched with the vehicle to be measured is determined from the video stream image frame of the current field of view; when it is determined that the vehicle drives away from the current field of view according to the vehicle tracking data, the target field of view corresponding to the current driving position of the vehicle is determined according to the speed measurement area path planning data, and new vehicle tracking data is determined again in the target field of view according to the current vehicle matching image; the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of each field of view are determined according to the vehicle tracking data of each field of view passed by the vehicle to be measured in the driving process, and the driving speed of the vehicle to be measured is determined according to the average driving speed and the driving speed of each field of view.

[0033] It should be noted that the above application scenarios are only shown for the purpose of facilitating the understanding of the ideas and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application can be applied to any applicable scenario. After introducing the technical solutions of the present application, various non-limiting embodiments of the present application will be described in detail below in combination with the drawings and specific embodiments. First, please refer to Figure 2 , Figure 2 A flowchart of a vehicle speed detection method provided for the present embodiment, which can include the following contents:

[0034] S201: Determine a vehicle matching image containing a vehicle to be speeded up according to a vehicle recognition result of a target reference area image corresponding to a target reference position of a speed measurement area.

[0035] The speed measurement area is an area in which the driving speed of a vehicle needs to be detected, and the target reference position is a position designated in advance in the speed measurement area, such as the position of the first camera through which a vehicle enters or passes in the speed measurement area. The target reference area image is the image data collected by the image collection device, such as a camera, installed at the target reference position. The vehicle recognition result is the result of vehicle recognition of the target reference area image, which can use any method capable of realizing target recognition in related technologies, such as using a target recognition algorithm or a target detection network model, which does not affect the implementation of the present application. Whether the currently collected target reference area image contains a vehicle, if not, continue to perform vehicle recognition on the next target reference area image, if yes, extract the vehicle information in the target reference area image, which includes an identifier uniquely identifying the vehicle, such as identifying the vehicle according to the license plate, of course, a unique ID can also be automatically generated for the vehicle, which is used for subsequent tracking of the vehicle, i.e. the vehicle needs to be speeded up. This step defines the vehicle as a vehicle to be speeded up. In order to facilitate subsequent image processing, the part of the image containing the vehicle tile can be extracted from the target reference area image as a vehicle matching image, which is used as a matching template for tracking the vehicle to be speeded up.

[0036] S202: Based on the vehicle matching image, determine vehicle tracking data matched with the vehicle to be speeded up from the video stream image frames of the current field of view.

[0037] When the previous step determines that there is a vehicle to be tested for speed, and the vehicle matching image of the previous step is the image of the starting position of the vehicle to be tested for speed, the current field of view is the vehicle to be tested for speed driving into the speed testing area, and the video stream image frame is the real-time video stream collected by the first image collection device in the driving process from the position of S201, vehicle detection is performed on each frame image of the video stream image frame to identify each image frame containing the vehicle to be tested for speed from the real-time video stream, the earliest image frame is the time when the vehicle to be tested for speed drives into the current field of view, and the position of the vehicle to be tested for speed in the current field of view and the driving time can be determined from each image frame to serve as vehicle tracking data of the vehicle to be tested for speed, that is, the vehicle tracking data at least includes the time when the vehicle to be tested for speed drives into the current field of view, the time when the vehicle to be tested for speed drives out of the current field of view, the total driving time in the current field of view, and the position information of the vehicle to be tested for speed in the current field of view, and the driving track of the vehicle to be tested for speed in the current field of view can be determined according to the position points. If the vehicle matching image of S201 is multiple, that is, multiple vehicles to be tested for speed need to be tracked, the corresponding current field of view of each vehicle to be tested for speed corresponding to each vehicle matching image needs to be determined, and then tracking is performed. If there are at least two vehicles to be tested for speed in the same field of view, different vehicles to be tested for speed need to be identified in the video stream image frame according to the vehicle matching image, after the type of the current vehicle is determined, corresponding tracking is performed, and vehicle tracking data is generated.

[0038] S203: When it is determined according to the vehicle tracking data that the vehicle drives out of the current field of view, the target field of view corresponding to the current driving position of the vehicle is determined according to the speed testing area path planning data, and new vehicle tracking data is determined again in the target field of view according to the current vehicle matching image.

[0039] It can be understood that the vehicle is driving forward, and the field of view of the image acquisition device is fixed, and the image acquisition device can only acquire image data within the field of view. In order to realize the speed detection of the vehicle to be measured, it is necessary to track the driving track of the vehicle to be measured in the speed measurement area in real time. When it is determined that the vehicle to be measured has driven out of the current field of view according to the vehicle tracking data of the current field of view in the last step, that is, when the vehicle to be measured cannot be detected in the real-time video stream data of the current field of view, it is considered that the vehicle to be measured has driven out of the current field of view, and the next field of view to be entered by the vehicle to be measured needs to be predicted. The target field of view is defined as the predicted next field of view in this step, so that the vehicle to be measured can be tracked through the real-time video stream of the next field of view. In this step, according to the path planning data of the speed measurement area, the path that the vehicle to be measured can drive can be known. According to the layout and field of view range of the image acquisition device of the speed measurement area, combined with the driving track of the vehicle determined by the vehicle tracking data of the current field of view, the next field of view of the vehicle can be predicted, and the image data collected by the image acquisition device corresponding to the next field of view is obtained. Through vehicle recognition on the real-time video stream data of the target field of view, each image frame containing the vehicle to be measured in S202 is determined. Similarly, in the image containing the vehicle to be measured, the image frame with the earliest time is the time when the vehicle to be measured enters the target field of view, and the last image is the last position of the vehicle to be measured in the target field of view. Similarly, the position point and driving time of the vehicle to be measured in the target field of view can be determined through each image frame, so as to generate the vehicle tracking data of the vehicle to be measured in the target field of view. When it is determined that the vehicle to be measured has driven away from the target field of view according to the vehicle tracking data of the target field of view, the next field of view of the target field of view is predicted according to the method of this step, and the process is repeated until the vehicle to be measured drives away from the speed measurement area, and the entire tracking process from the driving into the speed measurement area to the driving out of the speed measurement area is completed.

[0040] In this step, in the vehicle tracking process, in order to improve the vehicle matching accuracy, the vehicle matching image of S201 may be replaced by the latest and most accurate image. For example, if the license plate number of the vehicle matching image of S201 is not clear, it may be updated later. Therefore, the current vehicle matching image of this step may be the vehicle matching image of S201, or the latest updated vehicle matching image.

[0041] S204: According to the vehicle tracking data of each field of view passed by the vehicle to be measured during driving, the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of each field of view are determined, and the driving speed of the vehicle to be measured is determined according to the average driving speed and the driving speed of each field of view.

[0042] Since the vehicle tracking data of each field of view can determine the driving time and driving path of the vehicle to be tested in the corresponding field of view, the actual driving length corresponding to the driving path can be determined in combination with the path planning data of the speed testing area and the layout data of the image acquisition device, and then the average driving speed of the vehicle to be tested in each field of view, that is, the field of view driving speed of the present step, can be determined. Similarly, the driving path of the vehicle to be tested in the entire speed testing area can be determined according to the driving trajectory of each field of view passed by the vehicle to be tested during the driving process in the speed testing area, the actual total driving length corresponding to the area driving path can be determined in combination with the path planning data of the speed testing area and the layout data of the image acquisition device, the total driving time of the vehicle to be tested in the speed testing area can be determined according to the time when the vehicle to be tested enters the speed testing area, the time when the vehicle to be tested exits and the parking data, the average driving speed of the vehicle to be tested in the speed testing area can be determined according to the total driving time and the actual total driving length, and the parking data can be determined according to whether the vehicle tracking data exists at the same position at different times. Finally, the driving speed of the vehicle to be tested is determined according to the average driving speed and the driving speed of the single field of view in combination with the actual scene, for example, for the speed detection scene, as long as one of the average driving speed and the field of view driving speed exceeds the maximum speed limit value, it is considered that there is a speeding behavior.

[0043] In the technical scheme provided in the present embodiment, the vehicle recognition result of the image data collected at the reference position is taken as the vehicle matching image for subsequent tracking of the vehicle to be tested, and each image acquisition device passed by the vehicle to be tested during the driving process in the speed testing area is determined according to the vehicle driving trajectory and the path planning data of the speed testing area, so that vehicle detection is not required for all image acquisition devices in the speed testing area, which not only can improve the vehicle tracking efficiency and thus improve the speed detection efficiency, but also can effectively save the computing resources, is conducive to deployment on the edge device, and can determine the final driving speed of the vehicle according to the driving speed of the vehicle in each single camera area and the entire speed testing area, thereby effectively reducing the speed calculation error and improving the vehicle speed detection accuracy.

[0044] In the above embodiment, the driving speed of the vehicle to be tested is not limited, and an exemplary implementation manner is also given, which can include the following contents: for each field of view passed by the vehicle to be tested during the driving process, the appearance time and the exit time of the vehicle to be tested in the current field of view are determined according to the vehicle tracking data of the current field of view, of course, if the parking in the field of view is considered, the parking time of the vehicle in the field of view also needs to be detected, the driving time is obtained by subtracting the parking time from the time difference between the appearance time and the exit time of the previous field of view, and the driving speed of the vehicle to be tested in the current field of view is determined according to the path length of the current field of view according to the ratio of length to time. For example, the driving speed of the i-th field of view According to The calculation is li, which represents the road path length of the ith view, and ti represents the time difference of the vehicle to be measured from the ith view to leaving the ith view. When the vehicle to be measured is determined to leave the speed measurement area according to the vehicle tracking data of the last view passed by the vehicle to be measured during driving, the total driving time of the vehicle to be measured in the speed measurement area is determined according to the driving-in time of the first view and the driving-out time of the last view passed by the vehicle to be measured during driving. For the vehicle allowed to stop in the speed measurement area, the stopping time of the vehicle in the speed measurement area also needs to be detected, and the time difference between the driving-in time of the first view and the driving-out time of the last view is subtracted by the stopping time as the total driving time. The total driving length of the vehicle to be measured in the speed measurement area is determined according to the path planning data of the speed measurement area; and the average driving speed of the vehicle to be measured in the speed measurement area is determined according to the total driving length and the total driving time. If the average driving speed of the vehicle to be measured in the speed measurement area is greater than a preset speed threshold value, it is determined that the vehicle to be measured has a speeding behavior in the speed measurement area. If the average driving speed of the vehicle to be measured in the speed measurement area is less than the preset speed threshold value, it is determined that the vehicle to be measured has a low-speed driving behavior in the speed measurement area. According to The calculation is, which represents the length of the motion path of the vehicle to be measured in the speed measurement area, and t represents the total driving time of the vehicle to be measured in the speed measurement area.

[0045] After the average driving speed and the driving speed of each view are determined, the average speed in the view can be determined according to the driving speed of each view and the total number of views, and the speed difference between the average driving speed and the average speed in the view can be calculated, for example, according to where n represents the total number of views. When the speed difference between the average driving speed and the average speed in the view meets a preset speed threshold judgment condition, the preset speed threshold judgment condition is a condition set in advance, as a simple implementation, a threshold value When the speed difference is less than the threshold value, it is considered that the preset speed threshold judgment condition is met, that is, The average driving speed of the vehicle to be measured and the driving speed of each view meet the speed measurement error condition; when at least one speed value in the average driving speed and the driving speed of each view is greater than a preset speed limit value (maximum driving speed value), the vehicle to be measured has a speeding behavior in the speed measurement area. Of course, if at least one speed value is less than the preset speed limit value (minimum driving speed value), the vehicle to be measured has a low-speed driving behavior in the speed measurement area. In order to further improve the efficiency of speed detection, when the speeding driving or low-speed behavior is detected, an alarm information can be further generated, and the alarm information can include a vehicle picture and a motion trajectory and time of the vehicle in the speeding stage.

[0046] The above embodiments do not limit how to generate a vehicle matching image, and based on the above embodiments, the present application also provides an exemplary implementation, which can include the following contents:

[0047] The target reference position of the embodiment is the entrance position, and the target reference area image is entrance area image data. The entrance area image data of the entrance position of the speed measurement area is acquired. If the entrance area image data is from the same image acquisition source, or there is no vehicle in the overlapping field of view of different image acquisition sources in each entrance area image data, a target area image containing a vehicle tile is selected from each entrance area image data as a vehicle matching image of the vehicle to be measured.

[0048] The image acquisition source refers to whether the entrance area image data is acquired by one image acquisition device or multiple image acquisition devices, that is, whether multiple cameras are needed to identify the entrance position. In order to avoid the same vehicle being recorded by multiple cameras, resulting in multiple vehicles being considered, or the same vehicle being tracked multiple times, causing resource waste, the embodiment determines the final vehicle matching image according to different image acquisition devices.

[0049] For example, when there are multiple cameras, it is necessary to identify whether the vehicle is the same vehicle. The application also provides a simple same vehicle identification method: acquiring a first fixed pixel position of a fixed reference in a first overlapping area image, acquiring a first vehicle pixel position corresponding to a first vehicle contained in the first overlapping area, and determining a first pixel distance between the first vehicle and the fixed reference according to the first vehicle pixel position and the first fixed pixel position; acquiring a second fixed pixel position of the fixed reference in a second overlapping area image, acquiring a second vehicle pixel position corresponding to a second vehicle contained in the second overlapping area, and determining a second pixel distance between the second vehicle and the fixed reference according to the second vehicle pixel position and the second fixed pixel position; if the first pixel distance and the second pixel distance satisfy a preset same similarity condition, the first vehicle and the second vehicle are the same vehicle.

[0050] As can be seen from the above, the embodiment first detects the entrance vehicle, avoiding detection by all cameras in the park and saving computing resources. For multiple cameras, the same vehicle is identified according to the same or similar distance calculated in different field of view areas, effectively improving vehicle tracking accuracy and thus effectively improving speed detection accuracy.

[0051] Further, based on the above embodiment, the application further provides the following embodiment: target matching is performed on the video stream image frame and the vehicle matching image with the same size, to obtain a target vehicle with the same vehicle identification information as the vehicle to be tested, a target vehicle region of the target vehicle in the corresponding image frame and a corresponding score; if the score is less than or equal to a preset score threshold, the vehicle matching image is updated to the target vehicle region; if the score is greater than the preset score threshold, a driving trajectory is generated according to the position information of the vehicle to be tested in the video stream image frame, and it is determined that the vehicle to be tested drives out of the current field of view according to the driving trajectory, then the field of view driving speed is determined according to the driving time and the driving path of the vehicle to be tested in the current field of view.

[0052] In the embodiment, in order to improve the matching accuracy, the size of the vehicle matching image and the video stream image frame can be kept consistent. For example, the size of the vehicle matching image can be ensured to be the same as that of the video stream image frame by padding the vehicle matching image. For example, pixel points with pixel values of [114, 114, 114] are used to pad the edges of the vehicle matching image without changing the size of the vehicle matching image, so that the picture size of the vehicle matching image is consistent with that of the video stream image frame. If there are multiple vehicles with the same vehicle category, that is, multiple vehicles to be tested, the average score of the scores of each vehicle to be tested can be used as the final score. The preset score threshold is a preset value. In addition, a vehicle matching image quantity threshold of the same vehicle can be used. If the number of vehicle matching images of the vehicle is greater than the set threshold, the vehicle matching image is updated and the earliest vehicle matching image is discarded.

[0053] As can be seen from the above, the vehicle matching image is expanded to the same size as the video stream image frame to be detected in the embodiment, and the size of the vehicle matching image is not changed, which is more conducive to feature extraction by the original feature extraction network and retains the effective values of the feature part. Multiple vehicle matching images are matched, and the vehicle matching image is updated in a rolling manner, so that the latest template is more conducive to vehicle matching and improves the vehicle matching accuracy.

[0054] The above embodiment does not limit vehicle identification. Based on the above embodiment, the application further provides an exemplary implementation of vehicle identification, which can include the following content:

[0055] The target reference region image is input into the pre-trained vehicle detection network model, and the position of the vehicle tile contained in the target reference region image is determined according to the output result of the vehicle detection network model.

[0056] In the embodiment, the network structure of the vehicle detection network model is as follows: Figure 3As shown, the vehicle detection network model at least includes a first input layer, a feature extraction layer, a feature fusion layer, and a first output layer; the feature extraction layer includes a plurality of feature extraction sub-layers; the feature extraction layer performs image feature extraction on the target reference region image transmitted through the first input layer, and inputs the generated feature map to the feature fusion layer; the feature fusion layer outputs the position of the vehicle patch through the first output layer by multiple sampling and fusing the feature sub-map of at least one target feature extraction sub-layer.

[0057] After the network structure of the vehicle network model is determined, any one of the training sample sets for identifying the vehicle position and the vehicle category can be used to train the above vehicle network model by any one of the model training methods until the vehicle network model converges or the preset iteration number is reached or the accuracy reaches the preset accuracy threshold. In the training process of the vehicle network model, the total loss function loss can include the vehicle position loss lossp, the classification loss lossc, and the confidence loss losss, and the calculation relationship of each loss is as follows:

[0058] ;

[0059] ;

[0060] ;

[0061] .

[0062] wherein i represents the i-th training sample in the training sample set, j represents the j-th vehicle category, C represents the total number of categories, n represents the total number of training samples included in the training sample set, represents the Euclidean distance, c represents the normalization constant, is the predicted target frame, is the real target frame; is the real classification of the i-th training sample belonging to the j-th vehicle category, is the predicted classification of the i-th training sample belonging to the j-th vehicle category, is whether the i-th training sample belongs to the j-th vehicle category, and is 0 or 1, is the probability of the i-th training sample belonging to the j-th vehicle category, represents the respective weight coefficient, which can be an empirical value or can be determined in the training process.

[0063] For example, as shown in Figure 3 , the feature extraction layer is Figure 3The left part structure of the vehicle detection network model can include an initial feature extraction layer and a plurality of semantic feature extraction layers with the same structure; the initial feature extraction layer is connected with the first input layer, and the output thereof is connected with the input of the first semantic feature extraction layer, as shown in Figure 4 The initial feature extraction layer can include at least a two-dimensional convolution layer, a batch normalization layer and an activation function layer, such as a SiLU (function name) activation function. The output of each semantic feature extraction layer is connected with the input of the next semantic feature extraction layer, and the last semantic feature extraction layer is connected with the first output layer; each semantic feature extraction layer includes a first convolution layer and a multi-scale feature extraction layer with the same convolution kernel size, step length and padding parameters, and the number of channels increases; as shown in Figure 5 The multi-scale feature extraction layer includes at least a second convolution layer, a feature splitting layer, a local feature extraction layer, a semantic fusion layer and a third convolution layer. The network parameters of each layer of the vehicle detection network model can be as shown in Figures 3-6 The first convolution layer extracts the features of the received feature map and inputs the extracted new feature map to the second convolution layer, the second convolution layer is connected with the first convolution layer, extracts the features of the feature map output by the first convolution layer, and outputs a new feature map to the feature splitting layer, the feature splitting layer splits the received feature map into a first sub-feature map and a second sub-feature map, inputs the first sub-feature map to the local feature extraction layer, and inputs the second sub-feature map to the semantic fusion layer; the local feature extraction layer includes a plurality of feature extraction blocks with the same structure and connected in series, and the total number of the feature extraction blocks is determined according to the number of input channel dimensions, as shown in Figure 6 The feature extraction block includes two convolution layers with the same convolution kernel size and different step lengths; the output of the first feature extraction block of the local feature extraction layer is also connected to the semantic fusion layer, and the semantic fusion layer is connected with the last feature extraction block to output the fusion result of the received feature map to the third convolution layer.

[0064] As can be seen from the above, the semantic feature extraction layer splits the channels through the feature splitting layer first, and then merges them, which can retain the original features and the features after the feature extraction block. The feature extraction block is proportional to the number of input channels, and the more the channels, the more the stacked feature extraction blocks, thereby enhancing the feature extraction capability. When the number of input channels is small, the dimension is first increased and then decreased, which improves the nonlinear expression capability, reduces the parameter amount, improves the calculation efficiency, is suitable for lightweight models, and is more suitable for deployment on edge devices with limited resources. The multi-scale feature extraction layer does not change the shape of the input and output features, which is convenient for model building, and can better extract multi-scale features, improve the vehicle detection capability, and significantly reduce the model parameter amount and calculation amount while maintaining the detection accuracy, thereby realizing the lightweight and real-time of the target detection task.

[0065] For example, the feature fusion layer is Figure 3The right half of the structure includes at least the following sequentially connected layers: a first upsampling layer, a first splicing layer, a first multi-scale feature extraction layer, a second upsampling layer, a second splicing layer, a second multi-scale feature extraction layer, a fourth convolutional layer, a third splicing layer, a third multi-scale feature extraction layer, a fifth convolutional layer, and a fourth multi-scale feature extraction layer. Each multi-scale feature extraction layer has the same network result as the multi-scale feature extraction layer in the above embodiment, and the network parameters are shown in the corresponding diagram. Specifically, the first upsampling layer receives the output feature map from the feature extraction layer. The feature extraction layer includes, according to the data transmission direction, the first target multi-scale feature extraction layer and the second target multi-scale feature extraction layer with the largest parameter scale. The first splicing layer is also connected to the second target multi-scale feature extraction layer, the second splicing layer is also connected to the first target multi-scale feature extraction layer, and the third splicing layer is also connected to the first multi-scale feature extraction layer.

[0066] For example, such as Figure 7 As shown, the first output layer may include a confidence score calculation branch, a vehicle type calculation branch, and a bounding box calculation branch. Each of these branches includes a depthwise classifiable convolutional layer and a two-dimensional convolutional layer with identical kernel size, stride, and padding. The depthwise classifiable convolutional layer performs independent convolution operations on each channel of the feature fusion layer's output features, enabling parallel computation between channels. The confidence score calculation branch processes the feature fusion layer's output features through the depthwise classifiable convolutional layer and the two-dimensional convolutional layer, outputting the confidence score corresponding to the vehicle contained in the target reference region image. The vehicle type calculation branch processes the feature fusion layer's output features through the depthwise classifiable convolutional layer and the two-dimensional convolutional layer, outputting the vehicle category corresponding to the vehicle contained in the target reference region image, such as the vehicle ID. The bounding box calculation branch processes the feature fusion layer's output features through the depthwise classifiable convolutional layer and the two-dimensional convolutional layer, outputting the vehicle's location information in the target reference region image.

[0067] This embodiment calculates the confidence, category, and location of image features of the target reference region image by using depth-separable convolutional layers and two-dimensional convolutional layers. By using channel-wise convolution, the computational load and parameter count are significantly reduced while maintaining the receptive field, making it particularly suitable for resource-constrained devices such as mobile devices and edge servers. In addition, each convolutional kernel operates only on the data of a single channel, avoiding the complexity of inter-channel mixing operations in traditional convolution.

[0068] To enable those skilled in the art to better understand the processing procedure of the vehicle detection network model provided by this invention on the input target reference region image, this invention also provides... Figure 3 The structure of the vehicle detection network model shown is illustrated, and the image processing procedure of the vehicle detection network model can be described, including the following:

[0069] A1: input the vehicle target reference region image.

[0070] A2: the target reference region image enters the Conv layer, and is processed by Conv2d (two-dimensional convolution layer), BN (Batch Normalization, batch normalization) and SiLU respectively to obtain a new feature map.

[0071] In the formula, the size, stride s, padding p and channel number c of the convolution kernel K of Conv2d can be 6, 2, 2 and 64 respectively. s represents the distance of the convolution kernel sliding on the input image, p is used to add extra pixels at the edge of the original data to keep the output dimension unchanged or expand the receptive field, and c represents the dimension of the feature map after convolution.

[0072] A3: the feature map output by the A2 step is processed by the first Conv layer to obtain a new feature map.

[0073] A4: the feature map output by the A3 step is processed by the first multi-scale feature extraction layer to obtain a new feature map.

[0074] In this step, the number of feature extraction blocks of the multi-scale feature extraction layer is 3xd (input channel dimension).

[0075] A5: the feature map output by the A4 step is processed by the next first Conv layer to obtain a new feature map.

[0076] A6: the feature map output by the A5 step is processed by the second multi-scale feature extraction layer to obtain a new feature map, denoted as F1.

[0077] In the formula, the number of feature extraction blocks of the second multi-scale feature extraction layer is 6xd, which is also connected to the second splicing layer of the feature fusion layer.

[0078] A7: F1 is processed by the next first Conv layer to obtain a new feature map.

[0079] A8: the feature map output by the A7 step is processed by the third multi-scale feature extraction layer to obtain a new feature map, denoted as F2.

[0080] The number of feature extraction blocks of the third multi-scale feature extraction layer is 6xd, which is also connected to the first splicing layer of the feature fusion layer.

[0081] A9: F2 is processed by the next first Conv layer to obtain a new feature map.

[0082] A10: the feature map output by the A9 step is processed by the fourth multi-scale feature extraction layer to obtain a new feature map.

[0083] The step before A10 completes feature extraction, and the output feature map is input into the feature fusion layer.

[0084] A11: The feature map output by the step of A10 is input into the first up-sampling layer of the feature fusion layer, to obtain a new feature map, denoted as F3.

[0085] A12: F3 and F2 are merged by the first splicing layer.

[0086] Through the feature fusion in this step, the loss of feature information is prevented, and the feature extraction capability and detection capability of the model are increased.

[0087] A13: The merged feature is input into the multi-scale feature extraction layer, to obtain a new feature map, denoted as F4.

[0088] A14: F4 is up-sampled by the second up-sampling layer, to obtain a new feature map.

[0089] A15: The new feature map generated in the step of A14 and F1 are merged by the second splicing layer, to prevent the loss of feature information and increase the feature extraction capability and detection capability of the model.

[0090] A16: The second merged feature is input into the multi-scale feature extraction layer, to obtain a new feature map.

[0091] A17: The feature map generated in the step of A16 is input into the fourth Conv layer, to obtain a new feature map.

[0092] A18: The new feature map generated in A17 and F4 are merged by the third splicing layer, to prevent the loss of feature information and increase the feature extraction capability and detection capability of the model.

[0093] A19: The third merged feature is input into the multi-scale feature extraction layer again, to obtain a new feature map.

[0094] A20: The new feature map generated in the step of A19 is input into the fifth Conv layer, to obtain a new feature map.

[0095] A21: The new feature map generated in the step of A20 is input into the multi-scale feature extraction layer, to obtain a new feature map.

[0096] A22: The new feature map generated in the step of A21 is input into the first output layer, to obtain the position, classification and confidence of the vehicle.

[0097] The above embodiment does not limit the vehicle matching process, and based on the above embodiment, the present application further provides an exemplary implementation of vehicle matching, which can include the following contents:

[0098] The vehicle matching network model is used to complete the vehicle matching process between the vehicle matching image and the video stream image frames of each field of view. The vehicle matching network model can be as shown in FIG. 8, which includes three network structures: a matching feature extraction network, a video feature extraction network and a feature matching network.

[0099] In the embodiment, the feature extraction parameters of the feature extraction layer of the vehicle detection network model for vehicle recognition of the target reference area image can be acquired; according to the network structure of the feature extraction layer, the network structure of the matching feature extraction network for extracting the vehicle matching image features of the vehicle matching image and the network structure of the video feature extraction network for extracting the image features of each frame of the video stream image frames are determined, the feature extraction parameters of the matching feature extraction network and the video feature extraction network are consistent with the feature extraction parameters, and the parameter amount of the total model is reduced. That is, the network model structure of the vehicle matching network model for image feature extraction of the vehicle matching image and the video stream image frames is the same as the network model structure of the vehicle detection network model for feature extraction and the model parameters are the same. The input end of the feature matching network is connected with the outputs of the matching feature extraction network and the video feature extraction network respectively, is used to receive the vehicle matching image features and the image features of each frame, matches the vehicle matching image features and the image features of each frame, and outputs the vehicle categories, vehicle positions and corresponding confidence scores of each vehicle contained in the video stream image frames through the second output layer. Thus, the vehicle tracking data matched with the vehicle to be measured can be determined according to the output result of the vehicle matching network model.

[0100] Exemplarily, the feature matching network comprises a second input layer, a plurality of target matching layers, a detection head and a second output layer; each target matching layer comprises a position feature fusion layer, a plurality of simultaneously operated operation processing layers, a first multi-feature fusion layer, a layer normalization layer, a full connection layer and a second multi-feature fusion layer connected in sequence; the second input layer is connected with the video feature extraction network and the matching feature extraction network respectively, performs flattening operation on the received image features and vehicle matching image features respectively, and inputs the video image flattened features into the first operation processing layer and the matching image flattened features into the position feature fusion layer and the first multi-feature fusion layer; the vehicle matching image features are positionally encoded in a manner that the background block is set to infinity, and the positionally encoded features are input into the position feature fusion layer. Exemplarily, the second input layer comprises a first flattening layer, a second flattening layer and a position encoding layer, the first flattening layer is connected with the output of the video feature extraction network, performs flattening operation on the received image features, and inputs the video image flattened features into the first operation processing layer; the second flattening layer is connected with the output of the matching feature extraction network, performs flattening operation on the vehicle matching image features, and inputs the matching image flattened features into the position feature fusion layer and the first multi-feature fusion layer; the position encoding layer positionally encodes the vehicle matching image features in a manner that the background block is set to infinity, and inputs the positionally encoded features into the position feature fusion layer. Each operation processing layer performs matrix multiplication operation on the position feature fusion data and the video image flattened features, performs scaling operation on the first operation result, processes the scaling result using an activation function such as softmax, performs matrix multiplication operation on the processing result and the video image flattened features, and inputs the second operation result into the first multi-feature fusion layer, and the output end of the first multi-feature fusion layer is connected to the layer normalization layer and the second multi-feature fusion layer respectively.

[0101] In the embodiment, the position encoding layer sets the target-free region to negative infinity, avoids interference with vehicle matching, and improves vehicle detection accuracy. The operation processing layer can filter the invalid region in the vehicle matching image flattened features through feature addition, which not only reduces the matching data amount, but also avoids interference with vehicle matching, and improves vehicle detection accuracy. The scaling processing can prevent the dot product value from being too large, can avoid the situation that the dot product value may become very large with the increase of the vector dimension, causes the gradient of the softmax function to be very small, and further causes the gradient to disappear, thereby affecting the training effect of the model. The scaling operation makes the dot product result more stable, ensures that the softmax function works in a proper range, and thus ensures the training stability of the model.

[0102] In order to make the vehicle matching process of the vehicle matching network model provided by the present application more clear to those skilled in the art, the present application also provides a vehicle matching method using the vehicle matching network model. Figure 8The structure of the vehicle matching network model is shown, and the vehicle matching process of the vehicle matching network model is described, which can include the following contents:

[0103] B1: input video stream image frames or vehicle matching pictures.

[0104] The feature extraction process corresponding to the video stream image frames or the vehicle matching pictures can include: the video stream image frames or the vehicle matching pictures pass through the Conv layer to obtain a new feature map, denoted as FB1, FB1 passes through the Conv layer to obtain a new feature map, denoted as FB2, and then FB2 passes through the multi-scale feature extraction layer to obtain a new feature map, denoted as FB3. Then FB3 passes through the Conv layer to obtain a new feature map, denoted as FB4, FB4 passes through the multi-scale feature extraction layer to obtain a new feature map, denoted as FB5, FB5 passes through the Conv layer to obtain a new feature map, denoted as FB6, FB6 passes through the multi-scale feature extraction layer to obtain a new feature map, denoted as FB7; FB7 passes through the Conv layer to obtain a new feature map, denoted as FB8; FB8 passes through the multi-scale feature extraction layer to obtain a new feature map, denoted as FB9.

[0105] B2: flatten the video stream image frames to obtain features F11, and flatten the vehicle matching pictures to obtain features F22.

[0106] B3: perform a matching process on the features F11 and the features F22:

[0107] B3.1: F22 is positionally encoded according to the position of the vehicle matching picture, and the non-target region is set to negative infinity.

[0108] B3.2: perform a corresponding position addition operation on the position encoding and the F22 features to filter out invalid regions in F22.

[0109] B3.3: perform a matrix multiplication operation on the features F11 and the filtered features F22 to obtain new features F33. Scale the new features F33 by dividing by a specific value, and perform a softmax operation on the scaling result to obtain new features F44; perform a matrix multiplication operation on the new features F44 and F11 to obtain new features F55.

[0110] Wherein, this step is run simultaneously through M modules, and finally the obtained features are combined together to obtain new features F66.

[0111] B3.4: add the new features F66 and the features F22 to obtain new features F77.

[0112] Through the feature addition in this step and the following steps, the vehicle features can be highlighted, and the network can be prevented from being too deep to pass forward.

[0113] B3.5: Perform layer normalization on the new feature F77 to obtain the new feature F88.

[0114] B3.6: Pass the new feature F88 through a fully connected layer to obtain the new feature F99.

[0115] B3.7: Add the new feature F99 and feature F77 to obtain the new feature F0.

[0116] There are N target matching layers, which are connected in series. The value of N can be dynamically set according to the matching effect.

[0117] As can be seen from the above, the vehicle matching network model and the vehicle detection network model of this invention share the model parameters for feature extraction, reducing the total number of model parameters, improving vehicle detection efficiency, and making it easier to deploy on edge devices. Furthermore, a positional encoding method is adopted in the target search stage, setting the irrelevant regions of the vehicle matching image to negative infinity, avoiding interference from irrelevant parts, improving the search efficiency of the vehicle matching network model, and thus improving vehicle tracking efficiency and vehicle speed measurement efficiency.

[0118] For example, such as Figure 9 As shown, the detection head may include a video frame detection layer and a matching feature processing layer; wherein, the video frame detection layer is connected to the second multi-feature fusion layer, including a confidence calculation branch, a vehicle category recognition branch, and a vehicle position calculation branch; the confidence calculation branch, the vehicle category recognition branch, and the vehicle position calculation branch include convolutional layers with the same kernel size, stride value, and padding value, and two-dimensional convolutional layers with the same kernel size, stride value, and padding value but different number of channels; the matching feature processing layer receives vehicle matching image features, including a sixth convolutional layer, a feature addition layer, and a feature multiplication layer connected in sequence, the sixth convolutional layer... The output of the layer is also connected to the feature multiplication layer, and the feature addition layer is also connected to the output of the convolutional layer of the vehicle category recognition branch. The output of the feature multiplication layer is connected to the two-dimensional convolutional layer of the vehicle category recognition branch. The vehicle category recognition branch also includes a category determination layer. After the category determination layer is connected to the two-dimensional convolutional layer of the vehicle category recognition branch, when the vehicle category recognition result of the target vehicle corresponding to the current image frame is received, the distance between the target vehicle and other vehicle categories is calculated, and when the change in distance between the target vehicle and other vehicle categories meets the preset distance change condition, the vehicle category recognition result of the target vehicle is output.

[0119] In the embodiment, the features obtained by the feature matching process, such as the features F0 in the above embodiment, are input to the detection head part. The features F0 pass through the Conv of the confidence calculation branch to obtain new features, and the new features pass through the Conv2d to obtain the confidence. The features F0 pass through the Conv of the vehicle position calculation branch to obtain new features, and the new features pass through the Conv2d to obtain the target box in the image where the vehicle is located, that is, the position information. The features F0 pass through the Conv of the vehicle category recognition branch to obtain new features F31, and the vehicle matching image features pass through the Conv of the matching feature processing layer to obtain new features F41; the features F31 and the features F41 are added bit by bit to obtain new features F42, and the new features F42 and the features F41 are multiplied point by point to obtain new features F43, and the new features F43 pass through the Conv2d of the vehicle category recognition branch to obtain the ID of the vehicle; the distance between the vehicle and other categories is calculated; the calculated distance value can be stored in the database, and the change of the distance value is compared to determine whether the set threshold is met, so as to avoid vehicle detection errors and the situation that the distance between the ID and other targets is far or near. If a specific threshold is met, the predicted vehicle category is output.

[0120] As can be seen from the above, the detection head of the embodiment detects the correlation of the recognized vehicle and other categories different from the predicted category, that is, the distance detection of other models, which is beneficial to improve the recognition accuracy of the vehicle category and reduce misjudgment.

[0121] Finally, in order for those skilled in the art to more clearly understand the technical solutions of the present application, the present application further provides another implementation process for vehicle speed detection in a park, which can include the following contents:

[0122] S1: Obtain the video stream collected by the entrance camera of the park, and extract each frame image from the video stream.

[0123] S2: If the entrance area has multiple cameras and the fields of view of the cameras overlap, S3 is executed. If the entrance area has one camera or the fields of view of the cameras do not overlap, S9 is executed.

[0124] S3: Input all pictures of the park entrance to the vehicle detection network model.

[0125] S4: Determine whether there is a vehicle in the picture, if not, return to S1, if yes, continue to execute S4.

[0126] S5: Determine whether the vehicle is located in the overlapping area, if yes, execute S5, if not, execute S9.

[0127] S6: Extract the pictures in different fields of view in the area.

[0128] S7: Confirm whether the target in the region is the same vehicle, and mark the ID.

[0129] The same vehicle can be determined according to the distance between the vehicle and the fixed object in the field of view: the pixel coordinates of the same fixed object in different fields of view are extracted; the distance between the vehicle in the detection image and the fixed object in the image is calculated; the distance between the vehicle and the fixed object in different fields of view is sorted; the same target in different fields of view is determined according to the distance sorting, that is, the ID of the same vehicle is determined according to the distance sorting.

[0130] S8: Extract the target region containing the vehicle tile in different fields of view, and use the target region as the subsequent matching template, that is, the vehicle matching image.

[0131] The target region in one field of view is used as a matching template, and one target can use multiple templates for matching.

[0132] S9: Extract the target region as a matching template, that is, a vehicle matching image.

[0133] S10: Enter the target matching process:

[0134] S10.1: Continuously acquire the video stream of the current field of view, and fill the vehicle matching image with gray pixels until the size of the video stream picture.

[0135] S10.2: Input the vehicle matching image and the real-time video stream picture into the vehicle matching network model.

[0136] S10.3: Obtain the matched vehicle ID, position and confidence score from the real-time video stream picture.

[0137] When there are multiple vehicles with the same ID, the average score after matching multiple vehicles is calculated as the final score.

[0138] S10.4: Determine whether the score is greater than the set threshold, if not, execute S10.6, otherwise, execute S10.5.

[0139] S10.5: Extract the newly matched ID region according to the position information output by the vehicle matching network model, and use it as the vehicle matching image.

[0140] When the number of templates is greater than the set threshold, the templates are updated and the earliest template is discarded.

[0141] S10.6: Draw the trajectory of all IDs.

[0142] S10.7: judging whether there is a vehicle moving out of the current field of view according to the trajectory drawn in the above steps, if ID1 moves out of the field of view, S10.8 is executed, if not, S10.1 is returned to execute.

[0143] S10.8: calculating the time when ID1 moves out of the field of view.

[0144] S10.9: calculating the distance of all paths in the field of view and the average speed of ID1 in the field of view.

[0145] S11: extracting the video stream in the next field of view according to the path planning in the park and the trajectory of ID1.

[0146] S12: performing target matching on the video stream extracted in S11 according to the above steps S10.1-S11.

[0147] S13: judging whether ID1 moves out of the park, if not, S11 is returned to execute, if yes, S14 is executed.

[0148] S14: obtaining the time when ID1 moves in and the time when ID1 moves out, and calculating the time difference between the two.

[0149] S15: calculating the length of the moving path of ID1 in the park, and calculating the average speed according to the time difference in S14.

[0150] S16: calculating the average speed of ID1 in each field of view of the camera, and calculating the positive value of the difference between the average speed in each field of view and the average speed in the park in S15. When the positive value of the speed difference is less than a preset threshold, it is considered that the speed of ID1 is calculated correctly.

[0151] S17: judging whether the single field of view speed and the average speed of ID1 exceed a set threshold, if yes, an alarm information is generated, if not, S1 is executed.

[0152] S18: ending the program.

[0153] It should be noted that there is no strict execution sequence between the steps in the present application, as long as the logical sequence is met, the steps can be executed simultaneously, or executed according to a certain preset sequence, Figure 2 and Figure 10 it is only an illustrative way, and does not mean that only such an execution sequence can be used.

[0154] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.

[0155] The present application also provides a corresponding device for the vehicle speed detection method, further making the method more practical. The device can be described from the perspective of functional modules and the perspective of hardware. The vehicle speed detection device provided by the present application is introduced below, which is used to implement the vehicle speed detection method provided by the present application. In this embodiment, the vehicle speed detection device can include or be divided into one or more program modules, which are stored in a storage medium and executed by one or more processors to complete the vehicle speed detection method disclosed in embodiment one. The program module referred to in this embodiment refers to a series of computer program instruction segments capable of completing a specific function, which is more suitable for describing the execution process of the vehicle speed detection device in the storage medium than the program itself. The following description will specifically introduce the functions of each program module in this embodiment. The vehicle speed detection device described below can be referred to in conjunction with the vehicle speed detection method described above.

[0156] From the perspective of functional modules, refer to Figure 11 , Figure 11 The structure diagram of the vehicle speed detection device provided in this embodiment, which can include:

[0157] The template generation module 111 is configured to determine a vehicle matching image containing a vehicle to be measured based on a vehicle recognition result of a target reference area image corresponding to a target reference position of a speed measurement area.

[0158] The vehicle tracking module 112 is configured to determine vehicle tracking data matched with the vehicle to be measured from a video stream image frame of a current field of view based on the vehicle matching image; when it is determined that the vehicle drives away from the current field of view according to the vehicle tracking data, determine a target field of view corresponding to a current driving position of the vehicle according to speed measurement area path planning data, and determine new vehicle tracking data in the target field of view again based on the current vehicle matching image.

[0159] The speed measurement module 113 is configured to determine an average driving speed of the vehicle to be measured in the speed measurement area and a driving speed of each field of view based on the vehicle tracking data of each field of view passed by the vehicle to be measured during driving, and determine the driving speed of the vehicle to be measured based on the average driving speed and the driving speed of each field of view.

[0160] For example, in some embodiments of the present embodiment, the speed measurement module 113 can be further configured to: determine, for each field of view through which the vehicle to be measured passes during the driving process, the appearing time and the leaving time of the vehicle to be measured in the current field of view according to the vehicle tracking data of the current field of view, and determine the driving speed of the vehicle to be measured in the current field of view according to the path length of the current field of view; when it is determined that the vehicle to be measured leaves the speed measurement area according to the vehicle tracking data of the last field of view through which the vehicle to be measured passes during the driving process, determine the total driving time of the vehicle to be measured in the speed measurement area according to the entering time of the first field of view through which the vehicle to be measured passes during the driving process and the leaving time of the last field of view, and determine the total driving length of the vehicle to be measured in the speed measurement area according to the path planning data of the speed measurement area; and determine the average driving speed of the vehicle to be measured in the speed measurement area according to the total driving length and the total driving time.

[0161] For example, in some other embodiments of the present embodiment, the template generation module 111 can be further configured to: when the target reference position is an entrance position and the target reference area image is entrance area image data, obtain the entrance area image data of the entrance position of the speed measurement area; if the entrance area image data is from the same image acquisition source, or there is no vehicle located in the overlapping field of view of different image acquisition sources in each entrance area image data, select the target area image containing the vehicle tile from each entrance area image data as the vehicle matching image of the vehicle to be measured; if the entrance area image data is from different image acquisition sources, and at least one vehicle to be measured in the entrance area image data is located in the overlapping field of view of at least two image acquisition sources, obtain the overlapping area image obtained by image acquisition of the overlapping area from multiple fields of view; and if there are at least two target overlapping area images containing vehicle tiles that are not the same vehicle, select the target area image containing the vehicle tile from each target overlapping area image as the vehicle matching image corresponding to each vehicle to be measured.

[0162] For example, in some embodiments of the present embodiment, the speed measurement module 113 can be further configured to: determine, for each field of view through which the vehicle to be measured passes during the driving process, the appearing time and the leaving time of the vehicle to be measured in the current field of view according to the vehicle tracking data of the current field of view, and determine the driving speed of the vehicle to be measured in the current field of view according to the path length of the current field of view; when it is determined that the vehicle to be measured leaves the speed measurement area according to the vehicle tracking data of the last field of view through which the vehicle to be measured passes during the driving process, determine the total driving time of the vehicle to be measured in the speed measurement area according to the entering time of the first field of view through which the vehicle to be measured passes during the driving process and the leaving time of the last field of view, and determine the total driving length of the vehicle to be measured in the speed measurement area according to the path planning data of the speed measurement area; and determine the average driving speed of the vehicle to be measured in the speed measurement area according to the total driving length and the total driving time.

[0163] Exemplarily, in some other embodiments of the present embodiment, the speed measurement module 113 can also be configured to: determine the average speed in the field of view according to the field of view driving speed and the total number of fields of view; when the speed difference between the average driving speed and the average speed in the field of view meets the preset speed threshold judgment condition, the average driving speed of the vehicle to be measured and the field of view driving speed meet the speed measurement error-free condition; and when at least one of the average driving speed and the field of view driving speed is greater than the preset speed limit value, the vehicle to be measured has a speeding behavior in the speed measurement area.

[0164] Exemplarily, in some other embodiments of the present embodiment, the vehicle tracking module 112 can also be configured to: perform target matching on the video stream image frame and the vehicle matching image with the same size to obtain a target vehicle with the same vehicle identification information as the vehicle to be measured, a target vehicle region of the target vehicle in the corresponding image frame and a corresponding score; if the score is less than or equal to a preset score threshold, updating the vehicle matching image to the target vehicle region; if the score is greater than the preset score threshold, generating a driving trajectory according to the position information of the vehicle to be measured in the video stream image frame, and determining that the vehicle to be measured drives out of the current field of view according to the driving trajectory, then determining the field of view driving speed according to the driving time and the driving path of the vehicle to be measured in the current field of view.

[0165] Exemplarily, in some other embodiments of the present embodiment, the vehicle tracking module 112 can also be configured to: input the target reference region image into a pre-trained vehicle detection network model, and determine the position of the vehicle tile contained in the target reference region image according to the output result of the vehicle detection network model; wherein the vehicle detection network model at least includes a first input layer, a feature extraction layer, a feature fusion layer and a first output layer; the feature extraction layer includes a plurality of feature extraction sub-layers; the feature extraction layer performs image feature extraction on the target reference region image transmitted through the first input layer, and inputs the generated feature map into the feature fusion layer; the feature fusion layer fuses at least one target feature extraction sub-layer feature sub-map through multiple samplings, and outputs the position of the vehicle tile through the first output layer.

[0166] As an exemplary implementation of the above embodiment, the feature extraction layer includes an initial feature extraction layer and a plurality of groups of semantic feature extraction layers with the same structure; the initial feature extraction layer is connected to the first input layer, and the output thereof is connected to the input of the first semantic feature extraction layer; the output of each semantic feature extraction layer is connected to the input of the next semantic feature extraction layer; the last semantic feature extraction layer is connected to the first output layer; the initial feature extraction layer includes at least a two-dimensional convolution layer, a batch normalization layer and an activation function layer; each semantic feature extraction layer includes a first convolution layer with the same convolution kernel size, step and padding parameters and an increasing number of channels, and a multi-scale feature extraction layer; the multi-scale feature extraction layer includes at least a second convolution layer, a feature splitting layer, a local feature extraction layer, a semantic fusion layer and a third convolution layer; the second convolution layer is connected to the corresponding first convolution layer and outputs the extracted feature map to the feature splitting layer; the feature splitting layer splits the received feature map into a first sub-feature map and a second sub-feature map; the first sub-feature map is input to the local feature extraction layer, and the second sub-feature map is input to the semantic fusion layer; the local feature extraction layer includes a plurality of feature extraction blocks with the same structure and connected in series; the total number of the feature extraction blocks is determined according to the number of input channel dimensions; the feature extraction block includes two convolution layers with the same convolution kernel size and different steps; the output of the first feature extraction block of the local feature extraction layer is also connected to the semantic fusion layer; the semantic fusion layer is connected to the last feature extraction block and outputs the fusion result of the received feature map to the third convolution layer.

[0167] As an exemplary implementation of the above embodiment, the feature fusion layer includes at least a first up-sampling layer, a first splicing layer, a first multi-scale feature extraction layer, a second up-sampling layer, a second splicing layer, a second multi-scale feature extraction layer, a fourth convolution layer, a third splicing layer, a third multi-scale feature extraction layer, a fifth convolution layer and a fourth multi-scale feature extraction layer connected in sequence; the first up-sampling layer receives the output feature map of the feature extraction layer; the feature extraction layer includes, in the data transmission direction, a first target multi-scale feature extraction layer and a second target multi-scale feature extraction layer with the largest parameter scale; the first splicing layer is also connected to the second target multi-scale feature extraction layer; the second splicing layer is also connected to the first target multi-scale feature extraction layer; and the third splicing layer is also connected to the first multi-scale feature extraction layer.

[0168] In some other embodiments of the present embodiments, the vehicle tracking module 112 can further be configured to: obtain feature extraction parameters of a feature extraction layer of a vehicle detection network model for vehicle recognition on a target reference area image; determine a matching feature extraction network for extracting vehicle matching image features of the vehicle matching image and a video feature extraction network for extracting frame image features of the video stream image frames according to a network structure of the feature extraction layer, the matching feature extraction network and the video feature extraction network having the same feature extraction parameters; combine the matching feature extraction network, the video feature extraction network and a feature matching network into a vehicle matching network model; connect the input end of the feature matching network to the output ends of the matching feature extraction network and the video feature extraction network respectively, and match the vehicle matching image features and the frame image features to output vehicle categories, vehicle positions and corresponding confidence scores of each vehicle contained in the video stream image frames; and determine vehicle tracking data matched with the vehicle to be tested according to the output result of the vehicle matching network model.

[0169] As an exemplary implementation of the above embodiment, the feature matching network comprises a second input layer, a plurality of target matching layers, a detection head and a second output layer; each target matching layer comprises a position feature fusion layer, a plurality of simultaneously operated operation processing layers, a first multi-feature fusion layer, a layer normalization layer, a full connection layer and a second multi-feature fusion layer connected in sequence; the second input layer is connected to the video feature extraction network and the matching feature extraction network respectively, performs flattening operation on the received frame image features and vehicle matching image features respectively, and inputs the video image flattened features to the first operation processing layer and the matching image flattened features to the position feature fusion layer and the first multi-feature fusion layer; the vehicle matching image features are position encoded in a manner that the background block is set to infinity, and the position encoded features are input to the position feature fusion layer; each operation processing layer performs matrix multiplication operation on the position feature fusion data and the video image flattened features, performs scaling operation on the first operation result, processes the scaling result using an activation function, performs matrix multiplication operation on the processing result and the video image flattened features, and inputs the second operation result to the first multi-feature fusion layer, and the output end of the first multi-feature fusion layer is connected to the layer normalization layer and the second multi-feature fusion layer respectively.

[0170] As another exemplary implementation of the above embodiment, the detection head comprises a video frame detection layer and a matching feature processing layer; the video frame detection layer is connected with the second multi-feature fusion layer and comprises a confidence calculation branch, a vehicle category identification branch and a vehicle position calculation branch; the confidence calculation branch, the vehicle category identification branch and the vehicle position calculation branch comprise convolution layers with the same convolution kernel size, step value and padding value, and two-dimensional convolution layers with the same convolution kernel size, step value and padding value and different channel numbers; the matching feature processing layer receives vehicle matching image features and comprises a sixth convolution layer, a feature addition layer and a feature multiplication layer connected in sequence, the output of the sixth convolution layer is further connected to the feature multiplication layer, the feature addition layer is further connected to the output of the convolution layer of the vehicle category identification branch, and the output of the feature multiplication layer is connected to the two-dimensional convolution layer of the vehicle category identification branch; the vehicle category identification branch further comprises a category determination layer, which is connected after the two-dimensional convolution layer of the vehicle category identification branch, and when receiving the vehicle category identification result of the target vehicle corresponding to the current image frame, calculates the distance between the target vehicle and other vehicle categories, and when the distance change between the target vehicle and other vehicle categories meets the preset distance change condition, outputs the vehicle category identification result of the target vehicle.

[0171] The features of the vehicle speed detection device in the corresponding embodiment can be referred to the related description of the vehicle speed detection method in the corresponding embodiment, which will not be repeated here.

[0172] The vehicle speed detection device mentioned above is described from the perspective of functional modules, and further, the present application also provides an electronic device, which is described from the perspective of hardware. Figure 12 The structure schematic diagram of the electronic device provided by the embodiment of the present application in an implementation manner is shown. The electronic device comprises a memory 121 and a processor 122, the memory 121 stores a computer program, and the processor 122 is configured to run the computer program to execute the steps in any of the above vehicle speed detection method embodiments.

[0173] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above vehicle speed detection method embodiments when running.

[0174] In an exemplary embodiment, the above computer readable storage medium can include but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0175] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program realizes the steps in any of the vehicle speed detection method embodiments when executed by a processor.

[0176] The embodiment of the present application further provides another computer program product, which comprises a nonvolatile computer readable storage medium, and the nonvolatile computer readable storage medium stores a computer program, and the computer program realizes the steps in any of the vehicle speed detection method embodiments when executed by a processor.

[0177] The vehicle speed detection method, the electronic device, the computer readable storage medium and the computer program product provided by the present application are described in detail above. Each embodiment in the specification is described in a progressive manner, and each embodiment mainly describes the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. Whether the units and algorithm steps of each example described by each disclosed embodiment are executed by electronic hardware or computer software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, and such implementation should not be considered beyond the scope of the present application. Without departing from the principles of the present application, the present application can be improved and modified, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A vehicle speed detection method characterized by comprising: The application comprises the following steps: According to the vehicle identification result of the target reference area image corresponding to the target reference position of the speed measurement area, a vehicle matching image containing the vehicle to be measured is determined. Based on the vehicle matching image, vehicle tracking data matching the vehicle to be measured is determined from the video stream image frames of the current field of view. When the vehicle tracking data indicates that the vehicle has left the current field of view, the target field of view corresponding to the current driving position of the vehicle is determined according to the path planning data of the speed measurement area, and new vehicle tracking data is determined in the target field of view according to the current vehicle matching image. According to the vehicle tracking data of each field of view passed by the vehicle to be measured during driving, the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of each field of view are determined, and the driving speed of the vehicle to be measured is determined according to the average driving speed and the driving speed of each field of view. The target reference position is the position of the first camera through which the vehicle enters the speed measurement area or the entrance of the speed measurement area. The vehicle identification result is used as the vehicle matching image for subsequent tracking of the vehicle to be measured, and vehicle detection is no longer performed on all image acquisition devices in the speed measurement area. If the target reference area image is derived from different image acquisition sources and at least one vehicle to be measured is located in the field of view overlap area of at least two image acquisition sources, the overlap area images obtained by image acquisition of the field of view overlap area from multiple fields of view are obtained, the first fixed pixel position of the fixed reference object in the first overlap area image is obtained, the first vehicle pixel position corresponding to the first vehicle contained in the first overlap area is obtained, and the first pixel distance between the first vehicle and the fixed reference object is determined according to the first vehicle pixel position and the first fixed pixel position. The second fixed pixel position of the fixed reference object in the second overlap area image is obtained, the second vehicle pixel position corresponding to the second vehicle contained in the second overlap area is obtained, and the second pixel distance between the second vehicle and the fixed reference object is determined according to the second vehicle pixel position and the second fixed pixel position. If the first pixel distance and the second pixel distance satisfy the preset same similarity condition, the first vehicle and the second vehicle are the same vehicle. The determination process of the average driving speed of the vehicle to be measured and the driving speed of each field of view is as follows: for each field of view passed by the vehicle to be measured during driving, the appearance time and the departure time of the vehicle to be measured in the current field of view are determined according to the vehicle tracking data of the current field of view, and the driving speed of the vehicle to be measured in the current field of view is determined according to the path length of the current field of view. When the vehicle tracking data of the last field of view passed by the vehicle to be measured during driving indicates that the vehicle to be measured has left the speed measurement area, the total driving time of the vehicle to be measured in the speed measurement area is determined according to the driving-in time of the first field of view and the driving-out time of the last field of view, and the total driving length of the vehicle to be measured in the speed measurement area is determined according to the path planning data of the speed measurement area. The average driving speed of the vehicle to be measured in the speed measurement area is determined according to the total driving length and the total driving time. The driving speed determination process of the vehicle to be tested comprises: determining an average speed in the field of view according to each field of view driving speed and the total number of fields of view; when the speed difference between the average driving speed and the average speed in the field of view meets a preset speed threshold judgment condition, the average driving speed of the vehicle to be tested and each field of view driving speed meet the speed measurement error-free condition.

2. The vehicle speed detection method according to claim 1, characterized by, The driving speed of the vehicle to be tested is determined according to the average driving speed and each field of view driving speed, comprising: When at least one speed value in the average driving speed and each field of view driving speed is greater than a preset speed limit value, the vehicle to be tested has an overspeed behavior in the speed measurement area.

3. The vehicle speed detection method according to claim 1, characterized by, Based on the vehicle matching image, the vehicle tracking data matched with the vehicle to be tested is determined from the video stream image frame of the current field of view, comprising: Target matching is performed on the video stream image frame and the vehicle matching image with the same size to obtain a target vehicle with the same vehicle identification information as the vehicle to be tested, a target vehicle area of the target vehicle in the corresponding image frame and a corresponding score; If the score is less than or equal to a preset score threshold, the vehicle matching image is updated to the target vehicle area; If the score is greater than the preset score threshold, a driving trajectory is generated according to the position information of the vehicle to be tested in the video stream image frame, and when it is determined that the vehicle to be tested drives away from the current field of view according to the driving trajectory, a field of view driving speed is determined according to the driving time and driving path of the vehicle to be tested in the current field of view.

4. The vehicle speed detection method according to any one of claims 1 to 3, characterized by, According to the vehicle recognition result of the target reference area image corresponding to the target reference position of the speed measurement area, a vehicle matching image containing the vehicle to be tested is determined, comprising: The target reference area image is input into a pre-trained vehicle detection network model, and the position of the vehicle tile contained in the target reference area image is determined according to the output result of the vehicle detection network model. The vehicle detection network model at least comprises a first input layer, a feature extraction layer, a feature fusion layer and a first output layer; the feature extraction layer comprises a plurality of feature extraction sub-layers; the feature extraction layer extracts image features from the target reference area image transmitted through the first input layer, and inputs the generated feature map into the feature fusion layer; the feature fusion layer fuses the feature sub-map of at least one target feature extraction sub-layer through multiple sampling, and outputs the position of the vehicle tile through the first output layer.

5. The vehicle speed detection method according to claim 4, characterized by, The feature extraction layer comprises an initial feature extraction layer and a plurality of groups of semantic feature extraction layers with the same structure; The initial feature extraction layer is connected with the first input layer, and its output is connected with the input of the first semantic feature extraction layer; the output of each semantic feature extraction layer is connected with the input of the next semantic feature extraction layer; the last semantic feature extraction layer is connected with the first output layer; The initial feature extraction layer is connected with the first input layer, and its output is connected with the input of the first semantic feature extraction layer; the output of each semantic feature extraction layer is connected with the input of the next semantic feature extraction layer; the last semantic feature extraction layer is connected with the first output layer; The initial feature extraction layer at least comprises a two-dimensional convolution layer, a batch normalization layer and an activation function layer; each semantic feature extraction layer comprises a first convolution layer and a multi-scale feature extraction layer, which have the same convolution kernel size, step and padding parameters and an increasing number of channels; the multi-scale feature extraction layer at least comprises a second convolution layer, a feature splitting layer, a local feature extraction layer, a semantic fusion layer and a third convolution layer; The second convolution layer is connected with the corresponding first convolution layer and outputs the extracted feature map to the feature splitting layer; the feature splitting layer splits the received feature map into a first sub-feature map and a second sub-feature map, inputs the first sub-feature map to the local feature extraction layer and inputs the second sub-feature map to the semantic fusion layer; the local feature extraction layer comprises a plurality of feature extraction blocks which have the same structure and are connected in series, the total number of the feature extraction blocks is determined according to the number of input channel dimensions, and each feature extraction block comprises two convolution layers which have the same convolution kernel size and different steps; the output of the first feature extraction block of the local feature extraction layer is also connected to the semantic fusion layer, and the semantic fusion layer is connected with the last feature extraction block and outputs the fusion result of the received feature map to the third convolution layer.

6. The vehicle speed detection method according to claim 4, characterized by The feature fusion layer at least comprises a first up-sampling layer, a first splicing layer, a first multi-scale feature extraction layer, a second up-sampling layer, a second splicing layer, a second multi-scale feature extraction layer, a fourth convolution layer, a third splicing layer, a third multi-scale feature extraction layer, a fifth convolution layer and a fourth multi-scale feature extraction layer which are connected in sequence; The first up-sampling layer receives the output feature map of the feature extraction layer, the feature extraction layer comprises a first target multi-scale feature extraction layer and a second target multi-scale feature extraction layer which have the largest parameter scale in the data transmission direction; the first splicing layer is also connected with the second target multi-scale feature extraction layer, the second splicing layer is also connected with the first target multi-scale feature extraction layer, and the third splicing layer is also connected with the first multi-scale feature extraction layer.

7. The vehicle speed detection method according to any one of claims 1 to 3, characterized by, Based on the vehicle matching image, vehicle tracking data matched with the vehicle to be tested is determined from the video stream image frames of the current field of view, including: obtaining feature extraction parameters of a feature extraction layer of a vehicle detection network model for vehicle recognition on the target reference area image; determining network structures of a matching feature extraction network for extracting vehicle matching image features of the vehicle matching image and a video feature extraction network for extracting image features of each frame of the video stream image frames according to the network structure of the feature extraction layer, the feature extraction parameters of the matching feature extraction network and the video feature extraction network being consistent with the feature extraction parameters; combining the matching feature extraction network, the video feature extraction network and a feature matching network into a vehicle matching network model; the input end of the feature matching network is connected with the outputs of the matching feature extraction network and the video feature extraction network respectively, the vehicle matching image features and the image features of each frame are matched, and the vehicle categories, vehicle positions and corresponding confidence scores of each vehicle contained in the video stream image frames are output; According to an output result of the vehicle matching network model, vehicle tracking data matched with the vehicle to be tested is determined.

8. The vehicle speed detection method according to claim 7, characterized by, The feature matching network comprises a second input layer, a plurality of target matching layers, a detection head and a second output layer; each target matching layer comprises a position feature fusion layer, a plurality of simultaneously operated operation processing layers, a first multi-feature fusion layer, a layer normalization layer, a full connection layer and a second multi-feature fusion layer connected in sequence; The second input layer is connected with the video feature extraction network and the matching feature extraction network respectively, performs flattening operation on received image features and vehicle matching image features respectively, and inputs video image flattened features to the first operation processing layer and matching image flattened features to the position feature fusion layer and the first multi-feature fusion layer; the vehicle matching image features are position coded in a manner that the background block is set to infinity, and the position coded features are input to the position feature fusion layer; Each operation processing layer performs matrix multiplication operation on the position feature fusion data and the video image flattened features, performs scaling operation on the first operation result, processes the scaling result using an activation function, performs matrix multiplication operation on the processing result and the video image flattened features, and inputs the second operation result to the first multi-feature fusion layer, and the output end of the first multi-feature fusion layer is connected to the layer normalization layer and the second multi-feature fusion layer respectively.

9. The vehicle speed detection method according to claim 8, characterized by, The detection head comprises a video frame detection layer and a matching feature processing layer; The video frame detection layer is connected with the second multi-feature fusion layer, comprises a confidence calculation branch, a vehicle category identification branch and a vehicle position calculation branch; the confidence calculation branch, the vehicle category identification branch and the vehicle position calculation branch comprise convolution layers with same convolution kernel size, step value and padding value, and two-dimensional convolution layers with same convolution kernel size, step value and padding value and different channel numbers; The matching feature processing layer receives vehicle matching image features, comprises a sixth convolution layer, a feature addition layer and a feature multiplication layer connected in sequence, the output of the sixth convolution layer is also connected to the feature multiplication layer, the feature addition layer is also connected to the convolution layer output of the vehicle category identification branch, and the output of the feature multiplication layer is connected to the two-dimensional convolution layer of the vehicle category identification branch; The vehicle category identification branch further comprises a category determination layer connected after the two-dimensional convolution layer of the vehicle category identification branch, and when receiving the vehicle category identification result of the target vehicle corresponding to the current image frame, calculates the distance between the target vehicle and other vehicle categories, and when the distance change between the target vehicle and other vehicle categories meets a preset distance change condition, outputs the vehicle category identification result of the target vehicle.

10. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the vehicle speed detection method according to any one of claims 1 to 9. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the vehicle speed detection method according to any one of claims 1 to 9. ​ 11. A computer readable storage medium, characterized in that, ​ 12. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the vehicle speed detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and device for determining vehicle running speed and electronic equipment

    CN116311965A

  • Expressway interval speed calculation method based on vehicle feature matching

    CN118486175A