Vehicle speed detection method, electronic equipment, readable storage medium and program product

By collecting image data at a reference position for vehicle identification and trajectory tracking, combined with an edge server and a multi-camera system, the problem of insufficient accuracy of vehicle speed detection in complex environments in existing technologies is solved, and efficient and accurate vehicle speed detection is achieved.

CN120668955AActive Publication Date: 2025-09-19LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 15 Cites 0 Cited by

Patent Information

Application Number
CN202511188324.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-19
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing vehicle speed detection methods lack accuracy in complex environments and cannot meet high-precision requirements, especially in scenarios with complex road networks and frequent changes in vehicle speeds in closed areas. Existing technologies lack environmental robustness and deployment flexibility.

Method used

By collecting image data at a reference location for vehicle identification, determining the vehicle matching image, and based on the vehicle's driving trajectory and speed measurement area path planning data, combined with edge servers and multi-camera systems, the vehicle's driving speed in the speed measurement area can be tracked in real time, reducing computing resource requirements and improving detection efficiency and accuracy.

Benefits of technology

It achieves efficient and accurate vehicle speed detection in complex environments, reduces computing resource consumption, is suitable for edge device deployment, and improves the accuracy and efficiency of vehicle speed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668955A_ABST
    Figure CN120668955A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle speed detection method, electronic equipment, a readable storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: determining a vehicle matching image for follow-up vehicle tracking according to a vehicle identification result of a target reference position acquisition image of a speed measurement area; and determining vehicle tracking data matched with the vehicle to be subjected to speed measurement from the video stream image frame of the current view. After the vehicle leaves the current view, determining the view for extracting the next image data, and determining the vehicle tracking data of the new view according to the latest vehicle matching image; and determining the average driving speed of the vehicle in the whole speed measurement area and the visual field driving speed under each visual field according to all the vehicle tracking data, and determining the driving speed of the vehicle to be subjected to speed measurement according to the speed data. The problems of vehicle tracking failure and inaccurate speed measurement can be solved, and the vehicle speed detection accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a vehicle speed detection method, electronic equipment, computer-readable storage medium, and computer program product. Background Art

[0002] The accuracy of vehicle speed detection increases with the safety requirements of vehicle driving scenarios. During the vehicle speed measurement process, related technologies may have problems such as vehicle tracking failure and inaccurate speed measurement, which cannot meet users' high-precision vehicle speed detection needs. Summary of the Invention

[0003] The present invention provides a vehicle speed detection method, electronic equipment, computer-readable storage medium and computer program product, which effectively improve the accuracy of vehicle speed detection.

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: The present invention provides a vehicle speed detection method, comprising: Based on the vehicle recognition result of the target reference area image corresponding to the target reference position of the speed measurement area, a vehicle matching image containing the vehicle to be measured is determined; based on the vehicle matching image, vehicle tracking data matching the vehicle to be measured is determined from the video stream image frame of the current field of view; when it is determined according to the vehicle tracking data that the vehicle has left the current field of view, the target field of view corresponding to the current driving position of the vehicle is determined according to the speed measurement area path planning data, and new vehicle tracking data is again determined in the target field of view according to the current vehicle matching image; based on the vehicle tracking data of each field of view through which the vehicle to be measured passes during its driving process, the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of a single field of view are determined, and the driving speed of the vehicle to be measured is determined based on the average driving speed and the driving speed of each field of view.

[0005] The present invention also provides an electronic device, comprising a memory and a processor, wherein the processor is configured to implement the steps of any of the above-mentioned vehicle speed detection methods when executing a computer program stored in the memory.

[0006] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of any of the above-mentioned vehicle speed detection methods are implemented.

[0007] Finally, the present invention further provides a computer program product, comprising a computer program / instruction, which implements the steps of any of the above vehicle speed detection methods when executed by a processor.

[0008] The advantage of the technical solution provided by the present invention is that, by using the vehicle recognition result of the image data collected at the reference position as the vehicle matching image for subsequent tracking of the vehicle to be measured, and determining the image acquisition devices passed by the vehicle to be measured during driving in the speed measurement area according to the vehicle driving trajectory and the speed measurement area path planning data, there is no need to perform vehicle detection on all image acquisition devices in the speed measurement area. This not only improves the vehicle tracking efficiency and thus improves the speed measurement detection efficiency, but also effectively saves computing resources, which is conducive to deployment on edge devices. The final driving speed of the vehicle is determined according to the driving speed of the vehicle in each single camera area and the entire speed measurement area, effectively reducing the speed calculation error and improving the vehicle speed detection accuracy.

[0009] In addition, the present invention also provides corresponding electronic equipment, computer-readable storage medium and computer program product for implementing the vehicle speed detection method, further making the method more practical. The electronic equipment, computer-readable storage medium and computer program product have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 A schematic diagram of the hardware composition framework applicable to the vehicle speed detection method provided by the present invention; Figure 2 A schematic flow chart of a vehicle speed detection method provided by the present invention; Figure 3 A schematic diagram of the network model structure of the vehicle detection network model provided by the present invention in an exemplary application scenario; Figure 4 A schematic diagram of the network model structure of the initial feature extraction layer provided by the present invention in an exemplary application scenario; Figure 5 A schematic diagram of the network model structure of the multi-scale feature extraction layer provided by the present invention in an exemplary application scenario; Figure 6 A schematic diagram of the network model structure of the feature extraction block provided by the present invention in an exemplary application scenario; Figure 7 A schematic diagram of the network model structure of the first output layer provided by the present invention in an exemplary application scenario; Figure 8 A schematic diagram of the network model structure of the vehicle matching network model provided by the present invention in an exemplary application scenario; Figure 9 A schematic diagram of the network model structure of the detection head provided by the present invention in an exemplary application scenario; Figure 10 A schematic flow chart of another vehicle speed detection method provided by the present invention; Figure 11 A structural framework diagram of an exemplary embodiment of a vehicle speed detection device provided by the present invention; Figure 12 This is a schematic structural diagram of an exemplary embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. The terms "first," "second," "third," "fourth," etc. in the specification and the accompanying drawings are used to distinguish different objects rather than to describe a specific order. Furthermore, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. The term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0013] Vehicle speed detection accuracy increases with the safety requirements of driving scenarios. For enclosed areas with entrances and exits, such as logistics parks, campuses, and scenic areas, strict maximum vehicle speed control is necessary to prevent traffic accidents caused by speeding. Current vehicle speed detection methods include non-visual speed detection methods (such as radar speed measurement and ground sensor coil detection), camera-based speed measurement technologies (such as traditional video analysis and computer vision), and a combination of the two, such as vision and radar. However, these methods are costly and lack universal applicability.

[0014] Radar speed measurement technology first emits electromagnetic waves, then receives the reflected signal from the vehicle, and finally calculates vehicle speed based on the Doppler effect. This method can be linked with cameras (such as integrated radar cameras) to capture speeding violations. Its speed measurement accuracy reaches ±0.5 km / h and supports multi-lane coverage within a 100-meter range. However, metal interference can easily lead to misjudgments, and accuracy decreases in rainy and foggy weather. Furthermore, this method only outputs speed data and cannot identify vehicle characteristics (such as license plate and vehicle model). Ground-sensing coil detection technology embeds induction coils in the road surface. When a vehicle passes, the speed v is calculated based on the changes in the electromagnetic field. For example, speed can be calculated using v = s / Δt, where s is the coil spacing and Δt represents the transit time. This method is low-cost and highly reliable, but it requires disrupting road construction and requires re-embedding coils as the road surface changes. This leads to high maintenance costs and poor scalability. It cannot identify vehicle characteristics and is prone to speed errors when multiple vehicles are intersecting. Furthermore, infrared speed measurement utilizes infrared radiation from vehicles for monitoring, but its effective range is short (less than 50 meters) and is significantly affected by weather. Ultrasonic velocity measurement has strong resistance to rain and fog, but it has a large beam divergence angle, low resolution and significant error.

[0015] Traditional video analysis methods set virtual detection lines within video frames, triggering vehicle position analysis based on grayscale changes, and calculating speed based on the frame interval. This method analyzes the two-dimensional displacement of the vehicle in consecutive video frames and converts it into actual speed based on calibration parameters. However, due to increased image noise caused by lighting variations (such as at night and in shadows), false detection rates increase. Furthermore, manual calibration of reference objects is required, resulting in poor algorithm robustness and data processing relying on high-performance computing equipment, which lacks real-time performance. Computer vision technology utilizes deep learning algorithms and neural network models to achieve dynamic vehicle tracking. This enables license plate recognition and trajectory prediction, reduces reliance on manual calibration, improves adaptability to complex scenarios, and can simultaneously output structured data such as vehicle model and color. However, this method suffers from poor tracking performance and is prone to tracking failure when vehicles overtake or overlap, leading to speed measurement errors. For speed measurement methods that require the participation of cameras, since they rely on cameras to collect images, there will be problems with poor environmental adaptability: complex road conditions such as narrow roads and curves lead to poor vehicle tracking effects and frequent target occlusion; the difference in lighting between day and night, tree shadows, etc. reduce image quality and affect the overall speed measurement accuracy.

[0016] As can be seen from the above, although the relevant technologies can realize the speed measurement function, there are bottlenecks in environmental robustness, deployment flexibility and data intelligence, especially for scenarios with complex road networks and frequent changes in vehicle speed in closed areas. In view of this, in order to solve the problems existing in the above-mentioned relevant technologies, the present invention provides a low-dependence, high-compatibility and strong adaptability pure visual speed measurement method, which uses the vehicle recognition result of the image data collected at the reference position as the vehicle matching image for subsequent tracking of the vehicle to be measured, determines the various image acquisition devices passed by the vehicle to be tested during the driving process in the speed measurement area according to the vehicle driving trajectory and the speed measurement area path planning data, and determines the final driving speed of the vehicle according to the driving speed of the vehicle in each single camera area and the entire speed measurement area. On the basis of reducing the computing resources used in the speed measurement process, the vehicle speed in the speed measurement area is efficiently and accurately detected.

[0017] In conjunction with the specific application environment architecture or specific hardware architecture that the execution of the vehicle speed detection method relies on, the specific application environment architecture or specific hardware architecture is described here. Figure 1 Some possible application scenarios involved in the technical solution of the present invention are introduced by way of example, which may include the following: An entrance camera is deployed at the entrance of the speed measurement area, and multiple cameras are deployed around the vehicle's driving road according to the path planning of the speed measurement area. Each camera will transmit the image data collected within the field of view to the edge server in real time.

[0018] The edge server pre-deploys any target detection network model capable of identifying vehicles. When receiving image data from the entrance camera at the speed measurement area entrance, the image data is transmitted to the target detection network model to identify whether there is a vehicle. If so, the vehicle area image containing the vehicle block is output and used as the vehicle matching image for tracking the vehicle. Based on the vehicle matching image, the vehicle tracking data matching the vehicle to be measured is determined from the video stream image frame of the current field of view. When it is determined based on the vehicle tracking data that the vehicle has left the current field of view, the target field of view corresponding to the vehicle's current driving position is determined based on the speed measurement area path planning data, and new vehicle tracking data is again determined in the target field of view based on the current vehicle matching image. Based on the vehicle tracking data of each field of view that the vehicle to be measured passes through during its driving process, the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of each field of view are determined, and the driving speed of the vehicle to be measured is determined based on the average driving speed and the driving speed of each field of view.

[0019] It should be noted that the above application scenarios are only shown to facilitate understanding of the ideas and principles of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments. First, please refer to Figure 2 , Figure 2 This is a flow chart of a vehicle speed detection method provided in this embodiment. This embodiment may include the following contents: S201: Determine a vehicle matching image including a vehicle to be speed measured according to a vehicle recognition result of a target reference area image corresponding to a target reference position of the speed measurement area.

[0020] The speed measurement area is an area where the vehicle's speed needs to be detected. The target reference position is a pre-specified position in the speed measurement area, such as the entrance to the speed measurement area or the position of the first camera that a vehicle enters and passes through the speed measurement area. The target reference area image is the image data captured by an image acquisition device, such as a camera, installed at the target reference location. The vehicle recognition result is the result of vehicle recognition performed on the target reference area image. Any method capable of achieving target recognition in the relevant art, such as a target recognition algorithm or a target detection network model, may be used. This does not affect the implementation of the present invention. Whether the currently captured target reference area image contains a vehicle is determined. If not, vehicle recognition is performed on the next target reference area image. If so, vehicle information is extracted from the target reference area image. The vehicle information includes an identifier that uniquely identifies the vehicle, such as a license plate. Of course, a unique ID can also be automatically generated for the vehicle, and the vehicle is subsequently tracked using the ID, i.e., the vehicle needs to be speed-measured. In this step, the vehicle is defined as the vehicle to be speed-measured. To facilitate subsequent image processing, a portion of the image containing the vehicle block can be extracted from the target reference area image as a vehicle matching image. The vehicle matching image serves as a matching template for tracking the vehicle to be speed-measured.

[0021] S202: Based on the vehicle matching image, determine vehicle tracking data that matches the vehicle to be speed measured from the video stream image frame of the current field of view.

[0022] In which, after the previous step determines that there is a vehicle to be measured, and the vehicle matching image in the previous step is an image of the starting position for speed measurement of the vehicle to be measured, the current field of view is the image acquisition area of ​​the first image acquisition device passed by the vehicle to be measured when it enters the speed measurement area and moves forward from the position of S201. The video stream image frame is the real-time video stream captured by the image acquisition device. Vehicle detection is performed on each frame of the video stream image frame to identify each image frame containing the vehicle to be measured in S201 from the real-time video stream. The earliest image frame is the time when the vehicle to be measured enters the current field of view. The position and travel time of the vehicle to be measured in the current field of view can be determined based on each image frame to serve as vehicle tracking data of the vehicle to be measured. That is, the vehicle tracking data at least includes the time when the vehicle to be measured enters the current field of view, the time when it exits the current field of view, the total travel time in the current field of view, and various position information in the current field of view. The travel trajectory of the vehicle to be measured in the current field of view can be determined based on these position points. If there are multiple vehicle matching images in S201, meaning multiple vehicles to be tracked are required, this step involves determining the current field of view for each vehicle matching image, and then performing tracking. If there are at least two vehicles to be tracked within the same field of view, the vehicle matching images of the different vehicles to be tracked are individually identified within the video stream image frames. Once the current vehicle category is determined, the corresponding tracking is performed, and vehicle tracking data for each vehicle is generated.

[0023] S203: When it is determined based on the vehicle tracking data that the vehicle has left the current field of view, the target field of view corresponding to the vehicle's current driving position is determined based on the speed measurement area path planning data, and new vehicle tracking data is again determined in the target field of view based on the current vehicle matching image.

[0024] It can be understood that the vehicle is moving forward, while the field of view of the image acquisition device is fixed. The image acquisition device can only collect image data within the field of view. In order to realize the speed detection of the vehicle to be measured, it is necessary to track the driving trajectory of the vehicle to be measured in the speed measurement area in real time. When it is determined that the vehicle to be measured has left the current field of view based on the vehicle tracking data of the current field of view in the previous step, that is, when the vehicle to be measured cannot be detected in the real-time video stream data of the current field of view, it is considered that the vehicle to be measured has left the current field of view, and it is necessary to predict the next field of view that the vehicle to be measured will enter. This step defines the target field of view as the predicted next field of view, so that the vehicle to be measured can be tracked through the real-time video stream from the next field of view. In this step, the path that the vehicle to be measured can travel is known based on the path planning data of the speed measurement area. Based on the layout and field of view of the image acquisition equipment in the speed measurement area, combined with the vehicle's driving trajectory determined by the vehicle tracking data of the current field of view, the next field of view of the vehicle can be predicted, and the image data collected by the image acquisition equipment corresponding to the next field of view is obtained. By performing vehicle recognition on the real-time video stream data of the target field of view, each image frame containing the vehicle to be measured in S202 is determined. Similarly, among the images containing the vehicle to be measured, the earliest image frame is the time when the vehicle to be measured enters the target field of view, and the last image is the last position of the vehicle to be measured in the target field of view. Similarly, the position point and travel time of the vehicle to be measured in the target field of view can be determined from each image frame, thereby generating vehicle tracking data for the vehicle to be measured in the target field of view. When it is determined that the vehicle to be measured has left the target field of view based on the vehicle tracking data of the target field of view, the next field of view of the target field of view is predicted according to this step method. This process is repeated until the vehicle to be measured leaves the speed measurement area, completing the entire tracking process from the vehicle to be measured entering the speed measurement area to the vehicle to be measured leaving the speed measurement area.

[0025] In this step, during the vehicle tracking process, in order to improve the vehicle matching accuracy, the vehicle matching image of S201 may be replaced with the latest and most accurate image. For example, if the license plate number of the vehicle matching image of S201 is unclear, it may be updated later. Therefore, the current vehicle matching image of this step may be the vehicle matching image of S201 or the most recently updated vehicle matching image.

[0026] S204: Determine the average speed of the vehicle to be measured in the speed measurement area and the speed of each field of view based on the vehicle tracking data of each field of view that the vehicle to be measured passes through during its travel, and determine the travel speed of the vehicle to be measured based on the average speed and the speed of each field of view.

[0027] Since the vehicle tracking data for each field of view can determine the travel time and path of the vehicle to be measured in the corresponding field of view, the actual travel length corresponding to the travel path can be determined by combining the path planning data of the speed measurement area and the layout data of the image acquisition device, and then the average travel speed of the vehicle to be measured in each field of view can be determined, which is also the field of view travel speed of this step. Similarly, the travel path of the vehicle to be measured in the entire speed measurement area can be determined based on the travel trajectory of each field of view passed by the vehicle to be measured during the speed measurement area. The actual total travel length corresponding to the regional travel path can be determined by combining the path planning data of the speed measurement area and the layout data of the image acquisition device. The total travel time of the vehicle to be measured in the speed measurement area can be determined based on the time when the vehicle to be measured enters the speed measurement area, the time when it exits the speed measurement area, and whether it stops. The average travel speed of the vehicle to be measured in the speed measurement area can be determined based on the total travel time and actual total travel length. Whether it stops can be determined based on whether the vehicle tracking data is located at the same location at different times. Finally, the speed of the vehicle to be measured is determined based on the average driving speed and the driving speed of a single field of view combined with the actual scenario. For example, in the speeding detection scenario, as long as one of the speed values ​​of the average driving speed and the field of view speed exceeds the maximum speed limit, it is considered to be speeding.

[0028] In the technical solution provided in this embodiment, the vehicle recognition result of the image data collected at the reference position is used as the vehicle matching image for subsequent tracking of the vehicle to be measured, and the image acquisition devices passed by the vehicle to be measured in the speed measurement area are determined according to the vehicle driving trajectory and the speed measurement area path planning data. There is no need to perform vehicle detection on all image acquisition devices in the speed measurement area. This not only improves the vehicle tracking efficiency and thus improves the speed measurement detection efficiency, but also effectively saves computing resources, which is conducive to deployment on edge devices. The final driving speed of the vehicle is determined according to the driving speed of the vehicle in each single camera area and the entire speed measurement area, effectively reducing the speed calculation error and improving the vehicle speed detection accuracy.

[0029] In the above embodiment, there is no limitation on determining the speed of the vehicle to be measured. The present invention also provides an exemplary implementation method, which may include the following contents: for each field of view that the vehicle to be measured passes through during its travel, the appearance time and departure time of the vehicle to be measured in the current field of view are determined based on the vehicle tracking data of the current field of view. Of course, if the vehicle can stop in the field of view, it is also necessary to detect the stop time of the vehicle in the field of view. The time difference between the appearance time and the departure time of the front field of view minus the stop time is used as the travel time, and the travel speed of the vehicle to be measured in the current field of view is determined based on the length of the path of the current field of view and the ratio of length to time. For example, the travel speed of the i-th field of view is according to Calculation, li represents the road path length of the i-th field of view, ti represents the time difference between the vehicle to be measured entering the i-th field of view and leaving the i-th field of view. When the vehicle to be measured leaves the speed measurement area based on the vehicle tracking data of the last field of view passed by the vehicle to be measured, the total driving time of the vehicle to be measured in the speed measurement area is determined based on the entry time of the first field of view and the exit time of the last field of view passed by the vehicle to be measured. For vehicles that are allowed to stop in the speed measurement area, the stop time of the vehicle in the speed measurement area needs to be detected, and the total driving time is calculated as the time difference between the entry time of the first field of view and the exit time of the last field of view. The total length of travel of the vehicle to be measured in the speed measurement area is determined based on the path planning data of the speed measurement area; the average driving speed of the vehicle to be measured in the speed measurement area is determined based on the total driving length and the total driving time. If the average driving speed according to calculate, It represents the length of the moving path of the vehicle to be measured in the speed measurement area, and t represents the total driving time of the vehicle to be measured in the speed measurement area.

[0030] After the average driving speed and the driving speed of each field of view are determined, the average speed in the field of view can be determined according to the driving speed of each field of view and the total number of fields of view, and the speed difference between the average driving speed and the average speed in the field of view can be calculated. To calculate, n represents the total number of fields of view, when the speed difference between the average driving speed and the average speed in the field of view meets the preset speed threshold judgment condition, the preset speed threshold judgment condition is a pre-set condition. As a simple implementation method, a threshold can be set , when the speed difference is less than the threshold, it is considered that the preset speed threshold judgment condition is met, that is, , the average speed of the vehicle being measured and the speeds in each field of view meet the speed measurement accuracy conditions; if at least one of the average speed and the speeds in each field of view exceeds the preset speed limit (maximum speed), the vehicle being measured is speeding in the speed measurement area. Of course, if at least one speed is less than the preset speed limit (minimum speed), the vehicle being measured is underspeeding in the speed measurement area. To further improve the efficiency of speed detection, an alert message can be generated when speeding or underspeeding is detected. The alert message can include an image of the vehicle and the vehicle's movement trajectory and time during the speeding phase.

[0031] The above embodiment does not impose any limitation on how to generate a vehicle matching image. Based on the above embodiment, the present invention further provides an exemplary implementation method, which may include the following contents: In this embodiment, the target reference position is the entrance position, the target reference area image is the entrance area image data, and the entrance area image data of the entrance position of the speed measurement area is obtained; if the entrance area image data originate from the same image acquisition source, or there is no vehicle located in the overlapping field of view of different image acquisition sources in each entrance area image data, then the target area image containing the vehicle block is selected from each entrance area image data as the vehicle matching image of the vehicle to be measured; if the entrance area image data originate from different image acquisition sources, and at least one vehicle to be measured in the entrance area image data is located in the overlapping field of view of at least two image acquisition sources, then the overlapping area image obtained by image acquisition of the overlapping field of view from multiple fields of view is obtained; if there are at least two target overlapping area images whose vehicle blocks are not the same vehicle, then the target area image containing the vehicle block is selected from each target overlapping area image as the vehicle matching image corresponding to each vehicle to be measured.

[0032] Among them, the image acquisition source refers to whether the image data of the entrance area is collected by one image acquisition device or multiple image acquisition devices, that is, it is necessary to identify whether there are multiple cameras at the entrance position. In order to avoid the same vehicle being recorded by multiple cameras, resulting in being considered as multiple vehicles, or the same vehicle being tracked multiple times, resulting in a waste of resources, this embodiment will determine the final vehicle matching image according to the conditions of different image acquisition devices.

[0033] Exemplarily, when there are multiple cameras and it is necessary to identify whether the vehicles are the same vehicle, the present invention also provides a simple method for identifying the same vehicle: obtain the first fixed pixel position of the fixed reference object in the first overlapping area image, obtain the first vehicle pixel position corresponding to the first vehicle contained in the first overlapping area, and determine the first pixel distance between the first vehicle and the fixed reference object based on the first vehicle pixel position and the first fixed pixel position; obtain the second fixed pixel position of the fixed reference object in the second overlapping area image, obtain the second vehicle pixel position corresponding to the second vehicle contained in the second overlapping area, and determine the second pixel distance between the second vehicle and the fixed reference object based on the second vehicle pixel position and the second fixed pixel position; if the first pixel distance and the second pixel distance meet the preset identical similarity conditions, the first vehicle and the second vehicle are the same vehicle.

[0034] As can be seen from the above, this embodiment first detects the entrance vehicle, avoiding the need to detect all cameras in the park and saving computing resources; for the multi-camera situation, the same vehicle is identified as the same vehicle based on the same or similar distances calculated in different field of view areas, effectively improving the vehicle tracking accuracy and thus effectively improving the vehicle speed detection accuracy.

[0035] Furthermore, based on the above embodiment, the present invention also provides the following embodiment: target matching is performed on video stream image frames and vehicle matching images with the same size to obtain a target vehicle with the same vehicle identification information as the vehicle to be measured, a target vehicle area of ​​the target vehicle in the corresponding image frame, and a corresponding score; if the score is less than or equal to a preset score threshold, the vehicle matching image is updated to the target vehicle area; if the score is greater than the preset score threshold, a driving trajectory is generated according to the position information of the vehicle to be measured in the video stream image frame, and when it is determined that the vehicle to be measured has left the current field of view based on the driving trajectory, the field of view driving speed is determined based on the driving time and driving path of the vehicle to be measured in the current field of view.

[0036] In this embodiment, to improve matching accuracy, the size of the vehicle matching image and the video stream image frame can be kept consistent. For example, the vehicle matching image can be padded to ensure the same size between the two. For example, using pixel values ​​of [114, 114, 114], the edges of the vehicle matching image can be padded without changing the size of the vehicle matching image to ensure the image size of the vehicle matching image and the video stream image frame is consistent. If there are multiple vehicles of the same vehicle category, that is, multiple vehicles to be speed tested, the average score of each vehicle to be speed tested can be used as the final score. The preset score threshold is a pre-set value. In addition, a threshold for the number of vehicle matching images of the same vehicle can be set. If the number of vehicle matching images for a vehicle exceeds the set threshold, the vehicle matching images are updated in a rolling manner, and the oldest vehicle matching image is discarded.

[0037] As can be seen from the above, this embodiment expands the vehicle matching image to the same size as the video stream image frame to be detected, and does not change the size of the vehicle matching image, which is more conducive to feature extraction by the original feature extraction network and retains the valid value of the feature part; multiple vehicle matching images are used for matching, and the vehicle matching images are rolled updated, so that the latest template is more conducive to vehicle matching and the vehicle matching accuracy is improved.

[0038] The above embodiment does not limit vehicle identification in any way. Based on the above embodiment, the present invention further provides an exemplary implementation of vehicle identification, which may include the following: The target reference area image is input into a pre-trained vehicle detection network model, and the position of the vehicle block contained in the target reference area image is determined according to the output result of the vehicle detection network model.

[0039] In this embodiment, the network structure of the vehicle detection network model is as follows: Figure 3As shown, the vehicle detection network model includes at least a first input layer, a feature extraction layer, a feature fusion layer and a first output layer; the feature extraction layer includes multiple feature extraction sublayers; the feature extraction layer extracts image features from the target reference area image transmitted through the first input layer, and inputs the generated feature map into the feature fusion layer; the feature fusion layer samples multiple times and fuses the feature submaps of at least one target feature extraction sublayer, and outputs the position of the vehicle image block through the first output layer.

[0040] After the network structure of the vehicle network model is determined, any training sample set that identifies the vehicle location and vehicle category can be used to train the vehicle network model through any model training method until the vehicle network model converges, reaches a preset number of iterations, or reaches a preset accuracy threshold. The total loss function loss of the vehicle network model during training may include vehicle location loss lossp, classification loss lossc, and confidence loss losss. The calculation relationship of each loss is as follows: ; ; ; .

[0041] Where i represents the i-th training sample in the training sample set, j represents the j-th vehicle category, C represents the total number of categories, and n represents the total number of training samples contained in the training sample set. represents the Euclidean distance, c represents the normalization constant, is the predicted target box, is the real target frame; is the true classification of the i-th training sample belonging to the j-th vehicle category, is the predicted classification of the i-th training sample for the j-th vehicle category, Whether the i-th training sample is the j-th vehicle category exists, which is 0 or 1. The probability that the i-th training sample is predicted for the j-th vehicle category, Represents the respective weight coefficients, which can be empirical values ​​or determined during the training process.

[0042] For example, Figure 3 As shown, the feature extraction layer is Figure 3 The left part of the structure may include an initial feature extraction layer and multiple groups of semantic feature extraction layers with the same structure; the initial feature extraction layer is connected to the first input layer, and its output is connected to the input of the first semantic feature extraction layer, such as Figure 4As shown in , the initial feature extraction layer may include at least a two-dimensional convolution layer, a batch normalization layer, and an activation function layer. For example, the SiLU (function name) activation function may be used. The output of each semantic feature extraction layer is connected to the input of the next semantic feature extraction layer, and the last semantic feature extraction layer is connected to the first output layer. Each semantic feature extraction layer includes a first convolution layer and a multi-scale feature extraction layer with the same convolution kernel size, step size, and padding parameters and an increasing number of channels. Figure 5 As shown in the figure, the multi-scale feature extraction layer includes at least the second convolution layer, the feature splitting layer, the local feature extraction layer, the semantic fusion layer and the third convolution layer. The network parameters of each layer of the vehicle detection network model can be as follows: Figure 3-Figure 6 As shown. Among them, the first convolution layer extracts the features of the received feature map and inputs the extracted new feature map to the second convolution layer. The second convolution layer is connected to the corresponding first convolution layer, extracts features from the feature map output by the first convolution layer, and outputs the new feature map to the feature splitting layer. The feature splitting layer splits the received feature map into a first sub-feature map and a second sub-feature map, inputs the first sub-feature map to the local feature extraction layer, and inputs the second sub-feature map to the semantic fusion layer; the local feature extraction layer includes multiple feature extraction blocks with the same structure and connected in series. The total number of feature extraction blocks is determined according to the number of input channel dimensions, as shown in Figure 6 As shown in the figure, the feature extraction block includes two convolutional layers with the same convolution kernel size and different strides; the output of the first feature extraction block of the local feature extraction layer is also connected to the semantic fusion layer, and the semantic fusion layer is connected to the last feature extraction block, and outputs the fusion result of the received feature map to the third convolutional layer.

[0043] As can be seen above, the semantic feature extraction layer uses the feature splitting layer to first split the channels into two and then merge them, which can preserve both the original features and the features after passing through the feature extraction block. The feature extraction block is proportional to the number of input channels. The more channels, the more stacked feature extraction blocks, thereby enhancing feature extraction capabilities. When the number of input channels is small, first increasing the dimension and then reducing the dimension improves nonlinear expression capabilities, reduces the number of parameters, and improves computational efficiency. It is suitable for lightweight models and is more suitable for deployment on resource-limited edge devices. The multi-scale feature extraction layer does not change the feature shape of the input and output, facilitating model construction. In addition, it can better extract multi-scale features, improve vehicle detection capabilities, and significantly reduce the number of model parameters and computation while maintaining detection accuracy, achieving lightweight and real-time target detection tasks.

[0044] For example, the feature fusion layer is Figure 3The structure of the right half of the network comprises at least a first upsampling layer, a first splicing layer, a first multi-scale feature extraction layer, a second upsampling layer, a second splicing layer, a second multi-scale feature extraction layer, a fourth convolutional layer, a third splicing layer, a third multi-scale feature extraction layer, a fifth convolutional layer, and a fourth multi-scale feature extraction layer, which are connected in sequence. The network results of each multi-scale feature extraction layer are the same as those of the multi-scale feature extraction layer of the above embodiment, and the network parameters are shown in the corresponding figure. The first upsampling layer receives the output feature map of the feature extraction layer, and the feature extraction layer includes a first target multi-scale feature extraction layer and a second target multi-scale feature extraction layer with the largest parameter scale according to the data transmission direction. The first splicing layer is also connected to the second target multi-scale feature extraction layer, the second splicing layer is also connected to the first target multi-scale feature extraction layer, and the third splicing layer is also connected to the first multi-scale feature extraction layer.

[0045] For example, Figure 7 As shown, the first output layer may include a confidence score calculation branch, a vehicle type calculation branch, and a location box calculation branch. The confidence score calculation branch, vehicle type calculation branch, and location box calculation branch all include a deep classifiable convolution layer with the same convolution kernel size, stride, and padding parameters, and a two-dimensional convolution layer with the same convolution kernel size, stride, and padding parameters. The deep classifiable convolution layer performs independent convolution operations on each channel of the feature output from the feature fusion layer, enabling parallel computation between channels. The confidence score calculation branch processes the output features of the feature fusion layer through the deep classifiable convolution layer and the two-dimensional convolution layer, outputting the confidence score corresponding to the vehicle included in the target reference area image. The vehicle type calculation branch processes the output features of the feature fusion layer through the deep classifiable convolution layer and the two-dimensional convolution layer, outputting the vehicle category corresponding to the vehicle included in the target reference area image, such as the vehicle ID. The location box calculation branch processes the output features of the feature fusion layer through the deep classifiable convolution layer and the two-dimensional convolution layer, outputting the vehicle location information in the target reference area image.

[0046] This embodiment calculates the confidence, category, and position of the image features of the target reference area image through a depthwise separable convolution layer and a two-dimensional convolution layer. Through channel-by-channel convolution, while maintaining the receptive field, the amount of calculation and the number of parameters are significantly reduced. It is particularly suitable for resource-constrained devices such as mobile terminals and edge servers. In addition, each convolution kernel only acts on the data of a single channel, avoiding the complexity of inter-channel mixing operations in traditional convolution.

[0047] In order to make those skilled in the art more clear about the processing process of the vehicle detection network model provided by the present invention on the input target reference area image, the present invention also uses Figure 3 The structure of the vehicle detection network model shown in the figure illustrates the image processing process of the vehicle detection network model, which may include the following: A1: Input vehicle target reference area image.

[0048] A2: The target reference area image enters the Conv layer and is processed by Conv2d (two-dimensional convolution layer), BN (Batch Normalization) and SiLU to obtain a new feature map.

[0049] The size of the convolution kernel K, stride s, padding p, and number of channels c of Conv2d can be 6, 2, 2, and 64, respectively. s represents the distance the convolution kernel slides on the input image, p is used to add extra pixels to the edge of the original data to keep the output dimension unchanged or expand the receptive field, and c represents the dimension of the feature map after convolution.

[0050] A3: The feature map output in step A2 passes through the first Conv layer to obtain a new feature map.

[0051] A4: The feature map output by step A3 passes through the first multi-scale feature extraction layer to obtain a new feature map.

[0052] In this step, the number of feature extraction blocks in the multi-scale feature extraction layer is 3×d (input channel dimension).

[0053] A5: The feature map output by step A4 passes through the next first Conv layer to obtain a new feature map.

[0054] A6: The feature map output by step A5 passes through the second multi-scale feature extraction layer to obtain a new feature map, denoted as F1.

[0055] Among them, the number of feature extraction blocks of the second multi-scale feature extraction layer is 6×d, which is also connected to the second splicing layer of the feature fusion layer.

[0056] A7: F1 passes through the next first Conv layer to obtain a new feature map.

[0057] A8: The feature map output by step A7 passes through the third multi-scale feature extraction layer to obtain a new feature map, denoted as F2.

[0058] The number of feature extraction blocks in the third multi-scale feature extraction layer is 6×d, which is also connected to the first concatenation layer of the feature fusion layer.

[0059] A9: F2 passes through the next first Conv layer to obtain a new feature map.

[0060] A10: The feature map output by step A9 passes through the fourth multi-scale feature extraction layer to obtain a new feature map.

[0061] The steps before A10 complete feature extraction, and the feature maps output by them are input into the feature fusion layer.

[0062] A11: The feature map output by step A10 is subjected to the first upsampling layer of the feature fusion layer to obtain a new feature map, denoted as F3.

[0063] A12: Merge the features of F3 and F2 through the first concatenation layer.

[0064] The feature fusion in this step can prevent the loss of feature information and improve the feature extraction and detection capabilities of the model.

[0065] A13: The merged features are passed through a multi-scale feature extraction layer to obtain a new feature map, denoted as F4.

[0066] A14: Upsample F4 through the second upsampling layer to obtain a new feature map.

[0067] A15: Merge the new feature map generated in step A14 and F1 through the second splicing layer to prevent feature information loss and improve the model's feature extraction and detection capabilities.

[0068] A16: Pass the second merged feature through the multi-scale feature extraction layer to obtain a new feature map.

[0069] A17: Pass the feature map generated in step A16 through the fourth Conv layer to obtain a new feature map.

[0070] A18: Merges the new feature map generated by A17 and F4 through the third splicing layer to prevent feature information loss and improve the model's feature extraction and detection capabilities.

[0071] A19: Pass the third merged feature through the multi-scale feature extraction layer again to obtain a new feature map.

[0072] A20: Pass the new feature map generated in step A19 through the fifth Conv layer to obtain a new feature map.

[0073] A21: The new feature map generated in step A20 is passed through the multi-scale feature extraction layer to obtain a new feature map.

[0074] A22: Pass the new feature map generated in step A21 through the first output layer to obtain the vehicle's location, classification, and confidence.

[0075] The above embodiment does not limit the vehicle matching process. Based on the above embodiment, the present invention further provides an exemplary implementation of vehicle matching, which may include the following: In this embodiment, the vehicle matching process between the vehicle matching image and the video stream image frames of each field of view is completed through the vehicle matching network model. The vehicle matching network model can be shown in Figure 8, which includes three network structures: a matching feature extraction network, a video feature extraction network, and a feature matching network.

[0076] In this embodiment, feature extraction parameters of the feature extraction layer of a vehicle detection network model for vehicle identification in a target reference area image are obtained. Based on the network structure of the feature extraction layer, the network structures of a matching feature extraction network for extracting vehicle matching image features from a vehicle matching image and a video feature extraction network for extracting image features from each frame of a video stream image frame are determined. The feature extraction parameters of the matching feature extraction network and the video feature extraction network are consistent with the feature extraction parameters, thereby reducing the number of parameters in the overall model. In other words, the network model structure and model parameters of the vehicle matching network model for extracting image features from the vehicle matching image and the video stream image frame are identical to those of the vehicle detection network model. The input of the feature matching network is connected to the outputs of the matching feature extraction network and the video feature extraction network, respectively, to receive vehicle matching image features and image features from each frame, perform matching on the vehicle matching image features and image features from each frame, and output, via a second output layer, the vehicle category, vehicle location, and corresponding confidence score for each vehicle contained in the video stream image frame. Consequently, vehicle tracking data matching the vehicle to be speed-measured can be determined based on the output of the vehicle matching network model.

[0077] Exemplarily, the feature matching network includes a second input layer, multiple target matching layers, a detection head and a second output layer; each target matching layer includes a position feature fusion layer, multiple simultaneously running operation processing layers, a first multi-feature fusion layer connected in sequence, a layer normalization layer, a fully connected layer and a second multi-feature fusion layer; the second input layer is respectively connected to the video feature extraction network and the matching feature extraction network, and flattens the received image features of each frame and the vehicle matching image features, and inputs the flattened features of the video image into the first operation processing layer, and inputs the flattened features of the matching image into the position feature fusion layer and the first multi-feature fusion layer; the vehicle matching image features are position-encoded in a manner that the background block is set to infinity, and the position-encoded features are input into the position feature fusion layer. Exemplarily, the second input layer includes a first flattening layer, a second flattening layer, and a position encoding layer. The first flattening layer is connected to the output of the video feature extraction network, flattens the features of each received frame image, and inputs the flattened features of the video image into the first operation processing layer. The second flattening layer is connected to the output of the matching feature extraction network, flattens the vehicle matching image features, and inputs the flattened features of the matching image into the position feature fusion layer and the first multi-feature fusion layer. The position encoding layer position-encodes the vehicle matching image features by setting the background block to infinity, and inputs the position-encoded features into the position feature fusion layer. Each operation processing layer performs a matrix multiplication operation on the position feature fusion data and the flattened features of the video image, scales the first operation result, processes the scaled result using an activation function such as softmax, performs a matrix multiplication operation on the processing result and the flattened features of the video image, and inputs the second operation result into the first multi-feature fusion layer. The output end of the first multi-feature fusion layer is connected to the layer normalization layer and the second multi-feature fusion layer, respectively.

[0078] In this embodiment, the position encoding layer sets the target-free area to negative infinity to avoid interference with vehicle matching and improve vehicle detection accuracy. The computation processing layer can filter out invalid areas in the flattened features of the vehicle matching image through feature addition, which not only reduces the amount of matching data but also avoids interference with vehicle matching and improves vehicle detection accuracy. Scaling can prevent the dot product value from being too large. It can also prevent the dot product value from becoming very large as the vector dimension increases, resulting in the gradient of the softmax function being very small, which in turn causes the gradient to disappear and affects the training effect of the model. The scaling operation makes the dot product result more stable, ensuring that the softmax function operates within a more appropriate range, thereby ensuring the training stability of the model.

[0079] In order to make those skilled in the art more clear about the vehicle matching process of the vehicle matching network model provided by the present invention, the present invention also uses Figure 8The structure of the vehicle matching network model shown in the figure illustrates the vehicle matching process of the vehicle matching network model, which may include the following contents: B1: Input video stream image frame or vehicle matching picture.

[0080] The feature extraction process corresponding to the video stream image frame or vehicle matching image may include: passing the video stream image frame or vehicle matching image through the Conv layer to obtain a new feature map, represented as FB1, passing FB1 through the Conv layer to obtain a new feature map, represented as FB2, and then passing FB2 through the multi-scale feature extraction layer to obtain a new feature map, represented as FB3. Then, passing FB3 through the Conv layer to obtain a new feature map, represented as FB4, passing FB4 through the multi-scale feature extraction layer to obtain a new feature map, represented as FB5, passing FB5 through the Conv layer to obtain a new feature map, represented as FB6, passing FB6 through the multi-scale feature extraction layer to obtain a new feature map, represented as FB7; passing FB7 through the Conv layer to obtain a new feature map, represented as FB8; and passing FB8 through the multi-scale feature extraction layer to obtain a new feature map, represented as FB9.

[0081] B2: Flatten the video stream image frame to obtain feature F11, and flatten the vehicle matching image to obtain feature F22.

[0082] B3: Matching process of features F11 and F22: B3.1: F22 performs position encoding based on the position of the vehicle matching image and sets the target-free area to negative infinity.

[0083] B3.2: Add the position code and F22 features at the corresponding positions to filter out invalid areas in F22.

[0084] B3.3: Perform a matrix multiplication on feature F11 and the filtered feature F22 to obtain a new feature F33. Scale the new feature F33 by dividing it by a specific value and perform a softmax on the scaled result to obtain a new feature F44. Matrix multiply the new feature F44 by F11 to obtain a new feature F55.

[0085] Among them, this step is run simultaneously by M modules, and finally the obtained features are merged together to obtain a new feature F66.

[0086] B3.4: Add the new feature F66 and feature F22 to obtain the new feature F77.

[0087] By adding features in this step and the following steps, the vehicle features can be highlighted, preventing the network from being too deep and unable to pass forward.

[0088] B3.5: Perform layer normalization on the new feature F77 to obtain the new feature F88.

[0089] B3.6: Pass the new feature F88 through the fully connected layer to obtain the new feature F99.

[0090] B3.7: Add the new feature F99 and feature F77 to obtain the new feature F0.

[0091] Among them, there are N target matching layers, which are connected in series. The value of N can be dynamically set according to the matching effect.

[0092] As can be seen from the above, the vehicle matching network model and the vehicle detection network model of the present invention share feature extraction model parameters, reducing the total number of model parameters, improving vehicle detection efficiency, and making it more convenient for deployment on edge devices. Furthermore, during the target search phase, a position encoding method is used to set irrelevant areas of the vehicle matching image to negative infinity, avoiding interference from irrelevant parts, improving the search efficiency of the vehicle matching network model, and thus improving the efficiency of vehicle tracking and vehicle speed measurement.

[0093] For example, Figure 9 As shown, the detection head may include a video frame detection layer and a matching feature processing layer; wherein the video frame detection layer is connected to the second multi-feature fusion layer, including a confidence calculation branch, a vehicle category recognition branch and a vehicle position calculation branch; the confidence calculation branch, the vehicle category recognition branch and the vehicle position calculation branch include convolution layers with the same convolution kernel size, step value and padding value, and two-dimensional convolution layers with the same convolution kernel size, step value and padding value but different number of channels; the matching feature processing layer receives the vehicle matching image features, including a sixth convolution layer, a feature addition layer, a feature multiplication layer, and a sixth convolution layer connected in sequence. The output of the layer is also connected to the feature multiplication layer, the feature addition layer is also connected to the convolution layer output of the vehicle category identification branch, and the output of the feature multiplication layer is connected to the two-dimensional convolution layer of the vehicle category identification branch; the vehicle category identification branch also includes a category determination layer. After the category determination layer is connected to the two-dimensional convolution layer of the vehicle category identification branch, when the vehicle category identification result of the target vehicle corresponding to the current image frame is received, the distance between the target vehicle and other vehicle categories is calculated, and when the distance change between the target vehicle and other vehicle categories meets the preset distance change condition, the vehicle category identification result of the target vehicle is output.

[0094] In this embodiment, features obtained through the feature matching process, such as feature F0 in the above embodiment, are input into the detection head. Feature F0 undergoes Conv2d in the confidence calculation branch to obtain a new feature, which is then passed through Conv2d to obtain a confidence score. Feature F0 also undergoes Conv2d in the vehicle position calculation branch to obtain a new feature, which is then passed through Conv2d to obtain a target box within the image containing the vehicle, thereby obtaining its position information. Feature F0 passes through the Conv2d of the vehicle category identification branch to obtain a new feature F31. The vehicle matching image features pass through the Conv2d of the matching feature processing layer to obtain a new feature F41. Features F31 and F41 are added bit by bit to obtain a new feature F42. The new feature F42 is multiplied point by point by feature F41 to obtain a new feature F43. The new feature F43 passes through the Conv2d of the vehicle category identification branch to obtain the vehicle ID. The distance between the vehicle and other categories is calculated. The calculated distance value can be stored in a database, and the change in distance value can be compared to determine whether it meets the set threshold. This can avoid vehicle detection errors and the situation where the distance between the ID and other targets fluctuates. If the specific threshold is met, the predicted vehicle category is output.

[0095] As can be seen from the above, the detection head of this embodiment detects the correlation between the identified vehicle and other categories different from the predicted category, that is, the distance detection with other models, which is conducive to improving the recognition accuracy of vehicle categories and reducing misjudgments.

[0096] Finally, in order to make those skilled in the art more clearly understand the technical solution of the present invention, the present invention also provides another implementation process for detecting the speed of vehicles in the park, which may include the following contents: S1: Obtain the video stream captured by the camera at the entrance of the park and extract each frame image from the video stream.

[0097] S2: If there are multiple cameras at the entrance and their fields of view overlap, execute S3. If there is only one camera at the entrance or their fields of view do not overlap, execute S9.

[0098] S3: Input all images of the park entrance into the vehicle detection network model.

[0099] S4: Check whether there is a vehicle in the image. If not, return to S1. If yes, continue to execute S4.

[0100] S5: Determine whether the vehicle is located in the overlapping area. If yes, execute S5; if not, execute S9.

[0101] S6: Extract images of different fields of view within the area.

[0102] S7: Confirm whether the target in the area is the same vehicle and mark it with an ID.

[0103] The distance between the vehicle and the fixed object in the field of view can be used to determine whether it is the same vehicle: the pixel coordinates of the same fixed object are extracted in different fields of view; the distance between the vehicle in the detection image and the fixed object in the image is calculated; the vehicles are sorted according to the distance from the fixed object in different fields of view; the same target is determined in different fields of view based on the distance sorting, that is, the ID of the same vehicle is determined based on the distance sorting.

[0104] S8: Extract the target area containing the vehicle block under different fields of view, and use the target area as a subsequent matching template, that is, the vehicle matching image.

[0105] The target area under a field of view is used as a matching template, and one target can be matched using multiple templates.

[0106] S9: Extract the target area as a matching template, that is, the vehicle matching image.

[0107] S10: Enter the target matching process: S10.1: Continue to obtain the video stream of the current field of view, and fill the edge positions of the vehicle matching image with gray pixels until it reaches the same size as the video stream image.

[0108] S10.2: Input the vehicle matching image and the real-time video stream image into the vehicle matching network model.

[0109] S10.3: Obtain the matching vehicle ID, location, and confidence score from the real-time video stream image.

[0110] When there are multiple vehicles with the same ID, the average score after matching multiple vehicles is calculated as the final score.

[0111] S10.4: Determine whether the score is greater than the set threshold. If not, execute S10.6; otherwise, execute S10.5.

[0112] S10.5: Extract the newly matched ID area based on the position information output by the vehicle matching network model and use it as the vehicle matching image.

[0113] A template quantity threshold is set. When the number of templates is greater than the set threshold, the templates are rolled over and updated, and the earliest template is discarded.

[0114] S10.6: Draw the trajectories of all IDs.

[0115] S10.7: Determine whether there is a vehicle that has left the current field of view based on the trajectory drawn in the above steps. If vehicle ID1 has left the field of view, execute S10.8. If not, return to execute S10.1.

[0116] S10.8: Calculate the time when ID1 leaves the field of view.

[0117] S10.9: Calculate the distances of all paths within the field of view and calculate the average speed of ID1 within the field of view.

[0118] S11: Extract the video stream within the next field of view based on the path planning within the park and the motion trajectory of ID1.

[0119] S12: According to the above steps S10.1-S11, target matching is performed on the video stream extracted in S11.

[0120] S13: Determine whether ID1 has left the park. If not, return to execute S11. If yes, execute S14.

[0121] S14: Obtain the entry time and exit time of ID1, and calculate the time difference between the two.

[0122] S15: Calculate the length of the movement path of ID1 in the park, and calculate the average vehicle speed based on the time difference in S14.

[0123] S16: Calculate the average speed of ID1 within each camera's field of view and calculate the positive value of the difference between the average speed of ID1 and the average speed of the park in S15. If the positive value of the speed difference is less than a preset threshold, it is considered that the speed calculation of ID1 is correct.

[0124] S17: Determine whether the single-view speed and average speed of ID1 exceed the set threshold value. If yes, generate an alarm message; if not, execute S1.

[0125] S18: The procedure ends.

[0126] It should be noted that there is no strict order in which the steps in the present invention are performed. As long as they conform to a logical order, the steps can be performed simultaneously or in a predetermined order. Figure 2 and Figure 10 This is just a schematic and does not mean that this is the only execution order.

[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0128] The present invention also provides a corresponding device for the vehicle speed detection method, which further makes the method more practical. Among them, the device can be described from the perspective of functional modules and hardware. The vehicle speed detection device provided by the present invention is introduced below. The device is used to implement the vehicle speed detection method provided by the present invention. In this embodiment, the vehicle speed detection device may include or be divided into one or more program modules. The one or more program modules are stored in a storage medium and executed by one or more processors to complete the vehicle speed detection method disclosed in Example 1. The program module referred to in this embodiment refers to a series of computer program instruction segments that can complete specific functions, which is more suitable for describing the execution process of the vehicle speed detection device in the storage medium than the program itself. The following description will specifically introduce the functions of each program module of this embodiment. The vehicle speed detection device described below and the vehicle speed detection method described above can be referenced to each other.

[0129] From the perspective of functional modules, see Figure 11 , Figure 11 This is a structural diagram of a vehicle speed detection device provided in this embodiment under a specific implementation mode. The device may include: The template generation module 111 is configured to determine a vehicle matching image including a vehicle to be speed-measured based on a vehicle recognition result of an image of a target reference area corresponding to a target reference position of the speed-measurement area.

[0130] The vehicle tracking module 112 is used to determine vehicle tracking data that matches the vehicle to be measured based on the vehicle matching image from the video stream image frame of the current field of view; when it is determined based on the vehicle tracking data that the vehicle has left the current field of view, the target field of view corresponding to the vehicle's current driving position is determined based on the speed measurement area path planning data, and new vehicle tracking data is again determined in the target field of view based on the current vehicle matching image.

[0131] The speed measurement module 113 is used to determine the average driving speed of the vehicle to be measured in the speed measurement area and the driving speed of a single field of view based on the vehicle tracking data of each field of view passed by the vehicle to be measured during the driving process, and determine the driving speed of the vehicle to be measured based on the average driving speed and the driving speed of each field of view.

[0132] Exemplarily, in some implementation schemes of this embodiment, the speed measurement module 113 may also be used for: determining the appearance time and departure time of the vehicle to be measured in the current field of view according to the vehicle tracking data of the current field of view for each field of view passed by the vehicle to be measured during its travel, and determining the travel speed of the vehicle to be measured in the current field of view according to the path length of the current field of view; when it is determined that the vehicle to be measured has left the speed measurement area according to the vehicle tracking data of the last field of view passed by the vehicle to be measured during its travel, determining the total travel time of the vehicle to be measured in the speed measurement area according to the entry time of the first field of view passed by the vehicle to be measured and the exit time of the last field of view passed by the vehicle to be measured during its travel, and determining the total travel length of the vehicle to be measured in the speed measurement area according to the path planning data of the speed measurement area; and determining the average travel speed of the vehicle to be measured in the speed measurement area according to the total travel length and the total travel time.

[0133] Exemplarily, in some other implementations of this embodiment, the template generation module 111 may also be used for: the target reference position is the entrance position, the target reference area image is the entrance area image data, and the entrance area image data of the entrance position of the speed measurement area is obtained; if the entrance area image data originate from the same image acquisition source, or there is no vehicle located in the overlapping field of view of different image acquisition sources in each entrance area image data, then the target area image containing the vehicle block is selected from each entrance area image data as the vehicle matching image of the vehicle to be measured; if the entrance area image data originate from different image acquisition sources, and at least one vehicle to be measured in the entrance area image data is located in the overlapping field of view of at least two image acquisition sources, then the overlapping area image obtained by image acquisition of the overlapping field of view from multiple fields of view is obtained; if there are at least two target overlapping area images whose vehicle blocks are not the same vehicle, then the target area image containing the vehicle block is selected from each target overlapping area image as the vehicle matching image corresponding to each vehicle to be measured.

[0134] As an exemplary implementation of the above embodiment, the above template generation module 111 can also be further used to: obtain the first fixed pixel position of the fixed reference object in the first overlapping area image, obtain the first vehicle pixel position corresponding to the first vehicle included in the first overlapping area, and determine the first pixel distance between the first vehicle and the fixed reference object based on the first vehicle pixel position and the first fixed pixel position; obtain the second fixed pixel position of the fixed reference object in the second overlapping area image, obtain the second vehicle pixel position corresponding to the second vehicle included in the second overlapping area, and determine the second pixel distance between the second vehicle and the fixed reference object based on the second vehicle pixel position and the second fixed pixel position; if the first pixel distance and the second pixel distance meet the preset same similarity conditions, the first vehicle and the second vehicle are the same vehicle.

[0135] Illustratively, in some other implementations of this embodiment, the speed measurement module 113 may also be used to: determine the average speed within the field of view based on the driving speed of each field of view and the total number of fields of view; when the speed difference between the average driving speed and the average speed within the field of view meets the preset speed threshold judgment condition, the average driving speed of the vehicle to be measured and the driving speed of each field of view meet the speed measurement accuracy condition; when at least one speed value among the average driving speed and the driving speed of each field of view is greater than the preset speed limit value, the vehicle to be measured is speeding in the speed measurement area.

[0136] Exemplarily, in some other implementations of this embodiment, the above-mentioned vehicle tracking module 112 can also be used to: perform target matching on video stream image frames and vehicle matching images with the same size to obtain a target vehicle with the same vehicle identification information as the vehicle to be measured, a target vehicle area of ​​the target vehicle in the corresponding image frame, and a corresponding score; if the score is less than or equal to a preset score threshold, the vehicle matching image is updated to the target vehicle area; if the score is greater than the preset score threshold, a driving trajectory is generated according to the position information of the vehicle to be measured in the video stream image frame, and when it is determined that the vehicle to be measured has left the current field of view based on the driving trajectory, the field of view driving speed is determined according to the driving time and driving path of the vehicle to be measured in the current field of view.

[0137] Exemplarily, in some other implementations of this embodiment, the above-mentioned vehicle tracking module 112 can also be used to: input the target reference area image into a pre-trained vehicle detection network model, and determine the position of the vehicle image block contained in the target reference area image based on the output result of the vehicle detection network model; wherein the vehicle detection network model includes at least a first input layer, a feature extraction layer, a feature fusion layer and a first output layer; the feature extraction layer includes multiple feature extraction sublayers; the feature extraction layer extracts image features of the target reference area image transmitted through the first input layer, and inputs the generated feature map into the feature fusion layer; the feature fusion layer samples multiple times and fuses the feature submap of at least one target feature extraction sublayer, and outputs the position of the vehicle map block through the first output layer.

[0138] As an exemplary implementation of the above embodiment, the feature extraction layer includes an initial feature extraction layer and multiple groups of semantic feature extraction layers with the same structure; the initial feature extraction layer is connected to the first input layer, and its output is connected to the input of the first semantic feature extraction layer, the output of each semantic feature extraction layer is connected to the input of the next semantic feature extraction layer, and the last semantic feature extraction layer is connected to the first output layer; the initial feature extraction layer includes at least a two-dimensional convolution layer, a batch normalization layer and an activation function layer; each semantic feature extraction layer includes a first convolution layer with the same convolution kernel size, step size and padding parameters and an increasing number of channels, and a multi-scale feature extraction layer; the multi-scale feature extraction layer includes at least a second convolution layer, a feature splitting layer, a local feature extraction layer, a language layer, and a local feature extraction layer. The semantic fusion layer and the third convolutional layer are connected; wherein, the second convolutional layer is connected to the corresponding first convolutional layer, and outputs the extracted feature map to the feature splitting layer, the feature splitting layer splits the received feature map into a first sub-feature map and a second sub-feature map, inputs the first sub-feature map to the local feature extraction layer, and inputs the second sub-feature map to the semantic fusion layer; the local feature extraction layer includes multiple feature extraction blocks with the same structure and connected in series, the total number of feature extraction blocks is determined according to the number of input channel dimensions, and the feature extraction block includes two convolutional layers with the same convolution kernel size and different step sizes; the output of the first feature extraction block of the local feature extraction layer is also connected to the semantic fusion layer, the semantic fusion layer is connected to the last feature extraction block, and outputs the fusion result of the received feature map to the third convolutional layer.

[0139] As an exemplary implementation of the above embodiment, the feature fusion layer includes at least a first upsampling layer, a first splicing layer, a first multi-scale feature extraction layer, a second upsampling layer, a second splicing layer, a second multi-scale feature extraction layer, a fourth convolutional layer, a third splicing layer, a third multi-scale feature extraction layer, a fifth convolutional layer and a fourth multi-scale feature extraction layer, which are connected in sequence; wherein the first upsampling layer receives the output feature map of the feature extraction layer, and the feature extraction layer includes a first target multi-scale feature extraction layer and a second target multi-scale feature extraction layer with the largest parameter scale according to the data transmission direction; the first splicing layer is also connected to the second target multi-scale feature extraction layer, the second splicing layer is also connected to the first target multi-scale feature extraction layer, and the third splicing layer is also connected to the first multi-scale feature extraction layer.

[0140] Exemplarily, in some other implementations of this embodiment, the above-mentioned vehicle tracking module 112 can also be used to: obtain feature extraction parameters of the feature extraction layer of the vehicle detection network model for vehicle identification of the target reference area image; determine the matching feature extraction network for extracting the vehicle matching image features of the vehicle matching image, and the network structure of the video feature extraction network for extracting the image features of each frame of the video stream image frame according to the network structure of the feature extraction layer, and the feature extraction parameters of the matching feature extraction network and the video feature extraction network are consistent with the feature extraction parameters; combine the matching feature extraction network, the video feature extraction network and the feature matching network into a vehicle matching network model; the input end of the feature matching network is respectively connected to the output of the matching feature extraction network and the video feature extraction network, matches the vehicle matching image features and the image features of each frame, and outputs the vehicle category, vehicle position and corresponding confidence score of each vehicle contained in the video stream image frame; and determines the vehicle tracking data matching the vehicle to be measured based on the output result of the vehicle matching network model.

[0141] As an exemplary implementation of the above embodiment, the feature matching network includes a second input layer, multiple target matching layers, a detection head and a second output layer; each target matching layer includes a position feature fusion layer, multiple simultaneously running operation processing layers, a first multi-feature fusion layer connected in sequence, a layer normalization layer, a fully connected layer and a second multi-feature fusion layer; the second input layer is respectively connected to the video feature extraction network and the matching feature extraction network, and flattens the received image features of each frame and the vehicle matching image features, and inputs the flattened features of the video image into the first operation processing layer, and inputs the flattened features of the matching image into the Position feature fusion layer and the first multi-feature fusion layer; according to the background block being set to infinity, the vehicle matching image features are position-encoded, and the position-encoded features are input into the position feature fusion layer; each operation processing layer performs matrix multiplication on the position feature fusion data and the flattened features of the video image, scales the first operation result, and processes the scaling result using an activation function, performs matrix multiplication on the processing result and the flattened features of the video image, and inputs the second operation result into the first multi-feature fusion layer, and the output ends of the first multi-feature fusion layer are respectively connected to the layer normalization layer and the second multi-feature fusion layer.

[0142] As another exemplary implementation of the above embodiment, the detection head includes a video frame detection layer and a matching feature processing layer; wherein the video frame detection layer is connected to the second multi-feature fusion layer, including a confidence calculation branch, a vehicle category recognition branch and a vehicle position calculation branch; the confidence calculation branch, the vehicle category recognition branch and the vehicle position calculation branch include convolution layers with the same convolution kernel size, step value and padding value, and two-dimensional convolution layers with the same convolution kernel size, step value and padding value but different number of channels; the matching feature processing layer receives vehicle matching image features, including a sixth convolution layer, a feature addition layer, a feature addition layer and a feature addition layer connected in sequence. The output of the sixth convolutional layer is also connected to the feature multiplication layer, the feature addition layer is also connected to the convolution layer output of the vehicle category identification branch, and the output of the feature multiplication layer is connected to the two-dimensional convolution layer of the vehicle category identification branch; the vehicle category identification branch also includes a category determination layer. After the category determination layer is connected to the two-dimensional convolution layer of the vehicle category identification branch, when the vehicle category identification result of the target vehicle corresponding to the current image frame is received, the distance between the target vehicle and other vehicle categories is calculated, and when the distance change between the target vehicle and other vehicle categories meets the preset distance change condition, the vehicle category identification result of the target vehicle is output.

[0143] For the description of the features in the embodiment corresponding to the vehicle speed detection device, please refer to the relevant description of the embodiment corresponding to the vehicle speed detection method, and no further details will be given here.

[0144] The vehicle speed detection device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from the perspective of hardware. Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present invention in one implementation. The electronic device includes a memory 121 and a processor 122. The memory 121 stores a computer program, and the processor 122 is configured to run the computer program to perform the steps of any of the above vehicle speed detection method embodiments.

[0145] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned vehicle speed detection method embodiments when running.

[0146] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0147] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above vehicle speed detection method embodiments are implemented.

[0148] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned vehicle speed detection method embodiments.

[0149] The above is a detailed introduction to a vehicle speed detection method, electronic device, computer-readable storage medium and computer program product provided by the present invention. The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. Whether the units and algorithm steps of each example described in each disclosed embodiment are executed in electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A vehicle speed detection method, characterized in that: include: Determining a vehicle matching image containing the vehicle to be speed measured based on a vehicle recognition result of the target reference area image corresponding to the target reference position of the speed measurement area; Based on the vehicle matching image, determining vehicle tracking data matching the vehicle to be speed measured from the video stream image frame of the current field of view; When it is determined based on the vehicle tracking data that the vehicle has left the current field of view, the target field of view corresponding to the vehicle's current driving position is determined based on the speed measurement area path planning data, and new vehicle tracking data is again determined in the target field of view based on the current vehicle matching image; Based on the vehicle tracking data of each field of view passed by the vehicle to be measured during its travel, the average travel speed of the vehicle to be measured in the speed measurement area and the travel speed of a single field of view are determined, and the travel speed of the vehicle to be measured is determined based on the average travel speed and the travel speed of each field of view.

2. The vehicle speed detection method according to claim 1, wherein: Determining the average speed of the vehicle to be measured in the speed measurement area and the speed of each field of view according to the vehicle tracking data of each field of view passed by the vehicle to be measured during the movement of the vehicle includes: For each visual field that the speed-measured vehicle passes through during its travel, determining the appearance time and departure time of the speed-measured vehicle in the current visual field based on the vehicle tracking data of the current visual field, and determining the travel speed of the speed-measured vehicle in the current visual field based on the path length of the current visual field; When it is determined that the vehicle to be measured has left the speed measurement area based on the vehicle tracking data of the last field of view passed by the vehicle to be measured during its travel, the total travel time of the vehicle to be measured in the speed measurement area is determined based on the entry time of the first field of view and the exit time of the last field of view passed by the vehicle to be measured during its travel, and the total travel distance of the vehicle to be measured in the speed measurement area is determined based on the path planning data of the speed measurement area; The average driving speed of the vehicle to be measured in the speed measurement area is determined according to the total driving length and the total driving time.

3. The vehicle speed detection method according to claim 1, wherein: The target reference position is an entrance position, the target reference area image is entrance area image data, and according to the vehicle recognition result of the target reference area image corresponding to the target reference position of the speed measurement area, a vehicle matching image including each vehicle to be speed measured is determined, including: Acquiring entrance area image data of an entrance position of a speed measurement area; If the entrance area image data originate from the same image acquisition source, or if there is no vehicle located in the overlapping field of view of different image acquisition sources in the entrance area image data, then a target area image containing a vehicle image block is selected from the entrance area image data as the vehicle matching image of the vehicle to be speed measured; If the entrance area image data originates from different image acquisition sources, and at least one vehicle to be speed measured in the entrance area image data is located in an overlapping area of ​​fields of view of at least two image acquisition sources, acquiring an overlapping area image obtained by acquiring images of the overlapping area of ​​fields of view from multiple fields of view; If there are at least two target overlapping area images containing vehicle blocks that are not of the same vehicle, a target area image containing a vehicle block is selected from each target overlapping area image as a vehicle matching image corresponding to each vehicle to be speed measured.

4. The vehicle speed detection method according to claim 3, characterized in that: After acquiring the overlapping area image obtained by collecting images of the overlapping area of ​​the fields of view from multiple fields of view, the method further includes: Obtaining a first fixed pixel position of a fixed reference object in the first overlapping region image, obtaining a first vehicle pixel position corresponding to a first vehicle included in the first overlapping region, and determining a first pixel distance between the first vehicle and the fixed reference object based on the first vehicle pixel position and the first fixed pixel position; Obtaining a second fixed pixel position of the fixed reference object in the second overlapping area image, obtaining a second vehicle pixel position corresponding to a second vehicle included in the second overlapping area, and determining a second pixel distance between the second vehicle and the fixed reference object based on the second vehicle pixel position and the second fixed pixel position; If the first pixel distance and the second pixel distance satisfy a preset identical similarity condition, the first vehicle and the second vehicle are the same vehicle.

5. The vehicle speed detection method according to claim 1, characterized in that: Determining the driving speed of the vehicle to be measured according to the average driving speed and the driving speeds of each field of view includes: Based on the driving speed of each field of view and the total number of fields of view, the average speed within the field of view is determined; When the speed difference between the average speed and the average speed in the field of view meets the preset speed threshold judgment condition, the average speed of the vehicle to be measured and the speed of each field of view meet the speed measurement accuracy condition; When at least one of the average driving speed and the driving speeds in each field of view is greater than a preset speed limit, the vehicle to be measured is speeding in the speed measurement area.

6. The vehicle speed detection method according to claim 1, characterized in that: Determining vehicle tracking data matching the vehicle to be speed measured from the video stream image frame of the current field of view based on the vehicle matching image includes: Performing target matching on the video stream image frame and the vehicle matching image with the same size to obtain a target vehicle having the same vehicle identification information as the vehicle to be speed measured, a target vehicle area of ​​the target vehicle in the corresponding image frame, and a corresponding score; If the score is less than or equal to a preset score threshold, updating the vehicle matching image to the target vehicle area; If the score is greater than a preset score threshold, a driving trajectory is generated based on the position information of the vehicle to be measured in the video stream image frame, and when it is determined according to the driving trajectory that the vehicle to be measured has left the current field of view, the field of view driving speed is determined based on the driving time and driving path of the vehicle to be measured in the current field of view.

7. The vehicle speed detection method according to any one of claims 1 to 6, characterized in that: Determining a vehicle matching image containing a vehicle to be speed measured based on a vehicle recognition result of an image of a target reference area corresponding to a target reference position of the speed measurement area includes: Inputting the target reference area image into a pre-trained vehicle detection network model, and determining the position of the vehicle block contained in the target reference area image based on the output result of the vehicle detection network model; Among them, the vehicle detection network model includes at least a first input layer, a feature extraction layer, a feature fusion layer and a first output layer; the feature extraction layer includes multiple feature extraction sublayers; the feature extraction layer extracts image features of the target reference area image transmitted through the first input layer, and inputs the generated feature map into the feature fusion layer; the feature fusion layer fuses the feature submap of at least one target feature extraction sublayer through multiple sampling, and outputs the position of the vehicle image block through the first output layer.

8. The vehicle speed detection method according to claim 7, characterized in that: The feature extraction layer includes an initial feature extraction layer and multiple groups of semantic feature extraction layers with the same structure; The initial feature extraction layer is connected to the first input layer, and its output is connected to the input of the first semantic feature extraction layer, the output of each semantic feature extraction layer is connected to the input of the next semantic feature extraction layer, and the last semantic feature extraction layer is connected to the first output layer; The initial feature extraction layer includes at least a two-dimensional convolution layer, a batch normalization layer, and an activation function layer; each semantic feature extraction layer includes a first convolution layer and a multi-scale feature extraction layer with the same convolution kernel size, step size, and padding parameters and an increasing number of channels; the multi-scale feature extraction layer includes at least a second convolution layer, a feature splitting layer, a local feature extraction layer, a semantic fusion layer, and a third convolution layer; Among them, the second convolutional layer is connected to the corresponding first convolutional layer, and outputs the extracted feature map to the feature splitting layer, the feature splitting layer splits the received feature map into a first sub-feature map and a second sub-feature map, inputs the first sub-feature map to the local feature extraction layer, and inputs the second sub-feature map to the semantic fusion layer; the local feature extraction layer includes a plurality of feature extraction blocks with the same structure and connected in series, the total number of feature extraction blocks is determined according to the number of input channel dimensions, and the feature extraction block includes two convolution layers with the same convolution kernel size and different step sizes; the output of the first feature extraction block of the local feature extraction layer is also connected to the semantic fusion layer, the semantic fusion layer is connected to the last feature extraction block, and outputs the fusion result of the received feature map to the third convolutional layer.

9. The vehicle speed detection method according to claim 7, characterized in that: The feature fusion layer at least includes a first upsampling layer, a first splicing layer, a first multi-scale feature extraction layer, a second upsampling layer, a second splicing layer, a second multi-scale feature extraction layer, a fourth convolutional layer, a third splicing layer, a third multi-scale feature extraction layer, a fifth convolutional layer and a fourth multi-scale feature extraction layer, which are connected in sequence; Among them, the first upsampling layer receives the output feature map of the feature extraction layer, and the feature extraction layer includes a first target multi-scale feature extraction layer and a second target multi-scale feature extraction layer with the largest parameter scale according to the data transmission direction; the first splicing layer is also connected to the second target multi-scale feature extraction layer, the second splicing layer is also connected to the first target multi-scale feature extraction layer, and the third splicing layer is also connected to the first multi-scale feature extraction layer.

10. The vehicle speed detection method according to any one of claims 1 to 6, characterized in that: Determining vehicle tracking data matching the vehicle to be speed measured from the video stream image frame of the current field of view based on the vehicle matching image includes: Acquire feature extraction parameters of a feature extraction layer of a vehicle detection network model for performing vehicle recognition on the target reference area image; Determining, according to the network structure of the feature extraction layer, a matching feature extraction network for extracting vehicle matching image features of a vehicle matching image, and a network structure of a video feature extraction network for extracting image features of each frame of a video stream image frame, wherein feature extraction parameters of the matching feature extraction network and the video feature extraction network are consistent with the feature extraction parameters; The matching feature extraction network, the video feature extraction network, and the feature matching network are combined into a vehicle matching network model; the input end of the feature matching network is connected to the output of the matching feature extraction network and the output of the video feature extraction network, respectively, the vehicle matching image features are matched with the image features of each frame, and the vehicle category, vehicle position, and corresponding confidence score of each vehicle contained in the image frame of the video stream are output; According to the output result of the vehicle matching network model, vehicle tracking data matching the vehicle to be speed measured is determined.

11. The vehicle speed detection method according to claim 10, characterized in that: The feature matching network includes a second input layer, multiple target matching layers, a detection head, and a second output layer; each target matching layer includes a position feature fusion layer, multiple simultaneously running operation processing layers, a first multi-feature fusion layer connected in sequence, a layer normalization layer, a fully connected layer, and a second multi-feature fusion layer; The second input layer is connected to the video feature extraction network and the matching feature extraction network, respectively, and flattens the received image features of each frame and the vehicle matching image features, and inputs the flattened video image features to the first operation processing layer, and inputs the flattened matching image features to the position feature fusion layer and the first multi-feature fusion layer; position-encodes the vehicle matching image features in a manner that sets the background block to infinity, and inputs the position-encoded features to the position feature fusion layer; Each operation processing layer performs matrix multiplication on the position feature fusion data and the flattened features of the video image, scales the first operation result, processes the scaling result using an activation function, performs matrix multiplication on the processing result and the flattened features of the video image, and inputs the second operation result into the first multi-feature fusion layer, and the output ends of the first multi-feature fusion layer are respectively connected to the layer normalization layer and the second multi-feature fusion layer.

12. The vehicle speed detection method according to claim 11, characterized in that: The detection head includes a video frame detection layer and a matching feature processing layer; The video frame detection layer is connected to the second multi-feature fusion layer, and includes a confidence calculation branch, a vehicle category identification branch, and a vehicle position calculation branch; the confidence calculation branch, the vehicle category identification branch, and the vehicle position calculation branch include convolution layers with the same convolution kernel size, step value, and padding value, and two-dimensional convolution layers with the same convolution kernel size, step value, and padding value but different numbers of channels; The matching feature processing layer receives vehicle matching image features and includes a sixth convolutional layer, a feature addition layer, and a feature multiplication layer connected in sequence, wherein the output of the sixth convolutional layer is also connected to the feature multiplication layer, the feature addition layer is also connected to the output of the convolutional layer of the vehicle category identification branch, and the output of the feature multiplication layer is connected to the two-dimensional convolutional layer of the vehicle category identification branch; The vehicle category identification branch also includes a category determination layer. After the category determination layer is connected to the two-dimensional convolution layer of the vehicle category identification branch, when the vehicle category identification result of the target vehicle corresponding to the current image frame is received, the distance between the target vehicle and other vehicle categories is calculated, and when the change in the distance between the target vehicle and other vehicle categories meets the preset distance change condition, the vehicle category identification result of the target vehicle is output.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the vehicle speed detection method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the vehicle speed detection method according to any one of claims 1 to 12 are implemented.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the vehicle speed detection method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Traffic vehicle information acquisition method based on Mask R-CNN

    CN110379168A

  • Vehicle speed measurement method based on video analysis

    CN111753797A

  • Video stitching-based multi-camera collaborative swimming speed measurement method and system

    CN116309685A

  • Method and device for determining vehicle running speed and electronic equipment

    CN116311965A

  • Personnel trajectory detection method, device, system and equipment and storage medium

    CN116486438A