Traffic flow visual monitoring method and device based on AI full-automatic calibration

Through the AI-based vehicle flow visual monitoring method, the mapping estimation processing is used to use video stream data to achieve real-time accurate output of vehicle locations, solving the problem of insufficient efficiency and accuracy of traditional traffic monitoring methods, and is suitable for a variety of traffic monitoring scenarios and reducing costs.

CN119992839AInactive Publication Date: 2025-05-13QINGYUNARCH (BEIJING) INNOVATION TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510466176.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional traffic monitoring methods rely on fixed sensors, which are difficult to meet the needs of modern traffic management for efficient and accurate monitoring.

Method used

The vehicle flow visual monitoring method based on AI is adopted, and the video stream data is obtained and the mapping parameters from plane to space are obtained, and the actual position of the vehicle is output in real time.

Benefits of technology

There is no need to add new hardware and manual calibration, real-time and accurate output of vehicle locations is achieved, real-time and accuracy of monitoring is improved, and is suitable for a variety of traffic monitoring scenarios and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992839A_ABST
    Figure CN119992839A_ABST
Patent Text Reader

Abstract

The invention provides a traffic flow visual monitoring method and device based on AI full-automatic calibration, and the method comprises the steps: obtaining video stream data of a target road section, carrying out the mapping presumption processing, obtaining a mapping parameter from a plane to a space, and outputting the plane coordinates of a target vehicle in real time based on the video stream data, the space actual position of the target vehicle relative to the target road section is obtained in real time through the plane coordinates and the mapping parameters, and mapping from the two-dimensional image coordinates to the three-dimensional space coordinates can be achieved through the mapping parameters; according to the method, hardware does not need to be newly added, manual calibration is not needed, the vehicle position can be accurately output in real time, multiple scenes are adapted, and traffic intelligence is efficiently assisted at low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicle flow monitoring, and in particular to a vehicle flow visual monitoring method and device based on AI fully automatic calibration. Background Art

[0002] With the acceleration of urbanization and the continuous increase in the number of cars, traffic flow has increased dramatically, and problems such as traffic congestion and frequent accidents have become increasingly serious, bringing huge pressure to urban traffic management. In order to effectively respond to these challenges, it has become a top priority to achieve real-time monitoring and precise management of traffic conditions.

[0003] Traditional traffic monitoring methods mainly rely on fixed sensors, such as geomagnetic coils, radars, etc. These sensors have played a certain role in traffic monitoring and can obtain some traffic information, such as flow, speed, etc. However, in actual applications, they have gradually exposed many drawbacks and are unable to meet the needs of modern traffic management for efficient and accurate monitoring. Summary of the invention

[0004] In view of this, the embodiments of the present application provide a vehicle flow visual monitoring method, device, electronic device and storage medium based on AI fully automatic calibration. It does not require additional hardware or manual calibration, can accurately output vehicle positions in real time, adapt to multiple scenarios, and help promote intelligent transportation with high efficiency and low cost.

[0005] The technical solution of the embodiment of the present application is implemented as follows: In a first aspect, an embodiment of the present application provides a vehicle flow visual monitoring method based on AI fully automatic calibration, the method comprising: Obtain video stream data of the target road section; Performing mapping inference processing based on the video stream data to obtain mapping parameters of the target road section from plane to space; and outputting the plane coordinates of the target vehicle in the video image plane in real time based on the video stream data; wherein the mapping parameters are used to achieve mapping of two-dimensional image coordinates to three-dimensional space coordinates; The actual position of the target vehicle is output in real time based on the plane coordinates and the mapping parameters; wherein the actual position is the actual spatial position relative to the target road section.

[0006] In a second aspect, the embodiment of the present application further provides a vehicle flow visual monitoring device based on AI fully automatic calibration, the device comprising: An acquisition module, used to acquire video stream data of a target road section; A processing module, configured to perform mapping inference processing based on the video stream data to obtain mapping parameters of the target road section from plane to space; and output the plane coordinates of the target vehicle in the video image plane in real time based on the video stream data; wherein the mapping parameters are used to realize the mapping of two-dimensional image coordinates to three-dimensional space coordinates; An output module is used to output the actual position of the target vehicle in real time based on the plane coordinates and the mapping parameters; wherein the actual position is the actual spatial position relative to the target road section.

[0007] In the third aspect, an embodiment of the present application also provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to execute the vehicle flow visual monitoring method based on AI fully automatic calibration as described in any one of the first aspects.

[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the vehicle flow visual monitoring method based on AI fully automatic calibration as described in any one of the first aspects is executed.

[0009] The embodiments of the present application have the following beneficial effects: The embodiment of the present application only needs to obtain the video stream data of the target road section, without the need for additional hardware configuration, and can efficiently utilize existing monitoring resources to reduce costs; the plane-to-space mapping parameters are obtained through fully automatic mapping inference processing, without the need for complex manual calibration, simplifying the operation process; the plane coordinates and actual position of the target vehicle can be output in real time to enhance the real-time monitoring; the two-dimensional to three-dimensional mapping is used to improve the monitoring accuracy, which can more realistically reflect the actual status of the vehicle; and the actual position information of the vehicle obtained is suitable for traffic flow monitoring, speed measurement, digital twins and other scenarios, has wide practicality, and provides strong support for the development of intelligent transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0011] Figure 1 is a flowchart of steps S101-S103 provided in an embodiment of the present application; Figure 2It is a flowchart of steps S201-S204 provided in an embodiment of the present application; Figure 3 It is a schematic diagram of the overall method architecture provided by the embodiment of the present application; Figure 4 It is a flowchart of steps S401-S403 provided in an embodiment of the present application; Figure 5 It is a flowchart of steps S501-S503 provided in an embodiment of the present application; Figure 6 is a schematic diagram of a feature point grid provided in an embodiment of the present application; Figure 7 It is a structural schematic diagram of a vehicle flow visual monitoring device based on AI full-automatic calibration provided in an embodiment of the present application; Figure 8 It is a schematic diagram of the composition structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0012] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.

[0013] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0015] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0016] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0018] The embodiment of the present application aims to realize a kind of traffic flow visual monitoring based on AI automatic calibration, by acquiring the video stream data of the target road section, performing mapping inference processing to obtain the mapping parameters from plane to space, and combining the plane coordinates of the target vehicle in the video image plane output in real time, and finally outputting the actual position of the target vehicle in real time. Figure 1 , Figure 1 is a flow chart of steps S101-S103 of the vehicle flow visual monitoring method based on AI automatic calibration provided in the embodiment of the present application, which will be combined with Figure 1 Steps S101-S103 are shown for explanation.

[0019] In step S101, video stream data of a target road section is obtained.

[0020] Here, the camera equipment installed on the target road section is used to collect video stream data in real time. These video stream data contain the driving conditions of vehicles on the target road section, providing basic data for subsequent processing. Optionally, the collected video stream data can be pre-processed, such as denoising, contrast enhancement, etc., to improve the quality and availability of the data.

[0021] In step S102, mapping inference processing is performed based on the video stream data to obtain mapping parameters of the target road section from plane to space; and the plane coordinates of the target vehicle in the video image plane are output in real time based on the video stream data; wherein the mapping parameters are used to realize the mapping of two-dimensional image coordinates to three-dimensional space coordinates.

[0022] For mapping inference processing, key frames can be first extracted from the video stream data. Key frames should meet specific conditions, such as not being covered by traffic, including complete lane lines, and clarity greater than the clarity threshold. Then, lane marking target detection and semantic segmentation processing are performed on the key frames to obtain the location information, segmentation instances, distribution characteristics, and pixel coordinates of lane line corners. The feature point grid and coordinate pair construction is based on the selected calibration feature points to form a feature point grid, and the initial coordinate pair set from pixel coordinates to actual road marking coordinates is constructed by interpolation selection and other methods. The mapping parameter solution is based on the lane marking standard specification knowledge (such as highway width, lane line spacing, lane line width standard) and lane line distribution characteristics, combined with algorithm judgment, to automatically calibrate the actual position coordinates of the feature points and form the corresponding relationship between the actual coordinates and the pixel coordinates. Finally, using these coordinate pairs, through the inverse operation of spatial transformation, the least squares solution of the mapping matrix is ​​obtained to obtain the homography matrix H, which is the mapping parameter from plane to space, used to realize the mapping of two-dimensional image coordinates to three-dimensional space coordinates.

[0023] For real-time output of the plane coordinates of the target vehicle in the video image plane, an efficient and low-latency AI visual vehicle detection model based on the Transformer architecture can be used to process the video stream data in real time. The vehicle detection module can accurately detect the position of the target vehicle in the video image plane, and use the pixel coordinates of the vehicle target centroid in the image as the plane coordinates of the vehicle.

[0024] In step S103, the actual position of the target vehicle is output in real time based on the plane coordinates and the mapping parameters; wherein the actual position is the actual spatial position relative to the target road section.

[0025] Here, the obtained mapping parameters and the plane coordinates of the target vehicle are used to calculate the actual position of the target vehicle in real time through linear transformation operations. The actual position is the actual spatial position relative to the target road section. The latitude and longitude coordinates and distance of the vehicle can be further calculated by combining the actual horizontal coordinates of the vehicle with the latitude and longitude of the camera, the road direction and other information.

[0026] In some embodiments, see Figure 2 , Figure 2 It is a flow chart of steps S201-S204 provided in an embodiment of the present application. The mapping inference processing is performed based on the video stream data to obtain the mapping parameters of the target road section from the plane to the space. It can be achieved through steps S201-S204, which will be explained in combination with each step.

[0027] In step S201, key frame extraction is performed on the video stream data to obtain video key frames; wherein the video key frames represent video frames that meet a first specific condition, and the first specific condition includes at least one of not being covered by traffic, including complete lane lines, and clarity being greater than a clarity threshold.

[0028] In step S202, the target lane line is determined based on the video key frame, and the corner point of each lane line in the target lane line is determined; wherein the corner point of each lane line corresponds to a corner point coordinate information in the video image plane.

[0029] In step S203, the corner points of each lane line that meet the second specific condition are determined as calibration feature points, and a feature point grid is determined based on the calibration feature points; wherein the feature point grid includes multiple grid points, each of the multiple grid points corresponds to a plane coordinate information and an actual coordinate pair, the plane coordinate information represents the coordinate information of the grid point in the video image plane, the horizontal coordinate of the actual coordinate pair represents the actual width, and the vertical coordinate of the coordinate pair represents the actual length, and the actual coordinate pair is determined based on the lane marking rule.

[0030] In step S204, the mapping parameter is determined based on at least a specific number of grid points among the plurality of grid points.

[0031] As an example, see Figure 3 , Figure 3 is a schematic diagram of the overall method architecture provided by the embodiment of the present application, such as Figure 3 As shown, first, video frames that meet specific conditions are screened out from the video stream data as the basis for subsequent processing to reduce the amount of data processing and improve processing efficiency. Specifically, the video stream data is analyzed frame by frame, and video key frames are screened out according to the first specific condition (not covered by traffic, including complete lane lines, and clarity greater than a clarity threshold). For example, the traffic coverage area can be identified by a target detection algorithm, and the integrity and clarity of the lane lines can be determined by image processing technology, thereby determining the key frames that meet the conditions.

[0032] Next, based on the video keyframes, lane detection algorithms (such as Transformer-based object detection yolo v11 and instance segmentation Mask2Former model) are used to accurately identify and extract the target lane lines. Corner point detection is performed on each identified lane line to determine the corresponding corner point coordinate information in the video image plane. Corner points are usually the endpoints, intersections, or points with obvious feature changes of lane lines, which can be achieved through edge detection and corner point detection algorithms in image processing (such as Harris corner point detection, Shi-Tomasi corner point detection, etc.).

[0033] Then, the corner points of each lane line that meet the second specific condition are determined as calibration feature points. The second specific condition can be that the confidence of the corner point is higher than the set threshold, the distribution is uniform (not densely clustered, and relatively evenly tiled in the lane area), etc. For example, the response value of the corner point can be screened, and the judgment can be made based on the spatial distribution of the corner point in the image. Based on the calibration feature points, a feature point grid is constructed. The feature point grid consists of multiple grid points, and each grid point corresponds to a plane coordinate information and an actual coordinate pair. The plane coordinate information represents the coordinate information of the grid point in the video image plane, which can be directly obtained through the image pixel coordinate system. The horizontal coordinate of the actual coordinate pair represents the actual width, and the vertical coordinate represents the actual length. The actual coordinate pair is determined based on the lane marking rules (such as highway width, lane line spacing, lane line width standard, etc.). For example, based on the lane line spacing and the known lane width, the position of each grid point in the actual road can be calculated.

[0034] Finally, the plane coordinate information and actual coordinate pairs of at least a certain number of grid points among the multiple grid points are used to solve the mapping parameters through the spatial transformation model (such as homography matrix transformation) to achieve the mapping from plane to space. Select at least a certain number of grid points (such as 4 or more to ensure the stability of the solution, there is a solution when there are at least more than 4 points, and the error is basically within 1% for more than 8 points), and obtain their plane coordinates (xi, yi) and actual coordinates (Xi, Yi), i=1, 2, ⋯, n (n is the number of selected grid points). Establish a set of equations based on the spatial transformation model. For example, for the homography matrix H, there is the following relationship: ; Where (xi, yi) are plane coordinates, (x′i, y′i) are the corresponding coordinates on the image plane after the actual coordinate transformation (which can be converted through the relationship between the actual coordinates and the image resolution), and s is the scale factor. The least squares method or other optimization algorithms are used to solve the equations to obtain the mapping parameters (such as the specific parameter values ​​of the homography matrix H).

[0035] In some embodiments, see Figure 4 , Figure 4 It is a flow chart of steps S401-S403 provided in an embodiment of the present application. Determining the target lane line based on the video key frame and determining the corner point of each lane line in the target lane line can be achieved through steps S401-S403, which will be explained in combination with each step.

[0036] In step S401, target detection processing is performed on the video key frame to obtain at least one lane line; wherein each lane line in the at least one lane line corresponds to a confidence level.

[0037] In step S402, instance segmentation is performed on each lane line to obtain distribution features of each lane line in the video image plane.

[0038] In step S403, a lane line having a first confidence level higher than a first confidence level threshold is determined as the target lane line, and a corner point of each lane line in the target lane line is removed based on a distribution feature corresponding to the target lane line.

[0039] Here, a target detection algorithm is used, such as the YOLO v11 model based on deep learning, to quickly and accurately identify lane line targets in the image. The model will output at least one lane line information. Each lane line corresponds to a confidence level, which reflects the model's trust in the detection result as a true lane line. The confidence level is usually calculated based on the model's classification score. The higher the score, the more confident the model is that the detection result is a lane line.

[0040] Next, semantic / instance segmentation is used to identify the target objects (lane lines) in the image, accurately segment each target object from the background, and obtain the pixel-level distribution information of the target object. After segmenting each lane line, the distribution characteristics of the lane line in the video image plane are obtained. These distribution characteristics include the shape, length, width, curvature and other information of the lane line, which can be calculated and described by analyzing the pixel area obtained by segmentation. For example, the approximate length and width of the lane line can be obtained by calculating the bounding box of the pixel area, and the curvature information of the lane line can be obtained by fitting the pixel points in the bounding box.

[0041] Next, a first confidence threshold is set, and lane lines with a first confidence higher than the threshold are determined as target lane lines. The setting of the first confidence threshold can be adjusted according to the actual application scenario and data characteristics to ensure that both real and reliable lane lines can be screened out and valid lane line information can be avoided from being missed. For example, in an autonomous driving scenario with high requirements for lane line detection accuracy, the threshold can be set higher; while in a traffic monitoring scenario with high requirements for real-time performance but relatively low requirements for accuracy, the threshold can be appropriately lowered.

[0042] Then, based on the distribution characteristics of the target lane lines, locate the corner points of each lane line in the target lane line. Corner points are located at the end points, intersection points, or locations with large curvature changes of lane lines. Corner point location can be performed by the following methods: Combination of edge detection and corner detection algorithms: First, perform edge detection on the lane segmentation result, such as using the Canny edge detection algorithm to obtain the edge contour of the lane. Then, apply a corner detection algorithm on the edge contour, such as Harris corner detection or Shi-Tomasi corner detection, to determine the position of the corner points.

[0043] Based on lane line shape features: According to the distribution characteristics of lane lines, for example, the corners of straight lane lines are usually located at the end points of the line segments, and the corners of curved lane lines can be determined by analyzing the curvature changes. For straight lane lines, the four vertices of the segmented area boundary box can be directly found as candidate corners, and then the real corners are screened out according to the actual direction of the lane lines; for curved lane lines, the curvature change rate can be calculated, and when the curvature change rate exceeds a certain threshold, the position is determined to be a corner point.

[0044] In some embodiments, see Figure 5 , Figure 5 It is a flow chart of steps S501-S503 provided in an embodiment of the present application, wherein the second specific condition includes at least one of a second confidence level being higher than a second confidence level threshold and a uniform distribution; determining the feature point grid based on the calibrated feature points can be achieved through steps S501-S503, which will be described in conjunction with each step.

[0045] In step S501, the feature point grid is constructed by using a specific division rule and the calibrated feature points, wherein the specific division rule is determined based on the direction and / or distribution of the lane lines.

[0046] In step S502, some feature points are selected from the feature point grid as grid points; wherein the grid points are evenly distributed and match the direction and / or distribution of the lane lines.

[0047] In step S503, actual parameters are determined based on the lane marking rule, and actual coordinate pairs of the grid points are determined based on the actual parameters.

[0048] Here, the second specific condition is used to screen the calibration feature points, including at least one of the second confidence being higher than the second confidence threshold and uniform distribution. In addition to the confidence requirement, the calibration feature points also need to be evenly distributed on the image plane. Evenly distributed feature points can more comprehensively reflect the spatial information of the lane line and avoid inaccurate mapping parameter solutions due to feature points being concentrated in certain areas. For example, the distance between corner points can be calculated to ensure that the distance between adjacent calibration feature points is within a certain range, thereby achieving the uniform distribution requirement.

[0049] Next, a specific division rule for the feature point grid is constructed through specific division rules and calibrated feature points. The rule is determined based on the direction and / or distribution of the lane line. For example, if the lane line is mainly straight, the grid division can be performed in a direction parallel to the lane line; if the lane line is curved, the direction and density of the grid division can be adjusted according to the curvature change of the lane line. At the same time, the distribution of the lane line, such as the spacing and width of the lane line, is considered to make the grid division more in line with the actual situation. Using the calibrated feature points as a reference, a feature point grid is constructed on the image plane according to a specific division rule. The calibrated feature points can be used as the initial nodes of the grid, and then the positions of other grid points are determined according to the division rules. For example, in the case of a straight lane line, the calibrated feature points are used as the starting point, and the positions of other grid points are determined by extending to both sides and the front at a fixed spacing.

[0050] Based on the initially constructed feature point grid, an interpolation selection operation is performed to ensure that the grid points are evenly distributed and match the direction and / or distribution of the lane line. For example, if the grid points in some areas of the preliminary grid are too dense and some areas are too sparse, the distribution of the grid points can be adjusted by interpolation selection to make the entire grid more uniform. The grid points after interpolation selection must also meet the requirements of uniform distribution, that is, the distance between adjacent grid points is relatively uniform. At the same time, the distribution of the grid points must match the direction and / or distribution of the lane line, and can accurately reflect the spatial characteristics of the lane line. For example, at a curved lane line, the grid points should be distributed along the curved direction of the lane line to better capture the shape changes of the lane line.

[0051] Finally, the actual parameters are determined based on the lane marking rules, which include information such as the standard width of the lane line, the lane line spacing, and the number of lanes. For example, according to traffic regulations, the lane width of a highway is generally 3.5 meters, and the distance between the starting points of the two dotted lines is 15 meters. Using the lane marking rules and the distribution of the lane lines in the image, the actual parameters can be determined. For example, by measuring the pixel width of the lane line in the image, and calculating the actual lane line width based on the ratio between the actual lane line width and the pixel width. Based on the determined actual parameters, the actual coordinate pair can be determined for each grid point. The abscissa of the actual coordinate pair represents the actual width, and the ordinate represents the actual length.

[0052] As an example, see Figure 6 , Figure 6 is a schematic diagram of a feature point grid provided in an embodiment of the present application, such as Figure 6 As shown, the feature point grid is located in the area where the road surface is relatively flat / horizontal. The grid coordinate unit is meter, the horizontal axis is width, and the vertical axis is length.

[0053] In some embodiments, determining the mapping parameter based on at least a specific number of grid points in the plurality of grid points comprises: Based on the plane coordinate information corresponding to the at least specific number of grid points and the actual coordinate pairs, a least square solution of a mapping matrix from two-dimensional image coordinates to three-dimensional space coordinates is calculated to obtain a homography matrix; wherein the specific number is greater than or equal to 4; The homography matrix is ​​used as the mapping parameter.

[0054] Here, the mapping from two-dimensional image coordinates to three-dimensional space coordinates can be solved by the homography matrix. When the coordinates of multiple grid points in the image plane (plane coordinate information) and their coordinates in the actual three-dimensional space (actual coordinate pairs) are known, the homography matrix can be solved by the least squares method to achieve two-dimensional to three-dimensional mapping.

[0055] Collect the plane coordinate information (xi, yi) and actual coordinate pairs (Xi, Yi, Zi) of at least a certain number (greater than or equal to 4, usually between 8 and 16) of grid points (in practical applications, if the main focus is on plane mapping, Zi can be set to a constant or simplified according to the specific scenario. Here, the focus is on the homography of 2D to 2D planes, that is, assuming the mapping on a specific plane, the Z coordinate can be normalized or its change effect can be ignored). In solving the homography matrix of common road scenes, the plane-to-plane mapping can be mainly considered, which can be simplified to the correspondence between (xi, yi) and (Xi, Yi).

[0056] The homography matrix H is a 3×3 matrix that satisfies the following relationship for plane-to-plane mapping: ; In the formula, ~ represents a proportional relationship, and the value of H is expressed as: .

[0057] Expanding the above formula gives two linear equations: ; ; Then we can introduce the normalization of homogeneous coordinates and transform the above equations into linear form. For n grid points (n≥4), we can get 2n linear equations, forming a linear equation system Ah=0, where A is a 2n×9 matrix and h is a 9×1 vector containing the elements of the homography matrix.

[0058] Due to the presence of noise and errors in actual data, the linear equation system Ah = 0 usually has no exact solution. Therefore, the least squares method is used to solve the approximate solution of the equation system and finally construct the homography matrix H.

[0059] In some embodiments, an AI visual vehicle detection model is constructed based on a Transformer architecture, and traffic is monitored through the AI ​​visual vehicle detection model, wherein the determination of the target lane line is implemented based on the Mask2Former model backbone of the Transformer architecture, and the detection of the target vehicle is implemented based on the YoloV11 model backbone of the Transformer architecture, and the AI ​​visual vehicle detection model is pre-trained based on the COCO dataset, and the training data of the AI ​​visual vehicle detection model is data augmented and enhanced data, and the augmentation means include but are not limited to horizontal flipping and ±12% saturation adjustment to enhance data generalization ability.

[0060] The embodiment of this application aims to use the Transformer architecture to build an AI visual vehicle detection model to achieve accurate monitoring of traffic. By using the Mask2Former model backbone to determine the target lane line, the yolo v11 model backbone to detect the target vehicle, and pre-training based on the COCO dataset, combined with data augmentation and enhancement technology to improve the generalization ability of the model, the accuracy and reliability of vehicle detection can be effectively improved.

[0061] Mask2Former is an instance segmentation model based on the Transformer architecture, which can accurately segment different target areas in the image. In the embodiment of the present application, its feature extraction and segmentation capabilities are used to determine the target lane line.

[0062] Specifically, the video keyframe is used as the input image, and after preprocessing (such as scaling and normalization), it is sent to the Mask2Former model. The model first extracts features from the input image through the Transformer encoder to capture global and local feature information in the image, and then uses the Transformer decoder combined with the mask prediction head to perform instance segmentation on the lane lines to obtain the precise distribution of each lane line in the image. Finally, the segmentation results are post-processed, such as removing noise and smoothing boundaries, to obtain clear and accurate target lane lines.

[0063] YoloV11 is an improved target detection model based on the Transformer architecture. It inherits the fast and accurate characteristics of the YOLO series models, and uses the Transformer's self-attention mechanism to enhance the ability of feature extraction and target detection. In the embodiment of the present application, the image after the target lane line is determined is used as input and sent to the YoloV11 model. The model extracts image features through the Transformer backbone network, performs target detection at multiple scales, and outputs the category, position and confidence of each detection box. Finally, the detection results are processed by NMS, overlapping detection boxes are removed, and the most accurate detection results are retained.

[0064] The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and key point detection dataset that contains rich image and annotation information.

[0065] Using the COCO dataset to pre-train the AI ​​visual vehicle detection model allows the model to learn a wealth of target features and detection patterns. During the pre-training process, an appropriate learning rate, batch size, and number of training rounds are used to ensure that the model can fully converge.

[0066] In the embodiment of the present application, the data augmentation means include: Horizontal flip: Flip the training images horizontally to increase data diversity and improve the model's ability to detect vehicles in different directions.

[0067] Saturation adjustment: Randomly adjust the saturation of the image within a range of ±12% so that the model can adapt to vehicle detection under different lighting and color conditions.

[0068] In addition to the above methods, it may also include: Random cropping: Randomly crop the image to increase the model's ability to learn different perspectives and local features.

[0069] Noise addition: Add a certain amount of Gaussian noise to the image to improve the robustness of the model.

[0070] In some embodiments, the plane coordinates of the target vehicle in the video image plane are determined based on the pixel coordinates of the target centroid of the target vehicle in the video image plane; The real-time output of the actual position of the target vehicle based on the plane coordinates and the mapping parameters comprises: The plane coordinates are transformed and calculated based on the mapping parameters to obtain the actual position; wherein the transformation calculation is optimized and accelerated in parallel through matrix operations to improve the response speed.

[0071] Here, the plane coordinates of the target vehicle in the video image plane are determined based on the pixel coordinates of the target vehicle's target centroid in the video image plane. Through the homography matrix, the linear space mapping transformation can be realized to obtain the actual position of the target vehicle. In order to improve the response speed of the conversion calculation, the matrix operation optimization and acceleration parallel method can be adopted, including but not limited to: using GPU (graphics processing unit) for matrix operations, using efficient matrix operation algorithms, such as Strassen algorithm, etc., to reduce the computational complexity of matrix multiplication, and at the same time, vectorizing the matrix operation, making full use of the CPU's SIMD (single instruction multiple data) instruction set to improve computing efficiency, using parallel computing frameworks, such as OpenMP, MPI, etc., to distribute the conversion computing tasks to multiple computing nodes for parallel execution.

[0072] In some embodiments, the method further comprises: Determine the actual longitude and latitude coordinates of the target vehicle based on the actual position, the longitude and latitude of the shooting device of the video stream data, and the road direction; Alternatively, determining distance information of the target vehicle compared to the photographing device based on the actual position; Alternatively, the speed information of the target vehicle is determined based on the change in the actual position of the same target vehicle in any two frames of video and the time difference.

[0073] Here, the actual latitude and longitude coordinates of the target vehicle can be calculated by combining the actual position of the target vehicle relative to the shooting device, the latitude and longitude of the shooting device itself, and the road direction information. The latitude and longitude of the shooting device provide a reference for the geographic location, the actual position describes the spatial relationship of the target vehicle relative to the shooting device, and the road direction is used to convert the relative position into a geographic coordinate direction.

[0074] The distance information of the target vehicle compared to the camera can be directly calculated from its actual position, and the distance between the target vehicle and the camera can be calculated by plane geometry methods.

[0075] In video monitoring, the speed information of the target vehicle can be determined by calculating the actual position change (displacement) of the same target vehicle in any two frames of video and the time difference between the two frames. Specifically, for any two frames of video, the actual position of the same target vehicle in the two frames and the time difference between the two frames are obtained respectively, and then the displacement of the target vehicle between the two frames is calculated, and then the speed can be obtained based on the displacement and time difference.

[0076] In summary, the embodiments of the present application have the following beneficial effects: (1) The embodiments of the present application do not require the configuration of new equipment hardware, and can directly achieve accurate grasp of the real-time traffic situation based on existing traffic monitoring cameras. In the field of traffic management, a large number of monitoring cameras have been deployed in many places. The embodiments of the present application can make full use of these existing resources and avoid the high hardware procurement, installation and maintenance costs caused by the addition of new equipment. By transforming the existing cameras with AI, they are equipped with intelligent vehicle detection and traffic flow monitoring capabilities, which maximizes the utilization of resources and saves a lot of capital investment for traffic management departments.

[0077] (2) Traditional camera calibration methods usually require personnel to be present to perform complex operations and obtain precise camera parameters. However, the embodiment of the present application does not require personnel to be present to calibrate the camera, nor does it require camera parameters. It directly extracts feature points from the image and uses advanced AI algorithms to automatically complete the mapping of video pixel coordinates to actual horizontal coordinates, breaking through the limitations of traditional methods, greatly simplifying the operation process, and lowering the technical threshold. Even personnel without professional camera calibration knowledge can easily implement system deployment and application, improving work efficiency and accuracy.

[0078] (3) The embodiments of the present application do not require a large amount of manual data annotation, and only require hundreds of data to achieve excellent recognition results. By adopting advanced data augmentation and enhancement technologies, such as horizontal flipping and saturation adjustment, the diversity of training data is effectively expanded and the generalization ability of the model is improved. This enables the model to converge quickly and obtain a high recognition accuracy rate under limited data resources, greatly shortening the model training cycle and reducing manpower and time costs.

[0079] (4) The architecture of the embodiment of the present application is highly efficient and uses models based on the Transformer architecture, such as Mask2Former and YoloV11. These models can process image data quickly and accurately. At the same time, through matrix operation optimization and acceleration parallelization, the computing speed is further improved and real-time response is achieved. The efficient architecture not only improves the performance of the system, but also reduces the demand for computing resources, so that the cost of the entire solution is effectively controlled and has a high cost-effectiveness.

[0080] (5) The embodiments of the present application have a wide range of application scenarios and are suitable for various traffic speed measurement, flow monitoring, digital twins, traffic jam monitoring and other scenarios. In particular, in the traffic digital twin system, this solution can transform the existing traffic monitoring cameras into AI, making them an important data source in the digital twin system. By acquiring traffic information in real time, accurate input is provided for the digital twin model, and virtual mapping and real-time monitoring of the traffic system are realized. This helps traffic management departments to better understand traffic conditions, formulate scientific and reasonable traffic management strategies, improve traffic operation efficiency, reduce traffic congestion, and provide strong support for the intelligent development of urban transportation.

[0081] To sum up, the embodiments of the present application have significant advantages in terms of equipment utilization, operating procedures, data labeling, architecture efficiency and application scenarios, and can bring efficient and low-cost solutions to the field of traffic management and promote the intelligent development of the transportation industry.

[0082] Based on the same inventive concept, the embodiments of the present application also provide a vehicle flow visual monitoring device based on AI fully automatic calibration corresponding to the vehicle flow visual monitoring method based on AI fully automatic calibration in the first embodiment. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the above-mentioned vehicle flow visual monitoring method based on AI fully automatic calibration, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0083] like Figure 7 As shown, Figure 7 : is a schematic diagram of the structure of a vehicle flow visual monitoring device 700 based on AI full-automatic calibration provided in an embodiment of the present application. The vehicle flow visual monitoring device 700 based on AI full-automatic calibration includes: An acquisition module 701 is used to acquire video stream data of a target road section; The processing module 702 is used to perform mapping inference processing based on the video stream data to obtain mapping parameters of the target road section from plane to space; and output the plane coordinates of the target vehicle in the video image plane in real time based on the video stream data; wherein the mapping parameters are used to realize the mapping of two-dimensional image coordinates to three-dimensional space coordinates; The output module 703 is used to output the actual position of the target vehicle in real time based on the plane coordinates and the mapping parameters; wherein the actual position is the actual spatial position relative to the target road section.

[0084] Those skilled in the art should understand that Figure 7 The implementation functions of each unit in the vehicle flow visual monitoring device 700 based on AI full-automatic calibration can be understood by referring to the relevant description of the aforementioned vehicle flow visual monitoring method based on AI full-automatic calibration. Figure 7The functions of each unit in the AI-based fully automatic calibration vehicle flow visual monitoring device 700 shown can be implemented through a program running on a processor, or can be implemented through a specific logic circuit.

[0085] In a possible implementation manner, the processing module 702 performs mapping inference processing based on the video stream data to obtain mapping parameters of the target road section from a plane to a space, including: Performing key frame extraction processing on the video stream data to obtain video key frames; wherein the video key frames represent video frames that meet a first specific condition, and the first specific condition includes at least one of not being covered by traffic, including complete lane lines, and clarity being greater than a clarity threshold; Determine the target lane line based on the video key frame, and determine the corner point of each lane line in the target lane line; wherein the corner point of each lane line corresponds to a corner point coordinate information in the video image plane; Determine the corner points of each lane line that meet the second specific condition as calibration feature points, and determine a feature point grid based on the calibration feature points; wherein the feature point grid includes a plurality of grid points, each of the plurality of grid points corresponds to a plane coordinate information and an actual coordinate pair, the plane coordinate information represents the coordinate information of the grid point in the video image plane, the abscissa of the actual coordinate pair represents the actual width, the ordinate of the coordinate pair represents the actual length, and the actual coordinate pair is determined based on the lane marking rule; The mapping parameter is determined based on at least a certain number of grid points among the plurality of grid points.

[0086] In a possible implementation, the processing module 702 determines the target lane line based on the video key frame, and determines the corner point of each lane line in the target lane line, including: Performing target detection processing on the video key frame to obtain at least one lane line; wherein each lane line in the at least one lane line corresponds to a confidence level; Performing instance segmentation on each lane line to obtain distribution features of each lane line in the video image plane; A lane line with a first confidence level higher than a first confidence level threshold is determined as the target lane line, and a corner point of each lane line in the target lane line is removed based on a distribution feature corresponding to the target lane line.

[0087] In a possible implementation manner, the second specific condition includes at least one of a second confidence level being higher than a second confidence level threshold and a uniform distribution; The processing module 702 determines a feature point grid based on the calibrated feature points, including: Constructing the feature point grid by using a specific division rule and the calibrated feature points, wherein the specific division rule is determined based on the direction and / or distribution of the lane line; Selecting some feature points as grid points by inserting gaps in the feature point grid; wherein the grid points are evenly distributed and match the direction and / or distribution of the lane line; Actual parameters are determined based on the lane marking rule, and actual coordinate pairs of the grid points are determined based on the actual parameters.

[0088] In a possible implementation manner, the processing module 702 determines the mapping parameter based on at least a specific number of grid points among the multiple grid points, including: Based on the plane coordinate information corresponding to the at least specific number of grid points and the actual coordinate pairs, a least square solution of a mapping matrix from two-dimensional image coordinates to three-dimensional space coordinates is calculated to obtain a homography matrix; wherein the specific number is greater than or equal to 4; The homography matrix is ​​used as the mapping parameter.

[0089] In one possible implementation, the processing module 702 constructs an AI visual vehicle detection model based on the Transformer architecture, and monitors traffic through the AI ​​visual vehicle detection model, wherein the determination of the target lane line is implemented based on the Mask2Former model backbone of the Transformer architecture, and the detection of the target vehicle is implemented based on the YoloV11 model backbone of the Transformer architecture, and the AI ​​visual vehicle detection model is pre-trained based on the COCO dataset, and the training data of the AI ​​visual vehicle detection model is data augmented and enhanced data, and the augmentation means include but are not limited to horizontal flipping and ±12% saturation adjustment to enhance data generalization ability.

[0090] In a possible implementation manner, the plane coordinates of the target vehicle in the video image plane are determined based on the pixel coordinates of the target centroid of the target vehicle in the video image plane; The output module 703 outputs the actual position of the target vehicle in real time based on the plane coordinates and the mapping parameters, including: The plane coordinates are transformed and calculated based on the mapping parameters to obtain the actual position; wherein the transformation calculation is optimized and accelerated in parallel through matrix operations to improve the response speed.

[0091] In a possible implementation, the output module 703 further includes: Determine the actual longitude and latitude coordinates of the target vehicle based on the actual position, the longitude and latitude of the shooting device of the video stream data, and the road direction; Alternatively, determining distance information of the target vehicle compared to the photographing device based on the actual position; Alternatively, the speed information of the target vehicle is determined based on the change in the actual position of the same target vehicle in any two frames of video and the time difference.

[0092] The above-mentioned vehicle flow visual monitoring device based on AI automatic calibration has the following beneficial effects: (1) The embodiments of the present application do not require the configuration of new equipment hardware, and can directly achieve accurate grasp of the real-time traffic situation based on existing traffic monitoring cameras. In the field of traffic management, a large number of monitoring cameras have been deployed in many places. The embodiments of the present application can make full use of these existing resources and avoid the high hardware procurement, installation and maintenance costs caused by the addition of new equipment. By transforming the existing cameras with AI, they are equipped with intelligent vehicle detection and traffic flow monitoring capabilities, which maximizes the utilization of resources and saves a lot of capital investment for traffic management departments.

[0093] (2) Traditional camera calibration methods usually require personnel to be present to perform complex operations and obtain precise camera parameters. However, the embodiment of the present application does not require personnel to be present to calibrate the camera, nor does it require camera parameters. It directly extracts feature points from the image and uses advanced AI algorithms to automatically complete the mapping of video pixel coordinates to actual horizontal coordinates, breaking through the limitations of traditional methods, greatly simplifying the operation process, and lowering the technical threshold. Even personnel without professional camera calibration knowledge can easily implement system deployment and application, improving work efficiency and accuracy.

[0094] (3) The embodiments of the present application do not require a large amount of manual data annotation, and only require hundreds of data to achieve excellent recognition results. By adopting advanced data augmentation and enhancement technologies, such as horizontal flipping and saturation adjustment, the diversity of training data is effectively expanded and the generalization ability of the model is improved. This enables the model to converge quickly and obtain a high recognition accuracy rate under limited data resources, greatly shortening the model training cycle and reducing manpower and time costs.

[0095] (4) The architecture of the embodiment of the present application is highly efficient and uses models based on the Transformer architecture, such as Mask2Former and YoloV11. These models can process image data quickly and accurately. At the same time, through matrix operation optimization and acceleration parallelization, the computing speed is further improved and real-time response is achieved. The efficient architecture not only improves the performance of the system, but also reduces the demand for computing resources, so that the cost of the entire solution is effectively controlled and has a high cost-effectiveness.

[0096] (5) The embodiments of the present application have a wide range of application scenarios and are suitable for various traffic speed measurement, flow monitoring, digital twins, traffic jam monitoring and other scenarios. In particular, in the traffic digital twin system, this solution can transform the existing traffic monitoring cameras into AI, making them an important data source in the digital twin system. By acquiring traffic information in real time, accurate input is provided for the digital twin model, and virtual mapping and real-time monitoring of the traffic system are realized. This helps traffic management departments to better understand traffic conditions, formulate scientific and reasonable traffic management strategies, improve traffic operation efficiency, reduce traffic congestion, and provide strong support for the intelligent development of urban transportation.

[0097] To sum up, the embodiments of the present application have significant advantages in terms of equipment utilization, operating procedures, data labeling, architecture efficiency and application scenarios, and can bring efficient and low-cost solutions to the field of traffic management and promote the intelligent development of the transportation industry.

[0098] like Figure 8 As shown, Figure 8 The present invention provides a schematic diagram of the structure of an electronic device 800 according to an embodiment of the present invention. The electronic device 800 includes: A processor 801, a storage medium 802 and a bus 803, wherein the storage medium 802 stores machine-readable instructions executable by the processor 801. When the electronic device 800 is running, the processor 801 communicates with the storage medium 802 via the bus 803, and the processor 801 executes the machine-readable instructions to perform the steps of the vehicle flow visual monitoring method based on AI fully automatic calibration described in the embodiment of the present application.

[0099] In actual application, the components in the electronic device 800 are coupled together via bus 803. It is understood that bus 803 is used to realize the connection and communication between these components. Bus 803 includes not only a data bus, but also a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 8 Various buses are labeled as bus 803.

[0100] The electronic device has the following beneficial effects: (1) The embodiments of the present application do not require the configuration of new equipment hardware, and can directly achieve accurate grasp of the real-time traffic situation based on existing traffic monitoring cameras. In the field of traffic management, a large number of monitoring cameras have been deployed in many places. The embodiments of the present application can make full use of these existing resources and avoid the high hardware procurement, installation and maintenance costs caused by the addition of new equipment. By transforming the existing cameras with AI, they are equipped with intelligent vehicle detection and traffic flow monitoring capabilities, which maximizes the utilization of resources and saves a lot of capital investment for traffic management departments.

[0101] (2) Traditional camera calibration methods usually require personnel to be present to perform complex operations and obtain precise camera parameters. However, the embodiment of the present application does not require personnel to be present to calibrate the camera, nor does it require camera parameters. It directly extracts feature points from the image and uses advanced AI algorithms to automatically complete the mapping of video pixel coordinates to actual horizontal coordinates, breaking through the limitations of traditional methods, greatly simplifying the operation process, and lowering the technical threshold. Even personnel without professional camera calibration knowledge can easily implement system deployment and application, improving work efficiency and accuracy.

[0102] (3) The embodiments of the present application do not require a large amount of manual data annotation, and only require hundreds of data to achieve excellent recognition results. By adopting advanced data augmentation and enhancement technologies, such as horizontal flipping and saturation adjustment, the diversity of training data is effectively expanded and the generalization ability of the model is improved. This enables the model to converge quickly and obtain a high recognition accuracy rate under limited data resources, greatly shortening the model training cycle and reducing manpower and time costs.

[0103] (4) The architecture of the embodiment of the present application is highly efficient and uses models based on the Transformer architecture, such as Mask2Former and YoloV11. These models can process image data quickly and accurately. At the same time, through matrix operation optimization and acceleration parallelization, the computing speed is further improved and real-time response is achieved. The efficient architecture not only improves the performance of the system, but also reduces the demand for computing resources, so that the cost of the entire solution is effectively controlled and has a high cost-effectiveness.

[0104] (5) The embodiments of the present application have a wide range of application scenarios and are suitable for various traffic speed measurement, flow monitoring, digital twins, traffic jam monitoring and other scenarios. In particular, in the traffic digital twin system, this solution can transform the existing traffic monitoring cameras into AI, making them an important data source in the digital twin system. By acquiring traffic information in real time, accurate input is provided for the digital twin model, and virtual mapping and real-time monitoring of the traffic system are realized. This helps traffic management departments to better understand traffic conditions, formulate scientific and reasonable traffic management strategies, improve traffic operation efficiency, reduce traffic congestion, and provide strong support for the intelligent development of urban transportation.

[0105] To sum up, the embodiments of the present application have significant advantages in terms of equipment utilization, operating procedures, data labeling, architecture efficiency and application scenarios, and can bring efficient and low-cost solutions to the field of traffic management and promote the intelligent development of the transportation industry.

[0106] The embodiment of the present application also provides a computer-readable storage medium, which stores executable instructions. When the executable instructions are executed by at least one processor 801, the vehicle flow visual monitoring method based on AI fully automatic calibration described in the embodiment of the present application is implemented.

[0107] In some embodiments, the storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface storage, an optical disk, or a compact disc read-only memory (CD-ROM); it can also be various devices including one or any combination of the above memories.

[0108] In some embodiments, executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0109] As an example, executable instructions may, but do not necessarily, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (for example, files storing one or more modules, subroutines, or code portions).

[0110] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0111] The computer-readable storage medium has the following advantages: (1) The embodiments of the present application do not require the configuration of new equipment hardware, and can directly achieve accurate grasp of the real-time traffic situation based on existing traffic monitoring cameras. In the field of traffic management, a large number of monitoring cameras have been deployed in many places. The embodiments of the present application can make full use of these existing resources and avoid the high hardware procurement, installation and maintenance costs caused by the addition of new equipment. By transforming the existing cameras with AI, they are equipped with intelligent vehicle detection and traffic flow monitoring capabilities, which maximizes the utilization of resources and saves a lot of capital investment for traffic management departments.

[0112] (2) Traditional camera calibration methods usually require personnel to be present to perform complex operations and obtain precise camera parameters. However, the embodiment of the present application does not require personnel to be present to calibrate the camera, nor does it require camera parameters. It directly extracts feature points from the image and uses advanced AI algorithms to automatically complete the mapping of video pixel coordinates to actual horizontal coordinates, breaking through the limitations of traditional methods, greatly simplifying the operation process, and lowering the technical threshold. Even personnel without professional camera calibration knowledge can easily implement system deployment and application, improving work efficiency and accuracy.

[0113] (3) The embodiments of the present application do not require a large amount of manual data annotation, and only require hundreds of data to achieve excellent recognition results. By adopting advanced data augmentation and enhancement technologies, such as horizontal flipping and saturation adjustment, the diversity of training data is effectively expanded and the generalization ability of the model is improved. This enables the model to converge quickly and obtain a high recognition accuracy rate under limited data resources, greatly shortening the model training cycle and reducing manpower and time costs.

[0114] (4) The architecture of the embodiment of the present application is highly efficient and uses models based on the Transformer architecture, such as Mask2Former and YoloV11. These models can process image data quickly and accurately. At the same time, through matrix operation optimization and acceleration parallelization, the computing speed is further improved and real-time response is achieved. The efficient architecture not only improves the performance of the system, but also reduces the demand for computing resources, so that the cost of the entire solution is effectively controlled and has a high cost-effectiveness.

[0115] (5) The embodiments of the present application have a wide range of application scenarios and are suitable for various traffic speed measurement, flow monitoring, digital twins, traffic jam monitoring and other scenarios. In particular, in the traffic digital twin system, this solution can transform the existing traffic monitoring cameras into AI, making them an important data source in the digital twin system. By acquiring traffic information in real time, accurate input is provided for the digital twin model, and virtual mapping and real-time monitoring of the traffic system are realized. This helps traffic management departments to better understand traffic conditions, formulate scientific and reasonable traffic management strategies, improve traffic operation efficiency, reduce traffic congestion, and provide strong support for the intelligent development of urban transportation.

[0116] To sum up, the embodiments of the present application have significant advantages in terms of equipment utilization, operating procedures, data labeling, architecture efficiency and application scenarios, and can bring efficient and low-cost solutions to the field of traffic management and promote the intelligent development of the transportation industry.

[0117] In the several embodiments provided in the present application, it should be understood that the disclosed methods and electronic devices can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0118] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0120] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a platform server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0121] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A vehicle flow visual monitoring method based on AI automatic calibration, characterized in that: The method comprises: Obtain video stream data of the target road section; Performing mapping inference processing based on the video stream data to obtain mapping parameters of the target road section from plane to space; and outputting the plane coordinates of the target vehicle in the video image plane in real time based on the video stream data; wherein the mapping parameters are used to achieve mapping of two-dimensional image coordinates to three-dimensional space coordinates; The actual position of the target vehicle is output in real time based on the plane coordinates and the mapping parameters; wherein the actual position is the actual spatial position relative to the target road section.

2. The method according to claim 1, characterized in that The mapping inference processing is performed based on the video stream data to obtain the mapping parameters of the target road section from the plane to the space, including: Performing key frame extraction processing on the video stream data to obtain video key frames; wherein the video key frames represent video frames that meet a first specific condition, and the first specific condition includes at least one of not being covered by traffic, including complete lane lines, and clarity being greater than a clarity threshold; Determine the target lane line based on the video key frame, and determine the corner point of each lane line in the target lane line; wherein the corner point of each lane line corresponds to a corner point coordinate information in the video image plane; Determine the corner points of each lane line that meet the second specific condition as calibration feature points, and determine a feature point grid based on the calibration feature points; wherein the feature point grid includes a plurality of grid points, each of the plurality of grid points corresponds to a plane coordinate information and an actual coordinate pair, the plane coordinate information represents the coordinate information of the grid point in the video image plane, the abscissa of the actual coordinate pair represents the actual width, the ordinate of the coordinate pair represents the actual length, and the actual coordinate pair is determined based on the lane marking rule; The mapping parameter is determined based on at least a certain number of grid points among the plurality of grid points.

3. The method according to claim 2, characterized in that The determining of the target lane line based on the video key frame and determining the corner point of each lane line in the target lane line includes: Performing target detection processing on the video key frame to obtain at least one lane line; wherein each lane line in the at least one lane line corresponds to a confidence level; Performing instance segmentation on each lane line to obtain distribution features of each lane line in the video image plane; A lane line with a first confidence level higher than a first confidence level threshold is determined as the target lane line, and a corner point of each lane line in the target lane line is removed based on a distribution feature corresponding to the target lane line.

4. The method according to claim 2, characterized in that: The second specific condition includes at least one of a second confidence level being higher than a second confidence level threshold and a uniform distribution; The determining of a feature point grid based on the calibrated feature points comprises: Constructing the feature point grid by using a specific division rule and the calibrated feature points, wherein the specific division rule is determined based on the direction and / or distribution of the lane line; Selecting some feature points as grid points by inserting spaces in the feature point grid; wherein the grid points are evenly distributed and match the direction and / or distribution of the lane line; Actual parameters are determined based on the lane marking rule, and actual coordinate pairs of the grid points are determined based on the actual parameters.

5. The method according to claim 2, characterized in that: The determining the mapping parameter based on at least a specific number of grid points among the plurality of grid points comprises: Based on the plane coordinate information corresponding to the at least specific number of grid points and the actual coordinate pairs, a least square solution of a mapping matrix from two-dimensional image coordinates to three-dimensional space coordinates is calculated to obtain a homography matrix; wherein the specific number is greater than or equal to 4; The homography matrix is ​​used as the mapping parameter.

6. The method according to claim 2, characterized in that An AI visual vehicle detection model is constructed based on the Transformer architecture, and the traffic flow is monitored by the AI ​​visual vehicle detection model, wherein the determination of the target lane line is implemented based on the Mask2Former model backbone of the Transformer architecture, and the detection of the target vehicle is implemented based on the YoloV11 model backbone of the Transformer architecture. The AI ​​visual vehicle detection model is pre-trained based on the COCO dataset, and the training data of the AI ​​visual vehicle detection model is data augmented and enhanced data, and the augmentation means include but are not limited to horizontal flipping and ±12% saturation adjustment to enhance data generalization ability.

7. The method according to claim 1, characterized in that The plane coordinates of the target vehicle in the video image plane are determined based on the pixel coordinates of the target centroid of the target vehicle in the video image plane; The real-time output of the actual position of the target vehicle based on the plane coordinates and the mapping parameters comprises: The plane coordinates are transformed and calculated based on the mapping parameters to obtain the actual position; wherein the transformation calculation is optimized and accelerated in parallel through matrix operations to improve the response speed.

8. The method according to claim 7, characterized in that The method further comprises: Determine the actual longitude and latitude coordinates of the target vehicle based on the actual position, the longitude and latitude of the shooting device of the video stream data, and the road direction; Alternatively, determining distance information of the target vehicle compared to the photographing device based on the actual position; Alternatively, the speed information of the target vehicle is determined based on the change in the actual position of the same target vehicle in any two frames of video and the time difference.

9. A vehicle flow visual monitoring device based on AI automatic calibration, characterized in that: The device comprises: An acquisition module, used to acquire video stream data of a target road section; A processing module, configured to perform mapping inference processing based on the video stream data to obtain mapping parameters of the target road section from plane to space; and output the plane coordinates of the target vehicle in the video image plane in real time based on the video stream data; wherein the mapping parameters are used to realize the mapping of two-dimensional image coordinates to three-dimensional space coordinates; An output module is used to output the actual position of the target vehicle in real time based on the plane coordinates and the mapping parameters; wherein the actual position is the actual spatial position relative to the target road section.

10. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to execute the vehicle flow visual monitoring method based on AI fully automatic calibration as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Shared bicycle auxiliary positioning method and device based on monocular vision

    CN111062986A

  • Spatial calibration method and system

    CN112950717A

  • Vehicle-mounted panoramic camera calibration method and device based on lane line and storage medium

    CN113963066A

  • Vehicle detection method and system based on monocular vision and deep learning

    CN114419547A

  • Railway vehicle speed measurement method and device based on binocular vision

    CN116840503A