A parking lot display method and device based on multi-camera fusion
By deploying multiple cameras in the parking lot, calibrating the parking area, and calculating the affine transformation matrix, the problem of scattered parking lot monitoring images in existing technologies is solved, achieving unified presentation of vehicle location information and improving management efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AIPARK TECHNOLOGY CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-06-09
AI Technical Summary
In existing parking lot monitoring systems, the monitoring images captured by each camera are independent of each other, making it difficult to integrate vehicle location information in a unified manner. This results in the overall parking status of the parking lot being difficult to obtain intuitively, leading to low management efficiency.
Multiple cameras are deployed in the parking lot to calibrate the parking area and calculate the affine transformation matrix. The coordinates of the vehicle area and wheel contact points are mapped from the scene image to the planar display image. Through intersection-union ratio calculation and threshold filtering for duplicate detection, the unified position information of all vehicles in the parking lot is obtained, and the vehicle display model is drawn on the planar display image.
It enables a unified presentation of vehicle location information, improves parking lot management efficiency, and allows for intuitive access to the real-time parking status of parking lots.
Smart Images

Figure CN122176062A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and specifically to a parking lot display method and apparatus based on multi-camera fusion. Background Technology
[0002] In the development of modern cities, with the continuous increase in car ownership, parking lots, as a key component of urban transportation infrastructure, face increasingly severe challenges in management and operation. Efficient and intelligent parking management systems play a crucial role in alleviating urban traffic pressure, improving user experience, and optimizing resource allocation. Existing parking lots typically use cameras to monitor vehicle entry, exit, and parking, assisting managers in understanding parking status. However, in practice, multiple cameras are often deployed to cover different areas, with each camera capturing independent footage, usually displayed in a multi-screen or polling manner. Due to significant differences in perspective and spatial coordinate systems between cameras, existing monitoring methods struggle to integrate vehicle location information from various feeds. Managers must frequently switch between multiple feeds to roughly determine the overall parking status. This approach not only fails to intuitively reflect the global distribution of vehicles within the parking lot but also easily leads to unclear vehicle location identification, duplicate or missed identification in situations with dense parking spaces or a large number of vehicles, thus affecting the efficiency of quickly acquiring parking lot operational status and making management decisions. Summary of the Invention
[0003] This application provides a parking lot display method and device based on multi-camera fusion, which solves the technical problems in the prior art where parking lot monitoring images are scattered and vehicle location information is difficult to present in a unified manner, resulting in difficulty in intuitively obtaining the overall parking status of the parking lot and low management efficiency.
[0004] The first aspect of this application provides a parking lot display method based on multi-camera fusion, the method comprising:
[0005] Multiple cameras are deployed within the parking lot, and parking areas are identified in the scene images captured by these cameras. Based on the positional relationship between key points of the identified parking areas in the scene images and corresponding points in the planar display image, an affine transformation matrix is calculated from the scene images to the planar display image. The parking areas are then drawn in the planar display image based on the affine transformation matrix, resulting in a target planar display image. The images acquired in real-time by the multiple cameras are detected to obtain the coordinates of vehicle areas and wheel contact points. Combining the wheel contact point coordinates and the vehicle areas in the target planar display image, position mapping is performed. Through intersection-union ratio (IU) calculation and threshold filtering for duplicate detection, unified position information for all vehicles in the parking lot is obtained. A vehicle display model is then drawn in the target planar display image based on the unified position information to visually present the real-time parking status of the parking lot.
[0006] A second aspect of this application provides a parking lot display device based on multi-camera fusion, the device comprising:
[0007] Parking Area Calibration Module: Multiple cameras are deployed within the parking lot to calibrate parking areas in the scene images captured by these cameras. Parking Area Drawing Module: Based on the positional relationship between key points of the calibrated parking areas in the scene images and their corresponding points in the planar display image, an affine transformation matrix is calculated from the scene images to the planar display image. The parking areas are then drawn in the planar display image based on this affine transformation matrix, resulting in a target planar display image. Image Detection Module: Images acquired in real-time by the multiple cameras are detected to obtain the coordinates of vehicle areas and wheel contact points. Duplicate Detection Module: The wheel contact point coordinates and vehicle areas are mapped in the target planar display image. Duplicate detection is performed using intersection-union ratio (IU) calculation and threshold filtering to obtain unified position information for all vehicles in the parking lot. Visualization Module: Vehicle display models are drawn in the target planar display image based on the unified position information, presenting the real-time parking status of the parking lot in a visual manner.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0009] First, multiple cameras are deployed within the parking lot, and parking areas are identified in the scene images captured by these cameras. Next, based on the positional relationship between key points of the identified parking areas in the scene images and their corresponding points in the planar display image, an affine transformation matrix is calculated from the scene images to the planar display image. The parking areas are then drawn on the planar display image based on this affine transformation matrix, resulting in the target planar display image. Then, images acquired in real-time by multiple cameras are inspected to obtain the coordinates of vehicle areas and wheel contact points. Further, combining the wheel contact point coordinates and the vehicle areas on the target planar display image, position mapping is performed. Through intersection-union ratio (IU / U) calculation and threshold filtering for duplicate detection, unified position information for all vehicles in the parking lot is obtained. Finally, a vehicle display model is drawn on the target planar display image based on the unified position information, visually presenting the real-time parking status of the parking lot. This solves the technical problems of existing technologies where scattered parking lot monitoring images and difficulty in uniformly presenting vehicle position information lead to a lack of intuitive understanding of the overall parking status and low management efficiency. By using multi-camera fusion to obtain unified vehicle position information and visually presenting it in a planar display image, the technical effect of improving parking lot management efficiency is achieved. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic flowchart of a parking lot display method based on multi-camera fusion is provided for an embodiment of this application;
[0012] Figure 2 This is a schematic diagram of a parking lot scene and multi-camera setup provided in an embodiment of this application;
[0013] Figure 3 This application provides a floor plan view of a parking area in a car-free state, as shown in the embodiments of this application.
[0014] Figure 4 This application provides a vehicle detection target frame and a diagram showing the detection results of the four wheel contact points, as shown in the embodiments of this application.
[0015] Figure 5 A diagram illustrating the parking status in a floor plan, as provided in an embodiment of this application.
[0016] Figure 6 This is a schematic diagram of a parking lot display device based on multi-camera fusion, provided as an embodiment of this application.
[0017] Figure labeling: Parking area calibration module 11, parking area drawing module 12, image detection module 13, duplicate detection module 14, visualization module 15. Detailed Implementation
[0018] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0019] Example 1, as Figure 1 As shown, this application provides a parking lot display method based on multi-camera fusion, wherein the method includes:
[0020] Multiple cameras are deployed in the parking lot, and the parking area is marked in the scene images captured by the multiple cameras.
[0021] In this embodiment, the installation location and number of cameras are determined based on the spatial structure and parking area distribution of the parking lot. Cameras are installed on the top, pillars, or walls of the parking lot according to a preset arrangement rule, ensuring that the field of view of each camera covers the corresponding parking area, and that there is partial overlap between the fields of view of adjacent cameras to guarantee complete coverage of the parking area. The camera installation angle is set so that the optical axis points towards the ground, avoiding large-angle pitch or tilt to reduce the impact of perspective distortion on subsequent mapping calculations. Figure 2 In the parking lot shown, cameras were installed at four selected azimuth angles, at a height of 6 to 8 meters, to cover the entire parking area.
[0022] After the cameras are installed, scene images of the parking lot are acquired using each camera. Scene images with no obstructions and stable lighting are selected as input calibration images. The parking area, including parking spaces, lanes, and traffic flow areas, is manually calibrated in the calibration images using image annotation tools. During calibration, polygon vertices are sequentially selected along the parking space boundaries, lane lines, and traffic flow lines. The pixel coordinates of each vertex in the corresponding scene image are recorded, thus forming calibration data describing the spatial extent of the parking area. This calibration data is stored as a calibration configuration file associated with the camera number and serves as input for subsequent scene image-to-planar display mapping calculations.
[0023] Furthermore, multiple cameras are deployed within the parking lot, and parking areas are marked in the scene images captured by the multiple cameras, including:
[0024] Multiple cameras are arranged in the parking lot according to preset rules, which are that the field of view of the multiple cameras covers the entire parking area and the angle of the cameras should be parallel to the ground; using an image annotation tool, the boundaries and positions of the parking area are marked in the scene images of the multiple cameras, and the annotation content includes the edge lines, lane lines and guide lines of each parking space.
[0025] Multiple cameras are deployed within the parking lot according to preset rules. These rules determine the number and location of cameras based on the overall layout and distribution of parking areas, ensuring that each camera's field of view covers the entire parking area, with at least partial overlap between adjacent cameras. During installation, the optical axis of each camera is set to be substantially parallel to the ground to guarantee a stable planar structure of the parking area in the scene image, reducing the impact of image distortion on subsequent processing. After the cameras are deployed and the corresponding scene images are acquired, image annotation tools are used to calibrate the parking area in the scene images captured by each camera. The calibration process includes sequentially marking parking space edges, lane lines, and vehicle flow guide lines along the actual structural boundaries of the parking area in the scene image, and recording the pixel coordinates of these edges, lane lines, and flow guide lines in the corresponding scene image. This forms calibration data describing the spatial relationship between the parking area boundaries and the location of the parking area. This calibration data is used for subsequent parking area mapping and display processing.
[0026] Based on the positional relationship between the key points of the calibrated parking area in the scene image and the corresponding points in the planar display image, the affine transformation matrix from the scene image to the planar display image is calculated, and the parking area is drawn in the planar display image based on the affine transformation matrix to obtain the target planar display image.
[0027] In the scene image where the parking area has been calibrated, multiple key points with clear geometric significance are selected from the calibration results of parking space edges, lane lines, or guide lines. These key points include corner points, line segment endpoints, or line segment intersections within the parking area. Simultaneously, target point locations corresponding one-to-one with these key points are determined in a pre-defined plan view of the parking lot, such as a top-down view. Figure 3 As shown. Based on the pixel coordinates of the key points in the scene image and the planar coordinates of the corresponding points in the planar display image, a set of key point pairs is constructed. The set of key point pairs is then solved using an affine transformation method to obtain an affine transformation matrix describing the mapping relationship between the scene image coordinate system and the planar display image coordinate system. Subsequently, the affine transformation matrix is applied to the pixel coordinates of each edge line, lane line, and guide line in the calibrated parking area to perform coordinate transformation. The corresponding parking area structure is then drawn in the planar display image, thereby generating a target planar display image containing the complete spatial layout of the parking area. This target planar display image serves as a unified spatial reference for subsequent vehicle position mapping and display.
[0028] Furthermore, based on the positional relationship between the key points of the calibrated parking area in the scene image and the corresponding points in the planar display image, an affine transformation matrix from the scene image to the planar display image is calculated, and the parking area is drawn in the planar display image based on the affine transformation matrix to obtain the target planar display image, including:
[0029] Select M key points in the scene image of the calibrated parking area. The key points include the corners of the parking area or the intersections of lane lines. Obtain the corresponding positional relationship of the M key points on the planar display map. Calculate the affine transformation matrix based on the corresponding positional relationship using the affine transformation calculation function in the computer vision library.
[0030] In the scene image where the parking area has been calibrated, M key points (no fewer than three) with stable geometric features are selected from the calibration results. These key points include corner points of the parking area or intersections of lane lines, and their pixel coordinates in the corresponding scene image are recorded. Simultaneously, target position coordinates corresponding to each key point are determined in the planar display image, forming a key point correspondence set. Based on the pixel coordinates of the key points in the scene image and their corresponding coordinates in the planar display image, an affine transformation calculation function from a computer vision library (such as the `getAffineTransform` function in OpenCV) is called to solve the key point correspondence set, calculating an affine transformation matrix describing the mapping relationship from the scene image coordinate system to the planar display image coordinate system. Subsequently, the affine transformation matrix is applied to the coordinate data of each edge and region contour in the calibrated parking area, and the corresponding parking area structure is drawn in the planar display image, thereby generating the target planar display image.
[0031] For example, in camera A, selecting key points in the scene map. and the corresponding points in the planar diagram. Ma is the affine transformation matrix corresponding to scene A. The affine transformation can be expressed as:
[0032] .
[0033] The images captured in real time by the multiple cameras are detected to obtain the coordinates of the vehicle area and the wheel contact point.
[0034] In this embodiment, each camera continuously acquires real-time scene images of the parking lot at a preset acquisition frequency, and inputs these scene images into a vehicle detection and processing unit. The vehicle detection and processing unit, based on a pre-trained vehicle detection model, performs frame-by-frame inference processing on the input scene images. During a single detection process, it simultaneously outputs the target region of the vehicle in the scene image and multiple key point information corresponding to the vehicle target. The target region is represented by a rectangular bounding box, used to describe the spatial range of the vehicle in the scene image. The key point information includes the wheel contact points corresponding to the contact positions of the vehicle tires with the ground. For each detected wheel contact point, its pixel coordinate values in the corresponding scene image are recorded. These pixel coordinate values are associated with and stored with the target region information of the corresponding vehicle, serving as input data for subsequent vehicle position mapping and multi-camera fusion processing.
[0035] Furthermore, the images captured in real time by the multiple cameras are detected to obtain the coordinates of the vehicle area and the wheel contact point, including:
[0036] The vehicle detection model based on YOLO11-Pose is used to process the images acquired in real time by the multiple cameras. The vehicle detection model outputs the coordinates of the vehicle area and the wheel contact point through single-stage detection.
[0037] Preferably, scene images acquired in real-time by multiple cameras at a preset acquisition frequency are input into a vehicle detection model based on YOLO11-Pose for processing. YOLO11-Pose is a real-time pose estimation method based on the YOLO framework, achieving keypoint localization of targets through single-stage detection, offering advantages such as high efficiency and end-to-end training. The vehicle detection model employs a single-stage detection architecture, simultaneously performing vehicle target detection and keypoint regression processing during a single forward inference operation on the input image, thereby synchronously outputting the target region of the vehicle in the scene image and the corresponding wheel contact point coordinates. The vehicle target region is represented by a rectangular bounding box, characterizing the spatial extent of the vehicle in the scene image; the wheel contact point coordinates are the pixel coordinates of the tire contact position with the ground in the scene image coordinate system. The detected vehicle region information and wheel contact point coordinates are associated and stored as the foundation data for subsequent vehicle position mapping and multi-camera fusion processing.
[0038] Furthermore, the training method for the vehicle detection model includes:
[0039] Collect an image dataset containing various vehicle models, colors, lighting conditions, and parking scenarios; for each image in the image dataset, label the vehicle target with a bounding box and the ground contact coordinates of the four tires to form a training set and a validation set; use the YOLO code framework to perform end-to-end training of the vehicle detection model based on the training set, adjust the learning rate and batch size parameters, and evaluate the model accuracy using the validation set.
[0040] Preferably, firstly, an image dataset for model training is collected, covering various vehicle types, vehicle colors, different lighting conditions, and different parking lot application scenarios to improve the model's adaptability to complex parking environments. Then, vehicle targets in each image of the dataset are manually labeled, including rectangular bounding boxes representing the vehicle's spatial range and the coordinates of the contact points of the four tires with the ground. The labeled image data is then divided into training and validation sets according to a preset ratio. After data preparation, the vehicle detection model is built based on the YOLO code framework. The model is trained end-to-end using the training set. By adjusting training parameters such as the learning rate and batch size, the model simultaneously learns vehicle target detection and wheel contact point position regression tasks. During training, the model's detection accuracy is evaluated using the validation set to obtain a vehicle detection model that meets the needs of practical applications.
[0041] Specifically, a trained vehicle detection model is used to detect vehicles in each frame of images captured by multiple cameras. The vehicle detection model outputs bounding boxes for each vehicle in the image, as well as the contact patch coordinates of the four tires. For example... Figure 4 As shown, the specific range of the vehicle area can be determined by the target box, and the grounding position information of the wheels can be obtained by the grounding point coordinates to locate the position of the vehicle on the ground plane.
[0042] By combining the wheel contact point coordinates and the vehicle area position mapping on the target plane display, and through intersection-union ratio calculation and threshold filtering for duplicate detection, the unified position information of all vehicles in the field is obtained.
[0043] For the vehicle region and wheel contact point coordinates acquired by multiple cameras, based on the affine transformation matrix corresponding to each camera, the bounding box coordinates of the vehicle region and the pixel coordinates of the wheel contact points are mapped from the scene image coordinate system to the planar coordinate system of the target planar display image, thus obtaining the mapped position of each vehicle in the target planar display image. For example, in camera A, a certain vehicle scene... Figure 4 Wheel grounding point is And the coordinates of the points in the corresponding planar diagram are: Ma is the affine transformation matrix corresponding to scene A. The affine transformation can be expressed as:
[0044] .
[0045] Because the coverage of the entire parking lot area by multiple cameras overlaps, to avoid duplicate detection and information redundancy, the vehicle mapping results from different cameras can be aggregated. Specifically, the intersection-union ratio (IUR) is calculated for any two mapped vehicle regions (the IUR is the ratio of the intersection area to the union area of two rectangles). When the IUR is greater than a preset fusion threshold, the two vehicle regions are determined to correspond to the same vehicle, and their location information is merged. When the IUR is less than or equal to the fusion threshold, they are determined to be different vehicles, and their location information is retained separately. Through the above IUR calculation and threshold filtering for duplicate detection, unified location information for all vehicles covering the entire parking lot area is obtained. This unified location information is used for subsequent vehicle display and status visualization.
[0046] Furthermore, mapping the position of the wheel contact point and the vehicle area on the target plane display diagram, including:
[0047] The coordinates of the wheel contact point and the coordinates of the vehicle area in the scene image are substituted into the affine transformation matrix of the corresponding camera to perform calculations, thereby obtaining the coordinates of the corresponding vehicle in the target plane display image, which serve as the position mapping result.
[0048] Preferably, for the vehicle detection results acquired by each camera, the rectangular bounding box coordinates of the vehicle region in the scene image and the corresponding wheel contact point coordinates of the vehicle are extracted respectively. The affine transformation matrix corresponding to the camera is then used to perform coordinate transformation operations on the rectangular bounding box coordinates and wheel contact point pixel coordinates, mapping them from the scene image coordinate system to the planar coordinate system of the target planar display image. Based on the result of the coordinate transformation operation, the position coordinates of each vehicle in the target planar display image are obtained, and the position coordinates are used as the position mapping result of the vehicle for subsequent multi-camera fusion processing and vehicle status display.
[0049] Furthermore, by calculating the intersection-union ratio and using threshold filtering to detect duplicates, unified location information for all vehicles in the field is obtained, including:
[0050] For all vehicle mapping positions from different cameras in the target planar display image, calculate the intersection-union ratio (IUGR) between any two vehicle rectangles; if the IUGR of any two vehicle rectangles exceeds a preset fusion threshold, they are determined to be the same vehicle, and the corresponding two location information are merged, and a unique identifier is assigned to each vehicle obtained after merging, which is then added to the unified location information.
[0051] After obtaining the vehicle position mapping results from different cameras in the target planar display image, the positions of the rectangular boxes corresponding to all vehicles are aggregated into the same planar coordinate system. The intersection-union ratio (IUGR) of any two vehicle rectangular boxes is calculated. The IUGR is used to characterize the degree of overlap between the two vehicle rectangular boxes in spatial position. When the IUGR of any two vehicle rectangular boxes exceeds a preset fusion threshold, it is determined that the two vehicle rectangular boxes correspond to the same vehicle. The corresponding position information is merged, and a unified vehicle position representation is generated based on the merged position information. At the same time, a unique identifier is assigned to each merged vehicle, and the unique identifier is associated with the corresponding vehicle position information and stored to form a unified position information set for all vehicles in the venue, which is used for subsequent vehicle display and parking status management.
[0052] A vehicle display model is drawn based on the unified location information in the target plan view to present the real-time parking status of the parking lot in a visual manner.
[0053] After obtaining the unified location information of all vehicles, the unique identifier corresponding to each vehicle and its position coordinates in the target planar display image are read from the unified location information. Based on the position coordinates, a vehicle display model is loaded and drawn in the target planar display image. The vehicle display model represents the vehicle's occupancy status and spatial position in the planar display image; wherein, as shown... Figure 5 As shown, the vehicle display model can be drawn using two-dimensional graphic symbols, simplified vehicle outlines, or rectangular placeholders, and its display position on the target planar display map is updated in real time based on the unified location information of the vehicles. Through continuous drawing and refreshing of the vehicle display model, a real-time visual display of the parking lot's vehicle distribution is formed on the target planar display map, thus intuitively presenting the current parking status of the parking lot and providing an intuitive reference for parking lot operation and management.
[0054] In summary, the embodiments of this application have at least the following technical effects:
[0055] First, multiple cameras are deployed within the parking lot, and parking areas are identified in the scene images captured by these cameras. Next, based on the positional relationship between key points of the identified parking areas in the scene images and their corresponding points in the planar display image, an affine transformation matrix is calculated from the scene images to the planar display image. The parking areas are then drawn on the planar display image based on this affine transformation matrix, resulting in the target planar display image. Then, images acquired in real-time by multiple cameras are inspected to obtain the coordinates of vehicle areas and wheel contact points. Further, combining the wheel contact point coordinates and the vehicle areas on the target planar display image, position mapping is performed. Through intersection-union ratio (IU / U) calculation and threshold filtering for duplicate detection, unified position information for all vehicles in the parking lot is obtained. Finally, a vehicle display model is drawn on the target planar display image based on the unified position information, visually presenting the real-time parking status of the parking lot. This solves the technical problems of existing technologies where scattered parking lot monitoring images and difficulty in uniformly presenting vehicle position information lead to a lack of intuitive understanding of the overall parking status and low management efficiency. By using multi-camera fusion to obtain unified vehicle position information and visually presenting it in a planar display image, the technical effect of improving parking lot management efficiency is achieved.
[0056] Example 2 is based on the same inventive concept as the multi-camera fusion-based parking lot display method in the foregoing examples, such as... Figure 6 As shown, this application provides a parking lot display device based on multi-camera fusion, wherein the device includes:
[0057] Parking area calibration module 11: Multiple cameras are deployed in the parking lot to calibrate parking areas in the scene images captured by the multiple cameras; Parking area drawing module 12: Based on the positional relationship between the key points of the calibrated parking area in the scene image and the corresponding points in the planar display image, the affine transformation matrix from the scene image to the planar display image is calculated, and the parking area is drawn in the planar display image based on the affine transformation matrix to obtain the target planar display image; Image detection module 13: The images captured in real time by the multiple cameras are detected to obtain the coordinates of the vehicle area and the wheel contact point; Duplicate detection module 14: The wheel contact point coordinates and the vehicle area are combined to perform position mapping in the target planar display image, and duplicate detection is performed by calculating the intersection-union ratio and threshold filtering to obtain the unified position information of all vehicles in the parking lot; Visualization module 15: A vehicle display model is drawn in the target planar display image based on the unified position information to present the real-time parking status of the parking lot in a visual manner.
[0058] Furthermore, the parking area calibration module 11 is used to perform the following method:
[0059] Multiple cameras are arranged in the parking lot according to preset rules, which are that the field of view of the multiple cameras covers the entire parking area and the angle of the cameras should be parallel to the ground; using an image annotation tool, the boundaries and positions of the parking area are marked in the scene images of the multiple cameras, and the annotation content includes the edge lines, lane lines and guide lines of each parking space.
[0060] Furthermore, the parking area drawing module 12 is used to perform the following method:
[0061] Select M key points in the scene image of the calibrated parking area. The key points include the corners of the parking area or the intersections of lane lines. Obtain the corresponding positional relationship of the M key points on the planar display map. Calculate the affine transformation matrix based on the corresponding positional relationship using the affine transformation calculation function in the computer vision library.
[0062] Furthermore, the image detection module 13 is used to perform the following method:
[0063] The vehicle detection model based on YOLO11-Pose is used to process the images acquired in real time by the multiple cameras. The vehicle detection model outputs the coordinates of the vehicle area and the wheel contact point through single-stage detection.
[0064] Furthermore, the image detection module 13 is used to perform the following method:
[0065] Collect an image dataset containing various vehicle models, colors, lighting conditions, and parking scenarios; for each image in the image dataset, label the vehicle target with a bounding box and the ground contact coordinates of the four tires to form a training set and a validation set; use the YOLO code framework to perform end-to-end training of the vehicle detection model based on the training set, adjust the learning rate and batch size parameters, and evaluate the model accuracy using the validation set.
[0066] Furthermore, the duplicate detection module 14 is used to perform the following method:
[0067] The coordinates of the wheel contact point and the coordinates of the vehicle area in the scene image are substituted into the affine transformation matrix of the corresponding camera to perform calculations, thereby obtaining the coordinates of the corresponding vehicle in the target plane display image, which serve as the position mapping result.
[0068] Furthermore, the duplicate detection module 14 is used to perform the following method:
[0069] For all vehicle mapping positions from different cameras in the target planar display image, calculate the intersection-union ratio (IUGR) between any two vehicle rectangles; if the IUGR of any two vehicle rectangles exceeds a preset fusion threshold, they are determined to be the same vehicle, and the corresponding two location information are merged, and a unique identifier is assigned to each vehicle obtained after merging, which is then added to the unified location information.
[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A parking lot display method based on multi-camera fusion, characterized in that, The method includes: Multiple cameras are deployed in the parking lot, and the parking area is marked in the scene images captured by the multiple cameras; Based on the positional relationship between the key points of the calibrated parking area in the scene image and the corresponding points in the planar display image, the affine transformation matrix from the scene image to the planar display image is calculated, and the parking area is drawn in the planar display image based on the affine transformation matrix to obtain the target planar display image. The images captured in real time by the multiple cameras are detected to obtain the coordinates of the vehicle area and the wheel contact point. By combining the wheel contact point coordinates and the vehicle area position mapping on the target plane display, and through intersection-union ratio calculation and threshold filtering for duplicate detection, the unified position information of all vehicles in the field is obtained. A vehicle display model is drawn based on the unified location information in the target plan view to present the real-time parking status of the parking lot in a visual manner.
2. The parking lot display method based on multi-camera fusion as described in claim 1, characterized in that, Multiple cameras are deployed within the parking lot, and parking areas are marked in the scene images captured by the multiple cameras, including: Multiple cameras are arranged in the parking lot according to a preset rule, which is that the field of view of the multiple cameras covers the entire parking area, and the angle of the cameras should be parallel to the ground. Using an image annotation tool, the boundaries and locations of the parking area are marked in the scene images from the multiple cameras. The annotation includes the edge lines, lane lines, and guide lines of each parking space.
3. The parking lot display method based on multi-camera fusion as described in claim 1, characterized in that, Based on the positional relationship between the key points of the calibrated parking area in the scene image and the corresponding points in the planar display image, an affine transformation matrix from the scene image to the planar display image is calculated, and the parking area is drawn in the planar display image based on the affine transformation matrix to obtain the target planar display image, including: Select M key points in the scene image of the calibrated parking area. The key points include the corner points of the parking area or the intersections of lane lines. Obtain the corresponding positional relationship of the M key points on the planar display diagram; The affine transformation matrix is calculated based on the corresponding positional relationship using the affine transformation calculation function in the computer vision library.
4. The parking lot display method based on multi-camera fusion as described in claim 3, characterized in that, Detecting images acquired in real time by the multiple cameras to obtain the coordinates of the vehicle area and wheel contact points includes: The vehicle detection model based on YOLO11-Pose is used to process the images acquired in real time by the multiple cameras. The vehicle detection model outputs the coordinates of the vehicle area and the wheel contact point through single-stage detection.
5. A parking lot display method based on multi-camera fusion as described in claim 4, characterized in that, The training method for the vehicle detection model includes: Collect image datasets containing various vehicle types, colors, lighting conditions, and parking scenarios; For each image in the image dataset, a bounding box for the vehicle target and the coordinates of the ground contact points of the four tires are labeled to form a training set and a validation set; The vehicle detection model is trained end-to-end using the YOLO code framework based on the training set, with the learning rate and batch size parameters adjusted, and the model accuracy is evaluated using a validation set.
6. The parking lot display method based on multi-camera fusion as described in claim 1, characterized in that, Based on the coordinates of the wheel contact point and the vehicle area, a positional mapping is performed on the target plane display, including: The coordinates of the wheel contact point and the coordinates of the vehicle area in the scene image are substituted into the affine transformation matrix of the corresponding camera to perform calculations, thereby obtaining the coordinates of the corresponding vehicle in the target plane display image, which serve as the position mapping result.
7. A parking lot display method based on multi-camera fusion as described in claim 6, characterized in that, By calculating the intersection-union ratio and using threshold filtering to detect duplicates, the unified location information of all vehicles in the field is obtained, including: For all vehicle mapping positions from different cameras in the target planar display image, calculate the intersection-union ratio between any two vehicle rectangles; If the intersection-union ratio of any two vehicle rectangles exceeds a preset fusion threshold, they are determined to be the same vehicle. The two corresponding location information are then merged, and a unique identifier is assigned to each vehicle obtained after merging, which is then added to the unified location information.
8. A parking lot display device based on multi-camera fusion, characterized in that, The apparatus for implementing the parking lot display method based on multi-camera fusion as described in any one of claims 1-7 comprises: Parking area calibration module: Multiple cameras are deployed in the parking lot, and the parking area is calibrated in the scene images captured by the multiple cameras; Parking area drawing module: Based on the positional relationship between the key points of the calibrated parking area in the scene image and the corresponding points in the planar display image, calculate the affine transformation matrix from the scene image to the planar display image, and draw the parking area in the planar display image based on the affine transformation matrix to obtain the target planar display image; Image detection module: Detects images acquired in real time by the multiple cameras to obtain the coordinates of the vehicle area and the wheel contact point. Duplicate detection module: Combines the wheel contact point coordinates and the vehicle area with the target plane display image to perform position mapping, and obtains the unified position information of all vehicles in the field by calculating the intersection-union ratio and threshold filtering for duplicate detection; Visualization module: Draws a vehicle display model based on the unified location information in the target planar display map, and presents the real-time parking status of the parking lot in a visual manner.