An AVM parking space visual perception data enhancement method

By constructing a multi-type, multi-quality parking marking database and combining it with multi-sensor fusion technology and feature transfer model, the shortcomings of dynamic scene interference processing and marking material fusion were solved, generating high-quality parking space visual perception data and improving the accuracy and generalization ability of the AVM parking space perception model.

CN121904517BActive Publication Date: 2026-06-02SICHUAN VOCATIONAL & TECHN COLLEGE OF COMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN VOCATIONAL & TECHN COLLEGE OF COMM
Filing Date
2026-03-25
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for AVM parking space visual perception data processing lack the ability to handle dynamic scene interference, resulting in inaccurate selection of background texture sample areas, poor visual consistency, and a lack of efficient model support for fusion of marking materials and semantic annotation mapping, thus failing to provide high-quality training data.

Method used

A multi-class, multi-quality datum line database is constructed. High-resolution cameras are used to collect data and perform semantic segmentation. By combining multi-sensor fusion technology and feature transfer models, the selection of background texture sample areas and the calculation of perspective transformation parameters are optimized to accurately simulate real datum-free scenes and improve the accuracy of semantic annotation mapping.

Benefits of technology

Generate large-scale, multi-scenario, high-quality augmented data to support the improvement of AVM parking space perception model accuracy and optimization of generalization ability, meeting the needs of autonomous driving intelligent parking systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904517B_ABST
    Figure CN121904517B_ABST
Patent Text Reader

Abstract

The application discloses an AVM parking space visual perception data enhancement method, comprising the following steps: collecting basic AVM parking space data and semantic annotation, classifying according to color, line type and width, combining quantitative indexes and subjective characteristics, and constructing a multi-class and multi-quality marking line database; after loading data, annotation and configuration parameters, positioning the original picture parking space marking line area, using a dynamic scene anti-interference detection algorithm to optimize the selection of background texture sample area, generating a marking line-free base map through image repair; retrieving matching marking line materials from the database, calculating perspective transformation parameters combined with a multi-sensor fusion positioning model, filling into the base map after adaptation; synchronously updating semantic annotation, optimizing mapping accuracy through a ViT parking space feature migration model, and assisting in database coordinate calibration through a visual-inertial fusion mapping engine. The method solves the problems of limited data coverage, poor anti-interference, low fusion accuracy and the like in the prior art, generates high-quality enhanced data, and supports the improvement of AVM parking space perception model accuracy and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving and parking technology, and in particular to a method for enhancing visual perception data of AVM parking spaces. Background Technology

[0002] With the rapid development of autonomous driving technology and intelligent parking systems, the AVM (Around View Monitor) system, as a core component of parking space perception, has significantly increased its demand for higher quality and greater diversity of parking space visual data. In real parking lot scenarios, parking space markings often exhibit different quality levels due to wear, dirt, redrawing, peeling, and other issues. Furthermore, they are affected by factors such as changes in lighting and dynamic interference. Relying solely on real-scene images collected by onboard surround view cameras is insufficient to meet the large-scale, multi-scenario, and high-quality data requirements for training AVM parking space perception models. Simultaneously, existing data acquisition methods suffer from limited scene coverage, high annotation costs, and difficulty in accurately matching the actual parking space structure and marking features during data augmentation. Therefore, there is an urgent need to construct a multi-quality marking database using a systematic approach, combining multi-model and multi-sensor fusion technologies to achieve efficient augmentation of AVM parking space visual perception data, thereby supporting improvements in perception model accuracy and optimization of generalization capabilities.

[0003] Existing technologies for AVM parking space visual perception data processing have two significant drawbacks: First, they lack the ability to handle dynamic scene interference during data augmentation. When generating a background image without parking space markings by removing the original markings, they fail to effectively combine local neighborhood pixel features with background sample differences for anti-interference detection. This easily leads to inaccurate selection of background texture sample areas, resulting in significant edge differences and poor visual consistency in the restored background image, making it impossible to accurately simulate real-world parking space scenes without markings. Second, they lack efficient model support in the marking material fusion and semantic annotation mapping stages. It is difficult to accurately calculate perspective transformation parameters through multi-sensor fusion technology to achieve spatial position adaptation of marking materials, and it also fails to optimize semantic annotation mapping accuracy through feature transfer models. This results in defects in parking space structure matching, marking feature consistency, and semantic annotation accuracy in the augmented data, making it impossible to provide high-quality training data support for the AVM parking space perception model. Summary of the Invention

[0004] In order to overcome the shortcomings and deficiencies of the existing technology, the present invention provides a method for enhancing visual perception data of AVM parking spaces.

[0005] A method for enhancing visual perception data of AVM parking spaces includes the following steps: S1, collecting basic AVM parking space data and corresponding pixel-level semantic annotations. The basic AVM parking space data consists of real-scene images acquired by an onboard surround-view camera, and the semantic annotations include semantic segmentation information of parking space markings; S2, constructing a multi-category, multi-quality marking database. Parking space marking images are captured in a real parking lot using a high-resolution camera, based on a category coverage and quality grading strategy. Semantic segmentation is performed on the images to extract the binary mask of the markings. Geometric topology and sub-pixel-level edges are obtained through contour tracking and skeletonization. Marking areas are extracted to generate materials with transparent channels. These materials are categorized by color, line type, and width, and combined with pixel intensity contrast, edge sharpness, texture continuity quantification indicators, and subjective features such as wear, dirt, redrawing, and peeling to label 10 quality levels. S3: Complete database construction; Load basic AVM parking space data, semantic annotations, and enhanced configuration parameters. Configuration parameters include target parking space structural geometric parameters, marking attributes, interference modes, and random seeds; S4: Locate the original image parking space marking area based on semantic annotations, use an image inpainting algorithm, take a 5-10 pixel outward ring area from the marking edge as the background texture sample area, analyze color distribution and texture mode, eliminate the original markings, and generate a marking-free base map; S5: Retrieve matching marking materials from the database according to the configuration parameters, determine the virtual parking space layout by combining the original image parking space edge position, and fill the corresponding area of ​​the base map with the marking materials after perspective transformation and lighting adaptation; S6: Synchronously update the semantic annotation map, map the marking filled area to the corresponding semantic label, add interference element markers, and generate an enhanced AVM image and corresponding semantic annotations.

[0006] Furthermore, in step S4, a dynamic scene anti-disturbance detection algorithm is used to optimize the selection of the background texture sample area. The algorithm expression is as follows:

[0007] ,

[0008] in, For pixels The anti-interference detection value, This represents the number of pixels in the local neighborhood. These are the weighting coefficients. The grayscale values ​​of the original image pixels. The standard deviation of local gray levels in the original image. These are the pixel values ​​after Gaussian filtering. This represents the local standard deviation after filtering. Background weight coefficient, The number of background sample pixels. The background sample pixel values, The local standard deviation is the background sample.

[0009] Furthermore, in step S5, a multi-sensor fusion positioning model is used to calculate the perspective transformation parameters. The model expression is as follows:

[0010]

[0011] The result of image coordinate transformation. These are the parameters of the perspective transformation matrix. The fusion coefficient is... For the number of sensors, For sensor weights, To locate the coordinates of the lidar, To locate coordinates for millimeter-wave radar, These are the coordinates of the original image.

[0012] Furthermore, in S6, the ViT parking space feature transfer model is used to optimize the semantic annotation mapping accuracy. The model expression is:

[0013] ,

[0014] in, For pixels Feature mapping values, For feature dimension, This represents the number of layers in the Transformer encoder. , These are query, key, and value matrices, respectively. For migration coefficient, For the number of feature neighborhoods, For the source domain eigenvalues, For the feature values ​​of the target domain, It is an L2 norm.

[0015] Furthermore, in S2, a visual-inertial fusion mapping engine is used to assist in the spatial coordinate calibration of the datum database. The engine expression is:

[0016] ,

[0017] in, For calibrated spatial coordinates, For the number of visual frames, For visual and inertial weights, For visual camera coordinates, For inertial measurement unit coordinates, To optimize the coefficients, For the number of observation data, Let be the error function. These are the observed values.

[0018] Furthermore, in S5, a lane marking fill intensity adjustment model is constructed by combining the AVM parking space visual perception data enhancement parameters. The model expression is:

[0019] ,

[0020] in, For the strength of the marking fill, These are the maximum and minimum quality levels, respectively. For the current quality level, For feature weights, For the color feature value of the grading line, For the texture feature value of the gradation line, For adjustment coefficients, For feature dimension, For the first 3D eigenvalues The average characteristic value.

[0021] Further, step S3 includes the following sub-steps: S31, when reading the basic AVM parking space data, the image resolution consistency is checked, and images of different resolutions are unified to a preset pixel size using a bilinear interpolation algorithm to ensure pixel-level alignment accuracy between the image and semantic annotation in subsequent processing; S32, the enhancement configuration parameters are parsed, and the target parking space structure parameters are converted into geometric parameters in a pixel coordinate system, including parking space length, width, arc radius, etc., and a mapping relationship between parameters and image coordinates is established; S33, a random number generator is initialized according to a random seed to determine the random selection sequence of samples during database retrieval, ensuring a balance between repeatability and randomness in sample selection during each enhancement process; S34, the semantic annotation data is format-converted, and the original annotation format is unified into a standard JSON format including the marking category ID and pixel coordinate range, which facilitates the location and processing of the marking area in subsequent steps.

[0022] Further, step S4 includes the following sub-steps: S41, based on the marking category ID in the semantic annotation, traverse the image pixels to determine the smallest bounding rectangle of all parking space marking areas, and use this as the initial area range for marking removal; S42, based on the initial area range, expand outward by 5-10 pixels to form a ring-shaped background texture sample area, and use the sliding window method to statistically analyze the gray-level histogram and texture feature matrix of the pixels in the sample area; S43, select the fast traversal method as the core of the image restoration algorithm, and use the gray-level and texture features of the sample area as a basis to perform filling calculations on the marking pixels in the initial area to generate a preliminary marking-free base map; S44, perform edge smoothing processing on the preliminary base map, and use the Gaussian filtering algorithm to eliminate the boundary differences between the restored area and the non-restored area to ensure the overall visual consistency of the base map.

[0023] Further, S5 includes the following sub-steps: S51, based on the marking attributes in the configuration parameters, perform multi-condition retrieval in a multi-type, multi-quality marking database to filter out a set of marking materials that meet the requirements of color, line type, and quality level; S52, extract the parking space edge position information from the original image, combine it with the target parking space structure parameters, and use a polygon fitting algorithm to determine the contour coordinates of the virtual parking space, dividing the specific area for marking filling; S53, perform perspective transformation processing on the selected marking materials, calculate the perspective projection matrix of the materials in the virtual parking space filling area based on the intrinsic and extrinsic parameters of the AVM camera, and complete the spatial position adaptation of the materials; S54, analyze the light intensity distribution around the filling area of ​​the original image, adjust the brightness and contrast of the marking materials through a grayscale adjustment algorithm to make the lighting conditions of the materials consistent with those of the base image, and then fill the corresponding area with the processed materials.

[0024] A method for enhancing visual perception data of parking spaces using AVM (Autonomous Vehicle Visualization) is implemented through different units, including: a multi-dimensional real-vehicle image acquisition and preprocessing unit, used to acquire parking space marking images using a high-resolution camera according to a preset strategy, perform noise reduction and white balance adjustment processing, and output raw image data that meets the requirements for database construction; a marking semantic segmentation and feature extraction unit, which receives the image data output by the acquisition unit, extracts the marking binary mask using the U-Net segmentation model, obtains the geometric topology through contour tracking and skeletonization algorithms, and transmits the processing results to the database construction unit; and a multi-class, multi-quality marking database storage and retrieval unit, which receives the processing results from the feature extraction unit, stores marking materials in a three-level structure of category-quality-instance, and responds to the detection of the enhancement algorithm unit. The system requests and outputs matching parking space marking materials; the AVM parking space enhancement algorithm operation unit receives output data from the basic AVM data unit, semantic annotation unit, and database retrieval unit, respectively, and performs marking removal, virtual parking space generation, and marking fusion operations, transmitting intermediate operation results to the annotation synchronization generation unit; the pixel-level semantic annotation synchronization generation unit receives intermediate results from the enhancement algorithm unit, synchronously updates the semantic annotation map, adds interference element markers, and generates final annotation data; the enhancement data output and format conversion unit receives enhanced image data from the enhancement algorithm unit and annotation data from the annotation generation unit, converts the data into a format that meets the model training requirements, and outputs it to external storage or the training system. Different units interact and collaborate through a data bus.

[0025] Beneficial Effects: This invention proposes a method for enhancing visual perception data of AVM parking spaces. By constructing a multi-class, multi-quality marking database, classifying it by color, line type, and width, and combining it with quantitative indicators to label quality levels, it solves the problems of limited data coverage and high labeling costs in existing methods, providing rich material support for data enhancement. The method employs a dynamic scene anti-interference detection algorithm to optimize the selection of background texture sample areas, combining local neighborhood pixel features with differences in background samples to improve anti-interference capability. This effectively avoids the problems of obvious edge differences and poor visual consistency after base image restoration, accurately simulating real-world unmarked scenes and compensating for the shortcomings of existing technologies in handling dynamic scene interference. Through multi-sensor... By fusing a positioning model to accurately calculate perspective transformation parameters and adapting the spatial location of parking marking materials, and combining this with the ViT parking space feature transfer model to optimize semantic annotation mapping accuracy, the system also uses a visual-inertial fusion mapping engine to assist in calibrating the spatial coordinates of the parking marking database. This ensures that the enhanced data has advantages in parking space structure matching, parking marking feature consistency, and semantic annotation accuracy. This solves the problem of existing technologies lacking efficient model support in the parking marking material fusion and semantic annotation mapping stages. The resulting large-scale, multi-scenario, high-quality enhanced data can effectively support the improvement of the accuracy and optimization of the generalization ability of the AVM parking space perception model, meeting the needs of autonomous driving intelligent parking systems for parking space visual data. Attached Figure Description

[0026] Figure 1 This is a flowchart of the method steps of the present invention;

[0027] Figure 2 This is a diagram showing the unit composition for implementing the method of the present invention. Detailed Implementation

[0028] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] like Figure 1 As shown, a method for enhancing visual perception data of AVM parking spaces includes the following steps:

[0030] S1, collect basic AVM parking space data and corresponding pixel-level semantic annotations. The basic AVM parking space data is real scene images obtained by the vehicle-mounted surround view camera. The semantic annotations include semantic segmentation information of parking space markings.

[0031] Specifically, step S1 is implemented as follows: Basic AVM parking space data in a real parking lot scene is collected using an in-vehicle surround-view camera. During collection, the camera resolution is ensured to be no less than 1920×1080 pixels, and the frame rate is maintained at 30fps to obtain clear and continuous parking space images. Simultaneously, pixel-level semantic annotation is performed on each frame of the collected image. The annotation content must accurately cover the semantic segmentation information of the parking space markings, specifically including the pixel range definition of different types of markings such as parking space lines, parking guide lines, and no-parking zone lines. The annotation accuracy must be controlled within a single pixel error to ensure accurate identification and positioning of various marking areas in subsequent steps, providing accurate raw data and semantic reference for the subsequent data augmentation process. This step requires the collection and annotation of at least 1000 frames of images from different parking lot scenes and under different lighting conditions to ensure the diversity and representativeness of the basic data.

[0032] S2. Construct a multi-category, multi-quality parking marking database. Use a high-resolution camera to capture images of parking markings in a real parking lot according to a category coverage and quality grading strategy. Perform semantic segmentation on the images to extract binary masks of the markings. Obtain geometric topology and sub-pixel level edges through contour tracking and skeletonization. Extract marking areas to generate materials with transparent channels. Classify by color, line type, and width. Combine pixel intensity contrast, edge sharpness, texture continuity quantitative indicators, and subjective features such as wear, dirt, redraw, and peeling to label 10 quality levels, thus completing the database construction.

[0033] Specifically, step S2 is implemented as follows: A high-resolution camera with 4K resolution is used to capture images of parking space markings in a real parking lot according to a coverage and quality grading strategy. The capture range must include different line types (straight, curved, and polygonal), different colors (white, yellow, blue), and different widths (10cm, 15cm, 20cm). The captured images are first processed through semantic segmentation to extract the binary mask of the markings. Then, a contour tracking algorithm is used to obtain the complete contour of the markings. Finally, a skeletonization algorithm is used to extract the geometric topology and sub-pixel edges of the markings. The edge positioning accuracy needs to be... The resolution reaches 0.1 pixels; then, the marking area is extracted to generate materials with alpha channels, which are classified and stored according to color, line type, and width; at the same time, combined with quantitative indicators such as pixel intensity contrast (value range 0-255), edge sharpness (calculated by gradient operator, threshold set to 50), texture continuity (measured by gray-level co-occurrence matrix, correlation threshold set to 0.8), as well as subjective features such as wear, dirt, redrawing, and peeling, the marking materials are labeled with quality levels 1-10, and finally a multi-type, multi-quality marking database including at least 5,000 different categories and quality levels of marking materials is constructed.

[0034] S3 loads basic AVM parking space data, semantic annotations, and enhanced configuration parameters. The configuration parameters include the target parking space structure geometry parameters, marking attributes, interference modes, and random seeds.

[0035] Specifically, the implementation process of step S3 is as follows: First, load the basic AVM parking space data collected in step S1, the corresponding pixel-level semantic annotations, and the pre-set enhancement configuration parameters; among them, the target parking space structural geometric parameters in the enhancement configuration parameters specifically include parking space length (range 4.5-5.5 meters), parking space width (range 2.2-2.5 meters), and arc radius (range 1.5-2.0 meters if it is an arc-shaped parking space); the marking attribute parameters include marking color (white, yellow, blue), line type (straight line, arc, broken line), and width (10cm, 15cm, 20cm); the interference mode parameters include shadow interference (…). The loading process includes several parameters: coverage area ratio (5%-20%), light spot interference (brightness value (200-255)), and occlusion interference (occlusion type: stone pillars, roadblocks, occlusion area ratio (10%-30%). The random seed parameter value range is 0-10000. During loading, the basic AVM parking space data must be checked for integrity to ensure no missing frames or damaged images. The semantic annotation data must be checked for format to ensure that the annotation information corresponds one-to-one with the image pixels. The enhanced configuration parameters must be checked for rationality to ensure that the values ​​of each parameter are within the preset reasonable range. If any data abnormality occurs, an error will be reported in time and the loading process will be terminated until all data meets the requirements and the loading is completed.

[0036] S4. Based on semantic annotation, locate the original parking space marking area. Use an image inpainting algorithm to take the 5-10 pixel outward ring area of ​​the marking edge as the background texture sample area, analyze the color distribution and texture pattern, and eliminate the original marking to generate a marking-free background image.

[0037] Specifically, the implementation process of step S4 is as follows: Based on the semantic annotation data loaded in step S3, the parking space marking area in the original image is accurately located by traversing the image pixels and defining the pixel range of various markings in the semantic annotation. The positioning error needs to be controlled within 2 pixels. Then, the fast traversal method is used as the core of the image inpainting algorithm. The annular area formed by expanding 5-10 pixels outward from the edge of the marking is used as the background texture sample area. The width of this annular area can be adaptively adjusted according to the width of the marking. When the width of the marking is 10cm, it expands outward by 5 pixels, and when the width of the marking is 20cm, it expands outward by 10 pixels. Within the background texture sample area, the sliding window method is used (the window size is set to 3×3). The background texture features are statistically analyzed to determine the color distribution of pixels (represented by histograms of the RGB three-color channels, with the number of histogram bins set to 256) and texture mode (calculated using a gray-level co-occurrence matrix, with a step size of 1 and angles of 0°, 45°, 90°, and 135°). Based on the statistically obtained background texture features, pixel filling calculations are performed on the located parking space marking areas. During the filling process, it is necessary to ensure that the color transition difference between adjacent pixels does not exceed 10, the texture coherence matches the texture mode of the sample area, and finally, the pixel information of the original marking area is eliminated to generate a marking-free background image. The background image has the same resolution as the original image and no obvious repair traces. The visual consistency error needs to be controlled within 5%.

[0038] S5: Retrieve matching marking materials from the database according to the configuration parameters, determine the virtual parking space layout by combining the parking space edge position of the original image, and fill the marking materials into the corresponding area of ​​the base image after perspective transformation and lighting adaptation.

[0039] Specifically, the implementation process of step S5 is as follows: Based on the target parking space structure geometric parameters and marking attribute parameters in the enhanced configuration parameters loaded in step S3, a multi-condition search is performed in the multi-type, multi-quality marking database constructed in step S2. During the search, a subset of materials that meet the attribute requirements are first filtered according to marking color, line type, and width. Then, the subset is further filtered according to the quality level parameters (if the configuration parameters require high-quality markings, level 8-10 is selected; if low-quality markings are required, level 1-3 is selected) to finally determine the matching marking materials. At the same time, the position information of the parking space edge lines in the original image of the unmarked base map generated in step S4 is extracted. The pixel coordinates of the edge lines are determined by the edge detection algorithm (using the Canny operator, with a threshold set to 50-150). Combined with the target parking space structure geometric parameters, a polygon fitting algorithm is used to determine the arrangement of virtual parking spaces (parallel, vertical, or diagonal). The process involves several steps: First, the retrieved marking materials are processed using perspective transformation. The perspective projection matrix is ​​calculated based on the intrinsic (focal length, principal point coordinates) and extrinsic (installation height, horizontal angle) parameters of the AVM camera. This transforms the marking materials from the original coordinate system to the base map coordinate system, ensuring that the spatial position of the materials in the base map matches the virtual parking space outline. Next, the light intensity distribution around the virtual parking space filling area in the base map is analyzed (using the average brightness value, with a sampling range 10 pixels outside the filling area). A grayscale adjustment algorithm is used to adjust the brightness value of the marking materials to a value no greater than 20% of the average brightness value of the surrounding area, and the contrast to a value no greater than 15%, completing the lighting adaptation. Finally, the processed marking materials are filled into the corresponding virtual parking space area in the base map. During filling, the pixel alignment accuracy between the materials and the base map must be ensured, with an edge overlap of no less than 95%.

[0040] S6 synchronously updates the semantic annotation map, maps the line-filled areas to corresponding semantic labels, adds interference element markers, and generates an enhanced AVM image and corresponding semantic annotations.

[0041] Specifically, the implementation process of step S6 is as follows: After the marking material is filled in step S5, the corresponding semantic annotation map is updated synchronously; first, the pixel coordinate range of the filling area of ​​the marking material in the base map is determined, and the semantic labels of the pixels in this range are updated to the labels of the corresponding type of marking (e.g., the label of parking space line is set to 1, the label of parking guide line is set to 2, and the label of no-parking area line is set to 3). The label update accuracy must reach 100%; then, according to the interference mode parameters in the enhanced configuration parameters, the corresponding interference element markers are added to the semantic annotation map, such as the shadow interference area is marked as 4, the light spot interference area is marked as 5, and the occlusion interference area is marked as 6. The marking range must be completely consistent with the pixel coordinates of the interference area actually added in the base map; During the update process, a pixel-by-pixel comparison is required to verify the consistency of semantic information between the labeled image and the corresponding area in the base image. If any inconsistency is found, it should be corrected promptly. After the semantic annotation update is completed, the base image with filled lines generated in step S5 is integrated with the updated semantic annotation image to form the final enhanced AVM image and corresponding semantic annotations. Finally, the quality of the generated enhanced data is checked, including image sharpness (measured by peak signal-to-noise ratio, with a threshold of 30dB) and semantic annotation accuracy (100 pixels are randomly selected for verification, and the accuracy rate must reach 100%). If the check passes, the data is stored; if it fails, it is returned to the corresponding step for reprocessing until enhanced data that meets the quality requirements is generated.

[0042] Preferably, in step S4, a dynamic scene anti-disturbance detection algorithm is used to optimize the selection of the background texture sample area. The algorithm expression is as follows:

[0043] ,

[0044] in, For pixels The anti-interference detection value, This represents the number of pixels in the local neighborhood. These are the weighting coefficients. The grayscale values ​​of the original image pixels. The standard deviation of local gray levels in the original image. These are the pixel values ​​after Gaussian filtering. This represents the local standard deviation after filtering. Background weight coefficient, The number of background sample pixels. The background sample pixel values, The local standard deviation is the background sample.

[0045] Specifically, the dynamic scene anti-interference detection algorithm is used to optimize the selection of background texture sample areas in step S4. During implementation, the calculation logic of pixel anti-interference detection values ​​is first determined. The number of local neighboring pixels is set to 9-25 (corresponding to a 3×3 to 5×5 pixel window) to ensure sufficient coverage of surrounding pixel features. The weight coefficients ω1 and ω2 are set to 0.6 and 0.4 respectively, giving priority to the impact of differences in pixel grayscale values ​​in the original image on the detection results. The Gaussian filter uses a 5×5 window size and the filter standard deviation is set to 1.2 to balance the smoothing effect and detail preservation. The background weight coefficient λ is set to 0.3 to avoid excessive interference from background sample differences in detection. The number of background sample pixels is set to 36-100 (corresponding to a 6×6 to 10×10 pixel window) to ensure the reliability of background feature statistics. During the calculation process, the grayscale difference ratio between the original image pixels and the filtered pixels is calculated first, and then the weighted value of the background sample pixel difference is subtracted to obtain the anti-interference detection value of each pixel. Pixels with a detection value higher than 0.8 are judged as high-quality background sample areas, and pixels with a detection value lower than 0.3 are excluded. In this way, the background texture area without interference is accurately screened out, providing high-quality samples for subsequent image restoration, reducing the edge difference of the restored base image, and improving visual consistency.

[0046] Preferably, in step S5, a multi-sensor fusion positioning model is used to calculate the perspective transformation parameters, and the model expression is:

[0047]

[0048] The result of image coordinate transformation. These are the parameters of the perspective transformation matrix. The fusion coefficient is... For the number of sensors, For sensor weights, To locate the coordinates of the lidar, To locate coordinates for millimeter-wave radar, These are the coordinates of the original image.

[0049] Specifically, the multi-sensor fusion positioning model is used to calculate the perspective transformation parameters in step S5. During implementation, it needs to integrate data from at least two types of sensors, such as LiDAR and millimeter-wave radar. The perspective transformation matrix parameters are initially obtained through camera calibration and then dynamically adjusted based on the sensor data. The fusion coefficient γ is set to 0.2-0.4 to control the correction magnitude of the transformation result by the sensor data. The number of sensors K is set to 2-4 based on the actual vehicle configuration (e.g., LiDAR + millimeter-wave radar + ultrasonic radar), and the weight α of each sensor is... k β kBased on accuracy allocation, the weight of the LiDAR is set to 0.5-0.7, and that of the millimeter-wave radar is set to 0.3-0.5. During calculation, the original image coordinates are first transformed using a perspective transformation matrix, and then a weighted sum of the positioning coordinates from each sensor is superimposed. The positioning coordinate accuracy of the LiDAR is controlled within ±3cm, and that of the millimeter-wave radar within ±5cm, ensuring that the transformed image coordinates accurately match the real spatial position. This model, through multi-sensor data complementarity, compensates for the positioning errors of a single sensor in scenarios with occlusion and changes in lighting, making the spatial position of the lane marking material more accurate after perspective transformation and improving the adaptation between the virtual parking space and the base map.

[0050] Preferably, in step S6, the ViT parking space feature transfer model is used to optimize the semantic annotation mapping accuracy. The model expression is:

[0051] ,

[0052] in, For pixels Feature mapping values, For feature dimension, This represents the number of layers in the Transformer encoder. , These are query, key, and value matrices, respectively. For migration coefficient, For the number of feature neighborhoods, For the source domain eigenvalues, For the feature values ​​of the target domain, It is an L2 norm.

[0053] Specifically, the ViT parking space feature transfer model is used to optimize the semantic annotation mapping accuracy in step S6. During implementation, the feature dimension is set to 256-512 dimensions to ensure sufficient representation of parking space marking features. The Transformer encoder has 6-12 layers to balance feature extraction capability and computational efficiency. The transfer coefficient δ is set to 0.1-0.3 to avoid excessive influence of source and target domain feature differences on the mapping result. The number of feature neighbors is set to 8-16 to cover sufficient surrounding features for difference calculation. During the calculation, the basic mapping values ​​of pixel features are first obtained through query, key, and value matrix operations of the Transformer encoder, then normalized using the Softmax function, and finally superimposed with weighted values ​​of the source and target domain feature differences (the difference calculation uses the L2 norm to ensure the non-negativity and reasonableness of the difference values) to obtain the final feature mapping value. The matching degree between the feature mapping value and the semantic label needs to reach above 95%. If it is below 90%, the model parameters are readjusted to improve the mapping accuracy of semantic annotation from the original image to the enhanced image, ensuring that the enhanced annotation accurately corresponds to the filled marking area and reducing annotation errors.

[0054] Preferably, in step S2, a visual-inertial fusion mapping engine is used to assist in the spatial coordinate calibration of the marking database. The engine expression is:

[0055] ,

[0056] in, For calibrated spatial coordinates, For the number of visual frames, For visual and inertial weights, For visual camera coordinates, For inertial measurement unit coordinates, To optimize the coefficients, For the number of observation data, Let be the error function. These are the observed values.

[0057] Specifically, in the spatial coordinate calibration of the map line database in step S2, which is assisted by the visual-inertial fusion mapping engine, the number of visual frames M is set to 30-60 frames (corresponding to 1-2 seconds of continuous data acquisition) to ensure coverage of spatial changes within a short period of time; the visual and inertial weights μ m ν m The values ​​are set to 0.6 and 0.4 respectively, prioritizing the reference to the visual camera coordinates; the inertial measurement unit (IMU) sampling frequency is set to 100Hz to ensure continuous motion capture; the optimization coefficient ξ is set to 0.15-0.25 to control the magnitude of coordinate correction due to error optimization; the number of observation data N is set to 50-100, including visual feature points, IMU angular velocity, and acceleration data. During calculation, the visual camera coordinates and IMU coordinates are first weighted and fused to obtain the initial spatial coordinates. Then, the coordinate correction is calculated using error functions (such as reprojection error and IMU pre-integration error). The error function value must be controlled within 0.5 pixels (visual error) and within 0.1 rad (IMU angular error), iteratively optimizing the spatial coordinates. This engine, through the fusion of visual and inertial data, corrects the coordinate drift of monocular vision in scale and motion blur scenarios, ensuring the spatial coordinate accuracy of the materials in the grading database is controlled within ±2cm, providing accurate spatial reference for subsequent grading material filling.

[0058] Preferably, in step S5, a lane marking fill intensity adjustment model is constructed by combining the AVM parking space visual perception data enhancement parameters, and the model expression is:

[0059] ,

[0060] in, For the strength of the marking fill, These are the maximum and minimum quality levels, respectively. For the current quality level, For feature weights, For the color feature value of the grading line, For the texture feature value of the gradation line, For adjustment coefficients, For feature dimension, For the first 3D eigenvalues The average characteristic value.

[0061] Specifically, the AVM parking space visual perception data enhancement parameter related model is used to adjust the marking fill intensity in step S5, with a maximum quality level q during implementation. max Set to 10, minimum quality level q min Set to 1, consistent with the database quality grading in step S2; feature weights θ1 and θ2 are set to 0.5 and 0.5 respectively, balancing the influence of color and texture features on fill intensity; adjustment coefficient φ is set to 0.1-0.2, controlling the adjustment range of feature differences on intensity; feature dimension R is set to 6-10 (such as color mean, color variance, texture energy, texture entropy, etc.). During calculation, first determine the base intensity coefficient based on the proportion of the difference between the current quality level and the maximum and minimum quality levels of the datum line. The base coefficient is 1 when the quality level is 10, and 0.3 when the quality level is 1; then, combine the datum line color feature values ​​(such as the RGB channel mean, with values ​​from 0-255) and texture feature values ​​(such as gray-level co-occurrence matrix energy, with values ​​from 0-1) to calculate the feature contribution value. Finally, superimpose the proportion of the difference between each dimension feature and the average feature to obtain the final fill intensity. The fill strength ranges from 0.3 to 1.0. The strength value is achieved by adjusting the transparency of the marking material (0.3 corresponds to 70% transparency, and 1.0 corresponds to complete opacity) or the grayscale value, ensuring that high-quality markings are clearer after filling, while low-quality markings appear worn and faded, thus enhancing the realism of the data.

[0062] Preferably, step S3 includes the following sub-steps: S31, when reading the basic AVM parking space data, the image resolution consistency is checked, and images of different resolutions are unified to a preset pixel size using a bilinear interpolation algorithm to ensure pixel-level alignment accuracy between the image and semantic annotation in subsequent processing; S32, the enhancement configuration parameters are parsed, and the target parking space structure parameters are converted into geometric parameters in a pixel coordinate system, including parking space length, width, arc radius, etc., and a mapping relationship between parameters and image coordinates is established; S33, a random number generator is initialized according to a random seed to determine the random selection sequence of samples during database retrieval, ensuring a balance between repeatability and randomness in sample selection during each enhancement process; S34, the semantic annotation data is format-converted, and the original annotation format is unified into a standard JSON format including the marking category ID and pixel coordinate range, which facilitates the location and processing of the marking area in subsequent steps.

[0063] Specifically, step S3 is executed in four steps. First, in S31, when reading the basic AVM parking space data, the image resolution is checked for consistency. If the image resolution is lower than the preset 1920×1080 pixels, a bilinear interpolation algorithm is used for magnification. The neighboring pixel selection range is set to 3×3 during interpolation to ensure a smooth transition of image pixel values ​​after interpolation. If the resolution is higher than the preset value, it is scaled down proportionally to two decimal places. Finally, all images are unified to a size of 1920×1080 pixels to ensure pixel-level alignment between the image and semantic annotation in subsequent processing, with the alignment error controlled within 1 pixel. In S32, the enhanced configuration parameters are parsed, and the physical parameters of the target parking space structure (such as length 4.5-5.5 meters, width 2.2-2.5 meters) are converted according to the conversion ratio between camera pixels and actual distance (1 pixel corresponds to 0.002 meters). For geometric parameters in the pixel coordinate system, establish a mapping table between parameters and image coordinates. The mapping table must include the pixel range value corresponding to each physical parameter. In S33, initialize the random number generator according to the random seed (value range 0-10000). The generated random number sequence is used to determine the selection order of materials when searching the database. The random seed is changed every 100 enhancement processes to balance the repeatability and randomness of sample selection. In S34, convert the semantic annotation data into a standard JSON format, converting the original XML or TXT format annotation file into a standard JSON format. The JSON file must clearly specify the annotation line category ID (e.g., parking line ID is 1, guide line ID is 2) and the corresponding pixel coordinate range (represented by the pixel coordinates of the upper left and lower right corners). After conversion, the number of pixels in each annotation area must be verified to ensure that there are no missing or incorrectly labeled areas. The verification pass rate must reach 100%.

[0064] Preferably, step S4 includes the following sub-steps: S41, based on the marking category ID in the semantic annotation, traverse the image pixels to determine the smallest bounding rectangle of all parking space marking areas, and use this as the initial area range for marking removal; S42, based on the initial area range, expand outward by 5-10 pixels to form a ring-shaped background texture sample area, and use the sliding window method to statistically analyze the gray-level histogram and texture feature matrix of the pixels in the sample area; S43, select the fast traversal method as the core of the image restoration algorithm, and use the gray-level and texture features of the sample area as the basis to perform filling calculations on the marking pixels in the initial area to generate a preliminary marking-free base map; S44, perform edge smoothing processing on the preliminary base map, and use the Gaussian filtering algorithm to eliminate the boundary differences between the restoration area and the non-restoration area to ensure the overall visual consistency of the base map.

[0065] Specifically, in the implementation process of step S4, the four sub-steps need to be closely linked. In S41, based on the marking category ID in the semantic annotation, all pixels in the image are traversed. The parking space marking area is determined by comparing the annotation value of each pixel. Then, the bounding box algorithm is used to calculate the minimum bounding rectangle of the area. The coordinates of the rectangle boundary are accurate to integer pixels, which is used as the initial area range for marking elimination. The initial area range must completely cover all marking pixels without omission. In S42, based on the initial area range, it is expanded outward by 5-10 pixels to form a ring-shaped background texture sample area. The number of expanded pixels is determined according to the marking width (5 pixels when the marking width is 10cm, and 10 pixels when the marking width is 20cm). Then, a 3×3 sliding window method is used to traverse the sample area and count the pixels in each window. The grayscale histogram (with 256 histogram bins) and texture feature matrix (including 6 feature values ​​such as energy, entropy, and contrast) are statistically analyzed and stored in matrix form. In S43, the fast traversal method is selected as the core of the image inpainting algorithm. During inpainting, the grayscale and texture features statistically analyzed in the sample area are used as the basis to iteratively calculate the fill value of each pixel in the initial area. The number of iterations is set to 50-80 times until the difference between the pixel value of the filled area and the pixel value of the surrounding background is less than 5, generating a preliminary unmarked base map. In S44, the preliminary base map is subjected to edge smoothing processing using a 5×5 Gaussian filtering algorithm with a filtering standard deviation of 1.2. After filtering, the boundary pixel values ​​of the inpainted area and the non-inpainted area are checked. The difference in pixel values ​​at the boundary needs to be controlled within 8 to ensure the overall visual consistency of the base map.

[0066] Preferably, step S5 includes the following sub-steps: S51, based on the marking attributes in the configuration parameters, perform multi-condition retrieval in a multi-type, multi-quality marking database to select a set of marking materials that meet the requirements of color, line type, and quality level; S52, extract the parking space edge position information from the original image, combine it with the target parking space structure parameters, and use a polygon fitting algorithm to determine the contour coordinates of the virtual parking space, dividing the specific area for marking filling; S53, perform perspective transformation processing on the selected marking materials, calculate the perspective projection matrix of the materials in the virtual parking space filling area based on the intrinsic and extrinsic parameters of the AVM camera, and complete the spatial position adaptation of the materials; S54, analyze the light intensity distribution around the filling area of ​​the original image, adjust the brightness and contrast of the marking materials through a grayscale adjustment algorithm to ensure that the lighting conditions of the materials are consistent with those of the base image, and then fill the corresponding area with the processed materials.

[0067] Specifically, in the implementation process of step S5, the four sub-steps need to be carried out in an orderly manner. In S51, based on the marking attributes (color, line type, quality level) in the configuration parameters, a multi-condition search is performed in a multi-category, multi-quality marking database. During the search, the data is first filtered by color (white, yellow, blue), then further filtered by line type (straight line, arc, polyline), and finally determined by quality level (levels 1-10) to determine the set of marking materials that meet the requirements. The number of materials in each category in the material set should not be less than 50 to ensure sufficient margin for subsequent selection. In S52, the parking space edge position information of the original image is extracted. The Canny edge detection algorithm (threshold set to 50-150) is used to detect the edge pixels, and then the edge pixels are fitted by a polygon fitting algorithm. The allowable error range during fitting is set to 2 pixels to determine the contour coordinates of the virtual parking space. The contour coordinates must include the precise pixel value of each vertex to divide the specific area for marking filling. In S53, the selected marking elements are processed... The material undergoes perspective transformation. Based on the AVM camera's intrinsic parameters (focal length 1200-1500 pixels, principal point coordinates (960, 540)) and extrinsic parameters (installation height 1.5-1.8 meters, horizontal angle 0-5 degrees), the perspective projection matrix is ​​calculated, with matrix elements accurate to four decimal places. The material is then transformed from the original coordinate system to the base map coordinate system, and the positional deviation of the transformed material must be controlled within 3 pixels. In S54, the light intensity distribution around the filling area of ​​the original image is analyzed. A sampling range of 10 pixels is extended outward from the filling area as the center, and the average brightness value of the pixels within this range (range 0-255) is calculated. The brightness value of the marking material is adjusted to no more than 20 from the average brightness value through a grayscale adjustment algorithm, and the contrast is adjusted to no more than 15 from the contrast of the surrounding area (range 0-100). After completing the lighting adaptation, the material is filled into the corresponding area, ensuring that the overlap between the edge of the material and the boundary of the filling area is no less than 95%.

[0068] The dynamic scene anti-interference detection algorithm is used in this invention to optimize the selection of background texture sample areas during the AVM parking space visual perception data enhancement process. It eliminates the influence of dynamic interference factors on background samples and improves image restoration accuracy. Its implementation process requires first determining the number of local neighboring pixels (9-25, corresponding to a 3×3 to 5×5 pixel window), weight coefficients (ω1=0.6, ω2=0.4), Gaussian filter parameters (5×5 window, standard deviation 1.2), background weight coefficient (λ=0.3), and the number of background sample pixels (36-100). Then, the anti-interference detection value of each pixel is calculated: first, the grayscale difference ratio between the original image pixels and the filtered pixels reflects the local texture stability; then, the weighted value of the background sample pixel difference is subtracted to eliminate background interference; finally, pixels with a detection value higher than 0.8 are judged as high-quality background sample areas. This algorithm accurately selects background texture areas free from interference, providing high-quality samples for image restoration in step S4. It avoids problems such as obvious edge differences and poor visual consistency in the restored base map, and solves the defect of inaccurate background sample selection in dynamic scenes in the prior art. It ensures that the generated caliper-free base map can realistically simulate the actual scene, laying a high-quality foundation for subsequent caliper material filling and improving the authenticity and reliability of the augmented data.

[0069] The multi-sensor fusion localization model is used to calculate the perspective transformation parameters of AVM parking space visual perception data enhancement marking material. It improves spatial positioning accuracy by integrating data from multiple sensors. The implementation first determines the initial value of the perspective transformation matrix (obtained through camera calibration), the fusion coefficient (γ=0.2-0.4), the number of sensors (2-4, such as LiDAR and millimeter-wave radar), and the weight of each sensor (LiDAR 0.5-0.7, millimeter-wave radar 0.3-0.5). Then, it combines sensor positioning data (LiDAR accuracy ±3cm, millimeter-wave radar ±5cm) for calculation: first, the original image coordinates are transformed using the perspective transformation matrix, then the weighted sum of the positioning coordinates from each sensor is superimposed, dynamically correcting the transformation result. This model provides precise parameters for the perspective transformation of the marking materials in step S5, ensuring that the materials can accurately match the virtual parking space outline after being transformed from the original coordinate system to the base map coordinate system, with the spatial position deviation controlled within 3 pixels. This compensates for the positioning error of a single sensor in scenarios with occlusion and changes in lighting, solves the problem of low spatial adaptability of marking materials in existing technologies, ensures the consistency between the enhanced parking space structure and the real scene, and improves the effectiveness of the AVM parking space perception model training data.

[0070] The ViT parking space feature transfer model is a feature mapping optimization model based on the Transformer architecture, used to improve the accuracy of semantic annotation mapping in AVM parking space visual perception data augmentation. The implementation process requires setting the feature dimension (256-512 dimensions), the number of Transformer encoder layers (6-12 layers), the transfer coefficient (δ=0.1-0.3), and the number of feature neighbors (8-16). During calculation, pixel features are first extracted through the encoder's query, key, and value matrix operations. After Softmax normalization, a basic mapping value is obtained. Then, a weighted value of the difference between the source and target domain features is superimposed (using the L2 norm to calculate the difference), finally obtaining the feature mapping value. The matching degree between the mapping value and the semantic label must reach over 95%. In step S6, the model optimizes the semantic annotation mapping process from the original image to the augmented image, ensuring that the augmented annotation accurately corresponds to the filled marking area, avoiding annotation errors, and solving the problem of low semantic annotation mapping accuracy in existing technologies. This guarantees the accuracy of the semantic information in the augmented data, provides high-quality labeled data for the AVM parking space perception model, helps the model accurately learn parking space features, and improves the model's recognition accuracy and generalization ability.

[0071] The visual-inertial fusion mapping engine is the core engine that integrates data from visual cameras and inertial measurement units (IMUs) for spatial coordinate calibration of the AVM parking space marking database. Implementation requires determining the number of visual frames (30-60 frames) and the visual and inertial weights (μ). m =0.6、ν m The calculation process involves several parameters: ξ = 0.4, IMU sampling frequency (100Hz), optimization coefficients (ξ = 0.15-0.25), and the number of observation data (50-100). First, the visual camera coordinates and IMU coordinates are weighted and fused to obtain initial spatial coordinates. Then, a correction is calculated using an error function (reprojection error < 0.5 pixels, IMU angle error < 0.1 rad), iteratively optimizing the coordinates. In step S2, the engine corrects the spatial coordinates of the marking database materials, controlling the coordinate accuracy to ±2cm. This provides a precise spatial reference for subsequent marking material filling, solving the coordinate drift problem in monocular vision under scale and motion blur scenarios. It ensures the spatial accuracy of the materials in the database, avoids misalignment of marking fill due to coordinate deviations, improves the spatial consistency of the augmented data, and provides training data that conforms to the real spatial scale for the AVM parking space perception model.

[0072] like Figure 2As shown, an AVM (Autonomous Vehicle Visual Perception) parking space data enhancement method is implemented through different units, including: a multi-dimensional real vehicle image acquisition and preprocessing unit, used to acquire parking space marking images using a high-resolution camera according to a preset strategy, perform noise reduction and white balance adjustment processing, and output raw image data that meets the database construction requirements; a marking semantic segmentation and feature extraction unit, which receives the image data output by the acquisition unit, extracts the marking binary mask using the U-Net segmentation model, obtains the geometric topology through contour tracking and skeletonization algorithms, and transmits the processing results to the database construction unit; a multi-class, multi-quality marking database storage and retrieval unit, which receives the processing results from the feature extraction unit, stores marking materials in a three-level structure of category-quality-instance, and responds to the enhancement algorithm unit. The system receives retrieval requests and outputs matching lane marking materials. The AVM parking space enhancement algorithm operation unit receives output data from the basic AVM data unit, semantic annotation unit, and database retrieval unit, respectively, and performs lane marking removal, virtual parking space generation, and lane marking fusion operations, transmitting intermediate operation results to the annotation synchronization generation unit. The pixel-level semantic annotation synchronization generation unit receives intermediate results from the enhancement algorithm unit, synchronously updates the semantic annotation map, adds interference element markers, and generates final annotation data. The enhancement data output and format conversion unit receives enhanced image data from the enhancement algorithm unit and annotation data from the annotation generation unit, converts the data into a format that meets the model training requirements, and outputs it to external storage or the training system. Different units interact and collaborate through a data bus.

[0073] A visual perception data enhancement method for parking spaces using AVM (Autonomous Visualization) is proposed. This method classifies parking spaces by color, line type, and width, and combines quantitative indicators such as pixel intensity contrast and edge sharpness with subjective features like wear and dirt to label quality levels. This constructs a multi-category, multi-quality marking database, enabling refined management of materials and allowing for on-demand matching of different scene requirements. It effectively solves the problems of limited data coverage and high labeling costs in existing technologies, providing ample and diverse material support for data enhancement. Furthermore, when generating the base map by removing parking space markings from the original image, a dynamic scene anti-interference detection algorithm is used to optimize the selection of background texture sample areas. This, combined with the differences between local neighboring pixel features and background samples, improves anti-interference capabilities, avoiding significant edge differences and poor visual consistency after base map restoration. It accurately simulates real-world parking space-without-marking scenes, overcoming the shortcomings of existing technologies in handling dynamic scene interference.

[0074] This method accurately calculates perspective transformation parameters through a multi-sensor fusion positioning model, ensuring the spatial adaptability of lane marking materials within the virtual parking space filling area. It also optimizes semantic annotation mapping accuracy using the ViT parking space feature transfer model, making the semantic information of the enhanced data more accurate. Simultaneously, a visual-inertial fusion mapping engine assists in calibrating the spatial coordinates of the lane marking database, further improving the accuracy of material spatial positioning. The synergistic application of these models solves the problem of insufficient efficient model support in the lane marking material fusion and semantic annotation mapping stages of existing technologies. This ensures that the enhanced data has significant advantages in parking space structure matching, lane marking feature consistency, and semantic annotation accuracy. The resulting large-scale, multi-scenario, high-quality enhanced data better supports the improvement of AVM parking space perception model accuracy and optimization of generalization capabilities, meeting the needs of autonomous driving intelligent parking systems.

[0075] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0076] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for enhancing visual perception data of AVM parking spaces, characterized in that, Includes the following steps: S1. Collect basic AVM parking space data and corresponding pixel-level semantic annotations. The basic AVM parking space data consists of real-scene images obtained from vehicle-mounted surround-view cameras. The semantic annotations include semantic segmentation information of parking space markings. S2. Construct a multi-category, multi-quality parking space marking database. Parking space marking images are captured in real parking lots using a high-resolution camera, following a category coverage and quality grading strategy. Semantic segmentation is performed on the images to extract the marking binary mask. Contour tracking and skeletonization are used to obtain the geometric topology and sub-pixel-level edges. Marking areas are extracted to generate materials with alpha channels. These materials are categorized by color, line type, and width. Ten quality levels are assigned based on pixel intensity contrast, edge sharpness, texture continuity quantification indicators, and subjective features such as wear, dirt, redrawing, and peeling. Database construction is then completed. S3. Load... The process involves: S4, S5, S6, S7, S8, S9, S10, S20, S11, S20, S30, S40, S60, S70, S80, S90, S110, S20, S60, S120, S60, S130, S60, S140, S60, S150, S60, S160, S60, S160, S60, S170, S60, S180, S60, S180, S60, S180, S60, S180, S60, S180, S60, S180, S60, S180, S60, S180, S60, S180, S180, S60, S180, S180, S190, S200, S180, S190, S200, S180, S190, S200, S180, S190, S200, S19 ... In S5, a lane marking fill intensity adjustment model is constructed by combining AVM parking space visual perception data enhancement parameters. The model expression is: , in, For the strength of the marking fill, These are the maximum and minimum quality levels, respectively. For the current quality level, For feature weights, For the color feature value of the grading line, For the texture feature value of the gradation line, For adjustment coefficients, For feature dimension, For the first 3D eigenvalues The average characteristic value; In S6, the ViT parking space feature transfer model is used to optimize the semantic annotation mapping accuracy. The model expression is: , in, For pixels Feature mapping values, For feature dimension, This represents the number of layers in the Transformer encoder. , These are query, key, and value matrices, respectively. For migration coefficient, For the number of feature neighborhoods, For the source domain eigenvalues, For the feature values ​​of the target domain, It is an L2 norm.

2. The method for enhancing visual perception data of AVM parking spaces according to claim 1, characterized in that, In step S4, a dynamic scene anti-disturbance detection algorithm is used to optimize the selection of background texture sample areas. The algorithm expression is as follows: , in, For pixels The anti-interference detection value, This represents the number of pixels in the local neighborhood. These are the weighting coefficients. The grayscale values ​​of the original image pixels. The standard deviation of local gray levels in the original image. These are the pixel values ​​after Gaussian filtering. This represents the local standard deviation after filtering. Background weight coefficient, The number of background sample pixels. The background sample pixel values, The local standard deviation is the background sample.

3. The method for enhancing visual perception data of AVM parking spaces according to claim 1, characterized in that, In S5, a multi-sensor fusion positioning model is used to calculate perspective transformation parameters. The model expression is: , The result of image coordinate transformation. These are the parameters of the perspective transformation matrix. The fusion coefficient is... For the number of sensors, For sensor weights, To locate the coordinates of the lidar, To locate coordinates for millimeter-wave radar, These are the coordinates of the original image.

4. The method for enhancing visual perception data of AVM parking spaces according to claim 1, characterized in that, In S2, a visual-inertial fusion mapping engine is used to assist in the spatial coordinate calibration of the datum database. The engine expression is: , in, For calibrated spatial coordinates, For the number of visual frames, For visual and inertial weights, For visual camera coordinates, For inertial measurement unit coordinates, To optimize the coefficients, For the number of observation data, Let be the error function. These are the observed values.

5. The method for enhancing visual perception data of AVM parking spaces according to claim 1, characterized in that, S3 includes the following sub-steps: S31, when reading the basic AVM parking space data, the resolution consistency of the image is checked, and the images of different resolutions are unified to the preset pixel size through bilinear interpolation algorithm to ensure the pixel-level alignment accuracy between the image and the semantic annotation in subsequent processing; S32, the enhancement configuration parameters are parsed, and the target parking space structure parameters are converted into geometric parameters in the pixel coordinate system, including parking space length, width, arc radius, etc., and a mapping relationship between parameters and image coordinates is established; S33, the random number generator is initialized according to the random seed to determine the random selection sequence of samples when searching the database, ensuring a balance between the repeatability and randomness of sample selection in each enhancement process; S34, the semantic annotation data is format converted, and the original annotation format is unified into a standard JSON format including the marking category ID and pixel coordinate range, which facilitates the positioning and processing of the marking area in subsequent steps.

6. The method for enhancing visual perception data of AVM parking spaces according to claim 1, characterized in that, S4 includes the following sub-steps: S41, based on the marking category ID in the semantic annotation, traverse the image pixels to determine the minimum bounding rectangle of all parking space marking areas, and use this as the initial area range for marking elimination; S42, based on the initial area range, expand outward by 5-10 pixels to form a ring-shaped background texture sample area, and use the sliding window method to statistically analyze the gray-level histogram and texture feature matrix of the pixels within the sample area; S4 3. The fast traversal method is selected as the core of the image restoration algorithm. Based on the grayscale and texture features of the sample area, the marking pixels in the initial area are filled to generate a preliminary marking-free base map; S44. The preliminary base map is processed for edge smoothing. The Gaussian filtering algorithm is used to eliminate the boundary difference between the restoration area and the non-restoration area to ensure the overall visual consistency of the base map.

7. The method for enhancing visual perception data of AVM parking spaces according to claim 1, characterized in that, S5 includes the following sub-steps: S51, based on the marking attributes in the configuration parameters, perform multi-condition retrieval in a multi-type, multi-quality marking database to filter out a set of marking materials that meet the requirements of color, line type, and quality level; S52, extract the parking space edge position information from the original image, combine it with the target parking space structure parameters, and use a polygon fitting algorithm to determine the contour coordinates of the virtual parking space, dividing the specific area for marking filling; S53, perform perspective transformation processing on the selected marking materials, calculate the perspective projection matrix of the materials in the virtual parking space filling area based on the intrinsic and extrinsic parameters of the AVM camera, and complete the spatial position adaptation of the materials; S54, analyze the light intensity distribution around the filling area of ​​the original image, adjust the brightness and contrast of the marking materials through a grayscale adjustment algorithm to make the lighting conditions of the materials consistent with those of the base image, and then fill the corresponding area with the processed materials.