A 3D Semantic Occupancy Prediction Method for Parking Space Occupancy Detection in Parking Lots

Through the 3D semantic occupation prediction method, the parking lot scene is reconstructed in combination with depth information and image features, the problem of inaccurate vehicle parking recognition in traditional methods is solved, efficient parking space occupation detection and abnormal recognition is achieved, and the accuracy and safety of parking lot management is improved.

CN119811130BActive Publication Date: 2025-07-18HUNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510287278.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-18
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing parking lot management system relies on two-dimensional image information to accurately judge the parking situation of a vehicle. Especially in the case of vehicle occlusion and light changes, the recognition accuracy is low and it is difficult to distinguish different types of target objects, resulting in insufficient management efficiency and safety.

Method used

The 3D semantic occupancy prediction method is adopted to collect image data through the camera, extract image features and predict depth information, reconstruct the 3D occupancy grid space, perform semantic classification under depth guidance, identify parking space abnormalities, and use the Transformer architecture to enhance the semantic information of the 3D feature map to achieve accurate object recognition and classification.

Benefits of technology

It improves the accuracy of vehicle parking position identification, reduces the impact of occlusion and lighting changes, can automatically identify abnormal parking spaces, improves management efficiency and safety, reduces hardware costs and installation complexity, and reduces false alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811130B_ABST
    Figure CN119811130B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D semantic occupancy prediction method for parking space occupancy detection in a parking lot, belonging to the technical field of vehicle parking detection. The method includes the following steps: collecting image data of the parking lot scene through a camera in the parking lot; extracting image features from the collected image data and predicting the depth information of each pixel point in the image; reconstructing a 3D occupancy grid of the parking lot scene according to the extracted image features and the predicted depth information; in the 3D occupancy grid guided by depth, performing accurate semantic classification / recognition on each voxel in the space, and reconstructing a complex three-dimensional scene with semantic labels including occluded and image-invisible areas; performing parking space occupancy prediction based on the reconstructed three-dimensional scene, and identifying abnormal parking space situations according to the occupancy prediction results. The present invention can effectively solve the problem of insufficient accuracy existing in the 3D reconstruction method based on monocular vision by combining depth information and image features for 3D reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle parking detection, and particularly relates to a 3D semantic occupancy prediction method for detecting vehicle occupancy in a parking lot. Background Art

[0002] Traditional parking lot management systems mainly rely on cameras for monitoring. However, limited by two-dimensional image information, it is difficult to accurately judge the parking situation of vehicles. For example, in cases of vehicle occlusion, light changes, etc., the recognition accuracy is low, which easily causes false alarms or missed alarms, affecting the management efficiency and safety of the parking lot.

[0003] To solve the above problems, researchers have proposed various deep learning-based vehicle detection and recognition methods. For example, object detection algorithms such as YOLO, SSD, and Faster R-CNN can identify vehicles in images, but it is difficult to obtain accurate position and size information of the vehicles. And point cloud processing algorithms such as PointNet and PointNet++ can process three-dimensional point cloud data, but it is difficult to obtain semantic information of the vehicles. Deep learning-based semantic segmentation algorithms can segment semantic categories such as vehicles, ground, and parking spaces in images, but it is difficult to obtain accurate position and size information of the vehicles.

[0004] Moreover, the cost of collecting point cloud data is high, and it is difficult to obtain semantic information of the targets. For example, using devices such as lidar to collect point cloud data requires a high cost, and it is difficult to distinguish different types of target objects, such as vehicles, pedestrians, ground, etc.

[0005] Existing deep learning-based vehicle detection and recognition methods, including object detection, point cloud processing, and semantic segmentation methods, have certain limitations in solving the problem of detecting abnormal vehicle situations in parking lot management. Object detection algorithms are difficult to obtain accurate position and size information of the targets, point cloud processing algorithms have a high collection cost and are difficult to obtain semantic information of the targets, and semantic segmentation algorithms are difficult to obtain accurate position and size information of the targets.

[0006] It can be seen that the existing technologies have the following disadvantages:

[0007] Existing methods are mainly based on two-dimensional image information, and it is difficult to accurately judge whether the parking position of the vehicle is regular, such as whether it exceeds the parking space range, whether it is parked in a no-parking area, etc. Limited by two-dimensional image information, it is difficult to accurately judge the parking situation of the vehicle, especially in cases of vehicle occlusion, light changes, etc., the recognition accuracy is low. And it is difficult to distinguish different types of target objects, such as vehicles, pedestrians, ground, etc., and it is difficult to detect abnormal situations such as vehicles parked in no-parking areas. Summary of the Invention

[0008] The purpose of the embodiments of the present invention is to provide a 3D semantic occupancy prediction method for parking space occupancy detection in a parking lot, which combines three-dimensional space data and deep learning technology to improve the recognition accuracy of vehicle parking positions, reduce the influence of occlusion and illumination changes, and thus can effectively solve the problem of detecting the occupancy situation of parking spaces in parking lot management, improve the management efficiency and safety of the parking lot, and further solve at least one technical problem involved in the background technology.

[0009] To solve the above technical problems, the present invention is implemented as follows:

[0010] In the first aspect, the embodiments of the present invention provide a 3D semantic occupancy prediction method for parking space occupancy detection in a parking lot, including the following steps:

[0011] Step S1, collect image data of the parking lot scene through a camera in the parking lot;

[0012] Step S2, extract image features from the collected image data, and introduce DepthAnything to predict the depth information of each pixel point in the image;

[0013] Step S3, reconstruct the 3D occupancy grid space of the parking lot scene according to the extracted image features and the predicted depth information;

[0014] Step S4, in the 3D occupancy grid space guided by depth, perform precise semantic classification / recognition on each voxel in the space, and reconstruct a complex three-dimensional scene with semantic labels including occlusion and image-invisible areas;

[0015] Step S5, perform parking space occupancy prediction based on the reconstructed complex three-dimensional scene, and then identify abnormal parking space situations according to the occupancy prediction results.

[0016] In the second aspect, the embodiments of the present invention provide an electronic device, including:

[0017] At least one processor;

[0018] At least one memory for storing at least one program;

[0019] When the at least one program is executed by the at least one processor, the at least one processor implements the steps of the method described in the first aspect.

[0020] In the third aspect, the embodiments of the present invention provide a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0021] Fourthly, an embodiment of the present invention provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the method as described in the first aspect.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0023] 1. By combining depth information and image features for 3D reconstruction, the present invention can effectively solve the problem of insufficient accuracy in traditional monocular vision-based 3D reconstruction methods;

[0024] 2. Compared with traditional 3D reconstruction methods that often rely on complex devices such as binocular vision or structured light, which are costly and difficult to deploy, the present invention can be implemented with only an ordinary camera, greatly reducing the hardware cost and installation complexity; in addition, through accurate 3D reconstruction, the actual scene of the parking lot can be more realistically restored, providing a reliable data basis for subsequent parking space status analysis, thereby improving the accuracy of parking space occupancy detection;

[0025] 3. The present invention uses 3D occupancy grids guided by depth for semantic segmentation and occupancy prediction, solving the problem that traditional methods are difficult to accurately identify occluded and invisible areas; by mapping 2D image features to 3D space and using the Transformer architecture for feature interaction, the semantic information of the 3D feature map is enhanced, enabling accurate object recognition and classification even in complex and changing parking environments;

[0026] 4. Based on the occupancy prediction results, the present invention can automatically identify abnormal conditions of parking spaces, such as vehicles parked beyond the range or obstacles in the parking space, and thus timely issue an alarm or notify the management personnel to take corresponding measures. This not only helps to improve the management efficiency and service quality of the parking lot, but also effectively avoids potential safety hazards caused by improper use of parking spaces; in addition, by setting a reasonable abnormal score threshold, the system can flexibly adjust the sensitivity to abnormal conditions, ensuring both the accuracy of detection and reducing the false alarm rate, and ensuring the stable operation of the system. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0028] Figure 1 is a schematic flow chart of a 3D semantic occupancy prediction method for parking space occupancy detection provided by an embodiment of the present invention;

[0029] Figure 2 is one of the schematic diagrams of the hardware structure of the electronic device provided by the embodiments of the present invention;

[0030] Figure 3 is the second schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0032] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0033] Next, in conjunction with the accompanying drawings, the 3D semantic occupancy prediction method for parking lot space occupancy detection provided by the embodiments of the present invention will be described in detail through specific embodiments and their application scenarios.

[0034] Please refer to Figure 1 , which is a 3D semantic occupancy prediction method for parking lot space occupancy detection provided by the embodiments of the present invention, including the following steps:

[0035] Step S1, collecting image data of the parking lot scene through a camera in the parking lot;

[0036] Step S2, extracting image features from the collected image data and introducing DepthAnything to predict the depth information of each pixel point in the image;

[0037] Step S3, reconstructing the 3D occupancy grid space of the parking lot scene according to the extracted image features and the predicted depth information;

[0038] Step S4, in the 3D occupied grid space under the guidance of depth, accurately semantically classify / recognize each voxel in the space, and reconstruct a complex 3D scene with semantic labels including occluded and image invisible areas;

[0039] Step S5, performing parking space occupancy prediction based on the reconstructed complex three-dimensional scene, and then identifying abnormal parking space conditions according to the occupancy prediction results.

[0040] In step S1, the image data is a two-dimensional color image I ( u , v ),in, ( u , v ) are the pixel coordinates on the image plane.

[0041] In step S2, ResNet-50 is used to extract image features, including:

[0042] Image I ( u , v ) Input ResNet-50, output image I ( u , v ) At each pixel coordinate ( u , v ) at the eigenvector F ( u , v ), whose dimensions are B × d × H × W , where B represents the number of images processed in one forward propagation; d represents the number of channels; H represents the height of the output feature map; and W represents the width of the output feature map.

[0043] In step S2, DepthAnything is introduced to predict the depth information of each pixel in the image, including:

[0044] The feature vector F ( u , v ) Input DepthAnything, output image I ( u , v ) At each pixel coordinate ( u , v ) The predicted depth information ( u , v ), whose dimensions are B ×1× H ×W , where 1 indicates that the number of channels is 1.

[0045] Step S3 specifically includes:

[0046] Step S31, using depth information ( u , v ) and the camera intrinsic matrix K to convert the pixel coordinates of the 2D image I ( u , v ) to 3D spatial coordinates u , v ), which is expressed by the following formula: ( X , Y , Z ), and is expressed as:

[0047] ;

[0048] In the formula, is the camera optical center coordinate, and are the focal lengths in the horizontal and vertical directions;

[0049] Step S32, using the Transformer architecture for feature propagation, to transfer the 2D feature vector F ( u , v ) to the 3D space and generate a 3D feature map ;

[0050] Step S33, using the Transformer architecture for feature interaction, to enhance the semantic information of the 3D feature map and generate a complete 3D semantic scene .

[0051] Step S32 specifically includes:

[0052] Performing cross-attention mechanism calculation on the predefined voxel query Q and the corresponding 2D feature vector F ( u , v ) to obtain the 3D feature vector of the voxel query, and then generate a 3D feature map , which is expressed by the following formula:

[0053] ;

[0054] ;

[0055] In the formula, is a set of voxel queries and the specific voxel query in; is a 2D image feature sequence; is a voxel query and the corresponding visible image set; P (p, t) is a function that projects the 3D point p onto the image t; is the 2D feature of the image t; g is an aggregation function; is to update the voxel query and the resulting set is the set of; m is a mask token; DCA is a deformable cross-attention mechanism; DA is a deformable attention mechanism.

[0056] Step S33 specifically includes:

[0057] Perform self-attention mechanism calculation on the voxel feature vectors of the 3D feature map to enhance the semantic association between feature vectors, and aggregate the interacted feature vectors to obtain a complete 3D semantic scene which is expressed by the following formula:

[0058] ;

[0059] In the formula, DSA is a deformable self-attention mechanism; f is a voxel query or mask token; p is the position of the voxel query or mask token.

[0060] In step S4, perform precise semantic classification / recognition on each voxel in the space to obtain voxel features represents a low-resolution feature map; perform upsampling and project it into the output space to obtain the output where M + 1 represents M semantic classes and an empty class, that is, the result of occupancy prediction is obtained; Z represents the depth dimension.

[0061] In step S5, identifying the parking space anomaly situation according to the occupancy prediction result specifically includes:

[0062] Step S51, calculate the anomaly score, input the voxel feature vector of the 3D feature map and output the 3D voxel anomaly score s(x), which is expressed by the following formula:

[0063]

[0064] ​Wherein, the anomaly score s(x) represents the possibility that the voxel belongs to an abnormal object; g is a non-linear transformation function; x is the input data; y is the corresponding label; is an estimate of the unnormalized joint probability density;

[0065] Step S52, threshold processing, compare the anomaly score s(x) with the set threshold τ. If s(x)>τ, then mark the voxel as an abnormal voxel, which is expressed by the following formula:

[0066] A(x)=s(x)>τ

[0067] Wherein, A(x) represents the 3D anomaly voxel map, A(x)=1 represents an abnormal voxel, A(x)=0 represents a normal voxel, and τ represents the threshold;

[0068] Step S53, fuse the 3D semantic scene, fuse the 3D anomaly voxel map A(x) with the complete 3D semantic scene For fusion, in the fused 3D semantic scene, the abnormal voxels are marked as "unknown", which is expressed by the following formula:

[0069]

[0070] Wherein, represents the predicted result of the fused 3D semantic occupancy; represents the complete 3D semantic scene.

[0071] As Figure 2 shown, an embodiment of the present invention also provides an electronic device 600. The electronic device 600 includes a processor 601, a memory 602, a program or instruction stored on the memory 602 and executable on the processor 601. When the program or instruction is executed by the processor 601, it implements each process of the above-mentioned 3D semantic occupancy prediction method embodiment for parking space occupancy detection, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0072] It should be noted that the electronic device in the embodiment of the present invention includes the above-mentioned mobile electronic device and non-mobile electronic device.

[0073] Figure 3 FIG. is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention.

[0074] The electronic device 700 includes, but is not limited to: a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710, etc.

[0075] Those skilled in the art can understand that the electronic device 700 may further include a power source (such as a battery) for powering each component. The power source can be logically connected to the processor 710 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. Figure 3 The structure of the electronic device shown in the figure does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0076] It should be understood that in the embodiments of the present invention, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The graphics processor 7041 processes the image data of static images or videos obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 707 includes a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here. The memory 709 can be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 710 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 710.

[0077] The embodiments of the present invention also provide a readable storage medium. Programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, the various processes of the above-mentioned 3D semantic occupancy prediction method embodiment for parking space occupancy detection are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0078] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disks, or optical discs, etc.

[0079] Another embodiment of the present invention provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above-mentioned 3D semantic occupancy prediction method embodiment for parking space occupancy detection, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0080] It should be understood that the chip mentioned in the embodiments of the present invention may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0081] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0083] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention and the claims, and all of them belong to the protection scope of the present invention.

Claims

1. A 3D semantic occupancy prediction method for parking space occupancy detection, characterized in that, Including the following steps: Step S1, collecting image data of the parking lot scene through cameras in the parking lot; Step S2, extracting image features from the collected image data and introducing DepthAnything to predict the depth information of each pixel point in the image; Step S3, reconstructing the 3D occupancy grid space of the parking lot scene according to the extracted image features and the predicted depth information; Step S4, in the 3D occupancy grid space guided by depth, performing precise semantic classification / recognition on each voxel in the space to reconstruct a complex three-dimensional scene with semantic labels including occluded and invisible regions in the image; Step S5, based on the reconstructed complex three-dimensional scene, predicting the occupancy of parking spaces, and then identifying abnormal parking space situations according to the occupancy prediction results, specifically including: Step S51, calculate the anomaly score, input the voxel feature vector of the 3D feature map , and output the 3D voxel anomaly score s(x), which is expressed by the following formula: In the formula, the anomaly score s(x) represents the possibility that the voxel belongs to an anomalous object; g is a non-linear transformation function; x is the input data; y is the corresponding label; is an estimate of the unnormalized joint probability density; represents the voxel feature vector of the 3D feature map; Step S52, threshold processing, comparing the abnormal score s(x) with the set threshold τ. If s(x)>τ, then mark the voxel as an abnormal voxel, which is expressed by the following formula: A(x)=s(x)>τ In the formula, A(x) represents the 3D abnormal voxel map, A(x)=1 represents an abnormal voxel, A(x)=0 represents a normal voxel, and τ represents the threshold; Step S53, fuse the 3D semantic scene, and fuse the 3D abnormal voxel map A(x) with the complete 3D semantic scene For the fused 3D semantic scene, the abnormal voxels are marked as "unknown", which is expressed by the following formula: In the formula, represents the fused 3D semantic occupancy prediction result; represents the complete 3D semantic scene.

2. The method according to claim 1, characterized in that, In step S1, the image data is a two-dimensional color image I ( u , v ), where ([[]] u , v ) are the pixel coordinates on the image plane.

3. The method according to claim 2, wherein In step S2, ResNet-50 is used to extract image features, specifically including: Input the image I ( u , v ) into ResNet-50 and output the image I ( u , v ) with the feature vector at each pixel coordinate ( u , v ), where the dimension of the feature vector F ( u , v ) is B × d × H × W , where B represents the number of images processed in one forward propagation; d represents the number of channels; H represents the height of the output feature map; and W represents the width of the output feature map.

4. The method according to claim 3, wherein In step S2, DepthAnything is introduced to predict the depth information of each pixel point in the image, specifically including: Input the feature vector F ( u , v ) into DepthAnything and output an image I ( u , v ) with the depth information predicted at each pixel coordinate( u , v ), where the dimension is ( u , v ), and 1 represents that the number of channels is 1, B ×1× H × W .

5. The method according to claim 4, characterized in that Step S3 specifically includes: Step S31, using depth information ( u , v ) and the camera intrinsic matrix K The 2D image I ( u , v ) pixel coordinates( u , v ) converted to 3D space coordinates ( X , Y , Z ), which is expressed by the following formula: ; wherein, ( , ) are the coordinates of the optical center of the camera, and are the focal lengths in the horizontal and vertical directions; Step S32, perform feature propagation using the Transformer architecture to transfer the 2D feature vectors F ( u , v ) into the 3D space and generate a 3D feature map ; Step S33: Use the Transformer architecture for feature interaction to enhance the semantic information of the 3D feature map and generate a complete 3D semantic scene and generate a complete 3D semantic scene .

6. The method according to claim 5, wherein Step S32 specifically includes: Perform a predefined voxel query with the corresponding 2D feature vector F ( u , v ) to perform cross-attention mechanism calculation to obtain the 3D voxel feature vector of the voxel query, and further generate a 3D feature map , which is represented by the following formula: ; ; In the formula, is the set of voxel queries in the specific voxel queries; is the 2D image feature sequence; is the voxel query corresponding visible image set; P (p, t) is a function that projects the 3D point p onto the image t; is the 2D feature of the image t; g is an aggregation function; is the set after updating the voxel query i.e., the set of ; m is a mask token; DCA is the deformable cross-attention mechanism; DA is the deformable attention mechanism.

7. The method according to claim 6, wherein Step S33 specifically includes: Perform self-attention mechanism calculation on the voxel feature vectors of the 3D feature map to enhance the semantic association between feature vectors, and aggregate the feature vectors after interaction to obtain a complete 3D semantic scene , which is expressed by the following formula: = In the formula, DSA is the deformable self-attention mechanism; f is the voxel query or mask token; p is the position of the voxel query or mask token.

8. The method according to claim 7, characterized in that, In step S4, each voxel in the space is precisely semantically classified / identified to obtain voxel features , which are upsampled and projected into the output space to obtain an output , where M + 1 represents M semantic classes and one empty class, that is, the result of occupancy prediction is obtained; Z represents the depth dimension.

Citation Information

Patent Citations

  • Roadside parking management method and system based on three-dimensional vehicle detection

    CN118334873A