Vehicle pose estimation method based on semantic feature point detection

Through the vehicle position estimation method based on semantic feature point detection, the YOLOv8 model is used to detect 2D and 3D key points and solve the vehicle position pose through a multi-point perspective algorithm, which solves the problem of sample labeling difficulties and insufficient model generalization capabilities in traditional methods, and realizes efficient and accurate vehicle position estimation in complex road scenarios.

CN119991810AActive Publication Date: 2025-05-13ANHUI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510175928.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The traditional vehicle position estimation method has problems such as difficulty in labeling samples, insufficient generalization capabilities of models, and high cost and high complexity. It is especially difficult to achieve efficient and accurate vehicle position estimation in complex actual road scenarios.

Method used

The vehicle position estimation method based on semantic feature point detection is adopted. By constructing a virtual vehicle 3D model, the 2D and 3D key points are detected using the YOLOv8 model, and the vehicle position pose is solved through a multi-point perspective algorithm to achieve efficient and accurate vehicle position estimation.

Benefits of technology

Without additional parameters and manual marking, efficient and accurate vehicle position estimation in complex actual road scenarios is achieved. It is suitable for multiple fields such as autonomous driving and smart transportation, and has broad prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991810A_ABST
    Figure CN119991810A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a semantic feature point detection-based vehicle pose estimation method, which comprises the following steps of S1, constructing a virtual vehicle 3D model, and selecting points on the vehicle model as 3D key points; s2, collecting a vehicle picture in the real scene road picture, detecting a 2D bounding box of the vehicle picture through a YOLOv8 model, and extracting the vehicle picture in the 2D bounding box; and a step S3 of classifying the vehicles of the extracted vehicle pictures through a VisionTransform technology. According to the method, efficient and accurate vehicle pose estimation can be realized in a complex actual road scene without additional parameters and manual annotation, and the method is suitable for multiple fields such as automatic driving, intelligent traffic and augmented reality and has a wide prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a vehicle posture estimation method based on semantic feature point detection. Background Art

[0002] In modern cities, traffic monitoring systems have become an important means to maintain traffic order and improve road safety. As an important part of smart transportation and vehicle autonomous driving, vehicle pose estimation has received widespread attention. The core goal of vehicle pose estimation is to accurately obtain the vehicle's position information and pose parameters, including the vehicle's position coordinates, orientation, and rotation angle, by analyzing image or sensor data.

[0003] Traditional vehicle pose estimation methods are usually divided into two types. One is based on deep learning technology. This type of technology mainly predicts vehicles by training deep learning models, and is usually divided into two categories. One type trains models based on labeled data to directly obtain vehicle pose information; the other type detects vehicle key points, estimates pose information through perspective projection based on camera parameters, or estimates pose by matching with vehicle models.

[0004] The other method is to estimate the posture based on precision instruments such as sensors, integrate multiple sensors or filter results, iteratively update the vehicle status according to the vehicle's dynamic model, and estimate the vehicle's posture.

[0005] Traditional pose estimation methods based on deep learning have certain disadvantages:

[0006] ①Currently, deep learning methods for pose estimation usually use supervised methods. This process requires a large number of samples for model training. However, it is very difficult to annotate three-dimensional pose information in 2D images, which often requires a lot of manpower and material resources.

[0007] ② Once the deep learning model is trained, the types of vehicles that the model can detect remain fixed. Vehicles that are beyond the training samples will not be correctly identified.

[0008] ③The key point detection method requires information support, such as camera parameters and actual aspect ratio.

[0009] The precision instrument-based method has the following disadvantages:

[0010] ① Vehicle pose estimation methods based on precision instruments have the advantage of high accuracy, but they are also accompanied by high costs, complex equipment setup and operation, sensitivity to environmental conditions, and possible impacts on vehicle weight, energy consumption, and data latency.

[0011] ②These methods often only target a single vehicle that needs to be observed, and are limited by factors such as visibility, hardware failure, and signal interference, which may lead to inaccurate or failed pose estimation under certain circumstances. Therefore, a vehicle pose estimation method based on semantic feature point detection is proposed. Summary of the invention

[0012] In order to solve the technical problems existing in the prior art, the present invention provides a vehicle posture estimation method based on semantic feature point detection.

[0013] The present invention is implemented by the following technical solution: a vehicle posture estimation method based on semantic feature point detection comprises the following steps:

[0014] Step S1: construct a virtual vehicle 3D model and select points on the vehicle model as 3D key points;

[0015] Step S2: Collect vehicle images in real-world road images, detect the 2D bounding box of the vehicle image through the YOLOv8 model, and extract the vehicle image in the 2D bounding box;

[0016] Step S3: classifying the vehicles in the extracted vehicle images using VisionTransformer technology;

[0017] Step S4: Selecting a corresponding vehicle 3D model according to the classified vehicle type;

[0018] Step S5: Improve the YOLOv8 model by integrating the SWS layer into the convolutional layer in the YOLOv8 backbone network and Neck, and adding the BSAM layer after the SPPF layer of the backbone network;

[0019] Step S6: Use the improved YOLOv8 model to detect the vehicle 3D model and select 3D key points;

[0020] Step S7: Solve the vehicle posture through a multi-point perspective algorithm:

[0021] First, the vehicle 3D model is rendered to obtain synthetic image samples;

[0022] Then, use the synthetic samples to train the improved YOLOv8 model;

[0023] According to the definition of 3D semantic feature points, the 3D semantic feature points of the real vehicle image are annotated, and the improved YOLOv8 model is fine-tuned. The fine-tuned YOLOv8 model detects and obtains 2D key points;

[0024] The filtered 3D key points are matched one by one with the detected 2D key points, and the pose is solved through the multi-point perspective algorithm.

[0025] As a further improvement of the above solution, in step S1, a 3D modeling software is used to model a virtual vehicle 3D model.

[0026] As a further improvement of the above solution, the steps of selecting corner points as 3D key points in step S1 are as follows:

[0027] From any vertex x0 on the virtual vehicle 3D model, there are k neighboring points {x1,x2,…,x k}, the covariance matrix of the point set is in

[0028] Calculate the eigenvalues ​​λ1, λ2, λ3 of M, and 0≤λ1≤λ2≤λ3, It can reflect the curvature around the vertex and set the threshold ε. Then x0 is determined to be a corner point.

[0029] As a further improvement of the above solution, the 2D bounding box in step S2 includes the category of the detected target, the 2D coordinates of the center point of the 2D bounding box, and the length and width of the bounding box.

[0030] As a further improvement of the above solution, the synthesized image sample in step S7 includes the vehicle picture, the 3D feature points corresponding to the 3D key points, and the coordinate information of the 2D feature points;

[0031] Among them, the 3D feature point is the three-dimensional coordinate of the 3D key point in the world coordinate system established with the center of the virtual vehicle 3D model as the origin, and the 2D feature point is the projection coordinate of the 3D feature point on the 2D image.

[0032] As a further improvement of the above scheme, the training steps of the improved YOLOv8 model in step S7 are:

[0033] Configure the conda environment required by YOLOv8 and install related dependency libraries;

[0034] Change the 2D point annotation file in the synthetic sample to the format required by YOLOv8;

[0035] The data set is divided into training set, validation set and test set in a ratio of 8:1:1;

[0036] Change the learning rate, batch, epoch and other training-related configurations in the YOLOv8 configuration file;

[0037] Run the training script and specify the input folder and output folder. The input folder contains training set images, validation set images, and 2D annotation files. The output folder is used to save the weights of the model output.

[0038] As a further improvement of the above solution, in step S7, semantic features are defined for the filtered 3D key points and the corresponding 3D feature points and 2D feature points, and the names and types of the filtered 3D key points and the corresponding 3D feature points and 2D feature points are defined.

[0039] As a further improvement of the above solution, the multi-point perspective algorithm in step S7 is:

[0040] Among them, X i , Y i and Z i is the coordinate of the 3D key point, u i and v i is the coordinate of the 2D key point, K is the camera internal parameter, so as to solve the R rotation matrix and t translation matrix.

[0041] As a further improvement of the above solution, the YOLOv8 model includes Backbone, Neck and Head.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] The present invention can achieve efficient and accurate vehicle pose estimation in complex actual road scenes without additional parameters and manual labeling. It is suitable for multiple fields such as autonomous driving, smart transportation, augmented reality, etc. and has broad prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flow chart of a vehicle pose estimation method based on semantic feature point detection provided by the present invention;

[0045] Figure 2 A schematic diagram of the structure of the YOLOv8 model provided by the present invention;

[0046] Figure 3 A schematic diagram of the structure of the improved YOLOv8 model provided by the present invention;

[0047] Figure 4 This is a schematic diagram of the structure of the Conv_SWS-1 layer to the Conv_SWS-6 layer provided by the present invention. DETAILED DESCRIPTION

[0048] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.

[0049] Embodiment 1:

[0050] Please combine Figure 1 The vehicle posture estimation method based on semantic feature point detection of this embodiment includes the following steps:

[0051] Step S1: construct a virtual vehicle 3D model and select points on the vehicle model as 3D key points;

[0052] The 3D modeling software is used to model a virtual vehicle 3D model. The modeled virtual vehicle 3D model can be imported into software such as Unity, and corner points are selected as 3D key points. The steps are as follows:

[0053] Select any vertex x0 on the 3D model of the virtual vehicle, and there are k neighboring points {x1,x2,…,x k}, the covariance matrix of the point set is in

[0054]

[0055] Among them, x0, x1, x2, …, x k is the vertex on the 3D model of the virtual vehicle, M is the vertex through x0, x1, x2,…, x k The covariance matrix calculated by the point set, k is the number of neighboring points, xi represents {x0,x1,…,x k}, u refers to the three-dimensional coordinates of any point in {x0,x1,…,x k}The average value of these neighbor point sets, T is represented by the transpose of the matrix;

[0056] Calculate the eigenvalues ​​λ1, λ2, λ3 of M, and 0≤λ1≤λ2≤λ3, It can reflect the curvature around the vertex and set the threshold ε. Then x0 is determined to be a corner point; the threshold ε can reflect the curvature characteristics around the vertex, and can effectively distinguish the vertices with curvature changes in the local structure from those with relatively flat surfaces. For flat areas, the spatial distribution of the vertex and its neighboring points is relatively uniform, resulting in a smaller difference in the eigenvalues ​​of the covariance matrix; if It is considered that there is a large curvature around the vertex, which meets the characteristics of a corner point, and it is selected as a 3D key point;

[0057] Among them, λ1, λ2, λ3 are the eigenvalues ​​after the covariance matrix is ​​calculated, λ1, λ2, λ3 represent the degree of dispersion of the local geometry of the vertex in different directions, and λ1 and λ3 represent the directions of minimum and maximum dispersion respectively;

[0058] Step S2: Collect vehicle images in real-world road images, detect the 2D bounding box of the vehicle image through the YOLOv8 model, and extract the vehicle image in the 2D bounding box;

[0059] The 2D bounding box includes the category of the detected target, the 2D coordinates of the center point of the 2D bounding box, and the length and width of the bounding box;

[0060] Step S3: Classify the vehicles in the extracted vehicle images by using the VisionTransformer technology, and divide them into sedans, SUVs, and trucks;

[0061] Among them, Vision Transformer (ViT) is an image classification model based on the Transformer architecture. The network structure of ViT can be divided into several key parts. First, the input image is divided into patches of fixed size (for example, 16x16 pixels), each patch is flattened and mapped to an embedding space of fixed dimension through a linear transformation. Next, the position embedding is added to the embedding vector of each patch, providing the spatial position information of the patch in the image. These patches are then fed into a standard Transformer encoder, which includes multiple self-attention layers and feedforward neural network layers. The self-attention layer can capture the global dependencies between different patches, while the feedforward neural network further processes this information. The core of ViT is multiple stacked Transformer encoder layers, which improve the performance of image classification by capturing the relationship between patches in the image.

[0062] After being processed by the Transformer encoder, ViT extracts a special [CLS] patch from the input sequence, which contains the global features of the entire image and serves as the basis for image classification. Next, the final representation of the [CLS] tag is processed through a classification head to generate the final category prediction. Through the softmax function, the model can output the probability distribution of each category and select the most likely category as the label of the image based on the probability;

[0063] Step S4: Selecting a corresponding vehicle 3D model according to the classified vehicle type;

[0064] Step S5: Improve the YOLOv8 model by integrating the SWS layer into the convolutional layer in the YOLOv8 backbone network and Neck, and adding the BSAM layer after the SPPF layer of the backbone network;

[0065] Step S6: Use the improved YOLOv8 model to detect the vehicle 3D model and select 3D key points;

[0066] Step S7: Solve the vehicle posture through a multi-point perspective algorithm:

[0067] First, the vehicle 3D model is rendered to obtain synthetic image samples;

[0068] Among them, the vehicle 3D model rendering is to put the virtual 3D model into the virtual environment of Unity for synthesis. The synthesized image samples are pictures of the vehicle at different positions and angles in the virtual environment;

[0069] The synthetic image sample includes a vehicle picture, 3D feature points corresponding to the 3D key points, and 2D feature point coordinate information;

[0070] Among them, the 3D feature point is the three-dimensional coordinate of the 3D key point in the world coordinate system established with the center of the virtual vehicle 3D model as the origin, and the 2D feature point is the projection coordinate of the 3D feature point on the 2D image, and the 3D feature point and the 2D feature point are numbered so that the 3D feature point and the 2D feature point correspond to each other;

[0071] Then, use the synthetic samples to train the improved YOLOv8 model;

[0072] Among them, the training steps of the improved YOLOv8 model are:

[0073] Configure the conda environment required by YOLOv8 and install related dependency libraries;

[0074] Change the 2D point annotation file in the synthetic sample to the format required by YOLOv8;

[0075] The data set is divided into training set, validation set and test set in a ratio of 8:1:1;

[0076] Change the learning rate, batch, epoch and other training-related configurations in the YOLOv8 configuration file;

[0077] Run the training script and specify the input folder and output folder. The input folder contains training set images, validation set images, and 2D annotation files. The output folder is used to save the model output weights.

[0078] According to the definition of 3D semantic feature points, the 3D semantic feature points of the real vehicle image are annotated, and the improved YOLOv8 model is fine-tuned. The learning rate, batch_size and epoch parameters of the improved YOLOv8 model are adjusted, and the fine-tuned YOLOv8 model detects and obtains 2D key points;

[0079] Among them, semantic features are defined for the filtered 3D key points and the corresponding 3D feature points and 2D feature points, and the names and types of the filtered 3D key points and the corresponding 3D feature points and 2D feature points are defined;

[0080] The filtered 3D key points are matched one by one with the detected 2D key points, and the pose is solved through the multi-point perspective algorithm.

[0081] Among them, the multi-point perspective algorithm is:

[0082] Among them, X i , Y i and Z i is the coordinate of the 3D key point, u i and v i is the coordinate of the 2D key point, K is the camera internal parameter, so as to solve the R rotation matrix and t translation matrix;

[0083] Among them, R and t represent the rotation matrix and translation matrix respectively. R is a 3×3 matrix used to describe the rotation transformation from the world coordinate system to the camera coordinate system; it shows how the object rotates in three-dimensional space so that the coordinates of the object can be aligned with the camera coordinate system; t is a 3×1 vector, which represents the translation transformation from the world coordinate system to the camera coordinate system; it describes the displacement of the camera relative to the object; the combination of R and t is the object's posture.

[0084] Embodiment 2:

[0085] like Figure 2 As shown, the YOLOv8 model includes Backbone, Neck and Head;

[0086] Among them, Backbone includes Conv layer, Conv-1 layer, C2f-1 layer, Conv-2 layer, C2f-2 layer, Conv-3 layer, C2f-3 layer, Conv-4 layer, C2f-4 layer and SPPF layer connected in sequence;

[0087] Neck includes Upsample-1 layer, Concat-1 layer, C2f-5 layer, Upsample-2 layer, Concat-2 layer, C2f-6 layer, Conv-5 layer, Concat-3 layer, C2f-7 layer, Conv-6 layer, Concat-4 layer and C2f-8 layer connected in sequence;

[0088] Head includes Detect-1 layer, Detect-2 layer and Detect-3 layer;

[0089] The C2f-2 layer is connected to the Concat-2 layer, the SPPF layer is connected to the Upsample-1 layer and the Concat-4 layer, the C2f-6 layer is connected to the Detect-1 layer, the C2f-7 layer is connected to the Detect-2 layer, and the C2f-8 layer is connected to the Detect-3 layer.

[0090] Embodiment 3:

[0091] like Figure 3 As shown in the figure, the SWS (SimAM With Slicing) layer is integrated into the Conv-1 layer to the Conv-6 layer of the YOLOv8 model, and the BSAM (BiLebel Spatial Attention Module) layer is added after the SPPF layer of the backbone network to form an improved YOLOv8 model, which improves the accuracy of model detection;

[0092] The structures of Conv_SWS-1 to Conv_SWS-6 are as follows: Figure 4 As shown;

[0093] The BSAM layer is composed of a BRA module (Bi-level Routing Attention) and a spatial attention module. The spatial attention module weights the features of different positions in the spatial dimension to capture the relationship between different areas. The BRA module achieves more flexible content-aware computational allocation through a two-layer routing mechanism. The BRA module includes two levels of attention operations: Few-to-Many Attention and Many-to-Few Attention.

[0094] In the first level (Few-to-Many Attention), the number of query vectors is small, while the number of key vectors is large. Each query vector only performs attention calculations with some key vectors, which can reduce the amount of calculations and reduce the complexity. The second level (Many-to-Few Attention) ensures that each query vector performs attention calculations with all key vectors, thereby ensuring the expressiveness of the model and capturing more feature information. The attention calculation formula in the BRA module is: Where Q is the query vector, K is the key vector matrix, V is the value vector matrix, K^T represents the transpose of the key vector matrix, and C is a scalar factor used to avoid weight concentration and gradient vanishing.

[0095] The Spatial Attention Module aims to enhance the spatial information of feature maps by focusing on salient regions in the image. The core goal of this module is to highlight the regions in the image that are useful for the task by adaptively assigning weights to different spatial locations. The Spatial Attention Module first receives the feature map from the convolutional network as input. The module compresses the feature map in the spatial dimension through global average pooling and global maximum pooling to generate two single-channel images that capture the average response and maximum response of each position in the image, respectively. Then, the two pooling results are concatenated in the channel dimension to form a new feature map. The convolutional layer performs a convolution operation on the concatenated feature map to generate a single-channel spatial attention map. In order to normalize the output and limit it to the range of [0,1], the spatial attention map is processed by the Sigmoid activation function. Finally, the generated attention map is multiplied element-by-element with the input feature map to weight different regions in the image, and high-weight regions are enhanced while low-weight regions are suppressed.

[0096] The SWS (SimAM With Slicing) layer is a SimAM module that introduces slicing operations. SimAM (Simple Attention Mechanism) is a lightweight feature enhancement module. Its design is inspired by the way humans process visual information, that is, the brain judges the importance of each pixel in the image by evaluating its relationship with the surrounding pixels. In SimAM, this importance is represented by 3-D weights. By calculating the Energy function, SimAM can measure the importance of each pixel and dynamically adjust the weight of the input image. Finally, the adjusted feature map is obtained by multiplying the calculated weight with the original input.

[0097] SWS introduces a slicing operation based on SimAM, which divides the feature map into multiple small blocks and independently calculates the average value of pixel differences in each slice. Large targets have obvious texture features, which will have a significant impact on the average value of the block they are in, so the weighted enhancement of the area where the large target is located is reduced. However, when these slices are merged into a complete feature map, large targets can still maintain a high degree of recognizability and may even be further enhanced. In contrast, the features of small targets are far away from the local average value of the slice they are in, so they will receive more weighting, thereby significantly enhancing the features of small targets. By introducing the slicing operation, the SWS module ensures that both large and small targets receive fair attention during the feature enhancement process. The reduced weighting of large targets helps reduce their excessive influence in the feature map, while small targets can better highlight their features because they receive more weighting.

[0098] The present invention can achieve efficient and accurate vehicle pose estimation in complex actual road scenes without additional parameters and manual labeling. It is suitable for multiple fields such as autonomous driving, smart transportation, augmented reality, etc. and has broad prospects.

[0099] The above-mentioned embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by technicians in this field on the basis of the present invention shall fall within the scope of protection required by the present invention.

Claims

1. A vehicle pose estimation method based on semantic feature point detection, characterized in that: The following steps are involved: Step S1: construct a virtual vehicle 3D model and select points on the vehicle model as 3D key points; Step S2: Collect vehicle images in real-world road images, detect the 2D bounding box of the vehicle image through the YOLOv8 model, and extract the vehicle image in the 2D bounding box; Step S3: classifying the vehicles in the extracted vehicle images using VisionTransformer technology; Step S4: Selecting a corresponding vehicle 3D model according to the classified vehicle type; Step S5: Improve the YOLOv8 model by integrating the SWS layer into the convolutional layer in the YOLOv8 backbone network and Neck, and adding the BSAM layer after the SPPF layer of the backbone network; Step S6: Use the improved YOLOv8 model to detect the vehicle 3D model and select 3D key points: Step S7: Solve the vehicle posture through a multi-point perspective algorithm: First, the vehicle 3D model is rendered to obtain synthetic image samples; Then, use the synthetic samples to train the improved YOLOv8 model; According to the definition of 3D semantic feature points, the 3D semantic feature points of the real vehicle image are annotated, and the improved YOLOv8 model is fine-tuned. The fine-tuned YOLOv8 model detects and obtains 2D key points; The filtered 3D key points are matched one by one with the detected 2D key points, and the pose is solved through the multi-point perspective algorithm.

2. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: In the step S1, a virtual vehicle 3D model is formed by using 3D modeling software.

3. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: In step S1, the corner points are selected as 3D key points, and the steps are as follows: From any vertex x0 on the virtual vehicle 3D model, there are k neighboring points {x1,x2,…,x k }, the covariance matrix of the point set is in Calculate the eigenvalues ​​λ1, λ2, λ3 of M, and 0≤λ1≤λ2≤λ3, It can reflect the curvature around the vertex and set the threshold ε. Then x0 is determined to be a corner point.

4. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: In step S2, the 2D bounding box includes the category of the detected target, the 2D coordinates of the center point of the 2D bounding box, and the length and width of the bounding box.

5. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: The synthesized image sample in step S7 includes the vehicle picture, the 3D feature points corresponding to the 3D key points, and the coordinate information of the 2D feature points; Among them, the 3D feature point is the three-dimensional coordinate of the 3D key point in the world coordinate system established with the center of the virtual vehicle 3D model as the origin, and the 2D feature point is the projection coordinate of the 3D feature point on the 2D image.

6. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: The training steps of the improved YOLOv8 model in step S7 are: Configure the conda environment required by YOLOv8 and install related dependency libraries; Change the 2D point annotation file in the synthetic sample to the format required by YOLOv8; The data set is divided into training set, validation set and test set in a ratio of 8:1:1; Change the learning rate, batch, epoch and other training-related configurations in the YOLOv8 configuration file; Run the training script and specify the input folder and output folder. The input folder contains training set images, validation set images, and 2D annotation files. The output folder is used to save the weights of the model output.

7. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: In the step S7, semantic features are defined for the filtered 3D key points and the corresponding 3D feature points and 2D feature points, and the names and types of the filtered 3D key points and the corresponding 3D feature points and 2D feature points are defined.

8. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: The multi-point perspective algorithm in step S7 is: Among them, X i , Y i and Z i is the coordinate of the 3D key point, u i and v i is the coordinate of the 2D key point, K is the camera internal parameter, so as to solve the R rotation matrix and t translation matrix.

9. The vehicle pose estimation method based on semantic feature point detection according to claim 1, characterized in that: The YOLOv8 model includes Backbone, Neck and Head.

Citation Information

Patent Citations

  • Vision-based vehicle target position and attitude angle detection method

    CN113436262A

  • Local refinement mapping system and method based on SLAM and semantic segmentation

    CN116772820A

  • Roadside pedestrian and vehicle detection method and device based on deep learning

    CN117409378A

  • Vehicle key point marking method and device, vehicle auxiliary driving method and system, electronic equipment, computer readable storage medium and vehicle

    CN119313738A

  • Detection and classification of moving objects in radar data using deep learning

    EP4365624A1