A method, device, medium and product for identifying 3D key points of vehicle body modeling surface

By adopting a two-stage key point recognition model in body shape recognition, combining the backbone network, Transformer model and multi-layer perception machine, the problem of low recognition accuracy under small data conditions is solved, and high-precision recognition within 30 pixels is achieved.

CN119559405BActive Publication Date: 2025-05-13YANGTZE DEITA GRADUATE SCHOOI OF BEIJING INST OF TECH (JIAXING) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510103791.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing body shape key point recognition algorithm has low recognition accuracy under small data conditions, and cannot accurately obtain 3D coordinates, and the relationship mining is insufficient, making it difficult to meet the accuracy requirements of engineering applications.

Method used

A two-stage key point recognition model is adopted, including a backbone network, a Transformer model and a multi-layer perceptron. Images are acquired through multiple angles, grayscale processing and image interpolation are performed, and the model is trained using a loss function to improve the accuracy of 3D key point recognition of the body shape surface.

Benefits of technology

Under the condition of small data samples, the accuracy of 3D key point recognition of the body shape surface is significantly improved, and it can meet the requirements of key point errors within 30 pixels in engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559405B_ABST
    Figure CN119559405B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, medium and product for identifying 3D key points of a vehicle body styling surface, and relates to the technical field of vehicle body exterior styling design. The method comprises: acquiring images of a vehicle body to be tested from multiple angles as a group of images to be tested; grayscale processing the group of images to be tested to obtain a grayscale image group to be tested; inputting the grayscale image group to be tested into a two-stage key point recognition model to obtain 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested. The two-stage key point recognition model is obtained by training an initial two-stage key point recognition model using multiple annotated historical image groups. The present application can improve the accuracy of 3D key point recognition of vehicle body styling surfaces under the condition of small data samples by constructing an initial two-stage key point recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicle body exterior styling design, and in particular to a method, device, medium and product for identifying 3D key points of a vehicle body styling surface. Background Art

[0002] At present, in the field of exterior body design, the design concept is evolving from a single dimension that focuses only on the aesthetics and functionality of the styling to a multi-dimensional design that takes into account aesthetics, functionality and structural performance. In the exterior surface design stage, introducing the vehicle performance evaluation of the body structure design stage in advance and changing the traditional design mode of first styling and then structure has become an important means to improve the accuracy and efficiency of exterior styling design. By putting the structural performance evaluation in advance, not only the number of rework iterations of styling designers is greatly reduced, the design cycle is shortened, and the effectiveness of the styling solution is improved.

[0003] To achieve this goal, a new technical route has been developed: input the image of the modeling surface, combine deep learning technology with accumulated historical vehicle model data, and predict the output of the body-in-white stiffness, mode, collision inclination and other performance. However, due to the limitation of the small data sample of the vehicle model, the evaluation accuracy and stability of the existing model are difficult to meet the prediction accuracy requirements of engineering applications. In-depth analysis found that this is mainly due to the following two aspects: 1) Low feature recognition accuracy: Under small data conditions, the recognition accuracy of the modeling surface features is insufficient, resulting in inaccurate description information of the modeling surface features. The recognition accuracy of the modeling surface of the body-in-white is low. Mining found that the existing research on the key point recognition algorithm of the body modeling surface based on image recognition has the following defects: First, the existing body feature point recognition algorithm generally only obtains the coordinates of the key points on the 2D plane of the body, and cannot support multi-view body modeling surface input to obtain the 3D coordinates of the key points, thus causing the problem of missing input feature information; the recognition accuracy of the existing feature point recognition algorithm is generally within 5% of the pixel scale of the input image (corresponding to the image size of this application is 100 pixels), and the accuracy of the model is difficult to meet the engineering requirements for achieving the deviation between the key point and the ideal point within 30 pixels. 2) Insufficient relationship mining: Also under the condition of small data, the complex relationship between the styling surface features and the body-in-white performance is not fully mined, making it difficult to accurately describe the performance indicators of the body-in-white. Summary of the invention

[0004] The purpose of this application is to provide a method, device, medium and product for identifying 3D key points of a vehicle body surface, which can improve the accuracy of 3D key point recognition of a vehicle body surface under the condition of a small data sample.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for identifying 3D key points of a vehicle body surface, comprising:

[0007] Acquire images of the vehicle body to be tested from multiple angles as a set of images to be tested;

[0008] Performing grayscale processing on the image group to be tested to obtain a grayscale image group to be tested;

[0009] The grayscale image group to be tested is input into the two-stage key point recognition model to obtain the 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested; the two-stage key point recognition model is obtained by training the initial two-stage key point recognition model using multiple annotated historical image groups; the initial two-stage key point recognition model includes a backbone network, a Transformer model and a multi-layer perceptron; the output end of the backbone network is respectively connected to the input end of the Transformer model and the input end of the multi-layer perceptron; the input end of the multi-layer perceptron is also connected to the output end of the Transformer model.

[0010] Optionally, the angles include a side view and a top view.

[0011] Optionally, the backbone network is a MobileNet network;

[0012] The Transformer model includes: an input embedding layer, a position encoding layer, a self-attention mechanism, an encoder and a decoder connected in sequence;

[0013] The multi-layer perceptron comprises an input layer, a hidden layer and an output layer which are connected in sequence.

[0014] Optionally, before acquiring images of the vehicle body to be tested from multiple angles as the image group to be tested, the method further includes:

[0015] Acquire multiple historical image groups; different historical image groups correspond to different vehicle bodies; any historical image group includes images corresponding to different angles of the vehicle body;

[0016] Mark multiple key points in each historical map group separately and determine the actual 3D coordinates of each key point;

[0017] Perform grayscale processing on each annotated historical image group to obtain multiple historical grayscale image groups;

[0018] Using bilinear interpolation method, image interpolation processing is performed on multiple historical grayscale image groups to obtain multiple historical interpolation image groups;

[0019] Taking the historical interpolation graph group as input and the multiple key points in the historical graph group as output, the initial two-stage key point recognition model is trained by using a loss function to obtain a two-stage key point recognition model.

[0020] Optionally, determine the actual 3D coordinates of each keypoint, including:

[0021] Let the key point number k=1;

[0022] Get the side view coordinates of the kth key point and the top-down coordinates ; The side view coordinates are the 2D coordinates of the key points in the side view image of the vehicle body; the top view coordinates are the 2D coordinates of the key points in the top view image of the vehicle body;

[0023] Determine the actual 3D coordinates of the kth key point based on the side view coordinates and top view coordinates of the kth key point ;in ;

[0024] Increase the value of the key point number k by 1 and return to step "Get the side view coordinates of the kth key point" and the top-down coordinates "Until the value of the key point number k reaches the total number of key points, the actual 3D coordinates of each key point are obtained.

[0025] Optionally, the loss function is:

[0026] ;

[0027] ;

[0028] ;

[0029] in, is the loss function; is the loss sum of the Euclidean distance difference between the predicted point and the target point; is the sum of the angular momentum differences; and All are weights; is the predicted value of the 3D coordinate of the kth key point; is the actual direction vector; is the prediction direction vector.

[0030] Optionally, the grayscale image group to be tested is input into a two-stage key point recognition model to obtain 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested, including:

[0031] Inputting the grayscale image group to be tested into the backbone network to obtain the grayscale image features to be tested;

[0032] Inputting the grayscale image features into the Transformer model to obtain a guide point distribution heat map;

[0033] Extracting predicted 3D coordinates of a plurality of guide points from a guide point distribution heat map; the guide points correspond one-to-one to the key points;

[0034] According to the preset step length, the local features corresponding to each guiding point are extracted from the grayscale image features to be tested with the guiding point as the center;

[0035] Inputting the local features corresponding to all the guiding points into the multi-layer perceptron to obtain the 3D coordinate fine-tuning amount of each guiding point;

[0036] The sum of the predicted 3D coordinates of the same guide point and the 3D coordinate fine-tuning amount is determined as the 3D coordinate prediction value of the corresponding key point.

[0037] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned vehicle body styling surface 3D key point recognition method.

[0038] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for identifying 3D key points of a vehicle body shaping surface.

[0039] In a fourth aspect, the present application provides a computer program product, including a computer program, which implements the above-mentioned vehicle body styling surface 3D key point recognition method when executed by a processor.

[0040] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0041] The present application provides a method, device, medium and product for identifying 3D key points of a vehicle body styling surface. The key point identification in the existing SUV vehicle body styling surface is taken as the research object, and an initial two-stage key point recognition model is constructed, including a backbone network, a Transformer model and a multi-layer perceptron. The output end of the backbone network is respectively connected to the input end of the Transformer model and the input end of the multi-layer perceptron. The input end of the multi-layer perceptron is also connected to the output end of the Transformer model. Using multiple annotated historical image groups, the initial two-stage key point recognition model is trained to obtain a two-stage key point recognition model, and images of the vehicle body to be tested are obtained from multiple angles as the image group to be tested. Grayscale processing is performed on the image group to be tested to obtain a grayscale image group to be tested. The grayscale image group to be tested is input into the two-stage key point recognition model to obtain the 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested, which solves the technical problem that the existing technology cannot achieve the error recognition accuracy within 30 pixels, can improve the feature recognition accuracy under small data sample conditions, and meet the needs of engineering applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 This is a flow chart of a method for identifying 3D key points of a vehicle body surface in one embodiment of the present application;

[0044] Figure 2 This is a technical schematic diagram of a method for identifying 3D key points of a vehicle body surface in an embodiment of the present application;

[0045] Figure 3 This is a network structure diagram of a dual-stage key point recognition model in an embodiment of the present application;

[0046] Figure 4 This is a schematic diagram of searching for an area near a 3D guide point on a vehicle body surface in an embodiment of the present application;

[0047] Figure 5 This is a flowchart of the 3D key point recognition model training for the vehicle body styling surface in one embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0049] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0050] In an exemplary embodiment, Figure 1 As shown, a method for identifying 3D key points of a vehicle body surface is provided, comprising:

[0051] Step 101: Acquire images of the vehicle body to be tested from multiple angles as a set of images to be tested, including side views and top views.

[0052] Step 102: grayscale processing is performed on the image group to be tested to obtain a grayscale image group to be tested.

[0053] Step 103: Input the grayscale image group to be tested into the dual-stage key point recognition model to obtain the 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested. The dual-stage key point recognition model is obtained by training the initial dual-stage key point recognition model using multiple annotated historical image groups.

[0054] like Figure 3 , the initial two-stage key point recognition model includes a backbone network, a Transformer model and a multi-layer perceptron. The output of the backbone network is connected to the input of the Transformer model and the input of the multi-layer perceptron respectively. The input of the multi-layer perceptron is also connected to the output of the Transformer model. The Transformer model includes: an input embedding layer, a position encoding layer, a self-attention mechanism, an encoder and a decoder connected in sequence. The multi-layer perceptron includes: an input layer, a hidden layer and an output layer connected in sequence.

[0055] Step 103 includes: inputting the grayscale image group to be tested into the backbone network to obtain the grayscale image features to be tested. Inputting the grayscale image features to be tested into the Transformer model to obtain a guide point distribution heat map. Extracting the predicted 3D coordinates of multiple guide points from the guide point distribution heat map. The guide points correspond to the key points one by one. According to the preset step size, extract the local features corresponding to each guide point from the grayscale image features to be tested with the guide point as the center. Input the local features corresponding to all guide points into the multi-layer perceptron to obtain the 3D coordinate fine-tuning amount of each guide point. Determine the sum of the predicted 3D coordinates and the 3D coordinate fine-tuning amount of the same guide point as the 3D coordinate prediction value of the corresponding key point.

[0056] like Figure 5 , before step 101, it also includes: an initial two-stage key point recognition model training process.

[0057] The initial two-stage keypoint recognition model training process includes:

[0058] Step 104: Acquire multiple historical image groups. Different historical image groups correspond to different vehicle bodies. Any historical image group includes images corresponding to different angles of the vehicle body.

[0059] Step 105: Mark multiple key points in each historical graph group respectively, and determine the actual 3D coordinates of each key point.

[0060] Determine the actual 3D coordinates of each key point, including: Set the key point number k = 1. Get the side view coordinates of the kth key point and the top-down coordinates ; The side view coordinates are the 2D coordinates of the key point in the side view image of the vehicle body; the top view coordinates are the 2D coordinates of the key point in the top view image of the vehicle body; according to the side view coordinates and top view coordinates of the kth key point, the actual 3D coordinates of the kth key point are determined ;in ; Increase the value of the key point number k by 1 and return to step "Get the side view coordinates of the kth key point" and the top-down coordinates "Until the value of the key point number k reaches the total number of key points, the actual 3D coordinates of each key point are obtained.

[0061] Step 106: Perform grayscale processing on each annotated historical image group to obtain multiple historical grayscale image groups.

[0062] Step 107: Use bilinear interpolation to perform image interpolation processing on multiple historical grayscale image groups to obtain multiple historical interpolation image groups.

[0063] Step 108: Taking the historical interpolation graph group as input and the multiple key points in the historical graph group as output, the initial two-stage key point recognition model is trained using the loss function to obtain a two-stage key point recognition model.

[0064] The loss function is:

[0065] .

[0066] .

[0067] .

[0068] in, is the loss function. is the sum of the losses of the Euclidean distance difference between the predicted point and the target point. is the difference in angular momentum and . and are weights. is the predicted value of the 3D coordinate of the kth key point is the actual direction vector; is the prediction direction vector.

[0069] As another implementation, Figure 2,The initial two-stage key point recognition model training process includes: S10, obtaining the side view and top view (overhead) images of C-type and D-type vehicles in existing SUV models, a total of 110 sets of image data, and manually annotating the positions of 22 key points in each image to obtain 3D coordinates. S20 stage 1, grayscale processing of all input images, removing the recognition error caused by different color blocks of vehicle modeling, and inputting the processed features into the Transformer architecture network layer to obtain the distribution heat map of the Heatmap where the guide point is located, and extracting the predicted coordinates of the guide point. Then connect S30 stage 2, obtain local features (Local feature) by transforming the input features, and increase the number of search points near the guide point, and finally input it into the multi-layer perceptron (MLP network layer) to predict the final key point fine-tuning amount. S40, combine the guide point coordinates in S20 and the key point fine-tuning amount in S30 to obtain the final key point coordinates. S50, two-stage network model training, by designing a composite loss function, training Transformer and MLP network parameters, to achieve the goal of gradually improving model accuracy.

[0070] In step S10: the coordinates of the key feature points in all the modeling surfaces are given by the designer through manual annotation, and the process is as follows:

[0071] 1) Then mark the coordinates of the 22 key feature points in the top view and obtain the 2D coordinates.

[0072] 2) Since the x-axis coordinates are repeatedly labeled in the side view and top view, the final , the key point coordinates are

[0073] In step S20: pre-process the input multi-view image data, and input the processed image features into the Transformer architecture network layer to obtain the predicted coordinates of the guide points, and finally connect the Heatmap layer to obtain the distribution heat map of the guide points to provide guidance for the subsequent generation of predicted key points. The specific steps include:

[0074] Step S21: Image grayscale processing.

[0075] For color images, three color channels, RGB (red, green, and blue), are usually used. Since grayscale images only have one color channel, they are easier to process and require less computation than three color channels in color images. Grayscale images are often used to extract important features in images, such as edges, textures, and shapes, which are essential for image recognition and analysis. The common formula for converting RGB images to grayscale images is:

[0076]

[0077] where are the values ​​of the red, green, and blue channels for any point in the styling graph.

[0078] Step S22: Image interpolation processing.

[0079] First find the four closest integer coordinate points in the original image , , ,and .here, and is the result of rounding down the sum, and and They are and .

[0080] The coordinates of the new interpolation points after bilinear interpolation are

[0081] .

[0082] .

[0083] .

[0084] .

[0085] Where a is the decimal part of the coordinate. b is the decimal part of the coordinate. Is the new image in The pixel value at the location.

[0086] Step S23: Backbone network.

[0087] The core function of the backbone network is to extract features from the input data. It consists of multiple convolutional layers, pooling layers, and activation functions. It can extract meaningful feature representations from the original data and simplify the model construction and training process. In order to achieve the best performance and speed, this application proposes to use a multi-dimensional backbone network as feature input. There are three types of network formats selected. Through a large number of experimental comparisons, MobileNet is used as the final backbone network model. Backbone network used:

[0088] ResNet series: including ResNet18, ResNet34, ResNet50, ResNet101, ResNet152, etc. These network structures are optimized and can extract rich features from image data.

[0089] MobileNet series: lightweight networks designed for mobile and embedded devices that reduce computation through depth-wise separable convolutions.

[0090] Transformer series: A network based on the self-attention mechanism, originally used for natural language processing, is now also applied to computer vision tasks.

[0091] Step S24: Transformer-based guidance point prediction.

[0092] The Transformer model of images is an architecture based on the self-attention mechanism, which was originally proposed in the field of natural language processing (NLP) and applied in the field of image processing (CV). The Transformer model consists of an encoder and a decoder, each of which contains multiple identical layers (blocks). The Transformer network hyperparameters are shown in Table 1.

[0093] 1. Input Embedding

[0094] In the image transformer, the input image is divided into multiple patches, and each patch is mapped to a fixed-dimensional vector space through a linear layer:

[0095] E(x)=Wx+b, where E is the embedding layer, x is the image patch, and W and b are learnable weights and biases.

[0096] 2. Positional Encoding.

[0097] Since Transformer itself does not have the ability to process sequence order, position encoding is required to provide the position information of each element in the sequence:

[0098] .

[0099] .

[0100] in, It's the location. is the dimension index, is the dimension of the model.

[0101] 3. Self-Attention Mechanism.

[0102] The self-attention mechanism allows the model to dynamically assign different attention weights between different positions in the sequence:

[0103] .

[0104] Among them, Q, K and V are query, key and value respectively. is the dimension of the key.

[0105] 4. Encoder.

[0106] The encoder consists of multiple identical layers, each including self-attention and feed-forward networks:

[0107] .

[0108] .

[0109] Where v represents the input vector, represents the output layer, represents the i-th coding layer.

[0110] Table 1 Transformer network hyperparameters

[0111]

[0112] 5. Extract the guide points.

[0113] For each keypoint in the heatmap, find the maximum value and the corresponding index position: Among them, max_valuesmax_values ​​is the maximum value in the heat map, and max_indices is the index of these maximum values, which are represented as follows in the side view and top view respectively: and , which together form the 3D coordinates of the guide point , the recognition results are as follows Figure 4 .

[0114] .

[0115] .

[0116] .

[0117] In step S30: local features are obtained by transforming the input features, and the number of search points is increased near the guide point. Finally, the search points are input into the multi-layer perceptron (MLP network layer) to predict the final key point fine-tuning amount. . Specifically include:

[0118] Step S31: With the guide point as the center, a local neighborhood (e.g., a fixed-size window or patch) can be defined. Various feature descriptors can be extracted within this neighborhood. The method of obtaining local features through the guide point can effectively capture local information in the image.

[0119] Guide point As the center, define a local neighborhood N. This neighborhood can be a fixed-size window or patch, for example, a k×k area. Mathematically, this neighborhood can be expressed as:

[0120] .

[0121] Step S32: To achieve flexible, efficient, and high-performance model fitting, this application proposes to use MLP (multi-layer perceptron) as the core model of the second stage. The basic structure of a multi-layer perceptron (MLP) consists of an input layer, a hidden layer, and an output layer, where there can be multiple hidden layers. Each layer consists of multiple neurons, and each neuron performs a weighted summation of the input and passes through a nonlinear activation function.

[0122] The formula is as follows:

[0123] .

[0124] in It is The weight from the i-th neuron in the layer to the j-th neuron in the l-th layer. It is The output of the i-th neuron in the layer. is the bias term of the jth neuron in the lth layer. n is the The number of neurons in the layer.

[0125] Neuron output can be calculated using ReLU or other nonlinear activation functions.

[0126] .

[0127] .

[0128] In step S40: Combine the guide point coordinates in S20 And S30 key point fine-tuning amount , get the key point coordinates predicted by the final model .

[0129] .

[0130] In step S50: collect data to train the model.

[0131] ,

[0132] is the loss sum of the Euclidean distance difference between the predicted point and the target point, is the difference in angular momentum and, and is the weight.

[0133] , , .

[0134] , Represents a direction vector.

[0135] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for recognizing 3D key points of a body styling surface is implemented.

[0136] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0137] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0138] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0139] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0140] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0141] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0142] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for identifying 3D key points of a vehicle body surface, characterized in that: include: Acquire images of the vehicle body to be tested from multiple angles as a set of images to be tested; Performing grayscale processing on the image group to be tested to obtain a grayscale image group to be tested; Inputting the grayscale image group to be tested into a two-stage key point recognition model to obtain 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested; The two-stage key point recognition model is obtained by training the initial two-stage key point recognition model using multiple annotated historical graph groups; the initial two-stage key point recognition model includes a backbone network, a Transformer model and a multi-layer perceptron; the output end of the backbone network is respectively connected to the input end of the Transformer model and the input end of the multi-layer perceptron; the input end of the multi-layer perceptron is also connected to the output end of the Transformer model; the backbone network is a MobileNet network; the Transformer model includes: an input embedding layer, a position encoding layer, a self-attention mechanism, an encoder and a decoder connected in sequence; the multi-layer perceptron includes: an input layer, a hidden layer and an output layer connected in sequence; The grayscale image group to be tested is input into the two-stage key point recognition model to obtain the 3D coordinate prediction values ​​of multiple key points of the vehicle body to be tested, including: Inputting the grayscale image group to be tested into the backbone network to obtain the grayscale image features to be tested; Inputting the grayscale image features into the Transformer model to obtain a guide point distribution heat map; Extracting predicted 3D coordinates of a plurality of guide points from a guide point distribution heat map; the guide points correspond one-to-one to the key points; According to the preset step length, the local features corresponding to each guiding point are extracted from the grayscale image features to be tested with the guiding point as the center; Inputting the local features corresponding to all the guiding points into the multi-layer perceptron to obtain the 3D coordinate fine-tuning amount of each guiding point; The sum of the predicted 3D coordinates of the same guide point and the 3D coordinate fine-tuning amount is determined as the 3D coordinate prediction value of the corresponding key point.

2. The method for identifying 3D key points of a vehicle body surface according to claim 1, characterized in that: The angles include side views and top views.

3. The method for identifying 3D key points of a vehicle body surface according to claim 1, characterized in that: Before acquiring images of the vehicle body to be tested from multiple angles as a set of images to be tested, the following steps are also included: Acquire multiple historical image groups; different historical image groups correspond to different vehicle bodies; any historical image group includes images corresponding to different angles of the vehicle body; Mark multiple key points in each historical map group separately and determine the actual 3D coordinates of each key point; Perform grayscale processing on each annotated historical image group to obtain multiple historical grayscale image groups; Using bilinear interpolation method, image interpolation processing is performed on multiple historical grayscale image groups to obtain multiple historical interpolation image groups; Taking the historical interpolation graph group as input and the multiple key points in the historical graph group as output, the initial two-stage key point recognition model is trained by using a loss function to obtain a two-stage key point recognition model.

4. The method for identifying 3D key points of a vehicle body surface according to claim 3, characterized in that: Determine the actual 3D coordinates of each keypoint, including: Let the key point number k=1; Get the side view coordinates of the kth key point and the top-down coordinates ; The side view coordinates are the 2D coordinates of the key points in the side view image of the vehicle body; the top view coordinates are the 2D coordinates of the key points in the top view image of the vehicle body; Determine the actual 3D coordinates of the kth key point based on the side view coordinates and top view coordinates of the kth key point ;in ; Increase the value of the key point number k by 1 and return to step "Get the side view coordinates of the kth key point and the top-down coordinates "Until the value of the key point number k reaches the total number of key points, the actual 3D coordinates of each key point are obtained.

5. The method for identifying 3D key points of a vehicle body surface according to claim 3, characterized in that: The loss function is: ; ; ; in, is the loss function; is the loss sum of the Euclidean distance difference between the predicted point and the target point; is the sum of the angular momentum differences; and All are weights; is the predicted value of the 3D coordinate of the kth key point; is the actual direction vector; is the prediction direction vector.

6. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for identifying 3D key points of a vehicle body shaping surface according to any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying 3D key points of a vehicle body shaping surface according to any one of claims 1 to 5 is implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying 3D key points of a vehicle body shaping surface according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Vehicle re-recognition method based on key point detection and local feature alignment

    CN112990152A

  • Target key point detection method and system

    CN118247519A