Object recognition method and device, computer equipment, storage medium and program product
By using point cloud data and object category data sets to train object recognition models in the driving environment of autonomous driving vehicles, the problem of low object recognition accuracy in traditional technology is solved and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202311755756.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-27
AI Technical Summary
In traditional technology, the accuracy of object recognition is low, making it difficult to accurately identify obstacles on the road in the driving environment of autonomous vehicles.
By obtaining point cloud data and object category data sets in the driving environment of the vehicle, combining the preset object recognition model, the sample training set is used to train the initial object recognition model to improve the accuracy of the recognition model.
It improves the accuracy of the recognition results of objects in point cloud data and enhances the object recognition capabilities of autonomous driving vehicles in driving environments.
Smart Images

Figure CN120220101A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of object recognition, and in particular, to an object recognition method, apparatus, computer device, storage medium, and program product. Background Art
[0002] Generally, in a driving scenario, in order to accurately plan the driving path of an autonomous vehicle and avoid obstacles on the driving road, it is necessary to recognize objects on the driving road of the autonomous vehicle.
[0003] In traditional technologies, point cloud data on the road is collected, and the collected point cloud data is used for object recognition.
[0004] However, there is a problem of low accuracy in object recognition in traditional technologies. Summary of the Invention
[0005] Embodiments of the present disclosure provide an object recognition method, apparatus, computer device, storage medium, and program product, which can be used to improve the accuracy of object recognition.
[0006] In a first aspect, embodiments of the present disclosure provide an object recognition method, the method comprising:
[0007] Obtain point cloud data and an object category dataset in the driving environment of a vehicle; the object category dataset includes text data of multiple different categories of objects;
[0008] According to the point cloud data, the object category dataset, and a preset object recognition model, obtain the recognition result of the object in the point cloud data; the object recognition model is obtained by training an initial object recognition model with a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and the object category dataset in the driving environment of the vehicle.
[0009] In a second aspect, embodiments of the present disclosure provide a method for training an object recognition model, the method comprising:
[0010] Obtain a training sample set; the sample training set includes sample image data, sample point cloud data, and an object category dataset in the driving environment of a vehicle; the object category dataset includes text data of multiple different categories of objects;
[0011] Train an initial object recognition model with the sample training set to obtain an object recognition model.
[0012] In a third aspect, embodiments of the present disclosure provide a training apparatus for an object recognition model, the apparatus comprising:
[0013] A first acquisition module, configured to acquire point cloud data and an object category data set in a traffic tool driving environment; the object category data set includes text data of multiple different category objects.
[0014] A second acquisition module, configured to obtain an identification result of an object in the point cloud data according to the point cloud data, the object category data set, and a preset object recognition model; the object recognition model is obtained by training an initial object recognition model with a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and the object category data set in the traffic tool driving environment.
[0015] Fourthly, an embodiment of the present disclosure provides a training device for an object recognition model, and the device includes:
[0016] A third acquisition module, configured to acquire a training sample set; the sample training set includes sample image data, sample point cloud data, and an object category data set in a traffic tool driving environment; the object category data set includes text data of multiple different category objects.
[0017] A fourth acquisition module, configured to train an initial object recognition model with the sample training set to obtain an object recognition model.
[0018] Fifthly, an embodiment of the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the methods described in the first aspect and the second aspect are implemented.
[0019] Sixthly, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the methods described in the first aspect and the second aspect are implemented.
[0020] Seventhly, an embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the methods described in the first aspect and the second aspect are implemented.
[0021] The object recognition method, device, computer device, storage medium, and program product provided by the embodiments of the present disclosure. Since the object recognition model is obtained by training an initial object recognition model with a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and object category data sets in the vehicle driving environment, during the training process of the initial object recognition model, the recognition process of the sample point cloud data can be trained according to the sample image data and the object category data sets, which can improve the accuracy of training the initial object recognition model, thereby enabling the recognition accuracy of the obtained object recognition model. Furthermore, by obtaining the point cloud data in the vehicle driving environment and the object category data set including text data of various different types of objects, according to the point cloud data, the object category data set, and the preset object recognition model, the accuracy of the recognition result of the objects in the obtained point cloud data can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic flowchart of the object recognition method in one embodiment;
[0023] Figure 2 It is a schematic flowchart of the object recognition method in another embodiment;
[0024] Figure 3 It is a schematic flowchart of the object recognition method in another embodiment;
[0025] Figure 4 It is a schematic flowchart of the object recognition method in another embodiment;
[0026] Figure 5 It is a schematic flowchart of the object recognition method in another embodiment;
[0027] Figure 6 It is a schematic flowchart of the object recognition method in another embodiment;
[0028] Figure 7 It is a schematic flowchart of the object recognition method in another embodiment;
[0029] Figure 8 It is a schematic flowchart of the object recognition method in another embodiment;
[0030] Figure 9 It is a schematic flowchart of the object recognition method in another embodiment;
[0031] Figure 10 It is a schematic flowchart of the training method of the object recognition model in one embodiment;
[0032] Figure 11 It is a schematic flowchart of the training method of the object recognition model in another embodiment;
[0033] Figure 12 It is an architecture diagram for training an initial object recognition model in an embodiment;
[0034] Figure 13 It is a structural block diagram of an object recognition device in an embodiment;
[0035] Figure 14 It is a structural block diagram of an object recognition device in another embodiment;
[0036] Figure 15 It is a structural block diagram of a training device for an object recognition model in an embodiment;
[0037] Figure 16 It is a structural block diagram of a training device for an object recognition model in another embodiment;
[0038] Figure 17 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0039] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following further describes the embodiments of the present disclosure in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present disclosure, and are not used to limit the embodiments of the present disclosure.
[0040] First, before specifically introducing the technical solutions of the embodiments of the present disclosure, the technical background or the technical evolution context on which the embodiments of the present disclosure are based is introduced. Generally, in the field of autonomous driving, the current technical background is as follows: converting the point cloud data collected by a vehicle into image data, and then using an image segmentation model to identify the image data, and using the recognition result as the object recognition result corresponding to the point cloud data, resulting in a problem of low accuracy in recognizing objects on the driving road of an autonomous driving vehicle. Based on this background, through long-term model simulation research and development, as well as the collection, demonstration, and verification of experimental data, the applicant found that aligning point cloud data and image data in a low-dimensional space of pixels and points and then using an image segmentation model for recognition will result in a problem of low accuracy of object recognition results. How to achieve the alignment of point cloud data and image data in a high-dimensional space has become a difficult problem to be solved urgently. In addition, it should be noted that from determining the alignment of point cloud data and image data in a high-dimensional space and the technical solutions introduced in the following embodiments, the applicant has put in a lot of creative labor.
[0041] Next, the technical solutions related to the embodiments of the present disclosure are introduced in combination with the scenarios to which the embodiments of the present disclosure are applied.
[0042] In one embodiment, as Figure 1As shown, an object recognition method is provided. In this embodiment, taking the application of this method to a computer device as an example, it can be understood that this method can also be applied to a server, or to a system including a computer device and a server, and is implemented through the interaction between the computer device and the server. In this embodiment, the method includes the following steps:
[0043] S201, obtain point cloud data and an object category dataset in the driving environment of a vehicle; the object category dataset includes text data of multiple different categories of objects.
[0044] In a vehicle driving scenario, in order to plan the driving path of a vehicle, the lidar on the vehicle can be controlled to obtain the point cloud data of the vehicle on the driving road. By identifying the point cloud data, the driving path of the vehicle can be planned according to the identification result. Among them, the vehicle can be an autonomous vehicle, or the vehicle can also be a non-autonomous vehicle. During the driving process of the vehicle, road signs, traffic lights, other vehicles, etc. on the driving road belong to the driving environment. In this embodiment, the point cloud data in the driving environment of the vehicle can be obtained through the lidar during the driving process of the vehicle. Optionally, the point cloud data corresponding to a time point can be obtained, or the point cloud data corresponding to multiple time points can also be obtained.
[0045] In this embodiment, the various object categories identified within a historical time period can be used as the object category dataset, or the various object categories involved in the driving environment of the vehicle can also be used as the object category dataset. Among them, the object category dataset includes text data corresponding to multiple different categories.
[0046] S202, obtain the recognition result of the object in the point cloud data according to the point cloud data, the object category dataset, and a preset object recognition model; the object recognition model is obtained by training an initial object recognition model with a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and an object category dataset in the driving environment of the vehicle.
[0047] Among them, the preset object recognition model is a neural network model for identifying point cloud data. Optionally, the object recognition model can be a Contrastive Language-Image Pretraining (CLIP) model, or the object recognition model can also be a model based on a convolutional neural network. In this embodiment, the object recognition model can be obtained by training the initial object recognition model with a preset sample training set.
[0048] It should be noted that the initial object recognition model is a neural network model for images. To improve the recognition accuracy of the initial object recognition model for point cloud data, during the training process, it is necessary to use sample image data, object category datasets, and sample point cloud data as a sample training set to train the initial object recognition model. Among them, the sample image data can be images corresponding to the driving environment of a vehicle obtained by using an in-vehicle camera device during a historical time period, and the sample point cloud data can be clouds corresponding to the vehicle's imaging environment obtained by using an in-vehicle lidar during a historical time period. It can be understood that during the training of the initial object recognition model, the same object category dataset as that used when obtaining the recognition result can be selected as the object category training set during the training process.
[0049] In this embodiment, the point cloud data and the object category dataset can be input into the object recognition model, and the object recognition model uses the object category dataset to perform object recognition on the point cloud data to obtain the recognition results of each object in the point cloud data. It can be understood that if there are multiple objects in the point cloud data, the object recognition model can obtain the recognition results corresponding to multiple objects.
[0050] In the above object recognition method, since the object recognition model is obtained by training the initial object recognition model using a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and object category datasets in the vehicle driving environment, during the training of the initial object recognition model, the recognition process of the sample point cloud data can be trained according to the sample image data and the object category datasets, which can improve the accuracy of training the initial object recognition model, thereby enabling the recognition accuracy of the obtained object recognition model. Furthermore, by obtaining the point cloud data in the vehicle driving environment and the object category dataset including text data of various different categories of objects, according to the point cloud data, the object category dataset, and the preset object recognition model, the accuracy of the recognition results of the objects in the obtained point cloud data can be improved.
[0051] In the scenario of obtaining the recognition results of the objects in the point cloud data according to the point cloud data, the object category dataset, and the preset object recognition model, the point cloud data and the object category dataset can be input into the object recognition model to obtain the recognition results of the objects in the point cloud data. In one embodiment, the above S202 includes: inputting the point cloud data and the object category dataset into the object recognition model to perform point cloud feature recognition and text feature recognition to obtain the recognition results of the objects in the point cloud data.
[0052] Among them, the point cloud feature is the feature corresponding to the point cloud data, and the text feature is the feature corresponding to the text data. In this embodiment, the point cloud feature can be obtained by mapping the input point cloud data through an object recognition model, and the text feature set can be obtained by mapping the input object category data set including text data. It can be understood that the text feature set includes text features corresponding to multiple object categories.
[0053] In this embodiment, the point cloud data and the object category data set can be respectively input into the object recognition model. The object recognition model maps the point cloud data into point cloud features and maps the text data into text features. Then, the text features are used as the recognition weights for the object recognition model to recognize the point cloud features, and the recognition result of the point cloud features is obtained. Thus, the recognition result of the point cloud features is used as the recognition result of the object in the point cloud data.
[0054] In this embodiment, since the text features corresponding to the object category data set and the point cloud features corresponding to the point cloud data are recognized, the point cloud data can be recognized from the dimension of the feature space. Therefore, when the point cloud data and the object category data set are input into the object recognition model, the accuracy of the recognition result of the object in the point cloud data can be improved.
[0055] Next, the specific process of inputting the point cloud data and the object category data set into the object recognition model to perform point cloud feature recognition and text feature recognition to obtain the recognition result of the object in the point cloud data will be described in detail. In one embodiment, the object recognition model includes a point cloud encoder, a text encoder, and a point cloud label generation network, as Figure 2 shown. The above "inputting the point cloud data and the object category data set into the object recognition model to perform point cloud feature recognition and text feature recognition to obtain the recognition result of the object in the point cloud data" includes:
[0056] S301, input the object category data set into the text encoder to obtain the text feature set corresponding to the object category data set.
[0057] Among them, the text encoder is a neural network used to map text data into a high-dimensional feature space to obtain text features. The text encoder may include an encoder, an encoding space, and a decoder, and the text encoder may be a CLIP model. In this embodiment, the object category data set can be input into the text encoder, and through the text encoder, the object category data set can be mapped into the feature space to obtain the text feature set corresponding to the object category data set.
[0058] As an alternative implementation, the object recognition model may further include a predefined category template. The object category dataset can be input into the category template to generate a prompt word corresponding to each object category, and then each prompt word is input into the text encoder to obtain the corresponding text feature, so as to determine multiple text features as the text feature set corresponding to the object category dataset.
[0059] S302: Input the point cloud data into the point cloud encoder to obtain the point cloud feature map corresponding to the point cloud data.
[0060] Among them, the point cloud encoder is a neural network used to map the point cloud data into a high-dimensional feature space to obtain the point cloud feature. The point cloud encoder may include an encoder, an encoding space, and a decoder. The point cloud encoder may be a CLIP model. In this embodiment, the point cloud data can be input into the point cloud encoder, and through the point cloud encoder, the input point cloud data can be mapped into the feature space to obtain the corresponding point cloud feature, and the obtained point cloud feature is determined as the point cloud feature map corresponding to the point cloud data.
[0061] S303: Input the point cloud feature map and the text feature set into the point cloud label generation network to obtain the recognition result of the object in the point cloud data.
[0062] Among them, the point cloud label generation network is a neural network used to recognize the object in the point cloud data and output a label including the recognition result. In this embodiment, the point cloud feature map and the text feature set can be input into the point cloud label generation network. The point cloud label generation network calculates the similarity between each text feature in the text feature set and the point cloud feature respectively, determines the maximum similarity as the recognition result corresponding to the point cloud feature, and uses the text feature corresponding to the maximum similarity as the label of the point cloud recognition result, so as to output the recognition result of the object in the point cloud data.
[0063] As an alternative implementation, the point cloud label generation network may include a recognition sub-network and a generation sub-network. Among them, the recognition sub-network is used to obtain the recognition result corresponding to the point cloud feature map, and the generation sub-network is used to obtain the label of the recognition result according to the text feature. Optionally, the recognition sub-network may be a semantic segmentation model, or the recognition sub-network may also be an instance segmentation model.
[0064] In this embodiment, by inputting the object category dataset into the text encoder to obtain the text feature set corresponding to the object category dataset, and inputting the point cloud data into the point cloud encoder to obtain the point cloud feature map corresponding to the point cloud data, the point cloud feature map and the text feature set can be input into the point cloud label generation network to obtain the recognition result of the object in the point cloud data. Since in the recognition process, the object recognition model can be used to recognize the point cloud data from the dimension of the feature space, the accuracy of the recognition result of the object in the obtained point cloud data can be improved.
[0065] The training process of obtaining an object recognition model by training according to an initial object recognition model will be described in detail below. In one embodiment, as Figure 3 shown, the training process of the object recognition model includes:
[0066] S401, input an object category data set into the text encoder of the initial object recognition model to obtain a text feature set corresponding to the object category data set.
[0067] Among them, the text encoder is a neural network used to map text data into a high-dimensional feature space to obtain text features. In this embodiment, the object category data set can be input into the text encoder, and through the text encoder, the object category data set can be mapped into the feature space to obtain a text feature set corresponding to the object category data set.
[0068] S402, input sample image data into the image encoder of the initial object recognition model to obtain a first feature map corresponding to the sample image data.
[0069] It should be noted that when using the object recognition model to recognize point cloud data, since there is no clear correlation between text data and point cloud data, it is necessary to align the image data and the point cloud data in the feature space, so as to establish the correlation between the point cloud data and the sample data, and then train the initial object recognition model according to the alignment loss, and adjust the parameters in the initial object recognition model to obtain an object recognition model that can be used for point cloud data. In this embodiment, an image encoder can be set in the initial object recognition model to map the image data into a high-dimensional feature space to obtain an image feature map.
[0070] In this embodiment, image data in the driving environment of a vehicle taken during a historical time period can be obtained as sample image data, and the sample image data is input into the image encoder of the initial object recognition model. Through the image encoder, the sample image data can be mapped into a sample image feature map, and the sample image feature is determined as the first feature map. Optionally, the sample image data can be image data corresponding to the same vehicle driving environment, or image data corresponding to different vehicle driving environments.
[0071] S403, input sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain a second feature map corresponding to the sample point cloud data.
[0072] In this embodiment, point cloud data collected in the driving environment of a vehicle during a historical time period can be obtained as sample point cloud data. The sample point cloud data is input into the initial point cloud encoder of the initial object recognition model. Through the initial point cloud encoder, the sample point cloud data can be mapped into sample point cloud features, and the sample point cloud features are determined as the second feature map. Optionally, the sample point cloud data can be point cloud data corresponding to the driving environment of the same vehicle, or point cloud data corresponding to the driving environments of different vehicles.
[0073] It should be noted that during the training process of the initial object recognition model, the obtained sample image data and sample point cloud data are corresponding. For example, the sample image data is the image data at a certain time point on Road A, and the sample point cloud data is also the point cloud data at the same time point on Road A.
[0074] S404. Train the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain an object recognition model.
[0075] In this embodiment, the text feature set can be used as the recognition weight for the initial object recognition model to recognize the first feature map and as the recognition weight for the initial object recognition model to recognize the second feature map respectively, so as to obtain the recognition result corresponding to the first feature map and the recognition result corresponding to the second feature map. Then, the initial point cloud encoder is trained according to the recognition results, and the trained initial point cloud encoder is used as the point cloud encoder in the object recognition model, thereby obtaining the object recognition model.
[0076] In this embodiment, during the training process, the object category data set is input into the text encoder of the initial object recognition model to obtain the text feature set corresponding to the object category data set. The sample image data is input into the image encoder of the initial object recognition model to obtain the first feature map corresponding to the sample image data. The sample point cloud data is input into the initial point cloud encoder of the initial object recognition model to obtain the second feature map corresponding to the sample point cloud data. The initial point cloud encoder can be trained according to the text feature set, the first feature map, and the second feature map. Since the features of the object category data set, the sample image, and the point cloud image in the high-dimensional feature space are used to train the initial point cloud encoder during the training process, the accuracy of the trained initial point cloud encoder is improved. Thus, a second feature map with higher accuracy can be obtained, and further, the recognition accuracy of the object recognition model for recognizing point cloud data according to the second feature map can be improved, and the accuracy of the obtained recognition result can be improved.
[0077] In one embodiment, the initial object recognition model further includes an image label generation network and a point cloud label generation network. As Figure 4 shown, the above S404 includes:
[0078] S501. Input the text feature set and the first feature map into the image label generation network to obtain the first label image corresponding to the sample image data; the first label image includes the labels of each object in the sample image data.
[0079] Among them, the image label generation network is a neural network used to identify image data based on image features and obtain labels including the recognition results. In this embodiment, the recognition result corresponding to the first feature map can be obtained using the image label generation network. It can be understood that if the sample image includes multiple objects, the first label image may include the labels of multiple objects in the sample image.
[0080] In this embodiment, the text feature set and the first feature map can be input into the image label generation network. In the image label generation network, the similarity between each text feature in the text feature set and the first feature map can be calculated respectively. The one with the maximum similarity is used as the recognition result of the first feature map, and the text feature corresponding to the maximum similarity is used as the label of the recognition result of the first feature map. The label of the recognition result of the first feature map is used as the first label image corresponding to the sample image.
[0081] S502. Map the first label image into the three-dimensional space to obtain the second label image corresponding to the first label image.
[0082] It should be noted that the first label image is an image in the two-dimensional space. To align the output results of the sample image data and the sample point cloud data, the first label image can be mapped into the three-dimensional space, and the mapped first label image is determined as the second label image.
[0083] In this embodiment, according to the principle of projection, each pixel point in the first label image can be projected into the three-dimensional space, and all the pixel points in the three-dimensional space obtained after projection are used as the second label image. It can be understood that
[0084] S503. Input the text feature set and the second feature map into the point cloud label generation network to obtain the third label image corresponding to the sample point cloud data; the third label image includes the labels of each object in the sample point cloud data.
[0085] Among them, the point cloud label generation network is a neural network used to identify objects in point cloud data and output labels including the recognition results. It can be understood that if the sample point cloud includes multiple objects, the third label image may include the labels of multiple objects in the sample point cloud.
[0086] In this embodiment, the second feature map of the text feature set can be input into the point cloud label generation network. In the point cloud label generation network, the similarity between each text feature in the text feature set and the second feature map can be calculated respectively. The one with the maximum similarity is used as the recognition result of the second feature map, and the text feature corresponding to the maximum similarity is used as the label of the recognition result of the second feature map. The label of the recognition result of the second feature map is used as the third label image corresponding to the sample point cloud.
[0087] S504. Train the initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image to obtain an object recognition model.
[0088] In this embodiment, according to a preset loss function, the loss function value between the first feature map and the second feature map can be calculated using the first label image and the third label image, and the loss function value between the second label image and the third label image can be calculated, so as to train the initial point cloud encoder according to the calculated loss function value to obtain a trained initial object recognition model, and further obtain an object recognition model.
[0089] In this embodiment, by inputting the text feature set and the first feature map into the image label generation network, the first label image corresponding to the sample image data and including each object in the sample image data is obtained, and the first label image is mapped into the three-dimensional space to obtain the second label image corresponding to the first label image. In addition, by inputting the text feature set and the second feature map into the point cloud label generation network, the third label image corresponding to the sample point cloud data and including each object in the sample point cloud data is obtained. It is possible to train the initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image to obtain an object recognition model. Since the initial point cloud encoder is trained according to the first feature map and the second feature map in the feature space, and the first label image, the second label image, and the third label image in the output result space, the accuracy of the initial point cloud encoder can be improved, thereby improving the accuracy of the obtained second feature map, and further improving the recognition accuracy of the object recognition model for identifying point cloud data according to the second feature map, and improving the accuracy of the obtained recognition result.
[0090] The following specifically describes the detailed process of how to train the initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image to obtain an object recognition model. In one embodiment, as Figure 5 shown, the above S504 includes:
[0091] S601. Obtain the value of the first loss function according to the first feature map, the second feature map, the first label image, and the third label image.
[0092] It should be noted that when aligning the feature dimensions of image data and point cloud data, feature alignment can be performed from two aspects: semantics and space. In this embodiment, the first loss function can be used as the feature alignment loss between the image data and the point cloud data from the semantic perspective. Optionally, the first loss function can be a contrast loss function, or the first loss function can also be a class prototype loss function.
[0093] In this embodiment, the contrast loss value corresponding to the sample image data can be calculated by using the first label image and the first feature map, and the contrast loss value corresponding to the sample point cloud data can be calculated by using the third label image and the second feature map. Thus, the sum of the contrast loss value corresponding to the sample image data and the contrast loss value corresponding to the sample point cloud data is used as the value of the first loss function.
[0094] As an alternative implementation, in this embodiment, the value of the first loss function can be obtained by performing masked pooling on the first feature map and the second feature map respectively. As Figure 6 shown, the above S601 can include:
[0095] S701, obtaining a two-dimensional mask map corresponding to the sample image data according to the first label image and the text feature set.
[0096] Among them, the two-dimensional mask map can be used to represent the correspondence between the sample image data and the object category dataset. It can be understood that if the first label image includes labels of multiple objects and the text feature set includes text features corresponding to multiple categories, the two-dimensional mask map can include multiple two-dimensional mask maps, respectively representing the correspondence between each object and each category.
[0097] In this embodiment, the two-dimensional mask map corresponding to the sample image data can be obtained by calculating the element-wise multiplication product between the first label image and the text feature set. Exemplarily, if the first label image is a 2*2 matrix [1, 1, 1, 2] and the first text feature is a 2*1 matrix [1, 1], then the two-dimensional mask map is a 2*2 matrix [1, 1, 1, 0].
[0098] S702, obtaining a three-dimensional mask map corresponding to the sample point cloud data according to the third label image and the text feature set.
[0099] Among them, the three-dimensional mask map can be used to represent the correspondence between the sample image data and the object category dataset. It can be understood that if the third label image includes labels of multiple objects and the text feature set includes text features corresponding to multiple categories, the three-dimensional mask map can include multiple three-dimensional mask maps, respectively representing the correspondence between each object and each category.
[0100] In this embodiment, the three-dimensional mask map corresponding to the sample point cloud data can be obtained by calculating the product of the element-wise multiplication between the third label image and the text feature set.
[0101] S703. Determine the first embedding values of the objects of various categories in the object category dataset in the first feature map according to the first feature map and the two-dimensional mask map.
[0102] Among them, the first embedding value refers to the mapping relationship between the first feature map corresponding to the sample image data and the two-dimensional mask map. According to the first embedding value, the proportion of the text features corresponding to each category in each object in the first feature map can be characterized.
[0103] In this embodiment, the first feature map and the two-dimensional mask map can be input into a preset mask pooling formula, and the calculation result of this formula can represent the first embedding value. This formula can be:
[0104]
[0105] In the formula, m can represent a two-dimensional space or a three-dimensional space. When m represents a two-dimensional space, represents the first embedding value, represents the first feature map, represents the two-dimensional mask map, and i represents the processing stage of the image encoder.
[0106] S704. Determine the second embedding values of the objects of various categories in the object category dataset in the second feature map according to the second feature map and the three-dimensional mask map.
[0107] Among them, the second embedding value refers to the mapping relationship between the second feature map corresponding to the sample point cloud data and the three-dimensional mask map. According to the second embedding value, the proportion of the text features corresponding to each category in each object in the second feature map can be characterized.
[0108] In this embodiment, the second feature map and the three-dimensional mask map can be input into a preset mask pooling formula, and the calculation result of this formula can represent the second embedding value. This formula can be:
[0109]
[0110] In the formula, m can represent a two-dimensional space or a three-dimensional space. When m represents a three-dimensional space, represents the second embedding value, represents the second feature map, represents the three-dimensional mask map, and i represents the processing stage of the initial point cloud encoder.
[0111] S705. Obtain the value of the first loss function according to the first embedding values of objects of various categories in the first feature map and the second embedding values of objects of various categories in the second feature map.
[0112] In this embodiment, the first embedding value can be input into a preset first loss function to obtain the loss value corresponding to the first embedding value, and the second embedding value can be input into the preset first loss function to obtain the loss value corresponding to the second embedding value. Thus, the sum of the loss value corresponding to the first embedding value and the loss value corresponding to the second embedding value is used as the value of the first loss function.
[0113] As an alternative implementation, as Figure 7 shown, the above S705 may include:
[0114] S801. Obtain the first similarity between the first embedding values of objects of various categories in the first feature map and the second embedding values of objects of various categories in the second feature map respectively.
[0115] In this implementation, the first similarity between the first embedding value corresponding to each category of object and each second embedding value can be calculated respectively. It can be understood that if both the first embedding value and the second embedding value include multiple embedding values, the first similarity includes multiple similarities. Exemplarily, if the object category dataset includes 3 categories, the first embedding value includes 3 embedding values: A1, B1, C1, and the second embedding value includes 3 embedding values: A2, B2, C2. The similarities between A1 in the first embedding value and A2, B2, C2, the similarities between B1 in the first embedding value and A2, B2, C2, and the similarities between C1 in the first embedding value and A2, B2, C2 can be calculated respectively, and all the similarities are determined as the first similarity.
[0116] S802. Obtain the value of the first loss function according to the sum of each first similarity and a preset contrast loss function.
[0117] In this implementation, the sum value of each first similarity can be calculated, and then the sum value is brought into the preset contrast loss function, and the result of the contrast loss function is determined as the value of the first loss function.
[0118] S602. Obtain the value of the second loss function according to the first feature map and the second feature map.
[0119] In this embodiment, the second loss function can be used as the feature alignment loss of the image data and the point cloud data in the spatial angle. Optionally, the second loss function can be a contrast loss function, or the second loss function can also be a grid loss function.
[0120] In this embodiment, features corresponding to the same spatial position of the first feature map and the second feature map can be extracted, so as to calculate the similarity between the features corresponding to the same spatial position of the two, and then determine the value of the second loss function according to the similarity and the preset second loss function.
[0121] S603. Obtain the value of the third loss function according to the second label image and the third label image.
[0122] It should be noted that both the second label image and the third label image are label images in a three-dimensional space. The second label image and the third label image may be the same or different. Therefore, the initial point cloud encoder can be trained according to the loss value between the second label image and the third label image.
[0123] Among them, the value of the third loss function can characterize the alignment loss of the sample image data and the sample point cloud data on the output result of the initial object recognition model. In this embodiment, the similarity between the second label image and the third label image corresponding to the same object can be calculated, and then the value of the third loss function can be determined according to the similarity and the preset third loss function.
[0124] S604. Train the initial point cloud encoder according to the value of the first loss function, the value of the second loss function, and the value of the third loss function to obtain an object recognition model.
[0125] In this embodiment, the sum value of the value of the first loss function, the value of the second loss function, and the value of the third loss function can be used to adjust the parameters of the initial point cloud encoder, so as to train the initial point cloud encoder. The trained initial point cloud encoder is determined as the point cloud encoder, and then the object recognition model can be obtained according to the determined point cloud encoder.
[0126] In this embodiment, by obtaining the value of the first loss function according to the first feature map, the second feature map, the first label image, and the third label image, obtaining the value of the second loss function according to the first feature map and the second feature map, and obtaining the value of the third loss function according to the second label image and the third label image, the initial point cloud encoder can be trained according to the value of the first loss function, the value of the second loss function, and the value of the third loss function. Since the initial point cloud encoder is trained respectively by the value of the loss function of the feature maps corresponding to the sample image data and the sample point cloud data, and the value of the loss function of the corresponding results, the initial point cloud encoder can be trained respectively from the feature dimension and the output result dimension, so as to improve the accuracy of the second feature map output by the initial point cloud encoder, and further improve the recognition accuracy of the object recognition model for recognizing point cloud data according to the second feature map, and improve the accuracy of the obtained recognition result.
[0127] In the scenario of obtaining the value of the second loss function based on the first feature map and the second feature map as described above, the value of the second loss function can be obtained by obtaining the similarity between the grid feature map corresponding to the first feature map and the grid feature map corresponding to the second feature map and the contrast loss function. In one embodiment, as Figure 8 shown, S602 described above may include:
[0128] S901, mapping the second feature map to a two-dimensional space to obtain a third feature map corresponding to the second feature map.
[0129] It should be noted that the second feature map is the feature map in the three-dimensional space corresponding to the point cloud data. In order to calculate the value of the second loss function, the second feature map can be mapped to a two-dimensional space to obtain a third feature map in the two-dimensional space.
[0130] In this embodiment, the second feature map can be mapped into the two-dimensional space according to the principle of projection to obtain a third feature map corresponding to the second feature map.
[0131] S902, respectively perform grid division on the first feature map and the third feature map to obtain a first grid feature map corresponding to the first feature map and a second grid feature map corresponding to the second feature map.
[0132] In this embodiment, the first feature map and the second feature map can be respectively subjected to grid division according to the same division rule, and the divided grid maps are respectively determined as the first grid feature map corresponding to the first feature map and the second grid feature map corresponding to the second feature map. Exemplarily, the first feature map and the second feature map can both be divided into an m*n grid map, then an m*n first grid feature map and an m*n second grid feature map can be obtained.
[0133] S903, obtain a second similarity between the first grid feature map and the second grid feature map.
[0134] In this embodiment, some grids can be selected from the first grid feature map according to a preset selection rule, and some grids can be selected from the second grid feature map according to the same selection rule, and the second similarity between the first grid feature map and the second grid feature map corresponding to the grids at the same position is calculated respectively. It should be noted that if the selected first grid feature map and second grid feature map include grids at multiple positions, the second similarity includes multiple similarities. For example, the grids in the first 5 rows * 5 columns of the first grid feature map and the second grid feature map can be respectively selected, where both m and n are greater than 5.
[0135] As an alternative embodiment, as Figure 9 shown, S903 described above may include:
[0136] S1001. Obtain the third similarity between each grid feature in the first grid feature map and each grid feature in the second grid feature map.
[0137] S1002. Determine the sum of each third similarity as the second similarity.
[0138] It should be noted that the third similarity includes multiple similarities. In this embodiment, in order to improve the accuracy of the obtained third similarity, the similarity between each grid in the first grid feature map and the second grid feature map can be calculated respectively, so as to obtain the third similarity between each grid feature in the first grid feature map and each grid feature in the second grid feature map; further, the sum of each third similarity can be calculated, and the obtained sum value can be determined as the second similarity.
[0139] S904. Obtain the value of the second loss function according to the second similarity and a preset contrast loss function.
[0140] In this embodiment, the second similarity can be substituted into the preset contrast loss function to calculate the result of the contrast loss function, and the result of the contrast loss function can be determined as the value of the second loss function.
[0141] In this embodiment, by mapping the second feature map to a two-dimensional space, a third feature map corresponding to the second feature map is obtained, and the first feature map and the third feature map are respectively divided into grids to obtain the first grid feature map corresponding to the first feature map and the second grid feature map corresponding to the second feature map. Thus, the second similarity between the first grid feature map and the second grid feature map can be obtained, and the value of the second loss function can be obtained according to the second similarity and the preset contrast loss function. By dividing the first feature map and the third feature map into grids, the loss value between the first feature map and the second feature map can be calculated from the perspective of space, so as to improve the accuracy of the determined value of the second loss function, and further improve the accuracy of the initial point cloud encoder trained according to the value of the second loss function.
[0142] In one embodiment, a training method for an object recognition model is further provided. As Figure 10 shown, taking the application of this method to a computer device as an example for description, it includes the following steps:
[0143] S1101. Obtain a training sample set; the sample training set includes sample image data, sample point cloud data, and an object category data set in the driving environment of a vehicle; the object category data set includes text data of multiple different categories of objects.
[0144] Among them, the training sample set is the sample set for training the initial object recognition model. The training sample set may include sample image data, sample point cloud data, and object category data set in the driving environment of the vehicle. In this embodiment, the point cloud data collected in the historical time period in the driving environment of the vehicle can be obtained as the sample point cloud data, and the image data collected in the historical time period in the driving environment of the vehicle can be obtained as the sample image data, and various object categories recognized in the historical time period can be obtained as the object category data set.
[0145] S1102. Train the initial object recognition model using the sample training set to obtain an object recognition model.
[0146] In this embodiment, the sample training set can be input into the initial object recognition model to obtain a sample recognition result, and the initial object recognition model can be iteratively trained according to the sample recognition result, so as to obtain a trained initial object recognition model, and then the trained initial object recognition model is determined as the object recognition model.
[0147] In this embodiment, by obtaining the training sample set and training the initial object recognition model using the sample training set, since the sample training set includes sample image data, sample point cloud data, and object category data set in the driving environment of the vehicle, it is possible to determine the recognition results of the sample image data and the sample point cloud data by combining the text data of various different category objects in the object category data set during the training process of the initial object recognition model, thereby improving the recognition accuracy of the initial object recognition model, and further improving the recognition accuracy of the obtained object recognition model.
[0148] The following specifically describes the detailed process of training the initial object recognition model using the sample training set to obtain an object recognition model. In one embodiment, as Figure 11 shown, the above S1102 includes:
[0149] S1201. Input the object category data set into the text encoder of the initial object recognition model to obtain the text feature set corresponding to the object category data set.
[0150] Among them, the text encoder is a neural network for mapping text data into a high-dimensional feature space to obtain text features. In this embodiment, the object category data set can be input into the text encoder, and the object category data set can be mapped into the feature space through the text encoder to obtain the text feature set corresponding to the object category data set.
[0151] S1202. Input the sample image data into the image encoder of the initial object recognition model to obtain the first feature map corresponding to the sample image data.
[0152] Among them, the image encoder is a neural network used to map the sample image data into a high-dimensional feature space to obtain the sample image feature map. In this embodiment, the sample image data can be input into the image encoder of the initial object recognition model, and through the image encoder, the sample image data can be mapped into the feature space to obtain the first feature map corresponding to the sample image data set.
[0153] S1203. Input the sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain the second feature map corresponding to the sample point cloud data.
[0154] Among them, the initial point cloud encoder is a neural network used to map the sample point cloud data into a high-dimensional feature space to obtain the sample point cloud features. In this embodiment, the sample point cloud data can be input into the initial point cloud encoder of the initial object recognition model, and through the initial point cloud encoder, the sample point cloud data can be mapped into the feature space to obtain the second feature map corresponding to the sample point cloud data set.
[0155] S1204. Train the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain the object recognition model.
[0156] In this embodiment, the text feature set can be used as the recognition weight for the initial object recognition model to recognize the first feature map and as the recognition weight for the initial object recognition model to recognize the second feature map respectively, to obtain the recognition result corresponding to the first feature map and the recognition result corresponding to the second feature map. Then, train the initial point cloud encoder according to the recognition results, and use the trained initial point cloud encoder as the point cloud encoder in the object recognition model, thereby obtaining the object recognition model.
[0157] Exemplarily, the structural schematic diagram of the initial object recognition model can be as Figure 12 shown. In the figure, the initial object recognition model can obtain the first label image and the third label image according to the object category data set, the sample image data, and the sample point cloud data respectively, and then train the initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image.
[0158] In this embodiment, during the training process, the object category dataset is input into the text encoder of the initial object recognition model to obtain the text feature set corresponding to the object category dataset. The sample image data is input into the image encoder of the initial object recognition model to obtain the first feature map corresponding to the sample image data. The sample point cloud data is input into the initial point cloud encoder of the initial object recognition model to obtain the second feature map corresponding to the sample point cloud data. The initial point cloud encoder can be trained based on the text feature set, the first feature map, and the second feature map. Since the features of the object category dataset, the sample image, and the point cloud image in the high-dimensional feature space are utilized during the training process to train the initial point cloud encoder, the accuracy of the trained initial point cloud encoder is improved. Thus, a second feature map with higher accuracy can be obtained, and further, the recognition accuracy of the object recognition model for recognizing the point cloud data based on the second feature map can be improved, and the accuracy of the obtained recognition result can be enhanced.
[0159] For the convenience of understanding by those skilled in the art, the object recognition method provided in this application will be introduced in detail below. The method may include:
[0160] S1. Obtain the sample image data, the sample point cloud data, and the object category dataset in the driving environment of the vehicle. The object category dataset includes the text data of multiple different categories of objects.
[0161] S2. Input the object category dataset into the text encoder of the initial object recognition model to obtain the text feature set corresponding to the object category dataset.
[0162] S3. Input the sample image data into the image encoder of the initial object recognition model to obtain the first feature map corresponding to the sample image data.
[0163] S4. Input the sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain the second feature map corresponding to the sample point cloud data.
[0164] S5. Input the text feature set and the first feature map into the image label generation network to obtain the first label image corresponding to the sample image data.
[0165] S6. Map the first label image into the three-dimensional space to obtain the second label image corresponding to the first label image.
[0166] S7. Input the text feature set and the second feature map into the point cloud label generation network to obtain the third label image corresponding to the sample point cloud data.
[0167] S8. Obtain the two-dimensional mask image corresponding to the sample image data according to the first label image and the text feature set.
[0168] S9. Obtain the three-dimensional mask map corresponding to the sample point cloud data according to the third label image and the text feature set.
[0169] S10. Determine the first embedding values of various objects in the object category dataset in the first feature map according to the first feature map and the two-dimensional mask map.
[0170] S11. Determine the second embedding values of various objects in the object category dataset in the second feature map according to the second feature map and the three-dimensional mask map.
[0171] S12. Obtain the first similarity between the first embedding value of each object category in the first feature map and the second embedding value of each object category in the second feature map.
[0172] S13. Obtain the value of the first loss function according to the sum of the first similarities and the preset contrast loss function.
[0173] S14. Map the second feature map to a two-dimensional space to obtain the third feature map corresponding to the second feature map.
[0174] S15. Perform grid division on the first feature map and the third feature map respectively to obtain the first grid feature map corresponding to the first feature map and the second grid feature map corresponding to the second feature map.
[0175] S16. Obtain the third similarity between each grid feature in the first grid feature map and each grid feature in the second grid feature map.
[0176] S17. Determine the sum of the third similarities as the second similarity, and obtain the value of the second loss function according to the second similarity and the preset contrast loss function.
[0177] S18. Obtain the value of the third loss function according to the second label image and the third label image.
[0178] S19. Train the initial point cloud encoder according to the values of the first loss function, the second loss function, and the third loss function to obtain an object recognition model.
[0179] S20. Obtain the point cloud data in the driving environment of the vehicle.
[0180] S21. Input the object category dataset into the text encoder to obtain the text feature set corresponding to the object category dataset.
[0181] S22. Input the point cloud data into the point cloud encoder to obtain the point cloud feature map corresponding to the point cloud data.
[0182] S23. Input the point cloud feature map and the text feature set into the point cloud label generation network to obtain the recognition result of the objects in the point cloud data.
[0183] It should be noted that for the descriptions in S1 - S23 above, reference can be made to the relevant descriptions in the above embodiments, and the effects are similar. Therefore, this embodiment will not be elaborated here.
[0184] It should be understood that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above figures may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0185] In one embodiment, as Figure 13 shown, an object recognition device is provided, including: a first acquisition module 10 and a second acquisition module 11, where:
[0186] The first acquisition module 10 is configured to acquire point cloud data and an object category data set in the driving environment of a vehicle; the object category data set includes text data of multiple different categories of objects;
[0187] The second acquisition module 11 is configured to acquire the recognition result of an object in the point cloud data according to the point cloud data, the object category data set, and a preset object recognition model; the object recognition model is obtained by training an initial object recognition model with a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and an object category data set in the driving environment of a vehicle.
[0188] The object recognition device provided in this embodiment can execute the above - mentioned method embodiment, and its implementation principle and technical effects are similar, so they will not be elaborated here.
[0189] In one embodiment, as Figure 14 shown, the above - mentioned second acquisition module 11 includes: a first acquisition unit 111, where:
[0190] The first acquisition unit 111 is configured to input the point cloud data and the object category data set into the object recognition model to perform point cloud feature recognition and text feature recognition, so as to obtain the recognition result of an object in the point cloud data.
[0191] The object recognition device provided in this embodiment can execute the above - mentioned method embodiment, and its implementation principle and technical effects are similar, so they will not be elaborated here.
[0192] In one embodiment, the object recognition model includes a point cloud encoder, a text encoder, and a point cloud label generation network. Please continue to refer to Figure 14 The above-mentioned first acquisition unit 111 is specifically configured to: input the object category data set into the text encoder to obtain a text feature set corresponding to the object category data set; input the point cloud data into the point cloud encoder to obtain a point cloud feature map corresponding to the point cloud data; input the point cloud feature map and the text feature set into the point cloud label generation network to obtain the recognition result of the object in the point cloud data.
[0193] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0194] In one embodiment, please continue to refer to Figure 14 The above-mentioned device further includes: a fifth acquisition module 12, a sixth acquisition module 13, a seventh acquisition module 14, and an eighth acquisition module 15, where:
[0195] The fifth acquisition module 12 is configured to input the object category data set into the text encoder of the initial object recognition model to obtain a text feature set corresponding to the object category data set;
[0196] The sixth acquisition module 13 is configured to input the sample image data into the image encoder of the initial object recognition model to obtain a first feature map corresponding to the sample image data;
[0197] The seventh acquisition module 14 is configured to input the sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain a second feature map corresponding to the sample point cloud data;
[0198] The eighth acquisition module 15 is configured to train the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain an object recognition model.
[0199] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0200] In one embodiment, the initial object recognition model further includes an image label generation network and a point cloud label generation network. Please continue to refer to Figure 14 The above-mentioned eighth acquisition module 15 includes: a second acquisition unit 151, a third acquisition unit 152, a fourth acquisition unit 153, and a fifth acquisition unit 154, where:
[0201] The second acquisition unit 151 is configured to input the text feature set and the first feature map into the image label generation network to obtain a first label image corresponding to the sample image data; the first label image includes the labels of each object in the sample image data;
[0202] A third acquisition unit 152, configured to map the first label image into a three-dimensional space to obtain a second label image corresponding to the first label image;
[0203] A fourth acquisition unit 153, configured to input the text feature set and the second feature map into a point cloud label generation network to obtain a third label image corresponding to the sample point cloud data; the third label image includes labels of each object in the sample point cloud data;
[0204] A fifth acquisition unit 154, configured to train an initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image to obtain an object recognition model.
[0205] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0206] In one embodiment, please continue to refer to Figure 14 , specifically, the above-mentioned fifth acquisition unit 154 is configured to: obtain the value of a first loss function according to the first feature map, the second feature map, the first label image, and the third label image; obtain the value of a second loss function according to the first feature map and the second feature map; obtain the value of a third loss function according to the second label image and the third label image; train the initial point cloud encoder according to the values of the first loss function, the second loss function, and the third loss function to obtain an object recognition model.
[0207] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0208] In one embodiment, please continue to refer to Figure 14 , specifically, the above-mentioned fifth acquisition unit 154 is configured to: obtain a two-dimensional mask map corresponding to the sample image data according to the first label image and the text feature set; obtain a three-dimensional mask map corresponding to the sample point cloud data according to the third label image and the text feature set; determine a first embedding value of each object in the object category dataset in the first feature map according to the first feature map and the two-dimensional mask map; determine a second embedding value of each object in the object category dataset in the second feature map according to the second feature map and the three-dimensional mask map; obtain the value of the first loss function according to the first embedding value of each object in the first feature map and the second embedding value of each object in the second feature map.
[0209] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0210] In one embodiment, please continue to refer toFigure 14 Specifically, the fifth obtaining unit 154 is configured to: obtain a first similarity between the first embedding values of objects of various categories in the first feature map and the second embedding values of objects of various categories in the second feature map; and obtain a value of a first loss function according to the sum of the first similarities and a preset contrast loss function.
[0211] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0212] In one embodiment, please continue to refer to Figure 14 Specifically, the fifth obtaining unit 154 is configured to: map the second feature map to a two-dimensional space to obtain a third feature map corresponding to the second feature map; perform grid division on the first feature map and the third feature map respectively to obtain a first grid feature map corresponding to the first feature map and a second grid feature map corresponding to the second feature map; obtain a second similarity between the first grid feature map and the second grid feature map; and obtain a value of a second loss function according to the second similarity and a preset contrast loss function.
[0213] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0214] In one embodiment, please continue to refer to Figure 14 Specifically, the fifth obtaining unit 154 is configured to: obtain a third similarity between each grid feature in the first grid feature map and each grid feature in the second grid feature map; and determine the sum of the third similarities as the second similarity.
[0215] The object recognition device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0216] In one embodiment, as Figure 15 shown, a training device for an object recognition model is provided. The above device includes: a third obtaining module 16 and a fourth obtaining module 17, where:
[0217] The third obtaining module 16 is configured to obtain a training sample set; the training sample set includes sample image data, sample point cloud data, and an object category data set in a vehicle driving environment; the object category data set includes text data of multiple different categories of objects;
[0218] The fourth obtaining module 17 is configured to train an initial object recognition model by using the training sample set to obtain an object recognition model.
[0219] The training device for the object recognition model provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0220] In one embodiment, as Figure 16 shown, the above fourth acquisition module 17 includes:
[0221] A sixth acquisition unit 171, configured to input an object category data set into the text encoder of the initial object recognition model to obtain a text feature set corresponding to the object category data set;
[0222] A seventh acquisition unit 172, which inputs sample image data into the image encoder of the initial object recognition model to obtain a first feature map corresponding to the sample image data;
[0223] An eighth acquisition unit 173, which inputs sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain a second feature map corresponding to the sample point cloud data;
[0224] A ninth acquisition unit 174, which trains the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain an object recognition model.
[0225] The training device for the object recognition model provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0226] For the specific limitations on the object recognition device and the training device for the object recognition model, reference can be made to the limitations on the object recognition method and the training method for the object recognition model in the above text, which will not be elaborated here. Each module in the above object recognition device and the training device for the object recognition model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in the form of hardware or be independent of it, or be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.
[0227] Figure 17 FIG. is a block diagram of a computer device 1300 shown according to an exemplary embodiment. For example, the computer device 1300 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0228] Refer to Figure 17, the computer device 1300 may include one or more of the following components: a processing component 1302, a memory 1304, a power supply component 1306, a multimedia component 1308, an audio component 1310, an input / output (I / O) interface 1312, a sensor component 1314, and a communication component 1316. Among them, computer programs or instructions that run on the processor are stored on the memory.
[0229] The processing component 1302 generally controls the overall operation of the computer device 1300, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1302 may include one or more processors 1320 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 1302 may include one or more modules to facilitate the interaction between the processing component 1302 and other components. For example, the processing component 1302 may include a multimedia module to facilitate the interaction between the multimedia component 1308 and the processing component 1302.
[0230] The memory 1304 is configured to store various types of data to support the operation of the computer device 1300. Examples of such data include instructions for any application or method operating on the computer device 1300, contact data, phone book data, messages, pictures, videos, etc. The memory 1304 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0231] The power supply component 1306 provides power to various components of the computer device 1300. The power supply component 1306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the computer device 1300.
[0232] The multimedia component 1308 includes a touch display screen that provides an output interface between the computer device 1300 and the user. In some embodiments, the touch display screen may include a liquid crystal display (LCD) and a touch panel (TP). The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1308 includes a front camera and / or a rear camera. When the computer device 1300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0233] The audio component 1310 is configured to output and / or input audio signals. For example, the audio component 1310 includes a microphone (MIC) that is configured to receive external audio signals when the computer device 1300 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1304 or transmitted via the communication component 1316. In some embodiments, the audio component 1310 further includes a speaker for outputting audio signals.
[0234] The I / O interface 1312 provides an interface between the processing component 1302 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0235] The sensor component 1314 includes one or more sensors for providing a status assessment of various aspects of the computer device 1300. For example, the sensor component 1314 can detect the on / off state of the computer device 1300, the relative positioning of components, such as the display and the keypad of the computer device 1300. The sensor component 1314 can also detect a change in the position of the computer device 1300 or a component of the computer device 1300, the presence or absence of user contact with the computer device 1300, the orientation or acceleration / deceleration of the computer device 1300, and the temperature change of the computer device 1300. The sensor component 1314 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1314 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1314 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0236] The communication component 1316 is configured to facilitate communication between the computer device 1300 and other devices in a wired or wireless manner. The computer device 1300 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1316 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0237] In an exemplary embodiment, the computer device 1300 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above object recognition method.
[0238] In an exemplary embodiment, a computer program product is further provided. When the computer program is executed by a processor, the above method can be implemented. The computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, part or all of the above method can be implemented in accordance with the process or function described in the embodiments of the present disclosure.
[0239] It should be noted that for the solutions described in this specification and embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. When a user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
[0240] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0241] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0242] The above-described embodiments merely represent several implementation manners of the embodiments of the present disclosure. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present disclosure, several modifications and improvements can be made, and these all belong to the protection scope of the embodiments of the present disclosure. Therefore, the protection scope of the patent of the embodiments of the present disclosure should be subject to the appended claims.
Claims
1. An object recognition method, characterized in that, The method includes: Obtaining point cloud data and an object category dataset in the driving environment of a vehicle; the object category dataset includes text data of multiple different categories of objects; Obtaining an identification result of an object in the point cloud data according to the point cloud data, the object category dataset, and a preset object recognition model; the object recognition model is obtained by training an initial object recognition model with a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and the object category dataset in the driving environment of the vehicle.
2. The method according to claim 1, characterized in that The obtaining an identification result of an object in the point cloud data according to the point cloud data, the object category dataset, and a preset object recognition model includes: Inputting the point cloud data and the object category dataset into the object recognition model to perform text point cloud feature recognition and text feature recognition, so as to obtain an identification result of an object in the point cloud data.
3. The method according to claim 2, wherein The object recognition model includes a point cloud encoder, a text encoder, and a point cloud label generation network; the inputting the point cloud data and the object category dataset into the object recognition model to perform point cloud feature recognition and text feature recognition, so as to obtain an identification result of an object in the point cloud data includes: Inputting the object category dataset into the text encoder to obtain a text feature set corresponding to the object category dataset; Inputting the point cloud data into the point cloud encoder to obtain a point cloud feature map corresponding to the point cloud data; Inputting the point cloud feature map and the text feature set into the point cloud label generation network to obtain an identification result of an object in the point cloud data.
4. The method according to any one of claims 1 to 3, characterized in that The training process of the object recognition model includes: Inputting the object category dataset into the text encoder of the initial object recognition model to obtain a text feature set corresponding to the object category dataset; Inputting the sample image data into the image encoder of the initial object recognition model to obtain a first feature map corresponding to the sample image data; Inputting the sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain a second feature map corresponding to the sample point cloud data; Training the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain the object recognition model.
5. The method according to claim 4, characterized in that, The initial object recognition model further includes an image label generation network and a point cloud label generation network; the training the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain the object recognition model includes: Inputting the text feature set and the first feature map into the image label generation network to obtain a first label image corresponding to the sample image data; the first label image includes labels of each object in the sample image data; Mapping the first label image into a three-dimensional space to obtain a second label image corresponding to the first label image; Input the text feature set and the second feature map into the point cloud label generation network to obtain a third label image corresponding to the sample point cloud data; the third label image includes labels of each object in the sample point cloud data. Train the initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image to obtain the object recognition model.
6. The method according to claim 5, characterized in that, The step of training the initial point cloud encoder according to the first feature map, the second feature map, the first label image, the second label image, and the third label image to obtain the object recognition model includes: Obtain the value of the first loss function according to the first feature map, the second feature map, the first label image, and the third label image. Obtain the value of the second loss function according to the first feature map and the second feature map. Obtain the value of the third loss function according to the second label image and the third label image. Train the initial point cloud encoder according to the value of the first loss function, the value of the second loss function, and the value of the third loss function to obtain the object recognition model.
7. The method according to claim 6, wherein The step of obtaining the value of the first loss function according to the first feature map, the second feature map, the first label image, and the third label image includes: Obtain a two-dimensional mask map corresponding to the sample image data according to the first label image and the text feature set. Obtain a three-dimensional mask map corresponding to the sample point cloud data according to the third label image and the text feature set. Determine the first embedding values of objects in each category in the object category dataset in the first feature map according to the first feature map and the two-dimensional mask map. Determine the second embedding values of objects in each category in the object category dataset in the second feature map according to the second feature map and the three-dimensional mask map. Obtain the value of the first loss function according to the first embedding values of objects in each category in the first feature map and the second embedding values of objects in each category in the second feature map.
8. The method according to claim 7, wherein The step of obtaining the value of the first loss function according to the first embedding values of objects in each category in the first feature map and the second embedding values of objects in each category in the second feature map includes: Obtain the first similarities between the first embedding values of objects in each category in the first feature map and the second embedding values of objects in each category in the second feature map. Obtain the value of the first loss function according to the sum of each of the first similarities and a preset contrast loss function.
9. The method according to claim 6, wherein The step of obtaining the value of the second loss function according to the first feature map and the second feature map includes: Map the second feature map to a two-dimensional space to obtain a third feature map corresponding to the second feature map. Perform grid division on the first feature map and the third feature map respectively to obtain a first grid feature map corresponding to the first feature map and a second grid feature map corresponding to the second feature map. Obtain the second similarity between the first grid feature map and the second grid feature map. Obtain the value of the second loss function according to the second similarity and a preset comparison loss function.
10. The method according to claim 9, characterized in that, Obtaining the second similarity between the first grid feature map and the second grid feature map includes: Obtaining a third similarity between each grid feature in the first grid feature map and each grid feature in the second grid feature map; Determine the sum of the third similarities as the second similarity.
11. A training method for an object recognition model, characterized in that, The method includes: Obtain a training sample set; the sample training set includes sample image data, sample point cloud data, and an object category data set in the vehicle driving environment; the object category data set includes text data of multiple different categories of objects; Train an initial object recognition model using the sample training set to obtain an object recognition model.
12. The method according to claim 11, wherein The training the initial object recognition model using the sample training set to obtain an object recognition model includes: Input the object category data set into the text encoder of the initial object recognition model to obtain a text feature set corresponding to the object category data set; Input the sample image data into the image encoder of the initial object recognition model to obtain a first feature map corresponding to the sample image data; Input the sample point cloud data into the initial point cloud encoder of the initial object recognition model to obtain a second feature map corresponding to the sample point cloud data; Train the initial point cloud encoder according to the text feature set, the first feature map, and the second feature map to obtain the object recognition model.
13. An object recognition device, characterized in that, The device includes: A first acquisition module, configured to acquire point cloud data and an object category data set in the vehicle driving environment; the object category data set includes text data of multiple different categories of objects; A second acquisition module, configured to obtain an identification result of an object in the point cloud data according to the point cloud data, the object category data set, and a preset object recognition model; the object recognition model is obtained by training an initial object recognition model using a preset sample training set, and the sample training set includes sample image data, sample point cloud data, and the object category data set in the vehicle driving environment.
14. A training device for an object recognition model, characterized in that The device includes: A third acquisition module, configured to acquire a training sample set; the sample training set includes sample image data, sample point cloud data, and an object category data set in the vehicle driving environment; the object category data set includes text data of multiple different categories of objects; A fourth acquisition module, configured to train an initial object recognition model using the sample training set to obtain an object recognition model.
15. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.