Model training method and device based on laser radar, equipment and medium
Through three-dimensional and two-dimensional convolutional neural networks, the feature extraction and fusion of point cloud data of different lidar sensors is solved, and the problem of model training of lidar sensors spanning different manufacturers and models is achieved, achieving higher generalization and accuracy.
Patent Information
- Application Number
- CN202510338973.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing lidar-based model training methods cannot span different manufacturers and models of lidar sensors, resulting in insufficient generalization and accuracy of the model.
A three-dimensional convolutional neural network is used to extract and splice the point cloud data collected by different lidar sensors, and density and feature fusion are performed through a two-dimensional convolutional neural network, and the network model is reversely updated with the loss function to obtain the target model.
The generalization and accuracy of the model are improved, and the point cloud data of different lidar sensors can be integrated across domains to achieve better environmental perception.
Smart Images

Figure CN120259841A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a model training method, device, equipment and medium based on lidar. Background Art
[0002] Lidar sensors are playing an increasingly significant role in three-dimensional environmental perception, and there are more and more requirements for different landing applications. Due to advantages such as high ranging accuracy, insensitivity to light, and strong laser beam penetration, lidar is widely used in the fields of drones, robots, and autonomous driving.
[0003] Further, lidar can be roughly divided into mechanical rotating lidar and solid-state lidar according to the working scanning method; among them, the mechanical rotating lidar usually includes a rotating component, which is used to scan the surrounding environment, and it consists of one or more laser emitters and receivers, and realizes 360-degree environmental perception through rotation, thereby forming dense 3D point cloud data. The solid-state lidar has no moving parts, and it mainly relies on electronic components (such as optical phased arrays, photonic integrated circuits, etc.) to control the laser emission angle, and obtains dense 3D point cloud data in a fixed orientation after multiple non-repetitive scans.
[0004] Most of the existing lidar-based model training methods are trained based on supervised ground truth labels. The training data is the point cloud data collected by a specific lidar sensor, and the training ground truth is the manually labeled target 3D box or point instance label. By designing a specific perception algorithm on the training data and ground truth, the perception function of a specific lidar sensor is realized.
[0005] However, due to different lidar sensor models and scanning methods released by each manufacturer, the number of point clouds scanned for the same object is different, and the point cloud shapes are different, resulting in the inability to use the models trained by the above methods across different scenarios. Summary of the Invention
[0006] This application provides a model training method, device, equipment and medium based on lidar to improve the generalization and accuracy of model use.
[0007] In a first aspect, a model training method based on lidar is provided, including:
[0008] Feature extraction is respectively performed on N point cloud data by using a three-dimensional convolutional neural network, and the N 3D point cloud features extracted are spliced to obtain M three-dimensional neighborhood features; wherein, the N point cloud data are collected by different lidar sensors, and both N and M are integers greater than 1;
[0009] Perform densification processing on the M three-dimensional neighborhood features respectively to obtain M two-dimensional dense features, and use a two-dimensional convolutional neural network to extract features from the M two-dimensional dense features respectively, and fuse the M two-dimensional features extracted to obtain S two-dimensional fused features; where S is an integer greater than 1;
[0010] Perform prediction on the S two-dimensional fused features to obtain a prediction result; where the prediction result at least includes their respective occupancy categories, point cloud coordinates, and point cloud quantities;
[0011] Calculate a loss function according to their respective occupancy categories, point cloud coordinates, and point cloud quantities, and update the three-dimensional convolutional neural network and the two-dimensional convolutional neural network in reverse according to the loss function to obtain a target model; where the loss function is used to evaluate the prediction result accuracy of the target model.
[0012] In the embodiments of the present application, first, use a three-dimensional convolutional neural network to extract features from N point cloud data respectively, and splice the N 3D point cloud features extracted to obtain M three-dimensional neighborhood features. The N point cloud data are collected by different lidar sensors. Perform densification processing on the M three-dimensional neighborhood features respectively to obtain M two-dimensional dense features, and use a two-dimensional convolutional neural network to extract features from the M two-dimensional dense features respectively, and fuse the M two-dimensional features extracted to obtain S two-dimensional fused features, which can extract the features of the point cloud data of multiple lidars in the three-dimensional space and the two-dimensional space, providing a basis for achieving the cross-domain fusion perception effect in the follow-up; then, perform prediction on the S two-dimensional fused features to obtain a prediction result, calculate a loss function according to their respective occupancy categories, point cloud coordinates, and point cloud quantities in the prediction result, and update the three-dimensional convolutional neural network and the two-dimensional convolutional neural network in reverse according to the loss function to obtain a target model, which can predict the results of the point cloud data from multiple dimensions, making the model have better generalization and accuracy.
[0013] In some embodiments, before using the three-dimensional convolutional neural network to extract features from N point cloud data respectively, it further includes:
[0014] After adjusting the three-dimensional coordinate systems of the respective lidar sensors to a three-dimensional coordinate system in the same direction, collect the respective point cloud data through the respective lidar sensors;
[0015] Based on the respective point cloud data, form N voxels, and perform masking processing on the N voxels to obtain N masked voxels;
[0016] The using a three-dimensional convolutional neural network to extract features from N point cloud data respectively, and splicing the N 3D point cloud features extracted to obtain M three-dimensional neighborhood features includes:
[0017] Use the three-dimensional convolutional neural network to extract features from the N masked voxels respectively, and splice the N 3D point cloud features extracted to obtain M three-dimensional neighborhood features.
[0018] In some embodiments, the three-dimensional convolutional neural network includes a plurality of three-dimensional convolutional layers with shared parameters and a dense layer connected in sequence;
[0019] The plurality of three-dimensional convolutional layers with shared parameters are used to extract features from the N masked voxels respectively, and splice the N 3D point cloud features extracted to obtain M three-dimensional neighborhood features;
[0020] The dense layer is used to perform two-dimensional coordinate mapping on the M three-dimensional neighborhood features to obtain M two-dimensional dense features.
[0021] In some embodiments, the method of fusing the M two-dimensional features obtained to obtain S two-dimensional fusion features includes:
[0022] According to any target two-dimensional feature among the M two-dimensional features, determine the weights between the target two-dimensional feature and other two-dimensional features respectively;
[0023] Fuse the target two-dimensional feature and the other two-dimensional features according to the weights between the target two-dimensional feature and the other two-dimensional features to obtain the S two-dimensional fusion features.
[0024] In some embodiments, the method further includes:
[0025] According to the target point cloud data to be tested, perform a fine-tuning operation on the target model until the initial learning rate of the target model is less than the initial learning rate in the model training stage.
[0026] In a second aspect, a model training device based on lidar is provided, including:
[0027] A three-dimensional feature extraction module, configured to use a three-dimensional convolutional neural network to extract features from N point cloud data respectively, and splice the N 3D point cloud features extracted to obtain M three-dimensional neighborhood features; wherein, the N point cloud data are collected by different lidar sensors, and both N and M are integers greater than 1;
[0028] A two-dimensional feature extraction module, configured to perform densification processing on the M three-dimensional neighborhood features respectively to obtain M two-dimensional dense features, and use a two-dimensional convolutional neural network to extract features from the M two-dimensional dense features respectively, and fuse the M two-dimensional features obtained to obtain S two-dimensional fusion features; wherein, S is an integer greater than 1;
[0029] A prediction module for predicting the S two-dimensional fusion features to obtain a prediction result, where the prediction result at least includes their respective occupancy categories, point cloud coordinates, and point cloud quantities;
[0030] A training module for calculating a loss function based on their respective occupancy categories, point cloud coordinates, and point cloud quantities, and reversely updating the three-dimensional convolutional neural network and the two-dimensional convolutional neural network according to the loss function to obtain a target model, where the loss function is used to evaluate the accuracy of the prediction result of the target model.
[0031] In some embodiments, the device further includes a preprocessing module;
[0032] The preprocessing module is configured to adjust the three-dimensional coordinate systems of the respective lidar sensors to a three-dimensional coordinate system in a unified direction, and then collect the respective point cloud data through the respective lidar sensors; based on the respective point cloud data, N voxels are formed, and the N voxels are subjected to masking processing to obtain N masked voxels.
[0033] In some embodiments, the two-dimensional feature extraction module is specifically configured to:
[0034] According to any target two-dimensional feature among the M two-dimensional features, respectively determine the weights between the target two-dimensional feature and other two-dimensional features;
[0035] Fuse the target two-dimensional feature and the other two-dimensional features according to the weights between the target two-dimensional feature and the other two-dimensional features to obtain the S two-dimensional fusion features.
[0036] In a third aspect, an electronic device is provided, including:
[0037] A memory for storing a computer program; a processor for implementing the method steps described in any one of the first aspect when executing the computer program stored on the memory.
[0038] In a fourth aspect, a computer-readable storage medium is provided, where a computer program is stored in the computer-readable storage medium, and the computer program implements the method steps described in any one of the first aspect when executed by a processor.
[0039] For the various aspects and the possible technical effects that can be achieved in the second to fourth aspects above, please refer to the description of the possible technical effects that can be achieved in the first aspect or various possible solutions in the first aspect above, and details will not be repeated here. Description of the Drawings
[0040] Figure 1 It is a flowchart of a lidar-based model training method provided by an embodiment of the present application;
[0041] Figure 2 A logical schematic diagram of model training based on lidar provided by an embodiment of the present application;
[0042] Figure 3 A structural schematic diagram of a model training device based on lidar provided by an embodiment of the present application;
[0043] Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The specific operation methods in the method embodiments can also be applied to the device embodiments or system embodiments. It should be noted that in the description of the present application, "a plurality of" is understood as "at least two". "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The connection between A and B may represent: the direct connection between A and B and the connection between A and B through C. In addition, in the description of the present application, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0045] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application will be described first below.
[0046] (1) The 3D point cloud data is specifically represented by multiple attributes in the lidar coordinate system, such as the x-axis coordinate, y-axis coordinate, z-axis coordinate, reflectivity, timestamp, scan beam, etc.
[0047] (2) Voxelization is to convert the geometric form representation of an object into a voxel representation form that is closest to the object, generating a voxel data set. A voxel is short for volume pixel and is a basic unit in three-dimensional space, similar to a pixel in a two-dimensional image but extended to three-dimensional space. Voxelization not only includes the surface information of the model but also can describe the internal attributes of the model.
[0048] (3) A mask is a concept in computer vision and image processing. It is usually a matrix or image with the same size as the original image, used to indicate or identify pixels or regions in specific areas of the original image. A mask can be binary, that is, it only contains numerical values of 0 or 1. Among them, a pixel value of 1 indicates that the corresponding pixel of the original image belongs to a specific area or target, while a pixel value of 0 indicates that it does not belong to that area or target. A mask can also be multi-channel, and each channel may represent different types of information or targets.
[0049] To further illustrate the technical solutions provided by the embodiments of the present application, the following will be described in detail with reference to the accompanying drawings and specific implementation manners. Although the embodiments of the present application provide method operation steps as shown in the following embodiments or drawings, based on routine or non-creative labor, more or fewer operation steps may be included in the method. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. When the method is actually processed or executed by the device, it can be executed in the order shown in the embodiments or drawings or executed concurrently.
[0050] Figure 1 It is a flowchart of a lidar-based model training method provided by an embodiment of the present application. This process can be executed by a lidar-based model training device to improve the generalization and accuracy of model usage. As Figure 1 shown, this process includes the following steps:
[0051] 101: Use a three-dimensional convolutional neural network to extract features from N point cloud data respectively, and splice the N 3D point cloud features extracted to obtain M three-dimensional neighborhood features.
[0052] In this step, the N point cloud data are collected by different lidar sensors (specifically, lidar sensors of different models and different manufacturers). Both N and M are integers greater than 1.
[0053] In some embodiments, before using a three-dimensional convolutional neural network to extract features from N point cloud data respectively, the three-dimensional coordinate systems of each lidar sensor can be adjusted to a three-dimensional coordinate system in the same direction (for example, the positive direction of the x-axis is forward, the positive direction of the y-axis is to the left, and the positive direction of the z-axis is upward), and then each lidar sensor collects its own point cloud data; based on each point cloud data, N voxels are formed, and mask processing is performed on the N voxels to obtain N masked voxels, so as to ensure that the features extracted subsequently are more referenceable and effective.
[0054] In some embodiments, the three-dimensional convolutional neural network may include a plurality of three-dimensional convolutional layers and dense layers with shared parameters connected in sequence.
[0055] Further, the 3D convolutional layer with multiple shared parameters is used to extract features from N masked voxels respectively, so as to extract the point cloud features of each domain as much as possible; and the N 3D point cloud features extracted are concatenated to obtain M 3D neighborhood features, so as to merge the features of multiple domains for normalization processing, and effectively extract the features that are adjacent in space or have a certain spatial relationship.
[0056] It should be noted that the number of 3D convolutional layers can be flexibly set according to actual needs, and the normalization processing method can be min-max normalization, Z-score normalization, etc., which are not limited in the embodiments of the present application.
[0057] Since the 3D convolutional layer (e.g., spconv or other sparse convolutional layers) is used to perform convolutional operations on the input N masked voxels, the features extracted are usually sparse and only exist in the non-empty positions of the point cloud; therefore, in order to convert the sparse N 3D point cloud features into dense M 2D dense features, further, a dense layer (e.g., Dense layer) is used to perform 2D coordinate mapping on the M 3D neighborhood features to obtain M 2D dense features, so as to facilitate the subsequent extraction of features in the 2D space and provide a basis for the subsequent model to achieve the cross-domain fusion perception effect.
[0058] 102: Densify the M 3D neighborhood features respectively to obtain M 2D dense features, and use a 2D convolutional neural network to extract features from the M 2D dense features respectively, and fuse the M 2D features extracted to obtain S 2D fusion features.
[0059] In some embodiments, fusing the M 2D features extracted to obtain S 2D fusion features may be: determining the weights between the target 2D feature and other 2D features respectively according to any one of the M 2D features, namely the target 2D feature;
[0060] Fuse the target 2D feature and other 2D features according to the weights between the target 2D feature and other 2D features to obtain S 2D fusion features.
[0061] Further, the two-dimensional convolutional neural network may include multiple two-dimensional convolutional blocks and transformer blocks; further, a two-dimensional convolutional block includes a convolutional layer, a batch normalization layer, an activation function layer, etc., and adopts the method of sharing weight parameters, focusing on the extraction of local features to obtain M two-dimensional features; the transformer block may include a self-attention layer and a feed-forward neural network layer, focusing on the feature interaction between the M two-dimensional features, and each two-dimensional feature can be used as a query to query the interaction attention weights with other two-dimensional features, and then fuse them using the calculation method of encoder-decoder attention (cross-attention) to obtain S two-dimensional fusion features.
[0062] 103: Predict the above S two-dimensional fusion features to obtain a prediction result, and the prediction result at least includes their respective occupancy categories, point cloud coordinates, and point cloud quantities.
[0063] In some embodiments, a pre-trained model may be used to predict the S two-dimensional fusion features to obtain a prediction result, and the pre-trained model may include three parts: semantic prediction (Class head), point cloud coordinate prediction (Points head), and point cloud density prediction (Density head). The semantic prediction can be used to predict the occupancy category of the two-dimensional fusion features, the point cloud coordinate prediction can be used to predict the point cloud coordinates of the two-dimensional fusion features, and the point cloud density prediction can be used to predict the point cloud quantity of the two-dimensional fusion features, so as to infer its density distribution.
[0064] 104: Calculate a loss function according to their respective occupancy categories, point cloud coordinates, and point cloud quantities, and update the three-dimensional convolutional neural network and the two-dimensional convolutional neural network backward according to the loss function to obtain a target model.
[0065] In some embodiments, the loss function is used to evaluate the accuracy of the prediction result of the target model. The cross-entropy classification loss function and the absolute error loss function can be calculated respectively according to their respective occupancy categories, point cloud coordinates, and point cloud quantities, and then the three-dimensional convolutional neural network and the two-dimensional convolutional neural network are updated backward according to the loss function values of the cross-entropy classification loss function and the absolute error loss function respectively until a target model that meets the loss function threshold is obtained.
[0066] In some other embodiments, according to the target point cloud data to be tested, fine-tuning operations can also be performed on the target model until the initial learning rate of the target model is less than the initial learning rate in the model training stage, and then the perception performance of the target model can be further improved by fine-tuning with a small amount of data. Further, a three-dimensional object detection (PV-RCNN) algorithm or a three-dimensional perception algorithm can be used to perform fine-tuning operations on the target model (parameters in a three-dimensional convolutional neural network, a two-dimensional convolutional neural network, a Classhead, a Points head, and a Density head can be fine-tuned). Among them, the PV-RCNN algorithm can efficiently combine the advantages of voxel and point cloud feature learning, thus significantly improving the performance of three-dimensional object detection.
[0067] In the embodiments of the present application, first, a three-dimensional convolutional neural network is used to extract features from N point cloud data respectively, and the N extracted 3D point cloud features are concatenated to obtain M three-dimensional neighborhood features. The N point cloud data are collected by different lidar sensors. Densification processing is performed on the M three-dimensional neighborhood features respectively to obtain M two-dimensional dense features, and a two-dimensional convolutional neural network is used to extract features from the M two-dimensional dense features respectively. The M extracted two-dimensional features are fused to obtain S two-dimensional fusion features, which can extract the features of point cloud data of multiple lidars in three-dimensional space and two-dimensional space, providing a basis for achieving a cross-domain fusion perception effect in the follow-up; then, the S two-dimensional fusion features are predicted to obtain a prediction result, and a loss function is calculated according to the occupancy category, point cloud coordinates, and point cloud quantity in the prediction result, and the three-dimensional convolutional neural network and the two-dimensional convolutional neural network are updated backward according to the loss function to obtain the target model, which can predict the result of point cloud data from multiple dimensions, making the model have better generalization and accuracy.
[0068] Based on the above Figure 1 shown method, Figure 2 is a logical schematic diagram of a lidar-based model training method provided by the embodiments of the present application. In Figure 2Among them, first, the point cloud data collected by N different types of lidars are preprocessed to obtain N masked voxels; then, the N masked voxels are input into a three-dimensional convolutional neural network for feature extraction and densification processing to obtain M two-dimensional dense features; next, the M two-dimensional dense features are input into a two-dimensional convolutional neural network for feature extraction and cross-domain feature fusion to obtain S two-dimensional fusion features; secondly, the S two-dimensional fusion features are respectively input into a Class head, a Points head, and a Density head for prediction to obtain their respective occupancy categories, point cloud coordinates, and point cloud quantities; finally, a loss function is calculated according to their respective occupancy categories, point cloud coordinates, and point cloud quantities, and the three-dimensional convolutional neural network and the two-dimensional convolutional neural network are updated backward according to the loss function, so as to obtain the target model.
[0069] Based on the same technical concept, an embodiment of the present application also provides a lidar-based model training device, which can implement the above-mentioned lidar-based model training method flow in the embodiment of the present application.
[0070] Figure 3 It is a schematic structural diagram of a lidar-based model training device provided by an embodiment of the present application. As Figure 3 shown, the device includes: a three-dimensional feature extraction module 301, a two-dimensional feature extraction module 302, a prediction module 303, and a training module 304; further, the device also includes a preprocessing module 305.
[0071] The three-dimensional feature extraction module 301 is configured to respectively perform feature extraction on N point cloud data by using a three-dimensional convolutional neural network, and splice the extracted N 3D point cloud features to obtain M three-dimensional neighborhood features; wherein, the N point cloud data are collected by different lidar sensors, and both N and M are integers greater than 1.
[0072] The two-dimensional feature extraction module 302 is configured to respectively perform densification processing on the M three-dimensional neighborhood features to obtain M two-dimensional dense features, and respectively perform feature extraction on the M two-dimensional dense features by using a two-dimensional convolutional neural network, and fuse the extracted M two-dimensional features to obtain S two-dimensional fusion features; wherein, S is an integer greater than 1.
[0073] The prediction module 303 is configured to predict the S two-dimensional fusion features to obtain a prediction result; wherein, the prediction result at least includes their respective occupancy categories, point cloud coordinates, and point cloud quantities.
[0074] A training module 304 is configured to calculate a loss function based on the respective occupancy categories, point cloud coordinates, and point cloud quantities, and update the 3D convolutional neural network and the 2D convolutional neural network in reverse according to the loss function to obtain a target model; wherein, the loss function is used to evaluate the accuracy of the prediction result of the target model.
[0075] A preprocessing module 305 is configured to adjust the 3D coordinate systems of the respective lidar sensors to a 3D coordinate system in a unified direction, and then collect the respective point cloud data through the respective lidar sensors; based on the respective point cloud data, N voxels are formed, and masking processing is performed on the N voxels to obtain N masked voxels; the 3D feature extraction module 301 is specifically configured to: perform feature extraction on the N masked voxels respectively by using the 3D convolutional neural network, and splice the extracted N 3D point cloud features to obtain M 3D neighborhood features.
[0076] In some embodiments, the 2D feature extraction module 302 is specifically configured to:
[0077] Determine the weights between the target 2D feature and other 2D features respectively according to any one of the M 2D features;
[0078] Fuse the target 2D feature and the other 2D features according to the weights between the target 2D feature and the other 2D features to obtain the S 2D fusion features.
[0079] In some embodiments, the training module 304 is further configured to:
[0080] Perform a fine-tuning operation on the target model according to the target point cloud data to be tested until the initial learning rate of the target model is less than the initial learning rate in the model training stage.
[0081] It should be noted here that the above device provided in the embodiments of the present application can implement all the method steps in the above method embodiments and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.
[0082] Based on the same technical concept, an electronic device is further provided in the embodiments of the present application, and the electronic device can implement the functions of the foregoing model training device based on lidar.
[0083] Figure 4 It is a schematic structural diagram of an electronic device provided in the embodiments of the present application.
[0084] At least one processor 401 and a memory 402 connected to the at least one processor 401. In the embodiments of the present application, the specific connection medium between the processor 401 and the memory 402 is not limited. Figure 4 In this example, the processor 401 and the memory 402 are connected through a bus 400. The bus 400 is Figure 4 represented by a thick line. The connection manners between other components are only for illustrative purposes and are not restrictive. The bus 400 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 4 it is only represented by a thick line, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 401 can also be referred to as a controller, and there is no limitation on the name.
[0085] In the embodiments of the present application, the memory 402 stores instructions executable by the at least one processor 401. By executing the instructions stored in the memory 402, the at least one processor 401 can execute a model training method based on lidar described above. The processor 401 can implement Figure 3 the functions of each module in the device shown.
[0086] Among them, the processor 401 is the control center of the device. It can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 402 and calling the data stored in the memory 402, various functions of the device and process data, so as to monitor the device as a whole.
[0087] In the embodiments of the present application, the processor 401 may include one or more processing units. The processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 401. In some embodiments, the processor 401 and the memory 402 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0088] The processor 401 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of a model training method based on lidar disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0089] The memory 402 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 402 may include at least one type of storage medium. For example, it may include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical discs, and so on. The memory 402 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 402 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0090] By programming the design of the processor 401, the code corresponding to the method for model training based on lidar introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 1 the method for model training based on lidar shown in the embodiments. How to program the design of the processor 401 is a well-known technology to those skilled in the art and will not be elaborated here.
[0091] It should be noted here that the above electronic device provided in the embodiments of the present application can implement all the method steps implemented in the above method embodiments and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically elaborated in this embodiment.
[0092] Based on the same technical concept, the embodiments of the present application provide a computer storage medium, which includes: computer program code. When the computer program code runs on a computer, it causes the computer to execute a method for model training based on lidar as described in any of the foregoing. Since the principle of the above computer storage medium for solving problems is similar to that of a method for model training based on lidar, the implementation of the above computer storage medium can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0093] In a specific implementation process, the computer storage medium may include: various storage media that can store program codes, such as a Universal Serial Bus Flash Drive (USB), a mobile hard disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc.
[0094] Based on the same technical concept, an embodiment of the present application further provides a computer program product, which includes: computer program code. When the computer program code runs on a computer, it causes the computer to execute a method for lidar-based model training as described in any of the foregoing discussions. Since the principle of solving problems by the above computer program product is similar to that of a method for lidar-based model training, the implementation of the above computer program product can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0095] The computer program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0096] The method in the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the present application is executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM, or other programmable devices.
[0097] The computer program or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0098] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0099] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for realizing the functions specified in one block or a plurality of blocks
[0101] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications therein.
Claims
1. A method for model training based on lidar, characterized in that, Including: Using a three-dimensional convolutional neural network to extract features from N point cloud data respectively, and splicing the extracted N 3D point cloud features to obtain M three-dimensional neighborhood features; wherein, the N point cloud data are collected by different lidar sensors, and both N and M are integers greater than 1; Performing densification processing on the M three-dimensional neighborhood features respectively to obtain M two-dimensional dense features, and using a two-dimensional convolutional neural network to extract features from the M two-dimensional dense features respectively, and fusing the extracted M two-dimensional features to obtain S two-dimensional fusion features; wherein, S is an integer greater than 1; Predicting the S two-dimensional fusion features to obtain a prediction result; wherein, the prediction result at least includes their respective occupancy categories, point cloud coordinates, and point cloud quantities; Calculating a loss function according to their respective occupancy categories, point cloud coordinates, and point cloud quantities, and reversely updating the three-dimensional convolutional neural network and the two-dimensional convolutional neural network according to the loss function to obtain a target model; wherein, the loss function is used to evaluate the prediction result accuracy of the target model.
2. The method according to claim 1, wherein Before using the three-dimensional convolutional neural network to extract features from N point cloud data respectively, it further includes: After adjusting the three-dimensional coordinate systems of the respective lidar sensors to a three-dimensional coordinate system in a unified direction, collecting the respective point cloud data through the respective lidar sensors; Based on the respective point cloud data, constructing N voxels and performing masking processing on the N voxels to obtain N masked voxels; The step of using a three-dimensional convolutional neural network to extract features from N point cloud data respectively and splicing the extracted N 3D point cloud features to obtain M three-dimensional neighborhood features includes: Using the three-dimensional convolutional neural network to extract features from the N masked voxels respectively, and splicing the extracted N 3D point cloud features to obtain M three-dimensional neighborhood features.
3. The method according to claim 2, wherein The three-dimensional convolutional neural network includes a plurality of three-dimensional convolutional layers and a dense layer with shared parameters connected in sequence; The plurality of three-dimensional convolutional layers with shared parameters are used to extract features from the N masked voxels respectively, and splicing the extracted N 3D point cloud features to obtain M three-dimensional neighborhood features; The dense layer is used to perform two-dimensional coordinate mapping on the M three-dimensional neighborhood features to obtain M two-dimensional dense features.
4. The method according to claim 1, characterized in that The step of fusing the extracted M two-dimensional features to obtain S two-dimensional fusion features includes: According to any target two-dimensional feature among the M two-dimensional features, respectively determining the weights between the target two-dimensional feature and other two-dimensional features; Fusing the target two-dimensional feature and the other two-dimensional features according to the weights between the target two-dimensional feature and the other two-dimensional features to obtain the S two-dimensional fusion features.
5. The method according to claim 1, characterized in that, The method further includes: According to the target point cloud data to be tested, performing a fine-tuning operation on the target model until the initial learning rate of the target model is less than the initial learning rate in the model training stage.
6. A lidar-based model training device, characterized in that, Including: A three-dimensional feature extraction module, which is used to perform feature extraction on N point cloud data respectively by using a three-dimensional convolutional neural network, and splice the extracted N 3D point cloud features to obtain M three-dimensional neighborhood features; wherein, the N point cloud data are collected by different lidar sensors, and both N and M are integers greater than 1; A two-dimensional feature extraction module, which is used to perform densification processing on the M three-dimensional neighborhood features respectively to obtain M two-dimensional dense features, and perform feature extraction on the M two-dimensional dense features respectively by using a two-dimensional convolutional neural network, and fuse the extracted M two-dimensional features to obtain S two-dimensional fusion features; wherein, S is an integer greater than 1; A prediction module, which is used to predict the S two-dimensional fusion features to obtain a prediction result; wherein, the prediction result at least includes their respective occupancy categories, point cloud coordinates, and point cloud quantities; A training module, which is used to calculate a loss function according to their respective occupancy categories, point cloud coordinates, and point cloud quantities, and update the three-dimensional convolutional neural network and the two-dimensional convolutional neural network in reverse according to the loss function to obtain a target model; wherein, the loss function is used to evaluate the prediction result accuracy of the target model.
7. The device according to claim 6, wherein The device further includes a preprocessing module; The preprocessing module is used to adjust the three-dimensional coordinate systems of the respective lidar sensors to a three-dimensional coordinate system in a unified direction, and then collect their respective point cloud data through the respective lidar sensors; Based on each point cloud data, N voxels are formed, and the N voxels are subjected to masking processing to obtain N masked voxels.
8. The device according to claim 6, characterized in that The two-dimensional feature extraction module is specifically used for: According to any target two-dimensional feature among the M two-dimensional features, respectively determine the weights between the target two-dimensional feature and other two-dimensional features; Fuse the target two-dimensional feature and the other two-dimensional features according to the weights between the target two-dimensional feature and the other two-dimensional features to obtain the S two-dimensional fusion features.
9. An electronic device, characterized in that, Comprising: A memory for storing a computer program; A processor, when executing the computer program stored on the memory, implements the method according to any one of claims 1-5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-5.
Citation Information
Cited By
Radar data storage method and device
CN120872253A
A radar data storage method and apparatus
CN120872253B