Target detection method and device, terminal equipment and storage medium

By generating a simple point cloud detection student model from the complex point cloud detection teacher model, and using knowledge distillation technology to guide training, the problem of point cloud object detection time is solved, and fast and accurate object detection is achieved.

CN120279504AInactive Publication Date: 2025-07-08VANJEE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311873387.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Although the existing point cloud object detection algorithm has good detection effect, its complex structure leads to a long detection time.

Method used

Through knowledge distillation technology, a point cloud detection student model with a simpler structure and shorter detection time is generated from the pre-trained complex point cloud detection teacher model, and the point cloud detection teacher model is used to guide the training of the student model.

Benefits of technology

While maintaining the detection effect, the time for target detection is significantly shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279504A_ABST
    Figure CN120279504A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computers, and provides a target detection method and device, terminal equipment and a storage medium, and the method comprises the steps: firstly obtaining the point cloud data of a to-be-detected target, and then inputting the point cloud data of the to-be-detected target into a preset point cloud detection student model, the point cloud detection student model outputs a target detection result of the to-be-detected target, and the point cloud detection student model is generated through guidance training of a pre-trained point cloud detection teacher model. Therefore, the point cloud detection teacher model which is more complex in training structure and better in detection effect is used for guiding the point cloud detection student model which is simpler in structure and shorter in detection time to train, and the point cloud detection student model is used for detection in practical application, so that the target detection time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to an object detection method, device, terminal device, and storage medium. Background Art

[0002] By installing devices such as lidar and cameras on the roadside, the perception of the intersection scene can be realized, and the intelligent management of the intersection can be achieved. The main sensor used in the roadside perception system is lidar. By performing object detection on the 3D point cloud data generated by lidar scanning, the object information in the scene can be obtained.

[0003] In related technologies, common point cloud object detection algorithms are methods based on voxel partitioning, such as algorithms like pointpillar and voxelnet. It is found in actual applications that the more complex the structure of the detection algorithm, the richer the extracted features and the better the detection effect. However, this also brings the problem of longer detection time. Summary of the Invention

[0004] Embodiments of this application provide an object detection method, device, terminal device, and storage medium, which can solve the problem of longer detection time for detection algorithms with complex structures.

[0005] In a first aspect of the embodiments of this application, an object detection method is provided, including: obtaining point cloud data of a target to be detected; inputting the point cloud data of the target to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs an object detection result of the target to be detected, where the point cloud detection student model is generated under the guidance of a pre-trained point cloud detection teacher model.

[0006] Optionally, in a possible implementation manner of the first aspect, the above point cloud detection teacher model includes multiple modules, and the point cloud detection student model is generated through the following steps:

[0007] Performing knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models;

[0008] Determining the point cloud detection student model from the multiple candidate student models.

[0009] Optionally, in another possible implementation manner of the first aspect, the above performing knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models includes:

[0010] Obtaining multiple candidate student models by removing the sparse convolution module of the point cloud detection teacher model.

[0011] Optionally, in another possible implementation manner of the first aspect, the above performing knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models includes:

[0012] By removing the convolutional modules of the Region Proposal Network (RPN) layer of the point cloud detection teacher model, multiple candidate student models are obtained.

[0013] Optionally, in another possible implementation manner of the first aspect, determining the point cloud detection student model from the multiple candidate student models includes:

[0014] Training each candidate student model using a preset point cloud training set to continuously make the results of each candidate student model approach the results of the point cloud detection teacher model;

[0015] After the training is completed, determining the loss value corresponding to each candidate student model;

[0016] Determining the candidate student model with the minimum loss value as the point cloud detection student model.

[0017] Optionally, in another possible implementation manner of the first aspect, the loss function corresponding to each candidate student model includes: the loss value between the prediction result of each candidate student model and the preset label, and the loss value between the prediction result of each candidate student model and the inference result of the point cloud detection teacher model.

[0018] Optionally, in yet another possible implementation manner of the first aspect, the various candidate student models are trained together.

[0019] Optionally, in another possible implementation manner of the first aspect, the above-mentioned point cloud detection teacher model is a voxel-based object detection model.

[0020] A second aspect of the embodiments of the present application provides an object detection device, including:

[0021] A third aspect of the embodiments of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the object detection method of the first aspect is implemented.

[0022] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the object detection method of the first aspect is implemented.

[0023] A fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute the object detection method of the first aspect.

[0024] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The embodiments of the present application disclose a target detection method, device, terminal device, and storage medium. Among them, the method first obtains the point cloud data of the target to be detected, and then inputs the point cloud data of the target to be detected into a preset point cloud detection student model. The point cloud detection student model outputs the target detection result of the target to be detected, where the point cloud detection student model is generated by being guided and trained by a pre-trained point cloud detection teacher model. Thus, by training the point cloud detection student model with a simpler structure and shorter detection time by guiding it with a point cloud detection teacher model with a more complex structure and better detection effect, and using the point cloud detection student model for detection in actual applications, the time for target detection is shortened. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0026] Figure 1 It is a schematic flowchart of a target detection method provided by an embodiment of the present application;

[0027] Figure 2 It is a schematic structural diagram of a target detection device provided by an embodiment of the present application;

[0028] Figure 3 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0030] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0031] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0032] As used in the specification and appended claims of this application, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".

[0033] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0034] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0035] It should be understood that the magnitudes of the sequence numbers of the steps in this embodiment do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of this application.

[0036] In the related art, the commonly used point cloud object detection algorithms are methods based on voxel partitioning, such as algorithms like pointpillar and voxelnet. It has been found in practical applications that the more complex the detection algorithm structure, the richer the extracted features and the better the detection effect, but it will also bring the problem of longer detection time.

[0037] In view of this, an embodiment of the present application provides an object detection method, apparatus, terminal device, and storage medium. First, point cloud data of the object to be detected is obtained, and then the point cloud data of the object to be detected is input into a preset point cloud detection student model, and the point cloud detection student model outputs the object detection result of the object to be detected. Among them, the point cloud detection student model is generated by guiding the training with a pre-trained point cloud detection teacher model. Thus, by training the point cloud detection student model with a simpler structure and shorter detection time under the guidance of the point cloud detection teacher model with a more complex structure and better detection effect, and using the point cloud detection student model for detection in actual applications, the time for object detection is shortened.

[0038] To illustrate the technical solution of the present application, specific embodiments are used for illustration below.

[0039] Refer to Figure 1 , which shows a schematic flowchart of an object detection method provided by an embodiment of the present application.

[0040] As Figure 1 shown, the object detection method may include the following steps:

[0041] Step 101, obtain point cloud data of the object to be detected.

[0042] Among them, the point cloud data of the object to be detected can be collected and obtained by a lidar installed on the roadside.

[0043] Step 102, input the point cloud data of the object to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs the object detection result of the object to be detected.

[0044] Among them, the point cloud detection student model is generated by guiding the training with a pre-trained point cloud detection teacher model.

[0045] In a possible implementation manner of the present application, knowledge distillation can be introduced into the point cloud detection algorithm. First, a teacher model with a more complex structure and better detection effect is trained, and then the training of a student model with a simpler structure and shorter detection time is guided. The student model is used for detection in actual applications to ensure fast and good performance in actual detection. That is to say, knowledge distillation can be performed based on the point cloud detection teacher model to obtain multiple candidate student models; from the multiple candidate student models, the point cloud detection student model is determined.

[0046] Among them, knowledge distillation is a model compression technology. Knowledge distillation realizes model compression and acceleration by "distilling" the knowledge of a large and complex model into a small and simple model.

[0047] As an example, the point cloud detection teacher model can be a voxel-based object detection model. A voxel is a basic unit in three-dimensional space, similar to a pixel in a two-dimensional image. A voxel can be regarded as a cube with a fixed size and shape. The principle of the voxel-based object detection model is to divide the point cloud data into voxel grids with spatial dependency relationships, introduce spatial dependency relationships into the point cloud data by segmenting the three-dimensional space, and then use methods such as 3D convolution for processing.

[0048] In the embodiments of the present application, the point cloud detection teacher model is pre-trained through a point cloud dataset to obtain the pre-trained weights of the point cloud detection teacher model. The pre-trained weights refer to the model parameters learned by the teacher model during the pre-training stage. During the pre-training stage, the teacher model usually uses a large-scale point cloud dataset for training to learn the feature representation of the point cloud data. The pre-training process is usually unsupervised, that is, it does not require labeled data, but learns the feature representation of the point cloud data through self-supervised learning or unsupervised tasks. The pre-trained weights can usually be used to initialize a new model to accelerate the training of the model and improve the performance of the model. In knowledge distillation, the pre-trained weights can be used as the initial parameters of the teacher model to generate soft targets and guide the training of the student model. The quality and effect of the pre-trained weights largely depend on the quality and scale of the pre-training data, as well as the design and optimization of the pre-training tasks.

[0049] It should be noted that the model can be simplified by removing time-consuming modules from the original detection model, that is, the point cloud detection teacher model, to obtain a candidate point cloud detection student model. There are mainly two methods. One is distillation based on the feature level, and the other is distillation based on the result level.

[0050] As an example, distillation can be performed based on the feature level. Since the voxel-based object detection model consumes a lot of time in the sparse convolution part, the student model can be simplified by reducing the number of sparse convolution layers. That is to say, multiple candidate student models can be obtained by removing the sparse convolution module of the point cloud detection teacher model.

[0051] For example, assuming there are N sparse convolution modules, by reducing the sparse convolution modules, candidate student models with 1, 2... N-1 convolution modules can be obtained respectively.

[0052] As another example, distillation can be performed based on the result level, that is, the student model can be simplified by reducing the number of RPN (Region Proposal Network) layers. That is to say, multiple candidate student models can be obtained by removing the convolution module of the RPN layer of the point cloud detection teacher model.

[0053] It should be noted that the main function of the RPN layer is to generate candidate bounding boxes, which are used to represent regions that may contain target objects. The RPN layer usually contains one or more convolutional modules, which are used to extract features from the input image and calculate the scores of each candidate bounding box. The RPN layer generally serves as the first layer of the object detection model, receiving the input image and sliding a window of a fixed size over the input image to generate multiple candidate bounding boxes for each window. The positions and sizes of these candidate bounding boxes are calculated by applying a set of predefined anchor points to each window. Anchor points are rectangular boxes with predefined sizes and aspect ratios, which are placed at each position on the input image to generate a set of candidate bounding boxes with fixed sizes and aspect ratios. In the RPN layer, each candidate bounding box is assigned a score to represent the likelihood that the candidate bounding box contains the target object. This score is calculated by the neural network in the RPN layer, which is usually a convolutional neural network (CNN). Finally, the RPN layer sorts all candidate bounding boxes according to the scores and selects some candidate bounding boxes with the highest scores as the final object detection results. These candidate bounding boxes will be passed to the subsequent layers of the object detection model for further object detection and classification.

[0054] For example, assume that the RPN layer has M modules. By reducing the convolutional modules in the RPN layer, candidate student models with 1, 2... M - 1 convolutional modules can be obtained respectively.

[0055] In a possible implementation manner of the embodiment of the present application, each candidate student model can be trained first using a preset point cloud training set to continuously make the results of each candidate student model tend to the results of the point cloud detection teacher model; after the training is completed, the loss value corresponding to each candidate student model is determined; the candidate student model with the smallest loss value is determined as the point cloud detection student model.

[0056] It should be noted that after obtaining the candidate student models, each candidate student model needs to be trained using the point cloud training set. After multiple rounds of training, the candidate student model with the smallest loss value can be selected as the finally used point cloud detection student model. The smallest loss value indicates that its simplified model structure is most suitable for this data set and has the best effect.

[0057] Among them, the loss function of each student model during the training process is: L = L1 + L2. Wherein, L1 represents the loss value between the prediction result of each candidate student model and the preset label, and L2 represents the loss value between the prediction result of each candidate student model and the inference result of the point cloud detection teacher model. During the training process, the results of the candidate student models are continuously made to tend to the results of the point cloud detection teacher model.

[0058] As a possible implementation, each candidate student model can be trained jointly to improve the training efficiency.

[0059] For the object detection method disclosed in the above embodiments of the present application, first, the point cloud data of the object to be detected is obtained, and then the point cloud data of the object to be detected is input into a preset point cloud detection student model, and the point cloud detection student model outputs the object detection result of the object to be detected, where the point cloud detection student model is generated by being guided by a pre-trained point cloud detection teacher model. Thus, by training the point cloud detection student model with a simpler structure and shorter detection time under the guidance of the point cloud detection teacher model with a more complex structure and better detection effect, the object detection time is shortened when using the point cloud detection student model in actual applications.

[0060] See Figure 2 , which shows a schematic structural diagram of an object detection device provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0061] The object detection device may specifically include the following modules:

[0062] A data acquisition module, configured to acquire the point cloud data of the object to be detected.

[0063] An object detection module, configured to input the point cloud data of the object to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs the object detection result of the object to be detected, where the point cloud detection student model is generated by being guided by a pre-trained point cloud detection teacher model.

[0064] For the object detection device disclosed in the above embodiments of the present application, first, the point cloud data of the object to be detected is obtained, and then the point cloud data of the object to be detected is input into a preset point cloud detection student model, and the point cloud detection student model outputs the object detection result of the object to be detected, where the point cloud detection student model is generated by being guided by a pre-trained point cloud detection teacher model. Thus, by training the point cloud detection student model with a simpler structure and shorter detection time under the guidance of the point cloud detection teacher model with a more complex structure and better detection effect, the object detection time is shortened when using the point cloud detection student model in actual applications.

[0065] Further, in a possible implementation manner of the embodiments of the present application, the above point cloud detection teacher model includes multiple modules, and the object detection device may specifically include the following modules:

[0066] A first processing module, configured to perform knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models.

[0067] A second processing module, configured to determine the point cloud detection student model from the multiple candidate student models.

[0068] Further, in another possible implementation manner of the embodiment of the present application, the above first processing module may specifically include the following sub-modules:

[0069] The first processing sub-module is configured to obtain a plurality of candidate student models by removing the sparse convolution module of the point cloud detection teacher model.

[0070] Further, in another possible implementation manner of the embodiment of the present application, the above first processing module may specifically include the following sub-modules:

[0071] The second processing sub-module is configured to obtain a plurality of candidate student models by removing the convolution module of the RPN layer of the point cloud detection teacher model.

[0072] Further, in another possible implementation manner of the embodiment of the present application, the above second processing module may specifically include the following sub-modules:

[0073] The third processing sub-module is configured to train each candidate student model by using a preset point cloud training set, so as to continuously make the results of each candidate student model tend to the results of the point cloud detection teacher model.

[0074] The fourth processing sub-module is configured to determine the loss value corresponding to each candidate student model after the training ends.

[0075] The fifth processing sub-module is configured to determine the candidate student model with the smallest loss value as the point cloud detection student model.

[0076] Further, in another possible implementation manner of the embodiment of the present application, the loss function corresponding to each candidate student model includes: the loss value between the prediction result of each candidate student model and a preset label, and the loss value between the prediction result of each candidate student model and the inference result of the point cloud detection teacher model.

[0077] Further, in another possible implementation manner of the embodiment of the present application, the various candidate student models are trained together.

[0078] Further, in another possible implementation manner of the embodiment of the present application, the above point cloud detection teacher model is a voxel-based object detection model.

[0079] The object detection device provided by the embodiment of the present application can be applied to the foregoing method embodiment. For details, please refer to the description of the foregoing method embodiment, which will not be repeated here.

[0080] Figure 3 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. As Figure 3 shown, the terminal device 300 of this embodiment includes: at least one processor 310 (Figure 3 Only one processor, a memory 320, and a computer program 321 stored in the memory 320 and executable on the at least one processor 310 are shown. When the processor 310 executes the computer program 321, the following steps are implemented: obtaining point cloud data of a target to be detected; inputting the point cloud data of the target to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs a target detection result of the target to be detected, wherein the point cloud detection student model is generated by being guided and trained by a pre-trained point cloud detection teacher model.

[0081] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing knowledge distillation based on the point cloud detection teacher model to obtain a plurality of candidate student models; determining the point cloud detection student model from the plurality of candidate student models.

[0082] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining a plurality of candidate student models by removing the sparse convolution module of the point cloud detection teacher model.

[0083] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining a plurality of candidate student models by removing the convolution module of the RPN layer of the point cloud detection teacher model.

[0084] In one embodiment, when the processor executes the computer program, the following steps are further implemented: training each candidate student model by using a preset point cloud training set to continuously make the results of each candidate student model tend to the results of the point cloud detection teacher model; after the training ends, determining the loss value corresponding to each candidate student model; determining the candidate student model with the smallest loss value as the point cloud detection student model.

[0085] In one embodiment, the loss function corresponding to each candidate student model includes: the loss value between the prediction result of each candidate student model and a preset label, and the loss value between the prediction result of each candidate student model and the inference result of the point cloud detection teacher model.

[0086] In one embodiment, the various candidate student models are trained together.

[0087] In one embodiment, the above-mentioned point cloud detection teacher model is a voxel-based target detection model.

[0088] For the terminal device provided in the above embodiment, its implementation principle and technical effect are similar to those of the above method embodiment, and will not be elaborated here.

[0089] The terminal device 300 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 310 and a memory 320. Those skilled in the art can understand that Figure 3 merely examples of the terminal device 300, and do not constitute a limitation on the terminal device 300. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0090] The so-called processor 310 may be a central processing unit (CPU), and the processor 310 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0091] The memory 320 may be an internal storage unit of the terminal device 300 in some embodiments, such as the hard disk or memory of the terminal device 300. The memory 320 may also be an external storage device of the terminal device 300 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the terminal device 300. Further, the memory 320 may also include both the internal storage unit and the external storage device of the terminal device 300. The memory 320 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 320 may also be used to temporarily store data that has been output or will be output.

[0092] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0093] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0094] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0095] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the above-mentioned module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0096] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0097] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0098] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the following steps are implemented: obtaining point cloud data of a target to be detected; inputting the point cloud data of the target to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs a target detection result of the target to be detected, where the point cloud detection student model is generated by being guided and trained by a pre-trained point cloud detection teacher model.

[0099] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models; determining the point cloud detection student model from the multiple candidate student models.

[0100] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining multiple candidate student models by removing the sparse convolution module of the point cloud detection teacher model.

[0101] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining multiple candidate student models by removing the convolution module of the RPN layer of the point cloud detection teacher model.

[0102] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: training each candidate student model by using a preset point cloud training set to continuously make the results of each candidate student model tend to the results of the point cloud detection teacher model; after the training is completed, determining the loss value corresponding to each candidate student model; determining the candidate student model with the smallest loss value as the point cloud detection student model.

[0103] In one embodiment, the loss function corresponding to each candidate student model includes: the loss value between the prediction result of each candidate student model and a preset label, and the loss value between the prediction result of each candidate student model and the inference result of the point cloud detection teacher model.

[0104] In one embodiment, the various candidate student models are trained together.

[0105] In one embodiment, the above-mentioned point cloud detection teacher model is a voxel-based object detection model.

[0106] The computer-readable storage medium provided by the above-mentioned embodiment has the same implementation principle and technical effect as the above-mentioned method embodiment, and will not be elaborated here.

[0107] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0108] All or part of the processes in the above-mentioned method embodiments of the present application can also be completed by a computer program product. When the computer program product runs on a terminal device, the terminal device can be enabled to execute the steps in the above-mentioned method embodiments.

[0109] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A target detection method, characterized in that, Including: Obtain the point cloud data of the target to be detected; Input the point cloud data of the target to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs the target detection result of the target to be detected, wherein the point cloud detection student model is generated under the guidance of a pre-trained point cloud detection teacher model.

2. The object detection method according to claim 1, wherein The point cloud detection teacher model includes multiple modules, and the point cloud detection student model is generated through the following steps: Perform knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models; Determine the point cloud detection student model from the multiple candidate student models.

3. The object detection method according to claim 2, wherein The performing knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models includes: Obtain the multiple candidate student models by removing the sparse convolution module of the point cloud detection teacher model.

4. The object detection method according to claim 2, wherein The performing knowledge distillation based on the point cloud detection teacher model to obtain multiple candidate student models includes: Obtain the multiple candidate student models by removing the convolution module of the RPN layer of the point cloud detection teacher model.

5. The object detection method according to claim 2, wherein The determining the point cloud detection student model from the multiple candidate student models includes: Use a preset point cloud training set to train each candidate student model to continuously make the results of each candidate student model approach the results of the point cloud detection teacher model; After the training ends, determine the loss value corresponding to each candidate student model; Determine the candidate student model with the smallest loss value as the point cloud detection student model.

6. The object detection method according to claim 5, wherein The loss function corresponding to each candidate student model includes: the loss value between the prediction result of each candidate student model and the preset label, and the loss value between the prediction result of each candidate student model and the inference result of the point cloud detection teacher model.

7. The object detection method according to claim 5, wherein All the candidate student models are trained together.

8. The object detection method according to any one of claims 1-7, characterized in that The point cloud detection teacher model is a voxel-based target detection model.

9. A target detection device, characterized in that, Including: A data acquisition module for obtaining the point cloud data of the target to be detected; A target detection module for inputting the point cloud data of the target to be detected into a preset point cloud detection student model, and the point cloud detection student model outputs the target detection result of the target to be detected, wherein the point cloud detection student model is generated under the guidance of a pre-trained point cloud detection teacher model.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 8 is implemented.