A target detection method, device, equipment, and storage medium
By restoring and spatial conversion processing of the distance image generated by the lidar signal, combined with the weighted non-maximum value suppression method, the problems of spatial misalignment and field of view loss in the prior art are solved, and the accuracy of target detection is improved.
Patent Information
- Application Number
- CN202311120713.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-08-31
AI Technical Summary
When processing distance images generated by lidar signals, the existing target detection methods ignore depth information, resulting in spatial misalignment and field of view loss, affecting detection accuracy.
Lidar is used to collect signals, determine distance images, and input them into the preset object detection model for recovery processing, spatial conversion processing, feature extraction processing, redundant feature pruning processing and perceptual prediction processing to generate a prediction matrix. Then, based on the weighted non-maximum suppression method, the object detection result is generated.
By performing recovery processing and spatial conversion processing on the distance image, field of view loss and spatial misalignment are avoided, and the accuracy of object detection is improved.
Smart Images

Figure CN116994222B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an object detection method, device, equipment, and storage medium. Background Art
[0002] In recent years, object detection perception based on lidar sensors, that is, object detection technology, has made great progress in the field of autonomous driving. However, the existing object detection process ignores the nature that the distance image contains rich depth information, and is prone to problems such as spatial misalignment and field of view loss.
[0003] Therefore, how to propose a new object detection framework to more comprehensively analyze the distance image generated by lidar signals and improve the accuracy of object detection is an urgent problem to be solved at present. Summary of the Invention
[0004] The present invention provides an object detection method, device, equipment, and storage medium, which can more comprehensively analyze the distance image generated by lidar signals and improve the accuracy of object detection.
[0005] According to one aspect of the present invention, an object detection method is provided, including:
[0006] Collecting signals using a lidar and determining a distance image corresponding to each frame of lidar signal collected;
[0007] Inputting the distance image into a preset object detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix;
[0008] Generating an object detection result for the distance image based on the weighted non-maximum suppression method according to the prediction matrix.
[0009] According to another aspect of the present invention, an object detection device is provided, including:
[0010] A determination module for collecting signals using a lidar and determining a distance image corresponding to each frame of lidar signal collected;
[0011] A processing module for inputting the distance image into a preset object detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix;
[0012] A generation module for generating an object detection result for the distance image based on the weighted non-maximum suppression method according to the prediction matrix.
[0013] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the object detection method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the object detection method according to any embodiment of the present invention when executed.
[0018] In the technical solution of the embodiment of the present invention, a lidar is used for signal acquisition, and a distance image corresponding to each frame of lidar signal collected is determined; the distance image is input into a preset object detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix; based on a weighted non-maximum suppression method, according to the prediction matrix, an object detection result of the distance image is generated. By performing restoration processing on the distance image, the problem of vision loss can be avoided. By performing spatial transformation processing, the situation of spatial disorder can be avoided. In this way, the distance image generated by the lidar signal can be analyzed more comprehensively, and the accuracy of object detection can be improved.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0021] Figure 1 is a flowchart of an object detection method provided in Embodiment 1 of the present invention;
[0022] Figure 2 is a schematic structural diagram of a framework of an object detection model provided in Embodiment 2 of the present invention;
[0023] Figure 3 It is a structural block diagram of an object detection device provided in Embodiment 3 of the present invention;
[0024] Figure 4 It is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0026] It should be noted that the terms "first", "second", "target", "candidate", "alternative", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] Embodiment 1
[0028] Figure 1 It is a flowchart of an object detection method provided in Embodiment 1 of the present invention; this embodiment is applicable to the situation of using an object detection model to perform object perception on the distance image corresponding to the radar acquisition signal and generating an object detection result, especially applicable to the situation of using a 3D object detection model to generate a 3D object detection result. This method can be executed by an object detection device, which can be implemented in the form of hardware and / or software, and the object detection device can be configured in an electronic device, such as an autonomous driving vehicle. As Figure 1 shown, the object detection method includes:
[0029] S101. Use a lidar to collect signals and determine the distance image corresponding to each frame of lidar signal collected.
[0030] Among them, each pixel in the range image contains at least information such as the response intensity and echo distance of the returned laser pulse. That is to say, each pixel in the range image corresponds to the optical response of the lidar beam.
[0031] Optionally, the collected lidar data can be represented as an m×n matrix, called a range image, where m represents the number of emitted beams and n represents the number of measurements within a scan cycle. For example, for a lidar with m beams and n measurements within a scan cycle, the return values of one scan form an m×n matrix, and a range image can be generated.
[0032] Exemplarily, each column of the range image corresponds to a shared azimuth angle, and each row corresponds to a shared inclination angle, indicating the relative vertical and horizontal angles of the return points with respect to the lidar origin. Each pixel in the range image contains at least three geometric values, namely the echo distance r, azimuth angle θ, and inclination angle φ, which define a spherical coordinate system. Represent the Cartesian coordinates of each pixel point as x, y, z:
[0033] x = rcos(φ)cos(θ), y = rcos(φ)sin(θ), z = rsin(η)
[0034] The magnitude of the returned laser pulse measured by the lidar sensor can be named the response intensity η and the elongation rate ρ. In summary, the range image of the lidar signal can be designed as where each pixel contains information such as the echo distance r, azimuth angle θ, inclination angle φ, response intensity η, and elongation rate ρ. The range image can also be represented by an m×n distance matrix representation.
[0035] Optionally, a lidar can be used to collect signals around the vehicle, and based on a preset rule, each frame of the collected lidar signal is parsed to determine the response intensity, echo distance, and echo angle of the signal, so as to generate a corresponding range image.
[0036] S102. Input the range image into a preset target detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the range image, and obtain a prediction matrix.
[0037] Optionally, the object detection model is a 3D object detection model; the object detection model includes: a Vision Restoration Module, a Range Aware Kernel, a backbone network, a Redundancy Pruner, and a Region Proposal Network. The backbone network can be a Deep Layer Aggregation (DLA-34). The Redundancy Pruner is used to eliminate redundancy in deep features, thereby minimizing the computational cost of the subsequent Region Proposal Network and post-processing layer.
[0038] Optionally, performing a restoration process, a spatial transformation process, a feature extraction process, a redundant feature pruning process, and a perception prediction process on the distance image to obtain a prediction matrix, including: using the Vision Restoration Module of the object detection model to perform a restoration process on the distance image to obtain a restored range image; using the Range Aware Kernel of the object detection model to spatially transform the restored range image into multiple subspaces, and using the backbone network to perform a feature extraction process in each subspace to obtain non-linear features; using the Redundancy Pruner to perform a redundant feature pruning process on the non-linear features, and using the Region Proposal Network to perform a perception prediction process based on the pruned non-linear features to obtain a prediction matrix.
[0039] Optionally, in the feature space, the image features of the area restored by the Vision Restoration Module can be used as redundant features and discarded to obtain non-linear features.
[0040] Optionally, using the Vision Restoration Module of the object detection model to perform a restoration process on the distance image to obtain a restored range image, including: using the Vision Restoration Module of the object detection model to determine a left area image and a right area image within a preset range at both ends of the distance image; performing a restoration process on the right area of the distance image according to the left area image, and performing a restoration process on the left area of the distance image according to the right area image; determining the distance image obtained after the left and right restoration processes as the restored range image. Among them, the restoration process refers to a padding process performed on both ends of the distance image.
[0041] It should be noted that when the target object is located in the edge region of the image, due to the limited receptive field of the convolutional neural network, when the two-dimensional convolutional neural network is used as a feature extractor, the features around the left boundary cannot be shared with the features of the right boundary, and vice versa, that is, there is no compensation for the damaged area. This phenomenon is called field of view loss, which will greatly affect the detection accuracy of objects near the edge of the distance image. The field of view restorer in the object detection model proposed by the present invention can restore the damaged area near the edge of the distance image, thereby expanding the receptive field of the subsequent backbone network and improving the accuracy of the final object detection.
[0042] It should be noted that by performing restoration processing at the image edge, the boundary information of the distance image can be effectively maintained. If no restoration processing is performed, the pixel point information at the outermost edge of the picture will only be operated on by the convolution kernel once, but the pixel points in the middle of the image will be scanned many times, which will reduce the boundary information to a certain extent.
[0043] Optionally, using the distance perception kernel of the object detection model, the restored range image space is converted into multiple subspaces, and a backbone network is used to perform feature extraction processing in each subspace to obtain non-linear features, including: using the distance perception kernel of the object detection model to convert the restored range image space into at least two subspaces, and generating a tensor matrix according to the subspace matrices corresponding to the at least two subspaces; using the backbone network to perform non-linear feature extraction processing in each subspace according to the tensor matrix to obtain non-linear features.
[0044] Exemplarily, the distance perception kernel has l predefined perception windows, expressed as W = {w 1 , w 2 , …, w l-1 , w l}, where each perception window is a distance interval w i = [r i1 , r i2 . Given a frame of distance image The binary mask of 1 , w 2 , …, w l-1 , w l} of the distance perception kernel can be calculated first. Each is the pixel-level mask of the distance image I, indicating whether each distance value R = I j,k is within the current perception window w j,k,1 i . Subsequently, the distance perception kernel defines a tensor to represent l subspaces, and calculates each subspace K i = Mi ⊙I. Through the above steps, the distance perception kernel divides the distance image I into multiple subspaces K, where each subspace K i contains all the spotlight radar points belonging to the perception window w i . Further, a tensor matrix K is generated according to each subspace Ki. The subspace K generated by the distance perception kernel belongs to R l×m×n×8 and is then sent to the backbone network for non-linear feature extraction.
[0045] Exemplarily, the calculation logic for the distance perception kernel to determine the tensor matrix K can be:
[0046]
[0047]
[0048] It should be noted that in the related art, the way of processing the distance image by the distance map-based detector is often to directly input it into the convolutional network for processing. This way ignores the property that the distance image contains rich depth information. Even if two distance pixels are adjacent in the distance map coordinates, their actual distance in the three-dimensional space may exceed 30 meters. For the above problem of spatial misalignment, directly processing such pixels that are not relevant in the three-dimensional space with a two-dimensional convolutional kernel can only produce features with large noise, which hinders the process of extracting geometric information from the edges of foreground objects. The distance perception kernel in the object detection model proposed by the present invention can decompose the two-dimensional distance image space into multiple subspaces, so that the subsequent backbone network can independently extract features from each subspace, thereby avoiding the influence of spatial misalignment.
[0049] Optionally, DLA-34 can be used as the backbone network to extract non-linear features from the subspace In addition, redundant features are eliminated by the redundant trimmer to generate trimmed features which facilitates subsequent learning of the dense prediction matrix based on the depth feature F through the region proposal network.
[0050] Optionally, the region proposal network is used for perception prediction processing to obtain the prediction matrix, including: using three parallel convolutional layers in the region proposal network to respectively train and learn the trimmed non-linear features to obtain a prediction matrix composed of three-dimensional information of a classification tensor (denoted as C), a position tensor (denoted as B), and the confidence score of the three-dimensional detection box (denoted as U). The essence of C, B, and U is Tensor (tensor).
[0051] Among them, C is the classification tensor, corresponding to the detection category of the 3D detection box. Each unit in C is activated by the sigmoid function, so its value range is [0, 1]. The larger the activation value, the more likely there is a target object. B is the position tensor, that is, Box (the position information of the 3D detection box). Box has a total of 8 dimensions: [center_x, center_y, center_z, box_dim_x, box_dim_y, box_dim_z, heading_cosine, heading_sine]. Among them, the first three dimensions are the 3D Cartesian coordinate positioning of the center of the 3D detection box. The middle three dimensions represent the lengths of the x, y, and z dimensions of the 3D detection box. The last two are the sine and cosine of the orientation angle of the 3D detection box. U is the intersection over union (IoU) between the predicted 3D detection box and the actual 3D detection box (i.e., the original labeled box), which is also the confidence score, mainly used for the subsequent statistical analysis process based on the weighted non-maximum suppression method.
[0052] Exemplarily, connecting three parallel convolutional layers on the feature space F can obtain C, B, and U of the prediction matrix. During training, the Focal loss is used to optimize C and U, and the L1 loss is used to optimize B.
[0053] S103. Based on the weighted non-maximum suppression method, generate the target detection result for the distance image according to the prediction matrix.
[0054] Among them, the weighted non-maximum suppression method (WNMS) takes into account that many 3D detection boxes are generated through the target detection model, and many of these 3D detection boxes will have repeated localization to the same target. Therefore, the weighted non-maximum suppression method is used to remove the duplicate boxes to achieve the final detection result of each target in the distance image.
[0055] Optionally, based on the weighted non-maximum suppression method, generating the target detection result for the distance image according to the prediction matrix includes: determining the category, position, and score of each detected 3D detection box according to the prediction matrix; based on the weighted non-maximum suppression method, summarizing and statistically analyzing the category, position, and score of each 3D detection box to generate the target detection result for the distance image.
[0056] Optionally, based on the weighted non-maximum suppression method, the 3D detection boxes located in the same position area range can be determined as the overlapping boxes of the same target. Further, according to the confidence scores of each 3D detection box, the final 3D detection box is determined from the overlapping boxes of the same target, and the target detection result of this target in the distance image is generated according to the category and position of the final 3D detection box.
[0057] In the technical solution of the embodiment of the present invention, a lidar is used to collect signals, and a distance image corresponding to each frame of lidar signal collected is determined; the distance image is input into a preset target detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image, and a prediction matrix is obtained; based on the weighted non-maximum suppression method, according to the prediction matrix, a target detection result of the distance image is generated. In this way, the distance image generated by the lidar signal can be analyzed more comprehensively, and the accuracy of target detection can be improved.
[0058] Embodiment 2
[0059] Figure 2 It is a schematic diagram of the framework structure of a target detection model provided in Embodiment 2 of the present invention; on the basis of the above embodiment, this embodiment gives a preferred example of 3D target detection using a vision restorer, distance perception kernel, backbone network, redundant trimmer, and region proposal network of the target detection model. As Figure 2 shown, the architecture of this target detection process includes a vision restorer, a distance perception kernel, a backbone network, a redundant trimmer, and a region proposal network, which are respectively used to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image.
[0060] Optionally, the target detection model proposed in this embodiment can be called Range Perception, which is essentially a single-stage detector. The specific target detection process is as follows:
[0061] (1) Given an image I within a certain range, the Range Perception framework first uses the vision restorer to compensate for the damaged area to obtain a restored image.
[0062] (2) Subsequently, the distance perception kernel converts the image into a subspace of K∈R^(l×m×μ×8), and decomposes the misaligned distance pixels.
[0063] (3) Use DLA-34 as the backbone network to extract non-linear features from the subspace. In addition, the redundant trimmer of RangePerception eliminates redundant features and generates trimmed features.
[0064] (4) The region proposal network learns a dense prediction matrix based on the depth features. Where U is regarded as the confidence score of each target three-dimensional detection box.
[0065] (5) Finally, the weighted non-maximum suppression strategy (Weighted NMS) is used to aggregate the dense predictions generated by the network.
[0066] The technical solution of the present invention designs and adopts a Range Aware Kernel, which solves the problem of spatial misalignment and improves the 3D detection accuracy. It also designs and adopts a Vision Restoration Module, which solves the problem of vision loss and improves the 3D detection accuracy. Compared with the widely used bird's-eye view-based method CenterPoint, Range Perception also shows superior performance. Experiments prove that the inference speed of Range Perception is 1.3 times that of CenterPoint, demonstrating its better suitability for deployment on autonomous driving vehicles for real-time inference.
[0067] Embodiment 3
[0068] Figure 3 It is a structural block diagram of an object detection device provided in Embodiment 3 of the present invention; this embodiment is applicable to the situation where an object detection model is used to perform object perception on a distance image corresponding to a radar acquisition signal to generate an object detection result, especially applicable to the situation of using a 3D object detection model to generate a 3D object detection result. The object detection device can be implemented in the form of hardware and / or software and configured in a device with object detection functions, such as an autonomous driving vehicle. As Figure 3 shown, the object detection device specifically includes:
[0069] A determination module 301, configured to collect signals using a lidar and determine a distance image corresponding to each frame of the collected lidar signal;
[0070] A processing module 302, configured to input the distance image into a preset object detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix;
[0071] A generation module 303, configured to generate an object detection result for the distance image based on a weighted non-maximum suppression method according to the prediction matrix.
[0072] The technical solution of the embodiment of the present invention uses a lidar for signal acquisition and determines a distance image corresponding to each frame of lidar signal collected; the distance image is input into a preset target detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix; based on the weighted non-maximum suppression method, according to the prediction matrix, a target detection result of the distance image is generated. By performing restoration processing on the distance image, the problem of field of view loss can be avoided. By performing spatial transformation processing, the situation of spatial disorder can be avoided. In this way, the distance image generated by the lidar signal can be analyzed more comprehensively, and the accuracy of target detection can be improved.
[0073] Further, the target detection model is a 3D target detection model; the target detection model includes: a field of view restorer, a distance perception kernel, a backbone network, a redundant pruner, and a region proposal network.
[0074] Further, the processing module 302 may include:
[0075] A restoration unit, configured to use the field of view restorer of the target detection model to perform restoration processing on the distance image to obtain a restored range image;
[0076] An extraction unit, configured to use the distance perception kernel of the target detection model to spatially transform the restored range image into multiple subspaces, and use the backbone network to perform feature extraction processing in each subspace to obtain non-linear features;
[0077] A prediction unit, configured to use the redundant pruner to perform redundant feature pruning processing on the non-linear features, and perform perception prediction processing using the region proposal network according to the pruned non-linear features to obtain a prediction matrix.
[0078] Further, the restoration unit is specifically configured to:
[0079] Use the field of view restorer of the target detection model to determine a left region image and a right region image within a preset range at both ends of the distance image;
[0080] According to the left region image, perform restoration processing on the right region of the distance image, and according to the right region image, perform restoration processing on the left region of the distance image;
[0081] Determine the distance image obtained after the left and right restoration processing as the restored range image.
[0082] Further, the extraction unit is specifically configured to:
[0083] Using the distance perception kernel of the object detection model, the restored range image space is converted into at least two subspaces, and a tensor matrix is generated according to the subspace matrices corresponding to the at least two subspaces;
[0084] Using a backbone network, according to the tensor matrix, non-linear feature extraction processing is performed in each subspace to obtain non-linear features.
[0085] Furthermore, the prediction unit is specifically used for:
[0086] Using three parallel convolutional layers in the region proposal network, the pruned non-linear features are respectively trained and learned to obtain a prediction matrix composed of three-dimensional information of a classification tensor, a position tensor, and the confidence score of a three-dimensional detection box.
[0087] Furthermore, the generation module 303 is specifically used for:
[0088] According to the prediction matrix, the categories, positions, and scores of the detected three-dimensional detection boxes are determined;
[0089] Based on the weighted non-maximum suppression method, the categories, positions, and scores of the three-dimensional detection boxes are summarized and statistically analyzed to generate the object detection result of the distance image.
[0090] Embodiment 4
[0091] Figure 4 It is a schematic structural diagram of the electronic device provided in Embodiment 4 of the present invention. Figure 4 A schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0092] Such as Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0093] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0094] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the object detection method.
[0095] In some embodiments, the object detection method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the object detection method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the object detection method by any other appropriate means (e.g., by means of firmware).
[0096] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0097] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0098] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0099] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0100] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0101] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0102] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0103] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A target detection method, characterized in that, comprising: Using a lidar to collect signals and determining a distance image corresponding to each frame of lidar signal collected; Inputting the distance image into a preset target detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix, including: using the field-of-view restorer of the target detection model to perform restoration processing on the distance image to obtain a restored range image; using the distance perception kernel of the target detection model to spatially transform the restored range image into multiple subspaces, and using a backbone network to perform feature extraction processing in each subspace to obtain non-linear features; using a redundant pruning device to perform redundant feature pruning processing on the non-linear features, and based on the pruned non-linear features, using a region proposal network to perform perception prediction processing to obtain a prediction matrix; Based on the weighted non-maximum suppression method, generating a target detection result for the distance image according to the prediction matrix; Among them, using the distance perception kernel of the target detection model to spatially transform the restored range image into multiple subspaces, and using a backbone network to perform feature extraction processing in each subspace to obtain non-linear features, including: using the distance perception kernel of the target detection model to spatially transform the restored range image into at least two subspaces, and generating a tensor matrix according to the subspace matrices corresponding to the at least two subspaces; using a backbone network to perform non-linear feature extraction processing in each subspace according to the tensor matrix to obtain non-linear features.
2. The method according to claim 1, characterized in that, wherein, the target detection model is a 3D target detection model; the target detection model includes: a field-of-view restorer, a distance perception kernel, a backbone network, a redundant pruning device, and a region proposal network.
3. The method according to claim 1, characterized in that, Using the field-of-view restorer of the target detection model to perform restoration processing on the distance image to obtain a restored range image, including: Using the field-of-view restorer of the target detection model to determine a left region image and a right region image within a preset range at both ends of the distance image; Performing restoration processing on the right region of the distance image according to the left region image, and performing restoration processing on the left region of the distance image according to the right region image; Determining the distance image obtained after performing left and right restoration processing as the restored range image.
4. The method according to claim 1, characterized in that, Using a region proposal network to perform perception prediction processing to obtain a prediction matrix, including: Using three parallel convolutional layers in the region proposal network to respectively perform training and learning on the pruned non-linear features to obtain a prediction matrix composed of three-dimensional information of a classification tensor, a position tensor, and a confidence score of a three-dimensional detection box.
5. The method according to claim 1, characterized in that, Based on the weighted non-maximum suppression method, generating a target detection result for the distance image according to the prediction matrix, including: According to the prediction matrix, determining the category, position, and score of each detected three-dimensional detection box; Based on the weighted non-maximum suppression method, the categories, positions, and scores of each 3D detection box are summarized and statistically analyzed to generate the object detection results for the distance image.
6. An object detection device, characterized in that, it includes: a determination module, configured to collect signals using a lidar and determine the distance image corresponding to each frame of the collected lidar signal; a processing module, configured to input the distance image into a preset object detection model to perform restoration processing, spatial transformation processing, feature extraction processing, redundant feature pruning processing, and perception prediction processing on the distance image to obtain a prediction matrix; a generation module, configured to generate the object detection results for the distance image based on the weighted non-maximum suppression method according to the prediction matrix; wherein, the processing module includes: a restoration unit, configured to use the field-of-view restorer of the object detection model to perform restoration processing on the distance image to obtain a restored range image; an extraction unit, configured to use the distance perception kernel of the object detection model to spatially transform the restored range image into multiple subspaces, and use a backbone network to perform feature extraction processing in each subspace to obtain non-linear features; a prediction unit, configured to use a redundant pruning device to perform redundant feature pruning processing on the non-linear features, and perform perception prediction processing using a region proposal network according to the pruned non-linear features to obtain a prediction matrix; wherein, the extraction unit is specifically configured to: use the distance perception kernel of the object detection model to spatially transform the restored range image into at least two subspaces, and generate a tensor matrix according to the subspace matrices corresponding to the at least two subspaces; use a backbone network to perform non-linear feature extraction processing in each subspace according to the tensor matrix to obtain non-linear features.
7. An electronic device, characterized in that, the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the object detection method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to implement the object detection method according to any one of claims 1-5 when executed.
Citation Information
Patent Citations
Three-dimensional scene image recovery method based on erasure codes
CN111739158A
Target detection method, target detection model training method and related equipment
CN114764778A