Moving target detection method, device, flying device and storage medium
By aligning the coordinate system and coding the point cloud data of the flight equipment, combined with a variety of mask detection, the error detection and real-time problems of traditional algorithms when detecting motion obstacles are solved, and effective identification and accurate detection of unlabeled objects are achieved.
Patent Information
- Application Number
- CN202211741208.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-12-30
AI Technical Summary
When traditional algorithms detect motion obstacles in low-altitude flights of flight equipment, there is a false detection phenomenon, which cannot effectively utilize timing information, and cannot identify unmarked objects, affecting the real-time and accuracy of detection.
By performing coordinate system alignment processing of the current frame point cloud data with historical multi-frame point cloud data, voxelized encoding and feature extraction, combined with multiple masks for detection, point cloud data is processed directly end-to-end, using timing information and reducing the error detection rate.
It improves the real-time and accuracy of detection, can effectively identify unmarked objects, reduce the error detection rate, and accurately determine the moving target.
Smart Images

Figure CN115937259B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target detection, and more specifically, to a moving target detection method, device, flying device, and storage medium. Background Art
[0002] During the low-altitude flight of a flying device (such as a manned aircraft), it will encounter moving objects in the air, such as kites, birds, small unmanned aircraft, etc. These moving obstacles will affect the flight safety of the flying device. Therefore, the detection of moving targets has become an important issue for the low-altitude flight safety of flying devices.
[0003] Traditional algorithms process single-frame point clouds and identify moving targets based on the speed threshold of obstacles. Deep learning algorithms usually use a dataset with labeled targets to design a corresponding target detection network and infer single-frame point clouds to detect obstacles.
[0004] However, traditional algorithms may produce false detections for irregular obstacles in the real world. The objects detected by deep learning algorithms are all known labeled categories, and static targets will also be detected simultaneously. In addition, the single-frame detection method of deep learning does not fully utilize temporal information, affecting the real-time performance of the algorithm. Summary of the Invention
[0005] In view of the above problems, the present application proposes a moving target detection method, device, flying device, and storage medium to improve the above problems.
[0006] In a first aspect, an embodiment of the present application provides a moving target detection method, which includes: performing coordinate system alignment processing on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data; performing voxelization encoding on the fused point cloud data to obtain an encoding matrix; performing feature extraction on the encoding matrix to obtain a target feature map; determining multiple prediction masks according to the target feature map; determining a non-empty mask according to the encoding matrix; determining a detection mask according to the multiple prediction masks and the non-empty mask; and performing target extraction on the current frame point cloud data based on the detection mask to determine moving targets in the current frame point cloud data.
[0007] Second aspect, an embodiment of the present application provides a moving target detection device, which includes: a point cloud fusion module, configured to perform coordinate system alignment processing according to the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data; a voxelization module, configured to perform voxelization encoding on the fused point cloud data to obtain an encoded matrix; a feature extraction module, configured to perform feature extraction on the encoded matrix to obtain a target feature map; a prediction mask acquisition module, configured to determine multiple prediction masks according to the target feature map; a non-empty mask acquisition module, configured to determine a non-empty mask according to the encoded matrix; a detection mask acquisition module, configured to determine a detection mask according to the multiple prediction masks and the non-empty mask; and a moving target detection module, configured to perform target extraction on the current frame point cloud data based on the detection mask to determine moving targets in the current frame point cloud data.
[0008] Third aspect, an embodiment of the present application provides a flying device, including one or more processors; a memory; and one or more application programs, where the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the moving target detection method described in the first aspect above.
[0009] Fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which program code is stored, and when the program code is run by a processor, it performs the moving target detection method described in the first aspect above.
[0010] The technical solution provided by the present invention is to perform coordinate system alignment processing according to the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data; perform voxelization encoding on the fused point cloud data to obtain an encoded matrix; perform feature extraction on the encoded matrix to obtain a target feature map; determine multiple prediction masks according to the target feature map; determine a non-empty mask according to the encoded matrix; determine a detection mask according to the multiple prediction masks and the non-empty mask; and perform target extraction on the current frame point cloud data based on the detection mask to determine moving targets in the current frame point cloud data. This method directly processes point cloud data end-to-end, avoiding error accumulation caused by the linear process of traditional algorithms, and fusing historical multi-frame point cloud data, which can make full use of the temporal information of point cloud data and enhance the real-time performance of the algorithm; this method utilizes the advantages of voxel representation and can detect unlabeled objects; in addition, the combined use of multiple masks reduces the probability of false detection of the algorithm and can accurately determine moving targets. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It shows a schematic flowchart of a moving target detection method proposed in an embodiment of the present application.
[0013] Figure 2 It shows a schematic flowchart of step 110 in the embodiment of the present application.
[0014] Figure 3 It shows a schematic flowchart of step 120 in the embodiment of the present application.
[0015] Figure 4 It shows a schematic diagram of a voxel space in the embodiment of the present application.
[0016] Figure 5 It shows a schematic flowchart of step 121 in the embodiment of the present application.
[0017] Figure 6 It shows a schematic diagram of a spatio-temporal network model in the embodiment of the present application.
[0018] Figure 7 It shows a structural block diagram of a moving target detection device proposed in an embodiment of the present application.
[0019] Figure 8 It shows a structural block diagram of a flying device proposed in an embodiment of the present application.
[0020] Figure 9 It shows a structural block diagram of a computer-readable storage medium proposed in an embodiment of the present application. Detailed implementation manners
[0021] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application.
[0022] During the low-altitude flight of a flying device (such as a manned aircraft), it will encounter moving objects in the air, such as kites, birds, small unmanned aircraft, etc. These moving obstacles will affect the flight safety of the flying device. Therefore, the detection of moving targets has become an important issue for the low-altitude flight safety of flying devices.
[0023] Traditional algorithms process single-frame point clouds to determine attributes such as the category and speed of obstacles, and finally distinguish moving targets from stationary targets based on speed thresholds. Deep learning algorithms usually utilize a dataset of labeled targets to design corresponding object detection networks, and finally directly use the network model to infer single-frame point clouds to detect obstacles.
[0024] However, traditional algorithms cannot cover all situations in the real world. For irregular obstacles in the real world, misdetection may occur, and many parameters in the algorithms need to be adjusted manually, resulting in low efficiency. The objects detected by deep learning algorithms are all known labeled categories, and unlabeled objects cannot be recognized. Moreover, static targets will be detected simultaneously. In addition, the single-frame detection method of deep learning does not fully utilize temporal information, affecting the real-time performance of the algorithm.
[0025] To address the above problems, the inventors proposed the motion target detection method, device, flying device, and storage medium provided in this application. By performing coordinate system alignment processing based on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data, fused point cloud data is obtained; the fused point cloud data is voxelized and encoded to obtain an encoded matrix; a detection mask is determined based on multiple prediction masks and a non-empty mask; feature extraction is performed on the encoded matrix to obtain a target feature map; multiple prediction masks are determined based on the target feature map; a non-empty mask is determined based on the encoded matrix; target extraction is performed on the current frame point cloud data based on the detection mask to determine the moving targets in the current frame point cloud data. This method directly processes point cloud data end-to-end, avoiding error accumulation caused by the linear process of traditional algorithms, and fusing historical multi-frame point cloud data, which can fully utilize the temporal information of point cloud data and enhance the real-time performance of the algorithm; this method utilizes the advantages of voxel representation and can detect unlabeled objects; in addition, the combined use of multiple masks reduces the probability of misdetection of the algorithm and can accurately determine moving targets.
[0026] The following will specifically describe the embodiments of the present application in conjunction with the drawings.
[0027] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of a motion target detection method proposed in an embodiment of the present application. This method can be applied to a flying device. In the embodiments of the present application, the flying device may include a processor, and the processor can execute the motion target detection method provided in the embodiments of the present application. The motion target detection method of the embodiments of the present application includes steps 110 to 170.
[0028] Step 110: Perform coordinate system alignment processing based on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, the historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data.
[0029] In an embodiment of the present application, the flying device may include a point cloud acquisition device, which is arranged around the fuselage of the aircraft to collect point cloud images of the environment around the aircraft.
[0030] In some embodiments, the point cloud acquisition device may employ devices such as panoramic cameras, laser scanners (Laser Scanner / LiDAR Light Detection And Ranging), etc. to obtain richer environmental information around the aircraft.
[0031] It can be understood that the point cloud acquisition device can collect the point cloud data of the current time frame and the point cloud data of the past N frames relative to the current time frame. Among them, the value of N is related to the running efficiency of the method. The larger the value of N, the higher the detection accuracy of the method, but the lower the running efficiency. Preferably, when N is taken as 4, the detection accuracy and running efficiency can be balanced.
[0032] In some embodiments, the flying device further includes a radar device, where the radar device may be an infrared forward-looking radar or a lidar, used to obtain the combined navigation data corresponding to the time stamps of the current frame point cloud data and the historical N-frame point cloud data.
[0033] Since the coordinate systems of the point cloud data of different frames are different, therefore, taking the coordinate system where the current frame point cloud data is located as the reference coordinate system, the point cloud data of the past N frames is converted to the reference coordinate system corresponding to the current frame point cloud data, that is, the coordinate system alignment processing is completed. The conversion of coordinates in different coordinate systems may include translation transformation, rotation transformation, and composite transformation.
[0034] At the same time, the point cloud data of the current frame itself will also be merged with the past N frames of point clouds after the coordinate system conversion. In this way, a (N + 1)-frame fused point cloud data is formed in the coordinate system of the current frame point cloud data. This method encrypts the point cloud data of the current frame by incorporating the point cloud data of the past N frames. At the same time, the point cloud information of the moving objects in the past N frames is introduced, making full use of the temporal information, facilitating the detection of moving targets, and enhancing the real-time performance of the algorithm.
[0035] Step 120: Perform voxelization encoding on the fused point cloud data to obtain an encoding matrix.
[0036] In an embodiment of the present application, a voxel space needs to be preset. The voxel space includes a plurality of voxels, and the size of each voxel is preset. In addition, a preset point cloud boundary value is included. According to the voxelization size and the point cloud boundary value, the index (i, j, k) of each point of the fused point cloud data in the voxel space S is calculated to complete the voxelization of the fused point cloud data.
[0037] Traverse the index (i, j, k) of the fused point cloud data in three dimensions of the preset voxel space S. Among them, the voxel with point cloud data at the index (i, j, k) is marked as 1, and the voxel without point cloud data at the index is marked as 0. Thus, a three-dimensional 0-1 coding matrix can be obtained, and the size of the coding matrix is related to the point cloud boundary value and the voxelization size.
[0038] The voxelized point cloud data is stored in memory in an orderly manner, which is beneficial to reducing random memory access, increasing the efficiency of data operation, and facilitating the processing of a large number of point clouds of an order of magnitude. Moreover, the voxelized data can efficiently use spatial convolution, which is beneficial to extracting multi-scale and multi-level local feature information.
[0039] Step 130: Determine a non-empty mask according to the coding matrix.
[0040] In an embodiment of the present application, the voxel space preset in step 120 is three-dimensional. Flatten the three-dimensional voxel space to obtain a two-dimensional image.
[0041] In some embodiments, the index (i, j, k) of the fused point cloud data is transformed through a coordinate system to remove the height dimension k, and a two-dimensional coordinate (i, j) is obtained. (i, j) is the pixel coordinate corresponding to the index (i, j, k) of the fused point cloud data in the voxel space in the two-dimensional image.
[0042] Traverse each pixel point in the two-dimensional image, mark the pixel point with point cloud as 1, and mark the pixel without point cloud as 0. Thus, a 0-1 image is obtained, and this image is used as the non-empty mask.
[0043] Generally, the fused point cloud data corresponding to the pixel points occupied by the point cloud is a possible obstacle, including various obstacles that may pose a danger to the flying device, such as the ground, houses, birds, or small drones. The pixel points not occupied by the point cloud usually correspond to the area without obstacles. Therefore, the non-empty mask can be used to determine whether there are obstacles in a point cloud image, that is, in a certain area.
[0044] Step 140: Extract features from the coding matrix to obtain a target feature map.
[0045] In an embodiment of the present application, the encoder part includes four spatio-temporal coding sub-modules. The coding matrix is input into the spatio-temporal coding sub-modules. Each time a spatio-temporal coding sub-module processes it, a corresponding coded feature map is generated. Each coded feature map includes five data dimensions B, S, C, H, and W, where B represents the batch, S represents the number of consecutive frames sequence, the initial value of S is 5, that is, the current frame plus the number of the past 4 frames, C represents the number of channels channel, H represents the image height of the initial coded feature map, and W represents the image width of the initial coded feature map.
[0046] Each time a spatio-temporal coding sub-module is passed through, the data dimensions of the corresponding coded feature map undergo a transformation, changing from the original (B, S, C, H, W) to (B, S, 2C, H / 2, W / 2). That is, each time a spatio-temporal coding sub-module is passed through, the number of channels of the coded feature map generated by this spatio-temporal coding sub-module becomes twice that of the previous coded feature map, and the length and width of each coded feature map become half of those of the previous coded feature map. After being processed by multiple spatio-temporal coding sub-modules, the mapping size of the pixel points on each coded feature map in the original coded feature image gradually becomes larger, and features with a larger receptive field can be learned, and the generated coded feature map contains more global and higher semantic-level features.
[0047] In some embodiments, each spatio-temporal coding sub-module contains 2 sub-convolutional modules. Each sub-convolutional module is composed of a 2D convolutional layer, a normalization layer, and a Relu (Rectified Linear Unit) layer. The 2D sub-convolutional layer can learn the spatial features of the image, that is, the motion features. A 3D convolutional layer is added after the first two spatio-temporal coding sub-modules. Through 3D convolution and 3D pooling, the time can be modeled and the temporal features of the image can be retained. Therefore, the finally generated coded feature map includes motion features and temporal features.
[0048] The coded feature maps generated by each spatio-temporal coding sub-module are input into the decoder module. In some embodiments, the decoder module includes 4 spatio-temporal decoding sub-modules. Each coded feature map is first subjected to an interpolation process before being respectively input into the corresponding spatio-temporal decoding sub-module. Each time an interpolation process is passed through, the length H and width W of the corresponding coded feature map become twice those of the adjacent past decoded feature map. The coded feature map after the interpolation process is sent into the corresponding spatio-temporal decoding sub-module, that is, decoded to generate the corresponding decoded feature map.
[0049] The decoding process can effectively fuse the features of different receptive fields, thereby enhancing the detection accuracy.
[0050] At this point, after being processed by the decoder module and the encoder module, the target feature map is generated, and the size of the target feature map is the same as the H and W dimensions of the initial encoded feature map. Usually, for efficiency considerations, the dimension of the number of channels C is set to an integer multiple of 2. In this solution, the initial value of C can be 2. Finally, the C dimension of the target feature map becomes 32, S changes from 5 to 1, and the target feature map is generated.
[0051] Step 150: Determine multiple prediction masks according to the target feature map.
[0052] In an embodiment of the present application, the multiple prediction masks may include a category prediction mask, a state estimation mask, and a motion prediction mask.
[0053] In some embodiments, the category prediction mask can classify obstacles separated by the non-empty mask. The separated obstacles usually include background, such as the ground, runway, and open space, etc., as well as other obstacles. The category prediction mask can separate the background from other obstacles.
[0054] In some embodiments, the state prediction mask can further distinguish the other obstacles separated by the type prediction mask to detect moving targets and stationary targets among the other obstacles.
[0055] In some embodiments, the motion prediction mask predicts a two-dimensional motion vector (x, y) of each pixel, and when the predicted vector value exceeds a preset threshold, the point cloud data corresponding to the pixel is considered to be a moving target point cloud.
[0056] Step 160: Determine a detection mask according to the multiple prediction masks and the non-empty masks.
[0057] In an embodiment of the present application, the non-empty masks and multiple prediction masks obtained in step 130 and step 150 are all 0-1 matrices. The above non-empty masks and multiple prediction masks are correspondingly multiplied, and the final mask obtained is the detection mask, that is, the mask of the motion point cloud.
[0058] In some embodiments, the detection mask may be determined according to the following formula:
[0059] M=M no_empty ⊙M cat ⊙M state ⊙M motion
[0060] Among them, M represents the matrix corresponding to the detection mask, M no_empty 、M cat 、M state and M motionrespectively represent the matrices corresponding to the non-empty mask, the category prediction mask, the state estimation mask, and the motion prediction mask. ⊙ represents the element-wise multiplication of two matrices to obtain the Hadamard product.
[0061] Step 170: Extract the targets from the current frame point cloud data based on the detection mask to determine the moving targets in the current frame point cloud data.
[0062] In the embodiment of the present application, the Hadamard product obtained in step 160 is a 0-1 matrix, that is, the matrix corresponding to the detection mask. Extract the positions where the elements in this matrix are 1, which are the positions of the moving targets in the corresponding point cloud data, so as to realize the detection of the moving targets.
[0063] Please refer to Figure 2 , Figure 2 which shows the flowchart of step 110 in the embodiment of the present application. Step 110 further includes steps 111 to 113.
[0064] In the embodiment of the present application, the fused point cloud data is determined according to the following formula:
[0065]
[0066] where P is the fused point cloud after coordinate alignment. represents the inverse matrix of the current point cloud frame transformation matrix, that is, the integrated navigation data of the current frame. T -i represents the transformation matrix of the i-th past frame, that is, the corresponding integrated navigation data of the i-th past frame. Among them, the transformation matrix is the matrix obtained by the rotation transformation of the point cloud coordinate system of the i-th past frame relative to the reference coordinate system. Similarly, represents the matrix obtained by the rotation transformation of the current frame relative to the reference coordinate system, P -i represents the point cloud data of the i-th past frame.
[0067] Step 111: Determine the first matrix according to the integrated navigation data corresponding to the current frame point cloud data.
[0068] In the embodiment of the present application, is the first matrix obtained by combining the current frame point cloud data with the integrated navigation data.
[0069] Step 112: Respectively determine the second matrix corresponding to each frame of point cloud data according to each frame of point cloud data in the current frame point cloud data, the historical multi-frame point cloud data, and the corresponding integrated navigation data, and in combination with the first matrix.
[0070] In the embodiment of the present application, when N takes 4, the operation efficiency and detection accuracy can be balanced. When i takes 0, 1, 2, 3, or 4 respectively, 5 different That is, it is the second matrix obtained by combining the point cloud data of each of the current frame and the historical 4 frames with the corresponding integrated navigation data. The second matrix includes five.
[0071] Step 113: Determine the fused point cloud data based on all the second matrices.
[0072] In the embodiments of the present application, according to Equation 1, it can be known that the fused point cloud data can be obtained by summing all the second matrices.
[0073] Since the movement of the object between each frame of the point cloud is mainly affected by two factors: one is the movement of the object itself (such as the flight movement of a bird or a small unmanned aerial vehicle), and the other is the movement from one frame to another of each frame of the point cloud due to the flight movement of the aircraft itself. Therefore, in this solution, all the past N frames of point cloud data are transformed into the coordinate system of the current frame and superimposed, that is, all the second matrices are summed, self-motion compensation is completed, the influence of the aircraft's own movement is eliminated, the misjudgment caused by the aircraft's own movement in the detected moving target is avoided, and the detection accuracy of the moving target is improved.
[0074] Please refer to Figure 3 , Figure 3 which shows the flow diagram of Step 120 in the embodiments of the present application. Step 120 further includes Step 121 to Step 122.
[0075] Step 121: Determine the index of each point in the fused point cloud data in the voxel space respectively according to the coordinates of each point in the fused point cloud data, the preset voxelization size, and the point cloud boundary value.
[0076] Please refer to Figure 4 , Figure 4 which shows a schematic diagram of a voxel space Space in the embodiments of the present application. Set the voxel space Space in the preset reference coordinate system. The voxel space Space is evenly divided into a plurality of voxels V. The size of the preset voxel space Space is related to the point cloud boundary value, and the voxelization size of each voxel V is preset as cell.
[0077] In some embodiments, the point cloud boundary values are X max , X min , Y max , Y min , Z max and Z min respectively. Then the side lengths of the voxel space are X max -X min , Y max -Y min and Z max -Z min .
[0078] Experiments have shown that if the voxelization size is too large, the size of the feature map is small, and the features of the feature map cannot be fully learned in the subsequent feature extraction part. If the voxelization size is too small, the size of the feature map is large, which will affect the running speed of the subsequent network inference.
[0079] Preferably, when the side length of each voxel V is set to 0.2 m, better operating efficiency can be achieved.
[0080] Each point in the fused point cloud data has a corresponding true coordinate in the reference coordinate system of the point cloud in the current frame. Map each point in the fused point cloud data to the reference coordinate system where the preset voxel space Space is located, and calculate the index of each point in the fused point cloud data in the voxel space Space, that is, the voxelization process is completed.
[0081] In some embodiments, there may be one or more points under each voxel V. According to the index of each voxel V, the true coordinate data corresponding to the point cloud under the coordinates of each voxel V can be found.
[0082] Step 122: Determine the encoding matrix according to the index of each point in the fused point cloud data in the voxel space Space.
[0083] In the implementation of the present application, the voxelized point cloud data is encoded with 0-1. Traverse the index of each voxel V in three dimensions of the reference coordinate system of the voxel space Space respectively. Mark the voxel V with points in the index as 1, and mark the voxel V without points in the index as 0, so that a three-dimensional 0-1 encoding matrix can be obtained.
[0084] The coordinates of each point in the fused point cloud data include the first coordinate in the first direction, the second coordinate in the second direction, and the third coordinate in the third direction; the preset voxelization size includes the first size corresponding to the first direction, the second size corresponding to the second direction, and the third size corresponding to the third direction; the point cloud boundary values include the first boundary value corresponding to the first direction, the second boundary value corresponding to the second direction, and the third boundary value corresponding to the third direction; the index of each point in the voxel space Space includes the first sub-index corresponding to the first direction, the second sub-index corresponding to the second direction, and the third sub-index corresponding to the third direction.
[0085] Reference Figure 4 , in the embodiment of the present application, the reference coordinate system where the current frame where the fused point cloud data is located includes three mutually perpendicular directions of X, Y, and Z, corresponding to the first direction, the second direction, and the third direction respectively. The true coordinates corresponding to the coordinates of each point in the fused point cloud data in this reference coordinate system are (x, y, z), and x, y, and z correspond to the first coordinate, the second coordinate, and the third coordinate respectively.
[0086] In some embodiments, the preset voxelization size cell includes data in three dimensions, i.e., cell = {c x , c y , c z}, where c x , c y , c z correspond to the first dimension, the second dimension, and the third dimension respectively.
[0087] In some embodiments, the point cloud boundary values include a first boundary value X max and X min in the X direction, a second boundary value Y max and Y min in the Y direction, and a third boundary value Z max and Z min .
[0088] In some embodiments, the index of each point in the voxel space Space includes a first sub-index i corresponding to the X direction, a second sub-index j corresponding to the Y direction, and a third sub-index k corresponding to the Z direction, thereby obtaining the index (i, j, k) of each point in the fused point cloud data within the voxel space Space.
[0089] Please refer to Figure 5 , Figure 5 which shows a schematic flowchart of step 121. Step 121 further includes steps 1211 to 1213.
[0090] Step 1211: Determine the first sub-index corresponding to each point in the fused point cloud data according to the first coordinate, the first boundary value, and the first dimension of each point in the fused point cloud data.
[0091] Step 1212: Determine the second sub-index corresponding to each point in the fused point cloud data according to the second coordinate, the second boundary value, and the second dimension of each point in the fused point cloud data.
[0092] Step 1213: Determine the third sub-index corresponding to each point in the fused point cloud data according to the third coordinate, the third boundary value, and the third dimension of each point in the fused point cloud data.
[0093] In the embodiments of the present application, the first sub-index, the second sub-index, and the third sub-index are calculated respectively according to the following formulas:
[0094] i = floor((x - X min ) / c x )
[0095] j = floor((y - Ymin ) / c y )
[0096] k = floor((z - Z min ) / c z )
[0097] where X min , Y min , and Z min respectively represent the minimum values of the point cloud boundaries in the X, Y, and Z directions. Among them, the way to calculate i is to calculate the distance between the first coordinate x of each point cloud in the X direction and the minimum value of the point cloud boundary, including how many voxels V (implemented by rounding down, i.e., the floor function). Similarly, the values of j and k can be calculated respectively to obtain the index (i, j, k) of each point in the fused point cloud data.
[0098] Step 140 further includes the following steps.
[0099] (1) Extract features from the encoding matrix based on a pre-trained spatio-temporal network model to obtain a target feature map.
[0100] (2) Generate multiple prediction masks according to the target feature map.
[0101] In the embodiments of the present application, features are extracted from the encoding matrix based on a pre-trained spatio-temporal network model to obtain a target feature map; and multiple prediction masks are generated according to the target feature map. The target feature map contains motion features and temporal features.
[0102] Using a pre-trained spatio-temporal network model can avoid the problem of manually adjusting parameters in traditional algorithms and enhance the operation efficiency of the algorithm.
[0103] In the embodiments of the present application, the pre-trained spatio-temporal network model includes an encoder, a decoder, and a multi-task processor. Among them, the encoder includes multiple sequentially connected spatio-temporal encoding sub-modules. The encoding matrix is sequentially input into multiple spatio-temporal encoding sub-modules, and motion features and temporal features are extracted through multiple spatio-temporal encoding sub-modules in sequence to obtain multiple encoded feature maps; the multiple encoded feature maps correspond to the multiple spatio-temporal encoding sub-modules one by one, and each encoded feature map is obtained by the corresponding spatio-temporal encoding sub-module.
[0104] The step of extracting features from the encoding matrix based on a pre-trained spatio-temporal network model to obtain a target feature map includes the following steps.
[0105] (1) Input the encoding matrix into the encoder to obtain an encoded feature map.
[0106] Please refer to Figure 6 , Figure 6The figure shows a schematic diagram of a spatio-temporal network in an embodiment of the present application. In the encoder part, it includes four spatio-temporal encoding sub-modules A1, A2, A3, and A4, and each spatio-temporal encoding sub-module contains multiple sub-convolution modules (not shown in the figure). Preferably, it is set to contain 2 sub-convolution modules. Experiments show that when each spatio-temporal encoding sub-module contains 2 sub-convolution modules, the running efficiency of the spatio-temporal network model is higher.
[0107] Each sub-convolution module is composed of a 2D convolution layer, a normalization layer, and a ReLU layer. The 2D convolution layer is used to process image data. It has two dimensions, width and height. The purpose of the convolution operation is to extract the input motion features. The normalization layer stabilizes the input distribution of the intermediate layer of the model within an appropriate range, speeds up the convergence rate of the model training process, and enhances the anti-interference ability of the model to input changes. The ReLU function is used to remove the negative values in the convolution result and keep the positive values unchanged.
[0108] After being processed by each spatio-temporal encoding sub-module, a corresponding encoded feature map will be generated. Each encoded feature map includes five data dimensions B, S, C, H, and W, where B represents the batch, S represents the number of consecutive frames sequence, the initial value of S is 5, that is, the current frame plus the number of the past 4 frames, C represents the number of channels channel, H represents the picture height of the first encoded feature map, and W represents the picture width of the first encoded feature map.
[0109] After passing through each spatio-temporal encoding sub-module, the data dimensions of the corresponding encoded feature map change once, from the original (B, S, C, H, W) to (B, S, 2C, H / 2, W / 2). That is, after passing through each spatio-temporal encoding sub-module, the number of channels of the encoded feature map generated by this spatio-temporal encoding sub-module becomes twice that of the adjacent past encoded feature map, and the length and width of each encoded feature map become half of those of the adjacent past encoded feature map. After being processed by multiple spatio-temporal encoding sub-modules, the mapping size of the pixel points on each encoded feature map in the original encoded feature image gradually becomes larger, that is, it can learn features with a larger receptive field, and the generated encoded feature map contains more global and higher semantic-level features.
[0110] In some embodiments, a 3D convolution layer can be added after the first two spatio-temporal encoding sub-modules. Experiments show that adding a 3D convolution layer after the first two spatio-temporal encoding sub-modules has a higher prediction accuracy than adding it at other positions. The 3D convolution layer has an additional time dimension compared to the 2D convolution layer, so it can extract temporal features. After the encoded feature map is processed by the 3D convolution layer, its H and W dimensions do not change.
[0111] As Figure 6 shown, Figure 6 The figure shows a schematic diagram of a spatio-temporal network in an embodiment of the present application.
[0112] The encoded matrix generated in step 120 is the initial encoded feature map, and the dimension of the initial encoded feature map is (B, S, C, H, W). The initial encoded feature map is input into the spatio-temporal encoding sub-module A1 to generate an encoded feature map M1, and the dimension of the encoded feature map M1 is (B, S, 2C, W / 2, H / 2).
[0113] The encoded feature map M1 is input into the spatio-temporal encoding sub-module A2 to generate an encoded feature map M2, and the data dimension of the encoded feature map M2 is (B, S, 4C, H / 4, W / 4).
[0114] The encoded feature map M2 is input into the 3D convolution module to extract temporal features, generating an encoded feature map M3, and the data dimension of the encoded feature map M3 is (B, S', 4C, H / 4, W / 4). Among them, the initial value of S can be taken as 5. After extracting the temporal features, S becomes S', and S' = 1.
[0115] The encoded feature map M3 is input into the spatio-temporal encoding sub-module A3 to generate an encoded feature map M4, and the data dimension of the encoded feature map M4 is (B, S', 8C, H / 8, W / 8).
[0116] The encoded feature map M4 is input into the spatio-temporal encoding sub-module A4 to generate an encoded feature map M5, and the data dimension of the encoded feature map M5 is (B, S', 16C, H / 16, W / 16).
[0117] In some embodiments, usually for efficiency considerations, C is set to an integer multiple of 2. Here, C can be taken as 2, so the C dimension of the finally generated encoded feature map M5 is 32.
[0118] (2) Input the encoded feature map into the decoder to obtain the target feature map.
[0119] As Figure 6 shown, in the embodiments provided in the present application, the decoder includes a plurality of sequentially connected spatio-temporal decoding sub-modules; the plurality of spatio-temporal decoding sub-modules correspond one-to-one with the plurality of spatio-temporal encoding sub-modules, and each spatio-temporal decoding sub-module is connected to the corresponding spatio-temporal encoding sub-module.
[0120] This step (2) further includes performing interpolation processing on the plurality of encoded feature maps and respectively inputting them into the corresponding spatio-temporal decoding sub-modules, and sequentially performing feature fusion through the plurality of spatio-temporal decoding sub-modules to obtain the target feature map. After each spatio-temporal decoding sub-module processes the encoded feature map, its H and W dimensions both become 2 times that of the previous one.
[0121] In some embodiments, the height and width of the corresponding encoded feature map become twice that before the interpolation process. During the image magnification process, the pixels also increase accordingly, and the process of increase is where interpolation takes effect. At this time, there will be some points that cannot be directly mapped. For these points, their values can be determined through interpolation. Interpolation automatically selects pixels with better information as the space to increase and fill in the blank pixels, rather than only using adjacent pixels. When magnifying the image, the image will look smoother and cleaner.
[0122] In some embodiments, the interpolation processing method may further include mean / median / mode interpolation method, fixed value processing method, sliding average window method, etc., and the present application does not limit this.
[0123] In the embodiments of the present application, the decoder module includes four spatio-temporal decoding sub-modules B1, B2, B3, and B4. Each encoded feature map is first subjected to an interpolation process before being input into the corresponding spatio-temporal decoding sub-module. After one interpolation process, the H and W dimensions of the encoded feature map become twice that before the interpolation process.
[0124] In some embodiments, the encoded feature map M5 is input into the spatio-temporal decoding sub-module B1. After decoding processing, the decoded feature map M1' is obtained. The dimension of the decoded feature map M1' is (B, S', 16C, H / 8, W / 8). The encoded feature map M5 undergoes one interpolation process to obtain the feature map N1. The height and width of the feature map N1 are H / 8 and W / 8 respectively.
[0125] After fusing the feature map N1 and the decoded feature map M1', it is input into the spatio-temporal decoding sub-module B2. After decoding processing, the decoded feature map M2' is obtained. The dimension of the decoded feature map M2' is (B, S', 16C, H / 4, W / 4). The encoded feature map M4 undergoes one interpolation process to obtain the feature map N2. The height and width of the feature map N2 are H / 4 and W / 4 respectively.
[0126] After fusing the feature map N2 and the decoded feature map M2', it is input into the spatio-temporal decoding sub-module B3. After decoding processing, the decoded feature map M3' is obtained. The dimension of the decoded feature map M3' is (B, S', 16C, H / 2, W / 2). The encoded feature map M3 undergoes one interpolation process to obtain the feature map N3. The height and width of the feature map N3 are H / 2 and W / 2 respectively.
[0127] After fusing the feature map N3 and the decoded feature map M3', it is input into the spatio-temporal decoding sub-module B4. After decoding processing, the decoded feature map M4' is obtained. The dimension of the decoded feature map M4' is (B, S', 16C, H, W). The encoded feature map M1 undergoes one interpolation process to obtain the feature map N4. The height and width of the feature map N2 are H and W respectively.
[0128] After fusing the feature map N4 and the decoded feature map M4', the target feature map M5' is obtained, and the dimension of the target feature map M5' is (B, S', 16C, H, W).
[0129] Step 150 determines multiple prediction masks according to the target feature map, including: inputting the target feature map into a multi-task processor to obtain multiple prediction masks.
[0130] In the embodiment provided by this application, the multi-task processor includes: a category prediction sub-module, a state estimation sub-module, and a motion prediction sub-module; the multiple prediction masks include: a category prediction mask, a state estimation mask, and a motion prediction mask.
[0131] The step of inputting the target feature map into the multi-task processor to obtain multiple prediction masks includes: inputting the target feature map into the category prediction sub-module to obtain the category prediction mask.
[0132] In some embodiments, the category prediction head further classifies the two-dimensional image containing obstacles after classifying the non-empty masks. Through the inference of the category prediction head, the classification of obstacles in multiple categories such as background, pedestrians, and vehicles can be obtained. The pixels where the background is located are marked as 0 (the background is usually static, such as the ground, an empty field, etc.), and the pixels where the obstacles in the remaining categories are located are marked as 1. Thus, the binary classification between the background and the obstacles in the remaining categories is completed, and the category prediction mask in the form of a 0-1 matrix is generated.
[0133] Input the target feature map into the state estimation sub-module to obtain the state estimation mask.
[0134] In some embodiments, the state estimation sub-module further performs a binary classification of static or moving on the obstacles in the remaining categories classified by the category prediction mask. Through the inference of the state estimation sub-module, the classification of moving obstacles and static obstacles can be obtained. The pixels where the moving obstacles are located are marked as 1, and the pixels where the static obstacles are located are marked as 0. Thus, the binary classification between moving and static is completed, and the state estimation mask in the form of a 0-1 matrix is generated.
[0135] Input the target feature map into the motion prediction sub-module to obtain the motion prediction mask.
[0136] In some embodiments, the motion prediction sub-module performs vector prediction on each pixel value in the two-dimensional image of the current frame point cloud data. The two-dimensional motion vector of each pixel is (x, y). If the vector exceeds a preset threshold, the pixel is determined to be a pixel where a potentially moving obstacle is located, and the pixel where the potentially moving obstacle is located is marked as 1. If the vector is less than or equal to the preset threshold, the pixel is determined to be a pixel where a potentially stationary obstacle is located, and the pixel where the potentially stationary obstacle is located is marked as 0. Thus, the prediction of whether a pixel may be in a moving state is completed. Experiments show that when the preset threshold is 0.4, the detection effect is better.
[0137] In addition, the motion target detection method provided in the embodiments of the present application has high real-time performance during actual operation. For example, when running on a GPU, it only takes 30 ms, which can meet the real-time requirements.
[0138] In the embodiments provided by the present application, the state prediction head completes the binary classification task of moving and static, but there may be some misjudgment situations, such as a stationary obstacle being judged as a moving obstacle, and the motion prediction head may also have some misjudgment situations, such as a stationary point cloud predicting a motion vector. Therefore, combining the state estimation mask and the motion prediction mask, using them simultaneously and compensating each other, can reduce the probability of misjudgment of the model and enhance the detection accuracy of moving targets.
[0139] Please refer to Figure 7 , Figure 7 FIG. shows a motion target detection device provided by an embodiment of the present application. The motion target detection device includes: a point cloud fusion module 210, a voxelization module 220, a non-empty mask acquisition module 230, a feature extraction module 240, a prediction mask acquisition module 250, a detection mask acquisition module 260, and a motion target detection module 270.
[0140] Among them, the point cloud fusion module 210 is used to perform coordinate system alignment processing on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, the historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data.
[0141] The voxelization module 220 is used to perform voxelization encoding on the fused point cloud data to obtain an encoding matrix.
[0142] The non-empty mask acquisition module 230 is used to determine a non-empty mask according to the encoding matrix.
[0143] The feature extraction module 240 is used to extract features from the encoding matrix to obtain a target feature map.
[0144] The prediction mask acquisition module 250 is used to determine multiple prediction masks according to the target feature map.
[0145] The detection mask acquisition module 260 is configured to determine a detection mask based on multiple prediction masks and non-empty masks.
[0146] The moving object detection module 270 is configured to perform object extraction on the current frame point cloud data based on the detection mask to determine the moving objects in the current frame point cloud data.
[0147] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0148] It should be noted that each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. The relevant parts can refer to the corresponding descriptions in the method embodiments. For any processing method described in the method embodiments, it can be implemented by the corresponding processing module in the device embodiments, and will not be described in detail in the device embodiments.
[0149] In addition, in each embodiment of the present application, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0150] Please refer to Figure 8 , based on the above moving object detection method, an embodiment of the present application further provides another flying device 300 including a processor 310 that can execute the foregoing moving object detection method. The flying device 300 further includes one or more processors 310, a memory 320, and one or more application programs. Among them, the memory 320 stores a program that can execute the content in the foregoing embodiments, and the processor 310 can execute the program stored in the memory 320.
[0151] Among them, the processor 310 may include one or more cores for processing data and a message matrix unit. The processor 310 connects various parts within the entire flight device through various interfaces and lines, and executes various functions of the flight device 300 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 320, and by calling data stored in the memory 320. Optionally, the processor 310 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 310 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 310 and may be implemented separately through a communication chip.
[0152] The memory 320 may include random access memory (RAM) and may also include read-only memory. The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as the function of obtaining fused point clouds, the function of obtaining prediction masks, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal (such as fused point cloud data, encoding matrices, encoded feature maps, and decoded feature maps, etc.).
[0153] Please refer to Figure 9 , which shows a structural block diagram of a computer-readable storage medium 400 provided by an embodiment of the present application. Program code 410 is stored in the computer-readable medium 400, and the program code 410 can be called by the processor 310 to execute the moving target detection method described in the above method embodiments.
[0154] The computer-readable storage medium 400 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has a storage space for program code 410 that executes any of the method steps in the above method. These program codes can be read from or written into one or more computer program products. The program code 410 can be compressed in a suitable form.
[0155] In summary, a method, device, flying device, and storage medium for detecting a moving target provided in this application perform coordinate system alignment processing based on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, the historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data; perform voxelization encoding on the fused point cloud data to obtain an encoding matrix; determine a non-empty mask according to the encoding matrix; extract features from the encoding matrix to obtain a target feature map; determine multiple prediction masks according to the target feature map; determine a detection mask according to the multiple prediction masks and the non-empty mask; and perform target extraction on the current frame point cloud data based on the detection mask to determine the moving target in the current frame point cloud data. This method directly processes the point cloud data end-to-end, avoiding the error accumulation caused by the linear process of traditional algorithms, and fusing the historical multi-frame point cloud data, which can make full use of the temporal information of the point cloud data and enhance the real-time performance of the algorithm; this method takes advantage of the voxel representation to detect unlabeled objects; in addition, the combined use of multiple masks reduces the probability of false detection of the algorithm and can accurately determine the moving target.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A moving target detection method, characterized in that, Applied to a flight device, the method includes: Performing coordinate system alignment processing based on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data; Performing voxelization encoding on the fused point cloud data to obtain an encoding matrix; Determining a non-empty mask based on the encoding matrix; Performing feature extraction on the encoding matrix to obtain a target feature map; Determining multiple prediction masks based on the target feature map; Determining a detection mask based on the multiple prediction masks and the non-empty mask; Performing target extraction on the current frame point cloud data based on the detection mask to determine moving targets in the current frame point cloud data.
2. The motion target detection method according to claim 1, characterized in that The performing coordinate system alignment processing based on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data includes: Determining a first matrix based on the combined navigation data corresponding to the current frame point cloud data; Respectively based on the current frame point cloud data, each frame of point cloud data in the historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data, and in combination with the first matrix, determining a second matrix corresponding to each frame of point cloud data; Determining the fused point cloud data based on all the second matrices.
3. The motion target detection method according to claim 1, characterized in that, The performing voxelization encoding on the fused point cloud data to obtain an encoding matrix includes: Respectively determining the index of each point in the fused point cloud data in the voxel space S according to the coordinates of each point in the fused point cloud data, a preset voxelization size, and point cloud boundary values; Determining the encoding matrix according to the index of each point in the fused point cloud data in the voxel space S.
4. The motion target detection method according to claim 3, wherein The coordinates of each point in the fused point cloud data include a first coordinate in a first direction, a second coordinate in a second direction, and a third coordinate in a third direction; The preset voxelization size includes a first size corresponding to the first direction, a second size corresponding to the second direction, and a third size corresponding to the third direction; The point cloud boundary values include a first boundary value corresponding to the first direction, a second boundary value corresponding to the second direction, and a third boundary value corresponding to the third direction; The index of each point in the voxel space S includes a first sub-index corresponding to the first direction, a second sub-index corresponding to the second direction, and a third sub-index corresponding to the third direction; The determining the index of each point in the fused point cloud data in the voxel space S according to the coordinates of each point in the fused point cloud data, a preset voxelization size, and point cloud boundary values includes: Determining the first sub-index corresponding to each point in the fused point cloud data according to the first coordinate of each point in the fused point cloud data, the first boundary value, and the first size; Determining the second sub-index corresponding to each point in the fused point cloud data according to the second coordinate of each point in the fused point cloud data, the second boundary value, and the second size; Determine the third sub-index corresponding to each point in the fused point cloud data according to the third coordinate of each point in the fused point cloud data, the third boundary value, and the third dimension.
5. The motion target detection method according to claim 1, characterized in that, Perform feature extraction processing on the encoding matrix to obtain a target feature map; Determine multiple prediction masks according to the target feature map, including: Perform feature extraction on the encoding matrix based on a pre-trained spatio-temporal network model to obtain a target feature map; And generate multiple prediction masks according to the target feature map.
6. The method for detecting a moving object according to claim 5, wherein The pre-trained spatio-temporal network model includes an encoder, a decoder, and a multi-task processor; Perform feature extraction on the encoding matrix based on a pre-trained spatio-temporal network model to obtain a target feature map; And generate multiple prediction masks according to the target feature map, including: Input the encoding matrix into the encoder to obtain an encoded feature map; Input the encoded feature map into the decoder to obtain the target feature map; Input the target feature map into the multi-task processor to obtain the multiple prediction masks.
7. The motion target detection method according to claim 6, wherein The encoder includes multiple sequentially connected spatio-temporal encoding sub-modules; the step of inputting the encoding matrix into the encoder to obtain an encoded feature map includes: Sequentially input the encoding matrix into the multiple spatio-temporal encoding sub-modules, and sequentially perform extraction of motion features and temporal features through the multiple spatio-temporal encoding sub-modules to obtain multiple encoded feature maps; the multiple encoded feature maps correspond one-to-one with the multiple spatio-temporal encoding sub-modules, and each encoded feature map is obtained by the corresponding spatio-temporal encoding sub-module.
8. The motion target detection method according to claim 7, wherein The decoder includes multiple sequentially connected spatio-temporal decoding sub-modules; the multiple spatio-temporal decoding sub-modules correspond one-to-one with the multiple spatio-temporal encoding sub-modules, and each spatio-temporal decoding sub-module is connected to the corresponding spatio-temporal encoding sub-module; The step of inputting the encoded feature map into the decoder to obtain the target feature map includes: Perform interpolation processing on the multiple encoded feature maps, and respectively input them into the corresponding spatio-temporal decoding sub-modules, and sequentially perform feature fusion through the multiple spatio-temporal decoding sub-modules to obtain a target feature map.
9. The motion target detection method according to claim 6, wherein, The multi-task processor includes: a category prediction sub-module, a state estimation sub-module, and a motion prediction sub-module; the multiple prediction masks include: a category prediction mask, a state estimation mask, and a motion prediction mask; The step of inputting the target feature map into the multi-task processor to obtain the multiple prediction masks includes: Input the target feature map into the category prediction sub-module to obtain the category prediction mask; Input the target feature map into the state estimation sub-module to obtain the state estimation mask; Input the target feature map into the motion prediction sub-module to obtain the motion prediction mask.
10. A moving target detection device, characterized in that, Including: A point cloud fusion module, configured to perform coordinate system alignment processing on the current frame point cloud data, the combined navigation data corresponding to the current frame point cloud data, the historical multi-frame point cloud data, and the combined navigation data corresponding to the historical multi-frame point cloud data to obtain fused point cloud data; A voxelization module, configured to perform voxelization encoding on the fused point cloud data to obtain an encoding matrix; A non-empty mask acquisition module, configured to determine a non-empty mask according to the encoding matrix; A feature extraction module, configured to perform feature extraction on the encoding matrix to obtain a target feature map; A predicted mask acquisition module, configured to determine multiple predicted masks according to the target feature map; A detection mask acquisition module, configured to determine a detection mask according to the multiple predicted masks and the non-empty mask; A moving target detection module, configured to perform target extraction on the current frame point cloud data based on the detection mask to determine a moving target in the current frame point cloud data.
11. A flying device, characterized in that, Comprising: One or more processors; A memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the moving target detection method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the moving target detection method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Target detection method and device, intelligent driving method, equipment and storage medium
CN112101066A
Three-dimensional target detection and neural network training method, device and equipment
CN112444784A