Method and device for detecting and identifying non-line-of-sight moving targets by a planar array time-of-flight camera
The method projects non-visual domain point cloud data from face-array time-of-flight cameras onto a two-dimensional plane for detection using Yolov5, integrating depth information to enhance moving target recognition, addressing the limitations of existing face-array cameras in non-visual domain imaging.
Patent Information
- Application Number
- CN202510191923.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The non-sight imaging system based on the surface array time-of-flight camera in the prior art lacks the ability to detect and recognize motion objects, and is difficult to adapt to non-sight imaging three-dimensional point cloud data, which limits its application range.
By projecting the three-dimensional point cloud data of non-sighted moving targets acquired by the plane array time-of-flight camera onto the two-dimensional plane, the trained Yolov5 model is used for object detection, and combining the depth-direction distribution characteristics of the three-dimensional point cloud, the three-dimensional detection and recognition results of the moving targets are determined.
It realizes intelligent detection and identification of non-sighted motion targets, broadens the non-sighted application capabilities of the surface array time-of-flight camera, retains the speed advantages of the Yolov5 model and completes the three-dimensional position information.
Smart Images

Figure CN119693630B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of non-line-of-sight three-dimensional point cloud data processing, and particularly relates to a method and device for detecting and identifying non-line-of-sight moving targets by a matrix time-of-flight camera. Background Art
[0002] Non-line-of-sight imaging technology can achieve "seeing through walls" and is of great significance in fields such as medical imaging and disaster rescue. Time-of-Flight (ToF) technology can determine the distance of a target object by measuring the time it takes for an optical signal to travel from a transmitter to the target object and then reflect back to a receiver. Compared with streak tube time-of-flight cameras and single-photon avalanche diode time-of-flight cameras, matrix time-of-flight cameras with price advantages have attracted the attention of many scholars. However, existing research mainly focuses on the development of non-line-of-sight imaging systems based on matrix time-of-flight cameras and the imaging processing, detection, and identification of static non-line-of-sight targets, with insufficient ability to detect and identify moving targets.
[0003] In recent years, with the rapid development of deep learning technology, three-dimensional target recognition networks have emerged continuously. Existing projection-based target object detection methods usually first project a single view or multiple views of a three-dimensional point cloud to generate a two-dimensional grid, and then process the two-dimensional grid to achieve the detection and identification of target objects. This method achieves a good balance between time complexity and detection performance at the cost of losing spatial information. However, the above existing projection-based detection methods are difficult to be directly adapted to non-line-of-sight imaging three-dimensional point cloud data obtained by a matrix time-of-flight camera, which limits the application scope of non-line-of-sight imaging technology. Summary of the Invention
[0004] The present disclosure aims to at least solve one of the problems existing in the prior art, and provides a method and device for detecting and identifying non-line-of-sight moving targets by a matrix time-of-flight camera.
[0005] In one aspect of the present disclosure, a method for detecting and identifying non-line-of-sight moving targets by a matrix time-of-flight camera is provided. The method includes:
[0006] Obtaining imaging three-dimensional point cloud data of a non-line-of-sight moving target based on a matrix time-of-flight camera;
[0007] Projecting the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain a corresponding two-dimensional projection map;
[0008] Using a trained Yolov5 model to perform target detection on the two-dimensional projection map to obtain a two-dimensional detection and identification result corresponding to the moving target, where the two-dimensional detection and identification result includes the position coordinates, target category, and confidence level of a two-dimensional detection box;
[0009] Determine the depthwise position of the moving target according to the depthwise distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection frame;
[0010] Combine the two-dimensional detection and recognition result and the depthwise position to determine the three-dimensional detection and recognition result of the moving target.
[0011] Optionally, the projecting the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain a corresponding two-dimensional projection map includes:
[0012] Use the two-dimensional plane perpendicular to the depth direction of the area array time-of-flight camera as the projection plane;
[0013] Project the imaging three-dimensional point cloud data onto the projection plane for pixel coordinate system representation to obtain the two-dimensional projection map.
[0014] Optionally, the projecting the imaging three-dimensional point cloud data onto the projection plane for pixel coordinate system representation to obtain the two-dimensional projection map includes:
[0015] Represent the n th three-dimensional point cloud in the imaging three-dimensional point cloud data as , where respectively represent the lengthwise n -dimensional, widthwise x -dimensional, and depthwise y -dimensional corresponding coordinates of the z th three-dimensional point cloud in the camera coordinate system, ; N represents the total number of three-dimensional point clouds included in the imaging three-dimensional point cloud data;
[0016] Respectively determine the minimum coordinate x and the maximum coordinate of each three-dimensional point cloud in the lengthwise -dimensional of the imaging three-dimensional point cloud data, and the minimum coordinate y and the maximum coordinate of each three-dimensional point cloud in the widthwise -dimensional of the imaging three-dimensional point cloud data;
[0017] Use the quadrilateral defined by , , , as the image range boundary, and draw a scatter plot based on the x - y plane of the imaging three-dimensional point cloud data in the camera coordinate system to obtain the two-dimensional projection map.
[0018] Optionally, determining the depth - direction position of the moving target according to the depth - direction distribution characteristics of the three - dimensional point cloud within the range corresponding to the two - dimensional detection box includes:
[0019] Convert the position coordinates of the two - dimensional detection box to the camera coordinate system, and crop out all target point clouds within the range corresponding to the two - dimensional detection box in the camera coordinate system from the imaged three - dimensional point cloud data as the two - dimensional detection box point cloud;
[0020] Use a histogram to statistically analyze the depth - direction distribution of the two - dimensional detection box point cloud to obtain the depth - direction distribution histogram corresponding to the two - dimensional detection box point cloud;
[0021] According to the following formula, determine the minimum depth value and the maximum depth value of the target three - dimensional box corresponding to the moving target:
[0022] ;
[0023] ;
[0024] where and are both threshold coefficients, represents the minimum depth value corresponding to the two - dimensional detection box in the camera coordinate system, represents the maximum depth value corresponding to the two - dimensional detection box in the camera coordinate system, represents the depth value corresponding to the peak of the depth - direction distribution histogram.
[0025] Optionally, the conversion of the position coordinates of the two - dimensional detection box to the camera coordinate system includes:
[0026] Record the position coordinates of the two diagonals of the two - dimensional detection box in the pixel coordinate system as 、 respectively, where are the coordinates of the two diagonals of the two - dimensional detection box in the pixel coordinate system x d - dimensional coordinates and , are the coordinates of the two diagonals of the two - dimensional detection box in the pixel coordinate system y d - dimensional coordinates and ;
[0027] According to the following formula, convert the position coordinates of the two diagonals of the two - dimensional detection box in the pixel coordinate system 、 to the position coordinates in the camera coordinate system respectively:
[0028] ;
[0029] ;
[0030] Wherein, respectively represent the coordinates of the two-dimensional detection frame in the pixel coordinate system x dimensional coordinates corresponding coordinates in the camera coordinate system x dimensional coordinates, respectively represent the coordinates of the two-dimensional detection frame in the pixel coordinate system y dimensional coordinates corresponding coordinates in the camera coordinate system y dimensional coordinates, respectively represent the number of pixels in the x dimension, y dimension of the two-dimensional projection map in the pixel coordinate system.
[0031] Optionally, determining the three-dimensional detection and recognition result of the moving target by combining the two-dimensional detection and recognition result and the depth direction position includes:
[0032] Combining the position coordinates of the two-dimensional detection frame in the camera coordinate system with the depth direction position of the moving target to obtain the corresponding three-dimensional target frame position;
[0033] Marking the three-dimensional target frame, target category, and confidence level in the imaging three-dimensional point cloud data in the camera coordinate system according to the three-dimensional target frame position and the target category and confidence level included in the two-dimensional detection and recognition result as the three-dimensional detection and recognition result of the moving target.
[0034] Optionally, the Yolov5 model is trained according to the following steps:
[0035] Based on an area array time-of-flight camera, collect a variety of different pose data and speed data of the non-line-of-sight moving target, and perform imaging processing on the variety of different pose data and the speed data to obtain the non-line-of-sight point cloud data of the moving target in each different pose;
[0036] Label the non-line-of-sight point cloud data in each different pose respectively, determine the position information and category information of the moving target corresponding to the non-line-of-sight point cloud data in each different pose respectively, and construct a labeled data set;
[0037] Use the labeled data set to train and validate the Yolov5 model.
[0038] Another aspect of the present disclosure provides a non-line-of-sight moving target detection and recognition device for an area array time-of-flight camera, and the device includes:
[0039] An acquisition module, configured to acquire imaging three-dimensional point cloud data of a non-line-of-sight moving target based on an area array time-of-flight camera;
[0040] A projection module, configured to project the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain a corresponding two-dimensional projection map;
[0041] A two-dimensional detection module, configured to perform target detection on the two-dimensional projection map by using a trained Yolov5 model to obtain a two-dimensional detection and recognition result corresponding to the moving target, where the two-dimensional detection and recognition result includes the position coordinates, target category, and confidence of a two-dimensional detection box;
[0042] A depth direction determination module, configured to determine the depth direction position of the moving target according to the depth direction distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box;
[0043] A result determination module, configured to determine a three-dimensional detection and recognition result of the moving target by combining the two-dimensional detection and recognition result and the depth direction position.
[0044] Optionally, the apparatus further includes:
[0045] A training module, configured to train the Yolov5 model according to the following steps:
[0046] Based on an area array time-of-flight camera, collect a variety of different pose data and speed data of a non-line-of-sight moving target, and perform imaging processing on the variety of different pose data and the speed data to obtain non-line-of-sight point cloud data of the moving target in each different pose;
[0047] Label the non-line-of-sight point cloud data in each different pose respectively, determine the position information and category information of the moving target corresponding to the non-line-of-sight point cloud data in each different pose respectively, and construct a labeled data set;
[0048] Use the labeled data set to train and verify the Yolov5 model.
[0049] Another aspect of the present disclosure provides an electronic device, including:
[0050] At least one processor; and a memory communicatively connected to the at least one processor; wherein,
[0051] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the non-line-of-sight moving target detection and recognition method of the area array time-of-flight camera described above.
[0052] Another aspect of the present disclosure provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the non-line-of-sight moving target detection and recognition method of the area array time-of-flight camera described above.
[0053] Compared with the prior art, relying on the area array time-of-flight camera with cost advantages, the present disclosure not only retains the speed advantage of single-stage detection of the Yolov5 model, but also complements the three-dimensional position of the moving target, realizing the intelligent detection and recognition of the non-line-of-sight moving target imaging three-dimensional point cloud data obtained based on the area array time-of-flight camera, effectively broadening the non-line-of-sight application ability of the area array time-of-flight camera. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.
[0055] Figure 1 It is a flowchart of a non-line-of-sight moving target detection and recognition method for an area array time-of-flight camera provided by an embodiment of the present disclosure;
[0056] Figure 2 It is a flowchart of a non-line-of-sight moving target detection and recognition method for an area array time-of-flight camera provided by another embodiment of the present disclosure;
[0057] Figure 3 It is a schematic diagram of a non-line-of-sight imaging optical path provided by another embodiment of the present disclosure;
[0058] Figure 4 It is a schematic diagram of the two-dimensional target detection result of a moving target person provided by another embodiment of the present disclosure;
[0059] Figure 5 It is a schematic diagram of the two-dimensional target detection result of a moving target unmanned vehicle provided by another embodiment of the present disclosure;
[0060] Figure 6 It is a schematic diagram of the position of a two-dimensional detection frame in the pixel coordinate system provided by another embodiment of the present disclosure;
[0061] Figure 7 It is a schematic diagram of the three-dimensional detection and recognition result of a moving target person provided by another embodiment of the present disclosure;
[0062] Figure 8 It is a schematic diagram of the three-dimensional detection and recognition result of a moving target unmanned vehicle provided by another embodiment of the present disclosure;
[0063] Figure 9Schematic diagram of a structure of a non-line-of-sight moving target detection and recognition device for a matrix time-of-flight camera provided in another embodiment of the present disclosure;
[0064] Figure 10 Schematic diagram of a structure of an electronic device provided in another embodiment of the present disclosure. Specific embodiments
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present disclosure, many technical details are provided to help readers better understand the present disclosure. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present disclosure can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation to the specific implementation manner of the present disclosure. Each embodiment can be combined and cross-referenced with each other on the premise of no contradiction.
[0066] One embodiment of the present disclosure relates to a method for detecting and recognizing non-line-of-sight moving targets by a matrix time-of-flight camera, and its process is as Figure 1 shown, including steps S110 to S150. The following will describe steps S110 to S150 in detail with reference to Figure 2 to.
[0067] Step S110, acquiring imaging three-dimensional point cloud data of a non-line-of-sight moving target based on a matrix time-of-flight camera.
[0068] Specifically, a matrix time-of-flight camera is an imaging device that can acquire depth information of a scene in real time. It calculates the distance between an object and the camera by measuring the time difference of the optical signal from the camera to the object surface and then reflected back to the camera, thereby generating a three-dimensional depth image and converting it into three-dimensional point cloud data.
[0069] A non-line-of-sight moving target refers to a moving target that is not within the straight-line field of view of the matrix time-of-flight camera but can be detected by the matrix time-of-flight camera through other means such as reflection. In other words, there is an obstacle between the non-line-of-sight moving target and the matrix time-of-flight camera, resulting in the signal between the moving target and the matrix time-of-flight camera not being able to propagate along a straight line but only through means such as reflection.
[0070] For example, step S110 can use a matrix time-of-flight camera with the model number OPT8241 to acquire imaging three-dimensional point cloud data of a moving target based on the Figure 3 shown non-line-of-sight imaging optical path. In Figure 3In the non-line-of-sight imaging optical path shown, the optical signal emitted by the light source cannot reach the moving target through direct propagation due to the occlusion of the occlusion surface, but can only reach the moving target through reflection on the reflection wall. At the same time, the optical signal reaching the moving target cannot be reflected back to the area array time-of-flight camera through direct propagation due to the occlusion of the occlusion surface, but can only reach the area array time-of-flight camera through reflection on the reflection wall. At this time, the area array time-of-flight camera can receive the optical signal reflected by the moving target to complete the acquisition of non-line-of-sight moving target reflection data. By filtering, enhancing, and other processing of the non-line-of-sight moving target reflection data, the imaging three-dimensional point cloud data of the moving target can be obtained.
[0071] Step S120: Project the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain the corresponding two-dimensional projection map.
[0072] Exemplarily, step S120 may include: using the two-dimensional plane perpendicular to the depth direction of the area array time-of-flight camera as the projection plane; projecting the imaging three-dimensional point cloud data onto the projection plane for pixel coordinate system representation to obtain the two-dimensional projection map.
[0073] Specifically, step S120 may use the two-dimensional plane composed of the length direction and the width direction perpendicular to the depth direction of the area array time-of-flight camera as the projection plane to retain the length direction information and width direction information of the imaging three-dimensional point cloud data.
[0074] Exemplarily, in step S120, projecting the imaging three-dimensional point cloud data onto the projection plane for pixel coordinate system representation to obtain the two-dimensional projection map includes: representing the n th three-dimensional point cloud in the imaging three-dimensional point cloud data as , where respectively represent the length direction n -dimensional, width direction x -dimensional, and depth direction y -dimensional corresponding coordinates of the z th three-dimensional point cloud in the camera coordinate system, . N represents the total number of three-dimensional point clouds included in the imaging three-dimensional point cloud data. Determine the minimum coordinate x and the maximum coordinate of each three-dimensional point cloud in the length direction of the imaging three-dimensional point cloud data, as well as the minimum coordinate y and the maximum coordinate of each three-dimensional point cloud in the width direction of the imaging three-dimensional point cloud data. Use the quadrilateral defined by , , , as the image range boundary, based on thex - y Draw a scatter plot on a plane to obtain a two-dimensional projection map.
[0075] Specifically, the dataset corresponding to the imaged three-dimensional point cloud data P can be expressed as . , That is, the boundary point coordinates of the x -dimensional corresponding to the imaged three-dimensional point cloud data, , That is, the boundary point coordinates of the y -dimensional corresponding to the imaged three-dimensional point cloud data. The points , , , The defined quadrilateral is the boundary of the image range.
[0076] For example, when drawing a scatter plot, relevant drawing functions can be called. Based on the imaged three-dimensional point cloud data and the x - y coordinates in the camera coordinate system of the image range boundary, draw a scatter plot on a plane to obtain the corresponding two-dimensional projection map, and store the two-dimensional projection map in a specified folder for use in subsequent steps.
[0077] Step S130: Use the trained Yolov5 model to perform object detection on the two-dimensional projection map to obtain the two-dimensional detection and recognition results corresponding to the moving objects. The two-dimensional detection and recognition results include the position coordinates of the two-dimensional detection box, the object category, and the confidence level.
[0078] Specifically, as an object detection model, the Yolov5 model has high detection accuracy. The network structure of the Yolov5 model usually includes an input layer, a backbone network, a neck network, and a head network. Among them, the input layer is used to receive image data. The backbone network is used to extract the features of the image. The neck network is used to further process the features extracted by the backbone network to enhance the feature expression ability. The head network is used to generate the final detection results, including the position of the bounding box corresponding to the object, the object category, and the confidence level. Therefore, in step S130, the two-dimensional projection map can be input into the Yolov5 model, and two-dimensional detection and recognition can be performed based on the Yolov5 model, so as to obtain the two-dimensional detection and recognition results including the position coordinates of the two-dimensional detection box of the moving object, the object category, and the confidence level output by the Yolov5 model.
[0079] For example, step S130 can first set the hyperparameters of the Yolov5 model, then read the two-dimensional projection map stored in the specified folder, load the Yolov5 model, input the read two-dimensional projection map into the Yolov5 model for object detection, and obtain the position coordinates, object category, and confidence level of the two-dimensional detection box of the moving object output by the Yolov5 model. When the moving object is a person, the object detection result of the Yolov5 model, that is, the two-dimensional detection and recognition result, is as shown in Figure 4 . Figure 4 In, the object category is a person, corresponding to a person with outstretched arms and being detected and recognized with a confidence level of 0.69. When the moving object is an autonomous vehicle, the object detection result of the Yolov5 model, that is, the two-dimensional detection and recognition result, is as shown in Figure 5 . Figure 5 In, the object type is an autonomous vehicle, corresponding to the prominent outline and wheel features of the autonomous vehicle and being detected and recognized with a confidence level of 0.85.
[0080] Step S140 determines the depthwise position of the moving object according to the depthwise distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box.
[0081] Exemplarily, step S140 includes: converting the position coordinates of the two-dimensional detection box to the camera coordinate system, cropping out all the target point clouds within the range corresponding to the two-dimensional detection box in the camera coordinate system from the imaging three-dimensional point cloud data as the two-dimensional detection box point cloud. Using a histogram to statistically analyze the depthwise distribution of the two-dimensional detection box point cloud to obtain the depthwise distribution histogram corresponding to the two-dimensional detection box point cloud. According to the following formula, determine the minimum depth value and the maximum depth value of the target three-dimensional box corresponding to the moving object:
[0082] ;
[0083] .
[0084] Among them, and are both threshold coefficients, represents the minimum depth value corresponding to the two-dimensional detection box in the camera coordinate system, represents the maximum depth value corresponding to the two-dimensional detection box in the camera coordinate system, represents the depth value corresponding to the peak of the depthwise distribution histogram.
[0085] Step S140 mainly converts the position coordinates of the two-dimensional detection box in the pixel coordinate system to the camera coordinate system through coordinate transformation, so as to crop out all the target point clouds within the two-dimensional detection box according to the position coordinates of the two-dimensional detection box in the camera coordinate system, and determine the depthwise position of the moving object based on these target point clouds.
[0086] By determining the depth-direction position of the moving target, the depth-direction information of the moving target in the imaged three-dimensional point cloud data can be retained, avoiding the loss of the spatial information of the moving target, and enabling more accurate identification of the moving target.
[0087] Exemplarily, in step S140, converting the position coordinates of the two-dimensional detection box into the camera coordinate system includes: respectively denoting the position coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system as 、 , where are respectively the x -dimensional coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system and , are respectively the y -dimensional coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system and . According to the following formula, the position coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system 、 are respectively converted into the position coordinates in the camera coordinate system:
[0088] ;
[0089] .
[0090] Among them, respectively represent the x -dimensional coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system corresponding to the x -dimensional coordinates in the camera coordinate system, respectively represent the y -dimensional coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system corresponding to the y -dimensional coordinates in the camera coordinate system, respectively represent the number of pixels in the x -dimension and y -dimension of the two-dimensional projection map in the pixel coordinate system.
[0091] For example, as Figure 6 shown, in the pixel coordinate system, let the x -dimension point to the right, the y -dimension point downwards, 、 be the upper left corner and the lower right corner of the two-dimensional detection box in the pixel coordinate system respectively, then the two-dimensional detection box is within the rectangular range defined by the origin O and the point .
[0092] Step S150, combining the two-dimensional detection recognition result and the depth-direction position to determine the three-dimensional detection recognition result of the moving target.
[0093] Exemplarily, step S150 may include: combining the position coordinates of the two-dimensional detection box in the camera coordinate system with the depthwise position of the moving target to obtain the corresponding three-dimensional target box position. According to the three-dimensional target box position and the target category and confidence included in the two-dimensional detection and recognition result, mark the corresponding three-dimensional detection box, category, and confidence in the imaging three-dimensional point cloud data in the camera coordinate system as the three-dimensional detection and recognition result of the moving target.
[0094] For example, in and when Figure 4 the three-dimensional detection and recognition result of the moving target person in Figure 7 is as shown in Figure 5 the three-dimensional detection and recognition result of the moving target unmanned vehicle in Figure 8 is as shown. Among them, the X-axis, Y-axis, and Z-axis represent the dimensions of length, width, and height respectively, and the units of the X-axis, Y-axis, and Z-axis are all m (meters).
[0095] The non-line-of-sight moving target detection and recognition method for the area array time-of-flight camera provided by the embodiments of the present disclosure, compared with the prior art, relies on the area array time-of-flight camera with cost advantages, retains both the speed advantage of single-stage detection of the Yolov5 model and complements the three-dimensional position of the moving target, realizes the intelligent detection and recognition of the non-line-of-sight moving target imaging three-dimensional point cloud data obtained based on the area array time-of-flight camera, and effectively broadens the non-line-of-sight application ability of the area array time-of-flight camera.
[0096] Exemplarily, the Yolov5 model is trained according to the following steps: Based on the area array time-of-flight camera, collect various different pose data and speed data of the non-line-of-sight moving target respectively, and perform imaging processing on the various different pose data and speed data to obtain the non-line-of-sight point cloud data of the moving target in each different pose. Annotate the non-line-of-sight point cloud data in each different pose respectively, determine the position information and category information of the moving target corresponding to the non-line-of-sight point cloud data in each different pose, and construct an annotated data set. Use the annotated data set to train and validate the Yolov5 model.
[0097] For example, in this embodiment, it is possible to first build as Figure 3The optical path shown is used to collect data of moving target personnel in different postures such as swinging arms, walking, prone, etc. and at different speeds, or to collect data of moving target unmanned vehicles at 360° rotation, forward, backward, partial body and full body and at different speeds, based on which. At the same time, non-line-of-sight background data without moving targets is collected, and through further processing such as filtering and enhancement, the corresponding non-line-of-sight point cloud data is obtained. Then, Labelimg is used to label the non-line-of-sight point cloud data to determine the position information and category information of the moving targets corresponding to the non-line-of-sight point cloud data in different postures, and a labeled data set is constructed. Finally, the labeled data set is used to train and validate the Yolov5 model, so that the Yolov5 model can learn the features of specific targets such as two-dimensional position information and target category, and improve the detection performance.
[0098] Another embodiment of the present disclosure relates to a non-line-of-sight moving target detection and recognition device for an area array time-of-flight camera, as Figure 9 shown, including an acquisition module 910, a projection module 920, a two-dimensional detection module 930, a depth direction determination module 940, and a result determination module 950.
[0099] The acquisition module 910 is configured to acquire imaging three-dimensional point cloud data of non-line-of-sight moving targets based on an area array time-of-flight camera.
[0100] The projection module 920 is configured to project the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain a corresponding two-dimensional projection map.
[0101] The two-dimensional detection module 930 is configured to perform target detection on the two-dimensional projection map by using a trained Yolov5 model to obtain a two-dimensional detection and recognition result corresponding to the moving target. The two-dimensional detection and recognition result includes the position coordinates of the two-dimensional detection box, the target category, and the confidence level.
[0102] The depth direction determination module 940 is configured to determine the depth direction position of the moving target according to the depth direction distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box.
[0103] The result determination module 950 is configured to combine the two-dimensional detection and recognition result and the depth direction position to determine the three-dimensional detection and recognition result of the moving target.
[0104] Exemplarily, the non-line-of-sight moving target detection and recognition device for an area array time-of-flight camera further includes a training module.
[0105] The training module is used to train the Yolov5 model according to the following steps: Based on the area array time-of-flight camera, collect a variety of different pose data and speed data of the non-line-of-sight moving target respectively, and perform imaging processing on the variety of different pose data and speed data to obtain the non-line-of-sight point cloud data of the moving target in each different pose; label the non-line-of-sight point cloud data in each different pose respectively, determine the position information and category information of the moving target corresponding to the non-line-of-sight point cloud data in each different pose, and construct a labeled data set; use the labeled data set to train and validate the Yolov5 model.
[0106] For the specific implementation method of the non-line-of-sight moving target detection and recognition device provided by the embodiment of the present disclosure, reference may be made to the non-line-of-sight moving target detection and recognition method provided by the embodiment of the present disclosure, which will not be elaborated here.
[0107] The non-line-of-sight moving target detection and recognition device provided by the embodiment of the present disclosure, compared with the prior art, relies on the area array time-of-flight camera with cost advantages, retains the speed advantage of the single-stage detection of the Yolov5 model, and complements the three-dimensional position of the moving target, realizing the intelligent detection and recognition of the non-line-of-sight moving target imaging three-dimensional point cloud data obtained based on the area array time-of-flight camera, effectively broadening the non-line-of-sight application ability of the area array time-of-flight camera.
[0108] Another embodiment of the present disclosure relates to an electronic device, as Figure 10 shown, including:
[0109] At least one processor 1001; and, a memory 1002 communicatively connected to the at least one processor 1001; wherein, the memory 1002 stores instructions executable by the at least one processor 1001, and the instructions are executed by the at least one processor 1001 to enable the at least one processor 1001 to execute the non-line-of-sight moving target detection and recognition method described in the above embodiment.
[0110] Among them, the memory and the processor are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0111] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when performing operations.
[0112] Another embodiment of the present disclosure relates to a computer-readable storage medium storing a computer program, which when executed by a processor implements the method for non-line-of-sight moving target detection and recognition of the area array time-of-flight camera described in the above embodiment.
[0113] That is, those skilled in the art can understand that all or part of the steps in implementing the methods described in the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.
[0114] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present disclosure, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present disclosure.
Claims
1. A method for detecting and recognizing non-line-of-sight moving targets in a matrix time-of-flight camera, characterized in that, The method includes: Obtaining imaging three-dimensional point cloud data of a non-line-of-sight moving target based on an area array time-of-flight camera; Projecting the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain a corresponding two-dimensional projection map; Using the trained Yolov5 model to perform target detection on the two-dimensional projection map to obtain a two-dimensional detection and recognition result corresponding to the moving target, where the two-dimensional detection and recognition result includes the position coordinates, target category, and confidence level of the two-dimensional detection box; Determining the depth-wise position of the moving target according to the depth-wise distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box; Combining the two-dimensional detection and recognition result and the depth-wise position to determine the three-dimensional detection and recognition result of the moving target; The determining the depth-wise position of the moving target according to the depth-wise distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box includes: Converting the position coordinates of the two-dimensional detection box to the camera coordinate system, and cropping out all target point clouds within the range corresponding to the two-dimensional detection box in the camera coordinate system from the imaging three-dimensional point cloud data as the two-dimensional detection box point cloud; Using a histogram to statistically analyze the depth-wise distribution of the two-dimensional detection box point cloud to obtain a depth-wise distribution histogram corresponding to the two-dimensional detection box point cloud; Determine the minimum depth value and the maximum depth value of the target three-dimensional bounding box corresponding to the moving target according to the following formula and the maximum depth value : ; ; Among them, and are both threshold coefficients, represents the minimum depth value corresponding to the two-dimensional detection frame in the camera coordinate system, represents the maximum depth value corresponding to the two-dimensional detection frame in the camera coordinate system, represents the depth value corresponding to the peak of the depth-wise distribution histogram.
2. The method according to claim 1, wherein The projecting the imaging three-dimensional point cloud data onto a two-dimensional plane to obtain a corresponding two-dimensional projection map includes: Taking a two-dimensional plane perpendicular to the depth direction of the area array time-of-flight camera as the projection plane; Projecting the imaging three-dimensional point cloud data onto the projection plane for pixel coordinate system representation to obtain the two-dimensional projection map.
3. The method according to claim 2, wherein The projecting the imaging three-dimensional point cloud data onto the projection plane for pixel coordinate system representation to obtain the two-dimensional projection map includes: Represent the n th 3D point cloud in the imaged 3D point cloud data as , where respectively represent the corresponding coordinates of the n th 3D point cloud in the length x -dimensional, width y -dimensional, and depth z -dimensional directions in the camera coordinate system, ; N represents the total number of 3D point clouds included in the imaged 3D point cloud data; Determine the minimum coordinate and the maximum coordinate of each 3D point cloud in the length dimension of the imaging 3D point cloud data respectively x dimension and the maximum coordinate , and the minimum coordinate and the maximum coordinate of each 3D point cloud in the width dimension of the imaging 3D point cloud data y dimension and the maximum coordinate ; Take , , , the defined quadrilateral as the boundary of the image range, and draw a scatter plot on the x - y plane based on the three-dimensional imaging point cloud data in the camera coordinate system to obtain the two-dimensional projection map.
4. The method according to claim 3, wherein The converting the position coordinates of the two-dimensional detection box to the camera coordinate system includes: Denote the position coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system as , , where are the coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system in x dimensions respectively, and . are the coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system in y dimensions respectively, and ; According to the following formula, convert the position coordinates of the two diagonals of the two-dimensional detection box in the pixel coordinate system , into the position coordinates in the camera coordinate system respectively: ; ; Among them, respectively represent the x -dimensional coordinates of the two-dimensional detection frame in the pixel coordinate system and the corresponding x -dimensional coordinates in the camera coordinate system, respectively represent the y -dimensional coordinates of the two-dimensional detection frame in the pixel coordinate system and the corresponding y -dimensional coordinates in the camera coordinate system, respectively represent the number of pixels in the x -dimension and y -dimension of the two-dimensional projection map in the pixel coordinate system.
5. The method according to claim 4, wherein The combining the two-dimensional detection and recognition result and the depth-wise position to determine the three-dimensional detection and recognition result of the moving target includes: Combining the position coordinates of the two-dimensional detection box in the camera coordinate system with the depth-wise position of the moving target to obtain a corresponding three-dimensional target box position; Marking the three-dimensional target box, target category, and confidence level in the imaging three-dimensional point cloud data in the camera coordinate system according to the three-dimensional target box position and the target category and confidence level included in the two-dimensional detection and recognition result as the three-dimensional detection and recognition result of the moving target.
6. The method according to any one of claims 1 to 5, characterized in that, The Yolov5 model is trained according to the following steps: Based on an area array time-of-flight camera, respectively collecting various different pose data and speed data of a non-line-of-sight moving target, and performing imaging processing on the various different pose data and the speed data to obtain non-line-of-sight point cloud data of the moving target in each different pose; Respectively annotating the non-line-of-sight point cloud data in each different pose to determine the position information and category information of the moving target corresponding to the non-line-of-sight point cloud data in each different pose, and constructing an annotated data set; Using the annotated data set to train and validate the Yolov5 model.
7. An apparatus for detecting and recognizing non-line-of-sight moving targets in a matrix time-of-flight camera, characterized in that, The device includes: An acquisition module, configured to acquire three-dimensional point cloud data of a non-line-of-sight moving target based on an area array time-of-flight camera; A projection module, configured to project the three-dimensional point cloud data of the imaging onto a two-dimensional plane to obtain a corresponding two-dimensional projection map; A two-dimensional detection module, configured to perform target detection on the two-dimensional projection map by using a trained Yolov5 model to obtain a two-dimensional detection and recognition result corresponding to the moving target, where the two-dimensional detection and recognition result includes position coordinates, target category, and confidence of a two-dimensional detection box; A depth direction determination module, configured to determine the depth direction position of the moving target according to the depth direction distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box; A result determination module, configured to determine a three-dimensional detection and recognition result of the moving target by combining the two-dimensional detection and recognition result and the depth direction position; The determining the depth direction position of the moving target according to the depth direction distribution characteristics of the three-dimensional point cloud within the range corresponding to the two-dimensional detection box includes: Converting the position coordinates of the two-dimensional detection box into the camera coordinate system, and cropping out all target point clouds within the range corresponding to the two-dimensional detection box in the camera coordinate system from the three-dimensional point cloud data of the imaging as two-dimensional detection box point clouds; Using a histogram to statistically analyze the depth direction distribution of the two-dimensional detection box point clouds to obtain a depth direction distribution histogram corresponding to the two-dimensional detection box point clouds; Determine the minimum depth value and the maximum depth value of the target three-dimensional bounding box corresponding to the moving target according to the following formula and the maximum depth value : ; ; Among them, and are both threshold coefficients, represents the minimum depth value corresponding to the two-dimensional detection frame in the camera coordinate system, represents the maximum depth value corresponding to the two-dimensional detection frame in the camera coordinate system, represents the depth value corresponding to the peak of the depth-wise distribution histogram.
8. An electronic device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the non-line-of-sight moving target detection and recognition method for an area array time-of-flight camera according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the non-line-of-sight moving target detection and recognition method for an area array time-of-flight camera according to any one of claims 1 to 6.