Depth data processing method and related device
By intercepting the data of the second length from the depth data of the first length and mapping it into point cloud data, the problem of low depth data processing efficiency caused by insufficient computing power of the robot is solved, and efficient depth data processing and robot control task generation under limited computing power are realized.
Patent Information
- Application Number
- CN202411973396.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-02
AI Technical Summary
When robots process high-precision depth data, the control task generation efficiency is low due to insufficient computing power.
By intercepting the second depth data of the second length from the first depth data of the first length, and mapping the multiple second depth data into point cloud data within a single mapping instruction cycle, a robot control task is generated based on the point cloud data.
When the computing power of the robot terminal is limited, it efficiently processes deep data and quickly generates corresponding robot control tasks, improving data processing efficiency.
Smart Images

Figure CN119919480A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and more particularly to a method for processing depth data and related devices. Background Art
[0002] A robot is a mechanical device or system that can simulate human or other biological behaviors to a certain extent and perform specific tasks through programming control. Robots have certain perception, decision-making, control and execution capabilities, and can be used in many fields such as industry, daily life, medical treatment, education, scientific research, etc.
[0003] Generally speaking, robots need to perform tasks in complex and dynamic environments, so they need to respond quickly to external changes to ensure the safety and effectiveness of operations. Based on this, real-time performance is one of the necessary factors to be considered in robot design. In order to improve the real-time performance of data processing, low-power microcontrollers or custom chips with weaker computing power are usually configured in robots to ensure that the robots can process the data collected by sensors in real time and generate corresponding robot control tasks accordingly.
[0004] However, when a robot configured in the above manner processes high-precision depth data, there is a problem of low efficiency in generating robot control tasks due to insufficient computing power and inability to efficiently process the depth data. Summary of the invention
[0005] In view of this, multiple embodiments of the present application are dedicated to providing a method and related device for processing depth data, which can efficiently process depth data to quickly generate corresponding robot control tasks when the computing power of the robot terminal is limited.
[0006] One embodiment of the present application provides a method for processing depth data, which is applied to a robot terminal, comprising: intercepting second depth data of a second length from first depth data of a first length; wherein the second length is smaller than the first length; within a single mapping instruction cycle, mapping a plurality of second depth data to point cloud data; wherein the total length of the plurality of second depth data is smaller than or equal to the number of register bits;
[0007] A robot control task for driving the robot terminal to perform a specific operation is generated based on the point cloud data.
[0008] Optionally, the first length is equal to the number of bits of the register.
[0009] Optionally, the second depth data is stored as a one-dimensional array, and the data index corresponding to the one-dimensional array is used to indicate the positions of the second depth data in the one-dimensional array respectively; within a single mapping instruction cycle, the step of mapping multiple second depth data into point cloud data includes: within a single mapping instruction cycle, reading multiple unprocessed second depth data from the one-dimensional array according to the data index and writing them into a register, performing vectorized point cloud mapping calculation on the multiple second depth data in the register, and obtaining point cloud data corresponding to the multiple second depth data respectively.
[0010] Optionally, the first depth data is depth data for a target object in a depth map; the step of generating a robot control task for driving the robot terminal to perform a specific operation based on the point cloud data comprises: selecting a target point set from the point cloud data based on a point cloud distance between the point cloud data; generating a candidate pose set corresponding to the target point set based on a pose generation model; generating a target pose corresponding to the target object according to the candidate pose set; generating a robot control task for driving the robot terminal to perform a specific operation based on the target pose; wherein the robot control task comprises at least one of the following: a grasping task, an obstacle avoidance task, a clearing task, and a cutting task; and the specific operation is an operation performed on the target object.
[0011] Optionally, the step of selecting a target point set from the point cloud data based on the point cloud distance between the point cloud data includes: in a first distance calculation instruction cycle, calculating the spatial distance between the x-dimensional data and the y-dimensional data between the point cloud data to obtain the x-dimensional spatial distance and the y-dimensional spatial distance; in a second distance calculation instruction cycle, calculating the spatial distance between the z-dimensional data and the null value between the point cloud data to obtain the z-dimensional spatial distance; generating a corresponding point cloud distance based on the x-dimensional spatial distance, the y-dimensional spatial distance, and the z-dimensional spatial distance, and selecting a target point set from the point cloud data based on the point cloud distance.
[0012] Optionally, the operation process of the posture generation model includes function calculation, and when vectorized function calculation is performed in the register, the function calculation result is determined based on a preset function value lookup table.
[0013] Optionally, the step of generating a target pose corresponding to the target object based on the candidate pose set includes: vectorizing the calculation of pose distances between multiple candidate poses in the candidate pose set and other candidate poses in the candidate pose set within a single pose distance calculation instruction cycle; and determining the target pose of the target object based on the pose distance.
[0014] One embodiment of the present application provides a depth data processing device, including: a depth data truncation module, used to obtain second depth data of a second length from first depth data of a first length; wherein the second length is smaller than the first length; a vectorized mapping module, used to map multiple second depth data into point cloud data within a single mapping instruction cycle; wherein the total length of the multiple second depth data is less than or equal to the number of register bits; and a task generation module, used to generate a robot control task for driving the robot terminal to perform a specific operation based on the point cloud data.
[0015] One embodiment of the present application provides a computer device, which includes a memory and a processor, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the depth data processing method as described above.
[0016] One embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, the depth data processing method as described above can be implemented.
[0017] One embodiment of the present application provides a computer program product, which is used to implement the depth data processing method as described above.
[0018] In multiple embodiments provided in the present application, first depth data of a first length is truncated into second depth data of a relatively shorter second length, so that the total length of multiple second depth data is less than or equal to the number of register bits. This ensures that multiple second depth data are simultaneously mapped to point cloud data within a single mapping instruction cycle, and a robot control task is generated based on the point cloud data to drive the robot terminal to perform a specific operation, so as to efficiently process the depth data and quickly generate the corresponding robot control task when the computing power of the robot terminal is limited. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic diagram of the interaction between a CPU in a robot and a GPU applied in a method for processing depth data provided in one embodiment of the present application.
[0020] Figure 2 A flowchart of a method for processing depth data provided in accordance with an embodiment of the present application.
[0021] Figure 3 A schematic diagram of an application scenario provided for one embodiment of the present application.
[0022] Figure 4A schematic diagram of a module of a depth data processing device provided in accordance with an embodiment of the present application.
[0023] Figure 5 A schematic diagram of a computer device provided for one embodiment of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.
[0025] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0026] See also Figure 1 . In multiple embodiments provided in the present application, the embodiments of the present application provide a method for processing depth data that can be applied to a depth data processing device, and the depth data processing device can be applied in a robot terminal. The robot terminal includes at least a CPU for preprocessing and managing data collected by the sensor, a GPU for processing data transmitted by the CPU, and a sensor. The sensor includes at least a depth camera for collecting depth maps. The CPU is used to read the depth map collected by the depth camera and copy the depth data in the depth map to the GPU after preprocessing. The GPU is used to truncate the depth data to obtain shorter depth data, and simultaneously map multiple shorter depth data into point cloud data within a single mapping instruction cycle. The CPU is also used to generate a robot control task based on the point cloud data for driving the robot terminal to perform specific operations.
[0027] In some embodiments, the CPU can also be used to read the depth map captured by the depth camera, preprocess the depth data in the depth map, truncate the preprocessed depth data to obtain shorter depth data, and simultaneously map multiple shorter depth data into point cloud data within a single mapping instruction cycle.
[0028] See also Figure 2. An embodiment of the present application provides a method for processing depth data. The method for processing depth data can be applied to a depth data processing device, and the method for processing depth data can also be applied to a robot terminal. The depth data processing device can be an electronic device with certain computing power and network access capabilities. Of course, the depth data processing device can also refer to a software program running in an electronic device. In this embodiment, the method for processing depth data can include the following steps.
[0029] Step S110: extracting second depth data of a second length from the first depth data of the first length; wherein the second length is smaller than the first length.
[0030] Step S120: Mapping a plurality of second depth data into point cloud data within a single mapping instruction cycle; wherein the total length of the plurality of second depth data is less than or equal to the number of register bits.
[0031] Step S130: generating a robot control task for driving the robot terminal to perform a specific operation based on the point cloud data.
[0032] In this embodiment, the depth data processing device can obtain a depth map collected by a depth camera set on the robot terminal, wherein each pixel value in the depth map is used to represent the distance from the depth camera to the shooting environment, and the depth map is data in uint16 format, which is a 16-bit unsigned integer, and the data range that can be represented by uint16 format is [0-65535 mm]. The depth map can be used to generate point cloud data, and point cloud data needs to represent a larger dynamic range and a smaller resolution, so the representation accuracy of uint16 cannot directly meet the representation requirements of point cloud data.
[0033] Based on this, the depth data processing device can map each pixel value in the depth map in uint16 format to a floating point range to convert and obtain first depth data of a first length. Since the first depth data is a floating point number, it can have a higher representation accuracy and a wider representation range, which can meet the representation requirements of point cloud data. Among them, the first depth data can be single-precision floating point type (FP32) data, and the first length represents the number of bits occupied by the first depth data in the register, for example, 32 bits. Specifically, FP32 is a 32-bit floating point format, including a 1-bit sign bit (S) for indicating the positive and negative values, an 8-bit exponent bit (E) for indicating the position of the floating decimal point, and a 23-bit mantissa bit (M) for representing the actual number.
[0034] However, limited by the size and power consumption of the robot terminal, the hardware configuration of the depth data processing device is relatively low. Due to limited computing power, the depth data processing device can intercept the second depth data of the second length from the first depth data of the first length. When processing the second depth data with less space, multiple second depth data can be processed within one instruction cycle, which alleviates the problem of low data processing efficiency caused by limited computing power. Among them, the second depth data can be half-precision floating point type (FP16) data, the second length is less than the first length, and the second length represents the number of bits occupied by the second depth data in the register, for example, 16 bits. Specifically, FP16 is a 16-bit floating point format, including a 1-bit sign bit (S) for indicating the positive and negative values, a 5-bit exponent bit (E) for indicating the position of the floating decimal point, and a 10-bit mantissa bit (M) for representing the actual number.
[0035] Specifically, when the first depth data is FP32 data and the second depth data is FP16 data, the second depth data of the second length is obtained from the first depth data of the first length in the following manner: based on the expression FP32=(-1) S ×2 (E―127) ×(1+M / 2 23 ) extract the sign bit S of the first depth data of the first length FP32 、Exponent E FP32 and the last digit M FP32 ; S FP32 The value of S is determined FP16 The value of the exponent E FP32 Substitute into the expression: exponent E FP16 =E FP32 ―127+15, to calculate E FP16 ; Change the last digit to M FP32 The first 10 digits in the FP16 The value of the combined sign bit S FP16 、Exponent E FP16 and the last digit M FP16 , and obtain the second depth data of the second length. In the above manner, each first depth data can be truncated into the second depth data. In addition, when the first depth data exceeds the representation range of FP16, the first depth data is stored in the memory, and after the second depth data is processed, the first depth data is written into the register for point cloud mapping calculation.
[0036] Furthermore, the depth data processing device can process the second depth data based on a single instruction multiple data (SIMD) architecture. SIMD is a parallel computing architecture that allows the same operation to be performed on multiple data at the same time. Under the SIMD architecture, one instruction can simultaneously operate on multiple second depth data in a register, so that the second depth data can be efficiently processed using limited high-precision computing power. In a single mapping instruction cycle, the depth data processing device can vectorize and map multiple second depth data into point cloud data; wherein the total length of the multiple second depth data is less than or equal to the number of register bits, and vectorization refers to the use of hardware capabilities to perform the same calculation on multiple data simultaneously in one instruction cycle.
[0037] In this embodiment, the point cloud data can represent a point in a three-dimensional space in the form of (x, y, z). The register is used to temporarily store and quickly access data. The number of bits in the register refers to the maximum number of bits of data that can be stored, such as 32, 64. The mapping instruction is used to instruct the processor to map the second depth data to point cloud data. In one mapping instruction cycle, the second depth data that does not exceed the number of bits in the register can be written into the register and the second depth data in the register can be subjected to point cloud mapping calculation. The higher the number of bits in the register, the more second depth data can be stored in one mapping instruction cycle. For example, the number of bits in the register = 32, the number of bits occupied by the second depth data = 16, then, in one mapping instruction cycle, 2 second depth data can be written into the register, and the point cloud mapping calculation for the 2 second depth data can be implemented based on a single mapping instruction. For example, the number of bits in the register = 64, the number of bits occupied by the second depth data = 16, then, in one mapping instruction cycle, 4 second depth data can be written into the register, and the point cloud mapping calculation for the 4 second depth data can be implemented based on a single mapping instruction.
[0038] Furthermore, since the point cloud data can represent the second depth data in a three-dimensional environment, the point cloud data in the three-dimensional environment can help determine the distance between the robot terminal and each position point in the environment and perform some specific robot control tasks accordingly. Therefore, based on the point cloud data, it is also possible to generate a robot control task for driving the robot terminal to perform a specific operation. The robot control task can be any of the following tasks: a task for driving the robot terminal to perform a specific operation, a task for moving toward a destination, a task for deforming into a specified shape, a grasping task, an obstacle avoidance task, a clearing task, and a cutting task. The specific operation corresponding to any of the above robot control tasks may include one or more operations. After the robot terminal performs one or more operations, the corresponding robot control task can be completed. For example, the various operations corresponding to the grasping task include posture adjustment operation, travel operation, robot arm angle adjustment operation, grasping operation, etc.
[0039] The depth data processing method provided in the embodiment of the present application truncates the first depth data of the first length into the second depth data of the relatively shorter second length, so that the total length of the multiple second depth data is less than or equal to the number of register bits. This ensures that within a single mapping instruction cycle, the multiple second depth data are simultaneously mapped to point cloud data, so as to efficiently process the depth data when the computing power of the robot terminal is limited to quickly generate the corresponding robot control tasks.
[0040] In some implementations, the first length is equal to the number of register bits.
[0041] In this embodiment, since the first length (such as 32) is greater than the second length (such as 16), the number of first depth data that can be written to the register in one mapping instruction cycle is less than the number of second depth data that can be written. In one case, the first length (such as 32) is equal to the number of register bits (such as 32). In a mapping instruction cycle, only one first depth data can be written to the register. Due to the high-precision data computing power of the depth data processing device, processing a large amount of first depth data will be slow. Based on this, the first depth data of the first length can be truncated into the second depth data of the second length to increase the amount of data that can be processed in a single mapping instruction cycle.
[0042] In some implementations, the first length (eg, 32) is smaller than the number of register bits (eg, 64).
[0043] In some embodiments, the second depth data is stored as a one-dimensional array, and the data index corresponding to the one-dimensional array is used to indicate the positions of the second depth data in the one-dimensional array respectively; the depth data processing device can read multiple unprocessed second depth data from the one-dimensional array according to the data index and write them into a register within a single mapping instruction cycle, and perform vectorized point cloud mapping calculation on the multiple second depth data in the register to obtain point cloud data corresponding to the multiple second depth data respectively.
[0044] In this embodiment, in order to maximize the parallel computing performance of multi-threads and the combined access of data, the processing device of the depth data can store the second depth data as a matrix in the form of a one-dimensional array, and the data index = row*i+j records the position of the second depth data in the one-dimensional array, row represents the number of elements in each row of the matrix, i represents the i-th row in the matrix, and j represents the i-th column in the matrix. This can simplify the access difficulty of the second depth data and improve the access efficiency of the second depth data.
[0045] Based on this, within each mapping instruction cycle, multiple unprocessed second depth data can be read from the one-dimensional array each time and written into the register in the order specified by the data index, and then vectorized point cloud mapping calculations are performed on the multiple second depth data in the register to obtain multiple point cloud data corresponding to each mapping instruction cycle.
[0046] In addition, data indexes can be stored in a storage space for thread access, which can be a shared memory in the GPU. Shared memory has the characteristics of fast reading and writing. Data indexes are stored in the shared memory for efficient access by threads. After the thread reads the corresponding index from the shared memory, it can read the corresponding second depth data from the storage area accordingly. The storage area here can be any form of storage area for thread access, or it can be a shared memory.
[0047] In this embodiment, the depth data processing device performs vectorized point cloud mapping calculation on multiple second depth data in the register to obtain multiple point cloud data corresponding to each mapping instruction cycle by using the intrinsic parameter matrix corresponding to the depth camera to respectively map the multiple second depth data in the register into point cloud data, and performing the same above operation in each mapping instruction cycle to obtain multiple point cloud data corresponding to each mapping instruction cycle.
[0048] The internal parameter matrix includes the focal length (fx, fy) for representing the depth camera in the horizontal and vertical directions, the principal point coordinates (cx, cy) of the center point of the depth map, the coordinates (u, v) of each pixel in the depth map, and the second depth data is the value d of the corresponding coordinates (u, v). Based on the expression z=d, multiple second depth data can be mapped to point cloud data (x, y, z) in a single mapping instruction cycle, multiple point cloud data (x, y, z) can be obtained in a single mapping instruction cycle, and the processing of all second depth data can be completed by executing multiple mapping instruction cycles. The point cloud data (x, y, z) can be stored in a shared memory for thread access to perform post-processing steps.
[0049] See also Figure 3 In some embodiments, the first depth data is depth data for a target object in a depth map; the depth data processing device may further select a target point set from the point cloud data based on a point cloud distance between the point cloud data; generate a candidate pose set corresponding to the target point set based on a pose generation model; generate a target pose corresponding to the target object based on the candidate pose set; generate a robot control task for driving the robot terminal to perform a specific operation based on the target pose; wherein the robot control task includes at least one of the following: a grasping task, an obstacle avoidance task, a clearing task, and a cutting task; and the specific operation is an operation performed on the target object.
[0050] In this embodiment, in order to efficiently determine the target pose of the target object that needs to be interacted with, so that the robot can efficiently and accurately perform the corresponding robot control task, the depth data processing device can efficiently obtain the corresponding point cloud data based on the vectorized mapping of the second depth data, and determine the target pose of the target object accordingly. The target object can be an object, an animal, etc.
[0051] Specifically, the depth data processing device can obtain a color image (such as an RGB image) collected by a conventional camera, and generate a bounding box of a target object in the color image through a target detection model, and generate a mask image corresponding to the target object based on the target detection model, and then extract the first depth data for the target object from the depth map based on the mask image and the bounding box. Among them, the target detection model (such as YOLO, Faster R-CNN, etc.) is a deep learning model used to identify specific targets (such as people, cars, animals, etc.) and generate corresponding bounding boxes. The bounding box is used to identify the position of the target object in the color image, and the mask image can be a binary image that can mark the area where the target object is located in the color image.
[0052] Furthermore, the depth data processing device can select a representative target point set from the point cloud data as a basis for generating a candidate pose set based on the point cloud distance between the point cloud data, and the point cloud distance can be a Euclidean distance, etc. A candidate pose set corresponding to the target point set can be generated based on the pose generation model, which is a model for inferring the position and direction of the target object in three-dimensional space relative to the depth camera or robot terminal.
[0053] Furthermore, the processing device of the depth data can generate a target pose corresponding to the target object according to the candidate pose set and convert the target pose back to the first length and then copy it to the CPU. The CPU can generate at least one of a grasping task, an obstacle avoidance task, a clearing task, and a cutting task for driving the robot terminal to perform a specific operation based on the target pose, so as to realize the interaction between the robot terminal and the target object. Among them, the robot can complete the grasping of the target object after executing the specific operation indicated by the grasping task, the robot can avoid the target object on the route after executing the specific operation indicated by the obstacle avoidance task, the robot can clear the target object from the environment after executing the specific operation indicated by the clearing task, and the robot can complete the cutting of the target object after executing the specific operation indicated by the cutting task.
[0054] In some embodiments, the depth data processing device can calculate the spatial distance between the x-dimensional data and the y-dimensional data between the point cloud data in a first distance calculation instruction cycle to obtain the x-dimensional spatial distance and the y-dimensional spatial distance; calculate the spatial distance between the z-dimensional data and the null value between the point cloud data in a second distance calculation instruction cycle to obtain the z-dimensional spatial distance; generate the corresponding point cloud distance based on the x-dimensional spatial distance, the y-dimensional spatial distance, and the z-dimensional spatial distance, and select the target point set from the point cloud data based on the point cloud distance.
[0055] In this embodiment, in order to efficiently process point cloud data, the depth data processing device can calculate the point cloud distance based on vectorized calculation. The point cloud data (x, y, z) includes x-dimensional data, y-dimensional data and z-dimensional data. In the related art, the x-dimensional data, y-dimensional data and z-dimensional data are respectively consistent with the number of bits in the register. If the point cloud distance between two point cloud data needs to be calculated, the x-dimensional data can only be written into the register first and the x-dimensional data can be calculated (x 1 ―x 2 ) 2 , then write the y-dimensional data into the register and calculate (y 1 ―y 2 ) 2 , then write the z-dimensional data into the register and calculate (z 1 ―z 2 ) 2 , and then calculate (x 1 ―x 2 ) 2 +(y 1 ―y 2 ) 2 +(z 1 ―z 2 ) 2 , to get the point cloud distance, which has the problem of low efficiency.
[0056] In this embodiment, based on the above processing process, it can be known that the x-dimensional data, the y-dimensional data and the z-dimensional data are all of the second length. The processing device of the depth data can write the x-dimensional data and the y-dimensional data of the point cloud data into the register within the first distance calculation instruction cycle, so as to calculate the spatial distance between the x-dimensional data and the y-dimensional data and the x-dimensional data and the y-dimensional data of other point cloud data within a single distance calculation instruction cycle, and obtain the x-dimensional spatial distance and the y-dimensional spatial distance; then, write the z-dimensional data and the null value of the point cloud data into the register, so as to calculate the spatial distance between the z-dimensional data and the z-dimensional data of other point cloud data within a single distance calculation instruction cycle, and obtain the z-dimensional spatial distance, and then generate the corresponding point cloud distance based on the x-dimensional spatial distance, the y-dimensional spatial distance and the z-dimensional spatial distance. The above-mentioned other point cloud data refers to any point cloud data among the multiple point cloud data. The length of the z-dimensional data is smaller than the register, so when performing vectorized calculation on the register, it is necessary to fill the empty space in the register with the null value. Optionally, in addition to filling in the null value, the x-dimensional data of the next point cloud data to be calculated can also be written into the register.
[0057] In this embodiment, the depth data processing device selects the target point set from the point cloud data in the following manner: construct an empty pre-selected set, randomly select a point cloud data from the point cloud data set as the first point cloud data and write it into the pre-selected set and delete the point cloud data from the point cloud data set, calculate the point cloud distance between each data in the point cloud data set and the first point cloud data, write the point cloud data corresponding to the maximum point cloud distance into the pre-selected set as the second point cloud data and delete the point cloud data from the point cloud data set. Then, calculate the distance between each point cloud data in the pre-selected set and each point cloud data in the point cloud data set to obtain a distance set corresponding to each point cloud data in the point cloud data set, use the minimum value in each distance set as the distance between the corresponding point cloud data and the pre-selected set, then select the maximum value from each minimum value, write the point cloud data corresponding to the maximum value into the pre-selected set as the third point cloud data and delete the point cloud data from the point cloud data set. Execute the above steps repeatedly until the number of point cloud data in the pre-selected set is K, where K is a positive integer.
[0058] In some embodiments, the operation process of the posture generation model includes function calculation, and when the vectorized function calculation is performed in the register, the function calculation result is determined based on a preset function value lookup table.
[0059] In this embodiment, in order to improve the data processing effect, the depth data processing device can determine the function calculation result based on the preset function value lookup table. Specifically, when the candidate pose set corresponding to the target point set is generated by the pose generation model, the pose generation model can predict the candidate pose set by analyzing the shape features of the target object for geometric matching, or it can predict the candidate pose set by learning the spatial distribution of the object. In this process, function calculations such as softmax are usually involved. When performing vectorized function calculations on the data in the register, the function calculation result can be determined based on the preset function value lookup table, so that the function calculation results of multiple data in the register can be obtained more efficiently, wherein the data in the register may refer to the data involved in any function calculation step that needs to be executed during the operation of the pose generation model.
[0060] Among them, the preset function value lookup table is used to store the input and output results corresponding to functions such as trigonometric functions, exponential functions, and logarithmic functions. The preset function value lookup table can be constructed on the CPU, and the CPU can copy the preset function value lookup table to the GPU so that the GPU can quickly read the preset function value lookup table. The method for determining the function calculation result (such as sin(45°)) based on the preset function value lookup table can be: look up the input value and output value corresponding to the function from the preset function value lookup table, if the target input value to be calculated hits the preset function value lookup table, then the output value of the input value is used as the function calculation result of the target input value, if the target input value to be calculated does not hit the preset function value lookup table, then the output values corresponding to the input values adjacent to the target input value in the preset function value lookup table are interpolated and calculated to obtain the function calculation result of the target input value.
[0061] In some embodiments, the depth data processing device can vector-calculate the pose distances between multiple candidate poses in the candidate pose set and other candidate poses in the candidate pose set within a single pose distance calculation instruction cycle; and determine the target pose of the target object based on the pose distances.
[0062] In this embodiment, in order to improve the efficiency of generating pose distances, the depth data processing device can convert the candidate poses in the candidate pose set into quaternary data, and vectorize the pose distances between multiple candidate poses in the candidate pose set and other candidate poses in the candidate pose set within a single pose distance calculation instruction cycle, thereby obtaining the pose distance between every two candidate poses. The quaternary data is represented as [w, x, y, z], where w is the real part of the quaternary data, representing the amplitude of the rotation, and x, y, z are the imaginary parts of the quaternary data, which together represent the axis and angle of rotation. For example, in the candidate pose set, candidate pose q 1 =[w 1 ,x 1 ,y 1,z 1 ], candidate pose q 2 =[w 2 ,x 2 ,y 2 ,z 2 ], can be based on the expression: q 1 ,q 2 The inner product of 1 ,q 2 >=q 1 ·q 2 =w 1 w 2 +x 1 x 2 +y 1 y 2 +z 1 z 2 and d(q 1 ,q 2 )=1― 1 ,q 2 > 2 =1-cosθ, calculate the posture distance d. cosθ is q 1 ,q 2 The angle between 1 ,q 2 The rotation similarity of the candidate pose is the data of the second length.
[0063] In this embodiment, the depth data processing device can cluster candidate poses based on cosθ to obtain multiple pose clusters, and average the pose cluster containing the largest number of poses among the multiple pose clusters to obtain the target pose corresponding to the target object.
[0064] See also Figure 4 . One embodiment of the present application also provides a depth data processing device. The depth data processing device may include: a depth data truncation module, used to obtain second depth data of a second length from first depth data of a first length; wherein the second length is less than the first length; a vectorized mapping module, used to map multiple second depth data into point cloud data within a single mapping instruction cycle; wherein the total length of the multiple second depth data is less than or equal to the number of register bits; a task generation module, used to generate a robot control task for driving the robot terminal to perform a specific operation based on the point cloud data.
[0065] In this embodiment, the specific functions and effects achieved by the depth data processing device can be explained with reference to other embodiments of the present application and will not be repeated here.
[0066] The embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor implements the depth data processing method as described above.
[0067] The embodiment of the present application also provides a computer program product including instructions, and when the computer program product is executed by a processor, the depth data processing method as described above is implemented.
[0068] See also Figure 5 The present description embodiment may provide a computer device, the computer device comprising: a memory, and one or more processors communicatively connected to the memory; the memory stores instructions executable by the one or more processors, the instructions are executed by the one or more processors, so that the one or more processors implement the depth data processing method as described above.
[0069] In some embodiments, the computer device may include a processor, a non-volatile storage medium, an internal memory, a communication interface, a display device, and an input device connected by a system bus. The non-volatile storage medium may store an operating system and related computer programs.
[0070] It should be understood that the specific examples in this article are only intended to help those skilled in the art to better understand the embodiments of the present application, rather than to limit the scope of the present invention.
[0071] It can be understood that in the various implementations of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the implementation methods of the present application.
[0072] It can be understood that the various embodiments described in this application can be implemented individually or in combination, and the embodiments of this application are not limited to this.
[0073] Unless otherwise stated, all technical and scientific terms used in the embodiments of the present application have the same meaning as those generally understood by those skilled in the art of the technical field of the present application. The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the scope of the present application. The term "and / or" used in the present application includes any and all combinations of one or more related listed items. The singular forms of "a kind of", "above" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0074] It can be understood that the processor of the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method implementation can be completed by the hardware integrated logic circuit or software instructions in the processor. The above processor can be a general processor, a digital signal processor (DigitalSignal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiment of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor are combined and executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0075] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (programmable ROM, PROM), an erasable programmable read-only memory (erasable PROM, EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0076] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0077] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method implementation methods and will not be repeated here.
[0078] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0079] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0080] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0081] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0082] The above is only a specific implementation of the present application, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for processing depth data, characterized in that: The method is applied to a robot terminal, comprising: Intercepting the first depth data of the first length to obtain second depth data of the second length; wherein the second length is smaller than the first length; In a single mapping instruction cycle, a plurality of second depth data are mapped into point cloud data; wherein the total length of the plurality of second depth data is less than or equal to the number of register bits; A robot control task for driving the robot terminal to perform a specific operation is generated based on the point cloud data.
2. The method according to claim 1, characterized in that The first length is equal to the number of bits of the register.
3. The method according to claim 1, characterized in that The second depth data is stored as a one-dimensional array, and the data index corresponding to the one-dimensional array is used to indicate the position of the second depth data in the one-dimensional array respectively; within a single mapping instruction cycle, the step of mapping a plurality of second depth data into point cloud data includes: In a single mapping instruction cycle, unprocessed multiple second depth data are read from the one-dimensional array according to the data index and written into a register, and vectorized point cloud mapping calculation is performed on the multiple second depth data in the register to obtain point cloud data corresponding to the multiple second depth data respectively.
4. The method according to claim 1, characterized in that: The first depth data is depth data for a target object in the depth map; The step of generating a robot control task for driving the robot terminal to perform a specific operation based on the point cloud data comprises: Selecting a target point set from the point cloud data based on the point cloud distance between the point cloud data; Generate a candidate pose set corresponding to the target point set based on a pose generation model; generating a target pose corresponding to the target object according to the candidate pose set; Based on the target posture, a robot control task is generated to drive the robot terminal to perform a specific operation; wherein the robot control task includes at least one of the following: a grasping task, an obstacle avoidance task, a clearing task, and a cutting task; and the specific operation is an operation performed on the target object.
5. The method according to claim 4, characterized in that The step of selecting a target point set from the point cloud data based on the point cloud distance between the point cloud data includes: In the first distance calculation instruction cycle, the spatial distance between the x-dimensional data and the y-dimensional data between the point cloud data is calculated to obtain the x-dimensional spatial distance and the y-dimensional spatial distance; In the second distance calculation instruction cycle, the spatial distance between the z-dimensional data and the null value between the point cloud data is calculated to obtain the z-dimensional spatial distance; The corresponding point cloud distance is generated based on the x-dimensional space distance, the y-dimensional space distance, and the z-dimensional space distance, and the target point set is selected from the point cloud data based on the point cloud distance.
6. The method according to claim 4, characterized in that The operation process of the posture generation model includes function calculation. When the vectorized function calculation is performed in the register, the function calculation result is determined based on a preset function value lookup table.
7. The method according to claim 4, characterized in that The step of generating a target pose corresponding to the target object according to the candidate pose set comprises: In a single pose distance calculation instruction cycle, vectorizedly calculating pose distances between a plurality of candidate poses in the candidate pose set and other candidate poses in the candidate pose set; A target pose of the target object is determined based on the pose distance.
8. A computer device, characterized in that: The computer device includes a memory and a processor, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the depth data processing method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by a processor, the method for processing depth data according to any one of claims 1 to 7 can be implemented.
10. A computer program product, characterized in that The computer program product is used to implement the depth data processing method according to any one of claims 1 to 7.