Data processing method and related device
By marking extended constraint information on the point cloud bounding box, the problem of incomplete point cloud data being unable to be effectively used is solved, and the accuracy and utilization rate of point cloud data in autonomous driving systems are improved.
Patent Information
- Application Number
- CN202011613732.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-29
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2040-12-29
AI Technical Summary
In autonomous driving technology, the incomplete point cloud data collected by the acquisition equipment cannot restore the complete outline of the object, making the data difficult to use and affecting the accuracy of the perception and control model.
By marking the extended constraint information on the point cloud bounding box, including the key long side, wide side and high side, ensuring that the z direction is parallel to the z axis and perpendicular to the horizontal plane, the extended constraint is determined by combining the movement direction and the z direction to improve the utilization rate and accuracy of the point cloud data.
It improves the utilization rate of point cloud data and the accuracy of size processing, ensures the validity and integrity of the data, and is suitable for autonomous driving systems.
Smart Images

Figure CN114693865B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a data processing method and related devices. Background Art
[0002] With the development of autonomous driving technology, some unmanned vehicles or drones and other equipment have gradually been put into use. In order to obtain a good perception and control model, the algorithm team needs a rich and effective data set, which is inseparable from the support of data annotation information. However, when the acquisition device collects images (for example, point cloud data, two-dimensional images, etc.), for objects that are far away or obstructed, only partial point cloud data can be collected, that is, incomplete point cloud data. Since part of the point cloud data cannot restore the complete outline of the object, the point cloud data of this part is difficult to use. Summary of the Invention
[0003] An embodiment of the present application discloses a data processing method and related devices, which can mark the extended constraint information of the point cloud bounding box on the point cloud bounding box containing incomplete point cloud data of the target object, so as to improve the accuracy of size processing and the utilization rate of data.
[0004] In a first aspect, embodiments of the present application disclose a data processing method, wherein: incomplete point cloud data of a target object is annotated based on an image of the target object to obtain annotation information, the annotation information including the target object's movement direction and a point cloud bounding box containing the incomplete point cloud data, wherein the z-direction of the point cloud bounding box is parallel to the z-axis and the direction corresponding to the height of the point cloud bounding box, and the z-axis is perpendicular to the horizontal plane; extended constraint information for the point cloud bounding box is determined based on the movement direction and / or the z-direction, wherein the extended constraint information includes at least one of a critical long side, a critical wide side, and a critical high side intersecting at a critical vertex; and the extended constraint information is annotated on the point cloud bounding box. In other words, extended constraint information for sizing the point cloud bounding box is added to the point cloud bounding box, so that a target bounding box that meets actual requirements can be obtained based on the extended constraint information, thereby facilitating improved data utilization. The point cloud bounding box contains the incomplete point cloud data and maintains the z-direction corresponding to the height of the point cloud bounding box and the z-axis perpendicular to the horizontal plane. This avoids directional deviation of the cuboid due to acquisition angle and improves the accuracy of annotating the point cloud bounding box. The extended constraint information is determined according to the moving direction and / or the z-direction, so as to improve the accuracy of the size processing.
[0005] In one possible example, determining the extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction includes: determining the vertex confidence of each vertex in the point cloud bounding box, the vertex confidence being used to describe the probability that the vertex is a vertex of the object bounding box of the target object; determining the vertex corresponding to the maximum value of the vertex confidence as the key vertex; determining the three combined edges in the point cloud bounding box that intersect with the key vertex as three reference edges; and determining at least one of the key long edge, key wide edge, and key high edge among the three reference edges according to the moving direction and / or the z-direction. That is, first, the vertex in the point cloud bounding box that is most likely to coincide with the vertex in the object bounding box is taken as the key vertex, and then, according to the moving direction and / or the z-direction, the edges corresponding to the length, width, and height of the object bounding box among the three reference edges connected to the key vertex are determined as the key long edge, key wide edge, and key high edge, respectively, which can improve the accuracy of determining the extended constraint information.
[0006] In one possible example, determining the vertex confidence of each vertex in the point cloud bounding box includes: determining the vertex confidence of the vertex based on the number of point clouds corresponding to each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the vertex confidence; and / or determining the vertex confidence of the vertex based on the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance, the greater the vertex confidence, and the acquisition device has collected incomplete point cloud data. It can be understood that the point cloud can reflect the information collected by the target object, and the closer the distance between the acquisition device and the point cloud, the higher the accuracy of the collected point cloud. In this example, the probability that the vertex is a vertex of the object bounding box (i.e., the vertex confidence) is determined based on the number of point clouds corresponding to the vertex and / or the distance between the vertex and the acquisition device, which can improve the accuracy of determining the vertex confidence.
[0007] In one possible example, determining the extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction includes: determining the overall confidence of the three combined edges that intersect at each vertex in the point cloud bounding box, the overall confidence being used to describe the probability that the three combined edges are all edges of the object bounding box of the target object; determining the three combined edges corresponding to the maximum value of the overall confidence as three reference edges; determining the vertex where the three reference edges intersect as the key vertex; and determining at least one of the key long edge, key wide edge, and key high edge among the three reference edges according to the moving direction and / or the z-direction. That is, first, the three combined edges in the point cloud bounding box that are most likely to coincide with the edges in the object bounding box are used as the three reference edges, and then, according to the moving direction and / or the z-direction, the edges corresponding to the length, width, and height of the object bounding box among the three reference edges are determined as the key long edge, key wide edge, and key high edge, respectively, to improve the accuracy of determining the extended constraint information.
[0008] In one possible example, determining the overall confidence of three combined edges intersecting each vertex in a point cloud bounding box includes: determining the overall confidence of the three combined edges based on the number of point clouds corresponding to the three combined edges intersecting each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the overall confidence; and / or determining the overall confidence of the three combined edges intersecting each vertex in the point cloud bounding box based on the distance between each vertex in the point cloud bounding box and a capture device, wherein the smaller the distance, the greater the overall confidence, indicating that the capture device has captured incomplete point cloud data. It will be understood that a point cloud can reflect information about the target object being captured, and the closer the distance between the capture device and the point cloud, the greater the accuracy of the captured point cloud. In this example, determining the probability (i.e., the overall confidence) that the three combined edges intersecting a vertex in the point cloud bounding box are all edges of the object bounding box based on the number of point clouds corresponding to the three combined edges in the point cloud bounding box and / or the distance between the vertex where the three combined edges intersect and the capture device can improve the accuracy of determining the overall confidence.
[0009] In one possible example, the target object is a vehicle, and the annotation information also includes the vehicle type. The method further includes: determining a first size of the point cloud bounding box based on the vehicle type; and resizing the point cloud bounding box based on the extended constraint information and the first size to obtain a first target bounding box. This improves the accuracy of the point cloud bounding box resizing and can enhance the authenticity of the first target bounding box.
[0010] In one possible example, resizing the point cloud bounding box based on the extended constraint information and the first size to obtain the first target bounding box includes: determining at least one target edge among the key long edge, key wide edge, and key high edge, as well as a target length and target expansion direction for the at least one target edge, based on the extended constraint information and the first size; and resizing the target edge and the edge corresponding to the target edge in the point cloud bounding box based on the target length and target expansion direction to obtain the first target bounding box. In this way, obtaining the first target bounding box that satisfies the vehicle type based on the first size and the extended constraint information improves the accuracy of resizing the point cloud bounding box.
[0011] In a possible example, the method further includes storing the reference point cloud data obtained by marking the extended constraint information on the point cloud bounding box, thereby further improving the utilization rate of the data.
[0012] In one possible example, the target object is a vehicle, and the annotation information also includes the vehicle type. The method further includes: receiving an annotation instruction for the reference point cloud data; determining a second size of the point cloud bounding box based on the annotation instruction and the vehicle type; and resizing the point cloud bounding box based on the extended constraint information and the second size to obtain a second target bounding box. In this way, based on the second size and the extended constraint information, a second target bounding box that satisfies the annotation instruction and the vehicle type is obtained, thereby improving the accuracy of the point cloud bounding box resizing and increasing data utilization.
[0013] In the second aspect, an embodiment of the present application discloses a data processing device, wherein: a labeling unit is used to label the incomplete point cloud data of the target object based on the image of the target object to obtain labeling information, the labeling information includes the moving direction of the target object and the point cloud bounding box containing the incomplete point cloud data, the z direction of the point cloud bounding box is parallel to the z axis, and is parallel to the direction corresponding to the height of the point cloud bounding box, and the z axis is perpendicular to the horizontal plane; a determination unit is used to determine the extended constraint information of the point cloud bounding box based on the moving direction and / or the z direction, the extended constraint information includes at least one of the key long side, key wide side and key high side intersecting at the key vertex; the labeling unit is also used to label the extended constraint information on the point cloud bounding box. That is, the extended constraint information for size processing of the point cloud bounding box is added to the point cloud bounding box, so that the target bounding box that meets the actual needs can be obtained based on the extended constraint information, which is convenient for improving the utilization rate of the data. The point cloud bounding box contains incomplete point cloud data, and the z-axis remains parallel to the horizontal plane, avoiding directional deviations of the cuboid due to acquisition angles. This improves the accuracy of the point cloud bounding box annotation. Furthermore, the extended constraint information is determined based on the movement direction and / or the z-direction, which facilitates improved sizing accuracy.
[0014] In one possible example, the determination unit is specifically used to determine the vertex confidence of each vertex in the point cloud bounding box, wherein the vertex confidence is used to describe the probability that the vertex is a vertex of the object bounding box of the target object; determine the vertex corresponding to the maximum value of the vertex confidence as the key vertex; determine the three combined edges in the point cloud bounding box that intersect with the key vertex as three reference edges; and determine at least one of the key long edge, key wide edge, and key high edge among the three reference edges based on the moving direction and / or z direction. That is to say, the vertex in the point cloud bounding box that is most likely to coincide with the vertex in the object bounding box is first determined as the key vertex, and then the edges corresponding to the length, width, and height of the object bounding box among the three reference edges connected to the key vertex are determined as the key long edge, key wide edge, and key high edge respectively based on the moving direction and / or z direction, which can improve the accuracy of determining the extended constraint information.
[0015] In one possible example, the determination unit is specifically used to determine the vertex confidence of the vertex based on the number of point clouds corresponding to each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the vertex confidence; and / or, determine the vertex confidence of the vertex based on the distance between each vertex in the point cloud bounding box and the acquisition device, wherein, when the distance is smaller, the vertex confidence is greater, and the acquisition device has collected incomplete point cloud data. It can be understood that the point cloud can reflect the information collected by the target object, and the closer the distance between the acquisition device and the point cloud, the accuracy of the collected point cloud. In this example, the probability that the vertex is a vertex of the object bounding box (i.e., the vertex confidence) is determined based on the number of point clouds corresponding to the vertex and / or the distance between the vertex and the acquisition device, which can improve the accuracy of determining the vertex confidence.
[0016] In one possible example, the determination unit is specifically used to determine the overall confidence of three combined edges intersecting at each vertex in the point cloud bounding box, wherein the overall confidence is used to describe the probability that the three combined edges are all edges of the object bounding box of the target object; the three combined edges corresponding to the maximum value of the overall confidence are determined as three reference edges; the vertex where the three reference edges intersect is determined as the key vertex; and at least one of the key long edge, key wide edge, and key high edge among the three reference edges is determined according to the moving direction and / or the z-direction. In other words, the three combined edges in the point cloud bounding box that are most likely to coincide with the edges in the object bounding box are first used as the three reference edges, and then the edges corresponding to the length, width, and height of the object bounding box among the three reference edges are determined according to the moving direction and / or the z-direction as the key long edge, key wide edge, and key high edge, respectively, to improve the accuracy of determining the extended constraint information.
[0017] In one possible example, the determination unit is specifically configured to determine the overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box based on the number of point clouds corresponding to the three combined edges, wherein the greater the number of point clouds, the greater the overall confidence; and / or, to determine the overall confidence of the three combined edges intersecting at a vertex in the point cloud bounding box based on the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance, the greater the overall confidence, indicating that the acquisition device has captured incomplete point cloud data. It is understood that a point cloud can reflect information about the target object being captured, and the closer the distance between the acquisition device and the point cloud, the higher the accuracy of the captured point cloud. In this example, determining the probability (i.e., the overall confidence) that the three combined edges intersecting at a vertex in the point cloud bounding box are edges of the object bounding box based on the number of point clouds corresponding to the three combined edges, and / or the distance between the vertex where the three combined edges intersect and the acquisition device can improve the accuracy of determining the overall confidence.
[0018] In one possible example, the target object is a vehicle, and the annotation information also includes the vehicle type. The determination unit is further configured to determine a first size of the point cloud bounding box based on the vehicle type. The data processing device further includes a processing unit configured to resize the point cloud bounding box based on the extended constraint information and the first size to obtain a first target bounding box. This improves the accuracy of the point cloud bounding box resizing and enhances the authenticity of the first target bounding box.
[0019] In one possible example, the processing unit is specifically configured to determine, based on the first size and expansion constraint information, at least one target edge among the key long edge, key wide edge, and key height edge, as well as a target length and target expansion direction for the at least one target edge; and based on the target length and target expansion direction, perform resizing on the target edge and the edge corresponding to the target edge in the point cloud bounding box to obtain a first target bounding box. In this manner, the first target bounding box that satisfies the vehicle type is obtained based on the first size and expansion constraint information, thereby improving the accuracy of resizing of the point cloud bounding box.
[0020] In a possible example, the data processing device further includes: a storage unit configured to store reference point cloud data obtained by marking the extended constraint information on the point cloud bounding box, thereby further improving data utilization.
[0021] In one possible example, the target object is a vehicle, the annotation information also includes the vehicle type, and the data processing device further includes a communication unit and a processing unit, wherein: the communication unit is configured to receive annotation instructions for reference point cloud data; the determination unit is further configured to determine a second size of the point cloud bounding box based on the annotation instructions and the vehicle type; and the processing unit is configured to resize the point cloud bounding box based on the extended constraint information and the second size to obtain a second target bounding box. In this manner, obtaining a second target bounding box that satisfies the annotation instructions and the vehicle type based on the second size and the extended constraint information improves the accuracy of the point cloud bounding box resizing and increases data utilization.
[0022] In a third aspect, an embodiment of the present application discloses another data processing device, including a processor and a memory connected to the processor, the memory is used to store one or more programs, and is configured to execute the steps of the first aspect mentioned above by the processor.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the method of the first aspect.
[0024] In a fifth aspect, the present application provides a computer program product, which is used to store a computer program. When the computer program is run on a computer, it enables the computer to execute the method of the first aspect mentioned above.
[0025] In a sixth aspect, the present application provides a chip comprising a processor and a memory, wherein the processor is configured to call and execute instructions stored in the memory, so that a device equipped with the chip executes the method of the first aspect described above.
[0026] In the seventh aspect, the present application provides another chip, including: an input interface, an output interface and a processing circuit, the input interface, the output interface and the processing circuit are connected through an internal connection path, and the processing circuit is used to execute the method of the first aspect mentioned above.
[0027] In the eighth aspect, the present application provides another chip, including: an input interface, an output interface, a processor, and optionally, a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method in any of the above aspects.
[0028] In the ninth aspect, an embodiment of the present application provides a chip system comprising at least one processor, a memory and an interface circuit, wherein the memory, the transceiver and the at least one processor are interconnected through lines, and a computer program is stored in at least one memory; the computer program is executed by the processor according to the method in the first aspect mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The following is an introduction to the drawings used in the embodiments of this application.
[0030] Figure 1 This is a structural diagram of a data processing system provided in an embodiment of the present application;
[0031] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0032] Figure 3 is a schematic diagram of vehicle types available for selection in a marking platform provided in an embodiment of the present application;
[0033] Figure 4 It is a two-dimensional image and a point cloud image collected by a collection device provided in an embodiment of the present application;
[0034] Figure 5 This is a schematic diagram of a point cloud bounding box extended to an imaginary bounding box provided by the prior art;
[0035] Figure 6 This is a flow chart of a data processing method provided in an embodiment of the present application;
[0036] Figure 7 This is a schematic diagram of annotating a point cloud bounding box and a moving direction provided by an embodiment of the present application;
[0037] Figure 8 This is a schematic diagram of annotating extended constraint information provided by an embodiment of the present application;
[0038] Figure 9 It is a two-dimensional image and a point cloud image collected by another collection device provided in an embodiment of the present application;
[0039] Figure 10 This is a flow chart of another data processing method provided in an embodiment of the present application;
[0040] Figure 11 This is a schematic diagram of a point cloud bounding box size processing provided by an embodiment of the present application;
[0041] Figure 12 is a structural diagram of a data processing device provided in an embodiment of the present application;
[0042] Figure 13 It is a structural diagram of another data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0044] Please refer to Figure 1 , Figure 1 This is a system architecture diagram of the data transmission method used in the embodiment of the present application. Figure 1 As shown, the system includes an electronic device 10 and a collection device 20. The present application does not limit the number of the electronic device 10 and the collection device 20.
[0045] The electronic devices in the embodiments of the present application may include but are not limited to personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, mobile phones, tablet computers, personal digital assistants, media players, etc.), consumer electronic devices, minicomputers, mainframe computers, mobile robots, drones, etc. The electronic device may be an on-board device in a computer system (or an on-board system), or other devices, which are not limited here. Figure 1 In the description, the electronic device 10 is described as a personal computer.
[0046] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 2As shown, the electronic device 10 may include a display device 110, a processor 120, and a memory 130. The memory 130 may be used to store software programs and data, and the processor 120 may execute the software programs and data stored in the memory 130 to perform various functional applications and data processing of the electronic device 10.
[0047] The memory 130 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function (such as an image acquisition function), and the like; the data storage area may store data generated based on the use of the electronic device 10 (such as audio data, text information, image data, and the like). Furthermore, the memory 130 may include a high-speed random access memory and a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0048] The processor 120 is the control center of the electronic device 10. It connects the various parts of the entire electronic device 10 using various interfaces and lines. By running or executing software programs and / or data stored in the memory 130, it performs various functions of the electronic device 10 and processes data, thereby monitoring the electronic device 10 as a whole. The processor 120 may include one or more processing units. For example, the processor 120 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Among them, different processing units can be independent devices or integrated into one or more processors.
[0049] The NPU is a neural network (NN) computing processor that rapidly processes input information and continuously self-learns by drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain. The NPU can enable intelligent cognitive applications in electronic device 10, such as image recognition, face recognition, speech recognition, and text comprehension.
[0050] In some embodiments, the processor 120 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuits sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0051] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (CL). In some embodiments, the processor 120 may include multiple I2C bus lines. The processor 120 may be coupled to the touch sensor, charger, flash, camera 160, etc. via different I2C bus interfaces. For example, the processor 120 may be coupled to the touch sensor via the I2C interface, enabling communication between the processor 120 and the touch sensor via the I2C bus interface, thereby implementing the touch function of the electronic device 10.
[0052] The I2S interface can be used for audio communication. In some embodiments, the processor 120 can include multiple I2S buses. The processor 120 can be coupled to the audio module via the I2S bus to enable communication between the processor 120 and the audio module. In some embodiments, the audio module can transmit audio signals to the wireless fidelity (WiFi) module 190 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.
[0053] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module and WiFi module 190 can be coupled via a PCM bus interface. In some embodiments, the audio module can also transmit audio signals to WiFi module 190 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0054] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 120 and the WiFi module 190. For example, the processor 120 communicates with the Bluetooth module in the WiFi module 190 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module can transmit audio signals to the WiFi module 190 via the UART interface, enabling the playback of music via Bluetooth headphones.
[0055] The MIPI interface can be used to connect the processor 120 to peripheral devices such as the display device 110 and the camera 160. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 120 and the camera 160 communicate via the CSI interface to implement the camera function of the electronic device 10. The processor 120 and the display communicate via the DSI interface to implement the display function of the electronic device 10.
[0056] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 120 to the camera 160, the display device 110, the WiFi module 190, the audio module, the sensor module, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0057] A USB interface is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. A USB interface can be used to connect a charger to charge the electronic device 10, or to transfer data between the electronic device 10 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality (AR) devices.
[0058] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may also adopt a different interface connection method from the above embodiment, or a combination of multiple interface connection methods.
[0059] The electronic device 10 further includes a camera 160 for capturing images or videos. The camera 160 may be a common camera or a focus camera.
[0060] The electronic device 10 may further include an input device 140 for receiving input digital information, character information, or contact touch operations / contactless gestures, and generating signal inputs related to user settings and function control of the electronic device 10.
[0061] The display device 110 includes a display panel for displaying information input by a user or information provided to a user, as well as various menu interfaces of the electronic device 10. In the embodiment of the present application, it is mainly used to display an image to be detected captured by a camera or sensor in the electronic device 10. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0062] The electronic device 10 may also include one or more sensors 170, such as image sensors, infrared sensors, laser sensors (which may include laser displacement sensors and lidar sensors, etc.), pressure sensors, gyroscope sensors, air pressure sensors, magnetic sensors, acceleration sensors, distance sensors, proximity light sensors, ambient light sensors, fingerprint sensors, touch sensors, temperature sensors, bone conduction sensors, etc., among which the image sensor may be a time of flight (TOF) sensor, a structured light sensor, etc.
[0063] In addition, the electronic device 10 may also include a power supply 150 for powering other modules. The electronic device 10 may also include a radio frequency (RF) circuit 180 for communicating with wireless network devices, and a WiFi module 190 for communicating with other devices via WiFi, such as receiving images or data transmitted by other devices.
[0064] Although not in Figure 2 As shown in the figure, the electronic device 10 may further include other possible functional modules such as a flash, a Bluetooth module, an external interface, a button, a motor, etc., which are not described in detail here.
[0065] The collection equipment in the embodiments of the present application may be a movable device, which may include but is not limited to aircraft, ships, robots, vehicles, etc., and may also be equipment on the road, such as a road side unit (RSU). The aircraft, ships and vehicles described in the embodiments of the present application may be human-driven equipment or unmanned equipment, which is not limited here. Figure 1 In the description, the acquisition device 20 is described as a vehicle. The acquisition device may include a processor, a display device and a memory, and reference may be made to the description of the electronic device, which will not be repeated here. The acquisition device 20 may also include sensors, such as image pickup devices (for example, cameras, etc.) and lidar sensors. Among them, the image pickup device is used to acquire two-dimensional images. The lidar sensor is used to detect the reflected signal of the laser signal sent by the lidar, thereby obtaining a laser point cloud (or point cloud). It should be noted that the acquisition device for acquiring two-dimensional graphics and acquiring point clouds can be the same device or different devices, which is not limited here. The processor of the acquisition device may include a point cloud processing module for processing point cloud data.
[0066] In an embodiment of the present application, an acquisition device may be used to acquire data of a target object (e.g., at least one of a two-dimensional graphic, point cloud data, and the distance to the target object) and transmit the data to a corresponding electronic device. The electronic device may be used to receive the data transmitted from the acquisition device and execute the data processing method described in the embodiment of the present application based on the data. In one possible example, the data processing method described in the embodiment of the present application is directly executed by the acquisition device (or a processor in the acquisition device or a point cloud processing module in the processor).
[0067] The electronic device can run an application corresponding to the annotation platform, which can be used to display the data received from the collection device and provide it to the annotator for annotation. The data processing method described in the embodiments of the present application can also be executed by the application corresponding to the annotation platform, etc., without limitation here.
[0068] To facilitate understanding, several concepts and terms involved in this application are first introduced below.
[0069] (1) Target object, target object's moving direction and object type.
[0070] In the embodiments of the present application, the target object is an object that the acquisition device or electronic device needs to identify. The target object includes objects on the road and objects outside the road. Among them, the objects on the road include people, cars, traffic lights, traffic signs (such as speed limit signs, etc.), traffic sign poles and foreign objects on the road. Foreign objects refer to objects that should not appear on the road, such as cardboard boxes and tires left on the road. Objects outside the road include buildings and trees on both sides of the road and isolation belts between roads. The target object can also be equipment such as airplanes, ships, robots, etc., which are not limited here.
[0071] The moving direction refers to the moving direction of the target object. When the target object is a vehicle, since the moving direction of the vehicle is generally the forward direction, the moving direction can also be referred to as the vehicle head direction.
[0072] Object type refers to the classification of the target object, for example, aircraft type, ship type, robot type, vehicle type, etc. Object type can be further classified according to aircraft type, ship type, robot type, vehicle type, etc., or specifically classified into drone type, unmanned ship type, unmanned vehicle type, etc. Take vehicle type as an example, Figure 3 As shown, when the target object is a vehicle, the vehicle types in the annotation platform may include buses, motorcycles, bicycles, engineering vehicles, tricycles, tank trucks or pickup trucks for the annotation personnel to choose. The above vehicle types can also be provided to the electronic device for identification based on the image features of the vehicle type. It can be understood that the vehicle sizes are different between different vehicle types, and each vehicle corresponds to a size. The object sizes are also different between different object types. Therefore, the approximate size of the target object in the point cloud data can be determined based on the object type. The moving direction of the target object of different object types is different from the length and width of the rectangular parallelepiped corresponding to the object. For example, when the target object is a vehicle, the moving direction of the target object is usually consistent with the direction corresponding to the length of the rectangular parallelepiped corresponding to the vehicle. When the target object is a humanoid robot, the moving direction of the target object is usually the direction of upright walking, that is, consistent with the direction corresponding to the width of the rectangular parallelepiped corresponding to the humanoid robot.
[0073] (2) Point cloud data and point cloud images.
[0074] Point cloud data, also known as laser point cloud (PCD), three-dimensional point cloud or point cloud, is a collection of massive points that express the spatial distribution of the target and the surface characteristics of the target by using a laser to obtain the three-dimensional spatial coordinates of each sampling point on the surface of an object in the same spatial reference system (usually expressed in the form of x, y, z three-dimensional coordinates). Compared with images, although point clouds lack detailed texture information, they contain rich three-dimensional spatial information. In addition to three-dimensional spatial information, point cloud data can also include color information, grayscale values, depth, segmentation results, etc., which are not limited here. In the embodiment of the present application, the image obtained by projecting point cloud data onto a two-dimensional plane is called a point cloud image.
[0075] (3) Images and incomplete point cloud data of target objects.
[0076] The acquisition device can only collect part of the point cloud for objects that are far away or blocked. In the embodiment of this application, all the collected part of the point cloud data is called incomplete point cloud data. If the point cloud data of the target object is insufficient, all the collected point cloud data of the target object is called incomplete point cloud data of the target object. Please refer to Figure 4 , Figure 4 Take the target object as an example. Figure 4 (a) is a two-dimensional graphic collected by the collection device 20. Figure 4 (b) in the figure is the point cloud image corresponding to all the point cloud data collected by the acquisition device 20. Figure 4 As can be seen from (a) in the figure, there are four target objects (i.e., four vehicles) 21 in front of the acquisition device 20, and the distance between the acquisition device 20 and the target object 21 in front is relatively far, and the target object 21 near the roadside may be blocked by the leaves on the roadside, so the point cloud data of the target object 21 may be incomplete. Figure 4 As can be seen from (b), there is sparse point cloud data in the annotation box of the target object 21, so it can be determined that the point cloud data of the target object 21 is insufficient. All the point cloud data of the target object 21 collected by the acquisition device 20 is called incomplete point cloud data of the target object 21.
[0077] The image of the target object includes a two-dimensional image captured by the acquisition device for the target object, and may also include a three-dimensional image corresponding to incomplete point cloud data of the target object, etc., which is not limited here.
[0078] (4) Bounding box (BB), point cloud bounding box and object bounding box.
[0079] Bounding box is an algorithm for finding the optimal bounding space for a discrete set of points. The basic idea is to use a slightly larger geometric body with simpler characteristics (called a bounding box) to approximate complex geometric objects. The most common bounding boxes are spheres, axis-aligned bounding boxes (AABBs), spheres, oriented bounding boxes (OBBs), and fixed directions hulls (k-DOPs, FDHs). Among them, axis-aligned bounding boxes and directed bounding boxes are the bounding boxes corresponding to rectangular parallelepipeds. The axis-aligned bounding box of a given object is defined as the smallest hexahedron that contains the object and has sides parallel to the coordinate axes. The directed bounding box of a given object is defined as the smallest rectangular parallelepiped that contains the object and has an arbitrary direction relative to the coordinate axes. The most important feature of a directed bounding box is its arbitrary direction, which allows it to enclose the object as tightly as possible based on the shape characteristics of the enclosed object. For example, when the z-direction corresponding to the z-axis of the directed bounding box is perpendicular to the horizontal plane, the x-direction corresponding to the x-axis / the y-direction corresponding to the y-axis may have a certain angle with the x-axis / y-axis such that the area of the surface formed by the x-direction / y-direction is minimized. The z-axis described in the embodiments of the present application may be the z-axis direction in the geodetic coordinate system, etc., and is not limited here.
[0080] The point cloud bounding box in the embodiment of the present application is a cuboid that contains all the point clouds of a given object (i.e., incomplete point cloud data of the target object), and the z direction of the cuboid (i.e., the point cloud bounding box) is parallel to the z axis and the direction corresponding to the height of the point cloud bounding box, and the z axis is perpendicular to the horizontal plane. The axis of the coordinate axis corresponding to the point cloud bounding box can be located at the center of the point cloud bounding box, and the x-axis, y-axis, and z-axis directions of the coordinate axis can be parallel to the length, width, and height of the point cloud bounding box, respectively. For a cuboid, the height is perpendicular to the horizontal plane, and the length is greater than the width. Therefore, the z direction of the point cloud bounding box is parallel to the direction corresponding to the height of the point cloud bounding box, and is parallel to the x direction of the point cloud bounding box and the direction corresponding to the length of the point cloud bounding box, and the y direction is parallel to the direction corresponding to the width of the point cloud bounding box. The length of the side is related to the length of the point cloud data actually collected. That is to say, in addition to the edge corresponding to the height, the edge with the longer number of collected point clouds can be used as the length of the point cloud bounding box, and the other edge that intersects the length and height at a vertex or the edge parallel to the edge can be used as the width of the point cloud bounding box.
[0081] It can be understood that the z direction of the object bounding box corresponding to the target object is parallel to the z axis. When the z direction of the point cloud bounding box is limited to be parallel to the z axis, regardless of the acquisition angle, the height direction corresponding to the point cloud bounding box and the object bounding box can be guaranteed to be perpendicular to the horizontal plane. In other words, the z direction of the point cloud bounding box is parallel to the z direction of the target object, which can improve the accuracy of annotating the point cloud bounding box. Then, based on the mutually perpendicular directional relationship between the x-axis, y-axis and z-axis, the point cloud bounding box containing the incomplete point cloud data of the target object is annotated. Optionally, the type of the point cloud bounding box is a directed bounding box, and the z direction of the point cloud bounding box is parallel to the z axis. Since the directed bounding box is defined as the smallest cuboid that contains the object and is arbitrarily oriented relative to the coordinate axis, the compactness of the point cloud bounding box can be guaranteed, thereby further improving the accuracy of annotating the point cloud bounding box.
[0082] The object bounding box in the embodiment of the present application is a rectangular parallelepiped containing a given object (i.e., the target object), and the z direction of the rectangular parallelepiped is parallel to the z axis. Referring to the description of the point cloud bounding box, the direction corresponding to the height of the object bounding box is parallel to the z axis, the x direction of the object bounding box is parallel to the direction corresponding to the length of the object bounding box, and the y direction of the object bounding box is parallel to the direction corresponding to the width of the object bounding box. The length or width of the object bounding box can be determined by the moving direction and object type of the object bounding box. For example, when the object type of the target object is a vehicle, the moving direction of the target object is generally consistent with the direction corresponding to the length of the object bounding box, so that the side of the object bounding box that is parallel to the moving direction of the target object can be determined as the length of the object bounding box, and then the other side of the object bounding box that intersects with the length and height at a vertex or the side parallel to the side can be determined as the width of the object bounding box. When the target object is a humanoid robot, the moving direction of the target object is usually the direction of upright walking, that is, the direction corresponding to the width of the object bounding box is consistent. Therefore, the side of the object bounding box parallel to the moving direction of the target object can be determined as the length of the object bounding box, and the other side of the object bounding box that intersects the length and height at a vertex or the side parallel to the side can be used as the width of the object bounding box.
[0083] Currently, the commonly used methods for annotating incomplete point cloud data are as follows: Figure 5 As shown in (a) of FIG, when annotating data, the bounding box 30 corresponding to the incomplete point cloud data and the moving direction indicated by the arrow A1 are first annotated; then the bounding box 30 is expanded to obtain the imaginary bounding box 31 and the moving direction indicated by the arrow A2, so that the imaginary bounding box 31 conforms to the size of a normal vehicle (for example, Figure 5 The vehicle type marked in (b) is an engineering vehicle, and the length, width and height (in meters) of the imaginary bounding box 31 corresponding to the engineering vehicle are 1.77, 2.78 and 2.00 respectively.
[0084] In this method, the standard for expanding the bounding box is difficult to determine. Typically, the expansion is performed by annotators based on their subjective experience, which lacks objectivity. Furthermore, the expansion operation reduces annotation efficiency. Furthermore, the resulting hypothetical bounding box is difficult for other teams or individuals to use.
[0085] Based on this, a data processing method provided in the example of this application can be applied to a data processing device, which can be the above-mentioned electronic device or acquisition device. This embodiment of the application takes an electronic device as an example to describe the data processing method, please refer to Figure 6 , Figure 6 This is a flow chart of a data processing method used in an embodiment of the present application. The method may include the following steps S601 to S603, wherein:
[0086] S601: Annotate the incomplete point cloud data of the target object according to the image of the target object to obtain annotation information, wherein the annotation information includes the movement direction of the target object and a point cloud bounding box containing the incomplete point cloud data, the z direction of the point cloud bounding box is parallel to the z axis and parallel to the direction corresponding to the height of the point cloud bounding box, and the z axis is perpendicular to the horizontal plane.
[0087] Among them, the target object, the image of the target object and the incomplete point cloud data, the point cloud bounding box and the moving direction of the target object, the z direction and the height corresponding direction of the point cloud bounding box, and the z axis can refer to the above definitions and will not be repeated here. In the embodiment of the present application, the annotation information includes the point cloud bounding box containing the incomplete point cloud data, and the moving direction of the target object, such as Figure 7 As shown in (b), the point cloud bounding box 30 represented by a rectangular parallelepiped and the moving direction represented by arrow A1 may be included.
[0088] The annotation information may also include object type, occlusion, up and down direction, scaling ratio, etc., which are not limited here. Among them, the object type can be referred to above and will not be repeated here. When the target object is a vehicle, the annotation information may also include the vehicle type. The occlusion is used to describe whether the target object is occluded, as well as information such as the occluded part, which can determine the missing status of the incomplete point cloud data. When most of the data is missing, the accuracy of the target bounding box obtained by sizing the incomplete point cloud data is insufficient. When the object type of the target object is a vehicle, ship, aircraft or other manned equipment, the up and down direction corresponds to the width direction of the object bounding box. The width direction of the object bounding box can be directly determined based on the up and down direction. When the up and down direction is marked in the point cloud bounding box, the width direction corresponding to the point cloud bounding box and the object bounding box can be determined, which facilitates improving the rate of determining the expansion direction of the point cloud bounding box. The scaling ratio is used to describe the scaling between the point cloud data collected by the acquisition device and the real object. The scaling ratio can be understood as the scaling between the incomplete point cloud data corresponding to the target object and the actual target object.
[0089] The embodiments of the present application do not limit the method for annotating point cloud bounding boxes. In one possible example, a target area corresponding to a target object in point cloud data collected by an acquisition device is determined; and a point cloud bounding box corresponding to incomplete point cloud data in the target area is annotated according to a preset algorithm.
[0090] The target area is the location corresponding to the target object in the point cloud data collected by the acquisition device. In this application, the annotation personnel can compare the point cloud data with the two-dimensional graphics, and the obtained annotation information can determine the target area corresponding to the target object. The electronic device can also determine the location of the target object in the point cloud data as the target area based on the mapping relationship between the two-dimensional graphics and the point cloud data, thereby using the three-dimensional point cloud corresponding to the target area as the incomplete point cloud data corresponding to the target object. This application does not limit the method for determining the target area.
[0091] This application does not limit the preset algorithm. It can be obtained by shrinking the rectangle corresponding to the point cloud data in the target area according to the point cloud bounding box required to contain all the target objects' point clouds (i.e., incomplete point cloud data), and the z direction of the rectangle is parallel to the z axis.
[0092] Taking this example, Figure 7 As shown in (a), first determine the target area 32 corresponding to the target object in the point cloud data collected by the acquisition device, then determine the cuboid corresponding to the incomplete point cloud data in the target area 32 according to the preset algorithm, and then mark the cuboid, that is, Figure 7The point cloud bounding box 30 corresponding to the incomplete point cloud data shown in (b) in FIG. In this way, the accuracy of the annotated point cloud bounding box can be improved, and the point cloud bounding box can be expanded based on actual needs in the future.
[0093] Furthermore, the cuboid obtained by the above-mentioned preset algorithm can be fine-tuned by annotators to obtain a point cloud bounding box. It can be understood that the point cloud bounding box obtained by fine-tuning by annotators can further improve the accuracy of the point cloud bounding box.
[0094] S602: Determine extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, wherein the extended constraint information includes at least one of a key long side, a key wide side, and a key high side intersecting at a key vertex.
[0095] In an embodiment of the present application, the extended constraint information is used to limit the expansion direction of the point cloud bounding box, and may include key vertices and at least one of the key long sides, key wide sides and key high sides that intersect at the key vertices. It may also include the expansion directions corresponding to the key vertices and the key long sides, key wide sides and key high sides, etc., which are not limited here.
[0096] Among them, the key vertex is used to describe the vertex in the point cloud bounding box that is closest to the object bounding box. In other words, the key vertex is the vertex in the point cloud bounding box that is most likely to coincide with the vertex in the object bounding box. For example, Figure 8 The vertex d1 in the point cloud bounding box 30 shown in (b) is a key vertex.
[0097] The key long side, key wide side, and key high side are three combined edges that intersect at a key vertex, and are the edges corresponding to the length, width, and height of the object's bounding box, respectively. In other words, the key long side, key wide side, and key high side are most likely to coincide with the three edges of the object's bounding box that intersect at a vertex. For example, Figure 8 The middle arrow A1 indicates the moving direction of the target object. Line segments L1, L2 and L3 intersect at the key vertex d1. In the point cloud bounding box 30, the edges corresponding to line segments L1, L2 and L3 can be respectively called the key long edge, key wide edge and key high edge.
[0098] This application does not limit the method for determining the extended constraint information, and may include the following two implementations:
[0099] The first method is to determine the vertex confidence of each vertex in the point cloud bounding box; determine the vertex corresponding to the maximum value of the vertex confidence as the key vertex; determine the three combined edges in the point cloud bounding box that intersect with the key vertex as three reference edges; and determine at least one of the key long edge, key wide edge and key high edge among the three reference edges according to the moving direction and / or z direction.
[0100] The vertex confidence is used to describe the probability that a vertex of the point cloud bounding box is a vertex of the object bounding box of the target object. This application does not limit the method for determining vertex confidence. In a first possible example, the vertex confidence of each vertex is determined based on the number of point clouds corresponding to each vertex in the point cloud bounding box. The greater the number of point clouds, the greater the vertex confidence.
[0101] The number of point clouds corresponding to a vertex can be understood as the number of point clouds within a preset range corresponding to the vertex. The preset range can be a 1 / 4 sphere, a triangular pyramid, or a cube formed by connecting the points in the three combined edges connected to the vertex in the point cloud bounding box, whose distance from the vertex differs by the same threshold. It is not limited here. Optionally, the embodiment of the present application can select the vertex corresponding to a plane in the point cloud bounding box, such as Figure 8 As shown in (a), vertices d1, d2, d3, and d4 are selected from the point cloud bounding box. The relationship between the number of point clouds within the preset range corresponding to each of vertices d1, d2, d3, and d4 is d1 > d2 > d3 > d4. Since a larger number of point clouds indicates a higher vertex confidence, the relationship between the vertex confidences of vertices d1, d2, d3, and d4 is d1 > d2 > d3 > d4. Therefore, vertex d1 can be determined to be the key vertex.
[0102] It can be understood that a point cloud can reflect information about the target object being captured. The larger the number of point clouds, the greater the probability that the area corresponding to the point cloud data is captured. In this example, the probability that a vertex is a vertex of the object's bounding box (i.e., the vertex confidence) is determined based on the number of point clouds corresponding to the vertex, which can improve the accuracy of determining the vertex confidence.
[0103] In a second possible example, the vertex confidence of the vertex is determined according to the distance between the vertex of the point cloud bounding box and the acquisition device. When the distance is smaller, the vertex confidence is greater.
[0104] Among them, the acquisition device acquires incomplete point cloud data, that is, the acquisition device is a device for acquiring incomplete point cloud data. The acquisition device can also be a device for acquiring a two-dimensional image of a target object, which is not limited here. The distance between the vertices of the point cloud bounding box and the acquisition device can be calculated by the three-dimensional coordinates corresponding to the vertices of the point cloud bounding box and the three-dimensional coordinates corresponding to the acquisition device (which can be a lidar sensor in the acquisition device). It can also be calculated by the three-dimensional coordinates corresponding to the vertices of the object bounding box corresponding to the vertices of the point cloud bounding box and the three-dimensional coordinates corresponding to the acquisition device (which can be a lidar sensor in the acquisition device), etc., which is not limited here.
[0105] For example, Figure 8 As shown in (a), vertices d1, d2, d3, and d4 are selected from the point cloud bounding box. The distance between each of these vertices and the acquisition device is: d1 < d2 < d3 < d4. Since vertex confidence increases with closer distances, the confidence levels of vertices d1, d2, d3, and d4 are: d1 > d2 > d3 > d4. Therefore, vertex d1 is identified as the key vertex.
[0106] It is understood that the closer the acquisition distance of the acquisition device, the higher the accuracy of the acquired point cloud data. Therefore, in this example, the probability that a vertex of the point cloud bounding box is a vertex of the object bounding box (i.e., the vertex confidence) is determined based on the distance between the vertex and the acquisition device, which can improve the accuracy of determining the vertex confidence.
[0107] In a third possible example, the occlusion probability of each vertex in the point cloud bounding box is determined based on the two-dimensional image of the target object; the vertex confidence is determined based on the occlusion probability, and when the occlusion probability is greater, the vertex confidence is smaller.
[0108] The occlusion probability of a vertex is used to describe the probability that the vertex is occluded, that is, the probability that the vertex can be used to restore the contour. For example, Figure 8 As shown in (a) in the figure, it is assumed that the occlusion probabilities of vertices d1, d2, d3, and d4 can be determined from the two-dimensional image to be in the order of d1 < d2 < d3 < d4. Since a smaller occlusion probability indicates a higher vertex confidence, the confidence levels of vertices d1, d2, d3, and d4 are in the order of d1 > d2 > d3 > d4. Therefore, d1 can be determined to be the key vertex.
[0109] It can be understood that a two-dimensional image can reflect the occlusion of the target object. When the occlusion probability is greater, the probability that the vertex can be restored is lower, that is, the overall confidence level is lower. In this example, the probability of occlusion of each vertex (i.e., occlusion probability) can be determined based on the two-dimensional image, and the vertex confidence level is then determined based on the occlusion probability, which can improve the accuracy of determining the vertex confidence level.
[0110] It should be noted that the above three possible examples do not constitute a limitation on the embodiments of the present application. In practical applications, other implementation methods can also be used to determine vertex confidence or key vertices. For example, the occlusion probability of the vertices of the point cloud bounding box is determined based on the two-dimensional image and incomplete point cloud data, and the vertex confidence is determined based on the occlusion probability. The vertex confidence can also be determined based on the number of point clouds and the distance (for example, the weighted average corresponding to the number of point clouds and the distance is obtained, and the vertex confidence is determined by the weighted average, etc.); or the key vertices are determined based on the point cloud data and the distance (for example, when it is determined that the maximum number of point clouds corresponds to 2 or more vertices, the vertex with the smallest distance can be used as the key vertex based on the distance between the 2 or more vertices and the acquisition device; or when it is determined that the smallest distance corresponds to 2 or more vertices, the vertex with the smallest distance can be used as the key vertex; or determine the vertex confidence according to the number of point clouds and the occlusion probability (for example, obtain the weighted average value corresponding to the number of point clouds and the distance, and determine the vertex confidence according to the weighted average value); or determine the key vertex according to the number of point clouds and the occlusion probability (for example, when it is determined that the largest number of point clouds corresponds to 2 or more vertices, the vertex with the smallest occlusion probability can be used as the key vertex according to the occlusion probability of the 2 or more vertices here; or when it is determined that the smallest occlusion probability corresponds to 2 or more vertices, the vertex corresponding to the maximum number of point clouds can be used as the key vertex according to the number of point clouds corresponding to the 2 or more vertices here, etc.).
[0111] In a cuboid, each vertex can be connected to three edges. In the embodiment of the present application, the three combined edges that intersect with the key vertex in the point cloud bounding box can be called three reference edges. In other words, the reference edges are the edges in the point cloud bounding box that are connected to the key vertex.
[0112] In an embodiment of the present application, the key high side is the side corresponding to the height of the object bounding box, and the height of the object bounding box and the height of the point cloud bounding box are both parallel to the z direction of the point cloud bounding box. The side corresponding to the z direction among the three reference sides can be determined as the key high side. As mentioned above, the long side or wide side corresponding to the moving direction in the point cloud bounding box (object bounding box) can be determined according to the object type of the target object. Therefore, when the long side corresponding to the moving direction can be determined according to the object type, the long side corresponding to the moving direction can be selected from the three reference sides that intersect with the key vertex as the key long side of the point cloud bounding box. When the wide side corresponding to the moving direction can be determined according to the object type, the wide side corresponding to the moving direction can be selected from the three reference sides that intersect with the key vertex as the key wide side of the point cloud bounding box. It can be understood that each of the key long side, key wide side, and key high side is a reference side of the three reference sides excluding the other two reference sides. When the key long side corresponding to the movement direction can be determined based on the object type, the reference sides excluding the reference sides corresponding to the movement direction and the reference sides corresponding to the z-direction can be determined as the key wide sides. When the key wide side corresponding to the movement direction can be determined based on the object type, the reference sides excluding the reference sides corresponding to the movement direction and the reference sides corresponding to the z-direction can be determined as the key long sides.
[0113] Taking the target object as an example, the moving direction corresponds to the direction of the long side. In other words, the side corresponding to the moving direction among the three reference sides can be determined as the key long side. Figure 8 As shown in (b), when the key vertex is d1, the edge in the point cloud bounding box 30 that corresponds to the moving direction corresponding to arrow A1 and intersects with the key vertex d1 can be determined as the key long edge, that is, the key long edge is line segment L1. According to the z direction of the point cloud bounding box 30 corresponding to the z-axis direction, and the edge that intersects with the key vertex d1 is the key high edge, that is, the line segment L3 in the point cloud bounding box 30 is determined to be the key high edge. Finally, since the three reference edges have the reference edge corresponding to the moving direction and the reference edge corresponding to the z direction, the remaining reference edges are key wide edges. Therefore, the edge that intersects with the key vertex d1 in the point cloud bounding box still has line segment L2, which is the key wide edge.
[0114] This application does not limit which of the key long side, key wide side and key high side is determined. The key long side, key wide side and key high side can be determined separately, or the incomplete point cloud data can be analyzed to obtain the side that needs to be expanded. For example, the target size of the target object in the point cloud image is determined according to the object type or the specific object type (for example, vehicle type, aircraft type, ship type, robot type, etc.), and the target size includes the lengths corresponding to the length, width and height. The key long side and / or key wide side can be determined according to the moving direction and the z direction, and the key high side can be determined according to the z direction. Then, based on the lengths of the three sides in the target size and the sizes of the key long side, key wide side and key high side in the point cloud bounding box, the side that needs to be expanded among the key long side, key wide side and key high side can be determined.
[0115] It can be understood that in the first method of determining extended constraint information, the vertex confidence of each vertex in the point cloud bounding box is first determined, and then the vertex corresponding to the maximum vertex confidence is selected as the key vertex. In other words, the vertex in the point cloud bounding box that is most likely to overlap with a vertex in the object bounding box is first selected as the key vertex, and then the three combined edges that intersect with the key vertex are determined as the three reference edges. Then, based on the movement direction and / or z-direction, the edges of the three reference edges that correspond to the length, width, and height of the object bounding box are determined as the key length edge, key width edge, and key height edge, respectively. This can improve the accuracy of determining extended constraint information.
[0116] The second method is to determine the overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box; determine the three combined edges corresponding to the maximum value of the overall confidence as three reference edges; determine the vertex where the three reference edges intersect as the key vertex; and determine at least one of the key long edge, key wide edge and key high edge among the three reference edges according to the moving direction and / or z direction.
[0117] The overall confidence is used to describe the probability that all three combined edges are edges of the object bounding box of the target object. This application does not limit the method for determining the overall confidence. In a first possible example, the overall confidence of the three combined edges is determined based on the number of point clouds corresponding to the three combined edges intersecting at each vertex in the point cloud bounding box. The greater the number of point clouds, the greater the overall confidence.
[0118] Among them, the number of point clouds corresponding to the three combined edges intersecting at each vertex in the point cloud bounding box can be understood as the number of point clouds within a preset range corresponding to each vertex and the three combined edges. The preset range can be a 1 / 4 sphere, a triangular pyramid or a cube formed by connecting the points in the three combined edges whose distances from the vertices differ by the same threshold, and is not limited here. Optionally, the embodiment of the present application can select the three combined edges corresponding to the vertices corresponding to a plane in the point cloud bounding box, such as Figure 8As shown in (a), select vertices d1, d2, d3, and d4 in the point cloud bounding box. The three combined edges corresponding to vertex d1 are L1, L2, and L3, the three combined edges corresponding to vertex d2 are L2, L7, and L6, the three combined edges corresponding to vertex d3 are L1, L4, and L5, and the three combined edges corresponding to vertex d4 are L5, L6, and L8. The relationship between the number of point clouds corresponding to the three combined edges intersecting each vertex is: the three combined edges corresponding to vertex d1 > the three combined edges corresponding to vertex d2 > the three combined edges corresponding to vertex d3 > the three combined edges corresponding to vertex d4. According to the fact that the larger the number of point clouds, the greater the overall confidence, the size relationship between the overall confidence of the three combined edges corresponding to each vertex in vertex d1, vertex d2, vertex d3 and vertex d4 is: three combined edges corresponding to vertex d1 > three combined edges corresponding to vertex d2 > three combined edges corresponding to vertex d3 > three combined edges corresponding to vertex d4. Therefore, the three reference edges can be determined to be the three combined edges corresponding to vertex d1, namely L1, L2 and L3.
[0119] It can be understood that a point cloud can reflect information about the target object being captured. A larger number of point clouds indicates a greater probability that the area corresponding to the point cloud data has been captured. In this example, the probability that the three combined edges intersecting at each vertex in the point cloud bounding box are all edges of the object's bounding box (i.e., the overall confidence level) is determined based on the number of point clouds corresponding to the three combined edges. This can improve the accuracy of determining the overall confidence level.
[0120] In a second possible example, the overall confidence of three combined edges intersecting at the vertex in the point cloud bounding box is determined based on the distance between the vertex of the point cloud bounding box and the acquisition device. The smaller the distance, the greater the overall confidence.
[0121] Among them, the acquisition device acquires incomplete point cloud data, that is, the acquisition device is a device for acquiring incomplete point cloud data. The acquisition device can also be a device for acquiring a two-dimensional image of a target object, which is not limited here. The distance between the vertices of the point cloud bounding box and the acquisition device can be calculated by the three-dimensional coordinates corresponding to the vertices of the point cloud bounding box and the three-dimensional coordinates corresponding to the acquisition device (which can be a lidar sensor in the acquisition device). It can also be calculated by the three-dimensional coordinates corresponding to the vertices of the object bounding box corresponding to the vertices of the point cloud bounding box and the three-dimensional coordinates corresponding to the acquisition device (which can be a lidar sensor in the acquisition device), etc., which is not limited here.
[0122] For example, Figure 8As shown in (a), vertices d1, d2, d3, and d4 are selected from the point cloud bounding box. The distance between each of these vertices and the acquisition device is in the following order: vertex d1 < vertex d2 < vertex d3 < vertex d4. Since the closer the distance, the greater the overall confidence, the overall confidence of the three combined edges corresponding to each of vertices d1, d2, d3, and d4 is in the following order: three combined edges corresponding to vertex d1 > three combined edges corresponding to vertex d2 > three combined edges corresponding to vertex d3 > three combined edges corresponding to vertex d4. Therefore, the three reference edges can be determined to be the three combined edges corresponding to vertex d1, namely L1, L2, and L3.
[0123] It is understood that the closer the acquisition distance of the acquisition device, the higher the accuracy of the collected point cloud data. Therefore, in this example, the probability that the three combined edges corresponding to the vertex of the point cloud bounding box are all edges of the object bounding box (i.e., the overall confidence level) is determined based on the distance between the vertex and the acquisition device, which can improve the accuracy of determining the overall confidence level.
[0124] In a third possible example, the occlusion probability of three combined edges intersecting at each vertex in the point cloud bounding box is determined based on the two-dimensional image of the target object; the overall confidence is determined based on the occlusion probability, and the greater the occlusion probability, the smaller the overall confidence.
[0125] The occlusion probability of the three combined edges is used to describe the probability that the area corresponding to the three combined edges in the object bounding box is occluded, that is, the probability that the preset area corresponding to the three combined edges can be used to restore the contour. Figure 8 As shown in (a) in the figure, it is assumed that the occlusion probabilities of the three combined edges corresponding to each vertex d1, vertex d2, vertex d3, and vertex d4 in the two-dimensional image are determined to be: the three combined edges corresponding to vertex d1 < the three combined edges corresponding to vertex d2 < the three combined edges corresponding to vertex d3 < the three combined edges corresponding to vertex d4. Since the smaller the occlusion probability, the greater the overall confidence, the overall confidence of the three combined edges corresponding to each vertex d1, vertex d2, vertex d3, and vertex d4 is determined to be: the three combined edges corresponding to vertex d1 > the three combined edges corresponding to vertex d2 > the three combined edges corresponding to vertex d3 > the three combined edges corresponding to vertex d4. Therefore, the three reference edges can be determined to be the three combined edges corresponding to vertex d1, namely L1, L2, and L3.
[0126] It can be understood that a two-dimensional image can reflect the occlusion of the target object. When the occlusion probability is greater, the probability that the vertex can be restored is lower, that is, the overall confidence level is lower. In this example, the probability of occlusion of each of the three combined edges intersecting each vertex (i.e., the occlusion probability) can be determined based on the two-dimensional image. The overall confidence level is then determined based on the occlusion probability, which can improve the accuracy of the overall confidence level.
[0127] It should be noted that the above three possible examples do not constitute a limitation on the embodiments of the present application. In practical applications, other implementation methods can also be used to determine the overall confidence or determine three reference edges. For example, the occlusion probability of the vertices of the point cloud bounding box is determined based on the two-dimensional image and the incomplete point cloud, and the overall confidence is determined based on the occlusion probability. The overall confidence can also be determined based on the number of point clouds and the distance (for example, the weighted average corresponding to the number of point clouds and the distance is obtained, and the overall confidence is determined by the weighted average, etc.); or three key edges are determined based on the point cloud data and the distance (for example, when it is determined that the maximum number of point clouds corresponds to 2 or more vertices, the three combined edges that intersect with the vertex with the smallest distance can be used as the three key edges based on the distance between the 2 or more vertices and the acquisition device; or when it is determined that the smallest distance corresponds to 2 or more vertices, the three combined edges that intersect with the vertex with the smallest distance can be used based on the distance between the 2 or more vertices and the acquisition device; The number of point clouds corresponding to the 2 or more vertices here is used as the three combined edges that intersect with the vertex corresponding to the maximum number of point clouds as the three key edges, etc.); or the overall confidence is determined according to the number of point clouds and the occlusion probability (for example, the weighted average corresponding to the number of point clouds and the distance is obtained, and the overall confidence is determined by the weighted average, etc.); or the three key edges are determined according to the number of point clouds and the occlusion probability (for example, when it is determined that the maximum number of point clouds corresponds to 2 or more vertices, the three combined edges that intersect with the vertex with the smallest occlusion probability can be used as the three key edges according to the occlusion probability of the 2 or more vertices here; or when it is determined that the smallest occlusion probability corresponds to 2 or more vertices, the three combined edges that intersect with the vertex corresponding to the maximum number of point clouds can be used as the three key edges according to the number of point clouds corresponding to the 2 or more vertices here, etc.). When the edge confidence of each edge in the point cloud bounding box can be determined, the overall confidence corresponding to the three combined edges intersecting at a vertex can be determined based on the edge confidence of the three combined edges. The edge confidence here describes the probability that the edge corresponding to the edge confidence is an edge of the object bounding box. The overall confidence can be obtained by weighting the edge confidence of the three combined edges according to the preset weights corresponding to the length, width, and height respectively. The preset weights here can be set based on the importance of the length, width, and height to be expanded, etc., and are not limited here.
[0128] Determining the key long side, key wide side and key high side, and determining which side among the key long side, key wide side and key high side can refer to the description in the first method for determining extended constraint information, which will not be repeated here.
[0129] It can be understood that in the second method of determining extended constraint information, the overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box is first determined, and the three combined edges corresponding to the maximum overall confidence are used as the three reference edges. In other words, the edges in the point cloud bounding box that are most likely to overlap with the edges in the object bounding box are first used as reference edges, and then the three combined edges intersecting with the key vertices are determined as the three reference edges. Then, based on the movement direction and / or z-direction, the edges corresponding to the length, width, and height of the object bounding box among the three reference edges are determined as the key length edge, key width edge, and key height edge, respectively. This can improve the accuracy of determining extended constraint information.
[0130] Before step S602 , the method further includes: determining an unobstructed position of the target object including boundary information of the target object according to the two-dimensional image.
[0131] Boundary information includes the vertices or edges of the bounding box containing the target object. It is understood that when determining boundary information of the target object in incomplete point cloud data based on a two-dimensional image, expansion operations can be performed based on this boundary information. Furthermore, the fixed position of this boundary information facilitates improved accuracy in determining expanded constraint information and enhances data reuse.
[0132] In another possible example, if the unobstructed position of the target object determined according to the two-dimensional image does not include boundary information of the target object, step S602 is not performed.
[0133] like Figure 9 As shown in (a), the front part of the target object 21 in the elliptical frame is blocked by the trees on the road, and the body part of the target object is blocked by the vehicle next to it. Therefore, the point cloud data collected by the acquisition device is as follows Figure 9 In the elliptical box shown in (b), the incomplete point cloud data for the target object only includes point clouds corresponding to the unobstructed portions of the target object. Furthermore, the unobstructed portions do not include the boundary information of the target object, making it difficult to determine the key long side, key wide side, and key high side in the incomplete point cloud data for the target object. Therefore, step S602 is not performed. Otherwise, step S602 is performed to determine the extended constraint information for the point cloud bounding box.
[0134] S603: Marking extended constraint information on the point cloud bounding box.
[0135] In the embodiment of the present application, marking the extended constraint information on the point cloud bounding box can be understood as adding the extended constraint information to the marked information, for example, marking at least one of the key long side, key wide side and key high side on the marked information, or marking the key vertex and at least one of the key long side, key wide side and key high side, or the key vertex and at least one of the key long side, key wide side and key high side corresponding to the extended direction. Figure 8 (b) and Figure 11 As shown, arrows A3, A4 and A5 intersecting at the key vertex d1 are marked on the point cloud bounding box 30, and each of arrows A3, A4 and A5 is the expansion direction corresponding to the key long side L1, the key wide side L2 and the key high side L3 respectively.
[0136] exist Figure 6 In the described method, the incomplete point cloud data corresponding to the target object is annotated based on the image of the target object, thereby obtaining the annotation information of the point cloud bounding box and the moving direction of the target object. Then, the extended constraint information of the point cloud bounding box is determined based on the moving direction of the target object and / or the z direction of the point cloud bounding box, and the extended constraint information is annotated on the point cloud bounding box, so that the target bounding box that meets the actual needs can be obtained based on the extended constraint information, which is convenient for improving the utilization rate of the data. The point cloud bounding box contains incomplete point cloud data, and keeps the z direction corresponding to the height of the point cloud bounding box, and the z axis perpendicular to the horizontal plane parallel to the direction of the rectangular parallelepiped due to the acquisition angle, which improves the accuracy of the annotated point cloud bounding box. The extended constraint information is determined based on the moving direction and / or the z direction, which is convenient for improving the accuracy of the size processing.
[0137] In one possible example, the target object is a vehicle, and the annotation information includes the vehicle type. The method further includes: determining a first size of the point cloud bounding box according to the vehicle type; and performing size processing on the point cloud bounding box according to the extended constraint information and the first size to obtain a first target bounding box.
[0138] In the embodiment of the present application, the first size may include the length, width, and height of the target object as represented in the image, i.e., the length, width, and height of the cuboid corresponding to the first target bounding box. Since the vehicle is not a cuboid, the target object contained within the cuboid may include excess space. The first size may further include the lengths of the individual sides of the cube corresponding to the target object, etc., which are not limited here.
[0139] It is understandable that different vehicle types have different vehicle sizes. In an embodiment of the present application, the first size of the target object mapped in the point cloud data is determined based on the vehicle type, that is, the size that the point cloud bounding box needs to be processed to obtain. It should be noted that since the target object in the point cloud data is incomplete point cloud data, that is, the image of the target object is incomplete, for most cases, the size processing operation on the point cloud bounding box is an expansion operation. In addition, the first size can also be determined based on the scaling ratio in the annotation information. The scaling ratio can be used to obtain the size relationship between the target object in the point cloud data and the actual target object. Therefore, the size of the target object in the point cloud image can be obtained based on the scaling ratio and the vehicle type, which can improve the accuracy of determining the first size.
[0140] In this embodiment of the present application, the first target bounding box is a bounding box that is resized according to the vehicle type and extended constraint information of the point cloud bounding box. In this example, the first size of the point cloud bounding box is determined based on the vehicle type, and then the point cloud bounding box is resized according to the extended constraint information and the first size to obtain the first target bounding box. This improves the accuracy of the point cloud bounding box resizing and can enhance the authenticity of the first target bounding box.
[0141] The present application does not limit the method for obtaining the first target bounding box. In one possible example, at least one target edge among the key long edge, key wide edge, and key high edge, as well as the target length and target expansion direction of the target edge are determined based on the extension constraint information and the first size; the target edge and the edge corresponding to the target edge in the point cloud bounding box are sized according to the target length and the target expansion direction to obtain the first target bounding box.
[0142] The target edge is the edge that needs to be sized among the key long edge, key wide edge, and key high edge. The target length may include the length of the target edge to be sized, or the length corresponding to the target edge in the first size, etc., which is not limited here. The target extension direction is the extension direction corresponding to the key vertex and the target edge, such as Figure 11 As shown in (a) in the figure, the target expansion direction can be the direction indicated by at least one of the arrows that coincide with line segments L1, L2, and L3, that is, at least one of arrows A3, A4, and A5. The edge corresponding to the target edge in the point cloud bounding box refers to the edge that needs to be resized following the target edge. Since the point cloud bounding box is a rectangular parallelepiped, the edge corresponding to the target edge can be an edge in the rectangular parallelepiped that is parallel to the target edge. It should be noted that when the first size includes the length of each edge in the cube corresponding to the target object, the edge corresponding to the target edge can be another edge that needs to be reduced in size.
[0143] The following example takes the edge corresponding to the target edge as the edge parallel to the target edge. Figure 11As shown in (a) and (b1), assuming that line segment L1 is the key long side and the target side is line segment L1, the sides corresponding to line segment L1 in the point cloud bounding box 30 also include line segments L6, L9, and L10. Based on the first size and the direction corresponding to the key side length (i.e., the direction of arrow A3), the key long side (i.e., line segment L1) and the sides parallel to line segment L11 in the point cloud bounding box 30 (i.e., line segments L6, L9, and L10) are resized to obtain a first target bounding box 33. The point cloud data corresponding to the first target bounding box 33 and the movement direction indicated by arrow A1 are referred to as the target point cloud data.
[0144] like Figure 11 As shown in (a) and (b2) in FIG, assuming that line segment L2 is the key wide side and the target side is line segment L2, the edges corresponding to line segment L2 in the point cloud bounding box 30 also include line segments L5, L11, and L12. Based on the first size and the direction corresponding to the key wide side (i.e., the direction of arrow A4), the key wide side (i.e., line segment L2) and the edges parallel to line segment L2 in the point cloud bounding box 30 (i.e., line segments L5, L11, and L12) are resized to obtain a first target bounding box 34. The point cloud data corresponding to the first target bounding box 34 and the movement direction indicated by arrow A1 are referred to as the target point cloud data.
[0145] like Figure 11 As shown in (a) and (b3) in the figure, assuming that line segment L3 is the key high edge and the target edge is line segment L3, the edges corresponding to line segment L3 in the point cloud bounding box 30 also include line segments L4, L7, and L8. Based on the first size and the direction corresponding to the key high edge (i.e., the direction of arrow A5), the key high edge (i.e., line segment L3) and the edges parallel to line segment L3 in the point cloud bounding box 30 (i.e., line segments L4, L7, and L8) are resized to obtain a first target bounding box 35. The point cloud data corresponding to the first target bounding box 35 and the movement direction indicated by arrow A1 are referred to as the target point cloud data.
[0146] It should be noted that the above examples each use one target edge for size processing. In actual size processing, there may be two or three target edges. The target edges and the corresponding target edges are sized in accordance with the corresponding target sizes. This will not be described in detail here. The embodiment of this application uses a vehicle as an example for illustration. The processing method of the point cloud bounding box of other object types (for example, aircraft type, ship type, robot type, etc.) can refer to this method and will not be described in detail here.
[0147] It can be understood that in this example, the edges among the key long edges, key wide edges, and key high edges that do not meet the first size are first taken as target edges, and then the target length and target expansion direction of the first size are determined, so that each target edge in the point cloud bounding box and the edge corresponding to each target edge are size-processed according to the target length and target expansion direction, and the first target bounding box that meets the vehicle type is obtained, thereby improving the accuracy of size processing of the point cloud bounding box.
[0148] and Figure 6 For details on the embodiments shown in the examples, please refer to Figure 10 , Figure 10 This is a flow chart of another data processing method provided in an embodiment of the present application. The method is described using an electronic device as an example. The specific process may include the following steps S1001 to S1004, wherein:
[0149] S1001: Annotate incomplete point cloud data of the target object according to the image of the target object to obtain annotation information, wherein the annotation information includes a moving direction of the target object and a point cloud bounding box containing the incomplete point cloud data, a z direction of the point cloud bounding box is parallel to the z axis and to a direction corresponding to a height of the point cloud bounding box, and the z axis is perpendicular to the horizontal plane.
[0150] S1002: Determine extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, wherein the extended constraint information includes at least one of a key long side, a key wide side, and a key high side intersecting at a key vertex.
[0151] S1003: Marking the extended constraint information on the point cloud bounding box to obtain reference point cloud data.
[0152] S1004: Storing reference point cloud data.
[0153] Among them, steps S1001 to S1003 can refer to the description of steps S601 to S603, and will not be repeated here.
[0154] In the embodiment of the present application, the data obtained by annotating the extended constraint information on the point cloud bounding box of the annotated information can be referred to as reference point cloud data. That is, the reference point cloud data includes the annotation information obtained in step S601 or S1001 and the extended constraint information of the point cloud bounding box in the annotation information obtained in step S602 or S1002. Figure 8 As shown in (b), the point cloud image corresponding to the reference point cloud data is marked with the moving direction of the target object, the point cloud bounding box 30 and the extended constraint information.
[0155] exist Figure 10In the described method, the incomplete point cloud data corresponding to the target object is annotated based on the image of the target object, thereby obtaining the annotation information of the point cloud bounding box and the moving direction of the target object. Then, the extended constraint information of the point cloud bounding box is determined based on the moving direction of the target object and / or the z direction of the point cloud bounding box, the extended constraint information is annotated on the point cloud bounding box, reference point cloud data is obtained, and the reference point cloud data is stored, so that the target bounding box that meets the actual needs can be obtained based on the extended constraint information, which is convenient for further improving the utilization rate of the data. The point cloud bounding box contains incomplete point cloud data, and keeps the z direction corresponding to the height of the point cloud bounding box, and is parallel to the z axis perpendicular to the horizontal plane, which can avoid the directional deviation of the cuboid due to the acquisition angle, thereby improving the accuracy of the annotated point cloud bounding box. The extended constraint information is determined based on the moving direction and / or the z direction, which is convenient for improving the accuracy of size processing.
[0156] In one possible example, the target object is a vehicle, and the annotation information also includes the vehicle type. The method also includes: receiving an annotation instruction for the reference point cloud data; determining a second size of the point cloud bounding box based on the annotation instruction and the vehicle type; and performing size processing on the point cloud bounding box based on the extended constraint information and the second size to obtain a second target bounding box.
[0157] In the embodiment of the present application, the labeling instructions may include the size processing accuracy and size requirements corresponding to the point cloud bounding box. It is understandable that different teams have different requirements for size accuracy. For example, Team 1 requires the vehicle size to be accurate to decimeters, while Team 2 requires the vehicle size to be accurate to millimeters. In addition, different teams have different algorithms, and the required target sizes may also be different. For example, Team 1 requires the target size to be 5.2, 4.3, and 2.0 in length, width, and height, respectively, while Team 2 requires the vehicle size to be 5.25, 3.55, and 2.00, etc.
[0158] The labeling instruction may also include identification information of the electronic device that sends the labeling instruction, etc., which is not limited here. After obtaining the second target bounding box, the point cloud data corresponding to the second target bounding box is sent to the electronic device that sends the labeling instruction according to the identification information. The labeling instruction is used to instruct the electronic device to use the incomplete point cloud data in the reference point cloud data. It can be understood that the point cloud bounding box corresponding to the incomplete point cloud data is resized so that the incomplete point cloud data can be used as data. The labeling instruction can be an instruction obtained based on the information entered by the labeler in the electronic device, or it can be an instruction received from other electronic devices, which is not limited here.
[0159] It is understandable that different vehicle types have different vehicle sizes. In an embodiment of the present application, the second size of the target object mapped in the point cloud data, that is, the size that the point cloud bounding box needs to be processed to obtain, is determined based on the annotation instruction and the vehicle type. It should be noted that since the target object in the point cloud data is incomplete point cloud data, that is, the image of the target object is incomplete, for most cases, the size processing operation on the point cloud bounding box is an expansion operation, but there may also be a scaling operation based on different annotation instructions. In addition, the target size can also be determined based on the scaling ratio in the annotation information, and the scaling ratio can be used to obtain the size relationship between the target object in the point cloud data and the actual target object. Therefore, the size of the target object in the point cloud image can be obtained based on the scaling ratio, annotation instruction and vehicle type, which can improve the accuracy of obtaining the second size.
[0160] In the embodiment of the present application, the second target bounding box is a bounding box obtained by size processing according to the annotation instructions, vehicle type and extended constraint information of the point cloud bounding box. This application does not limit the method for obtaining the second target bounding box. Please refer to the description of the method for obtaining the first target bounding box, which will not be repeated here. The embodiment of the present application takes the target object as an example of a vehicle. The processing method of the point cloud bounding box of other object types (for example, aircraft type, ship type, robot type, etc.) can also refer to this method, which will not be repeated here.
[0161] It can be understood that in this example, when a labeling instruction is received, the second size of the point cloud bounding box can be determined first according to the labeling instruction and the vehicle type, and then the point cloud bounding box can be sized according to the extended constraint information and the second size, so as to obtain a second target bounding box that meets the vehicle type and the labeling instruction, thereby improving the accuracy of the point cloud bounding box sizing and improving the utilization rate of the data.
[0162] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.
[0163] See Figure 12 , Figure 12: is a structural diagram of a data processing device provided in an embodiment of the present application. The data processing device may include a labeling unit 1201, a determination unit 1202, a processing unit 1203, a storage unit 1204 and a communication unit 1205. When the data processing device is an electronic device, the communication unit 1205 can be used to receive information collected by an acquisition device, or receive labeling instructions sent by other electronic devices, and send data such as a target bounding box or point cloud data containing a target bounding box obtained after data processing to other electronic devices. When the data processing device is an acquisition device, the communication unit 1205 can be used to send data such as a target bounding box or point cloud data containing a target bounding box obtained after data processing to an electronic device. The embodiment of the present application takes the data processing device as an electronic device as an example, and a detailed description of each unit is as follows.
[0164] The labeling unit 1201 is configured to label the incomplete point cloud data of the target object according to the image of the target object to obtain labeling information, wherein the labeling information includes a movement direction of the target object and a point cloud bounding box containing the incomplete point cloud data, wherein the z direction of the point cloud bounding box is parallel to the z axis and to the direction corresponding to the height of the point cloud bounding box, and the z axis is perpendicular to the horizontal plane;
[0165] The determining unit 1202 is configured to determine extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, wherein the extended constraint information includes at least one of a key long side, a key wide side, and a key high side intersecting at a key vertex;
[0166] The labeling unit 1201 is further configured to label the extended constraint information on the point cloud bounding box.
[0167] In a possible example, the determination unit 1202 is specifically used to determine the vertex confidence of each vertex in the point cloud bounding box, wherein the vertex confidence is used to describe the probability that the vertex is a vertex of the object bounding box of the target object; determine the vertex corresponding to the maximum value of the vertex confidence as the key vertex; determine the three combined edges in the point cloud bounding box that intersect with the key vertex as three reference edges; and determine at least one of the key long edge, key wide edge, and key high edge among the three reference edges based on the moving direction and / or the z direction.
[0168] In one possible example, the determination unit 1202 is specifically used to determine the vertex confidence of the vertex based on the number of point clouds corresponding to each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the vertex confidence; and / or, determine the vertex confidence of the vertex based on the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance, the greater the vertex confidence, and the acquisition device has collected incomplete point cloud data.
[0169] In a possible example, the determination unit 1202 is specifically used to determine the overall confidence of three combined edges intersecting at each vertex in the point cloud bounding box, wherein the overall confidence is used to describe the probability that the three combined edges are all edges of the object bounding box of the target object; determine the three combined edges corresponding to the maximum value of the overall confidence as three reference edges; determine the vertex where the three reference edges intersect as the key vertex; and determine at least one of the key long edge, key wide edge and key high edge among the three reference edges based on the moving direction and / or z direction.
[0170] In one possible example, the determination unit 1202 is specifically used to determine the overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box based on the number of point clouds corresponding to the three combined edges intersecting at each vertex in the point cloud bounding box, wherein, when the number of point clouds is larger, the overall confidence is larger; and / or, determine the overall confidence of the three combined edges intersecting at the vertex in the point cloud bounding box based on the distance between each vertex in the point cloud bounding box and the acquisition device, wherein, when the distance is smaller, the overall confidence is larger, and the acquisition device has collected incomplete point cloud data.
[0171] In a possible example, the target object is a vehicle, the annotation information also includes the vehicle type, and the determination unit 1202 is further used to determine the first size of the point cloud bounding box according to the vehicle type; the data processing device also includes a processing unit 1203, which is used to size the point cloud bounding box according to the extended constraint information and the first size to obtain a first target bounding box.
[0172] In a possible example, the processing unit 1203 is specifically used to determine at least one target edge among the key long edge, the key wide edge and the key high edge, as well as the target length and target expansion direction of at least one target edge based on the first size and expansion constraint information; based on the target length and the target expansion direction, the target edge and the edge corresponding to the target edge in the point cloud bounding box are size-processed to obtain a first target bounding box.
[0173] In a possible example, the data processing apparatus further includes: a storage unit 1204 configured to store reference point cloud data obtained by marking the extended constraint information on the point cloud bounding box.
[0174] In a possible example, the target object is a vehicle, the annotation information also includes the vehicle type, and the data processing device also includes a communication unit 1205 and a processing unit 1203, wherein: the communication unit 1205 is used to receive annotation instructions for the reference point cloud data; the determination unit 1202 is also used to determine the second size of the point cloud bounding box according to the annotation instructions and the vehicle type; the processing unit 1203 is used to size the point cloud bounding box according to the extended constraint information and the second size to obtain a second target bounding box.
[0175] It should be noted that the implementation of each unit can also refer to Figure 6 and Figure 10 The corresponding description of the method embodiment shown.
[0176] See Figure 13 , Figure 13 A data processing device provided in an embodiment of the present application includes a processor 1301 , a memory 1302 and a communication interface 1303 , and the processor 1301 , the memory 1302 and the communication interface 1303 are interconnected via a bus 1304 . Figure 12 The related functions implemented by the communication unit 1205 shown can be implemented through the communication interface 1303. Figure 12 The related functions implemented by the storage unit 1204 can be implemented by the memory 1302. Figure 12 The related functions implemented by the marking unit 1201 , the determining unit 1202 and the processing unit 1203 shown can be implemented by the processor 1301 .
[0177] Memory 1302 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). Memory 1302 is used for storing computer programs and data. Communication interface 1303 is used to receive and send data.
[0178] The processor 1301 may be one or more central processing units (CPUs). When the processor 1301 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0179] The processor 1301 of the data processing device is configured to read the computer program code stored in the memory 1302 and perform the following operations:
[0180] Annotating the incomplete point cloud data of the target object according to the image of the target object to obtain annotation information, wherein the annotation information includes a movement direction of the target object and a point cloud bounding box containing the incomplete point cloud data, wherein the z direction of the point cloud bounding box is parallel to the z axis and the direction corresponding to the height of the point cloud bounding box, and the z axis is perpendicular to the horizontal plane;
[0181] Determining extended constraint information of the point cloud bounding box according to the movement direction and / or the z-direction, wherein the extended constraint information includes at least one of a key long side, a key wide side, and a key high side intersecting at the key vertex;
[0182] Annotate the extended constraint information on the point cloud bounding box.
[0183] In a possible example, in determining the extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, the processor 1301 is specifically configured to perform the following operations:
[0184] Determine a vertex confidence for each vertex in the point cloud bounding box, wherein the vertex confidence is used to describe the probability that the vertex is a vertex of the object bounding box of the target object;
[0185] Determine the vertex corresponding to the maximum vertex confidence as the key vertex;
[0186] Determine the three combined edges that intersect with the key vertex in the point cloud bounding box as three reference edges;
[0187] At least one of the key long side, the key wide side, and the key high side among the three reference sides is determined according to the moving direction and / or the z direction.
[0188] In a possible example, in determining the vertex confidence of each vertex in the point cloud bounding box, the processor 1301 is specifically configured to perform the following operations:
[0189] The vertex confidence of the vertex is determined according to the number of point clouds corresponding to each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the vertex confidence; and / or, the vertex confidence of the vertex is determined according to the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance, the greater the vertex confidence, and the acquisition device has collected incomplete point cloud data.
[0190] In a possible example, in determining the extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, the processor 1301 is specifically configured to perform the following operations:
[0191] Determine the overall confidence of the three combined edges that intersect at each vertex in the point cloud bounding box, where the overall confidence is used to describe the probability that the three combined edges are all edges of the object bounding box of the target object;
[0192] Determine the three combined edges corresponding to the maximum value of the overall confidence as the three reference edges;
[0193] Determine the vertex where the three reference edges intersect as the key vertex;
[0194] At least one of the key long side, the key wide side, and the key high side among the three reference sides is determined according to the moving direction and / or the z direction.
[0195] In one possible example, in determining the overall confidence of three combined edges intersecting each vertex in the point cloud bounding box, the processor 1301 is specifically configured to perform the following operations:
[0196] The overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box is determined according to the number of point clouds corresponding to the three combined edges, wherein the overall confidence is greater when the number of point clouds is greater; and / or, the overall confidence of the three combined edges intersecting at the vertex in the point cloud bounding box is determined according to the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance is, the greater the overall confidence is, and the acquisition device has collected incomplete point cloud data.
[0197] In a possible example, the target object is a vehicle, and the annotation information also includes the vehicle type. The processor 1301 is further configured to perform the following operations:
[0198] Determine a first size of the point cloud bounding box according to the vehicle type;
[0199] The point cloud bounding box is resized according to the extended constraint information and the first size to obtain a first target bounding box.
[0200] In a possible example, in terms of performing size processing on the point cloud bounding box according to the extended constraint information and the first size to obtain the first target bounding box, the processor 1301 is specifically configured to perform the following operations:
[0201] Determining at least one target edge among the key long edge, the key wide edge, and the key high edge, and a target length and a target expansion direction of the at least one target edge according to the expansion constraint information and the first size;
[0202] According to the target length and the target expansion direction, the target edge and the edge corresponding to the target edge in the point cloud bounding box are resized to obtain a first target bounding box.
[0203] In a possible example, the processor 1301 is further configured to perform the following operations:
[0204] The reference point cloud data obtained by marking the extended constraint information on the point cloud bounding box is stored.
[0205] In a possible example, the target object is a vehicle, and the annotation information also includes the vehicle type. The processor 1301 is further configured to perform the following operations:
[0206] receiving a labeling instruction for reference point cloud data;
[0207] Determine a second size of the point cloud bounding box according to the annotation instruction and the vehicle type;
[0208] The point cloud bounding box is resized according to the extended constraint information and the second size to obtain a second target bounding box.
[0209] It should be noted that the implementation of each operation can also refer to Figure 6 and Figure 10 The corresponding description of the method embodiment shown.
[0210] The embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed on an electronic device, Figure 6 and Figure 10 The method flow shown is realized.
[0211] The embodiment of the present application further provides a computer program product, which, when executed on an electronic device, Figure 6 and Figure 10 The method flow shown is realized.
[0212] The embodiment of the present application also provides a chip, including a processor, for calling and executing instructions stored in the memory from the memory, so that the terminal device equipped with the chip executes Figure 6 and Figure 10 The method shown.
[0213] The embodiment of the present application also provides another chip, which can be a chip in a terminal device or an access network device. The chip includes: an input interface, an output interface and a processing circuit. The input interface, the output interface and the circuit are connected through an internal connection path. The processing circuit is used to execute Figure 6 and Figure 10 The method shown.
[0214] The embodiment of the present application also provides another chip, including: an input interface, an output interface, a processor, and optionally, a memory, wherein the input interface, the output interface, the processor and the memory are connected through an internal connection path, and the processor is used to execute the code in the memory. When the code is executed, the processor is used to execute Figure 6 and Figure 10 The method shown.
[0215] The embodiment of the present application further provides a chip system, which includes at least one processor, a memory, and an interface circuit. The memory, the transceiver, and the at least one processor are interconnected via a line. A computer program is stored in the at least one memory. When the computer program is executed by the processor, Figure 6 and Figure 10 The method flow shown is realized.
[0216] In summary, by implementing the embodiments of the present application, extended constraint information for size processing is added to the point cloud bounding box, so that a target bounding box that meets actual needs can be obtained based on the extended constraint information, thereby improving data utilization. The point cloud bounding box contains incomplete point cloud data, and maintains the z direction corresponding to the height of the point cloud bounding box, as well as the z axis perpendicular to the horizontal plane. This can avoid directional deviation of the cuboid due to the acquisition angle, thereby improving the accuracy of annotating the point cloud bounding box. The extended constraint information is determined based on the movement direction and / or the z direction, which facilitates improving the accuracy of size processing.
[0217] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program or computer program-related hardware. The computer program can be stored in a computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing computer program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A data processing method, characterized in that: include: Annotating the incomplete point cloud data of the target object according to the image of the target object to obtain annotation information, wherein the annotation information includes a movement direction of the target object and a point cloud bounding box containing the incomplete point cloud data, wherein a z direction of the point cloud bounding box is parallel to a z axis and a direction corresponding to a height of the point cloud bounding box, and the z axis is perpendicular to a horizontal plane; Determining extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, wherein the extended constraint information includes at least one of a key long side, a key wide side, and a key high side intersecting at a key vertex, the key vertex being used to describe a vertex in the point cloud bounding box that is closest to a vertex in the object bounding box of the target object, the key long side being a side corresponding to the length of the object bounding box, the key wide side being a side corresponding to the width of the object bounding box, and the key high side being a side corresponding to the height of the object bounding box, the z-direction of the object bounding box being parallel to the z-axis, and the direction corresponding to the height of the object bounding box being parallel to the z-axis, and the object type of the target object and the moving direction being used to determine the length or width of the object bounding box; The extended constraint information is marked on the point cloud bounding box.
2. The data processing method according to claim 1, wherein: The determining of the extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction includes: Determining a vertex confidence of each vertex in the point cloud bounding box, wherein the vertex confidence is used to describe the probability that the vertex is a vertex of the object bounding box of the target object; Determine the vertex corresponding to the maximum value of the vertex confidence as the key vertex; Determining three combined edges in the point cloud bounding box that intersect with the key vertex as three reference edges; At least one of the key long side, the key wide side, and the key high side among the three reference sides is determined according to the moving direction and / or the z direction.
3. The data processing method according to claim 2, characterized in that: Determining the vertex confidence of each vertex in the point cloud bounding box includes: Determining a vertex confidence of each vertex according to the number of point clouds corresponding to each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the vertex confidence; and / or, The vertex confidence of each vertex in the point cloud bounding box is determined according to the distance between the vertex and the acquisition device, wherein the smaller the distance is, the greater the vertex confidence is, and the acquisition device has acquired the incomplete point cloud data.
4. The data processing method according to claim 1, wherein: The determining of the extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction includes: Determining an overall confidence of three combined edges intersecting at each vertex in the point cloud bounding box, wherein the overall confidence is used to describe the probability that the three combined edges are all edges of the object bounding box of the target object; Determine three combined edges corresponding to the maximum value of the overall confidence as three reference edges; Determine the vertex where the three reference edges intersect as the key vertex; At least one of the key long side, the key wide side, and the key high side among the three reference sides is determined according to the moving direction and / or the z direction.
5. The data processing method according to claim 4, characterized in that: Determining the overall confidence of three combined edges intersecting at each vertex in the point cloud bounding box includes: Determining the overall confidence of the three combined edges according to the number of point clouds corresponding to the three combined edges intersecting at each vertex in the point cloud bounding box, wherein the greater the number of point clouds, the greater the overall confidence; and / or, The overall confidence of the three combined edges in the point cloud bounding box that intersect at the vertex is determined based on the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance is, the greater the overall confidence, and the acquisition device has acquired the incomplete point cloud data.
6. The data processing method according to any one of claims 1 to 5, characterized in that: The target object is a vehicle, the annotation information further includes a vehicle type, and the method further includes: Determining a first size of the point cloud bounding box according to the vehicle type; The point cloud bounding box is resized according to the extended constraint information and the first size to obtain a first target bounding box.
7. The data processing method according to claim 6, characterized in that: The step of performing size processing on the point cloud bounding box according to the extended constraint information and the first size to obtain a first target bounding box includes: Determining at least one target side among the key long side, the key wide side, and the key high side, and a target length and a target expansion direction of the at least one target side according to the expansion constraint information and the first size; According to the target length and the target expansion direction, the target edge and the edge corresponding to the target edge in the point cloud bounding box are resized to obtain a first target bounding box.
8. The data processing method according to any one of claims 1 to 5, characterized in that: The method further comprises: The reference point cloud data obtained by marking the extended constraint information on the point cloud bounding box is stored.
9. The data processing method according to claim 8, characterized in that: The target object is a vehicle, the annotation information further includes a vehicle type, and the method further includes: receiving a labeling instruction for the reference point cloud data; determining a second size of the point cloud bounding box according to the labeling instruction and the vehicle type; The point cloud bounding box is resized according to the extended constraint information and the second size to obtain a second target bounding box.
10. A data processing device, characterized in that: include: a labeling unit, configured to label the incomplete point cloud data of the target object according to the image of the target object to obtain labeling information, wherein the labeling information includes a movement direction of the target object and a point cloud bounding box containing the incomplete point cloud data, wherein a z direction of the point cloud bounding box is parallel to a z axis and a direction corresponding to a height of the point cloud bounding box, and the z axis is perpendicular to a horizontal plane; a determining unit, configured to determine extended constraint information of the point cloud bounding box according to the moving direction and / or the z-direction, wherein the extended constraint information includes at least one of a key long side, a key wide side, and a key high side intersecting at a key vertex, the key vertex being used to describe a vertex in the point cloud bounding box that is closest to a vertex in the object bounding box of the target object, the key long side being a side corresponding to the length of the object bounding box, the key wide side being a side corresponding to the width of the object bounding box, and the key high side being a side corresponding to the height of the object bounding box; The labeling unit is further configured to label the extended constraint information on the point cloud bounding box.
11. The data processing device according to claim 10, characterized in that The determination unit is specifically used to determine the vertex confidence of each vertex in the point cloud bounding box, wherein the vertex confidence is used to describe the probability that the vertex is a vertex of the object bounding box of the target object; determine the vertex corresponding to the maximum value of the vertex confidence as the key vertex; determine three combined edges in the point cloud bounding box that intersect with the key vertex as three reference edges; and determine at least one of the key long edge, the key wide edge and the key high edge among the three reference edges according to the moving direction and / or the z direction.
12. The data processing device according to claim 11, characterized in that The determination unit is specifically used to determine the vertex confidence of the vertex according to the number of point clouds corresponding to each vertex in the point cloud bounding box, wherein, when the number of point clouds is larger, the vertex confidence is larger; and / or, determine the vertex confidence of the vertex according to the distance between each vertex in the point cloud bounding box and the acquisition device, wherein, when the distance is smaller, the vertex confidence is larger, and the acquisition device has acquired the incomplete point cloud data.
13. The data processing device according to claim 10, characterized in that The determination unit is specifically used to determine the overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box, wherein the overall confidence is used to describe the probability that the three combined edges are all edges of the object bounding box of the target object; determine the three combined edges corresponding to the maximum value of the overall confidence as three reference edges; determine the vertex where the three reference edges intersect as the key vertex; and determine at least one of the key long edge, the key wide edge and the key high edge among the three reference edges according to the moving direction and / or the z direction.
14. The data processing device according to claim 10, characterized in that The determination unit is specifically used to determine the overall confidence of the three combined edges intersecting at each vertex in the point cloud bounding box according to the number of point clouds corresponding to the three combined edges, wherein the overall confidence is greater when the number of point clouds is greater; and / or, determine the overall confidence of the three combined edges intersecting at the vertex in the point cloud bounding box according to the distance between each vertex in the point cloud bounding box and the acquisition device, wherein the smaller the distance is, the greater the overall confidence is, and the acquisition device has acquired the incomplete point cloud data.
15. The data processing device according to any one of claims 10 to 14, characterized in that: The target object is a vehicle, the annotation information also includes a vehicle type, and the determination unit is further used to determine a first size of the point cloud bounding box based on the vehicle type; the data processing device also includes a processing unit for performing size processing on the point cloud bounding box according to the extended constraint information and the first size to obtain a first target bounding box.
16. The data processing device according to claim 15, characterized in that The processing unit is specifically used to determine at least one target edge among the key long edge, the key wide edge and the key high edge according to the first size and the expansion constraint information, as well as the target length and target expansion direction of the at least one target edge; according to the target length and the target expansion direction, perform size processing on the target edge and the edge corresponding to the target edge in the point cloud bounding box to obtain a first target bounding box.
17. The data processing device according to any one of claims 10 to 14, characterized in that: The data processing device further includes: A storage unit is used to store the reference point cloud data obtained by marking the extended constraint information on the point cloud bounding box.
18. The data processing device according to claim 17, characterized in that The target object is a vehicle, the annotation information further includes a vehicle type, and the data processing device further includes a communication unit and a processing unit, wherein: The communication unit is configured to receive a labeling instruction for the reference point cloud data; The determining unit is further configured to determine a second size of the point cloud bounding box according to the labeling instruction and the vehicle type; The processing unit is configured to perform size processing on the point cloud bounding box according to the extended constraint information and the second size to obtain a second target bounding box.
19. A data processing device, characterized in that: The method comprises a processor and a memory connected to the processor, wherein the memory is used to store one or more programs and is configured to be executed by the processor, wherein the programs include instructions for executing the steps in the method according to any one of claims 1 to 9.
20. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the data processing method according to any one of claims 1 to 9.
21. A computer program product, characterized in that The computer program product is used to store a computer program, and when the computer program is run on a computer, the computer is caused to run the data processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Marking method and device for desultory points in point cloud data, electronic equipment and storage medium
CN108648156A
Three-dimensional object detection method based on view cone point cloud
CN109523552A