A point cloud processing method, apparatus, computing device, and storage medium

By employing point cloud processing methods in autonomous driving systems and optimizing point cloud frames by combining shape models and object states, the robustness problem of object tracking and shape reconstruction in outdoor scenes is solved, achieving more accurate pose estimation and 3D reconstruction.

CN115601386BActive Publication Date: 2026-07-03BEIJING TUSEN ZHITU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING TUSEN ZHITU TECH CO LTD
Filing Date
2021-07-07
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies lack robustness in object tracking and shape reconstruction in outdoor autonomous driving scenarios, especially due to missing observation data and noise, leading to inaccurate pose estimation and 3D reconstruction.

Method used

By employing a point cloud processing method, object point clouds are extracted from the object state in the initial point cloud frame, and shape codes are calculated using a shape model. The object state and shape codes in other frames of the point cloud sequence are then combined to optimize the object state and shape codes, thereby achieving joint optimization of object tracking and shape reconstruction.

Benefits of technology

It improves the accuracy and fidelity of object pose estimation and 3D reconstruction, enabling efficient object tracking and shape reconstruction in scenarios such as autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601386B_ABST
    Figure CN115601386B_ABST
Patent Text Reader

Abstract

This disclosure provides a point cloud processing method, apparatus, computing device, and storage medium. The point cloud processing method includes: acquiring object states in an initial point cloud frame; extracting object point clouds from the initial point cloud frame based on the object states; calculating a shape code of the initial point cloud frame based on a shape model and the object point clouds in the initial point cloud frame, wherein the shape code is used to characterize the surface shape of the object; and calculating the object states and / or shape codes of corresponding objects in other point cloud frames of the point cloud sequence based on the shape code and object states of the initial point cloud frame. Embodiments of this disclosure combine object tracking and shape reconstruction based on a shape model, improving the accuracy of object tracking and the fidelity of shape reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer vision, and more particularly to a point cloud processing method, apparatus, computing device, and storage medium. Background Technology

[0002] 3D object tracking and shape reconstruction are crucial tasks in computer vision and essential components of autonomous driving perception systems. Single-object tracking in 3D involves estimating the object's pose at all time points, given a series of continuous visual observations and its initial pose (position and orientation). Shape reconstruction involves recovering the object's geometric information (3D shape) based on observations. However, observation data from outdoor scenes in autonomous driving systems is often incomplete and noisy, affecting the robustness of pose estimation and 3D reconstruction. Summary of the Invention

[0003] Embodiments of this disclosure provide a point cloud processing method, apparatus, computing device, and storage medium to improve the accuracy and fidelity of object pose estimation and 3D reconstruction in point cloud data.

[0004] To achieve the above objectives, the embodiments of this disclosure adopt the following technical solutions:

[0005] A first aspect of this disclosure provides a point cloud processing method, comprising: acquiring an object state in an initial point cloud frame; extracting a corresponding object point cloud from the initial point cloud frame based on the object state; calculating a shape code of the corresponding object in the initial point cloud frame based on a shape model and the object point cloud in the initial point cloud frame, the shape code being used to characterize the surface shape of the object; and calculating the object state and / or shape code of the corresponding object in other point cloud frames of the point cloud sequence based on the shape code and the object state.

[0006] A second aspect of this disclosure provides a point cloud processing apparatus, comprising: an object state acquisition module, adapted to acquire the object state in an initial point cloud frame; an object point cloud extraction module, adapted to extract the corresponding object point cloud from the initial point cloud frame based on the object state in the initial point cloud frame; a shape encoding calculation module, adapted to calculate the shape encoding of the corresponding object in the initial point cloud frame based on a shape model and the object point cloud in the initial point cloud frame, wherein the shape encoding is used to characterize the surface shape of the object; and an iteration module, adapted to calculate the object state and / or shape encoding of the corresponding object in other point cloud frames of the point cloud sequence based on the shape encoding and the object state.

[0007] A third aspect of this disclosure provides a computing device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor; wherein, when the processor runs the computer program, it performs the point cloud processing method as described above.

[0008] A fourth aspect of this disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the point cloud processing method described above.

[0009] The point cloud processing scheme provided in this disclosure, given the object pose of an object in the initial frame of a point cloud sequence, can calculate the object's pose and shape in subsequent frames based on a shape model and continuous observations. Furthermore, this disclosure enables joint optimization of object tracking and shape reconstruction, improving shape reconstruction performance while using more accurate shapes to estimate object pose, thus achieving joint optimization. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A structural diagram of a vehicle 100 provided in an embodiment of this disclosure;

[0012] Figure 2 A flowchart of a point cloud processing method 200 provided in this embodiment of the disclosure;

[0013] Figure 3A and 3B These are schematic diagrams of the DeepSDF model and the occupancy network model, respectively, according to embodiments of this disclosure.

[0014] Figure 4 This is a schematic diagram illustrating the combined object tracking and shape reconstruction in an embodiment of this disclosure;

[0015] Figure 5 This is a diagram illustrating the vehicle state estimation and 3D reconstruction results in an embodiment of this disclosure.

[0016] Figure 6 A flowchart of another point cloud processing method 600 provided in this disclosure embodiment;

[0017] Figure 7 A structural diagram of a point cloud processing device 700 provided in an embodiment of this disclosure;

[0018] Figure 8 This is a structural diagram of a computing device 800 provided in an embodiment of the present disclosure. Detailed Implementation

[0019] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this disclosure.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] To enable those skilled in the art to better understand this disclosure, some technical terms appearing in the embodiments of this disclosure are explained below:

[0022] Point cloud: Data about the surrounding environment collected by radar (such as lidar, millimeter-wave radar, etc.) is represented by a set of sparse three-dimensional spatial points.

[0023] Frame: The measurement data received by a sensor after completing one observation. For example, a frame of data from a camera is an image, and a frame of data from a lidar is a set of laser point clouds.

[0024] Point cloud sequence: A series of consecutive point clouds over a period of time, which are point clouds collected by lidar over a continuous period of time.

[0025] Point cloud registration: By aligning two or more frames of point clouds of the same object, the motion of the object between the two or more frames of point clouds is calculated.

[0026] Objects: The target objects in each frame of data, which can include static and dynamic objects, such as pedestrians, vehicles, animals, obstacles, traffic lights, road signs, etc.

[0027] State: The position information of an object in the three-dimensional world, including position parameters and / or angle parameters, that is, position and / or orientation in pose.

[0028] State estimation: Calculating the state of an object in the three-dimensional world using known information.

[0029] Target detection: using algorithms to find the location of target objects in raw sensor data, typically represented by rectangles or cuboids to indicate the position of the object in 2D or 3D space.

[0030] Object tracking: Given sensor input data and a given target object over a period of time, calculate the state of the given target object at each moment.

[0031] 3D reconstruction: obtaining a 3D model of an object, usually expressed as a dense point cloud or CAD model.

[0032] Figure 1 This is a schematic diagram of a vehicle 100 in which the various technologies disclosed herein can be implemented. Vehicle 100 can be a car, truck, motorcycle, bus, boat, airplane, helicopter, lawnmower, excavator, snowmobile, aircraft, recreational vehicle, amusement park vehicle, farm equipment, construction equipment, tram, golf cart, train, trolleybus, or other vehicle. Vehicle 100 can operate fully or partially in an autonomous driving mode. In autonomous driving mode, vehicle 100 can control itself; for example, vehicle 100 can determine the current state of the vehicle and the current state of the environment in which the vehicle is located, determine the predicted behavior of at least one other vehicle in the environment, determine the trust level corresponding to the probability that the at least one other vehicle will perform the predicted behavior, and control vehicle 100 itself based on the determined information. In autonomous driving mode, vehicle 100 can operate without human interaction.

[0033] Vehicle 100 may include various vehicle systems, such as drive system 142, sensor system 144, control system 146, user interface system 148, control computer system 150, and communication system 152. Vehicle 100 may include more or fewer systems, and each system may include multiple units. Furthermore, each system and unit of vehicle 100 may be interconnected. For example, control computer system 150 is capable of data communication with one or more of vehicle systems 142-148 and 152. Thus, one or more of the described functions of vehicle 100 may be divided into additional functional components or physical components, or combined into a smaller number of functional components or physical components. In a further example, additional functional components or physical components may be increased to, for example... Figure 1 In the example shown.

[0034] The drive system 142 may include a plurality of operable components (or units) that provide kinetic energy to the vehicle 100. In one embodiment, the drive system 142 may include an engine or electric motor, wheels, a transmission, electronic systems, and a power source (or power source). The engine or electric motor may be any combination of internal combustion engines, electric motors, steam engines, fuel cell engines, propane engines, or other forms of engines or electric motors. In some embodiments, the engine may convert a power source into mechanical energy. In some embodiments, the drive system 142 may include multiple engines or electric motors. For example, a hybrid vehicle may include a gasoline engine and an electric motor, or other configurations may be included.

[0035] The wheels of vehicle 100 can be standard wheels. The wheels of vehicle 100 can be of various forms, including single-wheel, two-wheel, three-wheel, or four-wheeled, such as the four wheels on a car or truck. Other numbers of wheels are also possible, such as six or more wheels. One or more wheels of vehicle 100 can be operated to rotate in a different direction than the other wheels. A wheel can be at least one wheel fixedly connected to a transmission. The wheel can include a combination of metal and rubber, or other materials. The transmission can include units operable to transmit mechanical power from the engine to the wheels. For this purpose, the transmission can include a gearbox, clutch, differential gears, and driveshaft. The transmission can also include other units. The driveshaft can include one or more axles that match the wheels. The electronic system can include units for transmitting or controlling electronic signals of vehicle 100. These electronic signals can be used to activate multiple lights, multiple servo mechanisms, multiple electric motors, and other electronic drives or controls in vehicle 100. The power source can be an energy source that provides power to the engine or electric motor, either wholly or partially. That is, the engine or electric motor is capable of converting the power source into mechanical energy. For example, the power source may include gasoline, petroleum, petroleum-based fuels, propane, other compressed gaseous fuels, ethanol, fuel cells, solar panels, batteries, and other electrical energy sources. The power source may optionally include any combination of a fuel tank, battery, capacitor, or flywheel. The power source may also provide energy to other systems of vehicle 100.

[0036] Sensor system 144 may include multiple sensors for sensing information about the environment and conditions of vehicle 100. For example, sensor system 144 may include an inertial measurement unit (IMU), a global positioning system (GPS) transceiver, a radar (RADAR) unit, a laser rangefinder / LIDAR unit (or other distance measurement device), acoustic sensors, and cameras or image capture devices. Sensor system 144 may include multiple sensors for monitoring vehicle 100 (e.g., oxygen (O2) monitor, fuel gauge sensor, engine oil pressure sensor, etc.). Other sensors may also be configured. One or more sensors included in sensor system 144 may be driven individually or collectively to update the position, orientation, or both of the sensors.

[0037] The IMU may include a combination of sensors (e.g., accelerometers and gyroscopes) for sensing changes in the position and orientation of vehicle 100 based on inertial acceleration. The GPS transceiver may be any sensor used to estimate the geographic location of vehicle 100. For this purpose, the GPS transceiver may include a receiver / transmitter to provide position information of vehicle 100 relative to the Earth. It should be noted that GPS is an example of a Global Navigation Satellite System; therefore, in some embodiments, the GPS transceiver may be replaced with a BeiDou Navigation Satellite System transceiver or a Galileo Navigation Satellite System transceiver. The radar unit may use radio signals to sense objects in the environment in which vehicle 100 is located. In some embodiments, in addition to sensing objects, the radar unit may also be used to sense the speed and direction of travel of objects approaching vehicle 100. The laser rangefinder or LIDAR unit (or other distance measurement device) may be any sensor that uses lasers to sense objects in the environment in which vehicle 100 is located. In one embodiment, the laser rangefinder / LIDAR unit may include a laser source, a laser scanner, and a detector. The laser rangefinder / LIDAR unit is used to operate in continuous (e.g., using heterodyne detection) or discontinuous detection modes. The camera may include means for capturing multiple images of the environment in which the vehicle 100 is located. The camera may be a still image camera or a video camera.

[0038] The control system 146 is used to control the operation of the vehicle 100 and its components (or units). Accordingly, the control system 146 may include various units, such as a steering unit, a power control unit, a braking unit, and a navigation unit.

[0039] The steering unit may be a combination of mechanisms for adjusting the forward direction of vehicle 100. A power control unit (e.g., a throttle) may be used to control the engine speed, thereby controlling the speed of vehicle 100. The braking unit may include a combination of mechanisms for decelerating vehicle 100. The braking unit may utilize friction to decelerate the vehicle in a standard manner. In other embodiments, the braking unit may convert the kinetic energy of the wheels into electrical current. The braking unit may also take other forms. The navigation unit may be any system that determines a driving path or route for vehicle 100. The navigation unit may also dynamically update the driving path as vehicle 100 travels. The control system 146 may also additionally or optionally include other components (or units) not shown or described.

[0040] User interface system 148 can be used to allow vehicle 100 to interact with external sensors, other vehicles, other computer systems, and / or the user of vehicle 100. For example, user interface system 148 may include standard visual display devices (e.g., plasma displays, liquid crystal displays (LCDs), touchscreen displays, head-mounted displays, or other similar displays), speakers or other audio output devices, microphones or other audio input devices. For example, user interface system 148 may also include navigation interfaces and interfaces for controlling the internal environment of vehicle 100 (e.g., temperature, fan, etc.).

[0041] Communication system 152 can provide vehicle 100 with a means of communicating with one or more devices or other vehicles in the vicinity. In an exemplary embodiment, communication system 152 can communicate with one or more devices directly or through a communication network. Communication system 152 can be, for example, a wireless communication system. For example, the communication system can use 3G cellular communication (e.g., CDMA, EVDO, GSM / GPRS) or 4G cellular communication (e.g., WiMAX or LTE), and can also use 5G cellular communication. Optionally, the communication system can communicate with a wireless local area network (WLAN) (e.g., using...). In some embodiments, the communication system 152 can communicate directly with one or more devices or other vehicles in the vicinity, for example, using infrared light. Or ZigBee. Other wireless protocols, such as various vehicular communication systems, are also within the scope of this application. For example, the communication system may include one or more Dedicated Short Range Communication (DSRC) devices, V2V devices, or V2X devices that conduct public or private data communication with vehicles and / or roadside stations.

[0042] The control computer system 150 can control some or all of the functions of the vehicle 100. The autonomous driving control unit in the control computer system 150 can be used to identify, assess, and avoid or traverse potential obstacles in the environment in which the vehicle 100 is located. Typically, the autonomous driving control unit can be used to control the vehicle 100 without a driver or to assist a driver in controlling the vehicle. In some embodiments, the autonomous driving control unit is used to combine data from a GPS transceiver, radar data, LiDAR data, camera data, and data from other vehicle systems to determine the driving path or trajectory of the vehicle 100. The autonomous driving control unit can be activated to enable the vehicle 100 to be driven in autonomous driving mode.

[0043] The control computer system 150 may include at least one processor (which may include at least one microprocessor), which executes processing instructions (i.e., machine-executable instructions) stored in a non-volatile computer-readable medium (e.g., a data storage device or memory). The memory stores at least one machine-executable instruction, and the processor executes this instruction to implement functions including a map engine, a positioning module, a perception module, a navigation or path module, and an automatic control module. The map engine and positioning module provide map and positioning information. The perception module perceives objects in the vehicle's environment based on information acquired by the sensor system and map information provided by the map engine. The navigation or path module plans a driving path for the vehicle based on the processing results of the map engine, positioning module, and perception module. The automatic control module parses and converts the decision information input from modules such as the navigation or path module into control commands for the vehicle control system, and sends these commands to corresponding components in the vehicle control system via an in-vehicle network (e.g., an in-vehicle electronic network system implemented via CAN bus, local area network, multimedia orientation system transmission, etc.) to achieve automatic vehicle control; the automatic control module can also obtain information about various components in the vehicle via the in-vehicle network.

[0044] The control computer system 150 may also be multiple computing devices that distribute and control components or systems of the vehicle 100. In some embodiments, the memory may contain processing instructions (e.g., program logic) that are executed by a processor to implement various functions of the vehicle 100. In one embodiment, the control computer system 150 is capable of data communication with systems 142, 144, 146, 148, and / or 152. Interfaces within the control computer system facilitate data communication between the control computer system 150 and systems 142, 144, 146, 148, and 152.

[0045] The memory may also include other instructions, including instructions for sending data, instructions for receiving data, instructions for interaction, or instructions for controlling the drive system 140, sensor system 144, control system 146, or user interface system 148.

[0046] In addition to storing processing instructions, the memory can store various types of information or data, such as image processing parameters, road maps, and route information. This information can be used by the vehicle 100 and the control computer system 150 while the vehicle 100 is operating in automatic, semi-automatic, and / or manual mode.

[0047] Although the autonomous driving control unit is shown as separate from the processor and memory, it should be understood that in some embodiments, some or all of the functions of the autonomous driving control unit may be implemented using program code instructions residing in one or more memories (or data storage devices) and executed by one or more processors, and in some cases, the autonomous driving control unit may be implemented using the same processor and / or memory (or data storage device). In some embodiments, the autonomous driving control unit may be implemented at least in part using various special-purpose circuit logics, various processors, various field-programmable gate arrays (“FPGAs”), various application-specific integrated circuits (“ASICs”), various real-time controllers, and hardware.

[0048] The control computer system 150 can control the functions of the vehicle 100 based on inputs received from various vehicle systems (e.g., drive system 142, sensor system 144, and control system 146) or from the user interface system 148. For example, the control computer system 150 can use inputs from the control system 146 to control the steering unit to avoid obstacles detected by the sensor system 144. In one embodiment, the control computer system 150 can be used to control multiple aspects of the vehicle 100 and its systems.

[0049] Although Figure 1 The diagram shows various components (or units) integrated into vehicle 100, one or more of which may be mounted on or separately associated with vehicle 100. For example, a control computer system may exist partially or entirely independent of vehicle 100. Thus, vehicle 100 can exist as separate or integrated device units. The device units constituting vehicle 105 can communicate with each other via wired or wireless communication. In some embodiments, additional components or units may be added to or removed from various systems (e.g., ...). Figure 1 (LiDAR or radar shown).

[0050] As mentioned earlier, current object tracking and shape reconstruction technologies perform poorly in outdoor scenes, and existing solutions often separate object tracking and shape reconstruction, which are actually complementary: tracking an object allows for more observation of the object, resulting in better shape reconstruction; and having a more accurate shape allows for more accurate object tracking.

[0051] To address this, this disclosure proposes a high-fidelity, high-performance object tracking and 3D reconstruction scheme. This scheme can be used in online scenarios for autonomous vehicles, intelligent robots, drones, etc., to track objects and reconstruct shapes in their surroundings. It can also be used in offline scenarios to perform state estimation and shape reconstruction on each frame of a point cloud sequence using computing devices. Of course, any point cloud processing scenario that can apply object state estimation and shape reconstruction may be applicable to the embodiments of this disclosure, and these will not be listed individually here.

[0052] like Figure 2 As shown, this disclosure provides a point cloud processing method 200, including:

[0053] Step S201: Obtain the object state in the initial point cloud frame.

[0054] Step S202: Extract the corresponding object point cloud from the initial point cloud frame according to the object state.

[0055] Step S203: Based on the shape model and object point cloud computing, the shape code of the corresponding object in the initial point cloud frame is used to characterize the surface shape of the object.

[0056] Step S204: Based on the shape encoding and object point cloud computing, calculate the object state and / or shape encoding of the corresponding object in other point cloud frames of the point cloud sequence.

[0057] In some embodiments, the initial point cloud frame in step S201 can be the first frame of the point cloud sequence, a manually selected frame, or a frame from which the target detection algorithm achieves a predetermined accuracy (e.g., selecting the frame with the highest detection accuracy from a predetermined set of frames starting from the first frame as the initial point cloud frame). This disclosure does not impose any limitations on this. The initial point cloud frame contains the state of a given object, which can be obtained through manual annotation or calculated using a localization algorithm. Furthermore, this disclosure does not limit the form and acquisition method of the observation data; for example, it can be based on LiDAR point cloud data, Radar point cloud data, TOF (Time of Flight) data, etc.

[0058] In some embodiments, the object state includes at least one of the object's position parameters and angle parameters. The position parameters include at least one of the object's key points in a spatial coordinate system: a first coordinate, a second coordinate, and a third coordinate. The first, second, and third coordinates may correspond to the x-axis, y-axis, and z-axis coordinates in the spatial coordinate system. The angle parameters include at least one of pitch angle, yaw angle, and roll angle. There may be one or more key points, such as the center point of the object, the center point of a specific part of the object, or a set of these points. For example, for a vehicle, its key points may be the vehicle's center point, the center point of the front of the vehicle, the center points of the two rear wheels, or a set of multiple body points. This invention does not limit the number or location of these key points.

[0059] In some embodiments, step S202, extracting the corresponding object point cloud from the initial point cloud frame based on the object state, includes: extracting the corresponding object point cloud based on the object state and object size in the point cloud frame. Let the object pose be T and the object size be b (including length, width, and height parameters). Based on the object state and size, the object point cloud is extracted from a scene point cloud P in a frame, and this object point cloud is defined as X(T, b) = {x|Tx≤b, x∈P}. For example, this disclosure can pre-identify and remove ground point cloud points, and generate a detection box greater than or equal to the object size based on the current state of the object. The point cloud within this detection box is the object point cloud. In addition, considering that the object size does not change during the tracking process, for simplicity, this disclosure omits b and uses X(T) to represent the object point cloud extracted based on the object state T.

[0060] In some embodiments, step S203, which involves encoding the shape of the corresponding object in the initial point cloud frame based on the shape model and object point cloud computing, includes: obtaining the shape code of the object by fitting the point cloud of an object in the initial point cloud frame onto the surface of the object based on the shape model.

[0061] Here, the shape code z is used to encode the shape of an object, representing the shape of the object's surface. It can be represented as a hidden shape vector, such as a multi-dimensional vector, or a multi-dimensional matrix, but is not limited to these. The shape code represents the shape of the object corresponding to the three-dimensional coordinate point x.

[0062] In some embodiments, the shape model is a signed distance function (SDF), which is a continuous function used to represent the shape of an object. The input is any three-dimensional coordinate point x, and the output is the directed distance s from the three-dimensional coordinate point to the nearest point on the surface of the object. The sign of the value indicates whether the coordinate point is inside the object (negative) or outside the object (positive), i.e., SDF(x) = s.

[0063] Furthermore, the shape model is based on the deep learning-based symbolic distance function (DeepSDF), such as... Figure 3A As shown, the input to this shape model is a 3D spatial coordinate point x and a shape code z, and the output is the directed distance from the coordinate point to the nearest point on the object's surface. For a specific object represented by the shape code, the shape model f... θ It is a function that takes the hidden vector z and the query coordinate x as input and outputs the approximate signed distance value of the shape at this location, i.e., f. θ (z, x) ≈ SDF(x). DeepSDF model f θ (z, x) is a multilayer perceptron parameterized by learning weights θ, where the weights θ make the model f θ It can approximate the symbolic distance function SDF well.

[0064] Deep learning-based symbolic distance functions (SDFs) directly regress symbolic distance values ​​from 3D coordinates using neural networks. A trained network can directly predict the symbolic distance value for any given 3D coordinate. Therefore, by densely sampling in space and querying symbolic distance values, zero isosurfaces (i.e., object surfaces) can be extracted. Intuitively, this shape representation can be understood as performing a binary classification problem in space, with the decision boundary being the shape's surface. DeepSDF can theoretically represent continuous surfaces with arbitrarily high precision and is often used to extract common characteristics of different shapes and embed them into a low-dimensional space.

[0065] Those skilled in the art will understand that, unlike reconstruction on dense point clouds in the object coordinate system of synthetic datasets, the object state in outdoor scenes is often unavailable. Furthermore, unlike shape reconstruction alone, joint tracking and shape reconstruction require estimation of the object state.

[0066] In some embodiments, given the state and size of the frame objects in the initial frame, the joint tracking and shape reconstruction problem of this disclosure can be defined as the following optimization problem:

[0067]

[0068] In this context, the actual directed distance from the point cloud points of an object to the object's surface should be 0, since all point cloud points originate from the object's surface. Therefore, the optimization problem described above is to fit the object's point cloud points to the object's surface as closely as possible, which means moving the point cloud or the object's surface to make the predicted SDF value of the point cloud points as close to 0 as possible. However, in formula (1), the object state and shape encoding are two sets of variables with different scales and different solution spaces, which often leads to poor results when directly solving the formula. Therefore, this disclosure uses separate optimization to solve these two sets of variables.

[0069] In some embodiments, given the object state T0 of the initial point cloud frame in step S203, the shape encoding of the object in the initial point cloud frame can be obtained based on the shape model by fitting the point cloud of an object in the initial point cloud frame onto the surface of the object. Therefore, the point cloud points can be fitted to the object surface by minimizing the SDF value at the point cloud points. Those skilled in the art can determine the minimization function for fitting as needed, and this disclosure does not limit the form of the minimization function.

[0070] In some embodiments, the shape code of the initial point cloud frame can be calculated using the following formula:

[0071]

[0072] Where, argmin represents the value of the variable when the latter expression takes its minimum; smooth_l1 is the loss function, which in this formula represents the error between the predicted SDF value based on (x, z; θ) and the theoretical value 0; λ||z|| 2 λ is a regularization term used to avoid overfitting, and λ is the weight value. Of course, those skilled in the art can choose other loss functions as needed, such as the L2 loss function. This disclosure does not impose specific restrictions on the specific expression of the loss function.

[0073] After initialization, in step S204, the state and / or shape encoding of the corresponding object in subsequent frames can be iteratively optimized based on the shape encoding and object state of the initial frame. Further, the object point cloud of the next point cloud frame can be determined based on the object state of the current point cloud frame, and the object state and / or shape encoding of the corresponding object in the next point cloud frame can be determined based on the shape model. It should be noted that the corresponding object mentioned in this disclosure refers to the same object as in the current point cloud frame. For example, if the current point cloud frame calculates the state and shape of object A, the next point cloud frame will calculate the object state and shape of object A in the next point cloud frame based on the performance of object A in previous frames.

[0074] In some embodiments, the shape codes of other point cloud frames are uniformly calculated with reference to the shape codes of objects in the initial point cloud frame; in other embodiments, the shape code of an object in any (k+1)th point cloud frame is calculated based on the shape code of the kth point cloud frame, where k is an integer greater than or equal to 1; in yet another embodiment, the shape code of an object in the (k+1)th point cloud frame is calculated based on the shape codes of the corresponding objects in one or more sets of observation frames composed of the first (k+1) point cloud frames; in yet another embodiment, this disclosure omits the calculation of the shape codes of objects in other point cloud frames and uniformly uses the shape codes of objects in the initial point cloud frame to calculate the state of other point cloud frames.

[0075] In some embodiments, calculating the object state and / or shape code of other point cloud frames in the point cloud sequence based on the shape code in step S204 includes:

[0076] Based on the shape encoding of the initial point cloud frame and the object point cloud iteratively, the system performs m calculations of the object state for other point cloud frames and n calculations of the shape encoding for other point cloud frames. Here, n is an integer greater than or equal to 1, and m is an integer greater than or equal to 0. That is, the shape encoding of other point cloud frames can remain unchanged, uniformly adopting the shape encoding of the initial point cloud frame; while the object state is calculated based on the continuous changes in the point cloud frames.

[0077] Furthermore, both m and n are integers greater than or equal to 1, meaning that shape encoding calculations will be performed at least once for other point cloud frames. Object state calculations and shape encoding calculations are performed alternately. This alternation can be as follows: first optimize the object state of the current frame, then optimize the shape of the current frame; or first optimize the shape of the current frame, then optimize the state of the current frame; or first continuously optimize the object shape of one or more frames, then continuously optimize the object state of one or more frames; or first continuously optimize the object shape of one or more frames, then continuously optimize the object state of one or more frames; or only calculate the shape encoding of one or more frames, omitting the shape encoding calculations of other frames. Those skilled in the art can set the alternation calculation method for object state and shape optimization as needed, and this disclosure does not impose any restrictions on this.

[0078] Furthermore, this disclosure allows for shape encoding calculation of a point cloud frame every predetermined frame, uniformly using the shape encoding of the object in that frame to calculate the point cloud for the corresponding frame within a cycle starting from that frame. This maintains the accuracy of shape encoding over a period of time while reducing the amount of data computation. For example, if shape encoding calculation is performed every four frames, then the shape encoding will be calculated once for frames 1, 5, and 10, while the shape encoding of frame 1 will be uniformly used between frames 2 and 4, the shape encoding of frame 5 will be uniformly used between frames 6 and 9, and so on.

[0079] Those skilled in the art will understand that, with limited observations, it is difficult to maintain high fidelity in object shape reconstruction, which in turn affects object tracking performance. Therefore, as Figure 4 As shown, this disclosure proposes a unified framework for joint tracking and shape reconstruction of objects. Throughout the tracking process, this disclosure maintains a dynamic, changeable object shape, represented in the form of a shape code. Since the object pose at the initial moment is given, the object shape can be initialized using the point cloud aligned to the first frame based on the 3D reconstruction of the shape model.

[0080] Then, the object can be tracked and its shape reconstructed alternately. First, the object's pose is obtained through tracking. Then, the object's pose is aligned with the point cloud, and the object's shape is updated using the point cloud. This process is repeated until tracking is complete. By utilizing observations from different frames, the algorithm can obtain more object observations and aggregate historical information using shape encoding, thereby reconstructing a more accurate object shape. A more precise object shape will improve the performance of object tracking.

[0081] like Figure 4 As shown, the object shape is initialized using the point cloud data aligned to the first frame. Subsequently, at time t (corresponding to a frame in a series of observation frames), the shape of the point cloud in the current frame is first aligned with that of the previous frame by minimizing the SDF value at the point cloud. Then, the shape of the current frame is updated by similarly minimizing the SDF value at the point cloud (while also updating the shape of the previous frame), making it consistent with historical observations. Thus, given the previous object shape and the observation data at the current time, this disclosure uses joint object tracking and shape update to solve for the object state at the current time and update the object shape, improving the accuracy of state estimation and the fidelity of shape reconstruction.

[0082] Specifically, for a point cloud frame at time t, the object pose is first optimized based on the object shape reconstructed in the previous frame. The previous frame can be the frame preceding the current frame, or a set of one or more point cloud frames including the previous frame. These frames are used as reference frames. Based on the prior knowledge that the SDF value at each point cloud point should be 0 and a pre-trained shape model, the state of the corresponding object in the current frame is calculated using the object shape and object point cloud data in the reference frames. In some embodiments, the object state of the point cloud frame at time t can be optimized using the following formula.

[0083]

[0084] in, This is the shape encoding of the object in the point cloud frame at time t-1, where x is the point cloud X of the object at time t. t Points in (T), It is a one-way Chamfer Distance loss function applied to the object point cloud of the current frame and the aggregated historical point cloud (from time 0 to time t-1), providing more detail for pose estimation. Because the reconstructed shape always tends to be overly smoothed, this loss function can be used to improve object tracking performance. Figure 5 As shown in part (a), the initial point cloud is far from the object and has a high SDF value. By optimizing equation (3), the point cloud is pushed toward the object surface (i.e., the place where the SDF is 0).

[0085] After obtaining the object's state at time t, the algorithm will utilize historical observations to improve the fidelity of the object's shape. There are several methods to determine how to utilize historical observations χ. t ={X0, ..., X t In some embodiments, this disclosure uses a selection function Γ: x→2 χ This represents different selection strategies. The selection function represents the set of observation frames that can be formed within the historical observations, with each set containing at least one point cloud frame. For example, for historical χ... t ={X0, ..., X t The set of observation frames that can be formed includes the set of any single frame, the set of any two frames, ..., the set of any t-1 frames, and the set of complete t frames.

[0086] Then, the shape of the current frame can be optimized based on the shape encoding and point cloud of the same object within this set of one or more observation frames. In some embodiments, the update of the object shape in the point cloud frame at time t can be accomplished by optimizing the shape encoding, for example, by optimizing the shape encoding using the following formula.

[0087]

[0088] Where UΓ(χ) t ) is Γ(χ t The union of ). For example Figure 5 As shown in section (b), given the estimated pose of the object in the current frame, the algorithm will deform the SDF field and fit the boundary of the shape to the point cloud.

[0089] It can be seen that each time the object state or shape encoding is optimized based on the shape model, the required input includes the object state and shape encoding of the previous frame, and the corresponding object point cloud is determined according to the object state.

[0090] In other embodiments, the shape model is an occupancy network model, which is based on a deep learning occupancy function, specifically using a neural network to fit the occupancy function. The occupancy function is a continuous function, also used to represent the shape of an object. The input of the occupancy function is any three-dimensional coordinate point x, and the output is an occupancy probability between 0 and 1, representing the probability that the three-dimensional coordinate point is inside (1) or outside (0) the object.

[0091] The occupancy network model includes an encoder and a decoder. The encoder obtains the corresponding shape code based on the input set of 3D coordinate points, and the decoder outputs the occupancy probability of the 3D coordinate points based on the set of 3D coordinate points and the shape code.

[0092] Similar to DeepSDF, the shape representation of the occupancy network model can be understood as a binary classification problem in space, with the decision boundary being the surface of the shape. The occupancy network model receives specific object observations O (such as point clouds, images, etc.) and processes them through the corresponding encoder. (e.g., PointNet, ResNet, etc.) Generate shape encoding (e.g., hidden shape vector) z, i.e. The decoder, based on the shape encoding and the input set of 3D point coordinates, obtains the probability s, or f, of the corresponding object shape occupying that 3D coordinate. θ (z, x) = s. The specific structure of the occupied network is as follows: Figure 3B As shown, the occupancy network first uses an encoder to generate a hidden shape vector, and then uses a decoder to predict the approximate occupancy probability of a 3D coordinate point.

[0093] In some embodiments, for the occupancy network model, the optimization objective for joint tracking and shape reconstruction is as follows:

[0094]

[0095] Where f is the decoder in the occupancy network, and τ is a predefined adjustable parameter representing the occupancy probability of the object surface.

[0096] In the specific optimization process, the shape encoding initialized in step S203 It can be obtained through the following formula:

[0097]

[0098] The z in the above equation can be calculated by the encoder first, and then fine-tuned by the decoder:

[0099] In step S204, the state optimization of the object at time t can be performed using the following formula:

[0100]

[0101] In step S204, the shape encoding optimization of the object at time t can be performed using the following formula:

[0102]

[0103] Where UΓ(χ) t ) is Γ(χ t The union of ).

[0104] In some embodiments, the shape model is a shape blend, a shape representation based on Mesh Principal Component Analysis (PCA) used by GSNet. A mesh shape representation consists of vertices (a set of 3D coordinate points) and faces (a single mesh contains a set of faces, where each face consists of multiple 3D coordinate points). PCA is a basic shape representation that typically generates a series of topologically consistent (i.e., face-consistent) mesh representations, which are then reduced in dimensionality by running PCA on the mesh vertices. Taking the PCA representation in GSNet as an example: GSNet first aligns the 3D shape with camera observations using the SoffRas renderer, transforming an ellipsoidal mesh into different objects to obtain a topologically consistent mesh representation of the vehicle shape. Then, PCA is used to reduce the dimensionality of the mesh vertices.

[0105] The shape blender, on the other hand, uses a hybrid shape representation obtained by blending multiple mesh master-level analyses. Specifically, GSNet first uses a clustering algorithm (such as K-Means) to cluster all object meshes based on shape similarity, resulting in multiple subsets, for example, four subsets. For each subset, GSNet uses the PCA algorithm to obtain a low-dimensional (e.g., less than or equal to 10-dimensional) shape basis. For a specific input, Shape Blend first classifies each subclass and obtains the corresponding classification probability. Simultaneously, it predicts the PCA coefficient (correlation coefficient) for each subclass. The final shape is obtained by blending the shapes of different subclasses based on the classification probabilities.

[0106] Specifically, the hybrid deformer obtains a corresponding shape code based on the input shape features and outputs the corresponding object shape based on the shape code. This shape code can also be represented as a shape parameter s, which may include the classification probability and correlation coefficient of each class. Given the classification probability and correlation coefficient of each class, one can be selected as the shape code representing the object, thus obtaining the corresponding object shape. For example, the shape code corresponding to the highest classification probability, or the shape code corresponding to the highest correlation coefficient, or the shape code corresponding to the optimal combination of classification probability and correlation coefficient can be selected; this disclosure does not impose any limitations on this.

[0107] Of course, the classification probabilities and correlation coefficients of different classes can also be converted into object shapes M in the form of a grid through the BLEND operation, i.e., BLEND(s) = M.

[0108] In some embodiments, for a hybrid deformer, the optimization objective for its joint tracking and shape reconstruction is as follows:

[0109]

[0110] Here, `point_to_mesh` is the minimum distance from the point cloud to the mesh. This distance is obtained by iterating through the distances from the point cloud to each mesh surface and taking the minimum value.

[0111] In the specific optimization process, the shape encoding initialized in step S203 It can be obtained through the following formula:

[0112]

[0113] In step S204, the state optimization of the object at time t can be performed according to the following formula:

[0114]

[0115] In step S204, the shape encoding optimization of the object at time t can be performed using the following formula:

[0116]

[0117] It should be noted that those skilled in the art can choose other types of shape models, as well as corresponding objective functions and optimization formulas, as needed. This disclosure does not limit the specific form and expression formula of the shape model.

[0118] In some embodiments, method 200 may further include the step of: determining the corresponding object shape based on the shape encoding of the object in each point cloud frame, wherein the object shape is represented by a directed distance value, an occupancy probability value, or a voxel value. Points with a directed distance value, occupancy probability value, or voxel value that are preset values ​​(e.g., 0) constitute the object surface. Alternatively, points with directed distance values, occupancy probability values, or voxel values ​​within a preset range constitute the object surface. Here, the scene point cloud or object point cloud, along with the corresponding shape encoding in the calculated point cloud frame, are input into the shape model, and the output result of points with preset values ​​collectively constitutes the object surface.

[0119] In some embodiments, step S204, which calculates the shape code of the corresponding object in other point cloud frames of the point cloud sequence based on the shape code of the initial point cloud frame and the object state, includes:

[0120] Based on the shape model, the shape code of the corresponding object in the (k+1)th point cloud frame is calculated by combining the shape code of the k-th point cloud frame and the object point data of the (k+1)-th point cloud frame. For example, using the formula... To optimize the shape encoding of objects in the cloud frame at time t.

[0121] Where k is an integer greater than or equal to 1. In some embodiments, the k-th frame and the (k+1)-th frame are two adjacent frames in the point cloud sequence. In other embodiments, a predetermined number of frames are discarded between each k-th frame and the (k+1)-th frame, for example, a keyframe is determined every one or more frames, in which case the k-th frame and the (k+1)-th frame are two frames in the point cloud sequence that are separated by a predetermined number of frames.

[0122] Furthermore, the object point cloud in the (k+1)th point cloud frame is either extracted from the object state in the kth point cloud frame, or optimized from the object state in the (k+1)th point cloud frame. The former uses a rough object point cloud for shape optimization, usually before state optimization has been performed and the accurate object state in the (k+1)th frame is obtained. The latter uses an optimized object point cloud, usually after state optimization has been performed, and the point cloud is optimized based on the optimized state.

[0123] In some embodiments, calculating the shape code of the corresponding object in other point cloud frames of the point cloud sequence based on the shape code and the object state in step S204 further includes:

[0124] One or more observation frame sets are formed by the first k+1 point cloud frames, and each observation frame set includes at least one point cloud frame. Based on the shape model, the shape code of the corresponding object in the k+1 point cloud frame is obtained by calculating the shape code of the object and the object point cloud in each observation frame set. For example, formulas (4), (8), and (12) are used to optimize the shape code of the object in the point cloud frame at time t. In this process, the state of the object in each point cloud frame is also optimized at the same time.

[0125] In some embodiments, the object state of the corresponding object in other point cloud frames of the point cloud sequence based on shape encoding and object point cloud computing in step S204 includes:

[0126] Step A: Extract the corresponding object point cloud from the (k+1)th point cloud frame based on the object state of the kth point cloud frame.

[0127] Step B: Based on the shape model, calculate the corresponding object state in the k+1 point cloud frame according to the shape encoding of the k-th point cloud frame and the object point cloud frame in the k+1 point cloud frame. For example, use formulas (3), (7), and (11) to optimize the state of the object in the point cloud frame at time t.

[0128] Step C: Optimize the object point cloud in the (k+1)th point cloud frame based on the object state in the (k+1)th point cloud frame.

[0129] In some embodiments, step A, extracting the corresponding object point cloud from the (k+1)th point cloud frame based on the object state of the kth point cloud frame, includes:

[0130] Step A1: Based on the object state in the k-th point cloud frame and the inter-frame displacement value, calculate the estimated state of the object in the (k+1)-th point cloud frame; and

[0131] Step A2: Determine the estimated point cloud of the object in the (k+1)th point cloud frame based on the estimated state, and use it as the extracted object point cloud.

[0132] Wherein, the predicted state of the (k+1)th point cloud frame = the optimized state in the kth point cloud frame + the inter-frame displacement value ΔS prior (k). Corresponding to the state parameters, the inter-frame displacement values ​​also include at least one of the position parameter changes and the angle parameter changes. The position parameter changes include the coordinate changes of the object's key points in the three-dimensional coordinate system; the angle parameters include at least one of the pitch angle changes, yaw angle changes, and roll angle changes. The coordinate changes in the three-dimensional coordinate system can include xyz coordinate changes. Assuming the state parameters are represented as x, y, z, θ, the corresponding state parameter changes are Δx, Δy, Δz, and Δθ. Generally, this disclosure maintains the lifecycle of each object through a tracking algorithm. When the object is not detected for several consecutive frames, it is determined that the object has disappeared from the field of view, and the state prediction for the object can be stopped.

[0133] In some embodiments, if the k-th frame is the initial frame, there is limited reference information. Therefore, the inter-frame displacement value ΔS of the second frame relative to the initial frame is set. prior (1) = 0. If the k-th frame is not the initial frame, then the inter-frame displacement value is the average displacement value between multiple consecutive point cloud frames. For point clouds that are not the initial frame, the object states of some frames have already been obtained. Based on the object states of these known frames, the displacement value between each two frames can be calculated, and then the inter-frame displacement value corresponding to the k-th point cloud frame can be obtained through a weighted average algorithm, which is the inter-frame displacement value of the (k+1)-th frame relative to the k-th frame.

[0134] In some embodiments, step A2, determining the estimated point cloud of the object in the (k+1)th point cloud frame based on the estimated state, includes:

[0135] In the (k+1)th point cloud frame, a first detection box is generated based on the estimated state of the object, and point clouds related to the object are selected within the first detection box as estimated point clouds. The first detection box is greater than or equal to the object's bounding box, which can be obtained in the initial point cloud frame through manual annotation or object detection algorithms.

[0136] In some embodiments, step B, based on the shape encoding of the k-th point cloud frame and the object point cloud in the k+1-th point cloud frame, includes the following:

[0137] The optimal state of the object in the (k+1)th point cloud frame is obtained by calculating the shape encoding and point cloud of the same object within the observation frame set. Specifically, constraints are established for the object point cloud within the observation frame set, and the optimal state of the object in the (k+1)th point cloud frame is calculated based on the constraints.

[0138] It should be understood that before predicting the optimized state of the point cloud in the next frame, the optimized states of the object's point cloud in the previous frame may already be known, and are denoted as optimized states. Therefore, when optimizing the point cloud within the set of observation frames, this disclosure also updates the optimized states of the object in at least one previous frame to obtain the corresponding second, third, or nth optimized states, etc.

[0139] Here, the observation frame set can be a sliding window set with a preset frame length, such as four frames. When calculating the optimized state of each next point cloud frame, the point cloud optimization is performed by combining the point cloud data from the previous three frames. Simultaneously, while optimizing the next point cloud frame, the point clouds and states of the three previous frames are also updated. Furthermore, considering the relatively high confidence level of the initial frame data, the initial frame can be included in each optimization set to optimize the point clouds within the set in conjunction with this initial frame.

[0140] Generally, after obtaining the optimized state of an object each time, the optimized point cloud of the object in that frame of data can be determined based on that optimized state. Correspondingly, the first optimized state has a corresponding first optimized point cloud, the second optimized state has a corresponding second optimized point cloud, and so on. This may involve multiple filtering iterations. When the maximum number of iterations is reached, the current optimized point cloud and optimized state are determined. After each point cloud optimization, the latest optimized point cloud for each frame is obtained, and this latest optimized point cloud can then be used to participate in the point cloud optimization of subsequent frames. Generally, online scenarios have higher requirements for data real-time performance; therefore, online scenarios can output the first optimized state of each frame, while offline scenarios can output the last updated optimized state and optimized point cloud for each frame.

[0141] In some embodiments, the optimization of the object state in the (k+1)th point cloud frame comes from minimizing the loss function (constraints), which includes, but is not limited to, the following constraints:

[0142] 1) Constraints on the change of inter-frame displacement values ​​between multiple consecutive point cloud frames, that is, the change of inter-frame displacement values ​​between multiple consecutive frames cannot exceed a predetermined threshold.

[0143] 2) Constraints on the height difference between the object and the ground in the initial frame and the (k+1)th point cloud frame: The height difference between the object and the ground in the initial frame and the (k+1)th point cloud frame cannot exceed a predetermined threshold. Specifically, the ground point cloud is first determined, and the ground height is determined based on this ground point cloud. The height of the object above the ground is then determined based on the point cloud associated with the object. The point cloud associated with the object can be either a predicted point cloud or an optimized point cloud, depending on whether the frame has been optimized.

[0144] 3) Registration distance constraint for at least two object point clouds in the previous frame: the registration distance between any two objects in the previous frame cannot exceed a predetermined threshold.

[0145] 4) Consistency constraint of the object's orientation and direction of motion

[0146] 5) Constraints between the predicted SDF value of the object point cloud and the theoretical value of 0, i.e., the predicted SDF value of the object point cloud should be as close to 0 as possible.

[0147] 6) Registration distance constraint between the (k+1)th point cloud frame and at least one object point cloud in the previous frame: If the consistency constraint is satisfied after registration between the previous frame and the next frame, this may include:

[0148] a) First registration distance constraint between the optimized point cloud of the current point cloud frame and the predicted point cloud of the next point cloud frame.

[0149] b) Second registration distance constraint between the point cloud within the optimized shape of the initial point cloud frame and the predicted point cloud of the next point cloud frame.

[0150] c) Third registration distance constraint between the point cloud within the optimized shape of the current point cloud frame and the predicted point cloud of the next point cloud frame.

[0151] Each of the above constraints has its corresponding weight. By weighting these constraints, the loss function can be obtained, and this loss function can be used to optimize the point cloud within the observation frame set. The calculation formulas for some of the constraints are as follows:

[0152]

[0153]

[0154]

[0155]

[0156] Among them, O k-1 The optimized point cloud representing the point cloud in frame k-1. This represents the predicted point cloud of the k-th frame. Representing Ok-1 and The registration point set between them representative point set The number of points in This represents moving point p from state S in frame k-1. k-1 Move ΔS k The distance is q, where q is the point registered by point p in the point cloud of the kth frame, and ||2 represents the L2 norm.

[0157] M1 represents the optimized shape of the point cloud in the initial frame. Representing M1 and The registration point set between them This represents moving point p from the initial state of the point cloud in the initial frame. move The distance S k This represents the optimized state of the point cloud in frame k. M k-1 Represents the optimized shape of the point cloud in frame k-1. Representing M k-1 and The registration point set. v is the inter-frame displacement value of the point cloud from frame (k-1) to frame k, Δx k Let θ be one of the coordinate displacement values. k-1 and θ k These are the rotation angles of the point cloud in frame (k-1) and frame (k), respectively.

[0158] In some embodiments, step C, optimizing the object point cloud in the (k+1)th point cloud frame based on the object state, includes: determining the optimized point cloud of the object in the (k+1)th point cloud frame based on the optimized state. Specifically, a second detection box is generated in the (k+1)th point cloud frame based on the optimized state of the object, and point clouds related to the object are selected as optimized point clouds within the second detection box. The second detection box is less than or equal to the first detection box and greater than or equal to the bounding box of the object.

[0159] Let the second detection box be a times the size of the bounding box. Generally, for the initial point cloud frame, the inter-frame displacement value is 0, and the first detection box is c times the size of the bounding box. For non-initial point cloud frames, the inter-frame displacement value is the average displacement value between multiple consecutive point cloud frames, and the first detection box is b times the size of the bounding box; where c > b > a ≥ 1.

[0160] Here, when estimating the object state in the second frame from the initial frame, there is relatively little reference information, so it is necessary to filter the point cloud related to the object over a larger range. Therefore, c can be set to 2-4, for example, c=3. For non-initial frames, there is already some reference information from the previous frame, so it can be magnified by a slightly smaller factor, for example, b=1.3-1.7, and further, b=1.5 times. Since the optimized pose of the object can be known, the point cloud of the region where the object is located can be determined more accurately, so a=1-1.3, for example, a=1.1.

[0161] In some embodiments, method 200 may further include the steps of: calculating the displacement value of the object between the two point cloud frames based on the optimized state of the object in the k-th point cloud frame and the (k+1)-th point cloud frame; updating the inter-frame displacement value based on the calculated displacement value, so as to calculate the estimated state of the object in the next point cloud frame based on the updated inter-frame displacement value.

[0162] This section primarily focuses on updating based on the object's motion model, using this model to maintain a motion observation value ΔS. prior (k) represents the prediction of motion in the next frame. In actual operation, a moving average method is used to maintain it, and the inter-frame displacement value is automatically updated whenever a new inter-frame motion is predicted. Specifically, the inter-frame displacement value can be updated based on the optimized state of the object in the previous k+1 point cloud frames.

[0163] In some embodiments, it is assumed that the inter-frame displacement value corresponding to the point cloud in the k-th frame is ΔS. prior (k), the optimized states of the k-th frame and the (k+1)-th frame are S respectively. k and S k+1 Then the inter-frame displacement value corresponding to the point cloud in the (k+1)th frame is ΔS. prior (k+1)=αΔS prior (k)+(1-α)(S k+1 -S k Then, the object state in the (k+2)th point cloud frame can be estimated based on the updated inter-frame displacement value, and the object point cloud can be extracted based on the estimated state. The object state in the (k+2)th point cloud frame can be optimized based on the extracted object point cloud, and the object shape encoding can be updated.

[0164] In summary, this disclosure starts with an initial frame and uses an iterative optimization algorithm and shape model to calculate the state and shape in the point cloud frame by frame, thereby improving the accuracy of state estimation and the robustness of 3D reconstruction.

[0165] Figure 6 A flowchart of a point cloud processing method 600 according to another embodiment of this disclosure is shown, as follows: Figure 6 As shown, the point cloud processing method 600 includes:

[0166] Step S601: Obtain the object state in the initial point cloud frame and use the initial point cloud frame as the current point cloud frame.

[0167] Step S602: Extract the corresponding object point cloud from the current point cloud frame according to the object state of the current point cloud.

[0168] Step S603: Compute the shape code of the corresponding object in the current point cloud frame based on the shape model and the object point cloud of the current point cloud frame.

[0169] Step S604: Based on the shape encoding of the historical observation frame set including the current point cloud frame and the object point cloud, calculate the object state of the next point cloud frame.

[0170] Step S605: Determine whether the next point cloud frame is the end point cloud frame; if yes, end the process; otherwise, in step S606, update the current point cloud frame to the next point cloud frame and re-trigger step S602 until the end frame is processed. The end frame can be the last frame of the point cloud sequence or a preset frame.

[0171] In addition, such as Figure 7 As shown in the embodiments of this disclosure, a point cloud processing apparatus 700 is also provided, comprising:

[0172] The object state acquisition module 701 is adapted to acquire the object state in the initial point cloud frame;

[0173] The object point cloud extraction module 702 is adapted to extract the corresponding object point cloud from the initial point cloud frame according to the object state.

[0174] Shape encoding calculation module 703 is adapted to encode the shape of the corresponding object in the initial point cloud frame based on the shape model and the object point cloud computing, wherein the shape encoding is used to characterize the surface shape of the object; and

[0175] The iteration module 704 is adapted to calculate the object state and / or shape code of the corresponding object in other point cloud frames of the point cloud frame sequence based on the shape code and object state.

[0176] It should be noted that the point cloud processing device 700 provided in this disclosure is based on the... Figures 1-6 The details have been disclosed in the description and will not be repeated here.

[0177] In addition, embodiments of this disclosure also provide a computer-readable storage medium, including a program or instructions, which, when run on a computer, implement the point cloud processing method as described above.

[0178] In addition, this disclosure also provides an embodiment such as Figure 8The computing device 800 shown includes a memory 801 and one or more processors 802 communicatively connected to the memory. The memory 801 stores instructions executable by the one or more processors 802, which, when executed, cause the one or more processors 802 to implement the point cloud processing method described above. The computing device 800 may further include a communication interface 803, which can implement one or more communication protocols (LTE, Wi-Fi, etc.).

[0179] According to the technical solution of this disclosure, by utilizing the continuity of real observation data, the object is tracked while its 3D shape is continuously updated based on a shape model, thereby improving the fidelity of shape reconstruction. Simultaneously, this disclosure uses a more precise object shape to improve tracking performance, thus achieving joint optimization of object state and shape. This method eliminates the dependence on target detection, obtaining higher-precision state estimation by directly processing point cloud data; it achieves high-quality 3D reconstruction through multi-frame, multi-view information of point cloud sequences, without requiring supervised data or relying heavily on single-frame point cloud data.

[0180] This disclosure can be used in online perception modules or offline analysis modules for autonomous driving. In online scenarios, this disclosure utilizes general-purpose computing devices to process radar point clouds to obtain estimates of the absolute state of surrounding vehicles, which are then supplied to downstream planning and perception modules. In offline scenarios, this disclosure provides 3D model information and motion state information of surrounding vehicles through offline analysis data, providing benchmark data for planning, fusion, and other parts of the system. It can also be used to create virtual complex radar scenarios.

[0181] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0182] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0183] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0184] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0185] This disclosure uses specific embodiments to illustrate the principles and implementation methods of this disclosure. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.

Claims

1. A point cloud processing method, comprising: Obtain the object state in the initial point cloud frame; Extract the corresponding object point cloud from the initial point cloud frame based on the object state; Based on the shape model and the shape code of the corresponding object in the initial point cloud frame of the object point cloud computing, the shape code is used to characterize the surface shape of the object; as well as The object state and shape code of the corresponding object in other point cloud frames of the point cloud sequence are calculated based on the shape code and object state. The calculation of the object state and shape code of the corresponding object in other point cloud frames in the point cloud frame sequence based on the shape code includes: Based on the shape model, the object state calculation of the corresponding object in other point cloud frames is performed m times and the shape code calculation of the corresponding object in other point cloud frames is performed n times, according to the shape code and object state of the initial point cloud frame, where m and n are both integers greater than or equal to 1. The set of observation frames comprises one or more observation frames consisting of the first k+1 point cloud frames, and each observation frame set includes at least one point cloud frame. The calculation of the object state and shape code of the corresponding object in other point cloud frames of the point cloud sequence based on the shape code and object state includes: Based on the object state and inter-frame displacement value in the k-th point cloud frame, calculate the estimated state of the object in the (k+1)-th point cloud frame. In the (k+1)th point cloud frame, a first detection box is generated based on the estimated state of the object, and the first detection box is greater than or equal to the bounding box of the object. Within the first detection box, select the point cloud related to the object as the extracted object point cloud; The optimized state of the object in the (k+1)th point cloud frame is obtained by calculating the shape encoding and point cloud of the same object within the observation frame set. In the (k+1)th point cloud frame, a second detection box is generated based on the optimized state of the object. The second detection box is less than or equal to the first detection box and greater than or equal to the bounding box of the object. Within the second detection box, select the point cloud related to the object as the optimized point cloud.

2. The method according to claim 1, wherein, The object state of the corresponding object in other point cloud frames of the point cloud frame sequence is calculated based on the shape encoding and object state, including: Extract the corresponding object point cloud from the (k+1)th point cloud frame based on the object state of the kth point cloud frame, where k is an integer greater than or equal to 1. Based on the shape model, the object state in the (k+1)th point cloud frame is calculated by combining the shape encoding of the corresponding object in the k-th point cloud frame and the object point value in the (k+1)-th point cloud frame; and Optimize the object point cloud in the (k+1)th point cloud frame based on the object state in the (k+1)th point cloud frame.

3. The method according to claim 2, wherein, Calculating the shape encoding of the corresponding object in other point cloud frames of the point cloud sequence based on the shape encoding and object state includes: Based on the shape model, the shape code of the corresponding object in the (k+1)th point cloud frame is calculated by combining the shape code of the object in the k-th point cloud frame and the object point code of the (k+1)th point cloud frame. Among them, the object point cloud of the (k+1)th point cloud frame is the object point cloud extracted based on the object state of the kth point cloud frame, or the object point cloud optimized based on the object state in the (k+1)th point cloud frame.

4. The method according to claim 2, wherein, The shape encoding of the corresponding object in other point cloud frames of the point cloud frame sequence is calculated based on the shape encoding and object state, including: Based on the shape model, the shape code of the corresponding object in the (k+1)th point cloud frame is obtained by calculating the shape code and object point cloud in each set of observation frames.

5. The method according to claim 1, wherein, The shape model is any one of the following: A deep learning-based directed distance function outputs the directed distance from a 3D point to the nearest point on the surface of an object, based on the input 3D point coordinates and shape encoding. The deep learning-based occupancy network obtains the corresponding shape code based on the input set of three-dimensional coordinate points, and outputs the occupancy probability of the three-dimensional coordinate points based on the three-dimensional coordinate points and the shape code. The hybrid deformer obtains the classification probability and correlation coefficient of each class based on the input shape features, and outputs the corresponding object shape based on the classification probability and correlation coefficient.

6. The method according to claim 5, further comprising: The object shape is determined based on the shape encoding of the object in each point cloud frame. The object shape is represented by a directed distance value, an occupancy probability value, or a voxel value. Points with a preset directed distance value, occupancy probability value, or voxel value constitute the object surface.

7. The method according to claim 1, wherein, The second detection box is a times the size of the bounding box. For the initial point cloud frame, the inter-frame displacement value is 0, and the first detection box is c times the size of the bounding box; For non-initial point cloud frames, the inter-frame displacement value is the average displacement value between multiple consecutive point cloud frames, and the first detection box is b times the bounding box; where c > b > a ≥ 1.

8. The method according to claim 1, wherein, By calculating the shape encoding and point cloud of the same object within the observation frame set, the optimized state of the object in the (k+1)th point cloud frame is obtained, including: Establish constraints on the object point cloud within the set of observation frames; Calculate the optimized state of the object in the (k+1)th point cloud frame based on the constraints.

9. The method according to claim 8, wherein, The constraints include at least one of the following: Constraints on inter-frame displacement values ​​between multiple consecutive point cloud frames; Constraints on the change in the height difference between the object and the ground in the initial frame and the next cloud frame; Registration distance constraints for at least two object point clouds between previous frames; Registration distance constraint between the (k+1)th point cloud frame and at least one object point cloud in the previous frame.

10. The method according to claim 1, wherein, The object state includes at least one of position parameters and angle parameters. The position parameters include three-dimensional spatial coordinate values, and the angle parameters include at least one of pitch angle, yaw angle, and roll angle.

11. A point cloud processing apparatus, comprising: The object state acquisition module is suitable for acquiring the object state in the initial point cloud frame. An object point cloud extraction module is adapted to extract the corresponding object point cloud from the initial point cloud frame according to the object state. A shape encoding calculation module is adapted to calculate the shape encoding of the corresponding object in the initial point cloud frame based on the shape model and the object point cloud, wherein the shape encoding is used to characterize the surface shape of the object. as well as An iterative module is adapted to encode the object state and shape of the object based on the shape encoding and other point cloud frames in the object point cloud computing point cloud sequence. The iterative module is further adapted to perform, based on the shape model, m times the object state calculation of the corresponding object in other point cloud frames and n times the shape code calculation of the corresponding object in other point cloud frames, according to the shape code and object state of the initial point cloud frame, where m and n are both integers greater than or equal to 1. Wherein, one or more observation frame sets are composed of the first k+1 point cloud frames, and each observation frame set includes at least one point cloud frame, and the iterative module is further adapted to: Based on the object state and inter-frame displacement value in the k-th point cloud frame, calculate the estimated state of the object in the (k+1)-th point cloud frame. In the (k+1)th point cloud frame, a first detection box is generated based on the estimated state of the object, and the first detection box is greater than or equal to the bounding box of the object. Within the first detection box, select the point cloud related to the object as the extracted object point cloud; The optimized state of the object in the (k+1)th point cloud frame is obtained by calculating the shape encoding and point cloud of the same object within the observation frame set. In the (k+1)th point cloud frame, a second detection box is generated based on the optimized state of the object. The second detection box is less than or equal to the first detection box and greater than or equal to the bounding box of the object. Within the second detection box, select the point cloud related to the object as the optimized point cloud.

12. A computing device, comprising: Processor, memory, and computer programs stored in memory and capable of running on the processor; When the processor runs the computer program, it performs the method described in any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method of any one of claims 1-10.