Data processing method and device, electronic equipment, computer readable storage medium and computer program product
Through built-in sensors to collect motion state data sets and automatically identify dynamic obstacle point clouds with algorithms, the problem of time-consuming and labor-consuming manual editing in the existing technology is solved, and efficient and accurate point cloud data processing and model training are achieved.
Patent Information
- Application Number
- CN202510206488.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
When building a binocular depth prediction model, the prior art requires manual editing and cropping of point cloud data, which consumes a lot of human resources and has low accuracy, making it difficult to effectively remove dynamic obstacle point clouds.
The motion state data set is collected by the second sensor built into the first sensor, combined with the M-detector algorithm and the neighborhood search algorithm, and automatically identify and filter out the point cloud of dynamic obstacles, and build a high-quality training data set for binocular depth prediction model.
It improves the convenience and efficiency of data acquisition, enhances the accuracy of point cloud data sets, and improves the training efficiency and performance of binocular depth prediction models.
Smart Images

Figure CN120219871A_ABST
Abstract
Description
Technical Field
[0001] This application relates to data processing technologies, and in particular, to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] In the process of constructing a binocular depth prediction model, a high-quality training data set is required to train the model. The training data set usually includes binocular images and depth maps. Among them, it is necessary to collect point cloud data through a radar and convert it into a depth map in the training data set. Since the collected point cloud data often contains point clouds of dynamic obstacles, it is necessary to perform simultaneous localization and mapping (SLAM) using odometry data and point cloud data, and use point cloud editing software to crop the point clouds of dynamic obstacles. However, this process takes a long time and requires manual editing and cropping of point cloud data, consuming a large amount of human resources. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, apparatus, computer-readable storage medium, and computer program product, which can improve the efficiency and accuracy of obtaining a second point cloud data set and improve the efficiency of training a binocular depth prediction model.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] Embodiments of this application provide a data processing method, the method including:
[0006] Obtaining a motion state data set of a target object and a first point cloud data set of an environment where the target object is located, the first point cloud data set being collected by a first sensor, and the motion state data set being collected by a second sensor built in the first sensor;
[0007] Based on the motion state data set, determining target point clouds of dynamic obstacles in the environment from the first point cloud data set;
[0008] Determining supplementary point clouds of the dynamic obstacles based on the target point clouds;
[0009] Deleting the target point clouds and the supplementary point clouds from the first point cloud data set to obtain a second point cloud data set;
[0010] Constructing a training data set based on the second point cloud data set, the training data set being used to train a binocular depth prediction model.
[0011] Embodiments of this application provide a data processing apparatus, including:
[0012] An acquisition module, configured to acquire a motion state data set of a target object and a first point cloud data set of the environment where the target object is located, where the first point cloud data set is collected by a first sensor, and the motion state data set is collected by a second sensor built in the first sensor;
[0013] A determination module, configured to determine target point clouds of dynamic obstacles in the environment from the first point cloud data set based on the motion state data set;
[0014] The determination module is further configured to determine supplementary point clouds of the dynamic obstacles based on the target point clouds;
[0015] A deletion module, configured to delete the target point clouds and the supplementary point clouds from the first point cloud data set to obtain a second point cloud data set;
[0016] A construction module, configured to construct a training data set based on the second point cloud data set, where the training data set is used to train a binocular depth prediction model.
[0017] An embodiment of the present application provides an electronic device, where the electronic device includes:
[0018] A memory, configured to store computer-executable instructions or a computer program;
[0019] A processor, configured to implement the data processing method provided by the embodiment of the present application when executing the computer-executable instructions or the computer program stored in the memory.
[0020] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, which are used to implement the data processing method provided by the embodiment of the present application when being executed by a processor.
[0021] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where the computer program or computer-executable instructions implement the data processing method provided by the embodiment of the present application when being executed by a processor.
[0022] The embodiment of the present application has the following beneficial effects:
[0023] Through the embodiments of the present application, a motion state data set of a target object and a first point cloud data set of the environment where the target object is located are obtained. The first point cloud data set is collected by a first sensor, and the motion state data set is collected by a second sensor built in the first sensor. In this way, collecting the motion state data set by the second sensor built in the first sensor can reduce the number of sensors required for data collection, improving the convenience and efficiency of data collection. Moreover, since the first sensor and the second sensor are arranged together, it is possible to ensure the simultaneous collection of the motion state data set and the first point cloud data set, improving the coordination of data collection. Based on the motion state data set, target point clouds of dynamic obstacles in the environment are determined from the first point cloud data set; supplementary point clouds of the dynamic obstacles are determined based on the target point clouds; the target point clouds and the supplementary point clouds are deleted from the first point cloud data set to obtain a second point cloud data set. In this way, after determining the target point clouds of the dynamic obstacles in the environment, the supplementary point clouds are further determined, and the point clouds corresponding to the dynamic obstacles can be completely deleted from the first point cloud data set, thereby improving the accuracy of the obtained second point cloud data set. Constructing a training data set for training a binocular depth prediction model based on the second point cloud data set can improve the quality of the training data set, thereby improving the performance of the trained binocular depth prediction model. Therefore, through the embodiments of the present application, the convenience and efficiency of data collection can be improved, the efficiency and accuracy of obtaining the second point cloud data set can be improved, and the efficiency of training the binocular depth prediction model and the performance of the trained model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic diagram of the architecture of the data processing system provided by the embodiments of the present application;
[0025] Figure 2 is a schematic diagram of the structure of server 200 provided by the embodiments of the present application;
[0026] Figure 3A is a schematic flowchart of the data processing method provided by the embodiments of the present application;
[0027] Figure 3B is a schematic flowchart of determining target point clouds provided by the embodiments of the present application;
[0028] Figure 3C is a schematic flowchart of determining the first speed and the first position provided by the embodiments of the present application;
[0029] Figure 3D is a schematic flowchart of determining the first pose provided by the embodiments of the present application;
[0030] Figure 3E is another schematic flowchart of determining target point clouds provided by the embodiments of the present application;
[0031] Figure 3F It is a schematic flowchart of determining supplementary point cloud provided by an embodiment of the present application;
[0032] Figure 3G It is a schematic flowchart of constructing a training data set provided by an embodiment of the present application;
[0033] Figure 4 It is a schematic flowchart of training a binocular depth prediction model provided by an embodiment of the present application;
[0034] Figure 5 It is another schematic flowchart of data processing provided by an embodiment of the present application. Detailed implementation manners
[0035] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0036] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0037] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0038] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0040] In the embodiments of this application, when collecting and processing relevant data in practical applications, it should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope authorized by laws, regulations and the personal information subject.
[0041] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0042] 1) Binocular image acquisition device: That is, a binocular RGB camera, an imaging system that simulates human binocular stereoscopic vision, used to acquire binocular images. By the coordinated operation of two or more cameras, images are captured from different perspectives to restore the three-dimensional structure of the scene and obtain the depth information of the scene.
[0043] 2) Binocular image: That is, a pair of binocular RGB images, which are RGB color images captured by two cameras respectively. It is usually used to analyze the image data obtained from the same scene from two different perspectives to measure and deduce the depth information of the objects in the scene.
[0044] 3) Simultaneous localization and mapping: A technology that simultaneously and real-time builds a map and determines its own position in an unknown environment. For example, using a 3D lidar and an odometer to simultaneously build a map and locate its own position.
[0045] 4) 3D lidar: Also known as a lidar (Light Detection and Ranging, LiDAR) system, it is a device that measures the distance and shape of a target by emitting laser pulses towards the target and measuring the reflected light, and can quickly and accurately obtain three-dimensional space information of a large area.
[0046] 5) Odometer: A device used to measure and record the driving distance of a vehicle, usually installed on a car. In the field of robotics or automation, an odometer refers to a sensor or measurement system used to track the moving distance and direction of a robot or an object. An odometer can be mechanical or digital, and usually calculates the moving distance by detecting the number of wheel rotations.
[0047] 6) Inertial Measurement Unit (IMU): A combined unit of sensors that measures acceleration and rotational motion, usually including an accelerometer, a gyroscope, and sometimes a magnetometer, which can provide instant information about the motion and direction of an object in three-dimensional space.
[0048] 7) Dynamic Obstacle Detection Algorithm (M-detecor Algorithm): An algorithm specifically for dynamic obstacle detection, aiming to detect and track dynamic obstacles in a video stream in real time. By analyzing the changes between consecutive frames, it identifies and locates moving objects and marks them as obstacles.
[0049] 8) Neighborhood Search Algorithm (Radius Nearest Neighbors, Radius-NN): A nearest neighbor search algorithm based on a fixed radius, used to find all points within the neighborhood of each point in a given dataset. The size of the neighborhood is determined by a user-specified radius.
[0050] In the related art, it is necessary to continuously scan the environment using a 3D lidar to obtain a series of point cloud data. At the same time, an odometer is used to record the movement trajectory of the lidar during the acquisition process. Since the collected point cloud data often contains the point cloud of some dynamic obstacles, it is necessary to perform simultaneous localization and mapping (Simultaneous Localization and Mapping, Slam) using the odometer data and the point cloud data, and use point cloud editing software to crop the point cloud of the dynamic obstacles. However, this process requires manual editing and cropping of the point cloud data, which takes a long time and consumes a large amount of human resources. In addition, when cropping the point cloud of dynamic obstacles in the related art, it is easy to miss some of the point cloud of dynamic obstacles, resulting in low accuracy.
[0051] The embodiments of the present application provide a data processing method, device, equipment, computer-readable storage medium, and computer program product, which can improve the convenience and efficiency of data acquisition, completely delete the point cloud corresponding to the dynamic obstacle from the first point cloud dataset, and improve the efficiency of training the binocular depth prediction model and the performance of the trained model. The following describes the exemplary applications of the electronic equipment provided by the embodiments of the present application. The equipment provided by the embodiments of the present application can be implemented as various types of terminals such as in-vehicle terminals, laptop computers, tablet computers, desktop computers, set-top boxes, smartphones, smart speakers, smart watches, smart TVs, etc., or can also be implemented as a server. The following will describe the exemplary applications when the equipment is implemented as a server.
[0052] See Figure 1 , Figure 1 is a schematic diagram of the architecture of the data processing system provided by the embodiments of the present application. A first sensor 410 is provided in the terminal 400, and a second sensor 411 is built into the first sensor 410. The terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0053] The terminal 400 is configured to collect a motion state data set of a target object through a first sensor 410, collect a first point cloud data set of the environment where the target object is located through a second sensor 411 built in the first sensor 410, and send the motion state data set and the first point cloud data set to the server 200 through the network 300. The server 200 obtains the motion state data set of the target object and the first point cloud data set of the environment where the target object is located; determines target point clouds of dynamic obstacles in the environment from the first point cloud data set based on the motion state data set; determines supplementary point clouds of the dynamic obstacles based on the target point clouds; deletes the target point clouds and the supplementary point clouds from the first point cloud data set to obtain a second point cloud data set; constructs a training data set based on the second point cloud data set. The training data set is used to train a binocular depth prediction model and can be stored in the database 100. The database 100 can be independent of the server 200 or deployed on the server 200. It is exemplarily shown in Figure 1 that the database 100 is independent of the server 200. When model training is required, the server 200 obtains the training data set from the database 100, trains the binocular depth prediction model to be trained through the training data set, obtains the trained binocular depth prediction model, and stores it in the database 100. After that, when depth prediction is required, the server 200 calls the trained binocular depth prediction model from the database 100 for depth prediction.
[0054] Taking the electronic device for data processing as the server above as an example, refer to Figure 2 , Figure 2 which is a schematic structural diagram of the server 200 provided by an embodiment of the present application. Figure 2 The server 200 shown includes at least one processor 210, a memory 230, and at least one network interface 220. Each component in the server 200 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 240.
[0055] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0056] The memory 230 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 230 optionally includes one or more storage devices that are physically remote from the processor 210.
[0057] The memory 230 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.
[0058] In some embodiments, the memory 230 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.
[0059] The operating system 231 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0060] The network communication module 232 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wi-Fi (Wireless Fidelity), and Universal Serial Bus (USB), etc.
[0061] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is a data processing device 233 stored in the memory 230, which can be software in the form of programs and plugins, etc., and includes the following software modules: an acquisition module 2331, a determination module 2332, a deletion module 2333, and a construction module 2334. These modules are logical, and thus can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.
[0062] In some other embodiments, the device provided in the embodiments of the present application may be implemented in a hardware manner. As an example, the device provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the data processing method provided in the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0063] The exemplary applications and implementations of the terminal provided in the embodiments of the present application will be combined to illustrate the data processing method provided in the embodiments of the present application.
[0064] Next, the data processing method provided in the embodiments of the present application will be described. As mentioned above, the electronic device for implementing the data processing method in the embodiments of the present application may be a terminal, a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.
[0065] It should be noted that in the following examples of data processing, the object is taken as the face for illustration. Those skilled in the art can apply the data processing method provided in the embodiments of the present application to the processing of an image set including other types of objects according to the understanding of the following text.
[0066] See Figure 3A , Figure 3A is a schematic flowchart of the data processing method provided in the embodiments of the present application, and will be described in combination with Figure 3A the steps shown.
[0067] In step 101, a motion state data set of the target object and a first point cloud data set of the environment where the target object is located are obtained.
[0068] Here, the first point cloud dataset is collected by the first sensor, and the motion state dataset is collected by the second sensor built into the first sensor. The first sensor can be a sensor that can provide three-dimensional coordinate information, such as a lidar, an acoustic wave sensor, a structured light sensor, etc. For example, a 3D lidar determines the first point cloud dataset of the environment by emitting laser pulses into the environment where the target object is located and measuring the time difference or phase change of the reflected light. The second sensor is used to obtain information such as the acceleration and angular velocity of the target object, so as to provide the motion state and attitude information of the target object. The second sensor can be an IMU. By integrating the second sensor into the first sensor, the number of sensors required for data collection can be reduced, and the convenience and efficiency of data collection can be improved. Moreover, since the first sensor and the second sensor are arranged together, it is possible to ensure that the motion state dataset and the first point cloud dataset are collected simultaneously, improving the coordination of the collected data. The target object is a movable device, usually a vehicle. The environment where the target object is located refers to the spatial environment during the vehicle's driving, which may include static road surfaces, roads, plants, etc., as well as dynamic pedestrians and other vehicles. The motion state dataset includes the acceleration and angular velocity of the target object at different times, and is used to perceive the motion state and direction of the target object. The first point cloud dataset includes the 3D point cloud data of the environment where the target object is located at different times, and is a dataset composed of point clouds corresponding to a large number of points in the environment. Each point cloud represents the position of a certain point in the environment in three-dimensional space.
[0069] In step 102, based on the motion state dataset, the target point cloud of the dynamic obstacle in the environment is determined from the first point cloud dataset.
[0070] Here, the dynamic obstacle refers to an obstacle whose position moves in the environment, such as a pedestrian, an animal, another vehicle, a bicycle, etc. The target point cloud refers to the point cloud corresponding to these dynamic obstacles in the first point cloud dataset. Here, the motion state dataset and the first point cloud dataset can be input into the dynamic obstacle detection algorithm (M-detecor algorithm), and the target point cloud of the dynamic obstacle is determined through the M-detecor algorithm.
[0071] In some embodiments, the motion state dataset includes the motion state data collected by the second sensor at multiple acquisition times, and the motion state data includes acceleration and angular velocity.
[0072] Here, the second sensor can collect motion state data at preset time intervals. For example, 10 motion state data are collected on average within 1 second. Therefore, each motion state data corresponds to a collection moment. Among them, each motion state data includes the acceleration and angular velocity of the target object at this collection moment. Acceleration is a basic concept in physics, which describes the rate of change of an object's velocity, that is, the rate of change of velocity with respect to time. It belongs to a vector, having both magnitude and direction, and its unit is usually meters per second squared (m / s 2 ). Angular velocity is a physical quantity that describes the speed of an object's rotational motion. It is the angle by which an object rotates around the axis of rotation per unit time. It can be instantaneous angular velocity or average angular velocity, and its unit is radians per second (rad / s).
[0073] In some embodiments, referring to Figure 3B , step 102 can be implemented through steps 1021 to 1023, including:
[0074] In step 1021, based on the accelerations corresponding to multiple collection moments, determine the first velocity and the first position of the target object at the first moment.
[0075] Here, the first moment is any collection moment. For the first moment, discrete integration can be performed on the accelerations corresponding to each collection moment before the first moment to obtain the first velocity of the target object at the first moment. By performing discrete integration twice on the accelerations corresponding to each collection moment before the first moment, the first position of the target object at the first moment can be obtained.
[0076] In some embodiments, referring to Figure 3C , step 1021 can be implemented through steps 211 to 216, including:
[0077] In step 211, obtain the initial motion data of the target object.
[0078] Here, the initial motion data includes at least the initial velocity and the initial position. The initial motion velocity can be the motion state data corresponding to any moment before the first moment. For example, the initial motion velocity can be the initial motion data when the target object has not started moving and is about to start moving. At this time, the initial velocity is 0, and the initial position is the initial position of the target object. Or, the initial motion data can be the motion state data of the previous collection moment before the first moment, that is, the velocity and position of the target object at the previous collection moment. Or, the initial motion data can be the motion state data of a certain collection moment separated from the first moment by a period of time, such as the motion state data corresponding to 5 seconds before the first moment.
[0079] In step 212, based on the second moment and the first moment when the initial motion data is collected, N target accelerations are obtained from multiple accelerations.
[0080] Here, N is a positive integer, such as 1, 5, 10, etc. The second moment is the collection moment when the initial motion data is collected, and the time corresponding to the second moment is before the first time. Among them, the N target accelerations are determined according to the second moment and the first moment, including all the accelerations collected between the first moment and the second moment, and the acceleration collected at the second moment. For example, if the second moment is the first collection moment, that is, the moment when the target object has not started moving but is about to start moving, then the N target accelerations include all the accelerations collected before the first moment. Or, if the second moment is the previous collection moment before the first moment, then only 1 target acceleration needs to be obtained, that is, the acceleration corresponding to the first moment. Or, if the second moment is a certain collection moment separated from the first moment by a period of time, such as the second moment is 5 seconds before the first moment, then all the accelerations collected during these 5 seconds starting from the second moment need to be obtained as the N target accelerations.
[0081] In step 213, based on the first target acceleration, the initial velocity, and the collection interval, the first target velocity is determined.
[0082] Here, the collection interval is the time interval between two adjacent collection moments, that is, the time interval between two adjacent target accelerations. For the first target acceleration among the N target accelerations, the initial velocity is the velocity of the target object corresponding to the previous collection moment of the first target acceleration. The product of the first target acceleration and the collection interval is used as the velocity change amount corresponding to the first target acceleration. By adding the initial velocity and the velocity change amount corresponding to the first target acceleration, the first target velocity corresponding to the first target acceleration can be obtained. See formula (1):
[0083] V(t + Δt) = V(t) + a(t) × Δt (1)
[0084] Among them, V(t + Δt) represents the first target velocity, V(t) represents the initial velocity, Δt represents the collection interval, and a(t) represents the first target acceleration.
[0085] In step 214, based on the i-th target acceleration, the (i - 1)-th target velocity, and the collection interval, the i-th target velocity is determined.
[0086] Here, i = 2, …, N. Taking i = 2 as an example for illustration, the product of the second target acceleration and the acquisition interval is taken as the velocity change amount corresponding to the second target acceleration. By adding the first target velocity and the velocity change amount corresponding to the second target acceleration, the second target velocity is obtained. And so on, the target velocity corresponding to each target acceleration can be obtained. Among them, this process can continue to refer to formula (1). At this time, V(t + Δt) in formula (1) represents the i-th target velocity, V(t) represents the (i - 1)-th target velocity, Δt represents the acquisition interval, and a(t) represents the i-th target acceleration.
[0087] In step 215, the N-th target velocity is determined as the first velocity.
[0088] Here, since the N-th target velocity is the velocity corresponding to the N-th target acceleration, and the N-th acceleration is collected at the first moment, the N-th target velocity is the target velocity at the first moment, that is, the first velocity.
[0089] In step 216, based on each target velocity, the initial position, and the sampling interval, the first position is determined.
[0090] Here, for the first target velocity, multiplying the first target velocity by the sampling interval can obtain the position change amount of the target object at this acquisition moment. Adding the position change amount to the initial position can obtain the first position corresponding to the first target velocity; for the second target velocity, multiplying the second target velocity by the sampling interval can obtain the position change amount of the target object at this acquisition moment. Adding the position change amount to the previous position, that is, the first position, can obtain the second position corresponding to the second target velocity. And so on, the position corresponding to each target velocity can be obtained. Among them, the position corresponding to the last one, that is, the N-th target velocity, is the first position. This process can refer to formula (2):
[0091] P(t + Δt) = P(t) + V(t + Δt) × Δt (2)
[0092] Among them, P(t + Δt) represents the i-th position, P(t) represents the (i - 1)-th position, V(t + Δt) represents the i-th target velocity, and Δt represents the sampling interval.
[0093] In an embodiment of the present application, initial motion data of a target object is obtained, and the initial motion data includes at least an initial velocity and an initial position; based on a second moment and a first moment when the initial motion data is collected, N target accelerations are obtained from a plurality of accelerations, where N is a positive integer; based on a first target acceleration, the initial velocity, and a collection interval, a first target velocity is determined; based on an i-th target acceleration, an (i - 1)-th target velocity, and the collection interval, an i-th target velocity is determined, where the collection interval is a time interval between two adjacent collection moments, and i = 2, …, N; the N-th target velocity is determined as a first velocity; based on each target velocity, the initial position, and the sampling interval, a first position is determined. In this way, the velocity and position corresponding to each collection moment, including the first velocity and the first position, can be determined through a recursive process, and the calculation method is simple and fast, thereby improving the efficiency of data processing.
[0094] In step 1022, based on angular velocities corresponding to a plurality of collection moments, a first pose of the target object at the first moment is determined.
[0095] Here, the first moment is any collection moment. For the first moment, discrete integration can be performed on the angular velocities corresponding to each collection moment before the first moment to obtain the first pose of the target object at the first moment.
[0096] In some embodiments, the initial motion data further includes an initial pose. The initial pose may be the pose of the target object when it has not started moving and is about to start moving. Alternatively, the initial pose may be the pose at the previous collection moment before the first moment, that is, the pose of the target object at the previous collection moment. Or, the initial pose may be the pose at a certain collection moment separated from the first moment by a period of time, such as the pose of the target object corresponding to 5 seconds before the first moment.
[0097] Correspondingly, referring to Figure 3D , step 1022 can be implemented through steps 221 to 223, including:
[0098] In step 221, N target angular velocities are obtained from a plurality of angular velocities based on the second moment and the first moment.
[0099] Here, N is a positive integer, such as 1, 5, 10, etc. The second moment is the acquisition moment for collecting the initial motion data, and the time corresponding to the second moment is before the first time. Among them, the N target angular velocities are determined based on the second moment and the first moment, including all the angular velocities collected between the first moment and the second moment, and the angular velocity collected at the second moment. For example, if the second moment is the first acquisition moment, that is, the moment when the target object has not started moving but is about to start moving, then the N target angular velocities include all the angular velocities collected before the first moment. Or, if the second moment is the previous acquisition moment before the first moment, then only 1 target angular velocity needs to be obtained, that is, the angular velocity corresponding to the first moment. Or, if the second moment is a certain acquisition moment separated from the first moment by a period of time, such as the second moment is 5 seconds before the first moment, then all the angular velocities collected during these 5 seconds starting from the second moment need to be obtained as the N target angular velocities.
[0100] In step 222, based on the i-th target angular velocity and the acquisition interval, the i-th pose increment data is determined.
[0101] Here, the pose can be represented by a quaternion. A quaternion is an extended complex number structure, consisting of a real part and three imaginary parts, and is used to represent rotations and transformations in three-dimensional space. Therefore, the pose increment data can be represented by a quaternion. Among them, the i-th target angular velocity and the acquisition interval can be calculated through formula (3) to obtain the i-th pose increment data. See formula (3):
[0102]
[0103] Among them, ω represents the i-th target angular velocity, Δt represents the acquisition interval, and Δq represents the i-th pose increment data.
[0104] In step 223, the first pose is determined based on the initial pose data and each pose increment data.
[0105] Here, based on the product of the first pose increment data and the initial pose data, the first pose is determined; based on the product of the i-th pose increment data and the (i - 1)-th pose, the i-th pose is determined, where i = 2,..., N. For example, for the first pose increment data, the product between the first pose increment data and the initial pose data is used as the first pose corresponding to the first pose increment data; for the second pose increment data, the product between the second pose increment data and the first pose is used as the second pose corresponding to the second pose increment data; and so on, the N-th pose is the first pose. Among them, this process can refer to formula (4):
[0106]
[0107] Among them, q(t + Δt) represents the i-th pose, q(t) represents the (i - 1)-th pose, and Δq represents the incremental data of the i-th pose.
[0108] In the embodiment of the present application, the initial motion data further includes an initial pose. Based on the second moment and the first moment, N target angular velocities are obtained from multiple angular velocities; based on the i-th target angular velocity and the acquisition interval, the incremental data of the i-th pose is determined; based on the initial pose data and the incremental data of each pose, the first pose is determined. In this way, the speed and position corresponding to each acquisition moment can be determined through a recursive process, and the calculation method is simple and fast, thereby improving the efficiency of data processing. In this way, the pose corresponding to each acquisition moment, including the first pose, can be determined through a recursive process, and the calculation method is simple and fast, thereby improving the efficiency of data processing.
[0109] In step 1023, based on the first speed, the first position, and the first pose, the target point cloud of the dynamic obstacle in the environment at the first moment is determined from the first point cloud dataset.
[0110] Here, the first speed, the first position, and the first pose are used to describe the motion state of the target object at the first moment. According to the motion state of the target object at the first moment, the future position of each point in the environment can be predicted. If the predicted position does not match the actual position, the object corresponding to this point may be a dynamic obstacle, so that the target point cloud of the dynamic obstacle can be determined from the first point cloud dataset.
[0111] In the embodiment of the present application, the motion state dataset includes the motion state data collected by the second sensor at multiple acquisition moments. The motion state data includes acceleration and angular velocity; based on the accelerations corresponding to multiple acquisition moments, the first speed and the first position of the target object at the first moment are determined, and the first moment is any acquisition moment; based on the angular velocities corresponding to multiple acquisition moments, the first pose of the target object at the first moment is determined; based on the first speed, the first position, and the first pose, the target point cloud of the dynamic obstacle in the environment at the first moment is determined from the first point cloud dataset. In this way, according to the motion state of the target object at the first moment, the target point cloud of the dynamic obstacle at the first moment is determined, which can avoid the problem of determining the static object in the environment as a dynamic obstacle due to the movement of the target object itself, thereby improving the accuracy of determining the target point cloud.
[0112] In some embodiments, referring to Figure 3E , step 1023 can be implemented through steps 231 to 234, including:
[0113] In step 231, the first point cloud data corresponding to the first moment is obtained from the first point cloud dataset.
[0114] Here, the first point cloud dataset includes the point cloud data collected by the first sensor at multiple acquisition moments. The point cloud data collected by the first sensor at the first moment is obtained from the multiple point cloud data as the first point cloud data.
[0115] In step 232, for each first point cloud in the first point cloud data, based on the first velocity, the first position, and the first pose, the state of the first point cloud is predicted to obtain predicted state data.
[0116] Here, the first point cloud data includes the point cloud corresponding to all points in the environment where the target object is located at the first moment. Therefore, the first point cloud data includes multiple first point clouds. The first point cloud will move due to the movement of the target object in the point cloud data. For each first point cloud in the first point cloud data, according to the first velocity, the first position, and the first pose of the target object at the first moment, the position where the first point cloud should be located at the next acquisition moment is predicted, that is, the predicted state data is obtained. The predicted state data is used to represent the position where the first point cloud should be located at the next acquisition moment predicted according to the movement of the target object.
[0117] In step 233, the reference state data of the first point cloud is obtained from the first point cloud dataset, and the state difference between the predicted state data and the reference state data is determined.
[0118] Here, the reference state data of the first point cloud refers to the point cloud data collected at the next acquisition moment after the first moment in the first point cloud dataset, that is, the moments corresponding to the reference state data and the first moment are two adjacent moments, and the moment corresponding to the reference state data is after the first moment. The difference in the positions of the first point cloud in the two cases, or the difference in the poses of the first point cloud in the two cases, can be determined according to the predicted state data and the reference state data of the first point cloud as the state difference.
[0119] In step 234, when the state difference indicates that the first point cloud has moved, the first point cloud is determined as the target point cloud.
[0120] Here, even if the first point cloud is static, there may be slight errors between the predicted state data and the reference state data. Therefore, an error range can be preset. When the state difference is within the preset error range, it indicates that the first point cloud is static. When the state difference exceeds the preset error range, indicating that the first point cloud has moved, the first point cloud is determined as the target point cloud.
[0121] In the embodiments of the present application, the first point cloud dataset includes point cloud data collected by the first sensor at multiple acquisition moments. The first point cloud data corresponding to the first moment is obtained from the first point cloud dataset. For each first point cloud in the first point cloud data, based on the first speed, the first position, and the first pose, the state of the first point cloud is predicted to obtain predicted state data. The reference state data of the first point cloud is obtained from the first point cloud dataset, and the state difference between the predicted state data and the reference state data is determined. When the state difference indicates that the first point cloud has moved, the first point cloud is determined as the target point cloud. In this way, by predicting the state of the first point cloud and combining the existing reference state data in the first point cloud data, it is possible to simply and quickly determine whether the first point cloud has moved, and then determine the target point cloud, improving the efficiency and accuracy of determining the target point cloud.
[0122] In step 103, complementary point clouds of the dynamic obstacle are determined based on the target point cloud.
[0123] Here, the target point cloud is sometimes incomplete. For example, if the upper body of a pedestrian moves, only the point cloud of the upper body may be determined as the target point cloud, while in fact, the point clouds of the entire pedestrian belong to the point clouds of the dynamic obstacle. Therefore, it is necessary to determine the complementary point clouds of the dynamic obstacle based on the target point cloud to supplement the target point cloud. It can be understood that the complementary point clouds are the point clouds of the dynamic obstacle not included in the target point cloud. Among them, the neighborhood search algorithm (Radius Nearest Neighbors, Radius-NN) can be used to expand the target point cloud to achieve a more complete filtering of the dynamic point cloud.
[0124] In some embodiments, referring to Figure 3F , step 103 can be implemented through steps 1031 to 1034, including:
[0125] In step 1031, the edge point clouds located at the edge positions of the dynamic obstacle are determined from the target point cloud.
[0126] Here, for each target point cloud, it can be determined whether the surroundings of the target point cloud all belong to the target point cloud. If there is no target point cloud in any direction of the target point cloud, it means that this target point cloud is located at the edge position of the dynamic obstacle, and this target point cloud is determined as the edge point cloud.
[0127] In step 1032, for each edge point cloud, the point clouds whose distances from the edge point cloud are less than or equal to the preset distance are determined as candidate point clouds.
[0128] Here, a circular area centered on the edge point cloud can be determined with a preset distance as the radius, and the point cloud within the circular area is determined as the candidate point cloud for this edge point cloud. Among them, the preset distance can be flexibly set, and the preset distances corresponding to edge point clouds at different positions can be different.
[0129] In step 1033, the point cloud that does not belong to the target point cloud among the candidate point clouds is determined as the supplementary point cloud corresponding to the edge point cloud.
[0130] Here, since the candidate point clouds may include other target point clouds, it is necessary to determine the point clouds that do not belong to the target point cloud from the candidate point clouds and determine them as the supplementary point clouds corresponding to the edge point cloud.
[0131] In step 1034, the supplementary point cloud corresponding to each edge point cloud is determined as the supplementary point cloud of the dynamic obstacle.
[0132] Here, a corresponding supplementary point cloud will be determined for each edge point cloud, and the supplementary point cloud corresponding to each edge point cloud is determined as the supplementary point cloud of the dynamic obstacle.
[0133] In the embodiment of the present application, the edge point clouds located at the edge positions of the dynamic obstacles are determined from the target point clouds; for each edge point cloud, the point clouds whose distance from the edge point cloud is less than or equal to the preset distance are determined as candidate point clouds; the point clouds that do not belong to the target point cloud among the candidate point clouds are determined as the supplementary point clouds corresponding to the edge point cloud; the supplementary point cloud corresponding to each edge point cloud is determined as the supplementary point cloud of the dynamic obstacle. In this way, by supplementing the target point cloud with the supplementary point cloud, the point cloud corresponding to the complete dynamic obstacle can be determined, avoiding the problem of missing target point clouds.
[0134] In step 104, the target point cloud and the supplementary point cloud are deleted from the first point cloud dataset to obtain the second point cloud dataset.
[0135] Here, the target point cloud and the supplementary point cloud in the first point cloud dataset are deleted, that is, the point cloud corresponding to the dynamic obstacle is filtered out from the first point cloud dataset, and the obtained second point cloud dataset only includes the point cloud corresponding to static objects.
[0136] In step 105, a training dataset is constructed based on the second point cloud dataset.
[0137] Here, a training dataset is constructed according to the second point cloud dataset. The training dataset is used to train the binocular depth prediction model, and the trained binocular depth prediction model can accurately predict the depth information of the binocular image.
[0138] In the embodiments of the present application, a motion state data set of a target object and a first point cloud data set of the environment where the target object is located are obtained. The first point cloud data set is collected by a first sensor, and the motion state data set is collected by a second sensor built in the first sensor. In this way, by collecting the motion state data set with the second sensor built in the first sensor, the number of sensors required for data collection can be reduced, and the convenience and efficiency of data collection can be improved. Moreover, since the first sensor and the second sensor are arranged together, it is possible to ensure that the motion state data set and the first point cloud data set are collected simultaneously, improving the coordination of data collection. Based on the motion state data set, target point clouds of dynamic obstacles in the environment are determined from the first point cloud data set; supplementary point clouds of the dynamic obstacles are determined based on the target point clouds; the target point clouds and the supplementary point clouds are deleted from the first point cloud data set to obtain a second point cloud data set. In this way, after the target point clouds of the dynamic obstacles in the environment are determined, the supplementary point clouds are further determined, and the point clouds corresponding to the dynamic obstacles can be completely deleted from the first point cloud data set, thereby improving the accuracy of the obtained second point cloud data set. A training data set for training a binocular depth prediction model is constructed based on the second point cloud data set, which can improve the quality of the training data set, thereby improving the performance of the binocular depth prediction model after training. Therefore, through the embodiments of the present application, the convenience and efficiency of data collection can be improved, the efficiency and accuracy of obtaining the second point cloud data set can be improved, and the efficiency of training the binocular depth prediction model and the performance of the trained model can be improved.
[0139] In some embodiments, referring to Figure 3G , step 105 can be implemented through steps 1051 to 1055 and includes:
[0140] In step 1051, the camera internal parameters and the baseline distance of the binocular image acquisition device are obtained.
[0141] Here, the binocular image acquisition device is a binocular RGB camera, which is a device that uses two cameras to simulate the stereo vision of human eyes. It can capture images of the same scene from two different perspectives simultaneously to obtain binocular images. The camera internal parameters refer to the parameters related to the physical characteristics and imaging geometry of the camera itself, which are used to convert the image pixel coordinates into coordinates in the camera coordinate system or perform image distortion correction. For example, assume that the camera internal parameters of the two cameras in the binocular RGB camera are the same, namely fx, fy, cx, and cy. Among them, fx (focal length in the x-axis direction) and fy (focal length in the y-axis direction) are the focal lengths of the camera on the x-axis and y-axis, and cx (the image center in the x-axis direction, usually half of the image width) and cy (the image center in the y-axis direction, usually half of the image height) are the coordinates of the image center (optical center) on the x-axis and y-axis. The baseline distance refers to the horizontal distance between the two cameras, which is directly related to the depth accuracy that the system can perceive. The longer the baseline distance, the greater the depth change that the system can distinguish, but at the same time, it will reduce the depth of the field of view (i.e., the depth range where clear imaging can be achieved).
[0142] In step 1052, for each second point cloud data, based on the camera internal parameters and the baseline distance, determine the disparity map corresponding to the second point cloud data.
[0143] Here, the second point cloud data set includes multiple second point cloud data. For each second point cloud data, project the second point cloud data onto the image plane of the left camera according to the camera internal parameters to obtain pixel coordinates, and then calculate the horizontal position difference of the corresponding point of the second point cloud data on the right camera image according to the baseline distance, that is, the disparity. Fill the calculated disparity values into the corresponding positions in the disparity map. The disparity map is a two-dimensional matrix, and each value in it represents the disparity of the corresponding pixel point. The disparity map is usually expressed as the reciprocal of the depth value (the greater the disparity, the smaller the depth value).
[0144] In step 1053, obtain multiple binocular images collected by the binocular image acquisition device at multiple acquisition times.
[0145] Here, the acquisition times of the binocular image acquisition device correspond to the acquisition times of the first sensor and the second sensor. The binocular image acquisition device will collect multiple binocular images at multiple acquisition times.
[0146] In step 1054, for each binocular image, determine the disparity map corresponding to the binocular image, and determine the binocular image and the disparity map corresponding to the binocular image as the training data corresponding to the binocular image.
[0147] Here, align the binocular image and the disparity map according to the acquisition time, that is, determine the binocular image and the disparity map collected at the same acquisition time as a piece of training data.
[0148] In step 1055, a training data set is constructed based on each piece of training data.
[0149] Here, each piece of training data is constructed into a training data set. That is to say, the training data set includes multiple pieces of training data, and each piece of training data includes the binocular images and the disparity map corresponding to the same acquisition moment.
[0150] In the embodiment of the present application, the camera internal parameters and the baseline distance of the binocular image acquisition device are obtained; for each second point cloud data, based on the camera internal parameters and the baseline distance, the disparity map corresponding to the second point cloud data is determined; multiple binocular images collected by the binocular image acquisition device at multiple acquisition moments are obtained; for each binocular image, the disparity map corresponding to the binocular image is determined, and the binocular image and the disparity map corresponding to the binocular image are determined as the training data corresponding to the binocular image; based on each piece of training data, a training data set is constructed. In this way, by filtering the second point cloud data set of dynamic obstacles and the binocular images to construct the training data set, the quality of the training data set can be improved, thereby improving the model training effect.
[0151] In some embodiments, referring to Figure 4 , after constructing the training data set based on the second point cloud data set, the binocular depth prediction model can be trained through steps 201 to 204, including:
[0152] In step 201, for each piece of training data in the training data set, using the binocular depth prediction model to be trained, the binocular images in the training data are subjected to depth prediction to obtain a prediction result.
[0153] Here, the binocular depth prediction model to be trained refers to the binocular depth prediction model that has not been completed training. At this time, the prediction performance of the model is low and the prediction effect is inaccurate. Each piece of training data includes a binocular image and the disparity map corresponding to the binocular image. Using the binocular depth prediction model to be trained to perform depth prediction on the binocular images in the training data to obtain a prediction result. The prediction result can represent the predicted depth information of the binocular image. Since the depth value and the disparity are related, the prediction result can also be the predicted disparity map.
[0154] In step 202, based on the prediction result and the disparity map in the training data, the loss value of the binocular depth prediction model to be trained is determined.
[0155] Here, the loss value is used to measure the difference between the model prediction result and the true value. Taking the prediction result as the predicted disparity map as an example, the mean square error between the predicted disparity map and the disparity map in the training data can be calculated as the loss value.
[0156] In step 203, the binocular depth prediction model to be trained is trained based on each loss value until the training end condition is reached, and the trained binocular depth prediction model is obtained.
[0157] Here, based on each loss value, the loss value is backpropagated to the binocular depth prediction model to be trained. Based on the loss value, the parameters of the binocular depth prediction model to be trained are adjusted by the gradient descent algorithm. Then, the above steps are sequentially repeated for the next loss value until the training end condition is reached, and the trained binocular depth prediction model is obtained. Among them, the training end condition may be that the loss value is lower than the loss threshold or the difference between the current loss value and the previous adjacent loss value is less than the preset difference.
[0158] In the embodiment of the present application, for each piece of training data in the training dataset, the binocular depth prediction model to be trained is used to perform depth prediction on the binocular images in the training data to obtain a prediction result; based on the prediction result and the disparity map in the training data, the loss value of the binocular depth prediction model to be trained is determined; the binocular depth prediction model to be trained is trained based on each loss value until the training end condition is reached, and the trained binocular depth prediction model is obtained. By using a high-quality training dataset, the training efficiency of the binocular depth prediction model can be improved, and the prediction accuracy and performance of the trained binocular depth prediction model can be improved.
[0159] Next, an exemplary application of the embodiment of the present application in an actual application scenario will be described.
[0160] In the related art, it is necessary to continuously scan the environment using a 3D lidar to obtain a series of point cloud data. At the same time, an odometer is used to record the movement trajectory of the lidar during the acquisition process. Since the collected point cloud data often contains the point clouds of some dynamic obstacles, it is necessary to perform SLAM using the odometer data and the point cloud data, and use point cloud editing software to crop the point clouds of the dynamic obstacles. However, this process requires manual editing and cropping of the point cloud data, which takes a long time and consumes a large amount of human resources.
[0161] In the embodiment of the present application, only a laser radar and the IMU that comes with the laser radar are needed to automatically and efficiently collect data, using the point cloud data (corresponding to the first point cloud data set) collected by the radar (corresponding to the first sensor), and the IMU data (corresponding to the motion state data set) collected by the IMU sensor (corresponding to the second sensor). In addition, the dynamic obstacle point cloud is filtered using the M-detector algorithm in combination with the neighborhood search algorithm (Radius Nearest Neighbors, Radius-NN), which can avoid the process of manually cropping the dynamic obstacle point cloud and achieve a more complete dynamic point cloud filtering. In addition, combining the collected binocular RGB data (corresponding to the binocular image) as a training data set for binocular depth estimation can improve the quality of the generated training data set and improve the model training effect. The following will combine Figure 5 Specific instructions, including:
[0162] In step 501, a motion state data set and a first point cloud data set are collected.
[0163] Here, the first point cloud data set is collected by the laser radar, and the IMU data collected by the IMU sensor provided by the laser radar, that is, the motion state data set.
[0164] In step 502, odometer data is determined based on the motion state data.
[0165] Here, the motion state data set includes acceleration and angular velocity. By integrating acceleration and angular velocity, odometer data is calculated, including information such as position, velocity and attitude (corresponding to the first position, first velocity and first attitude). The integration refers to the integration of discrete time, that is, the accumulation is performed using discrete acquisition moments, including:
[0166] 1) Initialize variables: Initialize the position of the target object before it starts moving and is about to start moving: P = [x, y, z] (initialized to zero or a known starting position); Initialize the speed of the target object before it starts moving and is about to start moving: V = [vx, vy, vz] (initialized to zero); Initialize the posture of the target object before it starts moving and is about to start moving: usually represented by quaternion or rotation matrix, initialized to unit quaternion or unit matrix.
[0167] 2) Integrate the acceleration to get the velocity: Refer to formula (1) and update the velocity by integrating the acceleration data:
[0168] V(t+Δt)=V(t)+a(t)×Δt(1)
[0169] Among them, V(t) represents the velocity at the previous moment, a(t) represents the acceleration at the current moment, and Δt is the time interval.
[0170] 3) Integrate the velocity to obtain the position: Refer to Equation (2), update the position by integrating the velocity data:
[0171] P(t + Δt) = P(t) + V(t + Δt) × Δt (2)
[0172] Where P(t + Δt) represents the current position, P(t) represents the position at the previous moment, V(t + Δt) represents the velocity at the current moment, and Δt represents the acquisition interval.
[0173] 4) Integrate the angular velocity to obtain the attitude: Update the attitude by integrating the angular velocity data, usually using quaternions or rotation matrices. Refer to Equation (4), the attitude represented by quaternions can be updated by Equation (4):
[0174]
[0175] Where represents quaternion multiplication, and Δq is the quaternion increment calculated from the angular velocity (corresponding to the pose increment data). Refer to Equation (3), the quaternion increment is usually calculated using Equation (3):
[0176]
[0177] Where ω represents the angular velocity and Δt represents the acquisition interval.
[0178] In step 503, filter out the target point cloud and supplementary point cloud in the first point cloud dataset to obtain the second point cloud dataset.
[0179] Here, use the M_detector algorithm to process the odometer and the first point cloud dataset to filter out the target point cloud of dynamic obstacles. For the target point cloud calculated by M-detector, perform a neighborhood search on each target point cloud through Radius-NN to filter out the supplementary point cloud of dynamic obstacles not detected by M-detector, and retain the static point cloud to obtain the second point cloud dataset. This is because the dynamic point cloud calculated by the M-detector algorithm is sometimes incomplete (for example, for the point cloud of a pedestrian, only half of the body's point cloud may be calculated as the target point cloud by the M-detector algorithm). Through the Radius-NN algorithm, the remaining point cloud of a complete whole can be searched to supplement the target point cloud calculated by M-detector, solving the problem of missing the point cloud of dynamic obstacles.
[0180] In step 504, determine the disparity map based on the second point cloud dataset.
[0181] Here, assume that the internal parameters of the left RGB camera of the binocular RGB camera are \(f_x\), \(f_y\), \(c_x\), and \(c_y\). Among them, \(f_x\) (focal length in the x-axis direction) and \(f_y\) (focal length in the y-axis direction) are the focal lengths of the camera on the x-axis and y-axis, and \(c_x\) (image center in the x-axis direction, usually half of the image width) and \(c_y\) (image center in the y-axis direction, usually half of the image height) are the coordinates of the image center (optical center) on the x-axis and y-axis. Project the second point cloud dataset from 3D space onto a two-dimensional plane to generate a disparity map disparity.
[0182] Among them, for the left RGB camera, its internal parameter matrix is usually expressed as For a point \(P = [X, Y, Z]\) in the second point cloud dataset T , its homogeneous coordinates are \(P_h = [X, Y, Z, 1]\) T . The pixel coordinates \((u, v)\) of the projection of this point onto the image plane in the camera coordinate system can be calculated by the following formula (5):
[0183]
[0184] Since the depth map usually stores the depth value \(Z\) (or the disparity value inversely proportional to \(Z\), taking the depth value as an example here), the above formula (5) can be simplified to formula (6):
[0185]
[0186] After that, the corresponding pixel coordinates \((u, v)\) can be calculated, and the depth value \(Z\) is filled at the \((u, v)\) position of the depth map (or \(u\) and \(v\) are rounded to the nearest pixel position) to obtain the depth image.
[0187] Assume that the focal length \(f\) of the camera is known (usually assume that the focal lengths of the two cameras are the same) and the baseline distance \(B\), then the pixel value \(D(u, v)\) of the disparity map can be associated with the depth \(Z(u, v)\) by the following formula (7):
[0188]
[0189] Here, \(D(u, v)\) is the disparity value at the pixel \((u, v)\), indicating the horizontal pixel difference between the corresponding points of this pixel in the left and right images in the stereo image pair. Calculate \(D(u, v)\) for all corresponding pixels to obtain the disparity map disparity.
[0190] In step 505, collect binocular images.
[0191] Here, use a binocular camera to obtain binocular images, that is, a pair of binocular RGB images img_dual.
[0192] In step 506, frame synchronization is performed on the disparity map and the binocular images.
[0193] Here, frame synchronization is performed on the disparity map disparity and the binocular RGB image pair img_dual to ensure that the timestamps of the images are consistent, obtaining the synchronized disparity_syn and img_dual_syn.
[0194] In step 507, the synchronized disparity map and binocular images are output.
[0195] Here, the synchronized disparity map and binocular images are output, and the synchronized disparity map and binocular images are displayed or stored.
[0196] In the embodiment of the present application, the number of sensors required in the acquisition process is reduced. The traditional method requires an odometer, a 3D lidar, and a binocular RGB image acquisition device. The method in the embodiment of the present application only requires a 3D lidar with an IMU and a binocular RGB image acquisition device. Moreover, the embodiment of the present application simplifies the data acquisition steps and no longer requires prior map building, saving a large amount of manpower for later manual cropping of dynamic point clouds. In addition, in the embodiment of the present application, the M-detector algorithm is combined with the neighborhood search algorithm (Radius Nearest Neighbors, Radius-NN) to filter the dynamic obstacle point clouds, which can avoid the process of manually cropping dynamic obstacle point clouds and achieve relatively complete filtering of dynamic point clouds.
[0197] Next, the exemplary structure of the software module for implementing the data processing device 233 provided in the embodiment of the present application will be continued. In some embodiments, as Figure 2 shown, the software module stored in the data processing device 233 in the memory 230 may include:
[0198] An acquisition module 2331, configured to acquire a motion state data set of a target object and a first point cloud data set of the environment where the target object is located. The first point cloud data set is acquired by a first sensor, and the motion state data set is acquired by a second sensor built in the first sensor;
[0199] A determination module 2332, configured to determine, based on the motion state data set, target point clouds of dynamic obstacles in the environment from the first point cloud data set;
[0200] The determination module 2332 is further configured to determine supplementary point clouds of the dynamic obstacles based on the target point clouds;
[0201] A deletion module 2333, configured to delete the target point clouds and the supplementary point clouds from the first point cloud data set to obtain a second point cloud data set;
[0202] A building block 2334 for constructing a training dataset based on the second point cloud dataset, where the training dataset is used to train a binocular depth prediction model.
[0203] In some embodiments, the determining module 2332 is further configured to determine a first velocity and a first position of the target object at a first moment based on the accelerations corresponding to the multiple acquisition moments, where the first moment is any acquisition moment; determine a first pose of the target object at the first moment based on the angular velocities corresponding to the multiple acquisition moments; and determine target point clouds of dynamic obstacles in the environment at the first moment from the first point cloud dataset based on the first velocity, the first position, and the first pose.
[0204] In some embodiments, the determining module 2332 is further configured to obtain initial motion data of the target object, where the initial motion data includes at least an initial velocity and an initial position; obtain N target accelerations from multiple accelerations based on a second moment when the initial motion data is acquired and the first moment, where N is a positive integer; determine a first target velocity based on the first target acceleration, the initial velocity, and an acquisition interval; determine an i-th target velocity based on the i-th target acceleration, the (i - 1)-th target velocity, and the acquisition interval, where the acquisition interval is a time interval between two adjacent acquisition moments, and i = 2,..., N; determine the N-th target velocity as the first velocity; and determine the first position based on each of the target velocities, the initial position, and the sampling interval.
[0205] In some embodiments, the determining module 2332 is further configured to obtain N target angular velocities from multiple angular velocities based on the second moment and the first moment; determine an i-th pose increment data based on the i-th target angular velocity and the acquisition interval; and determine the first pose based on the initial pose data and each of the pose increment data.
[0206] In some embodiments, the determining module 2332 is further configured to obtain first point cloud data corresponding to the first moment in the first point cloud dataset; for each first point cloud in the first point cloud data, perform state prediction on the first point cloud based on the first velocity, the first position, and the first pose to obtain predicted state data; obtain reference state data of the first point cloud in the first point cloud dataset, and determine a state difference between the predicted state data and the reference state data; and when the state difference indicates that the first point cloud has moved, determine the first point cloud as a target point cloud.
[0207] In some embodiments, the determining module 2332 is further configured to determine, from the target point cloud, edge point clouds located at the edge positions of the dynamic obstacles; for each of the edge point clouds, determine point clouds whose distances from the edge point clouds are less than or equal to a preset distance as candidate point clouds; determine the point clouds that do not belong to the target point cloud among the candidate point clouds as supplementary point clouds corresponding to the edge point clouds; and determine the supplementary point clouds corresponding to each of the edge point clouds as supplementary point clouds of the dynamic obstacles.
[0208] In some embodiments, the constructing module 2334 is further configured to obtain the camera internal parameters and the baseline distance of the binocular image acquisition device; for each of the second point cloud data, determine a disparity map corresponding to the second point cloud data based on the camera internal parameters and the baseline distance; obtain a plurality of binocular images acquired by the binocular image acquisition device at a plurality of acquisition times; for each of the binocular images, determine the disparity map corresponding to the binocular image, and determine the binocular image and the disparity map corresponding to the binocular image as training data corresponding to the binocular image; and construct the training data set based on each of the training data.
[0209] In some embodiments, the constructing module 2334 is further configured to, for each piece of training data in the training data set, use a binocular depth prediction model to be trained to perform depth prediction on the binocular images in the training data to obtain a prediction result; determine a loss value of the binocular depth prediction model to be trained based on the prediction result and the disparity map in the training data; and train the binocular depth prediction model to be trained based on each of the loss values until a training end condition is reached to obtain a trained binocular depth prediction model.
[0210] An embodiment of the present application provides a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the data processing method in the above embodiments of the present application.
[0211] An embodiment of the present application provides a computer-readable storage medium, in which computer executable instructions or a computer program are stored. When the computer executable instructions or the computer program are executed by a processor, the processor will be caused to execute the data processing method provided in the embodiments of the present application, for example, Figure 3A the data processing method shown.
[0212] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0213] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and they may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0214] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored in a part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or code portions).
[0215] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected through a communication network.
[0216] In summary, through the embodiments of the present application, a motion state data set of a target object and a first point cloud data set of the environment where the target object is located are obtained. The first point cloud data set is collected by a first sensor, and the motion state data set is collected by a second sensor built into the first sensor. In this way, by collecting the motion state data set with the second sensor built into the first sensor, the number of sensors required for data collection can be reduced, and the convenience and efficiency of data collection can be improved. Moreover, since the first sensor and the second sensor are arranged together, it is possible to ensure that the motion state data set and the first point cloud data set are collected simultaneously, improving the coordination of data collection. Based on the motion state data set, target point clouds of dynamic obstacles in the environment are determined from the first point cloud data set; supplementary point clouds of the dynamic obstacles are determined based on the target point clouds; the target point clouds and the supplementary point clouds are deleted from the first point cloud data set to obtain a second point cloud data set. In this way, after determining the target point clouds of the dynamic obstacles in the environment, the supplementary point clouds are further determined, and the point clouds corresponding to the dynamic obstacles can be completely deleted from the first point cloud data set, thereby improving the accuracy of the obtained second point cloud data set. Constructing a training data set for training a binocular depth prediction model based on the second point cloud data set can improve the quality of the training data set, thereby improving the performance of the binocular depth prediction model after training. Therefore, through the embodiments of the present application, the convenience and efficiency of data collection can be improved, the efficiency and accuracy of obtaining the second point cloud data set can be improved, and the efficiency of training the binocular depth prediction model and the performance of the trained model can be improved.
[0217] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Acquire a motion state data set of a target object and a first point cloud data set of an environment where the target object is located, wherein the first point cloud data set is collected by a first sensor, and the motion state data set is collected by a second sensor built into the first sensor; Based on the motion state data set, determining a target point cloud of a dynamic obstacle in the environment from the first point cloud data set; Determine a supplementary point cloud of the dynamic obstacle based on the target point cloud; Deleting the target point cloud and the supplementary point cloud from the first point cloud data set to obtain a second point cloud data set; A training data set is constructed based on the second point cloud data set, and the training data set is used to train a binocular depth prediction model.
2. The method according to claim 1, characterized in that The motion state data set includes motion state data collected by the second sensor at multiple collection moments, and the motion state data includes acceleration and angular velocity; The determining, based on the motion state data set, from the first point cloud data set, a target point cloud of the dynamic obstacle in the environment comprises: Determine, based on the accelerations corresponding to the multiple acquisition moments, a first speed and a first position of the target object at a first moment, where the first moment is any acquisition moment; Determining a first pose of the target object at the first moment based on the angular velocities corresponding to the multiple acquisition moments; Based on the first speed, the first position, and the first posture, a target point cloud of the dynamic obstacle in the environment at the first moment is determined from the first point cloud data set.
3. The method according to claim 2, characterized in that The determining, based on the accelerations corresponding to the multiple acquisition moments, a first velocity and a first position of the target object at a first moment, comprises: Acquire initial motion data of the target object, wherein the initial motion data at least includes an initial speed and an initial position; Based on the second moment when the initial motion data is collected and the first moment, obtaining N target accelerations from a plurality of accelerations, where N is a positive integer; Determine a first target speed based on the first target acceleration, the initial speed and the acquisition interval; Determine the i-th target speed based on the i-th target acceleration, the i-1-th target speed and the acquisition interval, wherein the acquisition interval is the time interval between two adjacent acquisition moments, i=2, ..., N; determining the Nth target speed as the first speed; The first position is determined based on each of the target speed, the initial position, and the sampling interval.
4. The method according to claim 3, characterized in that The initial motion data also includes an initial posture, and determining the first posture of the target object at the first moment based on the angular velocities corresponding to the multiple acquisition moments includes: Based on the second moment and the first moment, acquiring N target angular velocities from a plurality of angular velocities; Determining the i-th position and posture incremental data based on the i-th target angular velocity and the acquisition interval; The first posture is determined based on the initial posture data and each of the posture incremental data.
5. The method according to claim 2, characterized in that: The first point cloud data set includes point cloud data collected by the first sensor at multiple collection times. The determining, based on the first speed, the first position, and the first posture, from the first point cloud data set, a target point cloud of the dynamic obstacle in the environment at the first moment comprises: Acquire first point cloud data corresponding to the first moment in the first point cloud data set; For each first point cloud in the first point cloud data, based on the first speed, the first position and the first posture, perform state prediction on the first point cloud to obtain predicted state data; Acquire reference state data of the first point cloud in the first point cloud data set, and determine a state difference between the predicted state data and the reference state data; When the state difference indicates that the first point cloud moves, the first point cloud is determined to be a target point cloud.
6. The method according to claim 1, characterized in that The determining the supplementary point cloud of the dynamic obstacle based on the target point cloud comprises: Determining an edge point cloud located at an edge position of the dynamic obstacle from the target point cloud; For each of the edge point clouds, determining a point cloud whose distance to the edge point cloud is less than or equal to a preset distance as a candidate point cloud; Determine the point cloud that does not belong to the target point cloud in the candidate point cloud as the supplementary point cloud corresponding to the edge point cloud; The supplementary point cloud corresponding to each of the edge point clouds is determined as the supplementary point cloud of the dynamic obstacle.
7. The method according to claim 1, characterized in that The second point cloud data set includes a plurality of second point cloud data, and the step of constructing a training data set based on the second point cloud data set includes: Obtain camera intrinsic parameters and baseline distance of a binocular image acquisition device; For each second point cloud data, determining a disparity map corresponding to the second point cloud data based on the camera intrinsic parameter and the baseline distance; Acquire a plurality of binocular images acquired by the binocular image acquisition device at a plurality of acquisition moments; For each of the binocular images, determining a disparity map corresponding to the binocular image, and determining the binocular image and the disparity map corresponding to the binocular image as training data corresponding to the binocular image; Based on each of the training data, the training data set is constructed.
8. The method according to claim 5, characterized in that After constructing the training data set based on the second point cloud data set, the method further includes: For each training data in the training data set, using the binocular depth prediction model to be trained, performing depth prediction on the binocular image in the training data to obtain a prediction result; Determining a loss value of the binocular depth prediction model to be trained based on the prediction result and the disparity map in the training data; The binocular depth prediction model to be trained is trained based on each of the loss values until a training end condition is reached, thereby obtaining a trained binocular depth prediction model.
9. A data processing device, characterized in that: The device comprises: An acquisition module, used to acquire a motion state data set of a target object and a first point cloud data set of an environment where the target object is located, wherein the first point cloud data set is acquired by a first sensor, and the motion state data set is acquired by a second sensor built into the first sensor; a determination module, configured to determine a target point cloud of a dynamic obstacle in the environment from the first point cloud data set based on the motion state data set; The determination module is further used to determine the supplementary point cloud of the dynamic obstacle based on the target point cloud; A deletion module, used for deleting the target point cloud and the supplementary point cloud from the first point cloud data set to obtain a second point cloud data set; A construction module is used to construct a training data set based on the second point cloud data set, and the training data set is used to train a binocular depth prediction model.
10. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, configured to implement the method according to any one of claims 1 to 8 when executing computer executable instructions or computer programs stored in the memory.
11. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer programs are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product comprising computer executable instructions or a computer program, characterized in that: When the computer executable instructions or computer programs are executed by a processor, the method according to any one of claims 1 to 8 is implemented.