A system, device and method for generating high-resolution and high-precision point clouds

By combining sensor data of low-resolution LiDAR and stereo cameras, high-resolution and high-precision point clouds are generated using machine learning algorithms, which solves the costly problems in the existing technology and achieves the effect of low-cost and high-precision point cloud generation.

CN113039579BActive Publication Date: 2025-07-01HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980075560.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-11-19
Filing Date
2019-11-19
Publication Date
2025-07-01
Estimated Expiration
2039-11-19

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-resolution and high-precision point clouds at low cost, especially when using expensive high-resolution LiDARs, which are expensive and unsuitable for commercial use.

Method used

By combining sensor data from low-resolution LiDAR and stereo cameras, high-resolution and high-precision point clouds are generated using machine learning algorithms. The specific steps include receiving camera and LiDAR point cloud data, determining the error of the camera point cloud, determining the correction function based on the error, generating a correction point cloud using the correction function, and updating the correction function repeatedly until the training error is less than the threshold.

Benefits of technology

It realizes the generation of high-resolution and high-precision point clouds under low-cost conditions, reduces the dependence on high-resolution LiDAR, is suitable for various point cloud application algorithms, and improves the accuracy and efficiency of environmental understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113039579B_ABST
    Figure CN113039579B_ABST
Patent Text Reader

Abstract

A system, device, and method for generating a high-resolution and high-precision point cloud. On the one hand, a computer vision system receives a camera point cloud from a camera system and a LiDAR (Light Detection and Ranging) point cloud from a LiDAR system; determines the error of the camera point cloud with the LiDAR point cloud as a reference; determines a correction function based on the determined error; uses the correction function to generate a corrected point cloud based on the camera point cloud; determines the training error of the corrected point cloud with a first LiDAR point cloud as a reference; updates the correction function based on the determined training error. At the end of training, the correction function can be used by the computer vision system to generate a generated high-resolution and high-precision point cloud based on the camera point cloud provided by the camera system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision, and more particularly to a system, device, and method for generating high-resolution and high-precision point clouds. Background Art

[0002] LiDAR (Light Detection and Ranging) point clouds can be used for point cloud map generation. Point cloud map generation includes: moving around in the environment to scan LiDAR (i.e., driving a car equipped with LiDAR along the block), collecting all the point clouds generated by LiDAR, and merging the generated point clouds together to generate a point cloud map. The generated point cloud map includes a larger set of 3D data points than traditional point clouds and has extended boundaries. The point cloud map can be used in vehicle driver assistance systems or autonomous vehicles to achieve map-based vehicle positioning in autonomous driving.

[0003] Computer vision systems rely on accurate and reliable sensor data and implement machine learning algorithms that can understand the sensor data. Computer vision systems are an important component in many applications of vehicle driver assistance systems and autonomous vehicles. There are various machine learning algorithms for generating two-dimensional (2D) or three-dimensional (3D) representations of the environment, including one or more of object detection algorithms, dynamic object removal algorithms, simultaneous localization and mapping (SLAM) algorithms, point cloud map generation algorithms, high-definition map creation algorithms, LiDAR-based positioning algorithms, semantic mapping algorithms, tracking algorithms, or scene reconstruction algorithms. These machine learning algorithms require accurate sensor data to provide efficient and reliable results. Point clouds are one of the most common data types for representing 3D environments. Point cloud maps can be generated when computer vision systems perform spatial perception. A point cloud is a collection of many data points in a coordinate system, especially in a 3D coordinate system. Each data point in the point cloud has three coordinates, namely, x, y, and z coordinates, which determine the positions of the data points along the x, y, and z axes respectively in the 3D coordinate system. LiDAR sensors or stereo cameras can be used to sense the environment and generate point clouds of the environment based on the sensor data captured by the LiDAR sensors or stereo cameras. Point clouds generated by individual sensors have some drawbacks.

[0004] The point clouds generated using LiDAR are very accurate, but there are two main problems. First, high-resolution LiDAR is usually very expensive compared to stereo cameras. Second, cheaper LiDARs such as 8-beam, 16-beam, and 32-beam LiDARs have sparsity (low resolution). On the other hand, stereo cameras are cheaper than LiDARs, but the sensor data captured by stereo cameras (e.g., the image data representing the images captured by stereo cameras) has high resolution. However, the limitation of stereo cameras is that the point clouds generated based on the sensor data captured by stereo cameras (e.g., the image data representing the images captured by stereo cameras) have low accuracy. That is, the spatial coordinates of the data points in the obtained point clouds are inaccurate. Therefore, the low resolution of the point clouds generated by LiDARs with fewer beams or the low accuracy of the point clouds generated based on the image data representing the images captured by stereo cameras will have an adverse impact on any machine learning algorithms that use these point clouds to generate an understanding of the environment.

[0005] A great deal of research has been conducted on how to overcome the limitations caused by low accuracy and low resolution of point clouds. One group of algorithms focuses on improving the resolution of low-resolution LiDAR point clouds. Another group of algorithms uses high-resolution LiDAR data and other sensor data, e.g., the images captured by stereo cameras, to generate point clouds with very high density. The latter group of algorithms combines and utilizes the point clouds generated based on the sensor data collected by high-resolution / high-accuracy LiDAR and the point clouds generated based on the sensor data collected by high-resolution stereo cameras. This is a quite technically challenging task and requires both sensors at the same time. Since high-resolution LiDAR is very expensive, this cannot be achieved in many cases.

[0006] Another technique for generating high-resolution / high-accuracy point clouds is to fuse the low-resolution point clouds generated by low-resolution LiDAR (with fewer beams) with the sensor data received from other sensors. In machine learning-based methods, since the sensors used during the training process are the same as those used in the test mode, the fusion of other sensing data is built into the solution. In machine learning algorithms, classical fusion methods are used to generate point clouds. The latter group of algorithms requires both sensors at the same time.

[0007] In view of this, there is still a need for an effective low-cost solution for generating high-resolution and high-accuracy point clouds. Summary of the Invention

[0008] The present invention provides a system, device, and method for generating high-resolution and high-precision point clouds based on sensor data, where the sensor data comes from at least two different types of sensors, for example, low-resolution LiDAR and stereo cameras, without using expensive sensors such as high-resolution LiDAR. Such systems are costly and not suitable for commercial use. Therefore, the purpose of the present invention is to provide an effective and / or efficient solution for generating high-resolution and high-precision point clouds using a combination of cheaper sensors. The resulting high-resolution and high-precision point clouds can be used in one or more point cloud application algorithms that use point clouds, especially those that require high-resolution and / or high-precision point clouds for effective operation. Examples of point cloud application algorithms that can use high-resolution and high-precision point clouds include one or more of object detection algorithms, dynamic object removal algorithms, simultaneous localization and mapping (SLAM) algorithms, point cloud map generation algorithms, high-definition map creation algorithms, LiDAR-based localization algorithms, semantic mapping algorithms, tracking algorithms, or scene reconstruction algorithms. These algorithms are typically used in 3D but may also be used in 2D. The output of the point cloud application algorithm can be used to generate a 3D representation of the environment. When the host device is a vehicle, the 3D representation of the environment can be displayed on a display of the vehicle's console or dashboard.

[0009] According to a first aspect of the present invention, a computer vision system is provided. According to an embodiment of the first aspect of the present invention, the computer vision system includes a processor system, a first LiDAR system coupled to the processor system and having a first resolution, a camera system coupled to the processor system, and a memory coupled to the processor system, where executable instructions are tangibly stored on the memory, and when the executable instructions are executed by the processor system, cause the computer vision system to perform a number of operations: receive camera point clouds from the camera system; receive first LiDAR point clouds from the first LiDAR system; determine the error of the camera point clouds with reference to the first LiDAR point clouds; determine a correction function based on the determined error; use the correction function to generate corrected point clouds based on the camera point clouds; determine the training error of the corrected point clouds with reference to the first LiDAR point clouds; and update the correction function based on the determined training error.

[0010] In some implementations of the foregoing aspects and embodiments, when the executable instructions are executed by the processor system, cause the computer vision system to repeatedly perform some of the above operations until the training error is less than an error threshold.

[0011] In some implementations of the foregoing aspects and embodiments, when the executable instructions for determining the error of the camera point cloud with reference to the first LiDAR point cloud and determining a correction function based on the determined error are executed by the processor system, the computer vision system is caused to: generate a mathematical representation (e.g., a model) of the error of the camera point cloud with reference to the first LiDAR point cloud; determine a correction function based on the mathematical representation of the error. In some implementations of the foregoing aspects and embodiments, when the executable instructions for determining the training error of the corrected point cloud with reference to the first LiDAR point cloud and updating the correction function based on the determined training error are executed by the processor system, the computer vision system is caused to: generate a mathematical representation of the training error of the corrected point cloud with reference to the first LiDAR point cloud; update the correction function based on the mathematical representation of the training error.

[0012] In some implementations of the foregoing aspects and embodiments, the computer vision system further includes a second LiDAR system coupled to the processor system and having a second resolution, wherein the second resolution of the second LiDAR system is higher than the first resolution of the first LiDAR system. When the executable instructions are executed by the processor system, the computer vision system is caused to: (a) receive a second LiDAR point cloud from the second LiDAR system; (b) determine the error of the corrected point cloud with reference to the second LiDAR point cloud; (c) determine a correction function based on the determined error; (d) generate a corrected point cloud using the correction function; (e) determine a second training error of the corrected point cloud with reference to the second LiDAR point cloud; (f) update the correction function based on the determined second training error.

[0013] In some implementations of the foregoing aspects and embodiments, when the executable instructions are executed by the processor system, the computer vision system repeatedly performs operations (d) to (f) until the second training error is less than an error threshold.

[0014] In some implementations of the foregoing aspects and embodiments, the second LiDAR system includes one or more 64-beam LiDAR units, and the first LiDAR system includes one or more 8-beam, 16-beam, or 32-beam LiDAR units.

[0015] In some implementations of the foregoing aspects and embodiments, the camera system is a stereo camera system, and the camera point cloud is a stereo camera point cloud.

[0016] In some implementations of the foregoing aspects and embodiments, the processor system includes a neural network.

[0017] According to another embodiment of the first aspect of the present invention, a computer vision system is provided, including a processor system, a first LiDAR system coupled to the processor system and having a first resolution, a camera system coupled to the processor system, and a memory coupled to the processor system. Wherein, executable instructions are tangibly stored on the memory, and when the executable instructions are executed by the processor system, the computer vision system is caused to perform a number of operations: receiving a camera point cloud from the camera system; receiving a first LiDAR point cloud from the first LiDAR system; determining an error of the camera point cloud with reference to the first LiDAR point cloud; determining a correction function based on the determined error; using the correction function to generate a corrected point cloud based on the camera point cloud; using the corrected point cloud to calculate an output of a point cloud application algorithm, and using the first LiDAR point cloud to calculate the output of the point cloud application algorithm; using the corrected point cloud to determine a loss of the error of the output of the point cloud application algorithm through a loss function; and determining the correction function based on the determined loss.

[0018] In some implementations of the foregoing aspects and embodiments, when the executable instructions are executed by the processor system, the computer vision system is caused to repeatedly execute some of the above operations until the loss is less than a loss threshold.

[0019] In some implementations of the foregoing aspects and embodiments, when the executable instructions for determining an error of the camera point cloud with reference to the first LiDAR point cloud and determining a correction function based on the determined error are executed by the processor system, the computer vision system is caused to: generate a mathematical representation (e.g., a model) of the error of the camera point cloud with reference to the first LiDAR point cloud; and determine a correction function based on the mathematical representation of the error.

[0020] In some implementations of the foregoing aspects and embodiments, when the executable instructions for determining a training error of the corrected point cloud with reference to the first LiDAR point cloud and updating the correction function based on the determined training error are executed by the processor system, the computer vision system is caused to: generate a mathematical representation (e.g., a model) of the training error of the corrected point cloud with reference to the first LiDAR point cloud; and update the correction function based on the mathematical representation of the training error.

[0021] In some implementations of the foregoing aspects and embodiments, the computer vision system further includes a second LiDAR system coupled to the processor system and having a second resolution, wherein the second resolution of the second LiDAR system is higher than the first resolution of the first LiDAR system. When the executable instructions are executed by the processor system, the computer vision system is caused to: (a) receive a second LiDAR point cloud from the second LiDAR system; (b) determine an error of the corrected point cloud with reference to the second LiDAR point cloud; (c) determine a correction function based on the determined error; (d) generate a corrected point cloud using the correction function; (e) calculate an output of a point cloud application algorithm using the corrected point cloud and calculate the output of the point cloud application algorithm using the second LiDAR point cloud; (f) determine a second loss of the error of the output of the point cloud application algorithm using the corrected point cloud through a loss function; (g) update the correction function based on the determined second loss.

[0022] In some implementations of the foregoing aspects and embodiments, when the executable instructions are executed by the processor system, the computer vision system is caused to repeatedly execute operations (d) to (g) until the second loss is less than a loss threshold.

[0023] In some implementations of the foregoing aspects and embodiments, the second LiDAR system includes one or more 64-beam LiDAR units, and the first LiDAR system includes one or more 8-beam, 16-beam, or 32-beam LiDAR units.

[0024] In some implementations of the foregoing aspects and embodiments, the camera system is a stereo camera system, and the camera point cloud is a stereo camera point cloud.

[0025] In some implementations of the foregoing aspects and embodiments, the processor system includes a neural network.

[0026] According to another embodiment of the first aspect of the present invention, there is provided a computer vision system, including a processor system, a camera system coupled to the processor system, and a memory coupled to the processor system, wherein executable instructions are tangibly stored on the memory. When the executable instructions are executed by the processor system, the computer vision system is caused to perform several operations: generate a camera point cloud by the camera system; generate a corrected point cloud with a resolution higher than that of the camera point cloud by applying a pre-trained correction function; calculate an output of a point cloud application algorithm using the corrected point cloud. In some implementations of the foregoing aspects and embodiments, a visual representation of the output of the point cloud application algorithm is output to a display of the computer vision system.

[0027] In some implementations of the foregoing aspects and embodiments, the point cloud application algorithms include one or more of an object detection algorithm, a dynamic object removal algorithm, a simultaneous localization and mapping (SLAM) algorithm, a point cloud map generation algorithm, a high-definition map creation algorithm, a LiDAR-based positioning algorithm, a semantic mapping algorithm, a tracking algorithm, or a scene reconstruction algorithm.

[0028] In some implementations of the foregoing aspects and embodiments, the camera system is a stereo camera system, and the camera point cloud is a stereo camera point cloud.

[0029] In some implementations of the foregoing aspects and embodiments, the processor system includes a neural network.

[0030] According to a second aspect of the present invention, there is provided a method for training a preprocessing module to generate a corrected point cloud. According to an embodiment of the second aspect of the present invention, the method includes: receiving a camera point cloud from a camera system; receiving a first LiDAR point cloud from a first LiDAR system; determining an error of the camera point cloud with reference to the first LiDAR point cloud; determining a correction function based on the determined error; using the correction function to generate a corrected point cloud based on the camera point cloud; determining a training error of the corrected point cloud with reference to the first LiDAR point cloud; and updating the correction function based on the determined training error.

[0031] According to another embodiment of the second aspect of the present invention, the method includes: receiving a camera point cloud from a camera system; receiving a first LiDAR point cloud from a first LiDAR system; determining an error of the camera point cloud with reference to the first LiDAR point cloud; determining a correction function based on the determined error; using the correction function to generate a corrected point cloud based on the camera point cloud; calculating an output of a point cloud application algorithm using the corrected point cloud, and calculating the output of the point cloud application algorithm using the first LiDAR point cloud; determining a loss of an error of the output of the point cloud application algorithm using the corrected point cloud through a loss function; and updating the correction function based on the determined loss.

[0032] According to an embodiment of another aspect of the present invention, there is provided a method for generating a corrected point cloud. The method includes: generating a point cloud based on a sensor system; applying a pre-trained correction function to generate a corrected point cloud with a resolution higher than that of the point cloud; and calculating an output of a point cloud application algorithm using the corrected point cloud. In some implementations of the foregoing aspects and embodiments, the method further includes: outputting a visual representation of the output of the point cloud application algorithm to a display.

[0033] According to another aspect of the present invention, there is provided a vehicle control system for a vehicle. The vehicle control system includes a computer vision system having the characteristics described above and in this embodiment.

[0034] According to another aspect of the present invention, there is provided a vehicle including a mechanical system for moving the vehicle, a drive control system coupled to the mechanical system for controlling the mechanical system, and a vehicle control system coupled to the drive control system, wherein the vehicle control system has the characteristics described above and in this embodiment.

[0035] According to still another aspect of the present invention, there is provided a non-transitory computer-readable medium having executable instructions tangibly stored thereon, wherein the executable instructions are executed by a processor system of a computer vision system having the characteristics described above and in this embodiment. When the executable instructions are executed by the processor system, the computer vision system is caused to execute the methods described above and in this embodiment. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Schematic diagram of a communication system suitable for implementing an exemplary embodiment of the present invention;

[0037] Figure 2 Block diagram of a vehicle including a vehicle control system according to an exemplary embodiment of the present invention;

[0038] Figure 3A and Figure 3B Simplified block diagram of a computer vision system for training a preprocessing module to generate a corrected point cloud according to an exemplary embodiment of the present invention;

[0039] Figure 4A and Figure 4B Flowchart of a method for training a preprocessing module to generate a corrected point cloud according to an exemplary embodiment of the present invention;

[0040] Figure 5A and Figure 5B Simplified block diagram of a computer vision system for training a preprocessing module to generate a corrected point cloud according to other exemplary embodiments of the present invention;

[0041] Figure 6A and Figure 6B Flowchart of a method for training a preprocessing module to generate a corrected point cloud according to other exemplary embodiments of the present invention;

[0042] Figure 7 Flowchart of a method for generating a high-resolution and high-precision point cloud according to an exemplary embodiment of the present invention;

[0043] Figure 8 Schematic diagram of a neural network;

[0044] Figure 9 Schematic diagram of epipolar geometry in a triangulation;

[0045] Figure 10 Schematic diagram of epipolar geometry when two image planes are parallel. Detailed implementation manners

[0046] The present invention is described with reference to the accompanying drawings showing embodiments. However, many different embodiments can be used. Therefore, this description should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to make the present invention more detailed and complete. In the drawings and the following description, the same reference numerals are used to indicate the same elements as much as possible, and in alternative embodiments, angular minute symbols are used to represent similar elements, operations, or steps. The shown division of separate boxes or functional elements of the shown systems and devices does not necessarily require a physical division of such functions, because these elements can communicate with each other through message passing, function calls, shared memory spaces, etc., without any physical isolation between them. Therefore, although each function is shown separately herein for convenience, these functions do not need to be implemented on physically or logically separate platforms. Different devices can have different designs. Therefore, some devices implement certain functions on fixed functional hardware, while other devices can implement these functions in a programmable processor using code obtained from a computer-readable medium. Finally, unless otherwise clearly stated or logically indicated by the context, elements in the singular form can be plural, and vice versa.

[0047] For convenience, the present invention describes exemplary embodiments of methods and systems in connection with motor vehicles, such as cars, trucks, buses, small boats or large ships, submarines, airplanes, storage equipment, construction equipment, tractors, or other agricultural equipment. The content of the present invention is not limited to any specific type of vehicle and is applicable to non-manned vehicles and manned vehicles. The content of the present invention can also be implemented in mobile robotic vehicles, including but not limited to, autonomous vacuum cleaners, exploration vehicles, lawn mowers, unmanned aerial vehicles (UAVs), and other objects.

[0048] Figure 1 Schematic diagram showing selected components of a communication system 100 according to an exemplary embodiment of the present invention. The communication system 100 includes a user equipment in the form of a vehicle control system 115 built into a vehicle 105. Details are as Figure 2As shown, the vehicle control system 115 is coupled to the drive control system 150 and the mechanical system 190 of the vehicle 105, as described below. In various embodiments, the vehicle control system 115 can operate the vehicle 105 in one or more of fully autonomous, semi-autonomous, or fully user-controlled modes.

[0049] The vehicle 105 includes a plurality of electromagnetic (EM)-wave-based sensors 110 and a plurality of vehicle sensors 111, where the sensors 110 collect data related to the external environment around the vehicle 105, and the vehicle sensors 111 collect data related to the operating conditions of the vehicle 105. The EM-wave-based sensors 110 can include, for example, one or more cameras 112, one or more LiDAR units 114, and one or more radar units such as synthetic aperture radar (SAR) units 116. The digital cameras 112, LiDAR units 114, and SAR units 116 are located around the vehicle 105 and are respectively coupled to the vehicle control system 115, as described below. In an exemplary embodiment, the cameras 112, LiDAR units 114, and SAR units 116 are located in front of, behind, to the left, and to the right of the vehicle 105 to capture data related to the environment in front of, behind, to the left, and to the right of the vehicle 105. For each EM-wave-based sensor 110, the units are installed or placed to have different fields of view (FOV) or coverage ranges to capture data related to the environment around the vehicle 105. In some examples, for each EM-wave-based sensor 110, the FOV or coverage ranges of some or all adjacent EM-wave-based sensors 110 partially overlap. Accordingly, the vehicle control system 115 receives the data related to the external environment of the vehicle 105 collected by the cameras 112, LiDAR units 114, and SAR units 116.

[0050] The vehicle sensor 111 may include: an inertial measurement unit (IMU) 118 that senses the specific force and angular velocity of the vehicle using a combination of an accelerator and a gyroscope, an electronic compass 119, and other vehicle sensors 120, such as a speedometer, a tachometer, a wheel traction sensor, a transmission gear sensor, a throttle and brake position sensor, and a steering angle sensor. The vehicle sensor 111 repeatedly (e.g., at regular intervals) senses the environment during operation and provides sensor data to the vehicle control system 115 in real time or near real time based on environmental conditions. The vehicle control system 115 may collect data related to the position and orientation of the vehicle 105 using signals received from the satellite receiver 132 and the IMU 118. The vehicle control system 115 may determine the linear velocity, angular velocity, acceleration, engine speed, transmission gear, and tire grip of the vehicle 105 and other factors using data from one or more of the satellite receiver 132, the IMU 118, and other vehicle sensors 120.

[0051] The vehicle control system 115 may further include one or more wireless transceivers 130, enabling the vehicle control system 115 to exchange data with the wireless wide area network (WAN) 210 of the communication system 100 and optionally perform voice communication. The vehicle control system 115 may use the wireless WAN 210 to access a server 240, such as a driving assistance server, through one or more communication networks 220, such as the Internet. The server 240 may be implemented as one or more server modules in a data center and is typically located behind a firewall 230. The server 240 is connected to network resources 250, such as supplementary data resources available for use by the vehicle control system 115.

[0052] In addition to the wireless WAN 210, the communication system 100 further includes a satellite network 260, which includes a plurality of satellites. The vehicle control system 115 includes the satellite receiver 132 (such as Figure 2) Its position can be determined using signals received by the satellite receiver 132 from multiple satellites in the satellite network 260. The satellite network 260 generally includes multiple satellites that are part of at least one Global Navigation Satellite System (GNSS) providing autonomous geospatial positioning through global coverage. For example, the satellite network 260 can be a collection of GNSS satellites. Exemplary GNSSs include the United States NAVSTAR Global Positioning System (GPS) or the Russian Global Navigation Satellite System (GLONASS). Other satellite navigation systems that have been deployed or are under development include the European Union's Galileo positioning system, China's Beidou Navigation Satellite System (BDS), India's Regional Navigation Satellite System, and Japan's satellite navigation system.

[0053] Figure 2Shows selected components of vehicle 105 according to an exemplary embodiment of the present invention. As described above, the vehicle 105 includes a vehicle control system 115 connected to a drive control system 150, a mechanical system 190, an EM wave-based sensor 110, and vehicle sensors 111. The vehicle 105 also includes various structural components well known in the art, such as a frame, doors, panels, seats, windows, mirrors, etc., but these components are omitted in the present invention to make the content of the present invention clear. The vehicle control system 115 includes a processor system 102 coupled to a plurality of components via a communication bus (not shown), where the communication bus provides a communication path between the components and the processor system 102. The processor system 102 is coupled to a drive control system 150, a random access memory (RAM) 122, a read only memory (ROM) 124, a permanent (non-volatile) memory 126 such as an erasable programmable read only memory (EPROM) (flash memory), one or more wireless transceivers 130 for exchanging radio frequency signals with a wireless WAN 210, a satellite receiver 132 for receiving satellite signals from a satellite network 260, a real-time clock 134, and a touch screen 136. The processor system 102 may include one or more processing units, for example, including one or more central processing units (CPUs), one or more graphical processing units (GPUs), one or more tensor processing units (TPUs), and other processing units.

[0054] The one or more wireless transceivers 130 may include one or more cellular (RF) transceivers that communicate with multiple different wireless access networks (such as cellular networks) through different wireless data communication protocols and standards. The vehicle control system 115 can communicate with any one of a plurality of fixed base stations ( Figure 1 shows one of them) within the geographical coverage of the wireless WAN 210 (such as a cellular network). The one or more wireless transceivers 130 can send and receive signals through the wireless WAN 210. The one or more wireless transceivers 130 may include multi-band cellular transceivers that support multiple radio frequency bands.

[0055] The one or more wireless transceivers 130 may also include a wireless local area network (WLAN) transceiver that communicates with a WLAN (not shown) via a WLAN access point (AP). The WLAN may include a Wi-Fi wireless network compliant with the IEEE 802.11x standard (sometimes referred to as ) or other communication protocols.

[0056] The one or more wireless transceivers 130 may also include a short-range wireless transceiver that communicates with a mobile computing device such as a smartphone or a tablet computer. For example, a transceiver. The one or more wireless transceivers 130 may also include other short-range wireless transceivers, including but not limited to Near field communication (NFC), IEEE 802.15.3a (also known as Ultrawideband (UWB)), Z-Wave, ZigBee, ANT / ANT+, or infrared (e.g., Infrared Data Association (IrDA) communication).

[0057] The real-time clock 134 may include an oscillator that provides accurate real-time time data. The time data may be periodically adjusted according to the time data received by the satellite receiver 132 or according to the time data received from a network resource 250 that implements the Network Time Protocol.

[0058] The touch screen 136 includes a display such as a color liquid crystal display (LCD), a light-emitting diode (LED) display, or an active-matrix organic light-emitting diode (AMOLED) display, and a touch-sensitive input surface or overlay coupled to an electronic controller. Other input devices (not shown) coupled to the processor system 102 may also be provided, including buttons, switches, and dials.

[0059] The vehicle control system 115 also includes one or more speakers 138, one or more microphones 140, and one or more data ports 142, such as a serial data port (e.g., a Universal Serial Bus (USB) data port). The vehicle control system 115 may also include other sensors, such as a tire pressure sensor (TPS), a door contact switch, a light sensor, a proximity sensor, etc.

[0060] The drive control system 150 is used to control the movement of the vehicle 105. The drive control system 150 includes a steering unit 152, a braking unit 154, and a throttle (or acceleration) unit 156. These units can all be implemented as software modules or control blocks in the drive control system 150. In a fully autonomous or semi-autonomous driving mode, the steering unit 152, the braking unit 154, and the throttle unit 156 process navigation instructions received from an autonomous driving system 170 (for the autonomous driving mode) or a driving assistance system 166 (for the semi-autonomous driving mode), and generate control signals to control at least one of the steering, braking, and throttling of the vehicle 105. The drive control system 150 may include other devices to control other aspects of the vehicle 105, including controlling turn signals and brake lights, etc.

[0061] The electromechanical system 190 receives control signals from the drive control system 150 to operate the electromechanical devices of the vehicle 105. The electromechanical system 190 causes the physical operation of the vehicle 105. The electromechanical system 190 includes an engine 192, a transmission 194, and wheels 196. The engine 192 can be a gasoline engine, a battery engine, or a hybrid engine, etc. The mechanical system 190 may include other devices, including turn signals, brake lights, fans, and windows, etc.

[0062] The processor system 102 renders the graphical user interface (GUI) of the vehicle control system 115 and displays it on the touch screen 136. The user can interact with the GUI through the touch screen 136 and optionally through other input devices (such as buttons, dials) to select the driving mode of the vehicle 105 (for example, fully autonomous driving mode or semi-autonomous driving mode) and display relevant data and / or information, such as navigation information, driving information, parking information, media player information, climate control information, etc. The GUI may include a series of traversal menus for specific content.

[0063] In addition to the GUI, a plurality of software systems 161 are also stored on the memory 126 of the vehicle control system 115, wherein each software system 161 includes instructions executable by the processor system 102. The software systems 161 include an operating system 160, the driving assistance system 166 for semi-autonomous driving, and the autonomous driving system 170 for fully autonomous driving. The driving assistance system 166 and the autonomous driving system 170 may each include one or more of a navigation planning and control module, a vehicle positioning module, a parking assistance module, and an autonomous parking module. A software module 168 that can be called by the driving assistance system 166 or the autonomous driving system 170 is also stored on the memory 126. The software module 168 includes a computer vision module 172. The computer vision module 172 is a software system that includes a learning-based preprocessing module 330 or 530, a point cloud processing module 340 or 540, and optionally includes a loss determination module 360 or 560. Other modules 176 include a mapping module, a navigation module, a climate control module, a media player module, a phone module, and a messaging module, etc. When the computer vision module 172 is executed by the processor system 102, the operations of the methods described herein are performed. The computer vision module 172 interacts in combination with the EM wave-based sensor 110 to provide a computer vision system, for example, the computer vision systems 300, 350, 500, or 550 detailed below( Figure 3A , Figure 3B , Figure 5A , Figure 5B ).

[0064] Although the computer vision module 172 is shown as a separate module that can be called by the driving assistance system 166 for semi-autonomous driving and / or the autonomous driving system 170, in some embodiments, one or more of the software modules 168, including the computer vision module 172, may be combined with one or more of the other modules 176.

[0065] The memory 126 also stores different data 180. The data 180 may include sensor data 182, user data 184, and a download cache 186. The sensor data 182 is received from the EM wave-based sensor 110; the user data 184 includes user preferences, settings, and optionally includes personal media files (such as music, videos, directions, etc.); the download cache 186 includes data downloaded through the wireless transceiver 130, for example, data downloaded from the network resource 250. The sensor data 182 may include image data from the camera 112, 3D data from the LiDAR unit 114, radar data from the SAR unit 116, IMU data from the IMU 118, compass data from the electronic compass 119, and other sensor data from other vehicle sensors 120. The download cache 186 may be deleted periodically, for example, after a predetermined time. System software, software modules, specific device applications, or portions thereof may be temporarily loaded into a volatile memory such as the RAM 122, where the volatile memory is used to store runtime data variables and other types of data and / or information. The data received by the vehicle control system 115 may also be stored in the RAM 122. Although specific functions of different types of memories are described, this is merely an example, and different functions may be assigned to different types of memories.

[0066] Generate a high-resolution and high-precision point cloud

[0067] Next, refer to Figure 3A 。 Figure 3A FIG. 10 is a simplified block diagram of a computer vision system 300 for training a preprocessing module to generate a corrected point cloud according to an exemplary embodiment of the present invention. The computer vision system 300 may be mounted on a test vehicle such as the vehicle 105 and may operate in an online mode or an offline mode. In the online mode, the EM wave-based sensor 110 actively receives sensor data; in the offline mode, the EM wave-based sensor 110 has previously received sensor data. Alternatively, the computer vision system 300 may be a separate system independent of the test vehicle, receiving and processing sensor data obtained from the test vehicle, and may operate in an offline mode. The computer vision system 300 receives sensor data from a high-resolution and high-precision sensor system such as a LiDAR system 310 including one or more LiDAR units 114, and receives sensor data from a low-precision sensor system such as a camera system 320 including one or more cameras 120.

[0068] The computer vision system 300 further includes a learning-based preprocessing module 330 (hereinafter referred to as the preprocessing module 330) and a point cloud processing module 340. The computer vision system 300 is used to train the preprocessing module 330 in a training mode. The training is intended to train the preprocessing module 330 to learn to generate a high-resolution and high-precision point cloud in real time or near real time using only sensor data received from a low-precision sensor system (e.g., the camera system 320) or a possible low-resolution system (e.g., one or more LiDAR units 114 including low-precision LiDAR sensors). The trained preprocessing module 330 can be used in a production vehicle including a low-precision sensor system such as a camera system and / or a low-resolution LiDAR subsystem 510 (as Figure 5A , Figure 5B shown), for example, on the vehicle 105, as described below. For example, the preprocessing module 330 can be set as a software module installed in the vehicle control system 115 of the production vehicle during production, before shipping to the customer, or at any suitable time thereafter.

[0069] The low-precision sensor system is generally a high-resolution sensor system. The high-resolution and high-precision sensor system provides a reference system. The reference system provides ground truth data for training the preprocessing module 330 to generate the corrected point cloud data from the low-precision sensor system, as described below.

[0070] The LiDAR system 310 includes one or more high-resolution LiDAR units, e.g., a LiDAR unit having 64 or more beams, and provides a high-resolution LiDAR point cloud using techniques well known in the art. In some examples, the LiDAR system 310 may include a number of 64-beam LiDAR units, e.g., 5 or more LiDAR units. When the desired resolution of the LiDAR point cloud to be generated by the LiDAR system 310 is higher than the resolution of the LiDAR point cloud originally generated by the available LiDAR units (e.g., the output of a single 64-beam LiDAR unit), the LiDAR system 310 may include multiple LiDAR units. Then, the LiDAR point clouds generated by each LiDAR unit in the multiple LiDAR units are merged to generate a super-resolution LiDAR point cloud. Those skilled in the art can understand that "super-resolution" refers to an enhanced resolution higher than the original resolution supported by the sensor system. One of several different merging algorithms, e.g., the K-Nearest Neighbors (KNN) algorithm, can be used to merge the LiDAR point clouds to generate the super-resolution LiDAR point cloud.

[0071] The camera system 320 includes one or more cameras, typically high-resolution cameras. In this embodiment, the one or more cameras are stereo cameras that provide a low-precision stereo camera point cloud, which is typically a high-resolution low-precision stereo camera point cloud. In other embodiments, a monocular camera may be used instead of a stereo camera.

[0072] The LiDAR system 310 and the camera system 320 are coupled to the preprocessing module 330 and provide a high-resolution LiDAR point cloud and a low-precision stereo camera point cloud. The preprocessing module 330 uses the high-resolution high-precision LiDAR point cloud as a reference, and the high-resolution high-precision LiDAR point cloud provides ground truth data for training the preprocessing module 330 to generate a corrected point cloud based on the low-precision stereo camera point cloud. The preprocessing module 330 outputs the corrected point cloud to the point cloud processing module 340, and the point cloud processing module 340 uses the corrected point cloud in one or more machine learning algorithms, such as object detection algorithms, dynamic object removal algorithms, simultaneous localization and mapping (SLAM) algorithms, point cloud map generation algorithms, high-definition map creation algorithms, localization algorithms, semantic mapping algorithms, tracking algorithms, or scene reconstruction algorithms, which can be used to generate a 3D representation of the environment. As described above, the preprocessing module 330 and the point cloud processing module 340 are software modules in the computer vision module 172.

[0073] The preprocessing module 330 is trained as described below. The training of the preprocessing module 330 can be performed simultaneously with other calibration procedures of the computer vision system 300. In at least some embodiments, the preprocessing module 330 may be or include a neural network. In other embodiments, the preprocessing module 330 may be or include other types of machine learning-based controllers, processors, or systems for implementing machine learning algorithms to train the preprocessing module to generate a corrected point cloud, as described below. A brief reference will be made below Figure 8 to describe an exemplary neural network 800. The neural network 800 includes a number of nodes (also referred to as neurons) 802, and the nodes 802 are arranged in multiple layers, including an input layer 810, one or more intermediate (hidden) layers 820 (only one is shown for simplicity), and an output layer 830. Each of the layers 810, 820, 830 is a group of one or more nodes that are independent of each other and support parallel computing.

[0074] The output of each node 802 in a given layer is connected to the output of one or more nodes 802 in the next layer, as shown by the connection 804 (in Figure 8Only one connection 804 is marked. Each node 802 is a logical programming unit that executes an activation function (also known as a transfer function) for transforming or operating on data based on its inputs, weights (if any), and bias factors (if any) to generate an output. Depending on the specific inputs, weights, and biases, the activation function of each node 802 will produce a specific output. The input of each node 802 can be a scalar, vector, matrix, object, data structure, and / or other item, or a reference thereto. Each node 802 can store its respective activation function, weights (if any), and bias factors (if any) independently of other nodes.

[0075] Examples of activation functions include mathematical functions (i.e., addition, subtraction, multiplication, division, convolution, etc.), object operation functions (i.e., creating an object, modifying an object, deleting an object, adding an object, etc.), data structure operation functions (i.e., creating a data structure, modifying a data structure, deleting a data structure, creating a data field, modifying a data field, deleting a data field, etc.), and / or other transformation functions, depending on the type of input. In some examples, the activation function includes at least one of a summation function or a mapping function.

[0076] Each node in the input layer 810 receives sensor data from the LiDAR system 310 and the camera system 320. Weights can be set for each input in one or more of the subsequent nodes in the intermediate layer 820 and the output layer 830 of the neural network 800, as well as for each input in the input layer 810. The weights are typically numerical values between 0 and 1, indicating the strength of the connection between a node in one layer and a node in the next layer. Biases can also be set for each input in the intermediate layer 820 and the output layer 830 of the neural network 800, as well as for each input in the input layer 810.

[0077] Determine the scalar product of each input of the input layer 810 with its respective weights and biases, and send it as an input to the nodes of the first intermediate layer 820. If there is more than one intermediate layer, link each of the scalar products to another vector, determine another scalar product of the input of the first intermediate layer with its respective weights and biases, and send it as an input to the nodes in the second intermediate layer 820. Repeat this process for each of the intermediate layers 820 in sequence until the output layer 830.

[0078] In different embodiments, the number of the intermediate layers 820, the number of nodes in each of the layers 810, 820, and 830, and the connections between the nodes of each layer may vary according to the different inputs (e.g., sensor data) and outputs (e.g., the corrected point cloud) provided to the processing module 340. The weights and biases of each node and even the activation function of the nodes of the neural network 800 may be determined through training (e.g., a learning process) to achieve the best performance, as described below.

[0079] A method 400 for training the preprocessing module 330 to generate a corrected point cloud according to an exemplary embodiment of the present invention will be described below with reference to Figure 4A The method 400 is at least partially implemented by software executed by the processor system 102 of the vehicle control system 115. The method 400 is executed when an operator activates (or adopts) the point cloud map learning mode of the computer vision system 300 of the vehicle 105. The point cloud map learning mode can be activated through interaction with a human-machine interface device (HMD) of the vehicle control system 115, such as voice activation by a predefined keyword combination or other user interactions, such as touch activation of the GUI of the computer vision system 300 displayed on the touch screen 136.

[0080] In operation 404, the computer vision system 300 receives a first point cloud obtained by a first sensor system, e.g., the camera system 320. The first point cloud is also referred to as the camera point cloud. The first point cloud is generally generated by the camera system 320 and received as an input by the processor system 102. Alternatively, the computer vision system 300 may generate the first point cloud based on digital images acquired from the camera system 320.

[0081] When the first point cloud is a camera-based point cloud, each point is defined by its spatial position represented by x, y, z coordinates and its color feature, where the color feature may be described by red green blue (RGB) values or other suitable data values. As described above, the camera system 320 is a low-precision system such as a high-resolution low-precision camera system, and thus, the first point cloud is a low-precision point cloud. The computer vision system 300 uses the camera system 320 to sense the environment of the vehicle 105 to generate the first point cloud using computer vision techniques well known in the art. Details of the computer vision techniques well known in the art are beyond the scope of the present invention. An example of a computer vision technique for generating a stereo vision point cloud will be briefly described below.

[0082] Stereo vision point clouds can be generated using epipolar geometry through triangulation (also known as reconstruction) techniques, based on images captured by a stereo camera or two monocular cameras. A stereo camera observes a 3D scene from two different positions, thereby creating several geometric relationships between 3D points and the projections of the 3D points on the 2D images captured by the stereo camera, resulting in constraints between the image points. These relationships are obtained by approximating each lens of the stereo camera (or each camera when using two monocular cameras) using the pinhole camera model. Figure 9 Schematic diagram of the intersection point of environmental characteristics observed by different camera lenses each having a different field of view.

[0083] As Figure 9 shown, in the standard triangulation method using epipolar geometry, two cameras observe the same 3D point P, i.e., the intersection point, where the projections of the intersection point in the respective image planes are located at p and p’. The camera centers are located at O1 and O2, and the line between the camera centers is called the baseline. The lines between the camera centers O1 and O2 and the intersection point P are projection lines. The plane defined by the two camera centers and P is the epipolar plane. The positions where the baseline intersects the two image planes are called the poles e and e'. The lines defined by the intersection of the epipolar plane and the two image planes are called epipolar lines. The epipolar lines intersect the baseline at the poles of the respective image planes. Based on the corresponding image points p and p’ and the geometry of the two cameras, the projection lines can be determined, and the projection lines intersect at the 3D intersection point P. The 3D intersection point P can be determined directly by linear algebra using known techniques.

[0084] Figure 10 Shows the epipolar geometry when the image planes are parallel to each other. When the image planes are parallel to each other, since the baseline connecting the centers O1 and O2 is parallel to the image planes and the epipolar lines are parallel to the axes of each image plane, the poles e and e' are located at infinity.

[0085] Performing triangulation requires the parameters of all 3D-to-2D camera projection functions for each camera involved. This can be represented using a camera matrix. The camera matrix or (camera) projection matrix is a 3×4 matrix that describes the mapping between 3D points in the real world and 2D points in the image for a pinhole camera. If x represents a 3D point (a 4D vector) in homogeneous coordinates and y represents the image of the point in the pinhole camera (a 3D vector), then the following relationship exists:

[0086] y = Cx (1)

[0087] where C is the camera matrix, and C is defined by the following equation:

[0088]

[0089] Among them, f is the focal length of the camera, and f > 0.

[0090] Since each point in the 2D image corresponds to a line in the 3D space, all points on the 3D line are projected onto the point in the 2D image. If a pair of corresponding points can be found in two or more images, this pair of corresponding points must be the projections of the same 3D point P, that is, the intersection point. The set of lines generated by these image points must intersect at P (3D point); and algebraic formulas for calculating the coordinates of P (3D point) can be obtained using various methods well-known in the art, such as the midpoint method and direct linear transformation.

[0091] In operation 406, the computer vision system 300 generates a second point cloud through a second sensor system, for example, the LiDAR system 310. The second point cloud is a high-resolution and high-precision point cloud. The second point cloud is also referred to as the LiDAR point cloud. The second point cloud is typically generated by the LiDAR system 310 and received as input by the processor system 102. The second point cloud can be generated simultaneously with the first point cloud. Alternatively, when the LiDAR system 310 generates the second point cloud in the same environment (which can be a test room or a reference room) as when the camera system 320 generates the first point cloud and based on the same reference position (such as the same position and orientation) in the environment, the second point cloud can be generated in advance and provided to the computer vision system 300.

[0092] In operation 408, the preprocessing module 330 receives the first point cloud and the second point cloud, and determines the error of the first point cloud with the second point cloud as the ground truth. At least in some embodiments, the preprocessing module 330 generates a mathematical representation (such as a model) of the error of the first point cloud. The mathematical representation of the error is based on the difference between the matching coordinates of the first point cloud and the second point cloud. The mathematical representation of the error can define a general error applicable to all data points in the first point cloud, or define a variable error that varies uniformly or non-uniformly throughout the first point cloud. As a preparatory step for generating the mathematical representation (such as a model) of the error, the data points of the LiDAR point cloud are matched (such as associated, mapped, etc.) with the data points of the stereo camera point cloud. For example, in some embodiments, the data points of the LiDAR point cloud can be matched with the data points of the stereo camera point cloud using the KNN algorithm described above. Two point clouds are provided to the preprocessing module 330: both point clouds have high resolution, but the accuracy of one is lower than that of the other. The KNN algorithm can be used to find the best KNN match between each point in the low-accuracy point cloud and its K nearest neighbors in the high-accuracy point cloud. In other embodiments, other matching, mapping, or interpolation algorithms / learning-based algorithms can be used.

[0093] In operation 410, the preprocessing module 330 determines a correction function with the second point cloud as the ground truth, based on the determined error, e.g., based on a mathematical representation of the error of the first point cloud. The correction function can be a correction vector applied to all data points in the first point cloud, several different correction vectors applied to different data points in the first point cloud, or a variable correction function that varies non-uniformly throughout the first point cloud, among many other possibilities.

[0094] In operation 412, the preprocessing module 330 generates a corrected point cloud based on the first point cloud using the correction function.

[0095] In operation 414, the preprocessing module 330 generates a mathematical representation of the training error between the corrected point cloud and the second point cloud with the second point cloud as the ground truth. The mathematical representation of the training error is based on the difference between the matching coordinates of the corrected point cloud and the second point cloud.

[0096] In operation 416, the preprocessing module 330 determines whether the training error is less than an error threshold. When the training error is less than the error threshold, it is considered that the correction function has been trained and method 400 ends. When the training error is greater than or equal to the error threshold, operation 418 is continued to recalculate the correction function. When the preprocessing module 330 is a neural network for training the correction function, the correction function can be recalculated by backpropagating the training error in the neural network by updating parameters such as the weights of the neural network to minimize the training error.

[0097] Next, refer to Figure 3B . Figure 3BFIG. 0 is a simplified block diagram of a computer vision system 350 for training a preprocessing module 330 to generate a corrected point cloud according to another exemplary embodiment of the present invention. The computer vision system 350 differs from the computer vision system 300 in that, in the computer vision system 350, combined training is performed based on one or more machine learning algorithms for processing point cloud data (e.g., one or more of an object detection algorithm, a dynamic object removal algorithm, a simultaneous localization and mapping (SLAM) algorithm, a point cloud map generation algorithm, a high-definition map creation algorithm, a LiDAR-based localization algorithm, a semantic mapping algorithm, a tracking algorithm, or a scene reconstruction algorithm). The output of the processing module 340 can be used to generate a 3D representation of the environment, which can be displayed on the touch screen 136. Since, according to a specific point cloud application algorithm, certain errors in the high-resolution / high-precision point cloud generated by the computer vision system are more or less more tolerant of errors than other errors, the combined training can further improve the final result (e.g., the 3D representation of the environment) of a specific point cloud application using the corrected point cloud, thereby improving the performance of the computer vision system for the intended application.

[0098] The computer vision system 350 includes the LiDAR system 310, the camera system 320, the preprocessing module 330, and the point cloud processing module 340 of the computer vision system 300, but also includes a loss determination module 360. The loss determination module 360 can be a hardware module or a software module 168, e.g., a software module of the computer vision module 172. The combined training is performed in an end-to-end manner, where the loss determination module 360 uses the data points in the corrected point cloud and the data points in the second point cloud to determine the "loss" associated with the training error based on the calculation of a specific point cloud application algorithm. The loss is determined by a loss function (or loss functions) for training the preprocessing module 330. The loss function is an objective function that maps the value of the output of a specific point cloud application algorithm to a real number representing the "loss" associated with the value. The loss function can be a point cloud application algorithm loss function or can be a combination of a point cloud application algorithm loss function and a preprocessing module loss function. The loss function can output a loss (or losses) based on a weighted combination of the losses of the preprocessing module and the point cloud application algorithm. The loss function is defined by an artificial intelligence / neural network designer, and the details are beyond the scope of the present invention. When it is desired to improve the performance through the specific point cloud application algorithm, the combined training can help improve the overall performance.

[0099] The following will refer to Figure 4BDescribe a method 420 for training the preprocessing module 330 to generate a corrected point cloud according to another exemplary embodiment of the present invention. At least part of the method 420 is implemented by software executed by the processor system 102 of the vehicle control system 115. The method 420 is similar to the method 400, except that combined training is performed in the method 420 as described above.

[0100] After generating the corrected point cloud in operation 412, in operation 422, the preprocessing module 330 uses the corrected point cloud and the second point cloud to calculate the output of a specific point cloud application algorithm.

[0101] In operation 424, the preprocessing module 330 determines (e.g., calculates) a loss as the output of the loss function based on the output of the specific point cloud application algorithm calculated using the corrected point cloud and the output of the specific point cloud application algorithm calculated with the second point cloud as the ground truth. The calculated "loss" represents the loss of the error of the output of the point cloud application algorithm calculated using the corrected point cloud instead of the second point cloud.

[0102] In operation 426, the preprocessing module 330 determines whether the loss is less than a loss threshold. When the loss is less than the loss threshold, it is considered that the loss has been minimized and the correction function has been trained, and the method 420 ends. When the loss is greater than or equal to the loss threshold, operation 428 is continued to recalculate the correction function. When the preprocessing module 330 is a neural network for training the correction function, the training error can be backpropagated in the neural network by updating parameters such as the weights of the neural network to recalculate the correction function to minimize the loss of the loss function. In some examples, the parameters such as the weights of the neural network can be updated to minimize the mean square error (MSE) between the output of the specific point cloud application algorithm calculated using the corrected point cloud and the output of the specific point cloud application algorithm calculated using the second point cloud.

[0103] Although the loss function that minimizes the loss to solve the optimization problem is used as the objective function in the above description, in other embodiments, the objective function can be a reward function, a profit function, a utility function, or an adaptation function, etc., in which the output of the objective function is maximized instead of minimized to solve the optimization problem.

[0104] Next, refer to Figure 5A . Figure 5AFIG. 5 is a simplified block diagram of a computer vision system 500 for training a learning-based preprocessing module 530 (hereinafter referred to as preprocessing module 530) to generate a corrected point cloud according to another exemplary embodiment of the present invention. The computer vision system 500 is similar to the computer vision system 300 described above, except that the computer vision system 500 further includes a low-resolution LiDAR system 510 having one or more LiDAR units 114.

[0105] The computer vision system 500 is used to train the preprocessing module 530. The training is to train the preprocessing module 530 to generate a high-resolution and high-precision point cloud in real time or near real time using the low-resolution LiDAR system 510 and the camera system 320. In some cases, the high-resolution and high-precision LiDAR system 310 may include one or more 64 (or more) beam LiDAR units, while the low-resolution LiDAR system 510 includes one or more 8-beam, 16-beam or 32-beam LiDAR units. Compared with the high-resolution and high-precision LiDAR system 310, the camera system 320 and the low-resolution LiDAR system 510 are less expensive. The trained preprocessing module 530 can be used in production vehicles including low-precision sensor systems such as camera systems and / or low-resolution systems such as the low-resolution LiDAR subsystem 510, for example, vehicle 105.

[0106] The preprocessing module 530 is trained in the stages described below. The training of the preprocessing module 530 can be carried out simultaneously with other calibration procedures of the computer vision system 500. At least in some embodiments, the preprocessing module 530 can be or include a neural network. In other embodiments, the preprocessing module 530 can be or include other types of machine learning-based controllers, processors or systems. First, the preprocessing module 530 is trained using low-resolution / high-precision data from the low-resolution LiDAR system 510, and then the preprocessing module 530 is trained using high-resolution / high-precision data from the high-resolution LiDAR system 310 to fine-tune the correction function of the preprocessing module 530.

[0107] Reference will be made below to Figure 6A Describe a method 600 for training a preprocessing module 530 to generate a corrected point cloud according to another exemplary embodiment of the present invention. At least part of the method 600 is implemented by software executed by the processor system 102 of the vehicle control system 115. The method 600 is similar to the method 400, except that combined training is performed in the method 600 as described above.

[0108] In operation 602, the computer vision system 500 generates a first point cloud using a first sensor system, e.g., the camera system 320.

[0109] In operation 604, the computer vision system 500 generates a second point cloud through a second sensor system, e.g., the LiDAR system 310. The second point cloud is a high-resolution and high-precision point cloud. The second point cloud can be generated simultaneously with the first point cloud. Alternatively, when the LiDAR system 310 generates the second point cloud in the same environment (which can be a test chamber or a reference chamber) as when the camera system 320 generates the first point cloud and based on the same reference position (e.g., the same position and orientation) in the environment, the second point cloud can be generated in advance and provided to the computer vision system 500.

[0110] In operation 606, the computer vision system 500 generates a third point cloud through a third sensor system, e.g., the LiDAR system 510. The third point cloud is a low-resolution and high-precision point cloud. The third point cloud is also referred to as the LiDAR point cloud. The third point cloud is typically generated by the LiDAR system 510 and received as an input by the processor system 102.

[0111] The third point cloud can be generated simultaneously with the first point cloud and the second point cloud. Alternatively, when the LiDAR system 510 generates the third point cloud in the same environment (which can be a test chamber or a reference chamber) as when the first point cloud and the second point cloud are generated and based on the same reference position (e.g., the same position and orientation) in the environment, the third point cloud can be generated in advance and provided to the computer vision system 500.

[0112] In operation 608, the preprocessing module 530 determines the error of the first point cloud using the third point cloud of the low-resolution LiDAR system 510 as the ground truth. In at least some embodiments, the preprocessing module 530 generates a mathematical representation (e.g., a model) of the error of the first point cloud. Similar to operation 408, the mathematical representation of the error is based on the difference between the matching coordinates of the first point cloud and the third point cloud.

[0113] In operation 610, similar to operation 410, the preprocessing module 330 determines a correction function using the third point cloud as the ground truth and based on the determined error, e.g., based on the mathematical representation of the error that defines the error of the first point cloud.

[0114] In operation 612, the preprocessing module 530 uses the correction function to generate a corrected point cloud based on the first point cloud.

[0115] In operation 614, the preprocessing module 530 generates a mathematical representation (e.g., a model) of the training error between the corrected point cloud and the third point cloud, with the third point cloud as the ground truth. The mathematical representation of the training error is based on the difference between the matching coordinates of the corrected point cloud and the third point cloud.

[0116] In operation 616, the preprocessing module 530 determines whether the training error is less than an error threshold. When the training error is less than the error threshold, it is considered that the first stage of training the correction function is complete, and the method 600 proceeds to operation 620 to enter the second stage of training. When the training error is greater than or equal to the error threshold, operation 618 is continued to recalculate the correction function. When the preprocessing module 530 is a neural network for training the correction function, the training error can be backpropagated in the neural network by updating parameters such as the weights of the neural network to recalculate the correction function to minimize the training error.

[0117] In operation 620, it is considered that the first stage of training the correction function is complete. The preprocessing module 530 generates a mathematical representation (e.g., a model) of the error of the corrected point cloud, with the second point cloud of the high-resolution high-precision LiDAR system 310 as the ground truth. Similar to operations 408 and 608, the mathematical representation of the error is based on the difference between the matching coordinates of the first point cloud and the second point cloud.

[0118] In operation 622, similar to operations 410 and 610, the preprocessing module 330 determines a correction function based on the mathematical representation of the error that defines the error of the first point cloud, with the second point cloud as the ground truth.

[0119] In operation 624, the preprocessing module 530 uses the correction function to generate a corrected point cloud based on the first point cloud.

[0120] In operation 626, the preprocessing module 530 generates a mathematical representation (e.g., a model) of the training error between the corrected point cloud and the second point cloud, with the second point cloud as the ground truth. The mathematical representation of the training error is based on the difference between the matching coordinates of the corrected point cloud and the second point cloud.

[0121] In operation 628, the preprocessing module 530 determines whether the training error is less than an error threshold. When the training error is less than the error threshold, the training of the correction function is considered complete, and method 600 ends. When the training error is greater than or equal to the error threshold, operation 630 is continued to recalculate the correction function. When the preprocessing module 530 is a neural network for training the correction function, the training error can be backpropagated in the neural network by updating parameters such as the weights of the neural network to recalculate the correction function to minimize the training error.

[0122] Next, refer to Figure 5B . Figure 5B FIG. is a simplified block diagram of a computer vision system 550 for training a preprocessing module 530 to generate a corrected point cloud according to another exemplary embodiment of the present invention. The computer vision system 550 is different from the computer vision system 500 in that in the computer vision system 550, combined training is performed based on one or more point cloud application algorithms (for example, one or more of an object detection algorithm, a dynamic object removal algorithm, a simultaneous localization and mapping (SLAM) algorithm, a point cloud map generation algorithm, a high-definition map creation algorithm, a LiDAR-based positioning algorithm, a semantic mapping algorithm, a tracking algorithm, or a scene reconstruction algorithm).

[0123] The computer vision system 550 includes the high-resolution and high-precision LiDAR system 310, the low-resolution LiDAR system 510, the camera system 320, the preprocessing module 530, and the point cloud processing module 340 of the computer vision system 500, but also includes a loss (or losses) determination module 560 similar to the loss determination module 360.

[0124] Next, refer to Figure 6B to describe a method 650 for training a preprocessing module 530 to generate a corrected point cloud according to another exemplary embodiment of the present invention. At least part of the method 650 is implemented by software executed by the processor system 102 of the vehicle control system 115. The method 650 is similar to the method 600, except that combined training is performed in the method 600.

[0125] After generating the corrected point cloud in operation 612, in operation 652, the preprocessing module 530 uses the corrected point cloud and the third point cloud to calculate the output of a specific point cloud application algorithm.

[0126] In operation 654, the preprocessing module 530 determines (e.g., calculates) the loss of the output of the loss function based on the output of the specific point cloud application algorithm that employs the corrected point cloud computing and the output of the specific point cloud application algorithm calculated with the third point cloud as the ground truth. The calculated "loss" represents the loss of the error of the output of the point cloud application algorithm that employs the corrected point cloud instead of the third point cloud.

[0127] In operation 656, the preprocessing module 530 determines whether the loss is less than a loss threshold. When the loss is less than the loss threshold, it is considered that the loss has been minimized and the first stage of training the correction function has been completed, and the method 650 proceeds to operation 660. When the loss is greater than or equal to the loss threshold, operation 658 is continued to recalculate the correction function. When the preprocessing module 530 is a neural network for training the correction function, the correction function can be recalculated by backpropagating the training error in the neural network by updating parameters such as the weights of the neural network to minimize the loss of the loss function. In some examples, the parameters such as the weights of the neural network can be updated to minimize the mean square error (MSE) between the output of the specific point cloud application algorithm that employs the corrected point cloud and the output of the specific point cloud application algorithm that employs the second point cloud.

[0128] In operation 660, it is considered that the first stage of training the correction function has been completed, and the preprocessing module 530 generates a mathematical representation (e.g., model) of the error of the corrected point cloud with the second point cloud of the high-resolution high-precision LiDAR system 310 as the ground truth. Similar to operations 408 and 608, the mathematical representation of the error is based on the difference between the matching coordinates of the first point cloud and the second point cloud.

[0129] In operation 662, similar to operations 410 and 610, the preprocessing module 530 determines the correction function based on the mathematical representation of the error that defines the error of the first point cloud with the second point cloud as the ground truth.

[0130] In operation 664, the preprocessing module 530 generates a corrected point cloud based on the first point cloud using the correction function.

[0131] In operation 666, the preprocessing module 530 calculates the output of the specific point cloud application algorithm using the corrected point cloud and the second point cloud.

[0132] In operation 668, the preprocessing module 530 determines (e.g., calculates) the loss of the output of the loss function based on the output of the specific point cloud application algorithm using the corrected point cloud computing and the output of the specific point cloud application algorithm calculated with the second point cloud as the ground truth. The calculated "loss" represents the loss of the error of the output of the point cloud application algorithm using the corrected point cloud instead of the second point cloud.

[0133] In operation 670, the preprocessing module 530 determines whether the loss is less than a loss threshold. When the loss is less than the loss threshold, it is considered that the loss has been minimized and the correction function has been trained, and method 650 ends. When the loss is greater than or equal to the loss threshold, operation 672 is continued to recalculate the correction function. When the preprocessing module 330 is a neural network for training the correction function, the correction function can be recalculated by backpropagating the training error in the neural network by updating parameters such as the weights of the neural network to minimize the loss of the loss function. In some examples, the parameters such as the weights of the neural network can be updated to minimize the mean square error (MSE) between the output of the specific point cloud application algorithm using the corrected point cloud and the output of the specific point cloud application algorithm using the second point cloud.

[0134] Although the loss function that minimizes the loss to solve the optimization problem is used as the objective function in the above description, in other embodiments, the objective function can be a reward function, a profit function, a utility function, or an adaptation function, etc., in which the output of the objective function is maximized instead of minimized to solve the optimization problem.

[0135] The following will refer to Figure 7 Describe a method 700 for generating a high-resolution and high-precision point cloud using the trained preprocessing module 330 or 530 according to an exemplary embodiment of the present invention. At least part of the method 700 is implemented by software executed by the processor system 102 of the vehicle control system 115.

[0136] In operation 702, the processor system 102 generates a first point cloud using a first sensor system, e.g., the camera system 320.

[0137] In operation 704, the processor system 102 generates a second point cloud using a second sensor system, e.g., the low-resolution LiDAR system 510. This step is optional. In other embodiments, the computer vision system only generates the first point cloud.

[0138] In operation 706, the processor system 102 generates a first corrected point cloud or optionally a second corrected point cloud based on the respective pre-trained calibration functions of the camera system 320 or the low-resolution LiDAR system 510 described above.

[0139] In operation 708, the processor system 102 computes the output of one or more specific point cloud application algorithms based on the corrected point cloud, e.g., one or more of an object detection algorithm, a dynamic object removal algorithm, a SLAM algorithm, a point cloud map generation algorithm, a high-definition map creation algorithm, a LiDAR-based localization algorithm, a semantic mapping algorithm, a tracking algorithm, or a scene reconstruction algorithm. These algorithms are typically used for 3D but may also be used for 2D.

[0140] In operation 710, optionally, the processor system 102 generates a representation of the environment based on the result / output of the one or more point cloud application algorithms, the representation typically being 3D, such as a 3D map.

[0141] In operation 712, the processor system 102 outputs the 3D representation. The output may include displaying the 3D representation on a display such as the touchscreen 136, outputting the 3D representation to a vehicle driver assistance system or an autonomous driving system, or a combination thereof. The vehicle driver assistance system or the autonomous driving system is typically part of the vehicle control system 115 and may be implemented in software as described above.

[0142] Using the trained preprocessing module 330 or 530, the high-resolution high-precision LiDAR system 310 can be removed from the test vehicle or omitted in the production vehicle, thereby helping to reduce the cost of the final product. Cheaper sensors such as the camera system 320 or optionally the low-resolution LiDAR system 510 can be used instead of the high-resolution high-precision LiDAR system 310, which is typically very expensive and not suitable for commercial use. The high-resolution high-precision point cloud generated by the trained preprocessing module 330 or 530 can be used by any machine learning algorithm that processes point cloud data to generate accurate results, which can be used to generate an accurate 3D representation of the environment.

[0143] The steps and / or operations in the flowcharts and diagrams described herein are merely examples. These steps and / or operations can have many variations without departing from the content of the present invention. For example, these steps can be executed in a different order, or steps can be added, deleted, or modified.

[0144] After those of ordinary skill in the art understand the present invention, they can know the coding of the software described for implementing the above method. The machine-readable code that can be executed by one or more processors of one or more devices respectively to implement the above method can be stored in a computer-readable medium such as the memory of a data manager. The terms "software" and "firmware" in the present invention can be interchanged and include any computer program stored in a memory for execution by a processor, and the memory includes random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and non-volatile RAM (NVRAM). The above memory types are only examples, and thus do not limit the memory types that can be used to store computer programs.

[0145] General rules

[0146] All values and subintervals in the disclosed intervals are also disclosed. In addition, although the systems, devices, and processes disclosed and illustrated herein may include a plurality of specific elements, the systems, devices, and components can be modified to include more or fewer such elements. Although several exemplary embodiments are described herein, they can be modified, adapted, or otherwise implemented. For example, the illustrated elements can be replaced, added, or modified, and the exemplary methods described herein can be changed by replacing, reordering, or adding steps of the disclosed methods. In addition, many specific details are set forth in order to provide a thorough understanding of the exemplary embodiments described herein. However, those of ordinary skill in the art can understand that the exemplary embodiments described herein can be practiced without these specific details. In addition, well-known methods, processes, and elements are not described in detail to make the exemplary embodiments described herein clear. The subject matter described herein is intended to cover and encompass all suitable technical variations.

[0147] Although part of the present invention is described as a method, those of ordinary skill in the art can understand that the present invention also provides various different elements for at least implementing certain aspects or features of the method, and the elements can be hardware, software, or a combination thereof. Accordingly, the technical solution of the present invention can be embodied as a non-volatile or non-transitory computer-readable medium (for example, an optical disc, a flash memory, etc.) storing executable instructions, wherein the executable instructions cause a processing device to execute the examples of the methods disclosed herein.

[0148] The term "processor" may include any programmable system, including systems using microprocessors / microcontrollers or nano-processors / nano-controllers, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), reduced instruction set circuits (RISC), logic circuits, and any other circuit or processor capable of performing the functions described herein. The term "database" may refer to a body of data, a relational database management system (RDBMS), or both. As used herein, a "database" may include any collection of data, including hierarchical databases, relational databases, flat file databases, object-relational databases, object-oriented databases, and any other structured collection of records or data stored in a computer system. The above examples are for illustrative purposes only and, therefore, do not limit the definition and / or meaning of the terms "processor" or "database" in any way.

[0149] The present invention may be embodied in other specific forms without departing from the subject matter of the claims. The described exemplary embodiments are illustrative in all respects and not restrictive. The present invention is intended to cover and embrace all applicable technical variations. Accordingly, the scope of the present invention is defined by the appended claims rather than the above description. The scope of the claims should not be limited by the embodiments in the examples, but should be given the broadest interpretation consistent with the overall description.

Claims

1. A computer vision system, characterized in that, Comprising: A processor system; A memory coupled to the processor system, wherein executable instructions are tangibly stored on the memory, and when the executable instructions are executed by the processor system, cause the computer vision system to: (i) Receive a camera point cloud from a camera system; (ii) Receive a first LiDAR point cloud from a first LiDAR system having a first resolution; (iii) Determine the error of the camera point cloud with reference to the first LiDAR point cloud; (iv) Determine a correction function based on the determined error; (v) Use the correction function to generate a corrected point cloud based on the camera point cloud; Wherein, in case A: (vi) Determine the training error of the corrected point cloud with reference to the first LiDAR point cloud; (vii) Update the correction function based on the determined training error; Or, in case B: (vi’) Calculate the output of a point cloud application algorithm using the corrected point cloud, and calculate the output of the point cloud application algorithm using the first LiDAR point cloud; (vii’) Determine the loss of the error of the output of the point cloud application algorithm using the corrected point cloud through a loss function; and (viii’) Update the correction function based on the determined loss; Wherein, when the executable instructions are executed by the processor system, further cause the computer vision system to: (a) Receive a second LiDAR point cloud from a second LiDAR system having a second resolution, wherein the second resolution of the second LiDAR system is higher than the first resolution of the first LiDAR system; (b) Determine the error of the corrected point cloud with reference to the second LiDAR point cloud; (c) Determine a correction function based on the determined error; (d) Generate a corrected point cloud using the correction function; Wherein, in case A: (e) Determine the second training error of the corrected point cloud with reference to the second LiDAR point cloud; and (f) Update the correction function based on the determined second training error; Or, in case B: (e’) Calculate the output of a point cloud application algorithm using the corrected point cloud, and calculate the output of the point cloud application algorithm using the second LiDAR point cloud; (f’) Determine the second loss of the error of the output of the point cloud application algorithm using the corrected point cloud through a loss function; and (g’) Update the correction function based on the determined second loss.

2. The computer vision system according to claim 1, characterized in that, When the executable instructions are executed by the processor system, cause the computer vision system to repeatedly execute operations (v) to (vii) in case A until the training error is less than an error threshold.

3. The computer vision system according to claim 1 or 2, characterized in that, When the executable instructions are executed by the processor system, cause the computer vision system to: When determining the error of the camera point cloud with reference to the first LiDAR point cloud and determining a correction function based on the determined error, specifically perform the following operations: Generate a mathematical representation of the error of the camera point cloud with reference to the first LiDAR point cloud; Determine a correction function based on the mathematical representation of the error; When determining the training error of the corrected point cloud with reference to the first LiDAR point cloud and updating the correction function based on the determined training error, the following operations are specifically performed: Generate a mathematical representation of the training error of the corrected point cloud with reference to the first LiDAR point cloud; Update the correction function based on the mathematical representation of the training error.

4. The computer vision system according to claim 1, wherein When the executable instructions are executed by the processor system, cause the computer vision system to: repeatedly execute operations (d) to (f) in case A until the second training error is less than the error threshold; repeatedly execute operations (d), (e’), (f’), and (g’) in case B until the second loss is less than the loss threshold.

5. The computer vision system according to claim 1, wherein The second LiDAR system includes one or more 64-beam LiDAR units, and the first LiDAR system includes one or more 8-beam, 16-beam, or 32-beam LiDAR units.

6. The computer vision system according to any one of claims 1 to 5, characterized in that, The camera system is a stereo camera system, and the camera point cloud is a stereo camera point cloud.

7. The computer vision system according to any one of claims 1 to 5, characterized in that, The processor system includes a neural network.

8. The computer vision system according to any one of claims 1 to 5, characterized in that, When the executable instructions are executed by the processor system, cause the computer vision system to repeatedly execute operations (v), (vi’), (vii’), and (viii’) in case B until the loss is less than the loss threshold.

9. A training method for generating a corrected point cloud, characterized in that, The method is applied to a preprocessing module, and the method includes: (i) Receive a camera point cloud from a camera system; (ii) Receive a first LiDAR point cloud from a first LiDAR system having a first resolution; (iii) Determine the error of the camera point cloud with reference to the first LiDAR point cloud; (iv) Determine a correction function based on the determined error; (v) Use the correction function to generate a corrected point cloud based on the camera point cloud; Wherein, in case A: (vi) Determine the training error of the corrected point cloud with reference to the first LiDAR point cloud; (vii) Update the correction function based on the determined training error; Or, in case B: (vi’) Calculate the output of the point cloud application algorithm using the corrected point cloud and calculate the output of the point cloud application algorithm using the first LiDAR point cloud; (vii’) Determine the loss of the error of the output of the point cloud application algorithm using the corrected point cloud through a loss function; and (viii’) Update the correction function based on the determined loss; Wherein, the method further includes: (a) Receive a second LiDAR point cloud from a second LiDAR system having a second resolution, wherein the second resolution of the second LiDAR system is higher than the first resolution of the first LiDAR system; (b) Determine the error of the corrected point cloud with reference to the second LiDAR point cloud; (c) Determine a correction function based on the determined error; (d) Generate a corrected point cloud using the correction function; Wherein, in case A: (e) Determine the second training error of the corrected point cloud with reference to the second LiDAR point cloud; and (f) Update the correction function based on the determined second training error; Alternatively, in Case B: (e’) Use the output of the point cloud application algorithm computed with the corrected point cloud, and use the second LiDAR point cloud to compute the output of the point cloud application algorithm; (f’) Use the corrected point cloud to determine a second loss of the error of the output of the point cloud application algorithm through a loss function; and (g’) Update the correction function based on the determined second loss.

10. The method according to claim 9, wherein Repeat operations (v) to (vii) in Case A until the training error is less than the error threshold.

11. The method according to claim 9 or 10, wherein determining the error of the camera point cloud with reference to the first LiDAR point cloud and determining a correction function based on the determined error includes: generating a mathematical representation of the error of the camera point cloud with reference to the first LiDAR point cloud; determining a correction function based on the mathematical representation of the error; determining the training error of the corrected point cloud with reference to the first LiDAR point cloud and updating the correction function based on the determined training error includes: generating a mathematical representation of the training error of the corrected point cloud with reference to the first LiDAR point cloud; updating the correction function based on the mathematical representation of the training error.

12. The method according to claim 9, characterized in that, Repeat operations (d) to (f) in Case A until the second training error is less than the error threshold; repeat operations (d), (e’), (f’), and (g’) in Case B until the second loss is less than the loss threshold.

13. The method according to claim 9, wherein The second LiDAR system includes one or more 64-beam LiDAR units, and the first LiDAR system includes one or more 8-beam, 16-beam, or 32-beam LiDAR units.

14. The method according to any one of claims 9 to 12, characterized in that, The camera system is a stereo camera system, and the camera point cloud is a stereo camera point cloud.

15. The method according to any one of claims 9 to 12, characterized in that The preprocessing module includes a neural network.

16. The method according to any one of claims 9 to 12, characterized in that, Repeat operations (v), (vi’), (vii’), and (viii’) in Case B until the loss is less than the loss threshold.

17. A computer-readable storage medium including instructions, characterized in that, When the instructions are executed by a processor of a computer vision system, cause the computer vision system to execute the method according to any one of claims 9 to 16.

Citation Information

Patent Citations

  • Low-wire harness laser radar and binocular camera-based fusion locating method and device

    CN108694731A

  • System and method for fusing outputs of sensors having different resolutions

    WO2017122529A1