Method, apparatus and computer program product for processing multi-echo data

By generating confidence and saliency maps in optical detection and ranging equipment, and combining neural networks and noise probability models, the effective echoes are distinguished and processed, thus solving the problem of noise echo interference and improving detection accuracy and resource utilization efficiency.

CN120405618APending Publication Date: 2025-08-01SONY GROUP CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410142027.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing optical detection and ranging equipment suffers from noise echo interference in the effective echo detection when processing multi-echo data, affecting the accuracy of target object detection and 3D modeling, and increasing resource requirements.

Method used

By generating a confidence map for each pixel in the depth map, the system assesses whether the echo data is noisy echo and generates a saliency map. It prioritizes the echo data of pixels with high effective information content and uses a neural network model and a pixel-specific noise probability model to distinguish between effective echoes and noisy echoes.

Benefits of technology

It reduces the interference of noise echoes on effective echoes, improves detection performance, reduces the demand for computing and storage resources, and improves the accuracy and efficiency of target object detection and 3D modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120405618A_ABST
    Figure CN120405618A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, an apparatus, and a computer program product for processing multi-echo data. Various embodiments for processing multi-echo data generated by a light detection and ranging device are described. In one embodiment, an example method includes: for each pixel point in a depth map of an environment generated by a light detection and ranging device, obtaining a plurality of echo data; determining a probability that a corresponding echo data of the plurality of echo data corresponds to a valid echo or a noisy echo, thereby forming a confidence map corresponding to a depth map of the environment; a saliency map corresponding to the depth map of the environment is generated based on the confidence map, wherein the saliency map is used to indicate an amount of effective information in a corresponding portion of the depth map of the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to multi-echo data processing, including techniques for processing multi-echo data generated by a light detection and ranging device. Background Art

[0002] A light detection and ranging device is a device for detecting and locating objects. Its working process includes emitting light through a transmitter and receiving the light reflected back after hitting a target object through a receiver. The distance to the target object is measured based on the time difference between light emission and reception (i.e., time of flight, ToF). The receiver records the number of photons received at different times, which is called echo data. After the light is reflected by the object, an echo is formed in the echo data, and the distance to the target object can be determined by the position of this echo. The light detection and ranging device can also present the three-dimensional structure information of the target object by analyzing information such as the magnitude of the reflected energy on the surface of the target object, the amplitude, frequency, and phase of the reflected spectrum.

[0003] In the case where the light source is a laser, the light detection and ranging device becomes a lidar (Light Detection and Ranging).

[0004] Through a light detection and ranging device such as lidar, multi-echo data corresponding to a single pixel can be obtained. In applications such as distance measurement and three-dimensional structure construction of a target object, it is desirable to effectively and efficiently utilize the multi-echo data. Summary of the Invention

[0005] A first aspect of the present disclosure relates to a method for processing multi-echo data. According to one embodiment, the method includes: for each pixel point in the depth map of the environment generated by a light detection and ranging device, obtaining a plurality of echo data; determining the probabilities that the corresponding echo data in the plurality of echo data correspond to valid echoes or noise echoes, thereby forming a confidence map corresponding to the depth map of the environment; and generating a saliency map corresponding to the depth map of the environment based on the confidence map, where the saliency map is used to indicate the amount of valid information in the corresponding part of the depth map of the environment. The first aspect of the present disclosure also relates to a light detection and ranging device for performing the method.

[0006] The second aspect of the present disclosure relates to a method for training a neural network model. According to one embodiment, the method includes: obtaining a set of data samples, where the set of data samples includes multiple echo data of each pixel point in depth maps generated by a plurality of light detection and ranging devices; for each pixel point, adding a label of a valid echo or a noise echo to the corresponding data sample using a pixel-point specific noise probability model; and using the set of data samples and the corresponding labels as training data to train the neural network model. The second aspect of the present disclosure also relates to an electronic device for performing the method.

[0007] The third aspect of the present disclosure relates to a method for establishing a noise probability model. According to one embodiment, the method includes: obtaining a set of data samples, where the set of data samples includes multiple echo data of a first pixel point in depth maps generated by a plurality of light detection and ranging devices; and determining parameters of a first pixel-point specific noise probability model based on the multiple echo data of the first pixel point. The third aspect of the present disclosure also relates to an electronic device for performing the method.

[0008] The present disclosure also relates to a computer program product including one or more instructions. When the one or more instructions are executed by a processor, various methods according to the embodiments of the present disclosure are implemented. For example, the methods include a method for processing echo data, a method for training a neural network model, and a method for establishing a noise probability model.

[0009] The above summary is provided to summarize some exemplary embodiments to provide a basic understanding of aspects of the subject matter described herein. Therefore, the above features are merely examples and should not be construed as narrowing the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the following detailed description in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] A better understanding of the present disclosure can be obtained when considering the following detailed description of the embodiments in conjunction with the accompanying drawings. The same or similar reference numerals are used in the various drawings to denote the same or similar components. The various drawings, together with the following detailed description, are included in this specification and form a part of the specification, and are used to illustrate the embodiments of the present disclosure and to explain the principles and advantages of the present disclosure. Among them:

[0011] Figure 1 An example of a light detection and ranging device according to an embodiment of the present disclosure is shown.

[0012] Figure 2A A schematic diagram of an application scenario for light detection and ranging according to an embodiment of the present disclosure is shown.

[0013] Figure 2BShows a schematic diagram of echo data obtained in the Figure 2A application scenario of

[0014] Figure 3 Shows an example method for processing multi-echo data according to an embodiment of the present disclosure.

[0015] Figure 4A and Figure 4B Shows an example of a confidence map corresponding to a depth map of the environment according to an embodiment of the present disclosure.

[0016] Figure 4C Shows an example of a saliency map corresponding to a depth map of the environment according to an embodiment of the present disclosure.

[0017] Figure 5 Shows an example method for training a neural network model according to an embodiment of the present disclosure.

[0018] Figure 6 Shows an example of a neural network model according to an embodiment of the present disclosure.

[0019] Figure 7 Shows an example of a convolutional neural network model according to an embodiment of the present disclosure.

[0020] Figure 8 Shows an example of a simulation model for optical detection and ranging according to an embodiment of the present disclosure.

[0021] ​ Shows an example method for establishing a noise probability model according to an embodiment of the present disclosure.

[0022] ​ Shows an example of a noise probability model according to an embodiment of the present disclosure.

[0023] ​ Shows an example block diagram of an electronic device for implementing various methods according to an embodiment of the present disclosure.

[0024] Although the embodiments described in the present disclosure may be susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and are described in detail herein. However, it should be understood that the drawings and the detailed description thereof are not intended to limit the embodiments to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternative schemes falling within the spirit and scope of the claims. Detailed Description

[0025] The following describes representative applications of aspects such as devices and methods according to the present disclosure. The description of these examples is only to add context and help understand the described embodiments. Thus, it will be clear to those skilled in the art that the embodiments described below can be implemented without some or all of the specific details. In other cases, well-known process steps are not described in detail to avoid unnecessarily obscuring the described embodiments. Other applications are possible, and the solutions of the present disclosure are not limited to these examples.

[0026] ​

[0027] ​ An example of a light detection and ranging device and its operation according to an embodiment of the present disclosure is shown.

[0028] In ​ the example, the light detection and ranging device 100 may include a transmitter 110, a receiver 120, a processor 130, and a memory 140. The light detection and ranging device 100 may be configured to detect and range objects in the surrounding environment by transmitting and receiving optical signals, and even perform three-dimensional modeling. For example, the transmitter 110 may be configured to transmit an optical signal 115 through a light source. The transmitted optical signal 115 propagates through a medium to reach a target object 160 in the environment and is reflected from the object 160. The reflected optical signal 115' propagates through the medium and returns to the light detection and ranging device 100 and is received by the receiver 120. In one embodiment, the transmitter 110 and the receiver 120 may each include an optical lens (not shown). The receiver 120 may include a detector to detect the received optical signal 115'. In one embodiment, the light source is configured to transmit a laser signal, and the laser has a plurality of pulses with specific sequences. In this way, the light detection and ranging device 100 will become a lidar (LiDAR).

[0029] In the example, the processor 130 may be configured to be coupled to the transmitter 110, the receiver 120, and the memory 140. The processor 130 may execute one or more modules and / or processes to enable the light detection and ranging device 100 to perform various functions. The functions include controlling the transceiver of optical signals (such as laser signals) and various functions described below. For example, the processor 130 may be configured to execute various functions by reading and executing computer programs, codes, or executable instructions stored in the memory 140. In some embodiments, the processor 130 may include a microprocessor, a microcontroller, a digital signal processor, a central processing unit (CPU), a graphics processing unit (GPU), etc.

[0030] The memory 140 may be a non-transitory computer-readable storage medium, including but not limited to electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanically encoded devices, and any suitable combination of the above.

[0031] ​ Only one target object 160 is shown. The target object can be any type of detection object, including but not limited to trees, furniture, roadblocks, vehicles, pedestrians, workpieces, etc. In actual situations, there are usually various types of multiple objects in the environment, making the detection, ranging, and three-dimensional modeling of the target object complex.

[0032] ​ A schematic diagram of an application scenario for optical detection and ranging according to an embodiment of the present disclosure is shown. Those skilled in the art should understand that ​ The application scenario shown is only an example, and the application scenarios for optical detection and ranging of the present disclosure are not limited thereto. ​ A schematic diagram of the echo data obtained in the application scenario of ​ is shown. The following describes this example application scenario in combination with ​ the optical detection and ranging device 100 in

[0033] As ​ shown, in the optical detection and ranging scheme, the transmitter 110 is configured to emit optical signals 210 and 220. The optical signal 210 propagates through the medium to reach the surface of the front object and is reflected from the object. The reflected optical signal 210' propagates through the medium and returns to the optical detection and ranging device 100 and is received by the receiver 120. The optical signal 220 propagates through the medium to reach the edge of the front object. Then, a part of the optical signal 220 is reflected from the edge of the front object, and the reflected optical signal 220' returns to the optical detection and ranging device 100 and is received by the receiver 120. Another part of the optical signal 220 continues to propagate through the medium to reach the rear object and is reflected from the surface of the object. The reflected optical signal 220'' returns to the optical detection and ranging device 100 and is received by the receiver 120. The receiver 120 may be configured to record the number of photons received at different times, thereby generating echo data.

[0034] The light detection and ranging device 100 may have a plurality of pixel points corresponding to objects in the field of view. Here, the meaning of a pixel point is similar to that of a pixel in image sensing, but the data recorded for a pixel point is not the RGB color components, but the number of received photons (i.e., echo data). Each pixel point may correspond to one or more echo data. In ​ In the example, after the optical signal 210 hits point A1 on the front object, it is received by the receiver 120, thereby generating one echo data for the corresponding pixel point. In the example, the pixel point corresponding to point A1 on the front object may be represented as pixel point A1. The optical signal 220 is received by the receiver 120 after hitting point A2 on the front object and point B2 on the rear object, thereby generating two echo data for the corresponding pixel point. In the example, the pixel point corresponding to point A2 on the front object and point B2 on the rear object may be represented as pixel point A2 / B2.

[0035] The pixel points may be in one-to-one correspondence with the field of view angles of the light detection and ranging device 100, and thus can be represented and distinguished by the field of view angles. In addition, the density of the pixel points can reflect the angular resolution of the light detection and ranging device 100. For the sake of simplicity of description, ​ only two pixel points A1 and A2 / B2 generated by the light detection and ranging device 100 are shown. In one or more embodiments, depending on the configuration of the light detection and ranging device 100, a greater number of pixel points may be generated.

[0036] In ​ only the optical signals emitted by the transmitter 110 and the real object echoes are shown. In the present disclosure, the real object echoes are referred to as valid echoes. However, generally, the receiver 120 will also receive optical signals caused by the propagation of ambient light (such as sunlight). Relative to the real object echoes, these optical signals constitute noise echoes. The noise echoes will interfere with the detection of the valid echoes, thereby affecting the detection accuracy of the valid echoes and even resulting in the inability to detect the valid echoes.

[0037] ​ Examples (1) and (2) respectively show a plurality of echo data corresponding to the two pixel points A1 and A2 / B2. In ​ In the example, the abscissa represents time and thus can indicate the flight time of the echo. The ordinate represents the number of received photons and thus can indicate the echo intensity. In the example, there may also be some lower peaks. Since these lower peaks do not meet the determination criteria of the echo (such as intensity and duration conditions), they may be disregarded.

[0038] Referring to Example (1), echo data a corresponds to the optical signal reflected by point A1 of the front object, and thus is a valid echo. Echo data b and c both correspond to noise echoes. Referring to Example (2), echo data d corresponds to the optical signal reflected by point A2 of the front object, and echo data e corresponds to the optical signal reflected by point B2 of the rear object, and thus both are valid echoes. Echo data f corresponds to a noise echo.

[0039] It should be understood that in the case of a higher peak, a noise echo may be misidentified as a valid echo, thus interfering with the detection of valid echoes. In Examples (1) and (2), the noise echoes b, c, or f may all affect the detection accuracy of the corresponding valid echoes. This can be significantly disadvantageous for the detection and ranging of target objects and three-dimensional modeling. It should also be understood that multiple valid echoes of a single pixel point can convey more useful information. For example, these multiple valid echoes may correspond to multiple target objects. Therefore, detecting and excluding noise echoes for a single pixel point, and / or detecting and identifying multiple valid echoes is beneficial for the detection and ranging of target objects and the surrounding environment and three-dimensional modeling.

[0040] In an embodiment of the present disclosure, the echo data of each pixel point in the depth map is formed into a confidence score to evaluate whether the echo data is a noise echo. Based on this evaluation, the noise echoes of the pixel points can be excluded, and only the valid echoes are retained. On the one hand, this can reduce the interference of noise echoes on valid echoes and improve the detection performance. On the other hand, the demand for computing and storage resources originally caused by the processing and storage of noise echoes can be reduced at the pixel point level.

[0041] In an embodiment of the present disclosure, a saliency score is formed for each pixel point or pixel point set in the depth map to evaluate the effective information amount (such as the number of valid echoes) of the pixel point or pixel point set. Based on this evaluation, the echo data of the pixel points or pixel point sets with a high effective information amount can be preferentially processed. By reducing the processing of the echo data of pixel points with a low effective information amount, the processing resources of echo data can be saved at the level of the entire depth map and even storage resources can be saved.

[0042] Still referring to ​ , according to an embodiment of the present disclosure, the memory 140 of the optical detection and ranging device 100 may include modules, such as an echo data acquisition module 142, a confidence analysis module 144, and a saliency analysis module 146, so as to improve the performance of the optical detection and ranging device 100 in view of the above and other aspects.

[0043] As an example, the echo data acquisition module 142 may be configured to obtain multiple echo data for each pixel point in the depth map of the environment generated by the optical detection and ranging device 100. The confidence analysis module 144 may be configured to determine the probability that the corresponding echo data among the multiple echo data corresponds to a valid echo or a noise echo, thereby forming a confidence map corresponding to the depth map of the environment. The saliency analysis module 146 may be configured to generate a saliency map corresponding to the depth map of the environment based on the confidence map, where the saliency map is used to indicate the amount of valid information in the corresponding part of the depth map of the environment. The detailed operations of each module can be understood in combination with the further description of the embodiments of the present disclosure.

[0044] In an embodiment, the term module is used to represent an example division of executable instructions for ease of discussion. It should be noted that the module division can be performed in different ways, and one or more functions can be arranged in different ways (for example, combined into a smaller number of modules, divided into a larger number of modules, etc.). In addition, the functions and modules described herein can be implemented in whole or in part by software and / or firmware executable on a processor, or can be implemented in whole or in part by hardware (such as dedicated processing circuits, etc.).

[0045] ​

[0046] ​ An example method for processing multi-echo data according to an embodiment of the present disclosure is shown. The example method 300 can be executed by various optical detection and ranging devices. The example operations of the method 300 are described below in conjunction with the optical detection and ranging device 100.

[0047] As ​ shown, the method 300 may include obtaining multiple echo data (310) for each pixel point in the depth map of the environment generated by the optical detection and ranging device 100. The pixel point is, for example, ​ pixel point A2 / B2 in, and the multiple echo data includes, for example, echo data d, e, and f. Generally speaking, the obtained echo data is echo data that meets specific determination criteria. For example, the determination criteria may require that the intensity of the echo is higher than a first predetermined threshold, and / or the duration of the echo is longer than a second predetermined threshold. It should be understood that any number of echo data can be obtained for a single pixel point, such as 4, 5, 6, or more. In one or more embodiments, the echo data may include information such as the echo start position, echo peak position, echo end position, and light intensity of the corresponding echo data.

[0048] Accordingly, method 300 may include determining the probabilities that respective echo data among the multiple echo data correspond to valid echoes or noise echoes, thereby forming a confidence map corresponding to multiple pixel points in the depth map (320). Further, the echo data may include the probabilities that respective echo data correspond to valid echoes or noise echoes, or may include flags indicating whether respective echo data are valid echoes or noise echoes. As described above, for a light detection and ranging device, optical signals caused by the propagation of ambient light (such as sunlight) will constitute noise echoes. Noise echoes will interfere with the detection of valid echoes, thereby affecting the detection accuracy of valid echoes and even resulting in the inability to detect valid echoes. In ​ the example, the peak difference between noise echo b and valid echo a or between noise echo f and valid echo e may not be sufficient to always pick out the noise echo. Accordingly, it is desirable to effectively distinguish between valid echoes and noise echoes. In one or more embodiments, a neural network model may be used to distinguish between valid echoes and noise echoes. For example, using a neural network model to distinguish between valid echoes and noise echoes may include using a unified neural network model to distinguish between valid echoes and noise echoes for multiple pixel points in the depth map. Alternatively, in one or more embodiments, a pixel-point specific noise probability model may be used to distinguish between valid echoes and noise echoes. For example, using a noise probability model to distinguish between valid echoes and noise echoes may include using multiple pixel-point specific noise probability models to distinguish between valid echoes and noise echoes for multiple pixel points in the depth map.

[0049] In an embodiment, the confidence map may be used to represent the probabilities or flags that the multiple echo data of each pixel point correspond to valid echoes or noise echoes, respectively. ​ and ​ shows an example of a confidence map corresponding to a depth map of the environment according to an embodiment of the present disclosure. In this example, it is assumed that the light detection and ranging device 100 has a pixel point configuration of 3 rows by 3 columns. Accordingly, the depth map may have a total of 9 pixel points of 3 rows by 3 columns. In this example, each row corresponds to a different pixel point. Pixel point (1, 1) corresponds to the pixel point in the first row and the first column, and so on. The columns correspond to multiple (such as 6) echo data. The cell at the intersection of the row and the column corresponds to a single echo data detected at the corresponding pixel point. It is easy to understand that the above pixel point configuration and the number of echo data are only examples. In one or more embodiments, a higher or lower pixel point configuration may be provided. The number of echo data for each pixel point may also be more or less, and the number of echo data between multiple pixel points may not have to be the same.

[0050] In ​ the example, the probability in the cell at the intersection of the row and the column represents the likelihood that the single echo data at the corresponding pixel point corresponds to a valid echo or a noise echo. As​ As shown, the probability that the echo data 1 at pixel point (1, 1) corresponds to a valid echo is 0.95 (or the probability that it corresponds to a noise echo is 0.05). The probability that the echo data 6 detected at pixel point (2, 3) corresponds to a valid echo is 0.13 (or the probability that it corresponds to a noise echo is 0.87). In one or more embodiments, one or more confidence thresholds may be preset, so as to determine the echo data with a probability of corresponding to a valid echo higher than the corresponding confidence threshold as a valid echo.

[0051] In ​ the example, the single echo data at the corresponding pixel point is represented as corresponding to a valid echo or a noise echo by the mark in the cell where the row and column intersect. ​ the example is based on ​ the probability information in

[0052] It should be understood that generating multiple echo data based on a single pixel point in the depth map will multiply the amount of data, and thus more resources are required for storage and processing. The increased resource requirements will be more obvious when there are more pixel points. In one or more embodiments, the demand for storage and processing resources of the multi-echo data is controlled at an appropriate level by excluding noise echoes from a single pixel point. For example, the echo data 3, 5, and 6 of pixel point (1, 1) can be excluded.

[0053] As ​ shown, method 300 may further include generating a saliency map (330) corresponding to the depth map based on the confidence map. The saliency map is used to indicate the saliency score of the corresponding pixel point or set of pixel points in the depth map of the environment. The saliency score is used to indicate the amount of valid information of the corresponding pixel point or set of pixel points. In one embodiment, the amount of valid information may refer to the number of valid echoes or the proportion of valid echoes in all the obtained echoes.

[0054] ​ shows an example of a saliency map corresponding to the depth map of the environment according to an embodiment of the present disclosure. This example is the same as ​ and ​ which are based on the pixel points of the same depth map, that is, there are a total of 9 pixel points with 3 rows and 3 columns. Each pixel point corresponds to 6 echo data. In the example, among the 6 echo data of pixel point (1, 1), there are 3 valid echoes, and its saliency score is 3 or 1 / 2. Among the 6 echo data of pixel point (2, 3), there are 5 valid echoes, and its saliency score is 5 or 5 / 6.

[0055] It should be understood that the higher the significance score, the more corresponding effective information. Therefore, the echo data of pixel points or sets of pixel points with high effective information can be preferentially processed. For example, in terms of allocating storage and processing resources, a higher priority can be given to pixel points or sets of pixel points with high effective information, and the priority of pixel points or sets of pixel points with low significance scores can be reduced. In this way, the processing resource requirements and even storage resource requirements of the echo data can be controlled at an appropriate level at the level of the entire depth map. For example, assuming that a threshold is preset to be 3, pixel points higher than or not lower than this threshold can be preferentially processed or stored. In ​ this case, there are 6 such pixel points, accounting for 2 / 3 of the total. The processing and even storage resources of the remaining 1 / 3 of the pixel points can be saved.

[0056] In one embodiment, the echo data obtained by the optical detection and ranging device 100 will be used for data fusion with the data obtained by other sensors. For example, in an autonomous driving application, it is generally necessary to fuse the data of lidar and the data of a camera to make up for the deficiencies of the two sensors and improve the accuracy and quality of the application. Through the significance score of a single pixel point or the entire significance map, only the effective echoes of pixel points with a significance score higher than a specific threshold can participate in data fusion. This can be beneficial for reducing the computational resource requirements of data fusion.

[0057] In one embodiment, a significance map can be output in real time based on the user's region of interest. For example, the user can determine their region of interest based on a specific application scenario. For example, for an indoor scene, the region of interest may be a cube with a side length of 3 meters. For an outdoor scene, the region of interest may be a cube with a side length of 10 meters. After selecting the region of interest, inside the optical detection and ranging device 100, a significance map within the region of interest is output in real time based on the global confidence map and multiple echo data. Then, based on the significance map within the region of interest output in real time, multi-echo data can be stored and processed.

[0058] ​

[0059] In the embodiments of the present disclosure, a neural network model can be used to distinguish effective echoes and noise echoes. When the performance of the neural network model is ensured through training, effective echoes and noise echoes can be quickly and accurately distinguished, achieving the real-time performance and accuracy of the operation. ​ An example method for training a neural network model according to an embodiment of the present disclosure is shown. The example method 500 can be executed by any computer or electronic device.

[0060] As ​As shown, method 500 may include obtaining a first set of data samples (510). The first set of data samples includes multiple echo data for each pixel point in depth maps generated by multiple light detection and ranging devices. In one or more embodiments, the first set of data samples is obtained during multiple frames. The continuous dynamic data of the multiple frames can enhance the diversity of the data samples.

[0061] As ​ shown, method 500 may include, for each pixel point, using a noise probability model to separately add labels of valid echoes or noise echoes to the corresponding data samples (520). Since the propagation environment of the optical signal for each pixel point and the characteristics of the object surface it strikes are independent of each other, in one or more embodiments, the noise probability model used is pixel-point specific. In this way, the accuracy of the labels added to the data samples can be improved.

[0062] As ​ shown, method 500 may further include using the first set of data samples and the corresponding labels as training data to train a neural network model (530). Specifically, the multi-echo data for each pixel point can be used as the input to the neural network model, and the output of the neural network model can be compared with the labels of the data samples. Based on the comparison results, the parameters and structure of the neural network model can be adjusted to achieve the training of the neural network model.

[0063] ​ shows an example of a neural network model according to an embodiment of the present disclosure. As ​ shown, the neural network model 600 includes multiple layers, including an input layer 601, an output layer 606, and intermediate layers (or hidden layers) 602 to 605. Each layer has a certain number of neurons, each neuron has a specific weight value, and there are connections between neurons in different layers. When the input value 620 is input into the neural network model 600, first, the neurons at the input layer 601 receive the corresponding numerical values and propagate the corresponding numerical values to the neurons in the intermediate layer 602 through the connections with the neurons in the next layer. The neurons in the intermediate layer 602 calculate the weighted sum of the output values of the neurons in the previous layer and output the weighted sum to the neurons in the next intermediate layer 603 through the connections with the neurons in the next layer. And so on, until the neurons in the output layer 606 calculate the weighted sum of the output values of the neurons in the previous layer and output the inference result 640 for the input value 620.

[0064] In ​In an example, the neural network model 600 has four intermediate layers 602 to 605. Depending on application requirements, the number of intermediate layers can be any number, and the present disclosure does not limit this. In the case where the number of intermediate layers is more than a certain number, the neural network model is also referred to as a deep neural network model. The neural network model 600 consists of a series of fully connected layers (i.e., all outputs are connected to all inputs) and is called a multi-layer perceptron (MLP) model. As a further example, the neural network model also includes a convolutional neural network (CNN) and a recurrent neural network (RNN) model, etc. ​ shows an example of a convolutional neural network (CNN) model according to an embodiment of the present disclosure. As ​ shown, the CNN model 700 is divided into four modules, namely an input module, a preprocessing module, a convolutional module, and an output module.

[0065] In one or more embodiments, a unified neural network model network can be trained based on data samples in a variety of environments. In this way, for pixel points of different depth maps or different pixel points in the same depth map, a unified neural network model can be used to distinguish valid echoes and noise echoes. In this way, the processing for distinguishing valid echoes and noise echoes can be simplified.

[0066] In one or more embodiments, a real light detection and ranging device and a simulation model can be used to obtain a first set of data samples. That is to say, the first set of data samples can include real data and simulation data obtained by the real device and the simulation model for one or more environments. For example, multi-echo data for each pixel point can be continuously generated at the same position using the real device or the simulation model during multiple frames. In the simulation, the device and object positions or environmental settings, etc., can be adjusted to obtain more diverse data samples.

[0067] ​ shows an example of a simulation model for light detection and ranging according to an embodiment of the present disclosure. This example simulation model can be established through tools such as Matlab, Python, Unreal Engine, and Autoware.

[0068] As ​ shown, the following physical entities of the components need to be simulated: a transmitter, a transmitting optical device, a photon detector, and a receiving optical device. The transmitting optical device includes a lens and a cover glass. The photon detector includes a sensing, filtering, and processing unit. The receiving optical device includes a lens, a filter, and glass. The background environment, target object, and other objects also need to be simulated. The background environment includes optical noise. Different ambient light conditions can be simulated by controlling the strength of the optical noise.

[0069] By simulating the light signal transmission, propagation, and reception stages of the light detection and ranging solution, a sufficient number and variety of data samples can be obtained. This helps to solve the problem that real data samples may be insufficient or it may be time-consuming to obtain a sufficient number and variety of data samples. At the same time, the data samples in the first set still include real data samples, which enables the neural network model training to still reflect the real environment.

[0070] In one embodiment, the neural network model can be trained offline and the trained neural network model can be stored for local use in the light detection and ranging device 100.

[0071] ​

[0072] As described above, in the embodiments of the present disclosure, a pixel-specific noise probability model can be used to distinguish the valid echoes and noise echoes of the corresponding pixel points. In the neural network model training, based on the discrimination result, labels of valid echoes or noise echoes can be added to the corresponding data samples. ​ An example method for establishing a noise probability model according to an embodiment of the present disclosure is shown. Since the propagation environment of the light signal for each pixel point and the characteristics of the object surface it hits are independent of each other, a noise probability model can be established for a single pixel point. The example method 900 can be executed by any computer or electronic device.

[0073] As ​ shown, the method 900 may include obtaining a second set of data samples for a first pixel point (910). The second set of data samples includes multiple echo data of the first pixel point in the depth maps generated by multiple light detection and ranging devices. In one or more embodiments, the second set of data samples is obtained during multiple frames. The continuous dynamic data of multiple frames can enhance the diversity of the data samples. The method 900 may further include determining the parameters of the noise probability model specific to the first pixel point based on the multiple echo data of the first pixel point (920).

[0074] In one or more embodiments, real light detection and ranging devices and simulation models can be used to obtain the second set of data samples. That is, the second set of data samples may include real data and simulation data obtained by real devices and simulation models for one or more environments. For example, multiple echo data of the first pixel point can be continuously generated at the same location using real devices or simulation models during multiple frames. In the simulation, the device and object positions or environmental settings, etc. can be adjusted to obtain more diverse data samples.

[0075] For a specific pixel point, through histogram decomposition, multiple echo data corresponding to the pixel point can be obtained from the original data. The waveform of the echo data can be approximated as a Gaussian waveform with a tail.​ An example of a noise probability model according to an embodiment of the present disclosure is shown. For a Gaussian distribution as ​ shown, the parameters to be determined include the mean μ and the standard deviation σ.

[0076] In a scheme for establishing a noise probability model based on data samples obtained from a real device, it is assumed that 5 echo data are obtained from the original data of a single pixel point, and based on experience, it is assumed that the last 3 echo data are noise echoes, or it is assumed that the 3 echo data with lower peaks are noise echoes. Based on this assumption and based on the data sample [144, 120, 114, 113, 112], it can be determined that the mean of the Gaussian distribution followed by the noise echoes is 113 (i.e., (114 + 113 + 112) / 3). Further, the standard deviation can be determined by Equation (1) based on the following empirical algorithm.

[0077]

[0078] If((Firstecho - Mean)>Minstd: σ = Firstecho - Mean,

[0079] Else: σ = Min std Equation (1)

[0080] In Equation (1), Min std represents the minimum standard deviation, Max peak represents the maximum peak among multiple echo data, Mean represents the peak mean of multiple echo data, and First echo represents the peak of the first echo data. In one or more embodiments, the mean and standard deviation of the Gaussian distribution can be estimated based on echo data of multiple frames, so as to fit the noise probability model of a specific pixel point.

[0081] In one or more embodiments, using the noise probability model includes determining whether the multiple echo data obtained are valid echo data or noise echo data based on the noise probability model. In the above example of the Gaussian distribution, the echo data with a peak greater than or equal to the sum of the mean and the standard deviation of the Gaussian distribution can be determined as valid echo data, and the echo data with a peak less than the sum of the mean and the standard deviation can be determined as noise echo data.

[0082] It should be noted that due to the limitations of conditions such as the bandwidth of the light detection and ranging device itself, the real device may not be able to output complete original data, but can only output multiple echo data obtained through histogram decomposition. The noise obtained in this way may be an extreme value and does not represent the general situation. Moreover, in the above process of establishing the noise probability model, assumptions are made for the noise echoes (i.e., the last 3 echoes), and the determination of the standard deviation is also more based on empirical algorithms. These factors will reduce the accuracy of the established noise probability model, and the corresponding parameters will be biased parameters. As a supplement or alternative, the original data in the form of a histogram of a single pixel can be obtained through simulation, thus making up for the limitation that the real device cannot output complete original data.

[0083] Therefore, in one or more embodiments, the light detection and ranging device can be simulated, and based on the physical model of the simulated light detection and ranging device, the original data in the form of a histogram of a single pixel can be obtained. Then, a set of data samples is obtained by repeatedly sampling from the histogram in order to obtain a more accurate noise probability model.

[0084] The ​ simulation model shown can be used to obtain a set of data samples. As described above, the simulation can involve the simulation of the light signal emission, propagation, and reception phases. Only the simulation examples of the light signal emission and propagation phases are specifically described below. Those skilled in the art can understand that the light signal reception phase can be simulated based on the specific principles and parameters of different types of receiving devices (such as PD, APD, SAPD).

[0085] In the example simulation, the illuminance caused by the signal light can be calculated by Equation (2):

[0086]

[0087] where and represent the object illuminance caused by the signal light, the object reflection parameter, and the lens parameter in sequence.

[0088] The illuminance caused by the ambient light is calculated by Equation (3):

[0089]

[0090] where E obj,amb represents the object illuminance caused by the ambient light. The meanings of other parameters are as follows:

[0091] P t : optical power; Ω TX : light projection angle; ρ: object reflectivity; F: F value of the lens;

[0092] η RX : optical system power; R: distance to the object.

[0093] The counting rate caused by the signal light is calculated by Equation (4), that is, the number of photons caused by the signal light detected by the detector per unit time.

[0094]

[0095] The counting rate caused by the signal ambient light is calculated by Equation (5), that is, the number of photons caused by the ambient light detected by the detector per unit time.

[0096]

[0097] Among them, the meanings of the parameters are as follows:

[0098] A pix : pixel area; h: Planck constant; c: speed of light; λ: light wavelength; PDE: photon detection efficiency.

[0099] The probability of detecting photons at least once within a time bin is calculated by Equation (6).

[0100]

[0101] Among them, the meanings of the parameters are as follows:

[0102] CR sig : ideal counting rate of signal light; CR amb : ideal counting rate of ambient light; T bin :

[0103] a time slot of a histogram.

[0104] p hiah obeys the Poisson distribution, representing the probability that a single detection diode is triggered within a time bin for each light pulse. Assuming that the light pulse is repeated N times, the triggering situation of the detection diode obeys the binomial distribution. The number of times the detection diode is triggered due to the overall signal light and ambient light and the number of times it is triggered due to the ambient light (i.e., noise) are represented by Equations (7) and (8) respectively.

[0105]

[0106]

[0107] The original echo data containing noise of a single pixel can be obtained through the above process. Then, the mean and variance of the Gaussian distribution can be determined based on the echo data. Example values of the parameters set in the simulation are as follows.

[0108]

[0109] In one embodiment, a noise probability model can be established offline and the established noise probability model can be stored locally in the light detection and ranging device 100 for use.

[0110] Embodiments of the present disclosure also provide an electronic device. The electronic device includes one or more processors and one or more memories storing one or more instructions thereon. When the one or more instructions are executed by the one or more processors, the one or more processors are caused to execute the method for processing multi-echo data, the method for training a neural network model, or the method for establishing a noise probability model according to the embodiments of the present disclosure.

[0111] Embodiments of the present disclosure also provide a computer-readable storage medium storing one or more instructions thereon. When the one or more instructions are executed by a processor, the processor is caused to execute the method for processing multi-echo data, the method for training a neural network model, or the method for establishing a noise probability model according to the embodiments of the present disclosure.

[0112] Embodiments of the present disclosure also provide a computer program product including one or more instructions. When the one or more instructions are executed by a processor, the processor is caused to execute the method for processing multi-echo data, the method for training a neural network model, or the method for establishing a noise probability model according to the embodiments of the present disclosure.

[0113] It should be understood that the instructions in the computer-readable storage medium according to the embodiments of the present disclosure can be configured to perform operations corresponding to the above device and method embodiments. When referring to the above device and method embodiments, the embodiments of the computer-readable storage medium are clear to those skilled in the art, and thus will not be described repeatedly. The computer-readable storage medium for carrying or including the above instructions also falls within the scope of the present disclosure. Such computer-readable storage media may include, but are not limited to, floppy disks, optical discs, magneto-optical discs, memory cards, memory sticks, and the like.

[0114] Embodiments of the present disclosure also provide various devices including components or units for performing the steps of the method for processing multi-echo data, the method for training a neural network model, or the method for establishing a noise probability model in the above embodiments.

[0115] It should be noted that the above-mentioned various components or units are only logical modules divided according to their specific functions, rather than limiting the specific implementation methods. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned various components or units can be implemented as independent physical entities, or can also be implemented by a single entity (such as a processor (CPU or DSP, etc.), integrated circuit, etc.). For example, in the above embodiments, multiple functions included in one unit can be implemented by separate devices. Alternatively, in the above embodiments, multiple functions implemented by multiple units can be respectively implemented by separate devices. In addition, one of the above functions can be implemented by multiple units.

[0116] In addition, it should be understood that the above-mentioned series of processes and devices can also be implemented by software and / or firmware. In the case of implementation by software and / or firmware, a program constituting the software is installed from a storage medium or network into a computer having a dedicated hardware structure, such as ​ the electronic device 1300 shown. When various programs are installed in the electronic device, it can perform various functions and so on. ​ The example block diagram of an electronic device for implementing various methods according to embodiments of the present disclosure is shown.

[0117] In ​ it, the central processing unit (CPU) 1301 executes various processes according to the program stored in the read-only memory (ROM) 1302 or the program loaded from the storage section 1308 into the random access memory (RAM) 1303. In the RAM 1303, data required when the CPU 1301 executes various processes, etc. is also stored as needed.

[0118] The CPU 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. The input / output interface 1305 is also connected to the bus 1304.

[0119] The following components are connected to the input / output interface 1305: an input section 1306, including a keyboard, mouse, etc.; an output section 1307, including a display, such as a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308, including a hard disk, etc.; and a communication section 1309, including a network interface card such as a LAN card, modem, etc. The communication section 1309 performs communication processing via a network such as the Internet.

[0120] As needed, a drive 1310 is also connected to the input / output interface 1305. A removable medium 1311, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1310 as needed, so that the computer program read therefrom is installed into the storage section 1308 as needed.

[0121] In the case where the above series of processes are implemented by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as a removable medium 1311.

[0122] Those skilled in the art should understand that such a storage medium is not limited to ​ the removable medium 1311 shown in which a program is stored and distributed separately from the device to provide the program to the user. Examples of the removable medium 1311 include a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including a compact disc read-only memory (CD-ROM) and a digital versatile disc (DVD)), a magneto-optical disk (including a mini disc (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be a ROM 1302, a hard disk included in the storage section 1308, etc., in which a program is stored and distributed to the user together with the device containing them.

[0123] It should be understood that the technical solution of the present disclosure can be implemented through the following exemplary embodiments.

[0124] 1. A method for processing multi-echo data, comprising:

[0125] For each pixel point in the depth map of the environment generated by a light detection and ranging device, obtaining a plurality of echo data;

[0126] Determining the probability that the corresponding echo data among the plurality of echo data corresponds to a valid echo or a noise echo, thereby forming a confidence map corresponding to the depth map of the environment; and

[0127] Generating a saliency map corresponding to the depth map of the environment based on the confidence map, where the saliency map is used to indicate the amount of valid information in the corresponding part of the depth map of the environment.

[0128] 2. The method according to clause 1, wherein determining the probability that the corresponding echo data corresponds to a valid echo or a noise echo includes using a neural network model to distinguish between valid echoes and noise echoes.

[0129] 3. The method according to clause 2, wherein using a neural network model to distinguish between valid echoes and noise echoes includes: for different pixel points in the depth map, using a unified neural network model to distinguish between valid echoes and noise echoes.

[0130] 4. The method according to clause 1, wherein the neural network model is trained as follows:

[0131] Obtaining a first set of data samples, where the first set of data samples includes a plurality of echo data for each pixel point in the depth maps generated by a plurality of light detection and ranging devices;

[0132] For each pixel, add a label of either a valid echo or a noise echo to the corresponding data sample using a pixel-specific noise probability model; and

[0133] Use the data samples in the first set and the corresponding labels as training data to train the neural network model.

[0134] 5. The method according to clause 4,

[0135] wherein the multiple light detection and ranging devices include real devices and simulation models, the data samples in the first set include real data and simulation data obtained by the real devices and simulation models for one or more environments, and

[0136] wherein the data samples in the first set are obtained during a first plurality of frames.

[0137] 6. The method according to clause 4, wherein the parameters of the pixel-specific noise probability model are determined as follows:

[0138] Obtain a data sample set, wherein the data sample set in the second set includes a plurality of echo data of a first pixel in a depth map generated by a plurality of light detection and ranging devices; and

[0139] Based on the plurality of echo data of the first pixel, determine the parameters of the first pixel-specific noise probability model.

[0140] 7. The method according to clause 6,

[0141] wherein the multiple light detection and ranging devices include real devices and simulation models, the data samples in the second set include real data and simulation data obtained by the real devices and simulation models for one or more environments, and

[0142] wherein the data samples in the second set are obtained during a second plurality of frames.

[0143] 8. The method according to clause 1, wherein determining the probability that the corresponding echo data corresponds to a valid echo or a noise echo includes using a pixel-specific noise probability model to distinguish between valid echoes and noise echoes.

[0144] 9. The method according to clause 1, further comprising:

[0145] Based on the saliency map, preferentially process the echo data of the part with high effective information content in the depth map of the environment; and / or

[0146] Output the saliency map in real time based on the user's region of interest.

[0147] 10. The method according to clause 1, wherein the neural network model is a convolutional neural network model, and / or the light detection and ranging device includes a lidar.

[0148] 11. A method for training a neural network model, comprising:

[0149] Obtaining a set of data samples, wherein the set of data samples includes multiple echo data of each pixel point in depth maps generated by multiple light detection and ranging devices;

[0150] For each pixel point, adding a label of a valid echo or a noise echo to the corresponding data sample respectively using a pixel-point specific noise probability model; and

[0151] Using the set of data samples and the corresponding labels as training data to train the neural network model.

[0152] 12. The method according to clause 11,

[0153] wherein the multiple light detection and ranging devices include real devices and simulation models, the set of data samples includes real data and simulation data obtained by the real devices and the simulation models for one or more environments, and

[0154] wherein the set of data samples is obtained during multiple frames.

[0155] 13. The method according to clause 11, wherein the neural network model is a convolutional neural network model, and / or the light detection and ranging device includes a lidar.

[0156] 14. A method for establishing a noise probability model, comprising:

[0157] Obtaining a set of data samples, wherein the set of data samples includes multiple echo data of a first pixel point in a depth map generated by multiple light detection and ranging devices; and

[0158] Based on the multiple echo data of the first pixel point, determining the parameters of a first pixel-point specific noise probability model.

[0159] 15. The method according to clause 14,

[0160] wherein the multiple light detection and ranging devices include real devices and simulation models, the set of data samples includes real data and simulation data obtained by the real devices and the simulation models for one or more environments, and

[0161] wherein the set of data samples is obtained during multiple frames.

[0162] 16. The method according to clause 14, wherein the noise probability model satisfies a Gaussian distribution, and the parameters include the mean and variance of the Gaussian distribution.

[0163] 17. A light detection and ranging device, comprising:

[0164] at least one processor; and

[0165] at least one memory, including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the light detection and ranging device to execute the method according to any one of clauses 1 - 10.

[0166] 18. An electronic device, comprising:

[0167] at least one processor; and

[0168] at least one memory, including computer program code, wherein the at least one memory and the computer program code are configured to, through the at least one processor, cause the electronic device to execute the method according to any one of clauses 11 - 13 or 14 - 16.

[0169] 19. A computer program product, comprising one or more instructions, which when executed by a processor, cause the implementation of the method according to any one of clauses 1 - 10, 11 - 13 or 14 - 16.

[0170] The exemplary embodiments of the present disclosure have been described above with reference to the accompanying drawings, but the present disclosure is of course not limited to the above examples. Those skilled in the art can obtain various changes and modifications within the scope of the appended claims, and it should be understood that these changes and modifications will naturally fall within the technical scope of the present disclosure.

[0171] For example, in the above embodiments, multiple functions included in one unit can be implemented by separate devices. Alternatively, multiple functions implemented by multiple units in the above embodiments can be respectively implemented by separate devices. Additionally, one of the above functions can be implemented by multiple units. Needless to say, such configurations are included in the technical scope of the present disclosure.

[0172] In this specification, the steps described in the flowcharts include not only the processes executed in the described order in a time series, but also processes executed in parallel or separately rather than necessarily in a time series. Moreover, even in the steps of processing in a time series, needless to say, the order can be appropriately changed.

[0173] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made without departing from the spirit and scope of the present disclosure as defined by the appended claims. Moreover, the term "comprising" in the embodiments of the present disclosure, "including" or any other variation thereof, is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

Claims

1. A method for processing multi - echo data, comprising: For each pixel point in the depth map of the environment generated by a light detection and ranging device, obtaining a plurality of echo data; Determining the probability that the corresponding echo data among the plurality of echo data corresponds to a valid echo or a noise echo, thereby forming a confidence map corresponding to the depth map of the environment; and Generating a saliency map corresponding to the depth map of the environment based on the confidence map, wherein the saliency map is used to indicate the amount of valid information in the corresponding part of the depth map of the environment.

2. The method according to claim 1, wherein, Determining the probability that the corresponding echo data corresponds to a valid echo or a noise echo includes using a neural network model to distinguish between valid echoes and noise echoes.

3. The method according to claim 2, wherein Using a neural network model to distinguish between valid echoes and noise echoes includes: for different pixel points in the depth map, using a unified neural network model to distinguish between valid echoes and noise echoes.

4. The method according to claim 1, wherein, The neural network model is trained as follows: Obtaining a first set of data samples, wherein the first set of data samples includes a plurality of echo data for each pixel point in the depth maps generated by a plurality of light detection and ranging devices; For each pixel point, adding a label of a valid echo or a noise echo to the corresponding data sample using a pixel - specific noise probability model; And Using the first set of data samples and the corresponding labels as training data to train the neural network model.

5. The method according to claim 4, Among them, The plurality of light detection and ranging devices include real devices and simulation models, the first set of data samples includes real data and simulation data obtained by the real devices and simulation models for one or more environments, and Wherein, the first set of data samples is obtained during a first plurality of frames.

6. The method according to claim 4, wherein, The parameters of the pixel - specific noise probability model are determined as follows: Obtaining a second set of data samples, wherein the second set of data samples includes a plurality of echo data for a first pixel point in the depth maps generated by a plurality of light detection and ranging devices; and Based on the plurality of echo data of the first pixel point, determining the parameters of the first pixel - specific noise probability model.

7. The method according to claim 6, Among them, The plurality of light detection and ranging devices include real devices and simulation models, the second set of data samples includes real data and simulation data obtained by the real devices and simulation models for one or more environments, and Wherein, the second set of data samples is obtained during a second plurality of frames.

8. The method according to claim 1, wherein Determining the probability that the corresponding echo data corresponds to a valid echo or a noise echo includes using a pixel - specific noise probability model to distinguish between valid echoes and noise echoes.

9. The method according to claim 1, further comprising: Based on the saliency map, preferentially processing the echo data of the part with high valid information amount in the depth map of the environment; And / or Real - time outputting the saliency map based on the user's region of interest.

10. The method according to claim 1, wherein, The neural network model is a convolutional neural network model, and / or the light detection and ranging device includes a lidar.