Indirect flight time expansion method and device, equipment and storage medium
By assigning orthogonal codes to multiple iTOF cameras and using optical diffraction elements to generate orthogonal light patterns, the problem of multiple iTOF cameras interfering with each other in the same scene is solved, and the depth estimation accuracy and 3D reconstruction quality are improved.
Patent Information
- Application Number
- CN202510674491.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-05
AI Technical Summary
When multiple iTOF cameras simultaneously emit modulated light signals in the same scene, they interfere with each other, affecting the accuracy of depth estimation.
Orthogonal coding is used to assign different spatial codes to each camera, and an orthogonal light pattern is generated through an optical diffraction element. The superimposed reflected light signals are received and decoded, the depth information of each camera is extracted, and a three-dimensional point cloud model is generated.
Ensure that the interference between cameras is minimized, achieve highly robust spatial separation of different signals, eliminate cross-interference caused by the superposition of multi-camera signals, and ensure the accuracy of depth information.
Smart Images

Figure CN120595255A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of laser radar technology, and in particular to an indirect time-of-flight expansion method, device, equipment and storage medium. Background Art
[0002] Color and depth (RGBD) cameras based on indirect time-of-flight (iTOF) technology can provide high-resolution real-time RGBD video streams, making them ideal tools for local 3D reconstruction. For rapid 3D reconstruction of large-scale scenes, multiple RGBD cameras are usually required to collect all-round information about the scene from different angles and fuse them to generate a complete 3D model. However, when the number of iTOF cameras deployed in a scene exceeds a certain threshold (for example, deploying 9 or more in the same scene), the modulated light signals emitted simultaneously by multiple iTOF cameras will cause severe interference, significantly reducing the accuracy of depth estimation, and causing crosstalk between devices, affecting their synchronous operation. The problem of mutual signal interference will significantly affect measurement performance.
[0003] The above content is only used to assist in understanding the technical solution of the present invention and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an indirect time-of-flight expansion method, device, equipment and storage medium, aiming to solve the technical problem in the prior art that when multiple iTOF cameras simultaneously emit modulated light signals in the same scene, they interfere with each other and affect the depth estimation accuracy.
[0005] To achieve the above objectives, the present application provides an indirect flight time expansion method, which includes:
[0006] Based on the coding matrix obtained by orthogonal coding, the corresponding spatial coding is assigned to each camera;
[0007] Based on the optical diffraction element, the spatial code is converted into a corresponding spatially coded light pattern and projected onto the current scene;
[0008] receiving a superimposed reflected light signal in the current scene, decoding the superimposed reflected light signal, and extracting depth information corresponding to each camera;
[0009] Based on the depth information, a three-dimensional point cloud model of the current scene is generated.
[0010] In one embodiment, the step of assigning corresponding spatial codes to each camera based on the coding matrix obtained by orthogonal coding includes:
[0011] Based on the coding matrix, determining row coding data of the coding matrix;
[0012] determining target row code data for each camera in the row code data based on the row number corresponding to the camera number;
[0013] The target row coded data is used as the spatial code of the corresponding camera.
[0014] In one embodiment, before the step of assigning a corresponding spatial code to each camera based on the coding matrix obtained by orthogonal coding, the step further includes:
[0015] Obtaining target sizes of an initial matrix and the encoding matrix;
[0016] Determining a target number of recursions based on the target size and a preset base number;
[0017] Based on the initial matrix and the first corresponding relationship between the initial matrix and the recursive matrix, determining the recursive matrix, and updating the current number of recursions;
[0018] When the current recursion number is less than the target recursion number, the initial matrix is updated to the recursion matrix, and the steps of determining the recursion matrix based on the initial matrix and the corresponding relationship between the initial matrix and the recursion matrix and updating the current recursion number are performed;
[0019] When the current recursion number is greater than or equal to the target recursion number, the recursion matrix is used as the encoding matrix.
[0020] In one embodiment, the step of converting the spatial code into a corresponding spatially coded light pattern based on an optical diffraction element and projecting it onto the current scene includes:
[0021] Calculating a phase distribution corresponding to the spatial encoding;
[0022] Based on the phase distribution corresponding to the spatial encoding, a corresponding optical diffraction element is assigned to each camera, where the optical diffraction element is deployed in front of a laser emitter or a light source of the corresponding camera;
[0023] When the camera emits a modulated light signal, the spatial code is converted into a corresponding spatially coded light pattern based on the optical diffraction element and projected onto the current scene.
[0024] In one embodiment, before the step of converting the spatial code into a corresponding spatially coded light pattern based on an optical diffraction element and projecting it into the current scene, the step further includes:
[0025] A corresponding time parameter is allocated to the camera, where the time parameter is any one of a signal transmission time slot and a signal transmission frequency, and different cameras have different time parameters.
[0026] In one embodiment, the step of receiving the superimposed reflected light signal in the current scene, decoding the superimposed reflected light signal, and extracting the depth information corresponding to each camera includes:
[0027] determining a second correspondence between the spatially coded light pattern, the superimposed reflected light signal, and the decoded signal based on an orthogonal decoding strategy;
[0028] Obtaining a decoded signal for each camera based on the spatially coded light pattern, the superimposed reflected light signal, and the second correspondence;
[0029] The decoded signal is used as the depth information of the corresponding camera.
[0030] In one embodiment, the step of generating a three-dimensional point cloud model of the current scene based on the depth information includes:
[0031] Based on the depth information, generating a corresponding depth map, and converting the depth map into point cloud data;
[0032] fusing the point cloud data to obtain fused point cloud data;
[0033] Aligning and filtering the fused point cloud data to obtain target point cloud data;
[0034] Based on the target point cloud data, a three-dimensional point cloud model of the current scene is constructed.
[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes an indirect flight time expansion device, which includes:
[0036] A coding allocation module is used to allocate corresponding spatial codes to each camera based on the coding matrix obtained by orthogonal coding;
[0037] A pattern generation module, configured to convert the spatial code into a corresponding spatially coded light pattern based on an optical diffraction element and project it onto the current scene;
[0038] an information decoding module, configured to receive the superimposed reflected light signals in the current scene, decode the superimposed reflected light signals, and extract depth information corresponding to each camera;
[0039] A three-dimensional reconstruction module is used to generate a three-dimensional point cloud model of the current scene based on the depth information.
[0040] In addition, to achieve the above-mentioned purpose, the present application also proposes an indirect flight time expansion device, which includes: a memory, a processor, and a computer program stored in the memory and runnable on the processor, and the computer program is configured to implement the steps of the indirect flight time expansion method as described above.
[0041] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the indirect flight time expansion method as described above are implemented.
[0042] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the indirect flight time expansion method as described above.
[0043] The present application provides an indirect time-of-flight expansion method. Based on the coding matrix obtained by orthogonal coding, a corresponding spatial coding is assigned to each camera; based on the optical diffraction element, the spatial coding is converted into a corresponding spatial coding light pattern and projected to the current scene; the superimposed reflected light signal in the current scene is received, the superimposed reflected light signal is decoded, and the depth information corresponding to each camera is extracted; based on the depth information, a three-dimensional point cloud model of the current scene is generated. The present application uses orthogonal spatial coding to assign different spatial codes to different cameras. The orthogonality of the spatial coding can ensure minimal interference between cameras, and uses optical diffraction elements to generate corresponding orthogonal light patterns to achieve high-robustness spatial separation of different signals. The zero cross-correlation of orthogonal coding is used during decoding to eliminate the cross-interference caused by the superposition of multiple camera signals, thereby ensuring the accuracy of depth information. This solves the technical problem that when multiple iTOF cameras simultaneously transmit modulated light signals in the same scene, they interfere with each other and affect the accuracy of depth estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0046] Figure 1 This is a flow chart of Example 1 of the indirect flight time expansion method of the present application;
[0047] Figure 2 A schematic diagram of the process of designing an optical diffraction element for the indirect time-of-flight expansion method provided in Example 2 of the present application;
[0048] Figure 3 A schematic diagram of the signal acquisition and decoding process of the indirect time-of-flight expansion method provided in Example 2 of the present application;
[0049] Figure 4 This is a flow chart of Example 2 of the indirect flight time expansion method of this application;
[0050] Figure 5 A schematic diagram of a simplified flow chart of the indirect flight time expansion method provided in Example 2 of the present application;
[0051] Figure 6 This is a schematic diagram of the module structure of the indirect flight time expansion device according to an embodiment of the present application;
[0052] Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the indirect flight time expansion method in the embodiment of the present application.
[0053] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0054] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0055] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0056] The main solutions of the embodiments of the present application are: based on the coding matrix obtained by orthogonal coding, assign corresponding spatial codes to each camera; based on the optical diffraction element, convert the spatial codes into corresponding spatially coded light patterns and project them into the current scene; receive the superimposed reflected light signals in the current scene, decode the superimposed reflected light signals, and extract the depth information corresponding to each camera; based on the depth information, generate a three-dimensional point cloud model of the current scene.
[0057] Currently, when the number of iTOF cameras deployed in a scene exceeds a certain threshold, the modulated light signals emitted simultaneously by multiple iTOF cameras will cause serious interference, significantly reducing the accuracy of depth estimation, and causing crosstalk between devices, affecting their synchronous operation. The problem of mutual signal interference will significantly affect measurement performance.
[0058] This application provides a solution that uses orthogonal spatial coding to assign different spatial codes to different cameras. The orthogonality of the spatial coding can ensure minimal interference between cameras, and uses optical diffraction elements to generate corresponding orthogonal light patterns to achieve highly robust spatial separation of different signals. During decoding, the zero cross-correlation of orthogonal coding is used to eliminate cross-interference caused by the superposition of multi-camera signals, thereby ensuring the accuracy of depth information. This solves the technical problem that when multiple iTOF cameras simultaneously emit modulated light signals in the same scene, they interfere with each other, affecting the accuracy of depth estimation.
[0059] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of performing the above functions, an indirect time-of-flight expansion device, etc., and this embodiment does not specifically limit this. The following uses the indirect time-of-flight expansion device as an example to illustrate this embodiment and the following embodiments.
[0060] The present application embodiment provides an indirect flight time expansion method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the indirect flight time expansion method of the present application.
[0061] In this embodiment, the indirect flight time expansion method includes steps S10 to S40:
[0062] Step S10, assigning corresponding spatial codes to each camera based on the coding matrix obtained by orthogonal coding;
[0063] It should be noted that the camera in this embodiment refers to an iTOF camera. When the iTOF camera is working, it usually calculates the depth information of the scene by emitting modulated light signals. The signal emitted by each camera estimates the flight time by measuring the phase change of light. When multiple cameras emit modulated light signals simultaneously in the same scene, these signals will interfere with each other, resulting in a decrease in the accuracy of depth estimation. This signal interference comes from the fact that multiple cameras emit similar modulated light signals, which are not clearly distinguished, making it difficult for the receiving end to distinguish signals from different cameras. Therefore, this embodiment uses orthogonal coding to assign different spatial codes to different cameras. The cameras emit corresponding light patterns according to their respective spatial codes. These light patterns do not overlap with each other and have orthogonal properties. Since these patterns are orthogonal, even if these signals overlap in space, they can still be distinguished, and the receiving end can accurately separate the information of each camera from the overlapping signals through a decoding algorithm, thereby avoiding signal interference and ensuring the accuracy of depth information.
[0064] In addition, it should be noted that the orthogonal coding used in this embodiment is Walsh-Hadamard coding, which is not specifically limited. Walsh-Hadamard coding generates orthogonal coding patterns by using Hadamard matrices (or Walsh matrices). This coding has orthogonality and is suitable for multiplexing and signal decoding applications. When deploying multiple cameras, the Walsh-Hadamard coding matrix can be used to generate spatial coding for each camera to ensure that the signals between different cameras do not interfere with each other.
[0065] In a feasible embodiment, before step S10, the steps include: obtaining target sizes of the initial matrix and the encoding matrix; determining a target number of recursions based on the target sizes and a preset base; determining a recursion matrix based on the initial matrix and a first correspondence between the initial matrix and the recursion matrix, and updating a current number of recursions; when the current number of recursions is less than the target number of recursions, updating the initial matrix to the recursion matrix, executing the steps of determining a recursion matrix based on the initial matrix and the correspondence between the initial matrix and the recursion matrix, and updating the current number of recursions; and when the current number of recursions is greater than or equal to the target number of recursions, using the recursion matrix as the encoding matrix.
[0066] It should be noted that the recursive matrix is the Hadamard matrix generated by recursive construction. Assuming that there is a 2 k ×2 k The initial matrix of , then a 2 can be generated by recursive construction k+1 ×2 k+1 The recursive matrix H k+1 The first correspondence between the initial matrix and the recursive matrix is the calculation relationship of the recursive matrix, as shown below:
[0067]
[0068] Where H k represents the initial matrix, H k+1 Represents the recursive matrix. The encoding matrix is the final matrix obtained by Walsh-Hadamard encoding. The target size of the encoding matrix is the size of the set encoding matrix. Assuming that the target size of the encoding matrix is N×N, since N=2 k , N must be a power of 2. After completing a recursive construction, the value of k increases by 1. The preset base is 2. According to N=2 k The value of k can be calculated to determine the number of times the recursive construction needs to be performed, that is, the target number of recursions. For example, the target size of the encoding matrix is 4×4, 4=2 2 , the target recursion number is 2. Generally speaking, the target size of the encoding matrix can be determined according to the number of deployed cameras. For example, if the number of cameras is N = 16, the target size is 16×16, so that spatial encoding can be allocated to the cameras later.
[0069] It can be understood that the current number of recursions is the number of times the recursive construction has been performed. Substituting the initial matrix into the above-mentioned first corresponding relationship, the corresponding recursive matrix is obtained, and the current number of recursions is increased by 1. If the current number of recursions is less than the target number of recursions, it means that the recursive construction needs to be continued. The initial matrix is updated to the recursive matrix, and the recursive construction is performed again to obtain a new recursive matrix until the current number of recursions is greater than or equal to the target number of recursions. The recursive matrix generated at this time is the final required encoding matrix. For example, the initial matrix is H1=[1], the target size of the encoding matrix is 4×4, the target number of recursions is 2, and the first recursive construction is performed, and the recursive matrix obtained is:
[0070]
[0071] Then the second recursive construction is performed, and the resulting recursive matrix is:
[0072]
[0073] Thus, the encoding matrix H4 is obtained.
[0074] It should be understood that encoding each row of the encoding matrix is regarded as spatial encoding of a different camera.
[0075] In a feasible implementation, step S10 includes: determining the row coding data of the coding matrix based on the coding matrix; determining the target row coding data of each camera in the row coding data based on the row number corresponding to the camera number; and using the target row coding data as the spatial coding of the corresponding camera.
[0076] It should be noted that the spatial code for each camera is obtained from the encoding matrix, where each row represents a unique code. The row code data is the code for each row of the encoding matrix, and the camera number is the camera number. By finding the corresponding row number based on the camera number, we can determine the target row code data for each camera, and thus the spatial code for each camera.
[0077] It can be understood that in this embodiment, the first row of codes in the coding matrix is assigned as the spatial code to the first camera, the second row of codes in the coding matrix is assigned as the spatial code to the second camera, and so on, the Nth row of codes in the coding matrix is assigned as the spatial code to the Nth camera. For example, for the above coding matrix H4, the spatial code of the first camera is [1, 1, 1, 1], the spatial code of the second camera is [1, -1, 1, -1], the spatial code of the third camera is [1, 1, -1, -1], and the spatial code of the fourth camera is [1, -1, -1, 1].
[0078] It should be understood that each row of the encoding matrix is orthogonal, that is, the inner product between each row in the matrix and every other row is zero, which prevents interference during signal processing and can be used to construct orthogonal spatial coding. This embodiment adopts an orthogonal coding design based on the Walsh-Hadamard matrix. An N×N orthogonal matrix is generated through recursive construction, and each row of the code is used as the spatial code of an independent camera to ensure the mathematical orthogonality of the multi-camera signals. At the same time, each row of the orthogonal matrix is uniquely bound to a different camera, achieving strict code differentiation and anti-overlapping interference capabilities.
[0079] Step S20: Based on the optical diffraction element, convert the spatial code into a corresponding spatially coded light pattern and project it onto the current scene;
[0080] It should be noted that, in order to convert the spatial coding into an actual optical pattern, this embodiment designs an optical diffraction element (DOE), which can generate a corresponding optical signal, namely, a spatially coded light pattern, according to the spatial coding of each camera.
[0081] In a feasible embodiment, step S20 may include: calculating the phase distribution corresponding to the spatial coding; based on the phase distribution corresponding to the spatial coding, assigning a corresponding optical diffraction element to each camera, and the optical diffraction element is deployed in front of the laser emitter or light source of the corresponding camera; when the camera emits a modulated light signal, based on the optical diffraction element, converting the spatial coding into a corresponding spatially coded light pattern and projecting it onto the current scene.
[0082] It should be noted that for each spatial code, the corresponding phase distribution can be calculated by an optimization algorithm (such as the Gerchberg-Saxton algorithm). The phase distribution determines the way in which the light wave changes its propagation path when passing through the optical diffraction element. The use of the optimization algorithm can ensure that each spatially coded light pattern can be correctly projected in the current scene. The current scene is usually a large three-dimensional scene, and there is no specific limitation on this. For example, assuming that the spatial code of camera A is [1,-1,1,-1], the phase distribution is designed according to [1,-1,1,-1] so that the light wave passing through the camera forms alternating light and dark stripes in the current scene.
[0083] It is understandable that the optimized phase distribution can be used to manufacture optical diffraction elements. Generally speaking, electron beam lithography or photolithography technology can be used to convert phase information (phase distribution) into a physical pattern to accurately control the phase of the light wave of each camera to ensure that its corresponding spatially encoded light pattern can be accurately projected into the current scene. According to the phase distribution corresponding to the spatial encoding of the camera, a corresponding optical diffraction element is assigned to each camera. Since all cameras use different spatial encodings, it can be seen that the optical diffraction elements required for different cameras are different. The optical diffraction element of each camera will be installed in front of the laser emitter or the light source to ensure that the light wave can project a specific optical pattern according to the predetermined phase distribution when passing through the optical diffraction element. Thus, each camera can generate its own unique, orthogonal optical pattern and perform depth acquisition in the scene.
[0084] In the specific implementation, refer to Figure 2 The steps of designing an optical diffraction element may include: inputting spatial encoding; designing phase distribution; manufacturing the optical diffraction element; and integrating it into a camera.
[0085] It should be understood that N iTOF cameras are mounted at different locations within the scene to ensure that their field of view covers the entire target area. A DOE is mounted in front of the iTOF camera's laser emitter or light source to ensure that the spatially coded light pattern is correctly projected into the scene. Each camera is assigned a unique pattern, ensuring orthogonality of the optical signals from each camera. Each camera's spatially coded light pattern is projected into the scene without interference from signals from other cameras. The physical integration of the DOE and the iTOF camera enables efficient projection and synchronous acquisition of the coded light field.
[0086] Step S30, receiving the superimposed reflected light signal in the current scene, decoding the superimposed reflected light signal, and extracting the depth information corresponding to each camera;
[0087] It should be noted that each camera emits a light signal into the current scene based on its unique spatially coded light pattern. Because each camera emits a different signal, the signals from multiple cameras will overlap in space. In this embodiment, the sensor receives the reflected light signals from the scene and records the superposition of all reflected light signals, i.e., the superimposed reflected light signal.
[0088] It's understandable that although the camera signals are transmitted via different spatially coded light patterns, they are mixed together at the same time, forming a composite signal containing multiple camera signals (superimposed with the reflected light signal). This composite signal contains the acquired data from all cameras. Decoding can be used to extract the independent signal from each camera and obtain the corresponding depth information. Once the independent depth information from each camera is obtained, this depth information can be used for 3D reconstruction and fusion to generate a complete 3D model.
[0089] In the specific implementation, refer to Figure 3 The steps of signal acquisition and decoding include: signal acquisition; signal superposition; signal decoding; information extraction; and result output.
[0090] It should be understood that by performing an inner product operation on the received signal and the coding pattern, the independent depth information of each camera can be separated from the composite signal, while the zero cross-correlation of orthogonal coding can be utilized to eliminate cross-interference caused by the superposition of multiple camera signals.
[0091] Step S40: generating a three-dimensional point cloud model of the current scene based on the depth information.
[0092] In a feasible implementation, step S40 may include: generating a corresponding depth map based on the depth information, and converting the depth map into point cloud data; fusing the point cloud data to obtain fused point cloud data; aligning and filtering the fused point cloud data to obtain target point cloud data; and constructing a three-dimensional point cloud model of the current scene based on the target point cloud data.
[0093] It should be noted that a depth map is the distance data from each pixel to the camera, representing the depth information of each point in the scene. Based on the depth information of each camera, a corresponding depth map is generated. Each camera's depth map data is converted into a point cloud, generating the corresponding point cloud data. Point cloud data consists of a large number of spatial points, each containing spatial coordinates and depth information. Each camera's point cloud data represents the three-dimensional structure of the scene from its perspective.
[0094] As you can understand, the point cloud data requires further processing to generate the final 3D model, or 3D point cloud model. The point cloud data from different cameras is fused, resulting in the fused data known as the fused point cloud data. Because each camera captures information from a different perspective, the point cloud data must be aligned to construct a complete 3D model. The alignment algorithm, which can use the Iterative Closest Point (ICP) algorithm, precisely aligns the point cloud data from different cameras, ensuring they fit together correctly in 3D space. During the point cloud fusion process, some noise or redundant points may appear. These can be removed using filtering algorithms, such as polarization filtering or statistical filtering (not specifically defined), to remove artifacts and noise caused by multipath or sensor errors. Finally, the point cloud quality is optimized to remove unwanted noise and smooth the point cloud, thereby improving the accuracy of 3D reconstruction. After alignment and filtering, the fused point cloud data yields the data required for model construction, known as the target point cloud data.
[0095] It should be understood that constructing a complete 3D point cloud model using target point cloud data can be used for visualization analysis or in applications such as virtual reality and augmented reality. The 3D point cloud model can display the complete structure of the scene, providing a more accurate and detailed 3D view. This embodiment significantly improves 3D reconstruction quality by reducing inter-camera interference and noise.
[0096] Furthermore, a corresponding time parameter is allocated to the camera, where the time parameter is any one of a signal transmission time slot and a signal transmission frequency, and different cameras have different time parameters.
[0097] It can be understood that different time slots or frequencies are allocated to each camera to achieve time multiplexing.
[0098] In addition, the theoretical upper limit of the number of cameras that can be deployed in this embodiment is: the number of orthogonal codes × the number of delayed exposure multiplexers, where the value of the delayed exposure multiplexer is 9. It can be seen that the number of iTOF cameras that can be supported by this embodiment is significantly higher than the traditional solution.
[0099] This embodiment provides an indirect time-of-flight expansion method. Based on a coding matrix derived from orthogonal coding, a corresponding spatial code is assigned to each camera. Using an optical diffraction element, the spatial code is converted into a corresponding spatially coded light pattern and projected onto the current scene. The superimposed reflected light signals in the current scene are received, decoded, and the depth information corresponding to each camera is extracted. Based on this depth information, a three-dimensional point cloud model of the current scene is generated. This embodiment utilizes orthogonal spatial coding to assign different spatial codes to different cameras. The orthogonality of the spatial codes minimizes interference between cameras. Optical diffraction elements are used to generate corresponding orthogonal light patterns, achieving highly robust spatial separation of different signals. During decoding, the zero cross-correlation of the orthogonal coding is utilized to eliminate cross-interference caused by the superposition of signals from multiple cameras, ensuring the accuracy of depth information.
[0100] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 4 , step S30 may include steps S301 to S303:
[0101] Step S301, determining a second correspondence between the spatially coded light pattern, the superimposed reflected light signal, and the decoded signal based on an orthogonal decoding strategy;
[0102] It should be noted that the received superimposed reflected light signal I(x,y) is the result of the superposition of spatially coded light patterns from multiple cameras. For example, assuming that four cameras transmit signals simultaneously, their signals superimpose in space. The received signal I(x,y) will be the weighted sum of the spatially coded light patterns emitted by each camera and the corresponding reflected light signals in the scene, as shown below:
[0103] I(x,y)=Φ1(x,y)·S1(x,y)+Φ2(x,y)·S2(x,y)+Φ3(x,y)·S3(x,y)
[0104] +Φ4(x,y)·S4(x,y)
[0105] Where I(z,y) represents the superimposed reflected light signal, S1(x,y), S2(x,y), S3(x,y), and S4(x,y) represent the reflections of the light signals emitted by each camera in the scene, that is, the reflected light signals of each camera are directly related to the depth information of the scene, and Φ1(x,y), Φ2(x,y), Φ3(x,y), and Φ4(x,y) represent the spatially coded light patterns of each camera.
[0106] It is understood that in order to extract the signal of each camera from the superimposed reflected light signal, the orthogonality of the spatially coded pattern can be exploited for decoding, i.e., an orthogonal decoding strategy can be adopted. The second correspondence between the spatially coded light pattern, the superimposed reflected light signal, and the decoded signal is the calculation relationship used for decoding, as shown below:
[0107] w i =∫I(x,y)·Φ i (x,y)dxdy
[0108] Where w i represents the decoded signal of the i-th camera, Φ i (x,y) represents the spatially coded light pattern of the i-th camera, and I(x,y) represents the superimposed reflected light signal.
[0109] Step S302: obtaining a decoding signal of each camera based on the spatially coded light pattern, the superimposed reflected light signal, and the second corresponding relationship;
[0110] It should be noted that since the spatially coded light patterns of different cameras are orthogonal, the decoding process separates the signal of each camera from the composite signal by calculating the inner product to obtain the decoded signal of each camera.
[0111] Step S303: Using the decoded signal as depth information of the corresponding camera.
[0112] It can be understood that the decoded signal of each camera obtained by decoding is the depth information of the camera, so that the depth information of each camera can be extracted from the superimposed reflected light signal. The depth information reflects the changes in the viewing angle and light signal reflection of each camera in the scene.
[0113] This embodiment provides an indirect time-of-flight expansion method. Based on an orthogonal decoding strategy, a second correspondence is determined between the spatially coded light pattern, the superimposed reflected light signal, and the decoded signal. Based on the spatially coded light pattern, the superimposed reflected light signal, and the second correspondence, a decoded signal for each camera is obtained. The decoded signal is used as the depth information for the corresponding camera. This embodiment utilizes orthogonal spatial coding to assign different spatial codes to different cameras. The orthogonality of the spatial coding minimizes interference between cameras. Optical diffraction elements are used to generate corresponding orthogonal light patterns, achieving highly robust spatial separation of different signals. During decoding, the zero cross-correlation of the orthogonal coding is utilized to eliminate cross-interference caused by the superposition of signals from multiple cameras, ensuring the accuracy of depth information.
[0114] For example, in order to help understand the implementation process of the indirect flight time expansion method obtained by combining this embodiment with the above-mentioned embodiment 2, please refer to Figure 5 , Figure 5A brief flowchart of an indirect time-of-flight expansion method is provided, specifically:
[0115] Step 1: Orthogonal spatial coding design. Input the number of cameras to be deployed (N), the size of the coding area (e.g., 16×128×128), and the coding resolution (e.g., 128×128). Generate a set of spatial codes using an orthogonal coding method. Normalize all codes to ensure orthogonality. Map the codes to continuous or binary values that can be generated using optical diffraction elements.
[0116] Step 2: Fabricate the optical diffraction element. Use an optimization algorithm to generate the corresponding DOE phase distribution to project the corresponding spatial encoding. Fabricate using electron beam lithography, photolithography, or 3D printing. Mount the DOE in front of the laser emitter or light source of the iTOF camera.
[0117] Step 3: iTOF camera deployment. Each camera group is assigned a unique DOE and its corresponding spatial code. Different DOEs are assigned to different camera groups.
[0118] Step 4: Signal Acquisition and Decoding. Each DOE projects its own unique coded light pattern into the scene, and the iTOF sensor receives the superimposed signals from all cameras. The received signals are decoded using the orthogonality of spatial encoding to extract the depth information corresponding to each camera.
[0119] Step 5: 3D scene processing and reconstruction. Each camera independently generates its own depth map. The depth maps from all cameras are fused into a global point cloud of the scene. A point cloud alignment algorithm is used to optimize the reconstruction results. A filtering algorithm is used to remove artifacts caused by multipath effects.
[0120] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the indirect flight time expansion method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0121] This application also provides an indirect flight time expansion device, please refer to Figure 6 , the indirect flight time expansion device includes:
[0122] A coding allocation module 10 is configured to allocate corresponding spatial coding to each camera based on a coding matrix obtained by orthogonal coding;
[0123] The pattern generation module 20 is configured to convert the spatial code into a corresponding spatially coded light pattern based on an optical diffraction element and project it onto the current scene;
[0124] An information decoding module 30 is configured to receive the superimposed reflected light signals in the current scene, decode the superimposed reflected light signals, and extract depth information corresponding to each camera;
[0125] The three-dimensional reconstruction module 40 is configured to generate a three-dimensional point cloud model of the current scene based on the depth information.
[0126] In a feasible implementation manner, the coding allocation module 10 is further configured to determine row coding data of the coding matrix based on the coding matrix;
[0127] determining target row code data for each camera in the row code data based on the row number corresponding to the camera number;
[0128] The target row coded data is used as the spatial code of the corresponding camera.
[0129] In a feasible implementation manner, the coding allocation module 10 is further configured to obtain target sizes of the initial matrix and the coding matrix;
[0130] Determining a target number of recursions based on the target size and a preset base number;
[0131] Based on the initial matrix and the first corresponding relationship between the initial matrix and the recursive matrix, determining the recursive matrix, and updating the current number of recursions;
[0132] When the current recursion number is less than the target recursion number, the initial matrix is updated to the recursion matrix, and the steps of determining the recursion matrix based on the initial matrix and the corresponding relationship between the initial matrix and the recursion matrix and updating the current recursion number are performed;
[0133] When the current recursion number is greater than or equal to the target recursion number, the recursion matrix is used as the encoding matrix.
[0134] In a feasible implementation manner, the pattern generation module 20 is further configured to calculate the phase distribution corresponding to the spatial encoding;
[0135] Based on the phase distribution corresponding to the spatial encoding, a corresponding optical diffraction element is assigned to each camera, where the optical diffraction element is deployed in front of a laser emitter or a light source of the corresponding camera;
[0136] When the camera emits a modulated light signal, the spatial code is converted into a corresponding spatially coded light pattern based on the optical diffraction element and projected onto the current scene.
[0137] In a feasible implementation manner, the pattern generation module 20 is further configured to allocate corresponding time parameters to the cameras, where the time parameters are any one of a signal transmission time slot and a signal transmission frequency, and different cameras have different time parameters.
[0138] In a feasible embodiment, the information decoding module 30 is further configured to determine a second correspondence between the spatially coded light pattern, the superimposed reflected light signal, and the decoded signal based on an orthogonal decoding strategy;
[0139] Obtaining a decoded signal for each camera based on the spatially coded light pattern, the superimposed reflected light signal, and the second correspondence;
[0140] The decoded signal is used as the depth information of the corresponding camera.
[0141] In a feasible implementation manner, the 3D reconstruction module 40 is further configured to generate a corresponding depth map based on the depth information, and convert the depth map into point cloud data;
[0142] fusing the point cloud data to obtain fused point cloud data;
[0143] Aligning and filtering the fused point cloud data to obtain target point cloud data;
[0144] Based on the target point cloud data, a three-dimensional point cloud model of the current scene is constructed.
[0145] The indirect time-of-flight expansion device provided by this application adopts the indirect time-of-flight expansion method of the above-mentioned embodiment, which can solve the technical problem that when multiple iTOF cameras simultaneously transmit modulated light signals in the same scene, they interfere with each other and affect the accuracy of depth estimation. Compared with the prior art, the beneficial effects of the indirect time-of-flight expansion device provided by this application are the same as the beneficial effects of the indirect time-of-flight expansion method provided by the above-mentioned embodiment, and the other technical features of the indirect time-of-flight expansion device are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.
[0146] The present application provides an indirect flight time expansion device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the indirect flight time expansion method in the above-mentioned embodiment one.
[0147] Reference below Figure 7, which shows a schematic diagram of the structure of an indirect time-of-flight expansion device suitable for implementing the embodiments of the present application. The indirect time-of-flight expansion device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The indirect time-of-flight expansion device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0148] like Figure 7 As shown, the indirect flight time expansion device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a ROM (Read Only Memory) 1002 or a program loaded from a storage device 1003 into a RAM (Random Access Memory) 1004. Various programs and data required for the operation of the indirect flight time expansion device are also stored in the RAM 1004. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. The communication devices 1009 can allow the indirect time-of-flight expansion device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an indirect time-of-flight expansion device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have instead.
[0149] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0150] The indirect time-of-flight expansion device provided by this application, which adopts the indirect time-of-flight expansion method of the above-mentioned embodiment, can solve the technical problem that when multiple iTOF cameras simultaneously transmit modulated light signals in the same scene, they interfere with each other and affect the accuracy of depth estimation. Compared with the prior art, the beneficial effects of the indirect time-of-flight expansion device provided by this application are the same as the beneficial effects of the indirect time-of-flight expansion method provided by the above-mentioned embodiment, and the other technical features of the indirect time-of-flight expansion device are the same as the features disclosed in the method of the previous embodiment, and are not further described here.
[0151] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0152] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0153] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the indirect flight time expansion method in the above-mentioned embodiment.
[0154] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0155] The computer-readable storage medium may be included in the indirect time-of-flight expansion device, or may exist independently without being assembled into the indirect time-of-flight expansion device.
[0156] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the indirect time-of-flight expansion device, the indirect time-of-flight expansion device: assigns corresponding spatial codes to each camera based on the coding matrix obtained by orthogonal coding; converts the spatial codes into corresponding spatially coded light patterns based on the optical diffraction element and projects them into the current scene; receives the superimposed reflected light signals in the current scene, decodes the superimposed reflected light signals, and extracts the depth information corresponding to each camera; and generates a three-dimensional point cloud model of the current scene based on the depth information.
[0157] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0158] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0159] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0160] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned indirect time-of-flight expansion method. It can solve the technical problem that when multiple iTOF cameras simultaneously emit modulated light signals in the same scene, they interfere with each other and affect the accuracy of depth estimation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the indirect time-of-flight expansion method provided in the above-mentioned embodiment, and will not be repeated here.
[0161] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned indirect flight time expansion method when executed by a processor.
[0162] The computer program product provided in this application can address the technical problem of interference between multiple iTOF cameras simultaneously transmitting modulated light signals in the same scene, affecting depth estimation accuracy. Compared to the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the indirect time-of-flight expansion method provided in the above-mentioned embodiments, and are not further elaborated here.
[0163] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An indirect flight time expansion method, characterized in that: The method comprises: Based on the coding matrix obtained by orthogonal coding, the corresponding spatial coding is assigned to each camera; Based on the optical diffraction element, the spatial code is converted into a corresponding spatially coded light pattern and projected onto the current scene; receiving a superimposed reflected light signal in the current scene, decoding the superimposed reflected light signal, and extracting depth information corresponding to each camera; Based on the depth information, a three-dimensional point cloud model of the current scene is generated.
2. The method according to claim 1, wherein The step of assigning corresponding spatial coding to each camera based on the coding matrix obtained by orthogonal coding includes: Based on the coding matrix, determining row coding data of the coding matrix; determining target row code data for each camera in the row code data based on the row number corresponding to the camera number; The target row coded data is used as the spatial code of the corresponding camera.
3. The method according to claim 1, wherein Before the step of assigning corresponding spatial coding to each camera based on the coding matrix obtained by orthogonal coding, the following step is further included: Obtaining target sizes of an initial matrix and the encoding matrix; Determining a target number of recursions based on the target size and a preset base number; Based on the initial matrix and the first corresponding relationship between the initial matrix and the recursive matrix, determining the recursive matrix, and updating the current number of recursions; When the current recursion number is less than the target recursion number, the initial matrix is updated to the recursion matrix, and the steps of determining the recursion matrix based on the initial matrix and the corresponding relationship between the initial matrix and the recursion matrix and updating the current recursion number are performed; When the current recursion number is greater than or equal to the target recursion number, the recursion matrix is used as the encoding matrix.
4. The method according to claim 1, wherein The step of converting the spatial code into a corresponding spatially coded light pattern based on the optical diffraction element and projecting it onto the current scene includes: Calculating a phase distribution corresponding to the spatial encoding; Based on the phase distribution corresponding to the spatial encoding, a corresponding optical diffraction element is assigned to each camera, where the optical diffraction element is deployed in front of a laser emitter or a light source of the corresponding camera; When the camera emits a modulated light signal, the spatial code is converted into a corresponding spatially coded light pattern based on the optical diffraction element and projected onto the current scene.
5. The method according to claim 1, wherein Before the step of converting the spatial code into a corresponding spatially coded light pattern based on the optical diffraction element and projecting it onto the current scene, the method further includes: A corresponding time parameter is allocated to the camera, where the time parameter is any one of a signal transmission time slot and a signal transmission frequency, and different cameras have different time parameters.
6. The method according to claim 1, wherein The steps of receiving the superimposed reflected light signal in the current scene, decoding the superimposed reflected light signal, and extracting the depth information corresponding to each camera include: determining a second correspondence between the spatially coded light pattern, the superimposed reflected light signal, and the decoded signal based on an orthogonal decoding strategy; Obtaining a decoded signal for each camera based on the spatially coded light pattern, the superimposed reflected light signal, and the second correspondence; The decoded signal is used as the depth information of the corresponding camera.
7. The method according to any one of claims 1 to 6, characterized in that The step of generating a three-dimensional point cloud model of the current scene based on the depth information includes: Based on the depth information, generating a corresponding depth map, and converting the depth map into point cloud data; fusing the point cloud data to obtain fused point cloud data; Aligning and filtering the fused point cloud data to obtain target point cloud data; Based on the target point cloud data, a three-dimensional point cloud model of the current scene is constructed.
8. An indirect flight time expansion device, characterized in that: The device comprises: A coding allocation module is used to allocate corresponding spatial codes to each camera based on the coding matrix obtained by orthogonal coding; A pattern generation module, configured to convert the spatial code into a corresponding spatially coded light pattern based on an optical diffraction element and project it onto the current scene; an information decoding module, configured to receive the superimposed reflected light signals in the current scene, decode the superimposed reflected light signals, and extract depth information corresponding to each camera; A three-dimensional reconstruction module is used to generate a three-dimensional point cloud model of the current scene based on the depth information.
9. An indirect time-of-flight expansion device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the indirect flight time expansion method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the indirect flight time expansion method according to any one of claims 1 to 7 are implemented.