A distributed real-time 3D spatial reconstruction system and method

The distributed real-time 3D spatial reconstruction system utilizes multiple 3D cameras and edge computing devices to acquire and align color depth information in real time, generate point cloud data, and render 3D models. This solves the field of view and synchronization problems in large-scale scene reconstruction and achieves efficient 3D reconstruction and data transmission.

CN119172515BActive Publication Date: 2025-10-31JITUO TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411179751.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2025-10-31
Estimated Expiration
2044-08-27

AI Technical Summary

Technical Problem

Existing 3D spatial reconstruction technologies suffer from problems such as limited camera field of view, difficulty in simultaneous acquisition of data from multiple cameras and data stitching, and low network transmission efficiency when reconstructing large-scale scenes or large-sized objects, making it difficult to achieve efficient real-time 3D reconstruction.

Method used

A distributed real-time 3D spatial reconstruction system is adopted, which combines multiple 3D cameras with edge computing devices to collect and align color depth information in real time, generate point cloud data and render 3D models using network transmission, and supports multi-camera spatial positioning and data stitching with a wide range of non-overlapping fields of view.

Benefits of technology

It enables real-time acquisition and reconstruction of large-scale 3D scenes, improves data transmission efficiency and camera installation flexibility, supports time synchronization of discrete camera data and 3D spatial stitching, and breaks free from the scene and distance limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119172515B_ABST
    Figure CN119172515B_ABST
Patent Text Reader

Abstract

This invention discloses a distributed real-time 3D spatial reconstruction system and method, capable of real-time acquisition and reconstruction of large-scale 3D scenes, emphasizing real-time performance and large scale, which is the main difference from traditional 3D reconstruction methods. It transmits encoded and compressed 3D data, including spatial location data and color image data, via real-time high frame rate streaming over a network. Traditional methods only transmit color data or only transmit spatial location data acquired in a single frame. Camera installation is flexible, allowing cameras to be installed in areas where data needs to be acquired, and the calibration method supports rapid deployment. Data acquired by discrete cameras is combined in real-time, undergoing time synchronization and 3D spatial stitching to form a large-scale, real-time changing 3D scene data, whereas traditional cameras acquire independent 2D images without spatial stitching. The 3D scene data supports real-time 3D interaction and display, allowing translation, rotation, and scaling according to user commands. Data processed and aligned in real-time in 3D space is still recorded, stored, and played back in 3D format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and in particular to a distributed real-time 3D spatial reconstruction system and method. Background Technology

[0002] In existing technologies, one typical method for 3D spatial reconstruction is to solve for extreme constraints from multiple color images, such as the scheme in Reference 1, which can calculate a 3D model from 2D image information. However, this method of obtaining a 3D model through inference rather than directly through 3D spatial perception is often inefficient and lacks sufficient modeling accuracy. A 3D camera, on the other hand, can directly acquire spatial depth information. Using a color sensor and depth sensor calibration method similar to that in Reference 2, the color image and depth image are spatially aligned to quickly reconstruct a color 3D model within the perceptible area.

[0003] However, the inventors discovered that during real-time 3D spatial reconstruction, due to the limited field of view of a single camera, multiple 3D cameras are often needed simultaneously to expand the camera coverage for large objects and large-scale scenes. Moreover, since the camera can only reconstruct the front surface of objects within its field of view, the back and sides of objects are inevitably missing, requiring multiple cameras to reconstruct them separately in real time from different angles before stitching them together.

[0004] According to the search, Reference 3 discloses a 3D panoramic shooting system that sets up multiple 3D image acquisition units around the shooting area to acquire color images containing color image information and depth images containing depth image information respectively. Based on the overlapping area of ​​adjacent acquisition units, the data acquired by multiple 3D image acquisition units are stitched together to generate a 3D panoramic image.

[0005] The inventors discovered the following shortcomings in this solution:

[0006] 1. All acquisition units are required to be equidistant from the center point, which is not applicable when the observed scene is further expanded or when multiple objects need to be observed simultaneously.

[0007] 2. This technology uses real-time camera positioning and data stitching based on image processing and feature point matching, which requires adjacent cameras to have overlapping fields of view. However, under conditions such as terrain limitations and scattered scenes, such camera layout conditions may not be met, rendering the method ineffective.

[0008] 3. This technology uses a local processing architecture that first stitches the data locally and then transmits it. However, when the distance between cameras is far or the distance between the camera and the local server is far, the direct connection between the camera and the server based on the data cable will fail. It is necessary to use network equipment for transmission. The working mode of processing first and then transmitting cannot be used. Instead, the working mode of transmitting first and then processing must be adopted.

[0009] In addition, when using multiple 3D cameras for real-time reconstruction, there are certain problems or deficiencies in the core technologies of real-time 3D spatial reconstruction, such as synchronous acquisition between cameras, synchronous processing of 3D data, comprehensive management and configuration of cameras, and real-time adjustment of reconstruction frame rate. It is also difficult to increase the number of cameras on a large scale.

[0010] Reference 1, Li Wenjing. Research on Multi-camera Calibration and 3D Reconstruction Technology [D]. Xi'an University of Technology, 2022. DOI:10.27398 / d.cnki.gxalu.2022.000869.

[0011] Reference 2, Patent Publication No. CN111311689B, Patent Title: A Method and System for Calibrating the Relative External Parameters of a LiDAR and a Camera.

[0012] Reference 3, Patent Publication No. CN108616742B, Patent Title: A 3D Panoramic Shooting System and Method. Summary of the Invention

[0013] This invention provides a distributed real-time 3D spatial reconstruction system and method, which can deploy and position multiple 3D cameras at any location for large-scale scenes or large-sized objects, and reconstruct the physical space covered in real time. It can also manage the 3D cameras, transmit reconstruction data, and present 3D models in real time through a network, thus eliminating the limitations of real-time 3D reconstruction on scene terrain, camera spacing, transmission distance, etc.

[0014] To achieve the above objectives, in a first aspect, the present invention provides a distributed real-time 3D spatial reconstruction system, the system comprising server nodes and multiple edge computing nodes, each edge computing node comprising a 3D camera and an edge computing device; wherein,

[0015] Each edge computing node is used for:

[0016] Color depth information is acquired using a 3D camera, and the color depth information includes both color information and depth information.

[0017] The edge computing device performs spatial alignment and temporal synchronization processing on the color depth information acquired by the corresponding 3D camera in real time; and after encoding the color depth information after spatial alignment and temporal synchronization processing, it is sent to the server node in real time.

[0018] Server nodes are used for:

[0019] After receiving the color depth information sent by each edge computing node, it is decoded in real time, and point cloud data is generated based on the decoded color depth information; the point cloud data can be rendered by server nodes or terminal nodes to obtain the corresponding three-dimensional model.

[0020] On the other hand, the present invention provides a distributed real-time three-dimensional space reconstruction method, applied to the above-mentioned system, the method comprising:

[0021] Each edge computing node acquires color depth information via a 3D camera; the color depth information includes color information and depth information.

[0022] Each edge computing node performs spatial alignment and temporal synchronization processing on the color depth information acquired by the corresponding 3D camera in real time through the edge computing device, and then encodes the color depth information after spatial alignment and temporal synchronization processing and sends it to the server node in real time.

[0023] After receiving the color depth information sent by each edge computing node, the server node decodes it in real time and generates point cloud data based on the decoded color depth information; the point cloud data can be rendered by the server node or terminal node to obtain the corresponding three-dimensional model.

[0024] The beneficial effects of this invention are:

[0025] 1. Real-time acquisition and reconstruction of large-scale 3D scenes, with an emphasis on real-time performance and large scale, is the main difference from traditional 3D reconstruction methods.

[0026] 2. Real-time high frame rate streaming transmission of encoded and compressed 3D data, including spatial location data and color image data, over the network. Traditional methods only transmit color data or only transmit spatial location data acquired in a single frame.

[0027] 3. The camera can be installed in a flexible location, and the calibration method supports rapid deployment of the equipment.

[0028] 4. Data acquired by discrete cameras are combined in real time, and after time synchronization and 3D spatial stitching, a large-scale, real-time changing 3D scene data is formed. In contrast, 2D images acquired by traditional cameras are independent and do not have 3D spatial stitching.

[0029] 5. Data that is aligned and stitched in real time in three-dimensional space is still recorded, stored, and played back in three-dimensional format. Attached Figure Description

[0030] Figure 1 This is a schematic diagram of the mounting assembly of the tracker, tracker base, rigid intermediate connector and camera in one embodiment;

[0031] Figure 2 This is a schematic diagram of a tracker and its base being mechanically mounted on a rigid intermediate connector in one embodiment.

[0032] Figure 3 This is a schematic diagram of a tracker and its base being magnetically mounted on a rigid intermediate connector in one embodiment.

[0033] Figure 4 This is a schematic diagram of a scenario in which the camera and the tracker mounted on it are calibrated in a coordinate system using a combination of a planar marker and the tracker in one embodiment.

[0034] Figure 5 A flowchart illustrating coordinate system calibration between a camera and a tracker mounted on it in one embodiment;

[0035] Figure 6 This is a schematic diagram of a single local tracker location in one embodiment;

[0036] Figure 7 Here is a flowchart of a single local tracker localization process in one embodiment;

[0037] Figure 8 This is a flowchart of global tracker expansion and global camera coordinate system unification in one embodiment;

[0038] Figure 9A This is a schematic diagram of global tracker expansion in one embodiment;

[0039] Figure 9B This is a schematic diagram of global tracker expansion in one embodiment;

[0040] Figure 9C This is a schematic diagram of global tracker expansion in one embodiment;

[0041] Figure 9D This is a schematic diagram of global tracker expansion in one embodiment;

[0042] Figure 9E This is a schematic diagram of global tracker expansion in one embodiment;

[0043] Figure 10 This is a schematic diagram of the structure of a distributed real-time 3D spatial reconstruction system in one embodiment;

[0044] Figure 11 This is a structural block diagram of a distributed real-time 3D spatial reconstruction system in one embodiment;

[0045] Figure 12 This is a topology diagram of module 2 in one embodiment;

[0046] Figure 13 This is a topology diagram of module 3 in one embodiment;

[0047] Figure 14 This is a topology diagram of module 4 in one embodiment;

[0048] Figure 15 This is a topology diagram of module 5 in one embodiment;

[0049] Figure 16 This is a topology diagram of module 5 in one embodiment;

[0050] Figure 17 This is a topology diagram of module 5 in one embodiment;

[0051] Figure 18 This is a topology diagram of module 6 in one embodiment;

[0052] Figure 19 This is a topology diagram of module 6 in one embodiment;

[0053] Figure 20 This is a topology diagram of module 6 in one embodiment;

[0054] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0055] This section will describe in detail specific embodiments of the present invention. Preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the drawings is to supplement the textual description with graphics, so that people can intuitively and vividly understand each technical feature and overall technical solution of the present invention, but they should not be construed as limiting the scope of protection of the present invention.

[0056] This invention relates to two major inventive concepts: the first is a distributed real-time 3D spatial reconstruction system and method, and the second is a multi-camera spatial positioning method and system with a large-area non-overlapping field of view.

[0057] First, the distributed real-time 3D spatial reconstruction system provided by the present invention is described. In one embodiment, a distributed real-time 3D spatial reconstruction system is provided, the system comprising multiple 3D cameras, edge computing devices deployed at each 3D camera, network communication devices, server nodes, and terminal nodes; wherein;

[0058] A 3D camera is used to acquire color depth information of the covered scene. Color depth information includes color information and depth information.

[0059] Edge computing devices, located next to a 3D camera, establish a direct data connection with the camera and possess information acquisition, data processing, and network communication functions. They are used for:

[0060] Establish a device management channel with the server node, receive instructions from the server node to adjust the operating parameters of the 3D camera, and send camera configuration to the server node;

[0061] The color depth information acquired by the corresponding 3D camera is spatially aligned and temporally synchronized in real time.

[0062] Receive instructions from the server node to adjust the processing method of the color depth information acquired by the 3D camera;

[0063] A data transmission channel is established with the server node, and the color depth information, after spatial alignment and time synchronization processing, is encoded and sent to the server node in real time.

[0064] Network communication equipment is used to establish network connections between edge computing devices, server nodes, and terminal devices.

[0065] Server nodes, which can be local or remote computing servers, are used for:

[0066] Establish network connections with terminal nodes and edge computing devices, and send control information received from terminal nodes to edge computing devices, or send 3D camera operating parameters and edge computing device operating status received from edge computing devices to terminal nodes.

[0067] After receiving color depth information sent by various edge computing devices, it decodes it in real time, and generates three-dimensional color point cloud data based on the restored color depth information.

[0068] Rendering is performed based on point cloud data to generate a 3D model, and the 3D model is sent to various terminals for display.

[0069] Store 3D model files when recording 3D scenes, or send 3D models to various terminals for display during playback.

[0070] Terminal nodes are used for:

[0071] Receive the 3D model sent by the server node and render and display it;

[0072] It receives user operation instructions and sends them to the server, thereby controlling edge computing devices and 3D cameras;

[0073] It receives user interaction commands for the presented 3D model and responds accordingly.

[0074] like Figure 10 and Figure 11 As shown, in the distributed real-time three-dimensional spatial reconstruction system of the present invention, a three-dimensional camera and an edge computing device next to it together form an edge computing node.

[0075] The edge computing node is the basic deployment unit of the acquisition unit. When deploying, the edge computing node is set up around the area to be reconstructed. There is no need to set up the field of view overlap area for the 3D camera, and there is no need to consider the physical distance between the edge computing node and the server node and the terminal node.

[0076] The network devices, including communication devices such as switches and routers, establish interconnections between edge computing nodes, server nodes, and terminal nodes. After the network connection is set up, edge computing nodes do not need to obtain the specific network address information of the server nodes, but they do need to obtain the network where the server nodes are located to register their own connection information therein; terminal nodes need to obtain the specific network address information of the server nodes in order to directly connect to the server nodes.

[0077] The terminal node can be independent of the server node or merged into the server node, with the functional settings and data transmission logic remaining unchanged.

[0078] Next, the method corresponding to the distributed real-time 3D spatial reconstruction system provided by this invention will be described. This method is implemented by dividing into six software modules, which will be introduced one by one below to facilitate understanding of the method corresponding to the distributed real-time 3D spatial reconstruction system provided by this invention by those skilled in the art.

[0079] Module 1: 3D Camera Spatial Position Calibration Module

[0080] The following method for multi-camera spatial localization with a large, non-overlapping field of view is used to obtain the relative positional relationship of each camera relative to a reference camera. The reference camera is the camera that provides the basic reference coordinate system among all cameras. The relative positional relationship is a 4×4 rotation-translation matrix of the other cameras relative to the reference camera, describing their azimuth distance and rotational attitude relative to the reference camera. The specific details of the multi-camera spatial localization method with a large, non-overlapping field of view can be found below and will not be repeated here.

[0081] Module 2: Synchronous Acquisition and Reception Module for Color Depth Information

[0082] like Figure 12 As shown, module 2 is used to synchronize color depth information acquired by multiple 3D cameras to the same time base.

[0083] The synchronous acquisition and reception of color depth information involves controlling multiple 3D cameras to acquire color depth information of the physical scene at the same or approximately the same time point. Edge computing nodes send color depth information with timestamps, and server nodes receive and rearrange the color depth information with timestamps to the same time reference.

[0084] When controlling multiple 3D cameras to acquire color depth information of a physical scene at the same or approximately the same time point, a network time server is first selected to provide a unified time reference accurate to milliseconds. Each edge computing node and the server node connects to this network time server to obtain the unified time reference and keep their system clock synchronized with the network server clock, ensuring the synchronization and consistency of data acquisition events. When a color depth information acquisition event is triggered, each edge computing node receives the acquisition signal and acquires color depth information at the same time point, making the color depth information acquisition time of each 3D camera the same or approximately the same.

[0085] Before sending color depth information, the edge computing nodes add the corresponding acquisition time as a timestamp to the collected color depth information, compress the combined information as a whole using a lossless encoding method, and then each edge computing node sends the encoded information to the server node.

[0086] After the data transmission channel is opened, the server node begins receiving time-stamped color depth information simultaneously sent by all edge computing nodes. The color depth information from each 3D camera enters the server node from different communication threads and is temporarily stored in the data queue corresponding to the edge computing node. The server node parses the timestamps corresponding to the color depth information from the received data and extracts the color depth information from all received data queues based on the acquisition time point. This ensures that the color depth information from all 3D cameras to be processed is on the same time base, with the same or similar acquisition time. The server then decodes the originally compressed color depth data for stitching and fusion of 3D spatial data. The time server can be a remote standard time server, a server node, or a single edge computing node.

[0087] The trigger event for acquiring color depth information can be an information acquisition signal set at a fixed time interval, corresponding to the automatic acquisition of color depth information of the physical scene at a fixed frame rate, serving the high-frequency real-time 3D reconstruction needs, or it can be an instantaneous acquisition signal sent by the terminal device, corresponding to the single acquisition of color depth information of the physical scene, serving the instantaneous real-time 3D reconstruction needs.

[0088] Module 3: 3D Spatial Data Stitching and Fusion Module

[0089] like Figure 13 As shown, module 3 is used to convert the color depth information of each camera into spatial 3D point cloud data in real time and stitch them together under a unified coordinate reference.

[0090] On each camera data processing thread on the server node, the decoded color and depth information is processed and transformed. After decoding the color and depth information, the two-dimensional color and depth images are preprocessed, including but not limited to using filtering algorithms to reduce measurement noise in the depth image and image processing algorithms to improve image quality in the color data. Subsequently, a stereo matching algorithm is used to extract high-precision 3D point clouds from the color and depth data. Each point has independent color and spatial location information, generating 3D point cloud data corresponding to each 3D camera. Under the same reference coordinate system, the extrinsic parameter matrix corresponding to each camera can be extracted from the extrinsic parameter matrix list of the camera calibration results and set to the corresponding processing thread. By transforming the spatial location coordinates of all point cloud data corresponding to each 3D camera according to the extrinsic parameter matrix of each 3D camera, the local physical scene acquired by all 3D cameras can be mapped to the corresponding position in the virtual space. The 3D point cloud data within the coverage area of ​​each camera are stitched together to form a larger reconstruction space.

[0091] The stitching and fusion of the three-dimensional spatial data is performed cyclically at a certain frame rate to reconstruct the physical scene in real time.

[0092] Although the 3D point cloud data comes from each 3D camera, they all use the same color space description method, and their color data values ​​remain consistent across the color data collected by different cameras.

[0093] When the 3D point cloud data is generated from color depth data, the point cloud position coordinates are all based on their respective camera coordinate systems. After transformation by the camera extrinsic matrix, the 3D point cloud data corresponding to all 3D cameras are transformed to the same coordinate space, using the coordinate system at camera calibration as a unified reference.

[0094] The point cloud data generation and coordinate position transformation are both differentiated by 3D camera and processed in an independent thread, receiving color and depth information from each 3D camera.

[0095] Module 4: 3D Point Cloud Data Synchronous Rendering Module

[0096] like Figure 14 As shown, module 4 is used to adjust the processing and rendering frame rate of the corresponding 3D point cloud data of each camera, so that each camera displays the point cloud data at the same or similar frame rate.

[0097] The data processing threads corresponding to each camera on the server node have different processing frame rates due to factors such as data volume, network latency, and uneven distribution of computing resources. That is, the number of times per second that color depth data is received, decoded, 3D point cloud is generated, and coordinate changes are processed varies, resulting in inconsistent reconstruction speeds for different local areas in the overall performance of the point cloud.

[0098] Terminal nodes can configure specific frame rate settings for each camera thread on the server node, adjusting the processing and rendering frame rate of the corresponding 3D point cloud data for each camera according to the actual situation, so that the processing speed of each camera thread tends to be the same or similar.

[0099] After setting a specific frame rate for a camera processing thread, the overall processing time T0 at that frame rate can be calculated first. This ensures that the corresponding 3D point cloud data is rendered in real-time at that frame rate while the thread processes each frame at time T0. After the processes of receiving color depth data, decoding, generating 3D point clouds, and handling coordinate changes for the current frame are completed, the time T1 consumed during these processes can be calculated. If T0 is greater than T1, it indicates that the current thread's data processing speed is too fast. In this case, T0-T1 is set as the waiting time interval for the current frame, and the thread pauses at this interval. After the waiting time, the corresponding point cloud data is submitted to the rendering device for image generation. If T0 is less than T1, it indicates that the current thread's processing speed is too slow and cannot reach the target frame rate. The target frame rate can be reduced, and this frame rate is synchronously updated to all camera data processing threads. Simultaneously, thread waiting is no longer set, and the corresponding point cloud data is directly submitted to the rendering device for image generation.

[0100] When calculating the waiting time interval according to the above process, if all processing threads encounter a situation where they need to wait, it indicates that the set target frame rate is too low, and the target frame rate can be gradually increased.

[0101] Module 5 Network Connectivity and Data Transmission Module

[0102] like Figure 15 , 16 As shown in Figure 17, module 5 is used to manage the 3D camera and transmit color depth information via the network.

[0103] The edge computing node has communication and network connectivity capabilities. After initiating a network connection, it registers its network address, port, and other device connection information with the network where the server node is located based on the network communication device, and waits for a connection signal from the terminal node.

[0104] The terminal node possesses communication and network connectivity capabilities. After initiating a network connection, it sends a connection signal to the server node based on the network communication device, establishing a network connection between the terminal node and the server node. Subsequently, the server node forwards the connection signals from the terminal devices to each edge computing node, establishing a network connection between the edge computing nodes and the terminal nodes.

[0105] Based on the network connection between edge computing nodes and terminal nodes, edge computing nodes are managed through terminal devices, establishing a device management channel. Specifically, edge computing nodes send configuration information of 3D cameras and edge computing devices to the outside world. This information is forwarded to terminal nodes via server nodes, displayed to users on the terminal devices, and edited by the users. After modifying the configuration information of the edge computing nodes, the terminal nodes send the updated configuration information to the outside world. This information is forwarded to the edge computing nodes via server nodes, and then the edge computing nodes update the configuration of the 3D cameras and edge computing devices.

[0106] Based on the network connection between edge computing nodes and terminal nodes, edge computing nodes send color depth information to server nodes in real time, which is then processed and sent to terminal devices for display, establishing a data transmission channel. Specifically, after acquiring color depth information through the data synchronization acquisition module, each edge computing device aligns the color depth information to the same time reference, encodes and compresses the color depth information using a lossless compression algorithm, and sends the compressed data to the server node via the network. The server node decodes the raw data to obtain the color depth information from each camera, processes it through the 3D spatial data stitching and fusion module to generate a 3D point cloud rendering image, and then sends this image to the terminal node. After receiving the 3D point cloud rendering image, the terminal node can interact with the point cloud data presented on the terminal node by collecting user interaction commands, or collect user-modified device configuration information and reset some edge computing node parameters that can be configured in the real-time data stream through the aforementioned device management channel.

[0107] In the aforementioned color and depth data transmission, it is also supported to transmit the color and depth data streams separately. The color data stream is compressed by an encoder and then transmitted over the network to the server node, where it is decoded and restored to its original color state. The depth data, however, is transmitted directly to the server node without an encoder, providing lossless data transmission. Both data streams are transmitted in parallel, but each stream includes a timestamp. Upon receiving the separated color and depth data, the server node extracts color and depth information collected at the same or nearby timestamps, merges them into a single frame, and achieves a realistic reconstruction of the 3D scene point cloud.

[0108] The data transmission access is established after the device management channel is set up and is activated when the user sends a start data transmission command from the terminal node to the edge computing node through the device management channel. Each completed transmission corresponds to the full processing of one frame of data from the 3D camera. In real-time 3D reconstruction, this data transmission channel must remain normal and continuously running until it is closed when the user sends a stop data transmission command from the terminal node to the edge computing node through the device management channel.

[0109] When server nodes connect to edge computing nodes, both the device management channel and the data transmission channel adopt a multi-threaded parallel approach. Each edge computing node establishes an independent network connection with the server node and simultaneously sends data to or receives data from the server node.

[0110] When the server node connects to the terminal node, both the device management channel and the data transmission channel use a single connection method. Device configuration and status information are transmitted serially between the two in a sequential manner. The color depth information sent by each camera is stitched together and processed into a 3D point cloud rendering image, which is then transmitted to the terminal node.

[0111] The network connection between the edge computing node and the cloud / local server, and the network transmission between the cloud / local server and the terminal node, can utilize a combination of wireless and wired transmission methods depending on the working environment, in conjunction with devices such as switches and routers to establish network connections. In specific embodiments, the network transmission technologies include, but are not limited to, one or more combinations of the following: Ethernet cables, optical communication, Wi-Fi, 4G / 5G, Bluetooth, ZigBee, etc.

[0112] In a specific embodiment, the preferred network connection between the edge computing node and the cloud / local server is implemented in two ways: first, each edge computing node is connected to the device-side network switch via a local Ethernet cable and / or Wi-Fi, and then connected to the cloud / local server via an Ethernet cable / optical communication method; second, each edge computing node is directly connected to the cloud server via a 4G / 5G network.

[0113] In a specific embodiment, the preferred network transmission between the cloud / local server and the terminal node is implemented in two ways: first, each terminal node is connected to the terminal-side network switch via a local Ethernet cable and / or Wi-Fi, and then connected to the cloud / local server via an Ethernet cable / optical communication method; second, each terminal node is directly connected to the cloud server via a 4G / 5G network.

[0114] Module 6: Recording and Playback of 3D Spatial Data

[0115] like Figure 18 , 19 As shown in Figure 20, module 6 is used to record real-time 3D reconstruction data and play back the recorded 3D reconstruction data.

[0116] In the real-time 3D spatial reconstruction system described in this invention, 3D reconstruction data can be displayed in real time and simultaneously stored in a file system for playback. The storage of 3D reconstruction data uses recording events as the basic descriptive unit. Each recording event includes a recording event message, configuration files for each camera, and color depth files for each camera.

[0117] The recorded event information is a general description of the recorded event, including at least the start time, end time, recording duration, encoding method, number of cameras, storage location of camera configuration files, and storage location of camera color depth files. This information can serve as a retrieval entry point for each recorded event, stored in a searchable database in tabular form with the recording start time as the primary key.

[0118] The camera configuration file contains the configuration and parameters of the 3D camera, including at least resolution, acquisition frame rate, intrinsic parameter matrix, and extrinsic parameter matrix. The configuration information for each camera is unique. This file is stored in a dedicated file storage system, and the storage path is recorded in the corresponding recording event.

[0119] The camera color depth file is a video file composed of camera color depth information arranged in a time sequence. It corresponds to the three-dimensional information of the physical scene captured by the camera during operation, and each camera's color depth file is unique. This file is stored in a dedicated file storage system, and the storage path is recorded in the corresponding recording event.

[0120] When the main thread of the server node receives the start recording signal from the terminal node, it creates a recording event with the current time as the start time point. Each camera processing thread, following the file naming rules, stores the received color depth information in the file storage system and then performs subsequent processing on the color depth data.

[0121] When the main thread of the server node receives a stop recording signal from the terminal node, it retrieves the configuration information of all cameras from the edge computing node and saves it to the corresponding location in the file system. Using the current time as the end time of the recording event, it completes the recording event information, including recording duration and encoding method, adds the full paths of the video file and camera configuration files, and saves the 3D reconstruction recording event to the database. Upon receiving the color depth information from the edge computing node, each camera processing thread no longer stores the color depth information and proceeds directly to subsequent processing.

[0122] The recording of the real-time 3D reconstruction data requires the data transmission channel of the real-time 3D reconstruction system to be opened, and the recording should not affect the continued operation of the reconstruction system when it is finished.

[0123] Compared to recording real-time 3D reconstruction data, the real-time 3D reconstruction system described in this invention also supports playback of the recorded 3D reconstruction data. After reading the recorded event sequence from the database, the terminal node can sort the recorded events according to their recording start time, allowing the user to select the events to be played back and send a playback command and the event start time to the server node. Upon receiving the playback command and the event start time, the server node retrieves the corresponding recorded event information from the recorded event database based on the event start time, then reads the configuration information of each camera from the file system, starts the corresponding processing thread for each camera according to the number of cameras, and sets the camera configuration information in the corresponding processing thread. Each thread begins reading the color depth information sequence from the corresponding color depth file and performs subsequent point cloud stitching processing and point cloud image rendering until the color depth file is completely read.

[0124] The recorded events can be retrieved from the recorded event database based on the event start time, or specific recorded events can be deleted from the recorded event database based on the event start time. At the same time, the camera configuration file and camera color depth file are also deleted from the file system.

[0125] The playback of the recorded events does not require enabling the edge computing node, enabling the device management channel, or connecting the terminal node and the edge computing node. On the data transmission channel, the edge computing node does not need to be started. Instead of receiving color depth information from the edge computing node, the corresponding camera processing threads on the server node read timestamped color depth information sequences from the corresponding camera files; subsequent processing remains unchanged.

[0126] Next, the multi-camera spatial localization method with a large-area non-overlapping field of view described in this invention will be described.

[0127] To better understand the technical contributions of this method, it is necessary to describe the background of the inventor's invention:

[0128] Computer vision, based on physical environment information acquired by cameras, can perform high-precision identification, tracking, and segmentation of people and objects within a scene, serving needs such as environmental perception, scene understanding, and intelligent interaction. It has created significant economic and social benefits and is a typical characteristic of the new generation of information technology. Among these, 3D cameras, due to their spatial perception capabilities, can obtain relative depth information of objects within a scene while perceiving color images, thus being widely used in 3D scene reconstruction, object segmentation, and spatial tracking. A typical method involves superimposing 3D physical scene information obtained from multiple cameras according to their actual spatial positions and orientations, supplementing the field of view of a single camera from different directions, and expanding the perceptible scene range through data stitching.

[0129] One of the key steps in this process is obtaining the relative position and attitude information between cameras, that is, obtaining the position and attitude of all cameras' local coordinate systems within a unified global coordinate system. A commonly used method is dual-camera calibration, which involves placing identifiable markers within the overlapping field of view between each pair of cameras. Image processing methods are then used to obtain the relative position and attitude relationships between the cameras and the markers, thereby calculating the relative position and attitude relationships between each pair of cameras. Finally, the coordinate system transfer is used to obtain the relative position and attitude relationships between all cameras. However, this method relies on overlapping fields of view between cameras, making it inefficient and failing in scenarios with large areas of non-overlapping fields of view, such as discretely distributed cameras, long distances, and large viewing angle deviations.

[0130] Based on research, existing multi-camera calibration schemes with non-overlapping fields of view can be broadly categorized into the following three technical approaches:

[0131] The first method involves unifying multiple calibration plates into the same coordinate system to obtain the transformation relationship between the coordinate systems, thus achieving coordinate system transfer between cameras with non-overlapping fields of view. As described in Reference 4, this method utilizes multiple small calibration plates. First, the intrinsic parameters of each camera are calibrated using the Zhang Zhengyou method. Then, the relative positional relationship of the small calibration plates is measured using a total station. These small calibration plates are then unified into a single large calibration plate, and the pose relationship of each camera relative to the unified calibration plate is calculated. Finally, the extrinsic parameters between the multiple cameras are obtained by unifying the coordinate system.

[0132] The second method involves moving the cameras so that they can all capture the same target, thereby obtaining the transformation relationship between coordinate systems and realizing the transfer of coordinate systems between cameras with non-overlapping fields of view. As described in Reference 5, a fixed planar calibration plate and a two-axis turntable are used to fix multiple cameras on the two-axis turntable. A world coordinate system is established with the upper left corner of the target as the origin. By rotating the turntable itself, multiple images are obtained from multiple positions using a single camera, and the relationship between the world coordinate system and the turntable coordinate system is determined. Next, the turntable is rotated so that the target enters the field of view of each camera to complete the extrinsic parameter calibration of the single camera. Finally, the relative positional relationship between each camera is solved based on the relationship between the world coordinate system and the turntable coordinate system, and the relationship between the world coordinate system and the coordinate systems of each camera.

[0133] The third method involves introducing an auxiliary camera to capture images of all calibration plates to obtain the transformation relationship between coordinate systems and realize the transfer of coordinate systems between cameras with non-overlapping fields of view. The principle is as described in reference 6, which involves obtaining the rotation matrix and translation vector between each of the cameras to be calibrated and the auxiliary camera.

[0134] Through analysis of existing multi-camera calibration techniques with non-overlapping fields of view, the inventors discovered that these three existing methods all require each camera to first take individual photos of the calibration board or target. Then, using methods such as the Zhang Zhengyou calibration method, the intrinsic and extrinsic parameters of each camera are determined from the photos. Only after these photos are used can a transformation matrix be constructed using the intrinsic and extrinsic parameters of each camera to obtain the coordinate system transformation relationship between the cameras, thus achieving the transfer of coordinate systems between cameras with non-overlapping fields of view (for detailed explanations of the formulas and principles, please refer to Chapter 3 of Reference 1 regarding multi-camera calibration). However, the method of constructing a transformation matrix between the coordinate systems of multiple cameras using the intrinsic and extrinsic parameters of each camera has the following shortcomings:

[0135] 1. Before performing multi-camera calibration, each camera must be calibrated separately to record its intrinsic and extrinsic parameters. Only then can these cameras be combined into a new system for multi-camera calibration without a common field of view. During the calibration process, the position and attitude of each camera must remain consistent with those during individual calibration. This is because changes in camera position and attitude will cause changes in extrinsic parameters, requiring individual recalibration of each camera to determine new extrinsic parameters, which is inconvenient.

[0136] 2. Due to manufacturing processes, lens distortion includes radial distortion, tangential distortion, eccentric distortion, and thin prism distortion. In single-camera calibration, for algorithmic complexity considerations, only radial distortion is typically considered, while other distortion parameters are ignored. For a single camera, distortion parameters have little impact on subsequent tasks such as 3D reconstruction, object segmentation, and spatial tracking. However, in multi-camera calibration, the distortion error of the camera furthest from the reference coordinate system is amplified during coordinate system transfer calculations. For example, if camera number 1 is used as the reference camera, and there are 40 cameras, the coordinate system transformation of camera number 40 to the reference camera requires parameters from 38 cameras (numbers 2 to 39) for coordinate transfer, leading to the accumulation and amplification of distortion error from camera number 40. Furthermore, if these 40 cameras are calibrated individually using different calibration boards, the parameters obtained from different camera calibrations may be inconsistent. These inconsistent parameters introduce uncertainty during coordinate system transfer in multi-camera calibration. For example, if the calibration accuracy of camera #5 is poor, then subsequent cameras that need to be converted from the parameters of camera #5 into the reference coordinate system will all have the accuracy shown in the tutorial. In summary, using the intrinsic and extrinsic parameters of the cameras as parameters for the coordinate system transfer matrix may propagate errors carried over from the calibration process to the multi-camera calibration system.

[0137] Reference 4, Wang Anran, Hao Xiangyang, Cheng Chuanqi, Jia Kaikai. A method for calibrating the extrinsic parameters of a multi-camera system using multiple small calibration plates [J]. Surveying and Mapping Spatial Geographic Information, 2019, 42(06):222-225+229.

[0138] Reference 5, Lu Yanan, Wan Zijing, Wang Xiangjun. A method for solving the positional relationship of cameras without a common field of view [J]. Applied Light

[0139] Journal of ...

[0140] Reference 6, Patent Publication No. CN116188591A, Patent Title: Multi-camera Global Calibration Method and Apparatus, Electronic Equipment.

[0141] To address this issue, the present invention provides a multi-camera calibration system and method with non-overlapping fields of view. This system can sequentially obtain the position and attitude parameters of each camera's local coordinate system under a common unified coordinate system, even when multiple cameras are discretely distributed, far apart, and have large viewing angle deviations over a large area without overlapping fields of view. Furthermore, it eliminates the need to calibrate the intrinsic and extrinsic parameters of each camera separately before multi-camera calibration, facilitating changes to the position and attitude of some cameras during the multi-camera calibration process and reducing the risk of propagation of intrinsic and extrinsic parameter errors from some cameras in the multi-camera system.

[0142] The method consists of four steps: mounting and positioning the tracker on the camera, calibrating the tracker coordinate system and the camera coordinate system, local tracker positioning, global tracker expansion, and global unification of the camera coordinate system.

[0143] It should be noted that since a 3D camera is an integration of a visible light image sensor (i.e., a camera) and a depth sensor (such as a laser scanner or a structured light scanner), in the multi-camera spatial localization method with a large-area non-overlapping field of view provided in the following embodiments, for the convenience of calibration between cameras, each 3D camera is actually treated as a camera with only a visible light image sensor (i.e., a camera). The calibration method below does not involve the calibration between the visible light image sensor (i.e., the camera) and the corresponding depth sensor on each 3D camera. The calibration of the two can be referred to in Reference 2, which is not an improvement of this application and will not be elaborated here. Therefore, the following 3D camera photographing of the checkerboard refers to taking pictures using a visible light image sensor.

[0144] Step one involves installing and positioning the tracker on the 3D camera.

[0145] The tracker has multiple signal receiving devices arranged in different directions on its outer surface. Working in conjunction with a base station transmitting signals, it can obtain the position and attitude information of the tracker's local coordinate system relative to the base station coordinate system based on the distribution of the received signal intensity, forming a corresponding position and attitude matrix. Its bottom can be a regular rectangular, circular, or other regular planar structure, and can be equipped with regular positioning structures such as square holes, round holes, V-grooves, and threaded holes. The center of its bottom is the origin of the coordinate system for spatial tracking. Two of the three coordinate axes are located on the bottom surface, and the third axis is perpendicular to the bottom surface, forming the local coordinate system on the tracker. The position and attitude information obtained during spatial tracking is the pose matrix of this coordinate system relative to the base station coordinate system.

[0146] The signal emitted by the base station is distributed outward from a single point, with a fixed coverage angle and range. Using its own coordinate system as a reference, it establishes real-time tracking relationships with trackers within the observable area. After adjusting the observation position and angle, the base station re-establishes a reference coordinate system at its current location and establishes tracking relationships with trackers within the new observable area.

[0147] The tracker is equipped with a dedicated base, which has a regular rectangular structure. It is precisely combined with the base through the positioning and connection structure at the bottom of the tracker. The rectangular edges of the base are parallel to the coordinate axes of the tracker. After the two are connected, they are not separated and are connected to the outside as a whole.

[0148] The tracker base can be equipped with positioning holes, positioning pins, and other structures. Except for the parts that connect to the tracker, the other structures are not restricted, and the positioning and connection structures can be specially designed according to the external connection method.

[0149] The tracking device and base are positioned on the 3D camera using a rigid intermediate connector with a special shape. First, this rigid intermediate connector is positioned on the outer surface of the 3D camera. Then, the tracking device and base are positioned as a whole and connected to this rigid intermediate connector, ensuring a fixed relative position between the tracking device and the 3D camera.

[0150] The rigid intermediate connector has a structure adapted to the shape of the 3D camera and is pre-fixed to each 3D camera by means of planar, end face contact positioning and hole-axis constraint positioning.

[0151] The tracker and base are positioned on the rigid intermediate connector using a mechanical method of end-face positioning and screw tightening. Alternatively, a magnetic method can be used, where a positioning pin and magnet are provided on the tracker base, and a positioning hole and a different type of magnet are provided at the corresponding position on the rigid intermediate connector. Positioning is achieved through hole-shaft cooperation, and connection and fixation are achieved through magnetic attraction.

[0152] The tracker and the base are fixed together on the rigid intermediate connector, forming a detachable structure. They are installed when the position of the 3D camera needs to be acquired, and disassembled after the position is acquired.

[0153] Step two involves calibrating the tracker coordinate system and the 3D camera coordinate system.

[0154] The 3D camera coordinate system is the reference coordinate system when the 3D camera acquires depth information (by first aligning the depth image with the color image in space, the depth information acquired by the depth sensor can be converted to the 3D camera coordinate system), and it is also the local coordinate system when stitching 3D physical scenes. It is defined on the color image sensor (i.e., the visible light image sensor) of the 3D camera.

[0155] The calibration of the tracker coordinate system and the 3D camera coordinate system involves obtaining the relative position and orientation relationship between the tracker coordinate system and the 3D camera coordinate system after the tracker and base are integrally positioned and installed onto the rigid intermediate connector. Specifically, a planar marker with a known relative position and orientation is set up along with the tracker assembly. The base station simultaneously acquires the coordinates of the tracker on the assembly and the tracker on the 3D camera. Coordinate system transformation is then used to obtain the position and orientation of the tracker mounted on the 3D camera relative to the local coordinate system of the planar marker. After the 3D camera obtains its position and orientation relative to the local coordinate system of the planar marker through image recognition, coordinate system transformation can be used to obtain the position and orientation of the 3D camera coordinate system relative to the local coordinate system of the tracker mounted on it.

[0156] The planar marker and tracker assembly includes a rectangular planar marker that can be located and identified via a color image, and a tracker and base placed at the center of the marker. The rectangular planar marker defines its own local coordinate system at its center. The coordinate axes of the tracker's coordinate system are parallel to the coordinate axes of the marker's local coordinate system, differing only in the vertical direction of the planar marker by the height of the tracker base. The relative transformation matrix between the two coordinate systems can be obtained by measuring the height of the tracker base.

[0157] The planar marker can be a regular checkerboard pattern, an ellipse, or any image with complex textures. Its relative positional relationship with the 3D camera can be obtained through feature point recognition methods.

[0158] The relative position and attitude relationship between the tracker coordinate system and the 3D camera coordinate system is consistent and universal when the shape of the tracker base, rigid intermediate connector and 3D camera does not change. When the same model of tracker and its base, rigid intermediate connector and 3D camera are connected, the relative position and attitude relationship between the tracker coordinate system and the 3D camera coordinate system does not change. The position and attitude parameters of the other coordinate system can be calculated from the position and attitude parameters of one coordinate system.

[0159] Step three is the local tracker location.

[0160] The local tracker localization involves using the base station to locate trackers within its observable area and obtain the coordinates of all trackers. Specifically, the position and observation angle of the base station are adjusted so that at least two 3D cameras are within its observable range. Trackers are then installed on all 3D cameras within this observable area, directly obtaining the position and attitude parameters of all trackers within the observable area—that is, their local coordinates relative to the base station's coordinate system. Subsequently, the local coordinates of all observed trackers are converted to relative coordinates relative to the coordinate system of a known object in this observation. Furthermore, based on the global coordinates of this known object, the coordinates of all observed trackers in the global reference coordinate system are calculated, and all trackers observed in this observation are recorded as new known trackers.

[0161] The known object refers to the object whose global coordinates have been acquired before a local localization. In the first local localization, this is the base station at the time of observation; in subsequent local observations, it is the first known tracker within the observation range. During each local localization, the local coordinates of this known object in that observation are acquired. When the base station is used as the known object, its position and rotation parameters for both local and global coordinates are 0, and its position and attitude matrices are identity matrices.

[0162] The coordinate system of the known object is a local coordinate system established by the base station or tracker itself.

[0163] The global reference coordinate system is the coordinate system of the base station when performing local positioning for the first time. The tracker position and attitude parameters observed in this instance are the global coordinates of the tracker, and there is no need to perform relative coordinate transformation.

[0164] Step four involves global tracker expansion and global unification of the 3D camera coordinate system.

[0165] The global tracker expansion involves sequentially adjusting the observation position and angle of the base station to change its observable area, gradually expanding until all trackers are covered. Specifically, the local tracker localization method is repeated multiple times, adjusting the position and angle of the base station each time. Except for the first localization, the new observable area covers at least one known tracker, until the position and attitude parameters of all trackers are obtained.

[0166] The global tracker expansion process can be performed simultaneously using one or more base stations, or the number of 3D cameras that the base station can locate each time can be increased by increasing the number of trackers.

[0167] The local tracker localization and global tracker extension require at least one base station and at least two trackers, regardless of the number of 3D cameras to be located.

[0168] Once the tracker becomes a known tracker, the tracker and base can be removed together from the rigid intermediate connector.

[0169] After the base station changes its observation position and angle, it can install trackers that have been removed and are no longer within the observable area onto the 3D camera within the current observable area.

[0170] When the number of trackers is insufficient to be installed on all 3D cameras covered by a single localization process, the localization process can be broken down into multiple localization processes, with trackers installed on only a small number of 3D cameras for localization each time.

[0171] When the distance between the 3D cameras is too far, causing the observable area of ​​the base station to not cover the trackers of at least two 3D cameras, one or more new relay trackers are set at the midpoint between the two 3D cameras. This allows the base station to change its position and observation angle so that its observable area covers the trackers installed on the 3D cameras and at least one relay tracker when the cameras are near their location, and simultaneously covers at least two relay trackers when the cameras are midpoint between them. Through multiple position transfers from the relay trackers, and following the local tracker localization and global tracker expansion method, the coordinate relationships of the trackers on multiple distant 3D cameras are established in the same coordinate system.

[0172] It can be observed that the coordinate system transfer between the various 3D cameras in this invention is not constructed using the intrinsic and extrinsic parameters obtained from the individual calibration of each 3D camera, but rather using the trackers fixed to the 3D cameras and the base station to construct the transformation matrix. Therefore, it eliminates the need to individually calibrate the intrinsic and extrinsic parameters of all 3D cameras before multi-camera calibration. This facilitates changes to the position and attitude of some 3D cameras during multi-camera calibration and reduces the risk of propagation of intrinsic and extrinsic parameter errors from some 3D cameras in the multi-camera system. Furthermore, in this method, each new tracker performs coordinate transformation with a known tracker, and the base station merely acts as an intermediary for coordinate system transformation. Therefore, there is no need to set up a dedicated module to control and record the rotation angle of the base station, which reduces the implementation difficulty.

[0173] On the other hand, assuming there are 40 3D cameras that need to be calibrated without a field of view, when the tracker is fixed to the 30th 3D camera, the position and orientation of the subsequent 10 3D cameras (numbers 31-40) can be adjusted. Unlike the prior art, which requires individual recalibration after adjustment before continuing multi-camera calibration, the present invention is more efficient in multi-camera calibration.

[0174] When the number of 3D cameras is small or they are close together, and the local positioning of all trackers on the 3D cameras can be completed through a single local positioning by the base station, the local tracker positioning process becomes the global positioning process, and the global tracker expansion is no longer performed.

[0175] Based on the calibration results of the tracker coordinate system and the 3D camera coordinate system, the position and orientation of the 3D camera coordinate system relative to the local coordinate system of the tracker mounted on it is a fixed constant. The position and orientation information of each 3D camera in the global coordinate system can be obtained by using the global coordinates of the tracker mounted on the corresponding 3D camera.

[0176] The method can further set the local coordinate system of one of the three-dimensional cameras as the reference, and through coordinate transformation, change the position and attitude parameters of all three-dimensional cameras to relative coordinates with respect to the local coordinate system of the three-dimensional camera, while the position and angle values ​​of the three-dimensional camera become 0, and the position and attitude matrix becomes the identity matrix.

[0177] The method can also introduce an additional reference coordinate system, changing the position and attitude of all 3D cameras to relative values ​​under this reference coordinate system, and performing an overall transformation on the acquired spatial position data.

[0178] This embodiment is for spatial positioning of multiple 3D spatial cameras in a large area with no overlapping field of view, but it is also applicable to spatial positioning of two or more ordinary 2D planar 3D cameras, as well as spatial positioning of two or more 3D cameras (i.e., 3D cameras) in close range with overlapping field of view.

[0179] Before implementing this embodiment, the installation positions of multiple 3D cameras have already been determined according to requirements, eliminating the need to pre-define overlapping field-of-view areas between the 3D cameras when designing the installation positions. According to this embodiment, the spacing between the 3D cameras can be set to an infinite distance, and they can face different observation angles.

[0180] When implementing this embodiment, the required operations are divided into four stages: installation and positioning of the tracker on the 3D camera, calibration of the tracker coordinate system and the 3D camera coordinate system, local tracker positioning, global tracker expansion and global unification of the 3D camera coordinate system.

[0181] Phase one of the implementation involves the installation and positioning of the tracker on the 3D camera.

[0182] refer to Figure 1 The tracker 102 and the tracker base 104 are assembled to form a tracker and base assembly 106. The rigid intermediate connector 108 and the 3D camera 110 are assembled to form a 3D camera and connector assembly 112. They are further combined to form an overall structure 114 containing the tracker and the 3D camera. This assembly is used in conjunction with the tracking base station 116 to obtain the spatial position and orientation of the overall structure 114 relative to the base station 116.

[0183] The tracker 102 has multiple signal receiving devices on its outer surface facing different directions. The base station 116 actively transmits specific signals into a certain area. When the tracker 102 is within the observable range of the base station 116, each signal receiving device calculates the position and attitude information of the tracker 102's local coordinate system relative to the coordinate system of the base station 116 in real time based on the distribution of signal sensing intensity, forming a corresponding position and attitude matrix. After adjusting the observation position and observation angle, the base station 116 will re-establish a reference coordinate system at the new position and establish a tracking relationship for the tracker within the new observable area.

[0184] In some embodiments, the tracker 102 can actively transmit signals in multiple directions, and the base station 116 can sense the signal strength and distribution to obtain their relative positions. The signal receiving device on the tracker 102 can also be configured as multiple marker points capable of independently tracking XYZ spatial coordinates. The base station 116 sequentially acquires the XYZ spatial coordinates of at least three marker points on the tracker 102, and combines them to obtain the spatial position and attitude of the tracker 102. The electromagnetic tracking technology constituted by the tracker and base station is a conventional technique in this field and will not be elaborated here. For specific implementation, please refer to the following patent documents: CN107452036A - A globally optimal optical tracker pose calculation method; CN102374847B - A six-degree-of-freedom pose dynamic measurement device and method for workspace. It should be noted that the present invention uses optical positioning technology, which has higher accuracy than UWB (ultra-wideband technology), reaching sub-millimeter level. In existing technologies, such as CN102374847B-Workspace Six-DOF Pose Dynamic Measurement Device and Method, the rotation angle of the transmitter (i.e., the base station) needs to be calculated in order to determine the position and attitude of each tracker. However, in this invention, each tracker that newly enters the current observation range performs coordinate transformation with the known trackers, and the base station only provides the relay function for the coordinate system transformation between the two. Therefore, it is not necessary to set up a special module to control and record the rotation angle of the base station, which can reduce the implementation difficulty.

[0185] In this embodiment, the bottom of the tracker 102 is a regular rectangular structure, with the center of the bottom serving as the origin of the coordinate system for spatial tracking. Two of the three coordinate axes are located on the bottom surface, and the third axis is perpendicular to the bottom surface, forming the local coordinate system of the tracker 102. The position and attitude information acquired during spatial tracking is the pose matrix of this coordinate system relative to the coordinate system of the base station 116. The dedicated base 104, corresponding to the bottom structure of the tracker 102, has a regular cuboid structure with edges parallel to the coordinate axes of the tracker 102. The two are precisely combined on the contact surface through positioning and connection structures to form the tracker and base assembly 106, which is not disassembled further and is connected externally as a whole through the base 104.

[0186] In this embodiment, a rigid intermediate connector 108 is used to position the overall structure 106 of the tracker 102 and the base 104 on the 3D camera 110. The rigid intermediate connector 108 has a structure adapted to the shape of the 3D camera 110 and is pre-fixed to the 3D camera 1120 by means of side plane and end face contact positioning, forming an integral 112 of the 3D camera and the connector, which does not affect the installation and normal use of the 3D camera 110 on the external support structure.

[0187] The tracker and base assembly 106 are positioned on the rigid intermediate connector 108 using... Figure 2The mechanical method shown involves end-face positioning and screw tightening. In this method, the tracker and base assembly 106 and the rigid intermediate connector 108 are provided with four pairs of corresponding positioning holes. The positioning holes of the rigid intermediate connector 108 are threaded. After alignment, the four bolts 202 are tightened. This method allows for quick installation and removal of the tracker and base assembly 106 from the 3D camera and connector assembly 112. This positioning can also be achieved using... Figure 3 The magnetic attraction method is shown. In this method, the tracker base is provided with three positioning pins 302 and one annular magnet 306, and the rigid intermediate connector 108 is provided with three positioning holes 304 and one annular magnet 308. The positioning holes 304 and positioning pins 302 are used to position the two, and the magnets 306 and 308 are connected and fixed by the attraction force between them.

[0188] In this embodiment, the positioning of the tracker and the base 106 on the rigid intermediate connector 108 can be achieved by two methods, both of which have the same positioning effect and do not affect the positioning method and final effect of the 3D camera.

[0189] The tracker and base assembly 106 is only mounted on the 3D camera and connector assembly 112 when the position of the 3D camera 110 needs to be obtained, and is disassembled after the position of the 3D camera 110 is obtained. The tracker base 106 and the rigid intermediate connector 108 need to have high machining precision and can be freely combined between different instances.

[0190] In some embodiments, the tracker base 104 and the rigid intermediate connector 108 are not limited in their structural design except for the area that directly contacts the tracker 102 and the 3D camera 110. They can be designed with matching appearance and positioning structures such as square holes, round holes, and V-grooves, as well as connection and force-bearing structures such as threaded holes, according to requirements.

[0191] Phase two involves the calibration of the tracker coordinate system and the 3D camera coordinate system.

[0192] After the overall structure 114 assembling the tracker and 3D camera is installed, the tracker 102 and the 3D camera 110 will have a fixed relative positional relationship. (Reference) Figure 4 The overall structure 114 of the tracker 102-1 and the 3D camera 110, together with a planar marker 402 and a tracker 106 with a base, is placed in the observable area of ​​the same base station 116. The overall structure 114 of the tracker and the 3D camera needs to be adjusted to the appropriate posture so that the 3D camera 110 can acquire images of the planar marker 402 at close range and is located as centrally as possible in the image.

[0193] refer to Figure 5The calibration of the coordinate system of tracker 102-1 and the coordinate system of 3D camera 110 may include steps 502 to 520. Steps 502 to 508 generate the coordinate transformation relationship between tracker 102-2 and planar marker 402, steps 510 to 512 generate the coordinate transformation relationship between tracker 102-1 and tracker 102-2 on the overall structure 114 of tracker and 3D camera, and steps 514 to 516 generate the coordinate transformation relationship between 3D camera 110 and planar marker 402.

[0194] Step 502 involves placing the assembly 106 of the base 104B and tracker 102-2 at the center of the planar marker 402, ensuring it is within the observable range of the base station 116 and that the planar marker 402 occupies most of the 2D image area of ​​the 3D camera 110. In step 504, the height of the tracker base 104B is measured using a high-precision measuring tool; this is the distance between the connection surface of the tracker base 104B and tracker 102-2 and the contact surface between the tracker base 104B and the planar marker 402. In step 506, the position of the base 104B on the planar marker 402 is adjusted so that the coordinate axes of the tracker 102-2 coordinate system are parallel to the coordinate axes of the local coordinate system of the marker 402, differing only in the vertical direction of the planar marker 402 by the height of the tracker base 104B. This distance represents the origin distance between the coordinate systems of the tracker 102-2 and the planar marker 402. In step 508, since the coordinate system of tracker 102-2 and the coordinate system of planar marker 402 differ only in the height distance of tracker base 104B in one coordinate axis direction, the transformation matrix between the two coordinate systems can be directly changed to the height distance of tracker base 104B based on the identity matrix.

[0195] Step 510 involves simultaneously observing trackers 102-1 and 102-2 on the overall structure 114 of the tracker and 3D camera via base station 116 to obtain their respective spatial positions. Then, in step 512, a coordinate system transformation is performed to obtain the relative transformation matrix between the two trackers 102-1 and 102-2.

[0196] In another embodiment, to avoid the mechanical error in measuring the height of the tracker base 104B affecting the transformation matrix between the planar marker coordinate system and the tracker coordinate system, a 3D camera calibration method can be used. Specifically, since the coordinates of each signal receiving device on the tracker are known in the tracker coordinate system, a 3D camera can be used to simultaneously capture images of the planar marker and the tracker on it. Then, using image feature point matching technology, the transformation relationship between the 3D camera and the tracker coordinate system, as well as the transformation relationship between the planar marker coordinate system and the planar marker coordinate system, can be obtained. The transformation matrix between the planar marker coordinate system and the tracker coordinate system can then be obtained by transferring the coordinates.

[0197] In step 514, the 2D planar 3D camera of the 3D camera 110 acquires the corresponding image and simultaneously records the image resolution and its own internal parameters, saving it as an electronic file. Step 516 involves the 3D camera 110 extracting the pixel coordinates of the marker points on the planar marker 402 image through image recognition, and obtaining its position and orientation relative to the coordinate system of the planar marker 402 through the mapping relationship between 2D and 3D positions; that is, the relative transformation matrix between the coordinate system of the 3D camera 110 and the coordinate system of the planar marker 402.

[0198] In step 518, by combining the coordinate transformation matrices of tracker 102-2 and planar marker 402, tracker 102-1 and tracker 102-2, and 3D camera 110 and planar marker 402, the coordinate transformation matrix between tracker 102-1 and 3D camera 110 on the overall structure 114 of the tracker and 3D camera is calculated. Step 520 involves storing the obtained coordinate transformation matrix between tracker 102-1 and 3D camera 110 as a known constant. Once the coordinate position of any object between tracker 102-1 and 3D camera 110 is obtained, the coordinate position of the other object can be calculated.

[0199] In this embodiment, the planar marker 402 is composed of black and white squares of equal size arranged according to certain rules. However, in some embodiments, the planar marker 402 can also be a regular circle or ellipse, or any image with complex textures. The relative positional relationship between the marker and the 3D camera can be obtained through feature point recognition methods.

[0200] In this embodiment, the relative position and orientation relationship between the coordinate system of the tracker 102 and the coordinate system of the 3D camera 110 are consistent and universal when the shape of the tracker base 104, the rigid intermediate connector 108 and the 3D camera 110 does not change. That is, after multiple trackers 102 and multiple 3D cameras 110 are combined through the tracker base 104 and the rigid intermediate connector 108, the relative position between the selected tracker 102 and the 3D camera 110 remains unchanged and the coordinate transformation matrix remains unchanged.

[0201] In some embodiments, when the same type of tracker 102 and its base 104, rigid intermediate connector 108 and 3D camera 110 are connected, the relative position and attitude relationship between the coordinate system of tracker 102 and the coordinate system of 3D camera 110 does not change, and the position and attitude parameters of the other coordinate system can be calculated from the position and attitude parameters of one coordinate system.

[0202] Phase three involves local tracking device location.

[0203] In this embodiment, each time a base station 116 is placed, the tracker 102 installed on the 3D camera 110 within the observable area is observed and located once, which is a local tracker localization process. This local tracker localization process only locates and obtains the position of the tracker 102, and does not calculate the position of the 3D camera 110. The position is only calculated after the corresponding trackers on all 3D cameras have been located.

[0204] refer to Figure 6 The position and observation angle of the base station 116 are adjusted so that at least two 3D cameras 110 are within its observable range. Trackers and bases are installed on all 3D cameras 110 within the observable range to obtain the overall structure 114 containing the trackers and 3D cameras. The position and attitude parameters of all trackers 102 within the observable range are directly obtained, which are the local coordinates relative to the coordinate system of the base station 116. The base station 116 then sends these position coordinates to the computer for storage and processing.

[0205] refer to Figure 7 The operation and algorithm flow for single-stage local tracker positioning may include steps 702 to 722. Step 702 involves adjusting the position and observation angle of the base station 116 so that at least two 3D cameras 110 are within their observation range, and installing trackers and the base assembly 106 on all 3D cameras 110. Step 704 involves simultaneously acquiring the coordinates of each tracker 102 relative to the coordinate system of the base station 116, and in step 706, sending the coordinate data to a computer in real time for storage and processing. Step 708 involves determining the positioning progress; steps 710 to 714 are performed during the first local positioning attempt, and steps 716 to 720 are performed when local positioning is not the first attempt.

[0206] In step 710, the first local positioning is initiated. No tracker 102 has yet obtained global coordinates, and no global coordinate system has been defined. At this point, base station 116 can be set as a known object, becoming the reference coordinate system for this observation. In step 712, the coordinate system of base station 116 is defined as the global coordinate system, and in step 714, the coordinates of all trackers 102 observed in this instance are set as global coordinates.

[0207] In step 716, at least one local localization has been performed and a global coordinate system has been established. Some trackers 102 have obtained their global coordinates. The coordinate system of the first known tracker 102 is selected as the reference coordinate system for this observation. In step 718, the relative position and attitude of all other trackers 102 observed in this observation relative to the reference coordinate system can be obtained through coordinate transformation, and multiple coordinate transformation matrices are obtained. In step 720, based on the obtained coordinates of the other trackers 102 relative to the known tracker 102 coordinate system, and the global coordinates of the known trackers 102, the global coordinates of all trackers 102 observed in this observation in the global reference coordinate system can be calculated.

[0208] Step 722 involves obtaining the global coordinates of all trackers 102 observed by base station 116 in this positioning, recording all trackers 102 observed in this operation as new known trackers 102, recording the data, and completing this local tracker positioning process.

[0209] In some embodiments, the base station 116 may acquire the coordinates of multiple known trackers 102 in a single observation, and may select one of the trackers 102 with known coordinates as the reference object for this positioning.

[0210] In some embodiments, relay trackers are arranged between 3D cameras that are far apart. In this case, the position of base station 116 can be adjusted so that its observation range covers the relay trackers. The relay trackers are only used as common positioning points for two local positioning operations.

[0211] Phase four of the implementation involves the global tracker expansion and the global unification of the 3D camera coordinate system.

[0212] In this embodiment, the observation position and observation angle of the base station 116 are adjusted sequentially to change its observable area. Except for the first positioning, the new observable area covers at least one known tracker. After multiple local tracker positioning processes, the observation area can be gradually expanded until it covers all trackers 102 on the three-dimensional camera. This process is the global tracker expansion process.

[0213] refer to Figure 8The global tracker expansion and global unification of the 3D camera coordinate system include steps 802 to 810. After updating the base station position according to step 702 so that its observation range covers at least two 3D cameras, in step 802, the tracker outside the current observation area is removed and installed on the 3D camera within the current observation area or used as a relay tracker according to step 804.

[0214] After performing local tracker localization according to the procedure at 700, step 806 determines whether the localization of trackers on all 3D cameras has been completed. If the localization of trackers on all 3D cameras has not been completed, steps 802 to 806 and 700 are repeated. If the localization of trackers on all 3D cameras has been completed, based on the relative transformation relationship between 3D camera 110 and the trackers 102 installed on it stored in step 520, the coordinates of all 3D cameras in a unified global coordinate system are calculated in step 808. In step 810, a certain 3D camera or an additional coordinate system can be set as a reference, and the coordinates of all 3D cameras are transformed to the new reference coordinate system, while the relative positional relationship between the 3D cameras remains unchanged.

[0215] Reference Figures 9A to 9E 3D cameras 110A, 110B, 110C, 110D, 110E, and 110F are the 3D cameras to be located. The 3D cameras are far apart and have no overlapping field of view. Using three trackers 102A, 102B, and 102C, the local tracker is located multiple times through the global tracker expansion process, and the position coordinates of the corresponding trackers on all 3D cameras can be obtained.

[0216] refer to Figure 9A The first local tracker localization was performed on 3D cameras 110A, 110B, and 110C. Base station 116, located at position 1, covers the observation range of 3D cameras 110A, 110B, and 110C. Trackers 102A, 102B, and 102C are installed on the corresponding 3D cameras, and their position coordinates can be acquired in one operation. In this local tracker localization, the coordinate system of base station 116 is the global coordinate system; no global coordinate system transformation is required for each tracker, and the trackers on 3D cameras 110A, 110B, and 110C are marked as known objects.

[0217] refer to Figure 9BA local tracker localization was performed on 3D cameras 110C and 110D. Base station 116, located at position 2, covers the observation range of 3D cameras 110C and 110D. Tracker 102A on 3D camera 110A was disassembled and reinstalled on 3D camera 110D, allowing the position coordinates of trackers 102C and 102A to be acquired in one operation. In this local tracker localization, tracker 102C is a known object. After calculating the relative position coordinates of tracker 102A with respect to tracker 102C, the global coordinates of tracker 102A are calculated based on the global coordinates of tracker 102C, and the tracker on 3D camera 110D is marked as a known object.

[0218] refer to Figure 9C Since the area of ​​3D camera 110D is far from the areas of 3D cameras 110E and 110F, tracker 102B on 3D camera 110B is disassembled and placed closer to 3D camera 110D, and tracker 102C on 3D camera 110C is disassembled and placed closer to 3D cameras 110E and 110F. Trackers 102B and 102C are set as relay trackers, and trackers 102A and 102B are localized once. Base station 116 is located at position 3, and its observation range covers tracker 102A on 3D camera 110D and relay tracker 102B, and the position coordinates of trackers 102A and 102B can be obtained at once. In this local tracker localization, tracker 102A is a known object. After calculating the relative position coordinates of relay tracker 102B with respect to tracker 102A, the global coordinates of relay tracker 102B are calculated based on the global coordinates of tracker 102A, and relay tracker 102B is marked as a known object.

[0219] refer to Figure 9D The first local positioning of relay trackers 102B and 102C was performed. Base station 116 is located at position 4, and its observation range covers relay trackers 102B and 102C, allowing its position coordinates to be obtained in one step. In this local tracker positioning, relay tracker 102B is a known object. After calculating the relative position coordinates of relay tracker 102C with respect to relay tracker 102B, the global coordinates of relay tracker 102C are calculated based on the global coordinates of relay tracker 102B, and relay tracker 102C is marked as a known object.

[0220] refer to Figure 9EA local tracker localization was performed on relay tracker 102C, 3D cameras 110E and 110F. Base station 116 is located at position 5, and its observation range covers 3D cameras 110E and 110F, as well as relay tracker 102C. Relay tracker 102B was disassembled and reinstalled on 3D camera 110E, and relay tracker 102A was disassembled and reinstalled on 3D camera 110F. The position coordinates of trackers 102C, 102B, and 102A can be obtained in one operation. In this local tracker localization, relay tracker 102C is a known object. After calculating the relative position coordinates of trackers 102B and 102A with respect to relay tracker 102C, the global coordinates of trackers 102B and 102A are calculated based on the global coordinates of relay tracker 102C, and the trackers on 3D cameras 110E and 110F are marked as known objects.

[0221] After calculating the positions of the trackers 102 installed on the 3D cameras 110A, 110B, 110C, 110D, 110E, and 110F, the coordinates of the 3D cameras 110A, 110B, 110C, 110D, 110E, and 110F relative to the base station at position 1 can be calculated based on the relative position transformation matrix between the 3D cameras 110 and the trackers 102 installed on them. This completes the spatial positioning of multiple 3D cameras at long distances with non-overlapping fields of view.

[0222] In some embodiments, the local coordinate system of one of the 3D cameras can be further set as a reference. Through coordinate transformation, the position and attitude parameters of all 3D cameras are changed to relative coordinates with respect to the local coordinate system of that 3D camera, while the position and angle values ​​of the 3D camera become 0, and the position and attitude matrices become identity matrices. Alternatively, an additional reference coordinate system can be introduced to transform the position and attitude of all 3D cameras to relative values ​​under that reference coordinate system, thereby performing a global transformation on the acquired spatial position data.

[0223] In this embodiment, a base station 116 and three trackers 102 are used for 3D camera localization. Once a tracker 102 becomes a known tracker, the tracker and its base assembly 106 can be removed from the rigid intermediate connector and installed on the new 3D camera 110 to be observed or at a relay location. In some embodiments, regardless of the number of 3D cameras to be localized, at least one base station and at least two trackers are required. The global tracker expansion process can be performed simultaneously using more than one base station, or the number of trackers can be increased to increase the number of 3D cameras that the base station can localize at each time, thereby speeding up the localization process.

[0224] In some embodiments, when the number of trackers is insufficient to be installed on all 3D cameras covered by a single localization, the localization process can be split into multiple localization processes, with trackers installed on only a small number of 3D cameras for localization each time.

[0225] In this embodiment, when the distance between the 3D cameras 110 is too great, causing the observable area of ​​the base station 116 to not cover the trackers 102 of at least two 3D cameras 110, one or more new relay trackers 102 are set at the middle position of the two 3D cameras 110. This allows the base station 116 to change its position and observation angle so that its observable area covers the trackers 102 installed on the 3D camera and at least one relay tracker 102 when the 3D camera 110 is near its location, and simultaneously covers at least two relay trackers 102 when the 3D camera 110 is located in the middle position. Through multiple position transfers of the relay trackers 102, the coordinate relationship of the trackers 102 on multiple distant 3D cameras 110 in the same coordinate system can be established.

[0226] In some embodiments, when the number of 3D cameras is small or they are close together, and the local positioning of all trackers on the 3D cameras can be completed by a single local positioning at the base station, the local tracker positioning process becomes the global positioning process, and global tracker expansion and coordinate system integration are no longer performed.

[0227] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A distributed real-time three-dimensional spatial reconstruction system, characterized in that, The system includes a 3D camera spatial positioning calibration module, a server node, and multiple edge computing nodes. Each edge computing node includes a 3D camera and an edge computing device. Each edge computing node is used for: Color depth information is acquired using a 3D camera, and the color depth information includes both color information and depth information. Edge computing devices perform real-time spatial alignment and temporal synchronization processing on the color depth information captured by the corresponding 3D cameras; the spatially aligned and temporally synchronized color depth information is then encoded and sent to the server node in real time; the server node is used for: After receiving the color depth information sent by each edge computing node, it is decoded in real time, and point cloud data is generated based on the decoded color depth information; the point cloud data can be rendered by server nodes or terminal nodes to obtain the corresponding three-dimensional model. The three-dimensional camera spatial position calibration module includes: At least two 3D cameras whose extrinsic parameters need to be calibrated; At least two trackers are provided, each tracker being detachably and rigidly connected to a 3D camera. Each tracker has multiple signal receiving devices arranged in different directions on its outer surface. These devices work in conjunction with a base station that transmits signals outward to obtain the position and attitude information of the tracker's local coordinate system relative to the base station coordinate system based on the distribution of the intensity of the received signals, thereby forming a corresponding position and attitude matrix. A base station, wherein the signal emitted by the base station is distributed outward from a single point, and within the signal coverage angle and range, it is used to establish a real-time tracking relationship with trackers within the observable area, using its own coordinate system as a reference; and The assembly includes a planar marker and a tracker with known relative positions and orientations, used to obtain the position and orientation relationship between the 3D camera coordinate system and the local coordinate system of the tracker mounted on it after the tracker is mounted on the 3D camera, through coordinate system transformation.

2. The distributed real-time three-dimensional spatial reconstruction system according to claim 1, characterized in that, The system also includes a terminal node, which is used for: The received point cloud data is rendered to obtain the corresponding 3D model and then displayed. Alternatively, it can receive user operation instructions and send them to the server node, thereby controlling the edge computing device and 3D camera; Alternatively, it can receive interactive commands from the user regarding the presented 3D model and respond accordingly.

3. The distributed real-time three-dimensional spatial reconstruction system according to claim 1, characterized in that, The server node is also used for: The control information received from the terminal node is sent to the edge computing device; wherein, the control information is used to adjust the operating parameters of the 3D camera; Alternatively, it may receive the 3D camera's operating parameters and the operating status of the edge computing device from the edge computing device and send them to the terminal node.

4. A distributed real-time three-dimensional spatial reconstruction method, characterized in that, The method, applied to the system according to any one of claims 1 to 3 and allowing non-overlapping fields of view between the three-dimensional cameras, comprises: Each edge computing node acquires color depth information via a 3D camera; the color depth information includes color information and depth information. Each edge computing node performs spatial alignment and temporal synchronization processing on the color depth information acquired by the corresponding 3D camera in real time through the edge computing device, and then encodes the color depth information after spatial alignment and temporal synchronization processing and sends it to the server node in real time. After receiving the color depth information sent by each edge computing node, the server node decodes it in real time and generates point cloud data based on the decoded color depth information; the point cloud data can be rendered by the server node or the terminal node to obtain the corresponding three-dimensional model. The method further includes: Based on a pre-defined multi-camera calibration process with no overlapping field of view, the relative positional relationship of each 3D camera relative to the reference camera is obtained. The reference camera is the 3D camera that provides the basic reference coordinate system among the various 3D cameras; the relative position relationship is the transformation relationship of other 3D cameras relative to the reference camera, which describes the orientation distance and rotation attitude of other 3D cameras relative to the reference camera. The preset multi-camera calibration process with non-overlapping fields of view includes: Camera and tracker calibration steps: Based on the assembly, determine the first transformation relationship between the camera coordinate system of each camera and the coordinate system of the tracker mounted on it; Local tracker localization steps: Adjust the position and observation angle of the base station to ensure that at least two cameras are within the observable range of the base station. After installing trackers on all cameras within the observable range, obtain the position and attitude parameters of all trackers relative to the base station within the observable range. Based on the observed position and attitude parameters of all trackers relative to the base station, calculate the transformation relationship between the tracker coordinate system of each tracker and the coordinate system of the known object in this observation. The coordinate system of the known object is a local coordinate system established for the base station or the tracker itself. The transformation relationship between the coordinate system of the known object and the global coordinate system is known. In the first local tracker localization step, the known object is the base station at the time of this observation. In subsequent local tracker localization steps, it is the first known tracker within the observation range. Global tracker extended positioning steps: Repeat the local tracker positioning steps multiple times, adjusting the position and observation angle of the base station each time. Except for the first positioning, ensure that the new observable area covers at least one known tracker, until the transformation relationship between the tracker coordinate system of all trackers and the tracker coordinate system of the known trackers in each observation is obtained. The steps for unifying the camera coordinate system globally are as follows: Based on the first transformation relationship and the transformation relationship between the coordinate system of the tracker installed on each camera and the global coordinate system, the second transformation relationship between the camera coordinate system of each camera and the global coordinate system is determined.

5. A distributed real-time three-dimensional spatial reconstruction method according to claim 4, characterized in that, The acquisition of color depth information via a 3D camera specifically includes: Control each 3D camera to acquire color depth information at the same or approximately the same time point; The encoding of the color depth information after spatial alignment and temporal synchronization specifically includes: The color depth information, after undergoing spatial alignment and time synchronization processing, is encoded by adding a corresponding timestamp; the timestamp is the time point when the color depth information was collected. After the server node receives the color depth information sent by each edge computing node, and before decoding, the method further includes: The server nodes rearrange the color depth information to the same time base based on the timestamp, so that the color depth information at the same time base can be decoded in the same batch. When transmitting color depth data separately, it is also necessary to merge the separate color depth data of each edge computing node.

6. The distributed real-time three-dimensional spatial reconstruction method according to claim 5, characterized in that, The server node rearranges the color depth information to the same time base based on timestamps, specifically including: The server node temporarily stores the color depth information received from each edge computing node in the data cache corresponding to each edge computing node; The server node parses out the timestamps corresponding to each color depth information; The server node extracts all color depth information with the same or nearly the same acquisition time point from all data caches based on the timestamp, so that all color depth information that needs to be decoded in the same batch is on the same time base.

7. A distributed real-time three-dimensional spatial reconstruction method according to claim 6, characterized in that, The generation of point cloud data based on the decoded color depth information specifically includes: Create corresponding camera processing threads for each 3D camera on the server node; On each camera processing thread, the color depth information corresponding to each 3D camera after decoding is processed and converted to generate point cloud data with the color depth information corresponding to each 3D camera under a unified standard system.

8. The distributed real-time three-dimensional spatial reconstruction method according to claim 7, characterized in that, The method further includes a process of controlling each camera's processing thread to submit point cloud data belonging to the same frame to the main thread, whereby the point cloud data corresponding to the color depth information acquired by each 3D camera at the same or approximately the same time point are considered as point cloud data belonging to the same frame. Each camera processing thread determines the target processing time for each frame based on the target frame rate; wherein, the target frame rate is the number of frames corresponding to point cloud data that each camera processing thread needs to submit to the main thread within one cycle; and the target processing time is the estimated time required for each camera processing thread to generate the point cloud data corresponding to that frame. Real-time monitoring of the actual time taken for each camera processing thread to generate point cloud data corresponding to the same frame; If the actual processing time of a camera processing thread is less than the target processing time, a waiting time is set for that camera processing thread. After the waiting time has elapsed, the camera processing thread submits the point cloud data to the main thread. The waiting time is the difference between the target processing time and the actual processing time. If the actual processing time of a camera processing thread is equal to the target processing time, then the camera processing thread will submit the point cloud data to the main thread. If the actual processing time of a camera processing thread exceeds the target processing time, the camera processing thread will reduce the target frame rate to increase the target processing time accordingly, and the reduced target frame rate will be synchronized to the other camera processing threads.

9. A distributed real-time three-dimensional spatial reconstruction method according to claim 4, characterized in that, The method further includes: When the main thread of the server node receives the start recording signal sent by the terminal node, it takes the current time as the start time point of the recording event and saves the three-dimensional spatial reconstruction data starting from the start time point as a camera color depth file. While storing the camera color depth files, the camera configuration files for each 3D camera are stored at the corresponding locations, using the start time point as a unique identifier. When the main thread of the server node receives a stop recording signal from the terminal node, it takes the current time as the end time of the recording event and stores the camera color depth file recorded from the start time to the end time in the corresponding location in the database.

10. A distributed real-time three-dimensional spatial reconstruction method according to claim 4, characterized in that, The method also includes playing back the recorded 3D reconstruction data, specifically including: Users select the recorded event to be played back from the terminal node based on information such as the recording start time and recording duration, and send the recording event playback command and the event start time to the server node; The server node retrieves the corresponding recording event information from the recording event database based on the event start time, then reads the configuration information of each camera from the file system, starts the processing thread corresponding to each camera according to the number of cameras, and sets the camera configuration information to the processing thread corresponding to the camera. Each thread begins reading the color depth information sequence from the corresponding color depth file and performs subsequent point cloud stitching and point cloud image rendering until the color depth file has been read.

Citation Information

Patent Citations

  • Work space six degree-of-freedom posture dynamic measurement equipment and method

    CN102374847B

  • Global optimum optical tracker pose calculation method

    CN107452036A

  • Method for jointly constructing 3D space and 3D object by multiple depth cameras

    CN111739080A

  • Real-time dynamic three-dimensional modeling method and system based on group of point cloud sensors

    CN115953538A