Scene multi-dimensional reconstruction method integrating data operation and space intelligence
By constructing a method that integrates 3D point cloud and business process knowledge graph, the problem of separation between business process and spatial perception information in existing technologies is solved, realizing unified representation and dynamic reflection of business logic and physical space, and supporting efficient business flow analysis and evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN MINGHOUTIAN INFORMATION TECH CORP LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-04-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing scene 3D reconstruction technologies cannot effectively integrate business process knowledge and spatial perception information, resulting in spatial models that cannot reflect the status and evolution of business activities and cannot handle the question of 'when, where, and what kind of business operation occurred' within a unified framework.
By fusing panoramic image sequences, multi-source business data streams, and environmental spatial point clouds, a semantically enhanced fused 3D point cloud and business process knowledge graph are constructed. The spatial operation areas of key operation nodes are marked, and the timestamps and business process status are encoded into time-varying attribute vectors of 3D spatial points. The 3D spatiotemporal voxel field is reconstructed, and the topological relationship of the business process knowledge graph is embedded to generate a multi-dimensional reconstruction model.
It achieves a unified representation of business logic and physical space, accurately maps abstract business processes to specific physical spaces, dynamically reflects business progress, and supports business flow analysis and efficiency evaluation based on spatial location.
Smart Images

Figure CN121837529A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of spatiotemporal semantic reconstruction technology, specifically a multi-dimensional scene reconstruction method that integrates data operation and spatial intelligence. Background Technology
[0002] Existing scene 3D reconstruction technologies primarily process the geometric and texture information of image sequences or point cloud data to generate static 3D models or models that only contain temporal changes in appearance. These models lack an intrinsic expression of specific business activities and process logic within the scene. Furthermore, multi-source data streams generated alongside business processes, such as operation records and status signals, typically exist independently of the spatial model. Existing methods often employ post-annotation or simple spatial coordinate binding to establish a shallow, static connection between the 3D model and business data, failing to automatically locate the corresponding physical operation area based on the dynamic progression of the business process.
[0003] This separation prevents spatial models from reflecting the state and evolution of business activities. Existing technologies struggle to address the multidimensional question of "when, where, what kind of business operation occurred, and what its state is" within a unified framework. Business logic and physical space are fragmented into two different representation systems, hindering in-depth analysis and optimization of business processes based on spatial location.
[0004] A method is needed to deeply integrate business process knowledge, multi-source data streams, and spatially aware information to construct a model that simultaneously encodes spatial structure, temporal evolution, and business semantics. This method should be able to automatically and accurately define operational areas in the physical space based on business process nodes, and encode abstract business state changes as dynamic attributes of the spatial model itself, thereby achieving a unified representation and computation of business logic and physical space. Summary of the Invention
[0005] This invention aims to solve at least one of the technical problems existing in the prior art; To this end, this invention proposes a multi-dimensional scene reconstruction method that integrates data operation and spatial intelligence, comprising: The collected panoramic image sequences, multi-source business data streams, and environmental spatial point clouds are fused to construct a semantically enhanced fused 3D point cloud and business process knowledge graph. Guided by the business process knowledge graph, spatial operation regions corresponding to key operation nodes are marked in the semantically enhanced fused 3D point cloud. The spatial operation regions are defined by a series of 3D spatial point clusters and their associated entity object pixel regions. Dynamic spatiotemporal encoding is performed on the fused 3D point cloud with semantic enhancement after labeling, and the timestamps of key operation nodes and business process status are encoded into time-varying attribute vectors of corresponding 3D spatial points. Based on the time-varying attribute vector, the semantically enhanced fused 3D point cloud is reconstructed in a spatiotemporal consistency, and a dynamically changing 3D spatiotemporal voxel field is generated in the 3D space with the spatial operation region as the core. The topological relationship of the business process knowledge graph is embedded in the three-dimensional spatiotemporal voxel field. The node type and edge connection relationship of the business process knowledge graph are attached to each three-dimensional spatiotemporal voxel to form a spatiotemporal semantic field with business logic semantics. The spatiotemporal semantic field is divided into multiple dimensions, and interrelated spatial grids, time slices and business state slices are generated along the spatial dimension, time dimension and business logic dimension to construct a multidimensional reconstruction model of the scene.
[0006] Furthermore, the construction of the semantically enhanced fused 3D point cloud and business process knowledge graph includes: The system acquires panoramic image sequences, multi-source business data streams, and environmental spatial point clouds of the target scene. The panoramic image sequences include visual frames with position and pose information, the multi-source business data streams include business operation records and business process status, and the environmental spatial point clouds include three-dimensional spatial points and their reflection intensity. Pixel-level semantic segmentation is performed on the panoramic image sequence to extract the pixel regions of entity objects in the image, and the image coordinates of the pixel regions of the entity objects are registered and associated with the corresponding three-dimensional points of the environmental spatial point cloud to generate a semantically enhanced fused three-dimensional point cloud. Business process modeling is performed on multi-source business data streams to identify key operation nodes and state transition paths in business operation records and to construct a business process knowledge graph. The step of performing pixel-level semantic segmentation on the panoramic image sequence, extracting the pixel regions of entity objects in the image, and registering and associating the image coordinates of the entity object pixel regions with the corresponding 3D points of the environmental spatial point cloud includes: The panoramic image sequence is input into a pre-trained panoramic semantic segmentation network to generate a pixel-level semantic label map for each visual frame. The pixel connected regions belonging to a predetermined category list are extracted from the semantic label map as the pixel regions of entity objects. Based on the position and pose information carried by the panoramic image sequence, the camera projection matrix corresponding to each visual frame is calculated. Using the camera projection matrix, the pixel coordinates of each entity object's pixel region are back-projected to three-dimensional space, and the set of three-dimensional spatial points that fall into the back-projected view frustum is found in the environmental spatial point cloud. The found set of three-dimensional spatial points is associated with the pixel region of the entity object, and the semantic labels of the pixel region of the entity object are assigned to the corresponding three-dimensional spatial points, thereby completing the registration and semantic enhancement.
[0007] Furthermore, the step of modeling business processes for multi-source business data streams, identifying key operation nodes and state transition paths in business operation records, and constructing a business process knowledge graph includes: Analyze each business operation record in the multi-source business data stream to extract the operation subject, operation object, operation type, operation timestamp, and operation result status; Based on the order of operation timestamps, multiple business operation records in the same business process are arranged into an operation sequence; In the operation sequence, identify the business operation records that trigger a fundamental change in the state of the business process, define them as key operation nodes, and define the state changes between adjacent key operation nodes as state transition paths. Using key operation nodes as graph nodes and state transition paths as directed edges, a business process knowledge graph representing the complete business process is constructed.
[0008] Furthermore, the step of marking spatial operation regions corresponding to key operation nodes in a semantically enhanced fused 3D point cloud, guided by a business process knowledge graph, includes: Traverse each key operation node in the business process knowledge graph and parse the operation object from the corresponding business operation record; In the semantically enhanced fused 3D point cloud, find 3D spatial point clusters whose semantic labels match the operation object; Using the centroid of the found three-dimensional point cluster as the center, a spatial region is grown outward in three-dimensional space until the density of the boundary point cloud is lower than a set threshold. The defined spatial range and the included three-dimensional point cluster are the spatial operation area corresponding to the key operation node.
[0009] Furthermore, the dynamic spatiotemporal encoding of the fused 3D point cloud with enhanced semantics after labeling, encoding the timestamps of key operation nodes and business process status into time-varying attribute vectors of corresponding 3D spatial points, includes: For each 3D spatial point in the semantically enhanced fused 3D point cloud, determine its corresponding spatial operation region; For a 3D spatial point belonging to any spatial operation region, obtain the timestamp and operation result status of the key operation node corresponding to its spatial operation region. Based on the timestamp, the normalized time offset of the three-dimensional spatial points is calculated, and the operation result state is mapped to a state encoding vector of a preset dimension. The normalized time offset is concatenated with the state encoding vector to generate the time-varying attribute vector of the three-dimensional spatial point.
[0010] Furthermore, the method of reconstructing the spatiotemporally consistent structure of the semantically enhanced fused 3D point cloud based on time-varying attribute vectors, and generating a dynamically changing 3D spatiotemporal voxel field with the spatial operation region as the core in 3D space, includes: Import the complete point cloud data, which includes 3D spatial points, semantic labels, and time-varying attribute vectors, into the spatiotemporal voxelization engine. In the spatiotemporal voxelization engine, three-dimensional space is discretized into spatial voxels along the spatial axis and into time frames along the time axis. For each spatial voxel, aggregate the time-varying attribute vectors of all three-dimensional spatial points within it, and calculate its aggregated vector sequence that changes with time frames. Based on the evolution pattern of the aggregated vector sequence in the time dimension, spatial voxels are classified into static voxels, periodic dynamic voxels, and aperiodic dynamic voxels, and their dynamic characteristics are recorded. All spatial voxels and their dynamic characteristics constitute a three-dimensional spatiotemporal voxel field.
[0011] Furthermore, the embedding of the topological relationships of the business process knowledge graph in the three-dimensional spatiotemporal voxel field, by attaching node types and edge connections from the business process knowledge graph to each three-dimensional spatiotemporal voxel, includes: Establish a mapping relationship between the set of spatial voxels contained in the spatial operation region in the three-dimensional spatiotemporal voxel field and the key operation nodes in the business process knowledge graph; Assign the node type of the key operation node as an attribute to all spatial voxels mapped to it; In a three-dimensional spatiotemporal voxel field, if two spatial voxel sets are mapped to two key operation nodes in a business process knowledge graph that are directly connected by directed edges, then virtual connection edges with the same direction and type are established between these two spatial voxel sets. After completing the connection of all voxels and sets, a spatiotemporal semantic field is formed that simultaneously contains spatial geometry, temporal dynamics, and business logic.
[0012] Furthermore, the multi-dimensional spatial partitioning of the spatiotemporal semantic field, generating interrelated spatial grids, temporal slices, and business state slices along the spatial, temporal, and business logic dimensions, includes: Along the spatial dimension, the spatiotemporal semantic field is divided into regular spatial grids, and each spatial grid contains one or more spatial voxels and all their attributes; Along the time dimension, the spatiotemporal semantic field is sliced at fixed time intervals, and each time slice contains a snapshot of the dynamic state of all spatial voxels within the corresponding time period. Along the business logic dimension, based on different business process branches or state types in the business process knowledge graph, the spatiotemporal semantic field is logically divided. Each business state slice contains all spatial voxels and connection relationships belonging to the same business branch or state. Establish an index relationship between spatial grids, time slices, and service status slices, so that any spatial voxel can be uniquely identified by its spatial grid number, time slice number, and service status slice number.
[0013] Furthermore, it also includes: Real-time acquisition of new business data streams and visual perception data, synchronization to the multidimensional reconstruction model, and updating of the internal states of the corresponding spatial grid, time slice, and business status slice; In response to the reconstruction command, the multidimensional spatial partitioning results under the specified spatial range, time span and business logic constraints are extracted from the multidimensional reconstruction model, and the 3D rendering engine is driven to generate a scene reconstruction view that is dynamically synchronized with the business flow. The real-time acquisition of new business data streams and visual perception data, and their synchronization to the multi-dimensional reconstruction model, updating the internal states of the corresponding spatial grid, time slice, and business state slice, includes: The newly acquired visual perception data is processed in real time to update the geometric and semantic information of the affected areas in the semantically enhanced fused 3D point cloud. Analyze the newly added business data flow, identify new key operation nodes and state transitions, and update the business process knowledge graph; Based on the updated business process knowledge graph and the semantically enhanced fused 3D point cloud, the affected spatial operation area and time-varying attribute vector are recalculated; Based on the recalculation results, the contents of the three-dimensional spatiotemporal voxel field, spatiotemporal semantic field, and corresponding spatial grid, time slice, and business status slice are updated incrementally.
[0014] Furthermore, in response to the reconstruction command, extracting multi-dimensional spatial partitioning results from the multi-dimensional reconstruction model under specified spatial range, time span, and business logic constraints, and driving the 3D rendering engine to generate a scene reconstruction view dynamically synchronized with the business flow, includes: Parse the spatial range parameters, time span parameters, and business logic constraints contained in the reconstruction instruction; The corresponding spatial grid is selected based on the spatial range parameter, the corresponding time slice sequence is selected based on the time span parameter, and the corresponding business state slice is selected based on the business logic constraints. The selected spatial grid, time slice sequence and business status slice are intersected to obtain a set of spatial voxels that satisfy all constraints and their complete attribute data. The spatial voxel set and its complete attribute data are converted into a scene graph data structure that can be recognized by the 3D rendering engine, and the 3D rendering engine is driven to render, generating an interactive scene reconstruction view that integrates spatial geometry, temporal evolution and business logic.
[0015] Compared with the prior art, the beneficial effects of the present invention are: This method employs a business process knowledge graph to guide the labeling of spatial operation areas, using non-spatial business logic nodes as prior knowledge to drive the search and definition of target regions within a fused 3D point cloud. This approach departs from the traditional approach of relying on purely visual or geometric features for segmentation, enabling the labeled regions to directly carry explicit business semantics. It achieves a precise mapping from abstract business processes to concrete physical spaces. The labeled spatial operation areas are inherently strongly correlated with business activities, providing a direct, semantic spatial carrier for location-based business flow analysis, compliance checks, and efficiency assessments.
[0016] By dynamically attaching time-varying attribute vectors that integrate business process states and timestamps to 3D spatial points, and reconstructing a 3D spatiotemporal voxel field based on these vectors, each spatial voxel not only records its position and appearance but also embeds information about the state evolution of business logic. The dynamic changes in this voxel field directly reflect the progress of the business process. This creates a unified data representation that can simultaneously express spatial existence, temporal continuity, and changes in business state, transforming the model's dynamic attributes from external geometric deformation to internally implied business state transitions. This allows for direct querying, slicing, and analysis of the spatiotemporal model from a business logic perspective. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the steps of the multi-dimensional scene reconstruction method that integrates data operation and spatial intelligence as described in this invention. Figure 2 A flowchart for constructing a business process knowledge graph. Detailed Implementation
[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See Figure 1The system fuses acquired panoramic image sequences, multi-source business data streams, and environmental spatial point clouds to construct a semantically enhanced fused 3D point cloud and a business process knowledge graph. Based on this construction, guided by the business process knowledge graph, spatial operation regions corresponding to key operation nodes are marked in the semantically enhanced fused 3D point cloud. These spatial operation regions are defined by a series of 3D spatial point clusters and their associated entity object pixel regions. Dynamic spatiotemporal encoding is performed on the marked semantically enhanced fused 3D point cloud, encoding the timestamps of key operation nodes and business process state information into time-varying attribute vectors for the corresponding 3D spatial points. Based on these time-varying attribute vectors, a spatiotemporal consistency structure reconstruction is performed on the semantically enhanced fused 3D point cloud, generating a dynamically changing 3D spatiotemporal voxel field centered on the spatial operation regions. The 3D spatiotemporal voxel field is further processed, embedding the topological relationships of the business process knowledge graph. Each 3D spatiotemporal voxel is appended with node type and edge connection attributes from the business process knowledge graph, thereby forming a spatiotemporal semantic field that simultaneously contains spatial, temporal, and business logic information. A multi-dimensional spatial partitioning is performed on the spatiotemporal semantic field, generating interrelated spatial grids, time slices, and business state slices along the spatial dimension, time dimension, and business logic dimension, thereby constructing a multi-dimensional reconstruction model of the scenario.
[0020] See Figure 2 In one embodiment of the present invention, the panoramic image sequence is continuously captured by 360-degree panoramic cameras deployed in the photovoltaic power station or substation area. Each visual frame of the panoramic image sequence is accompanied by precise position and attitude information provided by a synchronous positioning and mapping system. Multi-source business data streams are extracted in real time from the database of the power station monitoring system or production execution system, including operation records such as equipment inspection, regular maintenance, fault handling, and switching operations, as well as the equipment status corresponding to each operation. The environmental spatial point cloud is obtained by scanning the equipment area inside the station using airborne or ground-based lidar. The environmental spatial point cloud contains millions of three-dimensional spatial points and the reflection intensity value of each three-dimensional spatial point. In another example, a gas pipeline compressor station scenario, the panoramic image sequence is acquired by a high-point monitoring camera with a fixed viewing angle within the station. The multi-source business data stream comes from the station control system and includes operation logs for compressor start-up and shutdown, filter replacement, pressure regulation, and safety inspection. The environmental spatial point cloud is generated by a 3D scanner mounted on an inspection robot. By comparing the two scenarios, the panoramic image sequence must include position and attitude information to achieve spatial alignment, the multi-source business data stream must include structured operation records and status to support process modeling, and the environmental spatial point cloud must include 3D spatial point coordinates and reflection intensity to characterize geometry and material. However, the specific data content and format vary depending on the scenario's business requirements.
[0021] In practice, pixel-level semantic segmentation of the panoramic image sequence is performed using a pre-trained panoramic semantic segmentation network. This network employs a deep neural network with an encoder-decoder structure, outputting a pixel-level semantic label map for each input visual frame. Connected pixel regions belonging to a predetermined category list are extracted from the semantic label map as entity object pixel regions. In a photovoltaic power station scenario, this list includes photovoltaic module arrays, combiner boxes, inverters, box-type transformers, and inspection personnel; in a substation scenario, it includes circuit breakers, disconnect switches, transformers, instrument transformers, insulator strings, and maintenance personnel. Based on the position and attitude information carried by the panoramic image sequence, a camera projection matrix is calculated for each visual frame. This matrix maps 3D world coordinates to 2D image coordinates. Using the camera projection matrix, the pixel coordinates of each entity object pixel region are back-projected into 3D space. The set of 3D points falling within the back-projected view frustum is then searched in the environmental point cloud. The back-projected view frustum is defined by the view frustum space formed by the pixel region boundary and the camera optical center. It is understandable that the back projection process involves coordinate system one. An example formula describes the back projection calculation relationship from pixel coordinates to a ray in three-dimensional space: ; in: Indicates the scale factor. and Represents pixel coordinate values. This represents the camera intrinsic parameter matrix. This represents the rotation matrix from the world coordinate system to the camera coordinate system. Represents the translation vector. , , This represents the world coordinates of a point in 3D space. The found set of 3D points is associated with the pixel regions of entity objects, and the semantic labels of the entity object pixel regions are assigned to each corresponding 3D point, thereby completing registration and semantic enhancement, and generating a semantically enhanced fused 3D point cloud. In some embodiments, when multiple entity object pixel regions compete for the same 3D point, the final semantic label assignment is determined based on the category confidence score output by the semantic segmentation network. Optionally, the association process can employ a nearest neighbor search algorithm to accelerate the search for the set of 3D points.
[0022] In some embodiments, the process of modeling business processes for multi-source business data flows begins with parsing each business operation record. Parsing extracts the operation subject (e.g., employee ID or machine ID), the operation object (e.g., product ID or part ID), the operation type (e.g., "picking" or "installation"), the operation timestamp, and the operation result status (e.g., "success" or "failure"). Based on the order of the operation timestamps, multiple business operation records within the same business process are arranged into an operation sequence. In the example of periodic inspections of a photovoltaic power station, an inspection process's operation sequence includes "Inspection task start," "Arrival at designated array," "Inspect component appearance," "Record infrared temperature measurement data," "Inspect combiner box," and "Inspection task end." Within the operation sequence, business operation records that trigger fundamental changes in the business process state are identified and defined as key operation nodes, such as records where the state changes from "planned" to "executed," or where the equipment state changes from "normal" to "abnormal." State changes between adjacent key operation nodes are defined as state transition paths. Using key operation nodes as graph nodes and state transition paths as directed edges, a business process knowledge graph representing the complete business process is constructed. In practical implementation, the nodes of the business process knowledge graph store operation timestamps and operation result status attributes, while the directed edges store conditional descriptions of state transition paths. Optionally, the identification of key operation nodes can be achieved by comparing whether the status fields of adjacent operation records undergo a preset transition. It can be understood that the construction of the business process knowledge graph relies entirely on the timestamps and status logic in multi-source business data streams and does not involve external rules.
[0023] In one embodiment of the present invention, the process of traversing each key operation node in the business process knowledge graph is systematic. In the inspection example, a key operation node in the business process knowledge graph may contain the operation record "Infrared temperature measurement of the 3rd string of photovoltaic modules in the 5th array was performed, and an abnormal temperature was found." In the substation switching operation example, a key operation node may contain the operation record "Execute the operation of switching circuit breaker 1012 from operation to maintenance." The operation object is parsed from the business operation record corresponding to the key operation node. In the photovoltaic inspection scenario, the operation object is "the 3rd string of photovoltaic modules in the 5th array" or a specific "inverter number." In the manufacturing scenario, the operation object is "engine component." In the semantically enhanced fused 3D point cloud, all 3D spatial points whose semantic tags match the operation object "photovoltaic module" or "inverter" are searched. These 3D spatial points with the same semantic tags are clustered to form a 3D spatial point cluster. In some embodiments, the semantic tag matching is an exact string matching. The 3D spatial point cluster is preprocessed by a spatial clustering algorithm to eliminate noise points and form a more coherent point cloud cluster.
[0024] In practice, the process of growing a spatial region outward from the centroid of the found 3D point cluster is an iterative calculation process. The centroid is obtained by calculating the arithmetic mean of the coordinates of all 3D points in the cluster. Spatial region growth starts from the centroid and gradually incorporates neighboring 3D points into the current spatial operation area. Proximity is determined by the Euclidean distance between 3D points being less than a fixed threshold. The growth process continues until the point cloud density at the boundary is lower than a pre-set threshold. The defined spatial range and all included 3D point clusters are then defined as the spatial operation area corresponding to the key operation node. Optionally, the point cloud density threshold can be dynamically set based on the average spacing of the point cloud in the scene. In the inspection example, the spatial operation area generated around the "3rd string of photovoltaic modules in the 5th array" might be the 3D spatial range where that string of modules is located. In the maintenance example, the spatial operation area generated around the "#1 main transformer" might be the transformer itself and the surrounding safe working area. It can be understood that the boundary of the spatial operation area is not a geometrically regular shape, but a complex 3D surface determined by the actual point cloud distribution.
[0025] In some embodiments, the dynamic spatiotemporal encoding operation of the labeled semantically enhanced fused 3D point cloud is performed independently for each 3D spatial point in the point cloud. For each 3D spatial point in the semantically enhanced fused 3D point cloud, its corresponding spatial operation region is determined. The determination logic is based on whether the coordinates of the 3D spatial point fall within the 3D bounding box of any spatial operation region. For a 3D spatial point belonging to any spatial operation region, the timestamp and operation result status of the key operation node corresponding to its spatial operation region are obtained. Based on the timestamp of the key operation node, the normalized time offset of the 3D spatial point is calculated. The operation result status is then converted into a fixed-dimensional state encoding vector according to a preset mapping table. Optionally, the calculation of the normalized time offset can use the following example formula: ; in: This represents the normalized time offset. This represents the timestamp when a point in three-dimensional space was captured. This represents the timestamp of the key operation node corresponding to its spatial operation region. The total time span of the entire business process is used for normalization. It can be understood that for 3D spatial points that do not belong to any spatial operation region, their normalized time offset and state encoding vector can be set as default zero vectors. In specific implementations, the state encoding vector can use one-hot encoding or embedded encoding to represent discrete operation result states. Finally, the calculated normalized time offset scalar is... Concatenate the vector with the state encoding vector to generate a time-varying attribute vector representing the spatiotemporal and state attributes of the point in the three-dimensional space.
[0026] In one embodiment of the present invention, the process of importing complete point cloud data containing three-dimensional spatial point coordinates, semantic labels, and time-varying attribute vectors into the spatiotemporal voxelization engine is a data format conversion and import process. Each three-dimensional spatial point in the point cloud data received by the spatiotemporal voxelization engine carries spatial coordinates, semantic labels, and a time-varying attribute vector composed of a normalized time offset and a state encoding vector. In the spatiotemporal voxelization engine, the three-dimensional space is discretized along the spatial axis into uniform spatial voxels. The size of the spatial voxels is preset according to the scene scale; in a warehousing example, the side length of the spatial voxel may be set to 0.1 meters to finely represent the size of goods, while in a manufacturing example, it may be set to 0.05 meters to capture the details of precision parts. The total time span of the business process is discretized along the time axis into continuous time frames, with the interval between time frames set according to the minimum time interval of the key operation nodes in the business process.
[0027] In practice, the operation of aggregating the time-varying attribute vectors of all 3D spatial points within each spatial voxel and calculating the aggregated vector sequence is performed by traversing each voxel. Each spatial voxel contains zero or more 3D spatial points, and the aggregation function calculates the time-varying attribute vectors of all 3D spatial points falling within that voxel. An example formula for calculating the aggregated vector sequence is described below: ; in: Indicates the first The aggregated vector corresponding to each time frame Indicates the first The number of 3D spatial points falling into the spatial voxel within a time frame. Indicates the first A time-varying attribute vector of a three-dimensional spatial point. This can be understood as an aggregated vector sequence. This space voxel constitutes The dynamic attribute evolution is described over several time frames. Based on the evolution pattern of the aggregated vector sequence in the time dimension, spatial voxels are classified into static voxels, periodic dynamic voxels, or aperiodic dynamic voxels. The classification can be based on calculating the variance and autocorrelation characteristics of each component of the aggregated vector sequence. The aggregated vector sequence of static voxels is almost constant in the time dimension, the aggregated vector sequence of periodic dynamic voxels exhibits regular fluctuations, and the aggregated vector sequence of aperiodic dynamic voxels shows no fixed pattern of change. All spatial voxels and their classification results, together with dynamic feature parameters, constitute a three-dimensional spatiotemporal voxel field. In some embodiments, dynamic feature parameters may include the change period and the average change amplitude.
[0028] In some embodiments, embedding the topological relationships of a business process knowledge graph in a 3D spatiotemporal voxel field begins with establishing mapping relationships. A one-to-one mapping relationship is established between the set of spatial voxels contained in the spatial operation area of the 3D spatiotemporal voxel field and the key operation nodes in the business process knowledge graph. This mapping relationship is based on a predefined association between the spatial operation area and the key operation node. In the warehousing example, the set of spatial voxels representing the operation of "picking product SKU123" is mapped to the key operation node "picking operation" in the business process knowledge graph. The node type of the key operation node is assigned as an attribute to all the mapped spatial voxels, and category labels such as "shelf," "picking," "installation," and "inspection" are attached to the attribute list of each spatial voxel. In the 3D spatiotemporal voxel field, if two sets of spatial voxels are mapped to two key operation nodes in the business process knowledge graph that are directly connected by a directed edge, a virtual connection edge with the same direction and type is established between these two sets of spatial voxels. Optionally, the direction of the virtual connection edge can be implicitly contained in the spatiotemporal index order of the 3D spatiotemporal voxel field. After establishing the connections between all voxels and sets, a spatiotemporal semantic field is formed, simultaneously encompassing spatial geometry, temporal dynamics, and business logic topology. It can be understood that each spatial voxel in the spatiotemporal semantic field not only contains geometric, semantic, and spatiotemporal dynamic attributes, but also its corresponding node type in the business process knowledge graph and its logical connections with upstream and downstream nodes. In the manufacturing example, a set of spatial voxels representing the "tightening screws" operation not only includes the dynamic changes in the contact area between the screwdriver and the screw, but is also labeled as a "fastening" node type and has a virtual connection edge with the set of spatial voxels corresponding to the upstream "placing screws" node.
[0029] In one embodiment of the present invention, the process of regularly dividing the spatiotemporal semantic field along the spatial dimension divides the entire three-dimensional spatial range of the scene into uniformly sized cubic grids. Each spatial grid contains one or more spatial voxels and all their attributes. A spatial voxel is the basic spatial unit in the spatiotemporal voxel field. In the warehousing example, the spatiotemporal semantic field covering a large warehouse may be divided into cubic spatial grids with a side length of one meter. In the manufacturing example, the spatiotemporal semantic field of a precision assembly station may be divided into cubic spatial grids with a side length of 0.2 meters. Each spatial grid has a unique spatial grid number, which is calculated based on the grid's index position in three-dimensional space. An example formula describes how the spatial grid number is calculated: ; in: Indicates the spatial grid number, This indicates the index number of the spatial grid along the X-axis. This indicates the index number of the spatial grid along the Y-axis. This indicates the index number of the spatial grid along the Z-axis. This indicates the total number of spatial grid cells in the X-axis direction. This indicates the total number of spatial grid cells along the Y-axis. Spatial grid number. A consecutive integer is used to uniquely identify a spatial grid. All spatial voxels within the spatial grid inherit this spatial grid number as their spatial dimension identifier.
[0030] In practice, the process of slicing the spatiotemporal semantic field along the time dimension at fixed time intervals divides the entire business process's time span into continuous, equal-length time segments. Each time slice contains a snapshot of the dynamic state of all spatial voxels within the corresponding time segment. The fixed time interval is set according to the business rhythm. In the warehousing example, the picking process may be in minutes, and the fixed time interval could be set to sixty seconds. In the manufacturing example, the assembly cycle may be in seconds, and the fixed time interval could be set to ten seconds. Each time slice has a time slice number arranged chronologically, and the time slice number is a consecutive integer starting from zero or the first. Each time slice captures the attribute state of all spatial voxels in the spatiotemporal semantic field within that time segment, including aggregated vector sequence fragments of spatial voxels, dynamic classifications, and node type attributes of the business process knowledge graph. A time slice is essentially a discrete sampling along the time dimension. It can be understood that the fixed time interval of a time slice should be less than the minimum time interval of critical operation nodes in the business process to ensure that critical state changes are recorded by at least one time slice.
[0031] In some embodiments, the spatiotemporal semantic field is logically divided along the business logic dimension based on different business process branches or state types in the business process knowledge graph. Each business state slice contains all spatial voxels belonging to the same business branch or state and their connections. A business process branch in the business process knowledge graph refers to a path starting from the same initial state and reaching different termination states through different sequences of key operation nodes. A state type refers to the classification of the operation result state of key operation nodes, such as "ready," "in progress," "completed," or "abnormal." In the warehousing example, there may be two different business process branches: "normal order picking" and "urgent order picking." In the manufacturing example, there may be two business process branches: "normal assembly process" and "rework process." Based on these branches or state types, the spatiotemporal semantic field is divided into multiple logically non-overlapping business state slices. Each business state slice has a unique business state slice number, which can be a classification code. Optionally, a spatial voxel may belong to multiple business state slices simultaneously; for example, a spatial voxel may be associated with parallel branches simultaneously in the business process knowledge graph. The spatial voxels within the business state slice include not only the voxels themselves, but also other voxels connected by virtual connection edges, thus preserving the local business logic topology.
[0032] In practical implementation, the index relationship between spatial grids, time slices, and business state slices is established by assigning a triplet identifier to each spatial voxel in the spatiotemporal semantic field. Any spatial voxel can be uniquely identified by its spatial grid number, time slice number, and business state slice number; these three numbers form a composite key. The spatial grid number determines the spatial location of the spatial voxel, the time slice number determines its temporal segment, and the business state slice number determines its business logic. Refer to Table 1, an example table illustrating the mapping relationship between spatial voxels and their three dimensions.
[0033] Table 1: Examples of Spatial Voxel Multidimensional Index Representation In Table 1, the unique ID of a spatial voxel is its internal identifier within the system, and the spatial grid number (G) is determined by the formula... Calculations show that the time slice number (T) is an integer sequence number, and the business state slice number (B) is an encoding representing different business process branches or state types. This triplet index enables efficient querying and retrieval of arbitrary dimensional subsets in the spatiotemporal semantic field, such as querying all spatial voxels belonging to a specific business process within a specific time period for a specific spatial region. Optionally, the index relationships can be stored in a relational database or a spatiotemporal database for fast join queries. In some embodiments, a cross-reference table is also maintained between the spatial grids, time slices, and business state slices, recording which spatial grids appear in which time slices and business state slices to accelerate intersection operations. The final result of multidimensional spatial partitioning is the construction of a multidimensional reconstruction model of the scene. This model consists of an interrelated set of spatial grids, a sequence of time slices, and a set of business state slices, supporting the slicing and reconstruction of the scene from any or a combination of spatial, temporal, and business logic dimensions.
[0034] In one embodiment of the present invention, the process of acquiring new business data streams and visual perception data in real time is a continuous data stream monitoring and access process. In a warehousing scenario, new business data streams may come from newly generated order outbound records in the warehouse management system, and visual perception data may come from newly collected panoramic images of the shelf area by mobile robots. In a manufacturing scenario, new business data streams may come from newly reported inspection results from the production line's off-line workstations, and visual perception data may come from newly captured assembly images by fixed cameras. Synchronizing the new data to the multidimensional reconstruction model first requires real-time processing of the newly acquired visual perception data. The newly acquired visual perception data is updated with the geometric and semantic information of the affected areas in the semantically enhanced fused 3D point cloud through the deployed panoramic semantic segmentation network and point cloud registration process. For example, the 3D point cloud coordinates and semantic labels of the area where the newly stocked goods are located are added or modified.
[0035] In practical implementation, parsing new business data streams and updating the business process knowledge graph involves analyzing structured new data. Parsing the new business data streams identifies new key operation nodes and state transitions. For example, in the warehousing process, it identifies the new key operation node "transfer of product SKU456" and its state transition from "started" to "completed." In the manufacturing process, it identifies the new key operation node "final performance test" and its state transition from "testing" to "passed." Based on the updated business process knowledge graph and the updated semantically enhanced fused 3D point cloud, the affected spatial operation region and time-varying attribute vectors are recalculated. The process calls are recalculated using the same algorithm as the initial construction but only apply to the local regions where the data has changed. An example formula describes the update relationship of the affected voxel set: ; in: This represents the set of three-dimensional points that need to have their time-varying attribute vectors recalculated. This represents a local point cloud region in the updated semantically enhanced fused 3D point cloud that has undergone geometric or semantic changes. This represents the set of key operation nodes that have been added or modified in the updated business process knowledge graph. This represents the mapping function between the spatial operation region and the critical operation node. This is understandable. It is a selection function whose output is the 3D spatial points whose time-varying attribute vectors need to be recalculated due to changes in the point cloud or changes in associated critical operation nodes. Based on the recalculation results, the 3D spatiotemporal voxel field, spatiotemporal semantic field, and the corresponding spatial mesh, time slices, and business state slices are updated incrementally. Incremental updates mean that only the affected voxel field is updated. Instead of reconstructing the entire model, the affected spatial voxels and their associated multidimensional subdivision results are recalculated for attributes and indexed for updates.
[0036] In some embodiments, the process of generating a view in response to a reconstruction instruction and by extracting data from the model begins with instruction parsing. This parsing involves resolving the spatial extent parameters, time span parameters, and business logic constraints contained in the reconstruction instruction. Spatial extent parameters are typically given as 3D spatial bounding boxes or geographic coordinate ranges; time span parameters are given as start and end timestamps; and business logic constraints are given as business process branch names or state type names. Based on the spatial extent parameters, corresponding spatial grids are selected by comparing the spatial bounding boxes of the spatial grids with the input spatial extent parameters to determine if they intersect. Based on the time span parameters, corresponding time slice sequences are selected by comparing the time periods represented by the time slices with the input time span parameters to determine if they overlap. Based on the business logic constraints, corresponding business state slices are selected by matching the business state slice number with the input business process branch or state type name.
[0037] In practical implementation, the selected spatial mesh, time slice sequence, and business state slice are subjected to an intersection operation to obtain the target voxel set. The intersection operation is based on a fast set intersection based on the triplet index of spatial mesh number, time slice number, and business state slice number maintained for each spatial voxel. The spatial voxel set that satisfies all constraints and its complete attribute data are extracted. The complete attribute data includes the geometric position, semantic label, dynamic classification, node type, and virtual connection relationship with upstream and downstream voxels of the spatial voxel. It can be understood that the intersection operation ensures that the final result accurately corresponds to the intersection of the spatial, temporal, and business logic queries of the user. Optionally, the result of the intersection operation can be cached in the query index to improve the response speed to similar reconstruction instructions. The spatial voxel set and its complete attribute data are converted into a scene graph data structure that can be recognized by the 3D rendering engine. The conversion process includes converting the geometric attributes of each spatial voxel into triangular patches or voxel meshes, converting dynamic attributes into material or transparency parameters that change over time, and converting business logic attributes into object labels or connectors. The 3D rendering engine generates an interactive scene reconstruction view that integrates spatial geometry, temporal evolution, and business logic. In the photovoltaic power plant example view, users can observe a 3D dynamic reproduction of the inspection personnel's work path, equipment detection status, and changes within a specific photovoltaic array area over a specific time period, with abnormal equipment highlighted. In the substation example view, users can observe a 3D dynamic reproduction of the switching operation process and equipment status changes within a specific interval and time period, with the areas and equipment involved in the current operation step highlighted. Furthermore, the color or highlighting effect of objects in the view reflects their corresponding business status; for example, green indicates normal, red indicates alarm, and blue indicates operation in progress.
[0038] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence, characterized in that, include: The collected panoramic image sequences, multi-source business data streams, and environmental spatial point clouds are fused to construct a semantically enhanced fused 3D point cloud and business process knowledge graph. Guided by the business process knowledge graph, spatial operation regions corresponding to key operation nodes are marked in the semantically enhanced fused 3D point cloud. The spatial operation regions are defined by a series of 3D spatial point clusters and their associated entity object pixel regions. Dynamic spatiotemporal encoding is performed on the fused 3D point cloud with semantic enhancement after labeling, and the timestamps of key operation nodes and business process status are encoded into time-varying attribute vectors of corresponding 3D spatial points. Based on the time-varying attribute vector, the semantically enhanced fused 3D point cloud is reconstructed in a spatiotemporal consistency, and a dynamically changing 3D spatiotemporal voxel field is generated in the 3D space with the spatial operation region as the core. The topological relationship of the business process knowledge graph is embedded in the three-dimensional spatiotemporal voxel field. The node type and edge connection relationship of the business process knowledge graph are attached to each three-dimensional spatiotemporal voxel to form a spatiotemporal semantic field with business logic semantics. The spatiotemporal semantic field is divided into multiple dimensions, and interrelated spatial grids, time slices and business state slices are generated along the spatial dimension, time dimension and business logic dimension to construct a multidimensional reconstruction model of the scene.
2. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 1, characterized in that, The construction of the semantically enhanced fused 3D point cloud and business process knowledge graph includes: The system acquires panoramic image sequences, multi-source business data streams, and environmental spatial point clouds of the target scene. The panoramic image sequences include visual frames with position and pose information, the multi-source business data streams include business operation records and business process status, and the environmental spatial point clouds include three-dimensional spatial points and their reflection intensity. Pixel-level semantic segmentation is performed on the panoramic image sequence to extract the pixel regions of entity objects in the image, and the image coordinates of the pixel regions of the entity objects are registered and associated with the corresponding three-dimensional points of the environmental spatial point cloud to generate a semantically enhanced fused three-dimensional point cloud. Business process modeling is performed on multi-source business data streams to identify key operation nodes and state transition paths in business operation records and to construct a business process knowledge graph. The step of performing pixel-level semantic segmentation on the panoramic image sequence, extracting the pixel regions of entity objects in the image, and registering and associating the image coordinates of the entity object pixel regions with the corresponding 3D points of the environmental spatial point cloud includes: The panoramic image sequence is input into a pre-trained panoramic semantic segmentation network to generate a pixel-level semantic label map for each visual frame. The pixel connected regions belonging to a predetermined category list are extracted from the semantic label map as the pixel regions of entity objects. Based on the position and pose information carried by the panoramic image sequence, the camera projection matrix corresponding to each visual frame is calculated. Using the camera projection matrix, the pixel coordinates of each entity object's pixel region are back-projected to three-dimensional space, and the set of three-dimensional spatial points that fall into the back-projected view frustum is found in the environmental spatial point cloud. The found set of three-dimensional spatial points is associated with the pixel region of the entity object, and the semantic labels of the pixel region of the entity object are assigned to the corresponding three-dimensional spatial points, thereby completing the registration and semantic enhancement.
3. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 2, characterized in that, The process of modeling business processes from multiple sources of business data streams, identifying key operation nodes and state transition paths in business operation records, and constructing a business process knowledge graph includes: Analyze each business operation record in the multi-source business data stream to extract the operation subject, operation object, operation type, operation timestamp, and operation result status; Based on the order of operation timestamps, multiple business operation records in the same business process are arranged into an operation sequence; In the operation sequence, identify the business operation records that trigger a fundamental change in the state of the business process, define them as key operation nodes, and define the state changes between adjacent key operation nodes as state transition paths. Using key operation nodes as graph nodes and state transition paths as directed edges, a business process knowledge graph representing the complete business process is constructed.
4. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 3, characterized in that, The process of marking spatial operation regions corresponding to key operation nodes in a semantically enhanced fused 3D point cloud, guided by a business process knowledge graph, includes: Traverse each key operation node in the business process knowledge graph and parse the operation object from the corresponding business operation record; In the semantically enhanced fused 3D point cloud, find 3D spatial point clusters whose semantic labels match the operation object; Using the centroid of the found three-dimensional point cluster as the center, a spatial region is grown outward in three-dimensional space until the density of the boundary point cloud is lower than a set threshold. The defined spatial range and the included three-dimensional point cluster are the spatial operation area corresponding to the key operation node.
5. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 4, characterized in that, The dynamic spatiotemporal encoding of the fused 3D point cloud with semantic enhancement after labeling encodes the timestamps of key operation nodes and business process status into time-varying attribute vectors of corresponding 3D spatial points, including: For each 3D spatial point in the semantically enhanced fused 3D point cloud, determine its corresponding spatial operation region; For a 3D spatial point belonging to any spatial operation region, obtain the timestamp and operation result status of the key operation node corresponding to its spatial operation region. Based on the timestamp, the normalized time offset of the three-dimensional spatial points is calculated, and the operation result state is mapped to a state encoding vector of a preset dimension. The normalized time offset is concatenated with the state encoding vector to generate the time-varying attribute vector of the three-dimensional spatial point.
6. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 5, characterized in that, The method of reconstructing a spatiotemporally consistent structure of semantically enhanced fused 3D point clouds based on time-varying attribute vectors, generating a dynamically changing 3D spatiotemporal voxel field with the spatial operation region as the core in 3D space, includes: Import the complete point cloud data, which includes 3D spatial points, semantic labels, and time-varying attribute vectors, into the spatiotemporal voxelization engine. In the spatiotemporal voxelization engine, three-dimensional space is discretized into spatial voxels along the spatial axis and into time frames along the time axis. For each spatial voxel, aggregate the time-varying attribute vectors of all three-dimensional spatial points within it, and calculate its aggregated vector sequence that changes with time frames. Based on the evolution pattern of the aggregated vector sequence in the time dimension, spatial voxels are classified into static voxels, periodic dynamic voxels, and aperiodic dynamic voxels, and their dynamic characteristics are recorded. All spatial voxels and their dynamic characteristics constitute a three-dimensional spatiotemporal voxel field.
7. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 6, characterized in that, The embedding of the topological relationship of the business process knowledge graph in the three-dimensional spatiotemporal voxel field, which involves attaching node types and edge connections from the business process knowledge graph to each three-dimensional spatiotemporal voxel, includes: Establish a mapping relationship between the set of spatial voxels contained in the spatial operation region in the three-dimensional spatiotemporal voxel field and the key operation nodes in the business process knowledge graph; Assign the node type of the key operation node as an attribute to all spatial voxels mapped to it; In a three-dimensional spatiotemporal voxel field, if two spatial voxel sets are mapped to two key operation nodes in a business process knowledge graph that are directly connected by directed edges, then virtual connection edges with the same direction and type are established between these two spatial voxel sets. After completing the connection of all voxels and sets, a spatiotemporal semantic field is formed that simultaneously contains spatial geometry, temporal dynamics, and business logic.
8. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 7, characterized in that, The process of multi-dimensionally partitioning the spatiotemporal semantic field and generating interrelated spatial grids, time slices, and business state slices along spatial, temporal, and business logic dimensions includes: Along the spatial dimension, the spatiotemporal semantic field is divided into regular spatial grids, and each spatial grid contains one or more spatial voxels and all their attributes; Along the time dimension, the spatiotemporal semantic field is sliced at fixed time intervals, and each time slice contains a snapshot of the dynamic state of all spatial voxels within the corresponding time period. Along the business logic dimension, based on different business process branches or state types in the business process knowledge graph, the spatiotemporal semantic field is logically divided. Each business state slice contains all spatial voxels and connection relationships belonging to the same business branch or state. Establish an index relationship between spatial grids, time slices, and service status slices, so that any spatial voxel can be uniquely identified by its spatial grid number, time slice number, and service status slice number.
9. The method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 8, characterized in that, Also includes: Real-time acquisition of new business data streams and visual perception data, synchronization to the multidimensional reconstruction model, and updating of the internal states of the corresponding spatial grid, time slice, and business status slice; In response to the reconstruction command, the multidimensional spatial partitioning results under the specified spatial range, time span and business logic constraints are extracted from the multidimensional reconstruction model, and the 3D rendering engine is driven to generate a scene reconstruction view that is dynamically synchronized with the business flow. The real-time acquisition of new business data streams and visual perception data, and their synchronization to the multi-dimensional reconstruction model, updating the internal states of the corresponding spatial grid, time slice, and business state slice, includes: The newly acquired visual perception data is processed in real time to update the geometric and semantic information of the affected areas in the semantically enhanced fused 3D point cloud. Analyze the newly added business data flow, identify new key operation nodes and state transitions, and update the business process knowledge graph; Based on the updated business process knowledge graph and the semantically enhanced fused 3D point cloud, the affected spatial operation area and time-varying attribute vector are recalculated; Based on the recalculation results, the contents of the three-dimensional spatiotemporal voxel field, spatiotemporal semantic field, and corresponding spatial grid, time slice, and business status slice are updated incrementally.
10. A method for multi-dimensional scene reconstruction integrating data operation and spatial intelligence according to claim 9, characterized in that, The process of responding to a reconstruction command by extracting multidimensional spatial partitioning results from the multidimensional reconstruction model under specified spatial range, time span, and business logic constraints, and driving the 3D rendering engine to generate a scene reconstruction view dynamically synchronized with the business flow, includes: Parse the spatial range parameters, time span parameters, and business logic constraints contained in the reconstruction instruction; The corresponding spatial grid is selected based on the spatial range parameter, the corresponding time slice sequence is selected based on the time span parameter, and the corresponding business state slice is selected based on the business logic constraints. The selected spatial grid, time slice sequence and business status slice are intersected to obtain a set of spatial voxels that satisfy all constraints and their complete attribute data. The spatial voxel set and its complete attribute data are converted into a scene graph data structure that can be recognized by the 3D rendering engine, and the 3D rendering engine is driven to render, generating an interactive scene reconstruction view that integrates spatial geometry, temporal evolution and business logic.