Three-dimensional space data analysis method and system based on deep learning

Through a deep learning-based dynamic neural radiation field architecture, combining convolutional symbol distance field and spatiotemporal attention network to optimize static and dynamic object models, deploy lightweight models and use the federated learning framework to solve the computational complexity and accuracy problems of traditional three-dimensional spatial data analysis methods, and achieve efficient and accurate three-dimensional data processing.

CN120524831AInactive Publication Date: 2025-08-22ZHONGBO INFORMATION TECH RES INST CO LTD

Patent Information

Application Number
CN202511014450.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional three-dimensional spatial data analysis methods have high computational complexity and low accuracy when processing large-scale data, making it difficult to cope with complex and dynamic environments, especially in object recognition and object detection.

Method used

The space-time separation dynamic neural radiation field (DyNeRF) architecture based on deep learning is adopted to encode the static scene basis model through convolutional symbol distance field (ConvSDF) and multi-layer perceptron (MLP), and combine the space-time attention network and rigid body motion equation to optimize the trajectory of dynamic object, and deploy a lightweight model at the edge computing node. The federated learning framework aggregates feature difference data in the cloud to achieve efficient and secure three-dimensional data processing.

Benefits of technology

The efficiency and accuracy of three-dimensional data processing are improved, and the applicability is enhanced, especially the accuracy of object recognition and object detection in complex environments, solving the calculation complexity and accuracy problems of traditional methods, and achieving high-precision and low-latency three-dimensional spatial analysis capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524831A_ABST
    Figure CN120524831A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional space data analysis method and system based on deep learning, and relates to the field of computer vision, and the method comprises the steps: carrying out the modeling of a static scene fundamental model and a dynamic object motion track through a space-time separated dynamic nerve radiation field; a lightweight dynamic neural radiation field model is deployed at an edge computing node, multi-modal sensor data are processed in real time, local three-dimensional scene representation is generated, rendering and prediction computing of a neural radiation field are executed on the edge node, and the implicit feature difference quantity of scene change is uploaded to a cloud; and the cloud end aggregates feature difference data of multiple edge nodes through a federated learning framework, dynamically updates a global scene priori knowledge base and issues the global scene priori knowledge base to the edge nodes. The method can improve the processing efficiency, precision and applicability of the three-dimensional data, and is especially suitable for carrying out tasks such as object recognition, target detection and semantic segmentation in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a three-dimensional spatial data analysis method and system based on deep learning. Background Art

[0002] With the continuous development of artificial intelligence and sensor technology, three-dimensional spatial data is increasingly being used in various fields, such as autonomous driving, robot navigation, virtual reality, and augmented reality. However, traditional three-dimensional spatial data analysis methods face a series of challenges, such as the computational complexity of processing large-scale three-dimensional data, data noise interference, and object recognition accuracy.

[0003] Traditional point cloud data processing methods mostly rely on hand-crafted features or rule-based algorithms, often unable to cope with complex and dynamic environments. Especially when faced with large-scale data, they suffer from slow processing speeds and low accuracy. Deep learning technology, particularly convolutional neural networks (CNNs) and graph neural networks (GNNs), has achieved significant breakthroughs in computer vision and pattern recognition. It can automatically learn useful features from data and has stronger expressive power.

[0004] Deep learning-based 3D data analysis methods, particularly in the analysis of point cloud data and 3D models, can improve the accuracy and efficiency of data processing by automatically learning features. Therefore, this paper proposes a deep learning-based 3D data analysis method and system, aiming to improve the accuracy, speed, and applicability of 3D data analysis and address the shortcomings of existing technologies. Summary of the Invention

[0005] In light of this, the present invention aims to propose a deep learning-based 3D spatial data analysis method and system, enabling efficient processing, analysis, and application of 3D spatial data. This method improves the processing efficiency, accuracy, and applicability of 3D data, and is particularly suitable for tasks such as object recognition, target detection, and semantic segmentation in complex environments.

[0006] To achieve the above object, the present invention provides the following technical solutions: Based on the above objectives, in a first aspect, the present invention provides a three-dimensional spatial data analysis method based on deep learning, comprising the following steps: Static scene primitives and dynamic object motion trajectories are modeled using spatiotemporally separated dynamic neural radiance fields (DyNeRFs). The static scene primitives are encoded using convolutional signed distance fields (ConvSDFs), and the dynamic trajectories are jointly optimized with the rigid body motion equations using a spatiotemporal attention network. Deploy a lightweight dynamic neural radiation field model on edge computing nodes to process multimodal sensor data in real time and generate a local 3D scene representation. Rendering and prediction calculations of the neural radiation field are performed on the edge nodes, and implicit feature differences of scene changes are uploaded to the cloud. The cloud aggregates feature difference data from multiple edge nodes through a federated learning framework, dynamically updates the global scenario prior knowledge base, and sends it to the edge nodes.

[0007] As a further solution of the present invention, modeling the static scene primitives and dynamic object motion trajectories through a spatiotemporally separated dynamic neural radiation field includes the following steps: The static scene base model branch uses a multi-layer perceptron (MLP) to encode global illumination, material properties, and spatial structure information, and outputs a mapping between signed distance field and RGB color values; The dynamic trajectory branch predicts the motion trajectory of dynamic objects through a spatiotemporal attention network and calculates the rigid body motion constraint loss in combination with a differentiable physics engine; A spatiotemporal mask is generated based on the optical flow data of adjacent frames, and the output results of the static base model of the scene and the dynamic trajectory branch are dynamically fused to align the three-dimensional features across frames.

[0008] As a further solution of the present invention, when modeling the static scene base model and the dynamic object motion trajectory through the spatiotemporally separated dynamic neural radiation field, it also includes: Perform convolutional signed distance field encoding on static scenes to obtain a representation of the geometric structure; Perform spatiotemporal modeling of dynamic object motion and use spatiotemporal attention mechanism to model object motion; Based on the rigid body motion equation, the object's motion trajectory is jointly optimized in combination with the spatiotemporal attention network.

[0009] As a further solution of the present invention, modeling the static scene archetype also includes: extracting the features of lighting and material of the static scene; encoding the spatial structure information of the scene using a multi-layer perceptron model; and outputting a signed distance field (SDF) and mapping it to RGB color values ​​to form the static scene archetype.

[0010] As a further solution of the present invention, modeling the motion trajectory of a dynamic object also includes: using a spatiotemporal attention mechanism to model the motion trajectory of the object, and optimizing the trajectory in combination with the rigid body motion equation to ensure the physical consistency of the motion state.

[0011] As a further solution of the present invention, when a lightweight dynamic neural radiation field model is deployed on an edge computing node, lidar and cameras are used as sensors to obtain raw data, and the data is processed by the lightweight dynamic neural radiation field model to generate a local three-dimensional scene, extract the implicit features of the scene changes and upload them to the cloud.

[0012] As a further solution of the present invention, a lightweight dynamic neural radiation field model is deployed at the edge computing node. The lightweight dynamic neural radiation field model of the edge node is implemented in the following manner: Perform channel pruning on the original dynamic neural radiation field model, remove redundant convolutional layers in the static branch base model, and retain all parameters of the dynamic attention module; Knowledge distillation is used to use the output of the full model on the cloud as a teacher signal to guide the feature alignment of the edge model.

[0013] As a further solution of the present invention, before uploading the implicit feature difference of the scene change to the cloud, the method for generating the implicit feature difference includes the following steps: The implicit feature vector of the dynamic scene is extracted using a variational autoencoder to generate compressed feature differences. Compress the feature vector into a fixed-length binary stream through hash coding; Homomorphic encryption is used to encrypt the binary stream and then transmit it to the cloud.

[0014] As a further solution of the present invention, homomorphic encryption is used to encrypt the binary stream and then transmit it to the cloud, including: using quantum key distribution to generate an encryption key, encrypting the compressed implicit feature data, so that data privacy is protected during the transmission process between the cloud and the edge.

[0015] As a further solution of the present invention, the cloud aggregates feature difference data of multiple edge nodes through a federated learning framework, including the following steps: Store and update feature difference data collected from multiple edge computing nodes, dynamically optimize the prior knowledge of the global scene, and form a global scene knowledge base; By regularly aggregating feature difference data from edge nodes, the global model is optimized in the cloud, and the updated global scene knowledge is sent to the edge nodes.

[0016] In a second aspect, the present invention further provides a three-dimensional spatial data analysis system based on deep learning, comprising the following components: Multimodal acquisition module: used to synchronously acquire 3D spatial data, including lidar, multi-view cameras, and an inertial measurement unit (IMU). The lidar generates point cloud data, the camera captures multi-view images, and the IMU provides device motion parameters. Edge computing unit: Deployed on the sensor terminal side, it includes a reconfigurable computing array (FPGA) and a lightweight dynamic neural radiation field (LiteDyNeRF) model for real-time processing of multimodal data and generating a neural radiation field representation of the local 3D scene; Cloud-based optimization platform: This includes a federated learning server and a global scenario knowledge base, which aggregates feature difference data from multiple edge nodes and dynamically updates the global model. Encrypted Communication Interface: This uses an encrypted channel based on quantum key distribution for feature data transmission and model update instruction issuance between the edge computing unit and the cloud optimization platform as a further solution of the present invention.

[0017] As a further solution of the present invention, the edge computing unit includes: Point cloud preprocessing module: downsamples, denoises, and voxelizes the raw LiDAR point cloud to generate a sparse voxel representation; Dynamic Detection Accelerator: This uses a sparse convolutional neural network to identify dynamic object regions in real time and triggers a spatiotemporal attention network for trajectory prediction. Reconfigurable Rendering Engine: Supports parallel computation of ray casting and voxel rendering for neural radiation fields, achieving real-time rendering at 35ms / frame through FPGA hardware acceleration.

[0018] As a further solution of the present invention, the cloud optimization platform includes: Federated Learning Server: This server protects feature difference data uploaded by edge nodes through differential privacy technology and performs distributed gradient aggregation of model parameters. Global scene knowledge base: This base uses a graph neural network (GNN) to model the topological relationships between objects in the scene and uses a temporal memory network (TMC) to store implicit feature sequences of historical scenes. Model Update Distribution Module: A gradient masking strategy is designed to incrementally update only the model parameters related to dynamic objects, and the updates are distributed to edge nodes through encrypted channels.

[0019] As a further solution of the present invention, the lightweight dynamic neural radiation field model is implemented in the following manner: Model Compression Module: This module performs channel pruning on the original DyNeRF model, removes redundant 3D convolutional layers in the static branch, and retains the complete parameters of the dynamic attention module. Knowledge Distillation Module: The rendered output of the complete model on the cloud is used as the teacher signal, and the feature alignment loss function is used to guide the edge model to learn high-dimensional feature representations.

[0020] As a further embodiment of the present invention, the encrypted communication interface includes: Feature Difference Encoder: uses a variational autoencoder (VAE) to encode dynamic scene changes into a 0.5MB implicit feature vector; Quantum encryption module: Generates dynamic encryption keys through the quantum key distribution protocol, and transmits feature data after homomorphic encryption; ‌Decryption Verification Unit‌: Decrypts and verifies the integrity of the received encrypted data through the Trusted Execution Environment (TEE) in the cloud.

[0021] As a further solution of the present invention, the method for constructing the global scene knowledge base includes: Scene topology modeling: Analyze the spatial constraints between dynamic objects and static scenes based on graph neural networks (GNNs); Temporal feature storage: Storing implicit feature sequences of historical scenarios through a memory-augmented network. ‌Similarity retrieval algorithm‌: Design a fast retrieval mechanism based on locality sensitive hashing (LSH) to match the current scene with the prior patterns in the knowledge base.

[0022] As a further solution of the present invention, the collaborative method of the multimodal acquisition module includes: ‌Spatiotemporal Synchronization Unit‌: Aligns the acquisition timestamps of the LiDAR and camera through hardware trigger signals, with an error of less than 1ms; Cross-modal calibration module: Dynamically calibrates the extrinsic parameter matrices of the lidar and camera based on a checkerboard calibration plate; ‌Data Fusion Interface‌: A double-buffered queue mechanism is used to balance the processing rate differences between point cloud and image data.

[0023] In another aspect of the present invention, a computer device is provided, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, any one of the above-mentioned three-dimensional spatial data analysis methods based on deep learning according to the present invention is executed.

[0024] In another aspect of the present invention, a computer-readable storage medium is provided, which stores computer program instructions, and when the computer program instructions are executed, any one of the above-mentioned three-dimensional spatial data analysis methods based on deep learning according to the present invention is implemented.

[0025] Compared with the existing technology, the three-dimensional spatial data analysis method and system based on deep learning proposed in this invention have the following beneficial effects: 1. A dynamic neural radiation field architecture with spatiotemporal separation addresses geometric and motion distortion in dynamic scene modeling. In the static scene construction phase, the Convolutional Signed Distance Field (ConvSDF) and Multilayer Perceptron (MLP) collaborate to accurately reproduce high-precision geometric structures and material optical properties. The dynamic trajectory branch captures nonlinear motion patterns through a spatiotemporal attention mechanism and introduces rigid body motion constraints to ensure that predicted trajectories conform to physical laws, significantly reducing trajectory prediction errors in scenarios such as sharp vehicle turns.

[0026] 2. At the edge computing layer, lightweight model deployment and feature differentiation mechanisms overcome computing power and bandwidth limitations. Channel pruning removes redundant parameters in static branches. Knowledge distillation technology migrates the complex scene understanding capabilities of cloud models to the edge, enabling vehicles to maintain high-precision perception in complex environments and improving the efficiency and accuracy of three-dimensional data processing. At the same time, a variational autoencoder extracts implicit feature differences in scene changes, which are then uploaded after hash compression and quantum encryption, resolving security risks in edge-cloud data exchange and the conflict between data security and transmission efficiency.

[0027] 3. Cloud-based federated learning and the knowledge base co-evolution mechanism form a closed loop. By aggregating multi-node feature differences, the graph neural network constructs the topological relationship between dynamic objects and static scenes, and the temporal memory network stores the evolution patterns of historical scenes. The optimized global knowledge is refined into incremental parameters through gradient masking and distributed to edge nodes, realizing the adaptive evolution of scene prior knowledge. This enables devices or vehicles that have not experienced specific scenarios to instantly obtain advanced decision-making capabilities, improving the applicability of three-dimensional data processing. This invention can provide high-precision, low-latency three-dimensional spatial analysis capabilities for fields such as autonomous driving and industrial inspection.

[0028] These and other aspects of the present application will be more clearly understood in the following description of the embodiments. It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for the exemplary embodiments or related technical descriptions. The drawings are used to provide a further understanding of the present invention and constitute part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the drawings: Figure 1 This is a flowchart of a three-dimensional spatial data analysis method based on deep learning in an embodiment of the present invention.

[0030] Figure 2 This is a flowchart of modeling a dynamic neural radiation field separated by time and space in a three-dimensional spatial data analysis method based on deep learning in an embodiment of the present invention.

[0031] Figure 3 This is a flowchart of generating implicit feature difference amounts in a three-dimensional spatial data analysis method based on deep learning in an embodiment of the present invention.

[0032] Figure 4This is a flowchart of aggregating feature difference data of multiple edge nodes in the cloud through a federated learning framework in a three-dimensional spatial data analysis method based on deep learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] Below, the present application is further described in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0034] To make the purpose, technical solutions and advantages of the present invention more clearly understood, the following is a further detailed description of the embodiments of the present invention in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0035] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are intended to distinguish two non-identical entities or non-identical parameters with the same name. Therefore, "first" and "second" are used for convenience of expression only and should not be understood as limitations on the embodiments of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, other steps or units inherent to a process, method, system, product, or device that includes a series of steps or units.

[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0038] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0039] To improve the accuracy, speed, and applicability of 3D spatial data analysis, this paper proposes a 3D spatial data analysis method and system based on deep learning. This method effectively processes, analyzes, and applies 3D spatial data through deep learning. This method improves the processing efficiency, accuracy, and applicability of 3D data and is particularly suitable for tasks such as object recognition, target detection, and semantic segmentation in complex environments.

[0040] See also Figure 1 As shown, an embodiment of the present invention provides a three-dimensional spatial data analysis method based on deep learning, which includes the following steps: Step S10: Modeling the static scene primitives and dynamic object motion trajectories through spatiotemporal separation of dynamic neural radiance fields (DyNeRFs). The static scene primitives are encoded using convolutional signed distance fields (ConvSDFs), and the dynamic trajectories are jointly optimized with the rigid body motion equations through a spatiotemporal attention network.

[0041] Traditional Neural Radiance Field (NeRF) can produce "ghosting" artifacts when processing dynamic objects. For example, in autonomous driving scenarios, moving vehicles can leave afterimages when rendering consecutive frames. Dynamic Neural Radiance Field (DyNeRF) is a 3D scene representation technology that integrates spatiotemporal characteristics. Unlike traditional NeRF, which only processes static scenes, DyNeRF uses a dual-branch architecture to separately model static backgrounds and dynamic objects.

[0042] In this embodiment, processing 3D spatial data first requires constructing a dynamic neural radiation field, decoupling the complex environment into two parallel subsystems: a static base model branch focused on reconstructing the geometric structure and material properties of fixed background elements (such as roads and buildings), and a dynamic trajectory branch responsible for capturing the spatiotemporal motion patterns of moving objects (such as vehicles and pedestrians). These two subsystems work together through physical constraints and a cross-frame alignment mechanism to form a unified 3D scene representation.

[0043] In this step, see Figure 2 As shown, the modeling of the static scene base model and the dynamic object motion trajectory through the dynamic neural radiation field separated in time and space includes the following steps: Step S101: The static scene base model branch uses a multi-layer perceptron to encode global illumination, material properties, and spatial structure information, and outputs a mapping between a signed distance field and RGB color values; In this embodiment, the construction of the static scene base model begins with multimodal three-dimensional spatial data acquisition, including: lidar generates point cloud data, multi-view cameras collect multi-view images, and inertial measurement units (IMUs) provide device motion parameters.

[0044] In one embodiment, the collected point cloud data is preprocessed. First, downsampling is achieved through spatial grid filtering, and the discrete point cloud is divided into fixed voxel units. The core position points of each unit are retained to maintain key geometric features such as road contours. Dynamic noise filtering is then performed to identify and remove outlier noise clusters caused by rain and snow reflections or sensor errors based on the local density distribution of the point cloud. Finally, a structured transformation is performed to map the optimized point cloud to a three-dimensional grid space, and a sparse voxel representation is constructed through an octree indexing mechanism.

[0045] Based on these multi-source inputs, the system first performs convolutional signed distance field encoding on the static scene to obtain a geometric feature vector representing the geometric structure. The Convolutional Signed Distance Field (ConvSDF) is an enhanced implementation of the Signed Distance Field (SDF). It learns spatial continuity features through a three-dimensional convolutional network, addressing the geometric distortion issues inherent in traditional discrete SDF representations and making it more suitable for complex surface reconstruction. Optionally, the three-dimensional convolutional network can utilize a three-layer 3D convolutional network with a convolution kernel size of 3×3×3, a stride of 1, a LeakyReLU activation function, and 64 output channels.

[0046] The geometric feature vector is used as input data for a multilayer perceptron (MLP) to extract lighting and material features from the static scene. The MLP is then used to encode the scene's spatial structure information, outputting a signed distance field (SDF) mapped to RGB color values. This transforms the geometric features into renderable physical properties, forming a static scene primitive. The signed distance field value represents the geometric distance from the spatial point to the nearest object surface (positive values ​​represent the space outside the object, negative values ​​represent the space inside the object), and the RGB color value represents the visual performance characteristics of the location under specific lighting conditions. Optionally, the MLP includes four fully connected layers with a hidden layer dimension of 256.

[0047] The static scene model construction process uses convolutional signed distance fields and multi-layer perceptrons to achieve collaborative encoding of the scene's geometric and visual attributes. In autonomous driving applications, when processing a sample point on the road surface, the multi-layer perceptron first analyzes the spatial distribution characteristics of its three-dimensional coordinates. The network's hidden layer then uses feature transformation to identify the point as 0.5 cm below the asphalt pavement surface. Simultaneously, the grayscale value under current sunlight conditions is calculated based on a material reflectance model. This mapping process essentially establishes a mathematical connection between spatial topology and the laws of optical physics. Its technical advantage lies in eliminating geometric discontinuities caused by discrete sampling, resulting in smooth transitions on the road surface, adapting to changes in ambient lighting, and maintaining the physical accuracy of material reflectance properties, providing high-precision spatial data for navigation decisions.

[0048] Those skilled in the art will appreciate that this implementation, through an end-to-end neural network architecture, unifies the traditionally separate geometric modeling and material rendering processes into a single computational framework, ensuring spatial accuracy while also meeting the requirements of physical realism in complex lighting environments. In practice, neural network parameters can be dynamically adjusted based on scene complexity to achieve an optimal balance between computational efficiency and model accuracy.

[0049] Step S102: The dynamic trajectory branch predicts the motion trajectory of the dynamic object through the spatiotemporal attention network and calculates the rigid body motion constraint loss in combination with the differentiable physics engine; The dynamic trajectory branch is primarily responsible for predicting the trajectory of dynamic objects. The trajectory of dynamic objects is often complex and uncertain, requiring a powerful model to capture their motion patterns. A spatiotemporal attention network is an effective tool for this purpose. It simultaneously considers information in both temporal and spatial dimensions. By calculating the correlation between different time steps and spatial positions, it automatically focuses on information that is more important to the current prediction task, thereby more accurately predicting the object's trajectory. In this embodiment, dynamic trajectory prediction is based on a spatiotemporal attention network architecture to model moving objects. This network uses an adaptive weight allocation mechanism to capture the state dependencies of dynamic objects in a continuous time series. In specific implementation, the center coordinates of the target object's three-dimensional bounding box in a historical frame sequence are input into the network. The network's hidden layer uses a multi-head attention mechanism to calculate the displacement correlation between different time steps and establish a feature representation of the cross-frame trajectory. Optionally, the number of attention heads is eight. Taking a road vehicle as an example, the network first extracts the displacement vector sequence of the previous five frames. Using the self-attention weight matrix, it identifies key motion feature nodes (such as the sudden acceleration from the second to the third frame). Finally, it outputs the predicted coordinates of the trajectory for the next three frames.

[0050] In order to make the predicted motion trajectory more consistent with the laws of physics, this embodiment further introduces a differentiable physics engine to calculate the rigid body motion constraint loss. The differentiable physics engine can integrate the physical laws into the model and can calculate the gradient information so that the model can be optimized through the back-propagation algorithm. The differentiable physics engine models and simulates the motion of objects according to the physical laws of rigid body motion, such as Newton's laws of motion, the law of conservation of momentum, etc. When predicting the motion trajectory of an object, it is necessary to ensure that the trajectory meets the physical constraints of rigid body motion. For example, the motion of an object should conform to its dynamic characteristics and will not penetrate other objects or violate the laws of physics. The differentiable physics engine can calculate the difference between the predicted trajectory and the physical laws, that is, the rigid body motion constraint loss. By minimizing this loss function, the model can learn a motion trajectory that is more consistent with physical reality. The loss function is defined as follows: ; Where λ is the acceleration constraint weight coefficient, γ is the velocity continuity constraint weight coefficient, M is the object mass matrix, F_ext is the external force vector, and M is the object mass matrix (obtained through a preset object category prior knowledge base, such as vehicle mass distribution); ∂²x t / ∂t² is the acceleration of the object at time t, Δx t is the displacement prediction value, v t is the velocity observation value, Δt is the time interval, and T is the length of the prediction time window.

[0051] The loss function calculates the gradient through a differentiable physics engine (such as PyBullet or NVIDIA PhysX engine) and backpropagates to the spatiotemporal attention network to ensure that the predicted trajectory conforms to the momentum conservation and motion continuity laws of rigid body motion.

[0052] In specific implementation, the first-order derivative (velocity vector) and second-order derivative (acceleration vector) of the trajectory prediction result are first calculated, and then constraints are generated based on the law of conservation of momentum and the principle of collision avoidance. When the predicted trajectory conflicts with physical laws (for example, the vehicle trajectory penetrates the isolation zone), the system calculates the constraint loss gradient in real time and reversely optimizes the parameters of the spatiotemporal attention network. This mechanism can effectively suppress non-physical trajectory deviations. Through a collaborative optimization mechanism of spatiotemporal modeling and physical rules, including the spatiotemporal attention network capturing nonlinear motion characteristics (such as sudden acceleration and lane changes) through long-range dependency modeling, and rigid body constraints converting the collision detection function into a differentiable form to ensure that the predicted trajectory conforms to the impenetrability of physical objects, this collaborative optimization mechanism improves the physical credibility of motion prediction while ensuring real-time performance.

[0053] The spatiotemporal attention network captures complex motion patterns and contextual information, providing powerful feature representations for trajectory prediction. The differentiable physics engine, by introducing physical constraints, ensures the physical plausibility and stability of predictions. This combination not only improves the accuracy of trajectory predictions but also enhances the model's generalization and robustness, enabling it to produce reliable results in diverse scenarios and conditions. In autonomous driving scenarios, this combination enables autonomous vehicles to more accurately predict the motion trajectories of surrounding dynamic objects and make rational decisions based on physical laws, improving driving safety and comfort.

[0054] Step S103: Generate a spatiotemporal mask based on the optical flow data of adjacent frames, dynamically fuse the output results of the static scene base model and the dynamic trajectory branch, and align the three-dimensional features across frames.

[0055] In this embodiment, optical flow-driven spatiotemporal masking technology is used to achieve precise integration of dynamic and static scene elements. First, a binary spatiotemporal mask is generated based on the pixel-level motion vector field (i.e., optical flow data) between adjacent image frames. This mask divides the scene into static and dynamic regions. In specific implementation, the displacement of corresponding pixels in two consecutive image frames is calculated. When the displacement exceeds a preset threshold, the region is identified as dynamic (such as a moving vehicle), while the remaining region is classified as static background (such as the road surface). The spatiotemporal mask can accurately distinguish between moving vehicles and road markings, forming a precise region segmentation template. Guided by the mask, the system performs a fusion operation between the static base model and the dynamic trajectory branch: the signed distance field and RGB color values ​​output by the static base model branch are directly applied to the static region identified by the mask, while the motion attributes predicted by the dynamic trajectory branch cover the dynamic region.

[0056] To address spatial misalignment caused by object displacement across multiple frames, Deformable Feature Alignment (DFA) technology is used to correct for this problem. This technology maps the coordinates of dynamic objects in the current frame to the reference frame coordinate system using a three-dimensional coordinate transformation matrix. For example, in a vehicle turning scenario, when lateral displacement of the vehicle is detected across three consecutive frames, the system dynamically calculates a displacement compensation matrix to eliminate pixel-level offsets between the vehicle's outline and the road background. The coordinate transformation matrix is ​​derived based on the principles of rigid body kinematics to ensure that the correction process adheres to the laws of object motion.

[0057] In the technical solution of the embodiment of the present disclosure, through the dynamic neural radiation field architecture with time and space separation, a static scene base model is first constructed and the geometric structure and material properties are accurately encoded using a multi-layer perceptron to ensure high-fidelity restoration of static elements; the spatiotemporal attention network is further combined to predict dynamic trajectories, and rigid body motion constraints are incorporated to ensure physical rationality, so that the motion model remains stable in complex changes; finally, through optical flow-driven mask technology and cross-frame feature alignment, dynamic and static elements are dynamically integrated and spatiotemporal distortion is eliminated, thereby achieving high precision and high adaptability in three-dimensional scene construction, especially improving the reliability of object recognition and semantic segmentation in complex environments, and laying a foundation for efficient analysis for applications such as industrial inspection and virtual reality.

[0058] Step S20: Deploy a lightweight dynamic neural radiation field model on the edge computing node, process multimodal sensor data in real time and generate a local three-dimensional scene representation, perform rendering and prediction calculations of the neural radiation field on the edge node, and upload the implicit feature differences of the scene changes to the cloud.

[0059] When deploying a lightweight dynamic neural radiation field model on an edge computing node, lidar and cameras are used as sensors to obtain raw data. The data is processed by the lightweight dynamic neural radiation field model to generate a local three-dimensional scene, extract implicit features of the scene changes, and upload them to the cloud.

[0060] Optionally, to ensure the spatiotemporal consistency of multi-source data, multimodal data are collaboratively processed, including: using hardware trigger signals to precisely align the acquisition timing of the lidar and camera to eliminate time drift between sensors; continuously correcting the spatial mapping relationship between the lidar point cloud and the camera image based on a dynamically deployed checkerboard calibration plate to ensure accurate correspondence between three-dimensional coordinates and texture pixels; and using a double-buffered queue architecture to adaptively adjust the transmission rhythm of the point cloud stream and the image stream to resolve the frame mismatch problem caused by differences in data throughput rates.

[0061] Furthermore, the lightweight model performs end-to-end processing on the collaboratively processed fusion data. The implicit representation capability of the dynamic neural radiation field converts the multimodal input into a continuous three-dimensional scene representation, and simultaneously generates visual rendering results and motion trajectory prediction data. During this process, the system extracts scene evolution features through differential coding technology, including comparing the implicit representation differences between two consecutive frames, capturing only key change elements, and finally uploading the compressed feature differences to the cloud analysis platform through an encrypted channel. For example, the implicit representation capability of the dynamic neural radiation field is used to generate a continuous three-dimensional model of the drivable area, and the obstacle motion trajectory prediction and real-time rendering results of the drivable area are simultaneously output. At the same time, the system automatically captures the displacement trend of the suddenly cutting-in vehicle or the state transition of the traffic light, and only extracts the abstract features of these key dynamic changes and uploads them to the cloud platform.

[0062] In this embodiment, the lightweight dynamic neural radiation field model is deployed on the edge computing node. The lightweight dynamic neural radiation field model of the edge node is implemented in the following manner: Perform channel pruning on the original dynamic neural radiation field model, remove redundant convolutional layers in the static branch base model, and retain all parameters of the dynamic attention module; Knowledge distillation is used to use the output of the full model on the cloud as a teacher signal to guide the feature alignment of the edge model.

[0063] The method involves performing channel pruning on the original dynamic neural radiation field model, removing redundant convolutional layers in the static branch base model, and retaining all parameters of the dynamic attention module. This includes performing channel pruning on the original dynamic neural radiation field model. This method specifically removes redundant low-contribution convolution kernels in the static branch based on the feature importance evaluation of the convolution layer. For example, the deep convolution module used for low-frequency background details is deleted, while the key parameters for capturing motion trajectories in the dynamic attention module are fully retained. In autonomous driving scenarios, the pruned model successfully runs on an in-vehicle embedded platform, continuously constructing a static base model of the road (such as the geometric structure of traffic signs) while improving the speed of predicting pedestrian movement trajectories. Optionally, the pruning ratio is controlled in the range of 30%-50%, removing redundant 3D convolutional layers (channels with a contribution of less than 0.1) in the static branch, and retaining all parameters of the dynamic attention module.

[0064] In edge computing scenarios, knowledge distillation is a model compression and knowledge transfer technology. Its core idea is to efficiently transfer the reasoning capability of the complete dynamic neural radiation field model (teacher model) deployed in the cloud to the lightweight model (student model) at the edge. This embodiment further introduces knowledge distillation technology to construct a knowledge transfer channel for cloud-edge collaboration, uses the rendering output of the complete cloud model as the teacher signal, and guides the edge model to learn high-dimensional feature representations through the feature alignment loss function. Optionally, the complete cloud model performs forward reasoning on the same input scene, extracts the 128-dimensional feature map of the third layer of the dynamic attention module as the teacher signal, and uses the KL divergence loss function as the feature alignment loss function to measure the feature distribution difference L between the student model and the teacher model. KD , set the initial learning rate to 10 -3 , batch size 16, cosine annealing strategy is used to adjust the learning rate, freeze the teacher model parameters, and only update the student model. When the validation set feature alignment error L KD The training is terminated when the value <0.1 does not decrease for 5 consecutive epochs.

[0065] For example, in a low-contrast environment, the detailed contour features of a vehicle in a tunnel backlit environment are generated. The edge-end lightweight model simulates the deep representation distribution of the teacher model through the feature alignment loss function, and migrates the scene to the same low-contrast environment. For example, when dealing with a rainstorm interference scene, the penetrating rendering characteristics of the teacher model network for rain and fog are reproduced, thereby improving the reliability of obstacle recognition in low visibility.

[0066] Channel pruning reduces the computational load to meet the real-time processing requirements of the vehicle platform, while knowledge distillation compensates for information loss caused by pruning and maintains modeling accuracy. The organic integration of the two overcomes device computing bottlenecks while maintaining the ability to understand complex scenarios. This builds a knowledge evolution ecosystem that collaborates with the cloud and the edge. The teacher model continuously infuses knowledge of physical laws, while the edge model dynamically absorbs environmental adaptability, forming a closed-loop enhanced intelligent perception system.

[0067] See also Figure 3 As shown, before uploading the implicit feature difference of the scene change to the cloud, the method for generating the implicit feature difference includes the following steps: Step S201: Using a variational autoencoder to extract implicit feature vectors of dynamic scenes and generate compressed feature difference quantities; Generating and uploading implicit feature differences of scene changes is a key technical step in the implementation of edge computing nodes. This process uses a multi-stage process to ensure data compression efficiency and transmission security. These implicit feature differences represent the offset in the feature space of dynamic scene changes after being abstracted by a deep network. First, the system abstracts dynamic scene changes using a variational autoencoder (VAE). This encoder uses an encoder-decoder architecture to capture the core patterns of dynamic object motion or environmental evolution. For example, in an autonomous driving scenario, when a sudden lane change is detected by the vehicle ahead, the VAE analyzes its displacement vector and velocity changes to generate a highly compressed implicit feature vector. This feature vector not only significantly reduces data size but also preserves the essential information of the original scene change, avoiding interference from redundant details. Optionally, the VAE latent space dimension is fixed at 128, and the training dataset contains 100,000 dynamic scene samples.

[0068] Step S202: compress the feature vector into a binary stream of fixed length through hash coding; The feature vector is further processed using hash coding technology, converting it into a fixed-length binary stream. Hash coding uses a locality-sensitive hashing algorithm to ensure that similar scene changes (such as pedestrian movement trajectories in consecutive frames) generate similar binary sequences. This not only achieves secondary data compression but also facilitates rapid retrieval and comparison in the cloud. Optionally, hash coding can compress the feature vector into a 128-bit binary stream.

[0069] Step S203: Encrypt the binary stream using homomorphic encryption and transmit it to the cloud.

[0070] In this embodiment, the use of homomorphic encryption to encrypt a binary stream and transmit it to the cloud includes: encrypting the binary stream using homomorphic encryption, generating a dynamic encryption key using a quantum key distribution mechanism, and encrypting the compressed implicit feature data, thereby protecting data privacy during transmission between the cloud and the edge. Homomorphic encryption allows the cloud to directly process feature data (such as aggregating differential information from multiple edge nodes) without decryption, while quantum key distribution distributes one-time keys via quantum channels, completely eliminating the risk of man-in-the-middle attacks and ensuring privacy and security during transmission from the edge to the cloud. Optionally, a 256-bit dynamic key is generated based on the BB84 quantum key distribution protocol, and the binary stream is homomorphically encrypted using Paillier encryption before transmission.

[0071] In the technical solution of the disclosed embodiment, the variational autoencoder is responsible for feature abstraction and preliminary compression, hash coding ensures the uniformity and manageability of the data format, and homomorphic encryption combined with quantum mechanism provides privacy protection. This design not only adapts to the real-time needs in complex environments, but also lays a reliable foundation for large-scale data analysis in the cloud, while significantly reducing communication overhead and latency.

[0072] Step S30: The cloud aggregates feature difference data of multiple edge nodes through the federated learning framework, dynamically updates the global scene prior knowledge base and sends it to the edge nodes.

[0073] In this step, see Figure 4 As shown in the figure, the cloud aggregates feature difference data of multiple edge nodes through the federated learning framework, which includes the following steps: Step S301: Store and update feature difference data collected from multiple edge computing nodes, dynamically optimize prior knowledge of the global scene, and form a global scene knowledge base; The federated learning framework is a distributed machine learning architecture that allows multiple edge devices to process data and train models locally, uploading only model updates (not raw data) to the cloud for aggregation and optimization. The federated learning framework operates as a core collaborative mechanism, integrating feature difference data from multiple edge nodes through an encrypted gradient aggregation strategy. The framework first establishes a distributed computing pipeline to receive encrypted feature packets transmitted via quantum channels. These data packets are like cognitive fragments stripped of sensitive information, carrying the dynamic changes in scenarios in different geographical regions. The cloud server uses a weighted fusion algorithm to parse this fragmented knowledge. For example, it spatially aligns the overpass congestion pattern collected by the Shanghai node with the traffic characteristics of the foggy road section collected by the Chongqing node to generate a cross-regional road scene prior model.

[0074] The scene prior knowledge base is a dynamic database that stores cross-scenario cognitive experience. It uses a graph neural network to construct a spatial topological relationship model. This model abstracts road elements into nodes (such as traffic lights, crosswalks, and bus stops) and defines spatial constraints between them through connecting edges (for example, a no-travel relationship exists between red light nodes and pedestrian waiting area nodes). The temporal memory augmented network stores implicit feature sequences from historical scenes. When the current scene needs to match prior patterns in the knowledge base, the system triggers a similarity retrieval algorithm based on locality-sensitive hashing (LSH). This mechanism maps high-dimensional feature vectors to a low-dimensional hash space and uses Hamming distance calculations to quickly match similar patterns in the historical scene library (for example, identifying vehicle skidding during heavy rain), thus achieving scene pattern comparison. Furthermore, the temporal memory network captures the evolution of historical scenes, such as storing a curve of pedestrian density changes in school areas during the morning rush hour. When a vehicle enters an unfamiliar road section, the knowledge base proactively pushes relevant prior knowledge. If the system identifies a topological node in the hospital area, it automatically associates historical features that increase the probability of emergency ambulance passage and pre-loads a special vehicle avoidance strategy. This mechanism enables the autonomous driving system to have predictive decision-making capabilities.

[0075] Step S302: By regularly aggregating feature difference data of edge nodes, the global model is optimized in the cloud, and the updated global scene knowledge is sent to the edge nodes.

[0076] Cloud servers regularly receive encrypted feature difference packets uploaded by multiple edge nodes. These packets convey dynamic scene changes in distributed environments, such as vehicle lane changes at intersections in different urban areas. Using a federated learning aggregation engine, the cloud performs spatiotemporal alignment and weighted fusion on this feature difference data, distilling fragmented experience into universal physical laws, enabling the global model to continuously evolve with environmental adaptability.

[0077] The global model optimization process demonstrates a dual level of cognitive deepening: a graph neural network-based topological modeling engine analyzes spatial constraints between features, for example linking temporary construction roadblocks and vehicle detour trajectories into an "obstacle-avoidance path" topological chain. A temporal memory network captures the evolution of historical features. When detecting the morning rush hour characteristics of a school area, it automatically associates a spatiotemporal rule base based on "student crossing density-vehicle deceleration probability." The optimized knowledge is packaged into incremental update packages, using gradient masking technology to refine parameters related to dynamic objects (such as the trajectory prediction module for new vehicles). These packages are then distributed to edge nodes via quantum-encrypted channels.

[0078] The implementation of the present invention constructs the core value of the closed loop of group intelligence evolution-scenario cognition. By aggregating the encrypted feature difference data of multiple edge nodes through federated learning, the cloud breaks through the limitations of single-point perception, integrates scattered local experiences into universal physical laws, and forms a dynamically evolving global scenario knowledge base. Its essence is to establish a constraint network between objects and the environment through spatial relationship modeling technology, and use the historical scene memory function to capture the evolution context. This mechanism enables the system to have predictive decision-making capabilities, such as actively associating historical risk models to generate early warning strategies when identifying special areas. In the knowledge feedback link, the cloud uses precise parameter extraction technology to generate lightweight update packages, and combines quantum encryption channels to achieve edge node synchronization. The whole process forms a self-reinforcing technology ecosystem. The sudden environmental changes captured by the edge nodes trigger the immediate evolution of the cloud knowledge base. The updated cognitive capabilities are fed back to the edge through efficient channels, so that devices that have not experienced specific scenarios can also obtain advanced decision-making capabilities.

[0079] This embodiment addresses the geometric and motion distortion issues in dynamic scene modeling through a spatiotemporally separated dynamic neural radiation field architecture. In the static scene construction phase, the convolutional signed distance field (ConvSDF) and multi-layer perceptron (MLP) collaborate to accurately restore high-precision geometric structures and material optical properties. The dynamic trajectory branch captures nonlinear motion patterns through a spatiotemporal attention mechanism and introduces rigid body motion constraints to ensure that the predicted trajectory conforms to physical laws, significantly reducing trajectory prediction errors in scenarios such as sharp turns. At the edge computing layer, lightweight model deployment and feature differentiation mechanisms overcome computing power and bandwidth limitations. Channel pruning removes redundant parameters in static branches, and knowledge distillation technology migrates the complex scene understanding capabilities of the cloud model to the edge, enabling the vehicle to maintain high-precision perception in complex environments and improving the efficiency and accuracy of 3D data processing. Furthermore, a variational autoencoder extracts implicit feature differences in scene changes, which are then uploaded after hash compression and quantum encryption, addressing the security risks of edge-to-cloud data exchange and the conflict between data security and transmission efficiency. Cloud-based federated learning and the knowledge base's co-evolutionary mechanism form a closed loop. By aggregating multi-node feature differences, a graph neural network constructs the topological relationship between dynamic objects and static scenes, and a temporal memory network stores the evolution patterns of historical scenes. The optimized global knowledge is refined into incremental parameters using gradient masks and distributed to edge nodes, enabling the adaptive evolution of scene prior knowledge. This enables devices or vehicles that have not experienced specific scenarios to instantly acquire advanced decision-making capabilities, improving the applicability of three-dimensional data processing. This invention can provide high-precision, low-latency three-dimensional spatial analysis capabilities for fields such as autonomous driving and industrial inspection.

[0080] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0081] It should be understood that, although the above is described in a certain order, these steps are not necessarily performed in sequence according to the above order. Unless clearly stated herein, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, a part of the steps of the present embodiment may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.

[0082] In a second aspect of the embodiments of the present invention, the present invention further provides a three-dimensional spatial data analysis system based on deep learning, comprising: Multimodal acquisition module: used to synchronously acquire 3D spatial data, including lidar, multi-view cameras, and an inertial measurement unit (IMU). The lidar generates point cloud data, the camera captures multi-view images, and the IMU provides device motion parameters. The data collected by this module are directly input into the spatiotemporally separated dynamic neural radiation field to construct static scene archetypes and predict dynamic trajectories; Edge computing unit: Deployed on the sensor terminal side, it includes a reconfigurable computing array (FPGA) and a lightweight dynamic neural radiation field (LiteDyNeRF) model for real-time processing of multimodal data and generating a neural radiation field representation of the local 3D scene; This unit receives data from the multimodal acquisition module and processes it in real time to generate a local scene representation. For example, the point cloud preprocessing module optimizes the input, the LiteDyNeRF model performs rendering and prediction, and finally, a variational autoencoder extracts implicit feature differences that indicate scene changes. This entire process is completed at the edge, ensuring efficiency and privacy.

[0083] Cloud-based optimization platform: This includes a federated learning server and a global scenario knowledge base, which aggregates feature difference data from multiple edge nodes and dynamically updates the global model. The platform receives encrypted feature difference data uploaded by edge nodes and aggregates this information across multiple nodes through a federated learning framework. The federated learning server optimizes the global model, the knowledge base updates topological relationships, and the model update distribution module only distributes incremental parameters, enabling the adaptive evolution of scenario prior knowledge.

[0084] ‌‌Encrypted Communication Interface‌: This interface uses an encrypted channel based on quantum key distribution for transmitting feature data and issuing model update instructions between the edge computing unit and the cloud optimization platform.

[0085] This interface processes the implicit feature differences generated by the edge computing unit. After the VAE compresses the features, the quantum cryptography module encrypts the binary stream, and the cloud-based TEE decrypts and verifies it, ensuring data privacy and integrity during transmission and preventing man-in-the-middle attacks.

[0086] In this embodiment, the edge computing unit includes: Point cloud preprocessing module: downsamples, denoises, and voxelizes the raw LiDAR point cloud to generate a sparse voxel representation; Dynamic Detection Accelerator: This uses a sparse convolutional neural network to identify dynamic object regions in real time and triggers a spatiotemporal attention network for trajectory prediction. Reconfigurable Rendering Engine: Supports parallel computation of ray casting and voxel rendering for neural radiation fields, achieving real-time rendering at 35ms / frame through FPGA hardware acceleration.

[0087] In this embodiment, the cloud optimization platform includes: Federated Learning Server: This server protects feature difference data uploaded by edge nodes through differential privacy technology and performs distributed gradient aggregation of model parameters. Global scene knowledge base: This base uses a graph neural network (GNN) to model the topological relationships between objects in the scene and uses a temporal memory network (TMC) to store implicit feature sequences of historical scenes. Model Update Distribution Module: A gradient masking strategy is designed to incrementally update only the model parameters related to dynamic objects, and the updates are distributed to edge nodes through encrypted channels.

[0088] In this embodiment, the lightweight dynamic neural radiation field model is implemented in the following manner: Model Compression Module: This module performs channel pruning on the original DyNeRF model, removes redundant 3D convolutional layers in the static branch, and retains the complete parameters of the dynamic attention module. Knowledge Distillation Module: The rendered output of the complete model on the cloud is used as the teacher signal, and the feature alignment loss function is used to guide the edge model to learn high-dimensional feature representations.

[0089] In this embodiment, the encrypted communication interface includes: Feature Difference Encoder: uses a variational autoencoder (VAE) to encode dynamic scene changes into a 0.5MB implicit feature vector; Quantum encryption module: Generates dynamic encryption keys through the quantum key distribution protocol, and transmits feature data after homomorphic encryption; ‌Decryption Verification Unit‌: Decrypts and verifies the integrity of the received encrypted data through the Trusted Execution Environment (TEE) in the cloud.

[0090] In this embodiment, the method for constructing the global scenario knowledge base includes: Scene topology modeling: Analyze the spatial constraints between dynamic objects and static scenes based on graph neural networks (GNNs); Temporal feature storage: Storing implicit feature sequences of historical scenarios through a memory-augmented network. ‌Similarity retrieval algorithm‌: Design a fast retrieval mechanism based on locality sensitive hashing (LSH) to match the current scene with the prior patterns in the knowledge base.

[0091] In this embodiment, the collaboration method of the multimodal acquisition module includes: ‌Spatiotemporal Synchronization Unit‌: Aligns the acquisition timestamps of the LiDAR and camera through hardware trigger signals, with an error of less than 1ms; Cross-modal calibration module: Dynamically calibrates the extrinsic parameter matrices of the lidar and camera based on a checkerboard calibration plate; ‌Data Fusion Interface‌: A double-buffered queue mechanism is used to balance the processing rate differences between point cloud and image data.

[0092] Through the above detailed steps, the three-dimensional spatial data analysis system based on deep learning of the present invention is used to execute the steps of the three-dimensional spatial data analysis method based on deep learning in the above embodiment, which will not be repeated here.

[0093] In summary, the present invention addresses the geometric and motion distortion issues in dynamic scene modeling through a spatiotemporally separated dynamic neural radiation field architecture. In the static scene construction phase, the convolutional signed distance field (ConvSDF) and multi-layer perceptron (MLP) collaborate to achieve high-precision restoration of geometric structures and material optical properties. The dynamic trajectory branch captures nonlinear motion laws through a spatiotemporal attention mechanism and introduces rigid body motion constraints to ensure that the predicted trajectory conforms to physical laws, significantly reducing trajectory prediction errors in scenarios such as sharp vehicle turns. At the edge computing layer, lightweight model deployment and feature differentiation mechanisms overcome computing power and bandwidth limitations. Channel pruning removes redundant parameters in static branches, and knowledge distillation technology migrates the complex scene understanding capabilities of the cloud model to the edge, enabling the vehicle to maintain high-precision perception in complex environments and improving the efficiency and accuracy of three-dimensional data processing. Simultaneously, a variational autoencoder extracts implicit feature differences in scene changes, which are then uploaded after hash compression and quantum encryption, resolving security risks in edge-to-cloud data exchange and the conflict between data security and transmission efficiency. Cloud-based federated learning and the knowledge base's co-evolutionary mechanism form a closed loop. By aggregating multi-node feature differences, a graph neural network constructs the topological relationship between dynamic objects and static scenes, and a temporal memory network stores the evolution patterns of historical scenes. The optimized global knowledge is refined into incremental parameters using gradient masks and distributed to edge nodes, enabling the adaptive evolution of scene prior knowledge. This enables devices or vehicles that have not experienced specific scenarios to instantly acquire advanced decision-making capabilities, improving the applicability of three-dimensional data processing. This invention can provide high-precision, low-latency three-dimensional spatial analysis capabilities for fields such as autonomous driving and industrial inspection.

[0094] According to a third aspect of an embodiment of the present invention, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method of any one of the above embodiments is implemented.

[0095] The computer device includes a processor and a memory, and may also include an input system and an output system. The processor, memory, input system, and output system may be connected via a bus or other means. The input system may receive input digital or character information and generate signal input related to the migration of deep learning-based three-dimensional spatial data analysis. The output system may include a display device such as a display screen.

[0096] As a non-volatile computer-readable storage medium, the memory can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as the program instructions / modules corresponding to the three-dimensional spatial data analysis method based on deep learning in the embodiment of the present application. The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of the three-dimensional spatial data analysis method based on deep learning, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the local module via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0097] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data. The processors of the multiple computer devices of the computer device of this embodiment execute various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory, that is, implementing the steps of the three-dimensional spatial data analysis method based on deep learning of the above-mentioned method embodiment.

[0098] It should be understood that, to the extent that they do not conflict with each other, all the embodiments, features and advantages described above for the deep learning-based three-dimensional spatial data analysis method according to the present invention are also applicable to the deep learning-based three-dimensional spatial data analysis and storage medium according to the present invention.

[0099] It will also be appreciated by those skilled in the art that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.

[0100] Finally, it should be noted that the computer-readable storage medium (e.g., memory) herein may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. By way of example and not limitation, non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which may act as external cache memory. By way of example and not limitation, RAM is available in various forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memory.

[0101] The various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure herein may be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components, designed to perform the functions herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP, and / or any other such configuration.

[0102] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications may be made without departing from the scope of the embodiments disclosed in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention may be described or required in individual form, they may also be understood as multiple unless expressly limited to the singular.

[0103] It should be understood that, as used herein, the singular form "a" or "an" is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations of one or more of the items listed in association. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0104] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to limit the scope of the disclosure of the present invention (including the claims) to these examples. Within the spirit of the present invention, the technical features of the above embodiments or different embodiments may be combined, and many other variations exist in different aspects of the above embodiments, which are not provided in detail for the sake of clarity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A three-dimensional spatial data analysis method based on deep learning, characterized in that: The method comprises the following steps: Static scene primitives and dynamic object motion trajectories are modeled through spatiotemporally separated dynamic neural radiance fields. The static scene primitives are encoded using convolutional signed distance fields, and the dynamic trajectories are jointly optimized with the rigid body motion equations through a spatiotemporal attention network. Deploy a lightweight dynamic neural radiation field model on edge computing nodes to process multimodal sensor data in real time and generate a local 3D scene representation. Rendering and prediction calculations of the neural radiation field are performed on the edge nodes, and implicit feature differences of scene changes are uploaded to the cloud. The cloud aggregates feature difference data from multiple edge nodes through a federated learning framework, dynamically updates the global scenario prior knowledge base, and sends it to the edge nodes.

2. The three-dimensional spatial data analysis method based on deep learning according to claim 1, characterized in that: Modeling static scene primitives and dynamic object motion trajectories through spatiotemporally separated dynamic neural radiation fields includes the following steps: The static scene base model branch uses a multi-layer perceptron to encode global illumination, material properties, and spatial structure information, and outputs a mapping between signed distance field and RGB color values; The dynamic trajectory branch predicts the motion trajectory of dynamic objects through a spatiotemporal attention network and calculates the rigid body motion constraint loss in combination with a differentiable physics engine; A spatiotemporal mask is generated based on the optical flow data of adjacent frames, and the output results of the static scene base model and the dynamic trajectory branch are dynamically fused to align the three-dimensional features across frames.

3. The three-dimensional spatial data analysis method based on deep learning according to claim 2, characterized in that: When modeling static scene primitives and dynamic object motion trajectories through spatiotemporally separated dynamic neural radiation fields, it also includes: Perform convolutional signed distance field encoding on static scenes to obtain a representation of the geometric structure; Perform spatiotemporal modeling of dynamic object motion and use spatiotemporal attention mechanism to model object motion; Based on the rigid body motion equation, the object's motion trajectory is jointly optimized in combination with the spatiotemporal attention network.

4. The three-dimensional spatial data analysis method based on deep learning according to claim 1, characterized in that: When deploying a lightweight dynamic neural radiation field model on an edge computing node, lidar and cameras are used as sensors to obtain raw data. The data is processed by the lightweight dynamic neural radiation field model to generate a local three-dimensional scene, extract implicit features of the scene changes, and upload them to the cloud.

5. The three-dimensional spatial data analysis method based on deep learning according to claim 4, characterized in that: Deploy a lightweight dynamic neural radiation field model on the edge computing node. The lightweight dynamic neural radiation field model of the edge node is implemented in the following ways: Perform channel pruning on the original dynamic neural radiation field model, remove redundant convolutional layers in the static branch base model, and retain all parameters of the dynamic attention module; Knowledge distillation is used to use the output of the full model on the cloud as a teacher signal to guide the feature alignment of the edge model.

6. The three-dimensional spatial data analysis method based on deep learning according to claim 5, characterized in that: Before uploading the implicit feature difference of scene changes to the cloud, the method for generating the implicit feature difference includes the following steps: The implicit feature vector of the dynamic scene is extracted using a variational autoencoder to generate compressed feature differences. Compress the feature vector into a fixed-length binary stream through hash coding; Homomorphic encryption is used to encrypt the binary stream and then transmit it to the cloud.

7. The three-dimensional spatial data analysis method based on deep learning according to claim 6, characterized in that: The cloud aggregates feature difference data from multiple edge nodes through a federated learning framework, which includes the following steps: Store and update feature difference data collected from multiple edge computing nodes, dynamically optimize the prior knowledge of the global scene, and form a global scene knowledge base; By regularly aggregating feature difference data from edge nodes, the global model is optimized in the cloud, and the updated global scene knowledge is sent to the edge nodes.

8. A three-dimensional spatial data analysis system based on deep learning, characterized in that: The system is used to perform the three-dimensional spatial data analysis method based on deep learning according to any one of claims 1 to 7, comprising: Multimodal acquisition module: used to synchronously acquire 3D spatial data, including a laser radar (LiDAR), a multi-view camera, and an inertial measurement unit (IMU). The LiDAR generates point cloud data, the camera captures multi-view images, and the IMU provides device motion parameters. Edge computing unit: Deployed on the sensor terminal side, it includes a reconfigurable computing array and a lightweight dynamic neural radiation field model, which is used to process multimodal data in real time and generate a neural radiation field representation of the local three-dimensional scene; Cloud-based optimization platform: This includes a federated learning server and a global scenario knowledge base, which aggregates feature difference data from multiple edge nodes and dynamically updates the global model. Encrypted Communication Interface: This interface uses an encrypted channel based on quantum key distribution for transmitting feature data and issuing model update instructions between the edge computing unit and the cloud optimization platform.

9. The three-dimensional spatial data analysis system based on deep learning according to claim 8, characterized in that: The edge computing unit includes: Point cloud preprocessing module: downsamples, denoises, and voxelizes the raw LiDAR point cloud to generate a sparse voxel representation; Dynamic Detection Accelerator: This uses a sparse convolutional neural network to identify dynamic object regions in real time and triggers a spatiotemporal attention network for trajectory prediction. Reconfigurable Rendering Engine: Supports parallel computation of ray casting and voxel rendering for neural radiation fields, achieving real-time rendering at 35ms / frame through FPGA hardware acceleration.

10. The three-dimensional spatial data analysis system based on deep learning according to claim 9, characterized in that: The cloud optimization platform includes: Federated Learning Server: This server protects feature difference data uploaded by edge nodes through differential privacy technology and performs distributed gradient aggregation of model parameters. Global scene knowledge base: This base uses a graph neural network to model the topological relationships between objects in the scene and uses a temporal memory network to store implicit feature sequences of historical scenes. Model Update Distribution Module: A gradient masking strategy is designed to incrementally update only the model parameters related to dynamic objects, and the updates are distributed to edge nodes through encrypted channels.

Citation Information

Patent Citations

  • Scene space-time reconstruction method and system, electronic equipment and storage medium

    CN118397181A

  • 4D scene characterization method combining pose and radiation field optimization in complex mine environment

    CN119006687A

  • Scene space three-dimensional model dynamic modeling method based on multi-modal data

    CN119339008A

  • Building indoor and outdoor integrated three-dimensional model construction method, system and equipment

    CN120107488A

  • Data acquisition and reconstruction method and system for human body three-dimensional modeling based on single mobile phone

    US20240153213A1

Cited By

  • Article information pushing method based on knowledge distillation

    CN120812123A

  • Windows three-dimensional cloud rendering system, method and equipment

    CN120821548A

  • Trajectory prediction method and system based on multi-modal perception

    CN120832246A

  • Robot motion control system based on artificial intelligence

    CN120886275A

  • Kitchen air conditioner filtering device service life prediction method and system based on deep learning

    CN120950908A