A method and device for fusing laser point cloud and video based on 3D control ball
By using a 3D-based sphere deployment method, end-to-end security and spatiotemporal consistency processing of video surveillance data and laser point cloud data were achieved, generating intelligent and automated survey results. This solved the problem of incomplete scene representation in existing technologies and improved operational safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN HUITEST POWER TECH CO LTD
- Filing Date
- 2025-11-07
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, video surveillance equipment and laser point cloud mapping equipment have difficulty achieving end-to-end data security and spatiotemporal consistency processing in complex operating environments, resulting in incomplete scene representation and affecting operational safety.
The method based on 3D control ball is adopted. Through encryption processing, timestamp alignment and data format standardization, multiple computing nodes are used to extract features in parallel to generate a fused feature set and construct a multimodal 3D scene representation. This automatically generates an operation plan with spatial geometric constraints and operation safety boundaries.
It achieves end-to-end security and spatiotemporal consistency processing of video surveillance data and laser point cloud data, improves processing efficiency, and generates intelligent automated exploration results with spatial geometric constraints and operational safety boundaries.
Smart Images

Figure CN121639482B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a method and apparatus for fusing laser point clouds and video based on a 3D controlled sphere. Background Technology
[0002] Monitoring and analysis of work sites typically rely on video surveillance equipment or laser point cloud mapping equipment operating independently. Video surveillance equipment can provide two-dimensional image information for target recognition and scene perception, while laser point cloud mapping equipment can provide spatial coordinates and geometric structure information for building three-dimensional models. In practical applications, these two types of data are usually processed independently and are difficult to complement each other within a unified spatiotemporal framework. As a result, in complex work environments, it is often impossible to take into account texture details, geometric structure, and dynamic behavior characteristics, leading to incomplete scene representation and unreliable work plan generation.
[0003] A prominent shortcoming of existing technologies is that video surveillance data and laser point cloud data lack end-to-end security and spatiotemporal consistency guarantees in the data transmission, feature extraction and scene modeling stages. Features from the video side and the point cloud side often have temporal misalignment and spatial mismatch, which can easily lead to misjudgment of obstacle positions or errors in operation path planning, thereby affecting operation safety.
[0004] Therefore, a method is needed to achieve end-to-end security and spatiotemporal consistency processing of video surveillance data and laser point cloud data. Summary of the Invention
[0005] This invention provides a method and apparatus for fusing laser point clouds and video based on a 3D control sphere, which can achieve end-to-end security and spatiotemporal consistency processing of video surveillance data and laser point cloud data.
[0006] In a first aspect, the present invention provides a method for fusing laser point clouds and video based on a 3D controlled sphere, the method comprising:
[0007] After acquiring laser point cloud data and video data using a 3D surveillance sphere, the data is encrypted during transmission.
[0008] At the receiving end, the laser point cloud data and video data are decrypted, and the laser point cloud data and video data are respectively processed by timestamp alignment and data format standardization.
[0009] The standardized laser point cloud data and video data are distributed to multiple computing nodes, and each computing node performs feature extraction operations in parallel. Specifically, for the video data, multi-dimensional feature vectors based on image texture, edges and motion vectors are extracted, and for the laser point cloud data, three-dimensional geometric feature vectors based on spatial coordinates, surface normal vectors and point density are extracted.
[0010] The multidimensional feature vectors output by each computing node are subjected to cross-modal correlation calculation with the three-dimensional geometric feature vectors, and the multidimensional feature vectors are matched with the three-dimensional geometric feature vectors using the timestamp alignment results, thereby generating a fused feature set;
[0011] Based on the fused feature set, a multimodal 3D scene representation is constructed, and intelligent automated surveying and analysis of the work site are performed, outputting the work site analysis results;
[0012] Based on the multimodal 3D scene representation and the analysis results of the work site, an operation plan with spatial geometric constraints and work safety boundaries is automatically generated.
[0013] In a second aspect of the invention, an apparatus for fusing laser point clouds and video based on a 3D controlled sphere is provided. The apparatus is used to perform a method for fusing laser point clouds and video based on a 3D controlled sphere as described above. The apparatus includes an acquisition module, a processing module, and an output module, wherein:
[0014] The acquisition module is used to encrypt the laser point cloud data and video data acquired by the 3D control ball during the data transmission process.
[0015] The processing module is used to decrypt the laser point cloud data and video data at the receiving end, and to perform timestamp alignment and data format standardization on the laser point cloud data and the video data respectively.
[0016] The processing module is used to distribute the standardized laser point cloud data and video data to multiple computing nodes. Each computing node performs feature extraction operations in parallel. Specifically, for the video data, a multi-dimensional feature vector based on image texture, edge and motion vector is extracted, and for the laser point cloud data, a three-dimensional geometric feature vector based on spatial coordinates, surface normal vector and point density is extracted.
[0017] The processing module is used to perform cross-modal correlation calculation on the multi-dimensional feature vectors output by each computing node and the three-dimensional geometric feature vectors, and to match the multi-dimensional feature vectors and the three-dimensional geometric feature vectors using the timestamp alignment results, thereby generating a fused feature set;
[0018] The processing module is used to construct a multimodal 3D scene representation based on the fused feature set, and to perform intelligent and automated survey and analysis of the work site, and output the work site analysis results.
[0019] The output module is used to automatically generate a work plan with spatial geometric constraints and work safety boundaries based on the multimodal 3D scene representation and the work site analysis results.
[0020] In a third aspect of the invention, an electronic device is provided, including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the preceding embodiments.
[0021] In a fourth aspect, the present invention provides a computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.
[0022] In summary, one or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:
[0023] 1. This invention introduces an encryption mechanism based on device identity credentials and session key sets in the data transmission stage to ensure the confidentiality and integrity of laser point cloud data and video data in the link. At the receiving end, decryption, replay detection, and integrity verification are completed by combining access tokens, thereby achieving end-to-end security. At the same time, unified time-series labels are used for timestamp alignment and missing frame compensation. Extrinsic and intrinsic parameter matrices are used to establish cross-modal spatial mapping. After cross-node parallel feature extraction, a fused feature set is generated through time-series alignment and cross-modal matching to ensure that video data and point cloud data maintain a consistent time base and spatial reference throughout the entire process, thus achieving end-to-end spatiotemporal consistency processing.
[0024] 2. By generating time window batches and allocating them to multiple computing nodes, video data and laser point cloud data are processed in parallel within a unified time range. This ensures the consistency of cross-modal data and improves processing efficiency through parallel computing, thereby achieving efficient and synchronous feature extraction in large-scale scenarios.
[0025] 3. By establishing an encryption mechanism based on device identity credentials and hardware root keys at the 3D deployment ball terminal, and using point cloud session keys and video session keys to perform frame-level protection and authentication encryption on the data, and adding replay protection and encryption channels during transmission, the confidentiality, integrity and anti-replay capability of laser point cloud data and video data are guaranteed during transmission.
[0026] 4. By constructing a multimodal 3D scene representation based on a fusion feature set, terrain and landform segmentation, obstacle detection, and cross-object analysis are performed in the triangular mesh by combining geometric and visual attributes. The operation safety boundary is generated with risk indicators, and the final output is the analysis results with passable corridors, equipment placement positions, and operation path suggestions, thereby realizing intelligent and automated survey and risk assessment of the operation site.
[0027] 5. By establishing a joint optimization model of path and process on the task constraint graph, path planning, equipment placement and time scheduling are solved in a unified manner. Combined with the operation safety boundary, an executable instruction set with dynamic monitoring and emergency rollback mechanism is automatically generated, thereby realizing the generation of automated operation schemes that meet spatial geometric constraints and safety constraints. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating a method for fusing laser point clouds and video based on a 3D controlled sphere, as disclosed in an embodiment of the present invention.
[0029] Figure 2 This is a schematic diagram of a device for fusing laser point clouds and video based on a 3D control sphere, as disclosed in an embodiment of the present invention.
[0030] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.
[0031] Explanation of reference numerals in the attached drawings: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation
[0032] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0033] In the description of the embodiments of the present invention, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0034] In the description of the embodiments of the present invention, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0035] In existing technologies, video surveillance equipment and laser point cloud mapping equipment typically operate independently. The former focuses on two-dimensional image recognition, while the latter focuses on three-dimensional geometric modeling. The lack of unified spatiotemporal fusion between the two leads to incomplete scene representation, insufficient reliability of job plan generation, and a lack of end-to-end security and spatiotemporal consistency guarantees during data transmission and processing. This can easily cause feature misalignment, obstacle misjudgment, and path planning errors. Therefore, there is an urgent need for a method that can achieve end-to-end security and spatiotemporal consistency processing of video surveillance data and laser point cloud data.
[0036] This embodiment discloses a method for fusing laser point clouds and video based on a 3D controlled sphere, referring to... Figure 1 This includes the following steps S110-S160:
[0037] The S110 encrypts laser point cloud data and video data acquired by the 3D surveillance sphere during data transmission.
[0038] This invention discloses a method for fusing laser point clouds and video based on a 3D controlled sphere, which is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running a method for fusing laser point clouds and video based on a 3D controlled sphere. The server can be implemented using a standalone server or a server cluster composed of multiple servers.
[0039] In one possible implementation, after the 3D surveillance sphere acquires laser point cloud data and video data, encryption processing is performed during data transmission. Specifically, this includes: after the 3D surveillance sphere generates a device identity credential, binding the device identity credential to the hardware root key in the security chip, and outputting the device identity public key and device identity certificate; verifying the device identity certificate and issuing an access token and negotiation parameter set; wherein, the 3D surveillance sphere generates a session key set based on the negotiation parameter set and the device identity public key, the session key set including a point cloud session key for laser point cloud data and a session key for video data. Video session key; a frame-level protection structure is established for laser point cloud data using the point cloud session key, and a frame-level protection structure is established for video data using the video session key. A timing tag, frame sequence number, and source identifier are injected into the frame header, and the frame payload is subjected to authenticated encryption processing to generate a point cloud ciphertext frame sequence and a video ciphertext frame sequence. During transmission, a replay bitmap with a scrolling window and an accumulation counter are added to the point cloud ciphertext frame sequence and the video ciphertext frame sequence, and a connection-oriented encrypted channel is established based on the access token. The point cloud ciphertext frame sequence and the video ciphertext frame sequence are used as payloads for real-time transmission.
[0040] Specifically, after the 3D surveillance sphere generates its device identity credential, the process of binding this credential to the hardware root key in the security chip involves several steps. First, the security module is invoked within the 3D surveillance sphere's trusted execution environment to generate a unique device identity credential. This credential typically includes a device identifier, a timestamp, and a random challenge code, used to identify the uniqueness of the 3D surveillance sphere within the network. Subsequently, the device identity credential is encrypted and bound using the hardware root key within the security chip, ensuring that the credential cannot be forged or copied. The result of the hardware root key binding generates a device identity public key and a device identity certificate. The device identity public key is used for interaction with external authentication services, while the device identity certificate is used to prove the authenticity and integrity of the device's identity. The hardware root key here refers to the key material written during the security chip manufacturing stage and cannot be exported; it is the root of trust for all security operations.
[0041] During the process of verifying device identity and issuing access tokens and negotiation parameter sets, the receiving authentication service first verifies the validity of the received device identity, including signature verification, timestamp validity check, and certificate chain integrity verification. Upon successful verification, the authentication service issues an access token and a set of negotiation parameters. The access token serves as the unique credential for the device to subsequently establish an encrypted channel, while the negotiation parameter set includes elements such as the key negotiation algorithm, cipher suite selection, and random seed, guiding the device to generate a new session key set. Based on this, the 3D surveillance sphere uses the negotiation parameter set and the device's public key to execute a key exchange protocol to generate a session key set. This session key set includes a point cloud session key and a video session key, dedicated to the encryption protection of laser point cloud data and video data, respectively. The session key here refers to a dynamically generated temporary key during a session, used to ensure the confidentiality and forward security of communication.
[0042] In establishing frame-level protection structures for laser point cloud data using point cloud session keys and for video data using video session keys, each frame of laser point cloud data and each frame of video data is first organized according to a frame sequence. A timing tag, frame number, and source identifier are inserted into the frame header to ensure that data frames from different sources can be correctly reassembled and aligned at the decryption end. Subsequently, authenticated encryption is performed on the frame payload, simultaneously performing confidentiality encryption and integrity verification. Confidentiality encryption ensures that frame data is not stolen, while integrity verification ensures that frame data is not tampered with during transmission. The point cloud ciphertext frame sequence and the video ciphertext frame sequence are the encryption results obtained through the above process, respectively protecting 3D spatial coordinates and image texture information. The frame-level protection structure here refers to a secure encapsulation mechanism performed at the level of each frame of data, ensuring that each independent frame can be independently verified and decrypted.
[0043] During transmission, the system adds a replay bitmap with a scrolling window and an accumulator counter to the point cloud ciphertext frame sequence and the video ciphertext frame sequence. For each batch of transmitted data frames, a dynamically updated bitmap window is generated. This bitmap window records the sequence number of frames already received by the receiver, thus detecting replay attacks. The accumulator counter ensures the monotonically increasing nature of the sequence, preventing attackers from disrupting the timing by inserting old frames. Based on the access token, a connection-oriented encrypted channel is established between the 3D surveillance sphere and the receiver. In this encrypted channel, all transmitted data frames are protected within the context of a single session and undergo reliable handshake and key confirmation. Finally, the point cloud ciphertext frame sequence and the video ciphertext frame sequence are transmitted in real-time as payloads in the encrypted channel, ensuring that data is delivered to the receiver with low latency, security, and without tampering even in complex operating environments. The encrypted channel here refers to a secure communication link established through key negotiation and access token authentication, possessing anti-eavesdropping, anti-tampering, and anti-replay capabilities.
[0044] S120 decrypts the laser point cloud data and video data at the receiving end, and performs timestamp alignment and data format standardization on the laser point cloud data and video data respectively.
[0045] In one possible implementation, the decryption processing of laser point cloud data and video data at the receiving end specifically includes: after completing access token verification and locating the corresponding session key set, importing the point cloud session key and video session key, initializing the decryption context, and resetting the replay window, cumulative counter, and timing tag buffer; inputting the point cloud ciphertext frame sequence and video ciphertext frame sequence into the corresponding decryption context respectively, performing replay detection and order rearrangement based on timing tag, frame sequence number, and source identifier, and using the replay bitmap of the scrolling window and the cumulative counter to remove illegal ciphertext frames; performing authenticated integrity verification on the ciphertext frames that pass the replay detection, and verifying the integrity of the ciphertext frames that pass the verification. The encrypted frames are decrypted using the point cloud session key and the video session key respectively, outputting point cloud decrypted frame objects and video decrypted frame objects. The time sequence label, frame number, and source identifier are retained in the decrypted frame objects. Forward error correction and missing frame compensation are performed on the point cloud decrypted frame objects and video decrypted frame objects. Based on the clock mapping relationship constructed by the negotiated parameter set, the time sequence label correction and cross-modal delay compensation are completed, outputting a time-aligned sequence of point cloud decrypted frame objects and a sequence of video decrypted frame objects. The time-aligned sequence of point cloud decrypted frame objects is parsed into standardized laser point cloud data, and the time-aligned sequence of video decrypted frame objects is parsed into standardized video data.
[0046] Specifically, after the receiving end completes access token verification and locates the corresponding session key set, it imports the point cloud session key and the video session key. Within the security process, it constructs decryption contexts for the two types of data respectively, initializes the symmetric key, random number generator, and authentication integrity verification parameter set, and resets the replay window, cumulative counter, and time tag buffer. This allows the replay window to record the acceptable frame sequence number sliding range, the cumulative counter to provide a monotonically increasing arrival reference, and the time tag buffer to maintain a cross-modal arrival time reference. Finally, it outputs the point cloud decryption context and video decryption context, which are in an available state, as input for ciphertext frame reception and initial screening.
[0047] The point cloud ciphertext frame sequence and the video ciphertext frame sequence are respectively input into the corresponding decryption context. For each ciphertext frame, the timing tag, frame number, and source identifier in the frame header are parsed. The replay bitmap and cumulative counter of the replay window are used to perform replay detection and order reordering. The replay bitmap is used to mark the processed or seen frame number, and the cumulative counter is used to identify abnormal backtracking or skipping of frame sequences. Entries determined to be illegal ciphertext frames are discarded and the event is recorded. Ciphertext frames determined to be legitimate candidates are bucketed according to the source identifier and stably reordered according to the frame number within the bucket. The candidate ciphertext frame set is output as the input for the integrity verification and decryption calculation with authentication.
[0048] The system performs an authenticated integrity check on each candidate ciphertext frame set. For ciphertext frames that pass the check, the corresponding point cloud session key or video session key is invoked to decrypt the payload. The decryption output and authentication mark are encapsulated together into a decrypted frame object. The decrypted frame object retains the timing tag, frame number, and source identifier to maintain the traceability link. Ciphertext frames that fail authentication or have mismatched keys are rejected and an alarm count is triggered, thus forming point cloud decrypted frame objects and video decrypted frame objects. The system outputs two sets of decrypted frame objects as inputs for error recovery and timing correction.
[0049] Forward error correction recovery and missing frame compensation are performed on point cloud decryption frame objects and video decryption frame objects. In scenarios with forward error correction redundancy, the coding blocks are reassembled according to the source identifier and frame sequence number and erasure coding is solved. Equivalent decryption frame objects are generated for recoverable positions, and missing frame placeholders are registered for unrecoverable positions. Subsequently, based on the clock mapping relationship constructed by the negotiated parameter set, the local clock of the receiving end is aligned to the 3D deployment ball acquisition clock. The timing labels of the two types of decryption frame objects are corrected and jitter smoothed. Cross-modal delay compensation is completed using bounded hysteresis buffer and a timing-consistent frame object sequence is generated. The timing-aligned point cloud decryption frame object sequence and the timing-aligned video decryption frame object sequence are output as inputs for payload parsing and data standardization.
[0050] The temporally aligned point cloud decryption frame object sequence is parsed into fields such as spatial coordinates, surface normal vectors, and point density according to a predefined load structure. The field order and numerical encoding are unified to form standardized laser point cloud data. The temporally aligned video decryption frame object sequence is parsed into fields such as texture matrix, edge response, and motion vector. The color space, resolution, and temporal baseline are unified to form standardized video data. Both types of standardized results inherit the source identifier and the corrected temporal label, which serve as the unified input for subsequent distributed parallel feature extraction. Thus, after decryption, it directly enters the cross-node consistent feature calculation process.
[0051] Furthermore, after obtaining the point cloud decryption frame object sequence and the video decryption frame object sequence at the receiving end, a clock mapping sample set is first constructed using the time series labels and arrival times of the two types of decryption frame objects. Stable sample pairs are obtained by removing abnormal transition points, and the 3D deployment ball acquisition clock is mapped to the unified time axis of the receiving end using an affine mapping. The affine mapping is as follows:
[0052]
[0053] in, This indicates the timing tag carried by the decrypted frame object; 'a' represents the corrected timing label mapped to a unified time axis; 'a' represents the clock drift coefficient, a positive real number close to 1, estimated on the sample pair set using weighted least squares; 'b' represents the clock offset, a real number, estimated by... The simultaneous solution yields the following: The mapping principle is to synchronize the clocks at both ends using affine transformation, so that the same physical moment is aligned on a unified time axis, thereby providing a unified time reference for cross-modal alignment. The output point cloud timing labels corrected by affine mapping and the video timing labels corrected by affine mapping are used as inputs for jitter smoothing.
[0054] After obtaining the affine mapping-corrected time series labels, to suppress short-period jitter introduced by the network and decoding, first-order exponential smoothing is performed on both types of corrected time series labels, and the smoothing update is as follows:
[0055]
[0056] in, Indicates the first Smooth temporal labels for each frame; This represents the smoothing coefficient, which takes the value of a real number within an interval and is selected by statistically analyzing the jitter variance. Indicates the first The corrected timing labels for each frame are updated using a recursive low-pass filter to reduce high-frequency jitter and maintain timing monotonicity. The smoothed point cloud timing labels and the smoothed video timing labels are then used as inputs for unified timeline sampling.
[0057] A sampling time sequence is generated on a unified time axis based on the target frame rate and the start time. Two types of frames are then aligned to the same sampling time using a combination of nearest neighbor thresholding and small-gap linear interpolation. The sampling time sequence is as follows:
[0058]
[0059] in, Indicates the first A unified sampling time; Indicates the start time of the unified timeline; The target frame rate is represented by a positive real number, selected through downstream parallel processing constraints; the alignment principle is based on tolerance. The nearest neighbor matching is used within the range, if If direct pairing is not possible within the small gap threshold, linear interpolation is performed on adjacent frames to construct a placeholder frame. If the small gap threshold is exceeded, a missing frame placeholder mark is registered, and the output is processed accordingly. Aligned point cloud frame sequence and by Aligned video frame sequences are used as input for data format standardization.
[0060] In the standardization of video data formats, the first step is to standardize the format according to... The aligned video frame sequence undergoes decoding and color space unification, set to a preset color space and bit depth. Resolution normalization and aspect ratio preservation alignment padding are then performed. Finally, the time baseline is written to a unified timeline. The frame header field is rewritten with the source identifier and frame number to form standardized video data. To ensure the stability of subsequent texture and edge statistics, contrast-limited histogram equalization is performed on the grayscale or luminance channels of each frame, and the equivalent gain factor is recorded as a quality metric. The final output fields are in the following order: source identifier, correction timing label, frame number, color space identifier, resolution, pixel matrix, and quality metric, which serve as the unified input for the video side of distributed parallel feature extraction.
[0061] In the standardization of laser point cloud data formats, the first step is to standardize the data according to... The aligned point cloud frame sequence is spatially uniformized by mapping the point cloud coordinates of each frame to the global reference coordinate system according to the calibration extrinsic parameters, and uniform resampling is performed using a voxel grid to stabilize the point density. Subsequently, the surface normal vector and curvature are estimated in a fixed-radius neighborhood. The normal vector and curvature are obtained by eigenvalue decomposition of the neighborhood covariance, and the corresponding estimates are:
[0062]
[0063] Where n represents the surface normal vector, and its value is a unit vector; This represents the eigenvector corresponding to the smallest eigenvalue of the neighborhood covariance. A normalized measure of curvature, taking values of real numbers within an interval; The smallest eigenvalue representing the neighborhood covariance. , , The eigenvalues are sorted in ascending order. The principle is to characterize the local geometric principal axis with the principal component direction, the smallest eigenvector is perpendicular to the local fitting plane, and the eigenvalue ratio measures the strength of local curvature. The number of points per unit volume within each voxel is used as the point density, and the echo intensity is robustly normalized to suppress quantization bias. The final output fields are in the following order: source identifier, correction time sequence label, frame number, spatial coordinates in the global reference coordinate system, surface normal vector, point density, curvature, and normalized echo intensity when available, forming standardized laser point cloud data, which serves as the unified input to the point cloud side of distributed parallel feature extraction.
[0064] In cross-modal consistency verification, for the same The system checks the source identifier of the video frames and point cloud frames, corrects the bidirectional mapping relationship between the timing label and the frame number, and if there are missing frame placeholders, it synchronously registers the placeholders and time tolerances in the standardized video data and standardized laser point cloud data, and downloads the corresponding quality metric to the subsequent calculation stage. After the verification is completed, it outputs pairs of standardized laser point cloud data and standardized video data, and maintains a unified time axis, unified field order and unified source identifier, so that subsequent cross-node parallel feature extraction can be directly consumed without further timing or structural shaping.
[0065] S130 distributes the standardized laser point cloud data and video data to multiple computing nodes, and each computing node performs feature extraction operations in parallel.
[0066] In one possible implementation, standardized laser point cloud data and video data are distributed to multiple computing nodes, with each node performing feature extraction operations in parallel. Specifically, this includes: generating time-window batches containing standardized laser point cloud data and video data, and distributing these batches to multiple computing nodes; performing image preprocessing, texture enhancement, and motion estimation on the video data within each computing node, extracting texture statistics, edge orientation histograms, and motion vector fields, and concatenating them to form a multi-dimensional feature vector sequence with temporal labels, frame numbers, and source identifiers; performing spatial resampling and geometric estimation on the laser point cloud data within each computing node, calculating spatial coordinates, surface normals, and point density, and combining curvature, flatness, and linearity to form a three-dimensional geometric feature vector sequence with temporal labels, frame numbers, and source identifiers; and performing temporal aggregation and dimensionality reduction processing within the time window on the multi-dimensional feature vector sequence and the three-dimensional geometric feature vector sequence within each computing node, forming a node-level feature package with integrity flags and quality metrics, and outputting it through a distributed message queue.
[0067] Specifically, time window batches are generated with a fixed time window width and sliding step size under a unified time axis. The task scheduler packages standardized laser point cloud data and standardized video data into the same window, binding each time window batch with a source identifier, timing label range, and frame sequence number interval, and registering an integrity flag to indicate missing frames and space occupation. Subsequently, load balancing is performed based on the computing power, memory, and queue depth of the computing nodes, distributing the standardized laser point cloud data and standardized video data within the same time window to the target computing nodes to ensure cross-modal timing consistency. During this process, the time window width... With sliding step size Quota constraints are used to maintain a balance between node throughput and end-to-end latency, with frame counting within the time window. satisfy:
[0068]
[0069] in, Indicates the target frame rate; This indicates a floor function; the above constraint is used to ensure that the number of frames contained in each batch is stable and matches the parallel pipeline cycle time.
[0070] Within the computation node, image preprocessing, texture enhancement, and motion estimation are performed on the video data. First, denoising and color space homogenization are performed without altering the source identifier and temporal label. Then, contrast-limited histogram equalization is applied to the luminance channel to stabilize texture statistics. Subsequently, edge responses are calculated and quantized into edge orientation histograms. The first edge orientation histogram is defined as follows: The directional components are:
[0071]
[0072] in, This represents the set of pixels in the current video frame. Indicates the pixel gradient direction; Indicates the first One directional interval; This is an indicator function; it is 1 for true and 0 for false. Representing pixel gradient magnitude; then estimating dense motion vector field based on pyramid optical flow. The system then statistically analyzes velocity amplitude histograms and principal directions on the grid. A concatenated vector of texture statistics, edge direction histograms, and motion vector statistics is constructed, along with sequence labels, frame numbers, and source identifiers, to form a multidimensional feature vector sequence for subsequent cross-modal correlation. The aforementioned "multidimensional feature vector" refers to the set of statistical descriptions of a single frame of video in three dimensions: texture, edge, and motion; the "motion vector field" refers to the pixel-level displacement vector field obtained by optical flow calculation between adjacent frames.
[0073] Spatial resampling and geometric estimation are performed on the laser point cloud data within the computing node. First, a voxel grid is used with voxel side lengths... Uniform resampling is performed on the point cloud, and a voxel-represented point sampling strategy ensures stable point density; subsequently, in a fixed-radius neighborhood... The covariance matrix is calculated and eigenvalue decomposition is performed to obtain the surface normal vector and curvature. The geometric estimate is defined as:
[0074]
[0075] in, Represents the neighborhood covariance matrix; Indicates the neighborhood centroid; express eigenvalues; for The corresponding eigenvectors are used as surface normal vectors; Curvature is defined as a measure; flatness is then defined based on this. With linearity Used to depict local structures:
[0076]
[0077] in, A larger value indicates a better fit to the local planar structure. A larger value indicates a better fit to the local line structure; simultaneously, the point density is obtained by counting the number of points per unit volume. Spatial coordinates, surface normal vectors, point density, curvature, flatness, and linearity are combined with temporal tags, frame numbers, and source identifiers to form a three-dimensional geometric feature vector sequence. The aforementioned "spatial resampling" refers to reducing the impact of non-uniform sampling on statistics through voxelization; "surface normal vector" refers to the unit normal vector in the direction of the minimum eigenvalue, used to describe the local surface orientation; and "point density" refers to the number of points per unit volume.
[0078] Within a computing node, temporal aggregation and dimensionality reduction are performed on the multidimensional feature vector sequence and the three-dimensional geometric feature vector sequence within a time window. First, a bounded hysteresis buffer is used to register frames within the same time window according to their temporal labels, and placeholder markers and integrity flags are registered for missing frame locations. Then, the statistical measures are aggregated along the frame dimension to obtain the mean vector. Diagonal approximation with covariance and with a linearly interpretable mapping matrix To perform fidelity-preserving dimensionality reduction, node-level aggregation is defined as follows:
[0079]
[0080] in, Represents a node-level embedding vector; This indicates vector concatenation; Energy for preserving key statistics through offline least squares fitting or online recursive least squares; quality metrics Overall bad frame rate Occupancy rate With signal-to-noise ratio And given:
[0081]
[0082] in, The weights are non-negative and satisfy the normalization condition; This indicates the proportion of frames judged as bad. Indicates the proportion of interpolated frames; This represents the data signal-to-noise ratio based on gradient magnitude stability and local covariance condition number. Ultimately, it will include... Integrity flags and quality metrics The node-level feature packets are output through a distributed message queue. The message header carries the time window boundary, source identifier, and frame sequence number range to support downstream consistency verification. The aforementioned "time-series aggregation" refers to summarizing frame-level statistics in the time direction within a fixed time window; "dimensionality reduction" refers to compressing redundant dimensions with a linear mapping while retaining task-related energy; "distributed message queue" refers to middleware that provides an ordered, reusable, and traceable data transmission channel in a distributed environment, used to decouple computing nodes and fusion nodes and ensure data arrival order and recoverability.
[0083] S140 performs cross-modal correlation calculations on the multi-dimensional feature vectors output by each computing node and the three-dimensional geometric feature vectors, and uses the timestamp alignment results to match the multi-dimensional feature vectors and the three-dimensional geometric feature vectors, thereby generating a fused feature set.
[0084] In one possible implementation, the multidimensional feature vectors output by each computing node are subjected to cross-modal correlation calculation with the three-dimensional geometric feature vectors. The timestamp alignment results are then used to match the multidimensional feature vectors with the three-dimensional geometric feature vectors, thereby generating a fused feature set. Specifically, this includes: performing bounded hysteresis alignment and frame-missing interpolation compensation on the multidimensional feature vector sequence and the three-dimensional geometric feature vector sequence based on a unified temporal label, forming a temporally aligned multidimensional feature vector sequence and a three-dimensional geometric feature vector sequence; and using extrinsic and intrinsic parameter matrices to project the spatial coordinates in the temporally aligned three-dimensional geometric feature vectors onto a pixel plane, and combining this with the temporal alignment... The texture matrix and edge response map in the subsequent multidimensional feature vectors are used to establish a candidate cross-modal pairing set. The candidate cross-modal pairing set is matched and scored according to temporal compatibility, spatial proximity and feature similarity. A bipartite graph is constructed to perform maximum weight matching solution to obtain the cross-modal matching result set. The cross-modal matching result set is combined with the multidimensional feature vector and the three-dimensional geometric feature vector based on the matching score and quality metric to output the fused feature vector. In the case of one-to-many or many-to-one pairing, robust pooling and uncertainty screening are performed on the fused features in the same cluster to finally form the fused feature set as the input for the construction of the multimodal three-dimensional scene representation.
[0085] Specifically, in the distributed fusion unit, bounded hysteresis alignment is established using unified temporal labels. Multidimensional feature vector sequences and three-dimensional geometric feature vector sequences are simultaneously placed into a windowed temporal buffer. Window alignment is completed within the maximum hysteresis threshold, and frame-gap interpolation compensation is performed for small gaps caused by network jitter or decoding overhead. For the video side, when the temporal labels of two adjacent valid frames fall on either side of the same unified sampling time, placeholder features are generated through linear interpolation in the feature space. For the point cloud side, linear interpolation or nearest-neighbor placeholders are performed on the three-dimensional geometric feature vectors at the same unified sampling time, resulting in two types of temporally aligned sequences. The interpolation uses the following formula, which is calculated independently for the video side and the point cloud side:
[0086]
[0087] in, Indicates the time of uniform sampling The occupant feature vector at the location; and Indicated by time sequence label and Registered adjacent valid feature vectors; The linear interpolation coefficients are represented, and their values are real numbers within an interval. "Bounded hysteresis alignment" refers to buffering and rearranging within a fixed maximum hysteresis threshold to ensure cross-modal temporal consistency. "Frame missing interpolation compensation" refers to maintaining the integrity of the time axis by using feature space linear interpolation or nearest neighbor occupancy when short-term data is missing.
[0088] After temporal alignment is completed, pixel-plane projection is performed on the spatial coordinates in the 3D geometric feature vector based on the calibrated extrinsic and intrinsic parameter matrices. This projection is then used to establish a pixel-domain correspondence with the texture matrix and edge response map in the multidimensional feature vector, forming a candidate cross-modal pairing set. The pixel-plane projection process is expressed as follows:
[0089]
[0090] in, Represents a three-dimensional point in the global reference coordinate system; and This represents the rotation and translation of the extrinsic parameter matrix, with values obtained through offline or online calibration. This represents the intrinsic parameter matrix, which includes focal length and principal point parameters; Represents perspective division from homogeneous coordinates to the pixel plane; Represents pixel coordinates. Obtained by projection. Define a local pixel neighborhood centered on the texture matrix and edge response map, if If a feature vector falls into the neighborhood and satisfies the same temporal label and source identifier, then the 3D geometric feature vector and the corresponding multi-dimensional feature vector are paired to form a candidate cross-modal pair. The "extrinsic parameter matrix" refers to the rigid body transformation between cross-modal coordinate systems. The "intrinsic parameter matrix" refers to the internal parameters of the imaging geometry. The "pixel plane projection" refers to the mapping of a 3D point to a 2D pixel coordinate through the imaging model. The "candidate cross-modal pair set" refers to the set of cross-modal feature pairs that satisfy the temporal and spatial neighborhood consistency conditions.
[0091] The matching score is calculated for the candidate cross-modal pairing set, and a maximum weight matching of the bipartite graph is performed. The multidimensional feature vector and the three-dimensional geometric feature vector are respectively used as two vertex sets of the bipartite graph. The optimal one-to-one correspondence is solved using the matching score as the edge weight. The matching score is defined as:
[0092]
[0093] in, This represents the time-series label difference between cross-modal pairings, and its value is a non-negative real number. This represents the time-series scale parameter, and its value is a positive real number. This represents the Euclidean distance from the projection point on the pixel plane to the center of the local pixel neighborhood, and its value is a non-negative real number. This represents a spatial proximity scale parameter, and its value is a positive real number. and These represent the same-dimensional aligned feature sub-vectors derived from multidimensional feature vectors and three-dimensional geometric feature vectors, respectively; Represents the inner product operator; Represents the L2 norm; Describes non-negative weight coefficients that satisfy... Maximum weight matching with bipartite graph decision variables. The expression is as follows:
[0094]
[0095] in, Indicates the first The multidimensional feature vector and the first Matching score of three-dimensional geometric feature vectors; The matching selection variable takes a Boolean value; the two sets of inequality constraints each restrict each vertex to be matched at most once; "bipartite graph" refers to a graph in which the vertices are divided into two non-overlapping classes and the edges cross only the two classes; "maximum weight matching" refers to the matching selection that maximizes the sum of the edge weights under the matching constraints.
[0096] After obtaining the cross-modal matching result set, confidence-weighted synthesis is performed on the multi-dimensional feature vectors and three-dimensional geometric feature vectors of each matching pair to output a fused feature vector. In cases of one-to-many or many-to-one relationships, robust pooling and uncertainty filtering are applied to clustered fused features to suppress anomalies. The fusion follows a linear normalization rule.
[0097]
[0098] in, and Represents the eigenvectors after mapping to the same dimension; and This represents the confidence weight, with a value that is a non-negative real number, and can be determined by the match score and the quality measures of both sides. and The monotonic combinations are given, for example
[0099]
[0100] in, This represents the weighting coefficient, with values being real numbers within a range; the "quality metric" refers to the credibility index, which is a convergence of the bad frame ratio, occupancy ratio, and signal-to-noise ratio carried by the node-level feature packets. For robust pooling of multiple pairings within the same cluster, a weighted truncated mean can be used.
[0101]
[0102] in, Represents a set of matching pairs within the same cluster; Indicates the first Within each cluster, there are fusion candidates; This represents the estimate of the median within the cluster; The threshold value represents a positive real number. The weights of fusion candidates are adaptively reduced based on their outlier tendencies, and low-confidence candidates are filtered out using an upper bound on the covariance or an equivalent uncertainty index. The selected set of fusion feature vectors is retained as input for constructing the multimodal 3D scene representation. "Robust pooling" refers to statistically synthesizing multiple samples within a cluster while suppressing the influence of outliers. "Uncertainty filtering" refers to eliminating fusion results with insufficient confidence based on thresholds estimated using weights and dispersion.
[0103] S150, based on a fusion feature set, constructs a multimodal 3D scene representation, performs intelligent and automated surveying and analysis of the work site, and outputs the work site analysis results.
[0104] In one possible implementation, a multimodal 3D scene representation is constructed based on a fused feature set, and intelligent automated surveying and analysis of the work site is performed, outputting the work site analysis results. Specifically, this includes: constructing a voxel grid in a global reference coordinate system using spatial coordinates and surface normal vectors from the fused feature vectors, and performing voxel-level fusion updates based on a truncated signed distance function to form an attributed triangular mesh containing geometric and visual attributes; performing terrain segmentation on the attributed triangular mesh by combining elevation, normal vector tilt angle, point density, and motion vector statistics to generate terrain labels, initial obstacle labels, and column-type structure labels; and performing boundary compactness, voxel occupancy consistency, and texture boundary consistency calculations based on connected components of the initial obstacle labels, and introducing dynamic motion vector statistics. The discrimination probability is updated to update the obstacle probability, resulting in a multimodal 3D scene representation with obstacle probability and geometric circumscribed parameters. Crossing objects are identified from the multimodal 3D scene representation, and the minimum static gap and minimum dynamic gap are calculated. The gap parameters of the crossing objects, obstacle probability, and terrain slope are combined to form a risk index vector. Based on the weighted risk score, the spatial boundaries of the prohibited work area, the warning work area, and the permitted work area are generated, forming a multimodal 3D scene representation with work safety boundaries. Based on the multimodal 3D scene representation with work safety boundaries, the connectivity of the passable corridors, equipment placement locations, and work paths in the permitted work area is calculated. The output is a work site analysis result including terrain distribution, obstacle location and size parameters, crossing situations, work safety boundaries, and work path suggestions.
[0105] Specifically, a 3D voxel grid is constructed using the spatial coordinates in the fused feature vectors in the global reference coordinate system. The surface normal vectors in the fused feature vectors are used as geometric observations. Each fused feature vector is projected onto the voxel grid along the line of sight, and the truncation distance is calculated. The truncation signed distance function is used to incrementally fused and update the geometric state of the voxels. The voxel update follows a weighted average rule with normalized weights. The truncation signed distance of the voxels is iterated with the weights using the following formula:
[0106]
[0107] in, The sign distance representing the voxel history truncation is a bounded real number. The observation truncation sign distance is calculated from the fused feature vector and the projective geometry, and its value is a bounded real number. This represents the historical weight of the voxel, and its value is a non-negative real number. The observation weights, represented by non-negative real numbers, are obtained through monotonic mapping of the quality metric and incident angle consistency of the fused feature vector. Isosurface extraction is performed on the updated voxel raster to obtain triangular meshes. The surface normals and point densities in the fused feature vectors are propagated into geometric attributes through voxel-to-mesh vertex interpolation. The texture statistics and edge orientation histograms in the fused feature vectors are mapped onto the triangular meshes via keyframe backprojection, and confidence-weighted synthesis is performed in overlapping regions to form visual attributes, thus obtaining attributed triangular meshes containing both geometric and visual attributes.
[0108] On an attributed triangular mesh, a vertex feature vector is constructed for each vertex, consisting of elevation, normal vector tilt angle, point density, and motion vector statistics. Graph structure segmentation is performed on the vertex feature vectors, and terrain labels, initial obstacle labels, and column structure labels are output. The vertex feature vector is denoted as... ,in This represents the elevation in the global reference coordinate system, and its value is a real number. The angle between the normal vector and the direction of gravity is a real number within a certain interval. The point density is the number of points per unit volume, and its value is a non-negative real number. The amplitude statistics of the motion vector are represented, with values taking the form of non-negative real numbers. The Markov random field energy is used as the partitioning criterion, and the energy function is...
[0109]
[0110] in, A set of vertex labels; For the first Each vertex label can be one of the following: terrain label, initial obstacle label, or column structure label. For based on The single-point cost, which takes the value of a non-negative real number, is obtained through a logical mapping of elevation threshold, tilt angle threshold, and point density threshold. The cost of consistency between adjacent vertices is a non-negative real number, expressed as a function of the difference between the grid edge weights and the normal vector. This is the set of adjacent edges in the grid. The energy function is minimized to obtain the segmentation result, which is then assigned a corresponding label.
[0111] Boundary compactness, voxel occupancy consistency, and texture boundary consistency are calculated on the connected components of the initial obstacle labels. Dynamic discrimination probabilities are then obtained by combining motion vector statistics, and Bayesian updates are performed on the obstacle probabilities. Boundary compactness is represented by the perimeter-to-area ratio, defined as follows:
[0112]
[0113] in, Let be the perimeter of the connected component, and let its value be a non-negative real number. Let be the projected area of the connected component, and take the value of a non-negative real number. Smaller values indicate greater compactness. Voxel occupancy consistency is characterized by the stability of the sign histogram of voxel truncation sign distances, denoted as [symbol missing]. The value is a real number within an interval. Texture boundary consistency is characterized by the overlap between the edge response map and the boundary of the connected component, denoted as . The value is a real number within an interval. These three factors are combined using logical mapping to form the observed likelihood. , with historical obstacle probability Update:
[0114]
[0115] in, The updated obstacle probability takes the value of a real number within a range; The probability of a historical obstacle is represented, and its value is a real number within a range. For the reason , and The combined observation likelihood is taken as a real number within an interval. For each connected component, the geometric bounding parameters are simultaneously calculated, including the center, size, and orientation of the minimum bounding rectangle or oriented bounding box, with values being real numbers and unit vectors, respectively, for subsequent gap and risk assessment.
[0116] This paper identifies crossing objects on a multimodal 3D scene representation with obstacle probability and geometric circumference parameters. Approximately linear or overhanging components are selected as crossing objects based on their geometric shape and connectivity. Minimum static and minimum dynamic gaps are calculated. The minimum static gap is defined as the minimum distance from the crossing object to the terrain surface and the geometry enclosed by the obstacle, denoted as […].
[0117]
[0118] in, For the set of points spanning the object; Let be the set of points on the environmental surface, and let be the coordinates of these points in the global reference coordinate system. The minimum dynamic gap, considering relative motion, is defined as follows:
[0119]
[0120] in, and For time The set of locations of the objects to be crossed and the environment is obtained by statistical extrapolation of motion vectors. A risk index vector is constructed by combining the object gap parameter, obstacle probability, and terrain slope, and a weighted risk score is used to generate the spatial boundary.
[0121]
[0122] in, This is a weighted risk score, with values being non-negative real numbers; and The minimum gap is represented by a non-negative real number. It is a stability constant, taking the value of a positive real number and being much smaller than the conventional gap size; The probability of an obstacle is represented, and its value is a real number within a range. It is a normalized measure of topographic slope, and its value is a real number within an interval. , , , The weights are non-negative and satisfy normalization. Based on the threshold... By dividing the area into layers, the spatial boundaries of the prohibited work zone, the warning work zone, and the permitted work zone are obtained.
[0123] In a multimodal 3D scene representation with operational safety boundaries, a passable corridor map is constructed based on the work permit area. Node sets are generated for equipment placement locations and key operational points, and edge sets are generated for feasible corridors. Edge costs are defined using distance, slope, and risk score. Connectivity is evaluated using constrained shortest path methods, and operational path suggestions are generated. Edge cost is defined as...
[0124]
[0125] in, For the edge The cost is a non-negative real number; The length of the side is a positive real number. This represents the average slope along the edge, and its value is a non-negative real number. This represents the average weighted risk score along the edge, and its value is a non-negative real number. , , The weights are non-negative and normalized. A path search is performed on the entire map, constrained by the work safety boundary and the equipment work radius. The results are obtained by determining the connectivity of passable corridors, the accessibility of equipment placement locations, and the cost ranking of multiple candidate work paths. The output includes the work site analysis results, which include the terrain distribution, obstacle location and size parameters, crossing conditions, work safety boundary, and work path suggestions.
[0126] S160 automatically generates work plans with spatial geometric constraints and work safety boundaries based on multimodal 3D scene representation and work site analysis results.
[0127] In one possible implementation, based on the multimodal 3D scene representation and the results of the work site analysis, a work plan with spatial geometric constraints and work safety boundaries is automatically generated, specifically including:
[0128] A task constraint graph containing geometric entities, passable corridors, equipment placement locations, and suggested work paths is constructed in a global reference coordinate system. Work prohibition zones, work warning zones, and work permission zones are generated within the task constraint graph in conjunction with work safety boundaries. A joint optimization model of paths and processes is established on the task constraint graph, using path segment selection, equipment placement location selection, and work unit timelines as decision variables. A joint optimization problem containing path objective functions and scheduling objective functions is constructed, and spatial geometric constraints, work safety boundary constraints, and process sequence constraints are applied. The joint optimization problem is iteratively solved to generate path geometry, timeline curves, and resource allocation that satisfy the work safety boundaries. Dynamic gap monitoring instructions, emergency backtracking corridors, and speed and dwell time limits for warning zone boundaries are added to candidate paths. The path geometry, timeline curves, and resource allocation are instantiated into an executable instruction set, and real-time replanning trigger conditions and rollback points are set in risk areas to form a final work plan with spatial geometric constraints and work safety boundaries.
[0129] Specifically, in a global reference coordinate system, the multimodal 3D scene representation and the results of the on-site work analysis are mapped into a task constraint graph. First, geometric entities are extracted using attributed triangular meshes to generate a directed graph of passable corridors. Equipment placement locations and key work sites are registered as nodes, candidate work path segments are registered as edges, and the work safety boundary is interpreted as a three-part set. Then, each edge is assigned a comprehensive cost and a crossing permission flag. The comprehensive cost is a weighted sum of distance, slope, and risk score, defined as:
[0130]
[0131] in, Representing an edge The overall cost; Indicates the side length; This represents the average slope along the edge; This represents the average risk score along the edge; The weights are non-negative and satisfy normalization. Based on the work safety boundary, edges falling into the work prohibition zone are marked as unavailable, edges crossing the work warning zone boundary are recorded with speed and dwell limit, and edges located in the work permission zone are set as available and redundant corridor information is retained. This results in a task constraint graph that includes geometric entities, passable corridors, equipment placement positions and work path suggestions, and embeds work prohibition zone, work warning zone and work permission zone.
[0132] A joint optimization model of path and process is established on the task constraint graph. Path segment selection, equipment placement selection, and work unit schedule are used as decision variables. A weighted, multi-objective approach is adopted to simultaneously minimize the path objective function and the scheduling objective function. The joint objective and main constraints are defined as follows:
[0133]
[0134]
[0135] in, Indicates the path segment selection variable; Indicates the variable for selecting equipment placement; Indicates the time schedule of the work unit; For edge cost; The target weight takes values within the specified range; The scheduling weight has a non-negative and normalized value; For the maximum completion time; Work unit In pose trajectory Risk score at the location; For the set of flow conservation and start / end point constraints; This is the set of equipment placement reachability and line-of-sight constraints; It is a set of constraints including process sequence, non-overlapping resources, and time window. and These represent the feasible and infeasible subsets of the permitted and prohibited work areas in the task constraint graph. To mitigate risk exposure within the work warning area, velocity and dwell constraints are introduced for crossing the edge of the work warning area:
[0136]
[0137] in, For the border Execution speed; For the border Duration of stay; and These are the speed and loitering limits for the restricted area.
[0138] The joint optimization problem is solved iteratively, employing a hierarchical decomposition strategy to satisfy spatial geometric constraints, operational safety boundary constraints, and process sequence constraints. The first layer solves the problem on the traversable corridor subgraph. and The bounded shortest path constraint is used to obtain the initial path geometry; the second layer applies curvature-constrained smoothing to the initial path geometry in continuous space and fine-tunes the path with a gap gain term, making it feasible while reducing [the risk of failure]. The third layer uses mixed-integer programming or constraint propagation with priorities. Internal optimization of timeline curves and resource allocation, outputting results that satisfy both process sequence and non-overlapping resources. Parallelism configuration. For path segments with a high probability of crossing objects and obstacles, dynamic gap monitoring instructions and emergency backtracking corridors are automatically attached. For segments crossing the operation warning zone, speed limits and dwell limits are written, so that the iterative solution checks the operation safety boundary and risk exposure threshold after each update until the objective and constraints converge, and finally obtains the path geometry, time history curve and resource allocation that meet the operation safety boundary.
[0139] The path geometry, time-history curves, and resource allocation are instantiated into an executable instruction set. A task-oriented instruction template is used to generate a time-series of position, velocity, attitude, dwell, and sensor trigger instructions. Each instruction is bound to a source identifier, time-series label, and quality metric. For path segments marked as high-risk, real-time replanning trigger conditions are written, including dynamic gap thresholds, obstacle probability thresholds, and attitude drift thresholds. A rollback point and rollback corridor are configured for each trigger condition, allowing the executor to unambiguously roll back to a safe state when a threshold is triggered. To ensure online robustness, check instructions for rotating key verification and data consistency verification are appended to the end of the instruction set. A version number and hash digest are recorded for each job unit, achieving end-to-end traceability and auditability of the job plan. Finally, a job plan with spatial geometric constraints and job safety boundaries is formed and can be directly issued for execution.
[0140] This embodiment also discloses a device for fusing laser point clouds and video based on a 3D controlled sphere, referring to... Figure 2 The device includes an acquisition module 201, a processing module 202, and an output module 203, and is used to execute any of the above-described methods for fusing laser point clouds and video based on a 3D controlled sphere, wherein:
[0141] The acquisition module 201 is used to encrypt laser point cloud data and video data acquired by the 3D control ball during data transmission.
[0142] The processing module 202 is used to decrypt the laser point cloud data and video data at the receiving end, and to perform timestamp alignment and data format standardization processing on the laser point cloud data and video data respectively.
[0143] The processing module 202 is used to distribute the standardized laser point cloud data and video data to multiple computing nodes. Each computing node performs feature extraction operations in parallel. Specifically, it extracts multi-dimensional feature vectors based on image texture, edges and motion vectors for video data, and extracts three-dimensional geometric feature vectors based on spatial coordinates, surface normal vectors and point density for laser point cloud data.
[0144] The processing module 202 is used to perform cross-modal correlation calculations on the multi-dimensional feature vectors output by each computing node and the three-dimensional geometric feature vectors, and to match the multi-dimensional feature vectors with the three-dimensional geometric feature vectors using the timestamp alignment results, thereby generating a fused feature set.
[0145] The processing module 202 is used to construct a multimodal 3D scene representation based on the fused feature set, and to perform intelligent and automated survey and analysis of the work site, and output the work site analysis results.
[0146] The output module 203 is used to automatically generate a work plan with spatial geometric constraints and work safety boundaries based on the multimodal 3D scene representation and work site analysis results.
[0147] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0148] This embodiment also discloses an electronic device, as shown in the reference. Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.
[0149] The communication bus 302 is used to enable communication between these components.
[0150] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0151] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0152] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications. The GPU is responsible for rendering and drawing the content required for display. The modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0153] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc. The data storage area may store data involved in the various method embodiments described above. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. As a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface 303 module, and an application program for a method of laser point cloud and video fusion based on a 3D control sphere.
[0154] exist Figure 3 In the illustrated electronic device, the user interface 303 is primarily used to provide an input interface for the user and to acquire user input data. The processor 301 can be used to call an application stored in the memory 305 that describes a method for fusing laser point clouds and video based on a 3D controlled sphere. When executed by one or more processors 301, the electronic device performs one or more methods as described in the above embodiments.
[0155] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0156] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0157] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0158] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0159] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0160] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.
[0161] The present invention also discloses a computer-readable storage medium storing instructions. When executed by one or more processors 301, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.
[0162] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This invention is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for fusing laser point clouds and video based on a 3D controlled sphere, characterized in that, The method includes: After acquiring laser point cloud data and video data using a 3D surveillance sphere, the data is encrypted during transmission. At the receiving end, the laser point cloud data and video data are decrypted, and the laser point cloud data and video data are respectively processed by timestamp alignment and data format standardization. The standardized laser point cloud data and video data are distributed to multiple computing nodes, and each computing node performs feature extraction operations in parallel. Specifically, for the video data, multi-dimensional feature vectors based on image texture, edges and motion vectors are extracted, and for the laser point cloud data, three-dimensional geometric feature vectors based on spatial coordinates, surface normal vectors and point density are extracted. The multidimensional feature vectors output by each computing node are subjected to cross-modal correlation calculation with the three-dimensional geometric feature vectors, and the multidimensional feature vectors are matched with the three-dimensional geometric feature vectors using the timestamp alignment results, thereby generating a fused feature set; Based on the fused feature set, a multimodal 3D scene representation is constructed, and intelligent automated surveying and analysis of the work site are performed, outputting the work site analysis results; Based on the multimodal 3D scene representation and the analysis results of the work site, an operation plan with spatial geometric constraints and work safety boundaries is automatically generated. The step of performing cross-modal correlation calculations on the multidimensional feature vectors output by each computing node and the three-dimensional geometric feature vectors, and matching the multidimensional feature vectors and the three-dimensional geometric feature vectors using timestamp alignment results to generate a fused feature set, specifically includes: Based on a unified temporal label, bounded hysteresis alignment and frame missing interpolation compensation are performed on the multidimensional feature vector sequence and the three-dimensional geometric feature vector sequence to form a temporally aligned multidimensional feature vector sequence and a three-dimensional geometric feature vector sequence. The spatial coordinates in the temporally aligned 3D geometric feature vector are projected onto the pixel plane using the extrinsic and intrinsic parameter matrices, and a candidate cross-modal pairing set is established by combining the texture matrix and edge response map in the temporally aligned multi-dimensional feature vector. The candidate cross-modal pairing set is matched and scored according to temporal compatibility, spatial proximity and feature similarity, and a bipartite graph is constructed to perform maximum weight matching solution to obtain the cross-modal matching result set; The cross-modal matching result set is used to synthesize the multi-dimensional feature vector and the three-dimensional geometric feature vector with confidence weighting based on matching score and quality metric, and outputs a fused feature vector. In scenarios with one-to-many or many-to-one pairing, robust pooling and uncertainty screening are performed on the fused features in the same cluster to finally form a fused feature set as input for the construction of multimodal three-dimensional scene representation.
2. The method for fusing laser point clouds and video based on a 3D controlled sphere according to claim 1, characterized in that, The process of distributing standardized laser point cloud data and video data to multiple computing nodes, with each computing node performing feature extraction operations in parallel, specifically includes: Generate time window batches containing the standardized laser point cloud data and the video data, and distribute the time window batches to multiple computing nodes; Within each computing node, image preprocessing, texture enhancement, and motion estimation are performed on the video data. Texture statistics, edge orientation histograms, and motion vector fields are extracted and concatenated to form a multidimensional feature vector sequence with temporal labels, frame numbers, and source identifiers. Spatial resampling and geometric estimation are performed on the laser point cloud data within each computing node to calculate spatial coordinates, surface normal vectors and point density, and combine curvature, flatness and linearity to form a three-dimensional geometric feature vector sequence with time label, frame number and source identifier; Within each computing node, temporal aggregation and dimensionality reduction processing are performed on the multidimensional feature vector sequence and the three-dimensional geometric feature vector sequence within a time window to form a node-level feature package with integrity flags and quality metrics, which is then output through a distributed message queue.
3. The method for fusing laser point clouds and video based on a 3D controlled sphere according to claim 1, characterized in that, Based on the fused feature set, a multimodal 3D scene representation is constructed, and intelligent automated surveying and analysis of the work site are performed, outputting the work site analysis results, specifically including: Using the spatial coordinates and surface normal vectors in the fused feature vector, a voxel grid is constructed in the global reference coordinate system, and voxel-level fusion update is performed based on the truncated signed distance function to form an attributed triangular mesh containing geometric and visual attributes. On the attributed triangular mesh, terrain and landform segmentation is performed by combining elevation, normal vector tilt angle, point density and motion vector statistics to generate terrain and landform labels, initial obstacle labels and column structure labels; Based on the initial obstacle label, the connected component performs boundary compactness, voxel occupancy consistency and texture boundary consistency calculations, and introduces dynamic discrimination probability of motion vector statistics to update the obstacle probability, thus obtaining a multimodal 3D scene representation with obstacle probability and geometric circumference parameters. From the multimodal 3D scene representation, identify the crossing object and calculate the minimum static gap and minimum dynamic gap. Combine the gap parameter of the crossing object with the obstacle probability and the terrain slope to form a risk index vector. Based on the weighted risk score, generate the spatial boundaries of the work prohibition zone, the work warning zone and the work permission zone to form a multimodal 3D scene representation with work safety boundaries. Based on the multimodal 3D scene representation with operational safety boundaries, the connectivity of accessible corridors, equipment placement locations, and operational paths within the permitted operational area is calculated. The output includes on-site analysis results containing terrain distribution, obstacle location and size parameters, crossing conditions, operational safety boundaries, and operational path suggestions.
4. The method for fusing laser point clouds and video based on a 3D controlled sphere according to claim 1, characterized in that, The automatic generation of a work plan with spatial geometric constraints and work safety boundaries based on the multimodal 3D scene representation and the work site analysis results specifically includes: A task constraint map containing geometric entities, passable corridors, equipment placement locations, and work path suggestions is constructed in a global reference coordinate system. Work prohibition zones, work warning zones, and work permission zones are generated in the task constraint map by combining work safety boundaries. A joint optimization model of path and process is established on the task constraint graph. The path segment selection, equipment placement location selection and work unit schedule are used as decision variables to construct a joint optimization problem that includes path objective function and scheduling objective function. Spatial geometric constraints, work safety boundary constraints and process sequence constraints are applied. The joint optimization problem is solved iteratively to generate path geometry, time history curves and resource allocation that satisfy the operational safety boundary, and dynamic gap monitoring instructions, emergency back-off corridors and speed and dwell time limits of the warning zone boundary are added to the candidate paths; The path geometry, time-history curve, and resource allocation are instantiated into an executable instruction set, and real-time replanning trigger conditions and rollback points are set in the risk area to form a final operation plan with spatial geometric constraints and operational safety boundaries.
5. The method for fusing laser point clouds and video based on a 3D controlled sphere according to claim 1, characterized in that, After acquiring laser point cloud data and video data using a 3D surveillance sphere, the data is encrypted during transmission, specifically including: After the 3D surveillance ball generates the device identity certificate, the device identity certificate is bound to the hardware root key in the security chip, and the device identity public key and device identity certificate are output. The device identity is verified and an access token and negotiation parameter set are issued. The 3D surveillance sphere generates a session key set based on the negotiation parameter set and the device identity public key. The session key set includes a point cloud session key for the laser point cloud data and a video session key for the video data. The laser point cloud data is protected by the point cloud session key, and the video data is protected by the video session key. A time tag, frame number and source identifier are injected into the frame header, and the frame payload is encrypted with authentication to generate a point cloud ciphertext frame sequence and a video ciphertext frame sequence. During transmission, a replay bitmap with a scrolling window and an accumulation counter are added to the point cloud ciphertext frame sequence and the video ciphertext frame sequence, and a connection-oriented encrypted channel is established based on the access token, with the point cloud ciphertext frame sequence and the video ciphertext frame sequence used as payloads for real-time transmission.
6. The method for fusing laser point clouds and video based on a 3D controlled sphere according to claim 5, characterized in that, The decryption processing of laser point cloud data and video data at the receiving end specifically includes: After completing the access token verification and locating the corresponding session key set, import the point cloud session key and the video session key, initialize the decryption context and reset the replay window, cumulative counter and time tag buffer; The point cloud ciphertext frame sequence and the video ciphertext frame sequence are respectively input into the corresponding decryption context. Replay detection and order rearrangement are performed based on the time tag, the frame number and the source identifier. Illegal ciphertext frames are removed using the replay bitmap of the scrolling window and the cumulative counter. Perform an authenticated integrity check on the encrypted frames that pass the replay test. The encrypted frames that pass the check are decrypted using the point cloud session key and the video session key respectively, and point cloud decrypted frame objects and video decrypted frame objects are output. The time tag, frame number and source identifier are retained in the decrypted frame objects. Forward error correction recovery and missing frame compensation are performed on the point cloud decryption frame object and the video decryption frame object. Timing label correction and cross-modal delay compensation are completed based on the clock mapping relationship constructed based on the negotiated parameter set. The timing-aligned point cloud decryption frame object sequence and video decryption frame object sequence are output. The time-aligned point cloud decryption frame object sequence is parsed into standardized laser point cloud data, and the time-aligned video decryption frame object sequence is parsed into standardized video data.
7. A device for fusing laser point clouds and video based on a 3D controlled sphere, characterized in that, The device is used to perform a method for fusing laser point clouds and video based on a 3D controlled sphere as described in any one of claims 1-6. The device includes an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to encrypt the laser point cloud data and video data acquired by the 3D control ball during the data transmission process. The processing module is used to decrypt the laser point cloud data and video data at the receiving end, and to perform timestamp alignment and data format standardization on the laser point cloud data and the video data respectively. The processing module is used to distribute the standardized laser point cloud data and video data to multiple computing nodes. Each computing node performs feature extraction operations in parallel. Specifically, for the video data, a multi-dimensional feature vector based on image texture, edge and motion vector is extracted, and for the laser point cloud data, a three-dimensional geometric feature vector based on spatial coordinates, surface normal vector and point density is extracted. The processing module is used to perform cross-modal correlation calculation on the multi-dimensional feature vectors output by each computing node and the three-dimensional geometric feature vectors, and to match the multi-dimensional feature vectors and the three-dimensional geometric feature vectors using the timestamp alignment results, thereby generating a fused feature set; The processing module is used to construct a multimodal 3D scene representation based on the fused feature set, and to perform intelligent and automated survey and analysis of the work site, and output the work site analysis results. The output module is used to automatically generate a work plan with spatial geometric constraints and work safety boundaries based on the multimodal 3D scene representation and the work site analysis results.
8. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-6.