Real-time Processing Method for 3D Physical Internet of Things Perception Data Based on Edge Computing
By performing real-time target segmentation and three-dimensional video fusion on edge computing devices, the problems of high latency and cloud computing pressure in traditional monitoring systems are solved, and efficient and accurate monitoring data processing and three-dimensional visualization are achieved.
Patent Information
- Application Number
- CN202510648527.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In traditional monitoring systems, video data needs to be transmitted to the cloud for processing, resulting in high latency, which cannot meet the real-time monitoring needs, and cloud computing pressure is high, existing three-dimensional video fusion technology is prone to introduce errors, and edge computing power equipment lacks computing power, making it difficult to achieve real-time target segmentation and efficient operation.
Video processing and three-dimensional fusion computing are migrated from the central server to edge computing power equipment and user terminals, combined with lightweight deep learning algorithms to achieve real-time target segmentation, and lightweight semantic segmentation model based on the encoder-decoder architecture is used to perform foreground target and background segmentation. Multi-scale enhanced feature fusion module extracts context information of different scales, and matches and fusion of video and three-dimensional scenes on the user terminal device.
It reduces the delay in monitoring data processing, improves the efficiency of convergence, improves the real-time and accuracy of the monitoring system, reduces the pressure of cloud computing, and is suitable for scenarios where network bandwidth is limited or real-time requirements are high.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video fusion, and particularly relates to a real-time processing method for real-scene three-dimensional Internet of Things perception data based on edge computing. Background Art
[0002] In traditional monitoring systems, video data needs to be transmitted to the cloud or central server for processing, resulting in high latency and unable to meet the requirements of real-time monitoring (such as target tracking, response to abnormal events). At the same time, the transmission of a large amount of video data occupies a large amount of network bandwidth, and centralized processing causes heavy pressure on server resources and high costs.
[0003] Existing three-dimensional video fusion technologies rely on complex calibration processes or global perspective transformations, which are prone to introducing errors, resulting in target distortion or blurred model textures. Edge computing devices have insufficient computing power and are difficult to operate efficiently at the edge. Existing technologies are difficult to perform real-time segmentation of the monitoring area, generally requiring a large number of cloud servers for processing, resulting in inaccurate recognition and tracking of monitoring targets, affecting the intelligence level and user experience of the monitoring system. Summary of the Invention
[0004] Therefore, the technical problem to be solved by the present invention is to overcome the defects in the prior art, and proposes a real-time processing method for real-scene three-dimensional Internet of Things perception data based on edge computing devices. This method migrates video processing and three-dimensional fusion calculations from the central server to edge computing devices and user terminal devices, and combines lightweight deep learning algorithms to achieve real-time target segmentation, thereby effectively reducing the processing latency of monitoring data and improving the fusion efficiency.
[0005] To solve the above technical problems, the present invention provides a real-time processing method for real-scene three-dimensional Internet of Things perception data based on edge computing, including the following steps:
[0006] 1) Create a three-dimensional model and load it on the user terminal device;
[0007] 2) The edge computing device receives the monitoring video data, decodes it into a processable frame sequence using the hardware acceleration module, and then performs image preprocessing operations;
[0008] 3) The edge computing device uses a lightweight semantic segmentation model based on the encoder-decoder architecture to process the preprocessed image for foreground target and background segmentation. Among them, the encoder of the lightweight semantic segmentation model uses ResNet-34 as the backbone network, the decoder part uses a progressive upsampling strategy, and combines a multi-scale enhanced feature fusion module to generate a high-resolution segmentation result. The multi-scale enhanced feature fusion module uses a multi-branch structure combined with depthwise separable convolutions to extract context information of different scales;
[0009] 4) The edge computing device synthesizes the segmentation processing results of each frame of image into video data in real time and sends it to the user terminal device;
[0010] 5) The user terminal device maps the video data to the corresponding position of the 3D model to achieve the matching and fusion of the video data and the 3D scene.
[0011] As one of the preferred solutions, the encoder includes the following five steps,
[0012] Step 1, the input image passes through a 3×3 convolutional layer, and is equipped with batch normalization and ReLU activation function;
[0013] Step 2, the first residual block group of ResNet-34 is used for preliminary feature extraction;
[0014] Step 3, the second residual block group is used for further feature extraction;
[0015] Step 4, the third residual block group is used for in-depth feature mining;
[0016] Step 5, the fourth residual block group is used for final feature representation.
[0017] As one of the preferred solutions, the output of each step of the encoder is transmitted to the decoder through skip connections and fused with the output of the previous step of the decoder as the input of the multi-scale enhanced feature fusion module to achieve multi-scale feature fusion.
[0018] As one of the preferred solutions, the specific process of the decoder is as follows:
[0019] Step 6, use the multi-scale enhanced feature fusion module to process the high-level features output by step 5 of the encoder;
[0020] Step 7, fuse the output of decoder step 6 with the features output by encoder step 4, and then use the multi-scale enhanced feature fusion module for fusion processing;
[0021] Step 8, fuse the output of decoder step 7 with the features output by encoder step 3, and then use the multi-scale enhanced feature fusion module for fusion processing;
[0022] Step 9, fuse the output of decoder step 8 with the features output by encoder step 2, and then use the multi-scale enhanced feature fusion module for fusion processing;
[0023] Step 10, fuse the output of decoder step 9 with the features output by encoder step 1, and then use the multi-scale enhanced feature fusion module for fusion processing.
[0024] As one of the preferred solutions, the multi-scale enhanced feature fusion module constructs four independent branches for different processing methods. Among them,
[0025] The first branch uses a 1×3 depthwise separable convolution on the input feature map to extract horizontal features and obtain a first feature map;
[0026] The second branch uses a 3×1 depthwise separable convolution on the input feature map to extract vertical features and obtain a second feature map.
[0027] The third branch first applies a 3×3 standard convolution to the input feature map, then performs normalization and ReLU activation, and further adjusts the feature distribution through a 1×1 convolution to obtain a third feature map.
[0028] The fourth branch directly retains the input feature map as the fourth feature map;
[0029] Finally, the feature maps obtained from these four branches are concatenated in the channel dimension.
[0030] As one of the preferred solutions, the multi-scale enhanced feature fusion module further includes the step of performing channel compression on it through pointwise convolution to obtain the input feature map.
[0031] As one of the preferred solutions, it also includes feature point extraction and matching, including using the SIFT algorithm to extract the image features of the three-dimensional model and the video geometry, then performing mesh division on the video geometry, using the Delaunay triangulation method for triangular mesh division on the regular contour video geometry, and using the constrained Delaunay triangulation method for the non-regular contour geometry.
[0032] As one of the preferred solutions, it also includes using the Laplace mesh deformation algorithm to adjust the vertex coordinates of the video geometry to correct texture deformation.
[0033] The technical solution of the present invention has the following advantages:
[0034] Based on the monitoring visualization method of three-dimensional video fusion, the present invention innovatively introduces edge computing devices, enabling real-time target segmentation of monitoring videos to be completed at the edge, thereby reducing data transmission latency and cloud computing pressure. At the same time, a lightweight neural network is deployed on the edge computing device. Through efficient target segmentation and feature extraction, the accuracy and real-time performance of the fusion of monitoring videos and three-dimensional models are improved using locally deployed user terminal devices, ensuring that efficient intelligent monitoring processing can still be achieved on low-power devices. Detailed implementation
[0035] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0036] Traditionally, the conventional video processing method is usually adopted. First, the video stream data acquired by the monitoring device is transmitted to the cloud server or the local computer for centralized processing. Subsequently, image processing algorithms are used to analyze the video stream data to identify and detect the monitoring targets. For the fusion of video data and 3D models, generally, through a complex camera calibration and coordinate transformation process, the mapping relationship from the image to the 3D model is calculated to achieve texture mapping and rendering of the monitoring video to the 3D model. The traditional video processing method depends on the network transmission and the computing power of the cloud server, and is prone to data transmission and processing delays due to network latency and excessive server load. At the same time, the centralized processing mode is difficult to meet the requirements of real-time and high efficiency in large-scale monitoring systems. In terms of monitoring target recognition, the existing image processing algorithms have limited accuracy in identifying targets in complex scenarios, and cannot segment the monitoring area in real time, making it difficult to accurately distinguish the foreground target (segmented target) and the background, thus affecting the accurate positioning and tracking of the monitoring target. In addition, the scheme based on the complex perspective transformation matrix has a high calculation cost, is not applicable to edge computing devices, and is difficult to be extended to large-scale monitoring scenarios.
[0037] Therefore, the present invention proposes a real-time processing method for real-scene 3D IoT perception data based on edge computing, including the following steps:
[0038] 1) Create a 3D model and load it on the user terminal device; wherein, the 3D data of the monitoring area is mainly obtained through point cloud scanning technology. After the scanning device performs omnidirectional acquisition in the monitoring area and generates the original point cloud data, first, statistical filtering denoising processing is performed on the point cloud to remove abnormal points and noise. Subsequently, a registration and alignment operation is executed to ensure that the point cloud data obtained from multiple scans can be accurately fused, and the voxel grid downsampling technology is used to reduce redundant points, optimizing the data scale while maintaining the integrity of geometric features. The processed point cloud data is converted into a 3D mesh model through an improved greedy projection triangulation algorithm, which can effectively process irregular surfaces and improve the geometric accuracy of the model. To enhance the visual effect of the model, texture mapping technology can also be combined to map the environmental image information onto the mesh surface, making the 3D model have more realistic surface details.
[0039] 2) The edge computing device receives the monitored video data, decodes it into a processable frame sequence using the hardware acceleration module, and then performs image preprocessing operations. Specifically, the video stream data is directly collected by the monitoring camera device and transmitted to the edge computing device in real time. The present invention supports multiple input sources, including RTSP, HTTP streaming media, and MIPI CSI cameras, and adopts the H.264 / H.265 encoding standard to ensure the efficiency of data transmission. At the same time, the basic information metadata of video acquisition, such as resolution, frame rate, field of view angle, and timestamp, is recorded, providing a basis for subsequent data processing, such as providing the necessary original video stream attributes for subsequent video frame extraction and other processes. In the edge computing device (taking OrangePi AIpro as an example), the edge computing device first receives and processes the video stream data. The edge computing device is equipped with a 64-bit multi-core processor and has 8 TOPS of AI computing power, etc., which can fully support complex vision processing tasks. After receiving the video stream, the edge computing device uses a dedicated hardware acceleration module for efficient decoding, converts the original video data into a processable frame sequence and performs frame extraction processing. First, the video is subjected to frame extraction processing, and then the video is output by the frame combination method. When designing the frame extraction, it is necessary to consider the inter-frame difference and time consistency constraints, that is, to consider the video stream processing time and the image foreground image segmentation processing time, etc., which helps to improve the stability and continuity of the output video and enhance the quality performance of the output video. This process is particularly crucial when facing video streams captured under adverse weather conditions, because these images usually contain interference factors such as rain, fog, and low light. After decoding, a series of image preprocessing operations are performed on the frame sequence, including size adjustment, color correction, and noise filtering. These preprocessing steps can initially eliminate the visual interference caused by adverse weather and improve the accuracy and reliability of subsequent analysis. Especially in the noise filtering link, algorithms optimized for different weather conditions are adopted to effectively improve the image quality.
[0040] The entire processing flow makes full use of the multi-core CPU architecture of the edge computing device to construct an efficient parallel processing pipeline, and strictly controls the latency of the preprocessing link within 50 ms. This low-latency processing mechanism lays a solid foundation for subsequent real-time object segmentation, ensuring that the edge computing device can still maintain efficient and accurate visual analysis capabilities in the edge computing environment, and can maintain reliable performance even under poor environmental conditions.
[0041] 3) Process using a lightweight semantic segmentation model based on the Encoder-Decoder architecture to segment and extract foreground objects in the image to be processed. Among them, the encoder of the lightweight semantic segmentation model uses ResNet-34 as the backbone network, the decoder part adopts a progressive upsampling strategy, and combines a multi-scale enhanced feature fusion (MEFF) module to generate a high-resolution segmentation result. The multi-scale enhanced feature fusion (MEFF) module uses a multi-branch structure combined with depthwise separable convolutions to extract context information at different scales;
[0042] Specifically, the present invention proposes a lightweight semantic segmentation model based on the Encoder-Decoder architecture. The encoder part uses ResNet-34 as the backbone network and consists of five stages (Stage1 - Stage5). The encoder extracts features through layer-by-layer downsampling, uses residual blocks to enhance the feature representation ability, and performs feature fusion at different scales. The decoder part adopts a progressive upsampling strategy and combines a multi-scale enhanced feature fusion (MEFF) module to efficiently recover spatial information and finally generate a high-resolution segmentation result.
[0043] 4) The edge computing device synthesizes the segmented image of the target into video data in real time and sends it to the user terminal device. When synthesizing the video data, it is necessary to synchronously combine and record the basic information metadata of video acquisition, such as resolution, frame rate, field of view angle, and timestamp, etc.,
[0044] 5) The user terminal device maps the video data to the corresponding position of the 3D model to achieve the matching and fusion of the video and the 3D scene. The real-time video data is used as texture information and attached to the video geometry to match and fuse with the 3D model. In the final video data picture of the present invention, if a segmented target (such as a person) is detected, the target will be highlighted to distinguish the foreground target from the background; if no segmented target is detected in the picture, the video will be the same as the original picture. Since the edge computing device first performs frame extraction on the video, and then combines the processed images into a video and outputs the video data by frame combination, the consistency in frame difference, image segmentation processing time, and video overlay display time is ensured. The constraint of time consistency helps to improve the stability and continuity of the output video and enhance the quality performance of the output video.
[0045] This technical solution realizes the local processing and 3D visualization of monitoring data through an edge computing architecture and a lightweight neural network model, solves problems such as high latency, high bandwidth occupancy, and limited computing resources in traditional monitoring systems, improves the real-time performance, reliability, and adaptability of the system, and provides an efficient and feasible technical path for fields such as intelligent monitoring, industrial inspection, and smart cities. The present invention utilizes the high computing power of edge computing devices to perform real-time target segmentation on the monitoring video stream locally, transferring the computing load from the user terminal or cloud server to the edge side. This optimization strategy significantly reduces the bandwidth requirements for data backhaul and reduces the cloud computing cost, enabling the monitoring system to operate efficiently under limited network resources, especially suitable for scenarios with limited network bandwidth or high real-time requirements. Secondly, the lightweight semantic segmentation model of the present invention can achieve high-precision target segmentation on edge computing devices while maintaining the integrity of edge details. Compared with traditional cloud computing that relies on high-performance servers, this solution enables low-power edge computing devices to also undertake high-precision intelligent analysis tasks, improves the independent processing ability of the devices, and ensures that the monitoring data received by user terminal devices is more accurate and efficient. After the target segmentation is completed, the present invention supports the user terminal to receive the processed video data and complete the fusion of the 3D video and the monitoring video on the user terminal device, such as the user's computer. The user terminal device generally has a graphics card and has the ability to fuse 3D videos, while the edge computing device (OrangePi AIpro) only processes the video data. Through precise perspective transformation calculation and region adaptive fusion strategy, the segmented monitoring video data can be accurately matched to the corresponding positions in the 3D model, enabling users to intuitively view the real-time status of the monitoring area in the 3D environment. This fusion method improves the visualization ability of the monitoring system, making applications such as security monitoring, industrial inspection, and intelligent operation and maintenance more efficient and intuitive.
[0046] Among them, as a preferred method, the encoder includes the following five steps:
[0047] Step 1, the input image passes through a 3×3 convolutional layer with the number of channels set to 32, equipped with batch normalization (Batch Normalization) and ReLU activation function, and then passes through a max pooling layer (MaxPooling) with a stride of 2, reducing the feature map size to H / 2 × W / 2;
[0048] Step 2: The first residual block group of ResNet-34 (a total of 3 residual blocks) is used for preliminary feature extraction. The number of channels is expanded to 64, and downsampling is performed through a max-pooling layer with a stride of 2, further reducing the feature map size to H / 4 × W / 4. The preliminary extraction and refinement of the image features after the initial convolution and pooling are carried out, converting the image from a relatively simple feature space to a feature space with certain semantic information, preparing for subsequent more complex feature extraction. Moreover, through convolution and pooling operations, without losing key information, the spatial resolution of the feature map is reduced, the subsequent computational amount is decreased, and at the same time, the abstraction degree of the features is increased.
[0049] Step 3: The second residual block group (a total of 4 residual blocks) is used for further feature extraction. The number of channels is expanded to 128, and it continues to pass through a max-pooling layer with a stride of 2, reducing the feature map size to H / 8 × W / 8. Based on the features extracted by the first residual block group, deeper features in the image are further mined, capturing more complex semantic information and patterns, enabling the model to learn more representative feature representations. Through the stacking of multiple residual blocks, more non-linear transformations are introduced, enriching the diversity of features and improving the model's ability to express different image features, which helps to better identify and classify different image contents.
[0050] Step 4: The third residual block group (a total of 6 residual blocks) is used for in-depth feature mining. The number of channels is further increased to 256, and downsampling is performed through a max-pooling layer, reducing the feature map size to H / 16 × W / 16. This step continues to deeply mine the features, extracting more abstract and high-level features in the image, which can better distinguish different categories of images and play a key role in the image classification task. As the network depth increases, the feature maps of different layers have different receptive fields, capable of capturing image information at different scales. The third residual block group effectively fuses multi-scale information through residual connections and convolution operations, enabling the model to have better adaptability to image scale changes.
[0051] Step 5: The fourth residual block group (a total of 3 residual blocks) is used for the final feature representation. The number of channels is expanded to 512, and downsampling is performed through a max-pooling layer, reducing the feature map size to H / 32 × W / 32. The final feature representation of the image is generated, and these features contain the most discriminative information in the image, capable of providing accurate feature bases for subsequent classification or other tasks. Through fine-tuning the features by the last few residual blocks, the generalization ability of the model is further enhanced, enabling the model to output results more stably and accurately when facing different image data and reducing the occurrence of overfitting phenomena.
[0052] These 4 residual block groups cooperate with each other. Starting from the initial image feature extraction, they gradually dig deeper and abstract features, and finally provide high-quality feature representations for the classification or other tasks of the model, enabling ResNet-34 to achieve excellent performance in various image tasks.
[0053] During the feature extraction process, the outputs of each step are passed to the decoder through skip connections and fused with the output of the previous step of the decoder as an input to the multi-scale enhanced feature fusion (MEFF) module to achieve multi-scale feature fusion.
[0054] The decoder uses the step-by-step up-sampling method to restore the spatial information, and combines the multi-scale feature fusion module (MEFF, Multi-scale Enhanced Feature Fusion) to make full use of the features at different stages of the encoder. The specific process is as follows.
[0055] Step 6: Use the MEFF module to process the high-level features output by the encoder in step 5 to restore the size to H / 16×W / 16.
[0056] Step 7: Fuse the output decoded in step 6 with the features output by the encoder in step 4, and enhance the multi-scale feature expression through the MEFF module. The fused feature map is then up-sampled by a factor of 2 to restore the size to H / 8×W / 8.
[0057] Step 8: Fuse the output decoded in step 7 with the features output by the encoder in step 3, and process it through the MEFF, and then up-sample it by a factor of 2 to restore the feature map to H / 4×W / 4.
[0058] Step 9: Fuse the output decoded in step 8 with the features output by the encoder in step 2, and process it through the MEFF to enhance the feature expression ability, and up-sample it by a factor of 2 to restore the feature map to H / 2×W / 2.
[0059] Step 10: Fuse the output decoded in step 9 with the features output by the encoder in step 1, and process it through the MEFF module to ensure that the fused features have both global and local information, and up-sample it by a factor of 2 to restore the feature map to (H×W).
[0060] Finally, to generate the segmentation result, the decoder output is adjusted in channels through a 1×1 convolution and the Sigmoid activation function is used to meet the requirements of different segmentation tasks.
[0061] Specifically, the basic residual module consists of two consecutive 3×3 convolutional layers, and ReLU is used as the activation function after each convolutional layer to increase the non-linear expression ability of the model. The input features are used to extract local features through the first 3×3 convolutional layer, and after ReLU activation, they are further passed to the second 3×3 convolutional layer for deeper feature mapping learning. At the same time, the initial input features directly jump to the final feature fusion stage through the residual connection, and after the addition operation, they are processed by the ReLU activation function to enhance the information flow and alleviate the gradient vanishing problem. This module can effectively retain the original information and achieve better feature expression ability during the deep feature learning process, which has a positive effect on improving the convergence speed and optimization stability of the deep neural network.
[0062] Furthermore, in image segmentation tasks, the objects of interest usually have large-scale variations and irregular shapes. Especially under adverse weather conditions (such as rainy days), the object boundaries are vulnerable to noise interference, making accurate segmentation more difficult. Therefore, to be sufficiently robust to effectively analyze objects of different scales and adapt to the image features under complex environments. Although an intuitive way to improve network performance is to increase the model size, this strategy often has significant drawbacks. A larger network size usually means more parameters and computational complexity, which not only increases the risk of overfitting but also significantly increases the consumption of computing resources, making it difficult to be deployed on edge computing devices with limited computing power.
[0063] The multi-scale enhanced feature fusion module (MEFF) of the present invention can improve the network's perception ability of multi-scale targets while being lightweight. This module adopts a multi-branch structure combined with depthwise separable convolutions, which can efficiently extract context information of different scales while reducing the computational complexity, making it particularly suitable for edge computing scenarios. In addition, the multi-branch structure can capture richer local and global information under adverse weather conditions, thereby enhancing the network's segmentation performance for targets under complex environments such as rainy days. The following part will detail the specific implementation of the MEFF module. In view of the limited computing resources of adverse weather and edge computing devices, an enhanced feature fusion module with a multi-branch structure is designed in the network model to improve the feature extraction ability. Through this structure, the model can more effectively extract key features under different weather conditions, thereby improving the overall recognition performance.
[0064] Specifically, given the input feature map , represents the input feature map, represents the set of real numbers, , Batch size (the size of a batch), that is, the number of samples included in one training or calculation. , Channels (the number of channels); , Height, refers to the size of the feature map in the vertical direction; , Width, refers to the size of the feature map in the horizontal direction. First, pointwise convolution (with a kernel size of 1×1) is used to compress its channels to reduce the number of parameters and computational overhead, resulting in a compressed feature map . Subsequently, the step constructs four independent branches to process the compressed feature map in different ways to enhance the feature extraction ability. Among them, the first branch uses a depthwise separable convolution of 1×3 to extract features in the horizontal direction, resulting in a first feature map ; the second branch uses a depthwise separable convolution of 3×1, mainly for extracting features in the vertical direction, resulting in a second feature map . Compared with the traditional 3×3 depthwise separable convolution, this strip convolution can expand the receptive field while reducing the number of parameters, and is especially suitable for edge computing devices with limited computing resources. The third branch first applies a standard convolution of 3×3 to to enhance the local feature learning ability, followed by normalization and ReLU activation, and further adjusts the feature distribution through a 1×1 convolution to obtain a third feature map . The fourth branch directly retains the compressed feature map as the fourth feature map to preserve the original information. Finally, the step concatenates the feature maps obtained from these four branches in the channel dimension to fuse multi-scale information and enhance the feature expression ability of the network. The specific series of operations are as follows,
[0065]
[0066]
[0067]
[0068]
[0069]
[0070] Furthermore, the real-time fusion of video data and the 3D model by the user terminal device means accurately mapping the foreground video data after target segmentation to the corresponding positions of the 3D model to achieve the fusion display of the video and the 3D scene. The real-time video data is used as texture information and attached to the video geometry to match and fuse with the 3D model.
[0071] Among them, to ensure that the video geometry can accurately correspond to the 3D model, feature point extraction and matching are required. The SIFT (Scale-Invariant Feature Transform) algorithm is used to extract the image features of the 3D model and the video geometry, and redundant points are removed through manual correction to ensure the matching accuracy of the feature points. Due to the dynamic change characteristics of video data, it may be affected by lighting changes and environmental factors in different time frames. Therefore, during the feature point extraction process, it is necessary to optimize the screening criteria for feature points to improve the stability of matching.
[0072] Based on the feature point matching, it is necessary to perform mesh division on the video geometry to optimize the texture fitting effect. The Delaunay triangulation method is used to perform triangular mesh division on the video geometry to form a reasonable mesh structure. For geometries with irregular contours, the constrained Delaunay triangulation method is used to maintain the overall contour of the geometry, avoid deformation distortion in the boundary region, and provide topological support for subsequent texture optimization.
[0073] Due to factors such as the camera shooting angle, perspective effect, and 3D model accuracy, video texture may be distorted when mapped to the 3D model. Therefore, the Laplacian mesh deformation algorithm is used to adjust the vertex coordinates of the video geometry to correct texture deformation. Under the constraint of this algorithm, the local mesh structure of the video geometry can be optimized while maintaining the original feature relationship, making the texture mapping more conform to the geometric features of the 3D model, thereby improving the accuracy of video fusion.
[0074] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of this invention.
Claims
1. A real-time processing method for real-scene three-dimensional Internet of Things perception data based on edge computing, characterized in that, It includes the following steps: 1) Create a 3D model and load it on the user's terminal device; 2) The edge computing device receives the monitoring video data, decodes it into a processable frame sequence using the hardware acceleration module, and then performs image preprocessing operations; 3) The edge computing device uses a lightweight semantic segmentation model based on the encoder-decoder architecture to process the preprocessed image for foreground object and background segmentation. Among them, the encoder of the lightweight semantic segmentation model uses ResNet-34 as the backbone network, the decoder part adopts a progressive upsampling strategy, and combines a multi-scale enhanced feature fusion module to generate a high-resolution segmentation result. The multi-scale enhanced feature fusion module adopts a multi-branch structure combined with depthwise separable convolutions to extract context information of different scales; 4) The edge computing device synthesizes the segmentation processing results of each frame image into video data in real time and sends it to the user's terminal device; 5) The user's terminal device maps the video data to the corresponding position of the 3D model to achieve the matching and fusion of the video data and the 3D scene; Among them, the multi-scale enhanced feature fusion module constructs four independent branches to perform different processing methods. Among them, The first branch uses a 1×3 depthwise separable convolution on the input feature map to extract horizontal features and obtain the first feature map; The second branch uses a 3×1 depthwise separable convolution on the input feature map to extract vertical features and obtain the second feature map. The third branch first applies a 3×3 standard convolution to the input feature map, then performs normalization and ReLU activation, and further adjusts the feature distribution through a 1×1 convolution to obtain the third feature map. The fourth branch directly retains the input feature map as the fourth feature map; Finally, the feature maps obtained from these four branches are concatenated in the channel dimension.
2. The real-time processing method for panoramic 3D IoT perception data based on edge computing according to claim 1, wherein, The encoder includes the following five steps: Step 1, the input image passes through a 3×3 convolutional layer, and then batch normalization and the ReLU activation function are used; Step 2, the first residual block group of ResNet-34 performs preliminary feature extraction; Step 3, the second residual block group further performs feature extraction; Step 4, the third residual block group performs deep feature mining; Step 5, the fourth residual block group performs the final feature representation.
3. The real-time processing method for panoramic 3D IoT perception data based on edge computing according to claim 2, wherein The output of each step of the encoder is transmitted to the decoder through skip connections and fused with the output of the previous step of the decoder as the input of the multi-scale enhanced feature fusion module to achieve multi-scale feature fusion.
4. The real-time processing method for 3D physical Internet of Things perception data based on edge computing according to claim 3, wherein, The specific process of the decoder is as follows: Step 6, use the multi-scale enhanced feature fusion module to process the high-level features output by step 5 of the encoder; Step 7, fuse the output of decoder step 6 with the features output by encoder step 4, and then use the multi-scale enhanced feature fusion module for fusion processing; Step 8, fuse the output of decoder step 7 with the features output by encoder step 3, and then use the multi-scale enhanced feature fusion module for fusion processing; Step 9: Fuse the output of the decoder in Step 8 with the features output by the encoder in Step 2, and then perform fusion processing using a multi-scale enhanced feature fusion module; Step 10: Fuse the output of the decoder in Step 9 with the features output by the encoder in Step 1, and then perform fusion processing using a multi-scale enhanced feature fusion module.
5. The real-time processing method of the three-dimensional physical Internet of Things perception data based on edge computing according to claim 4, characterized in that The multi-scale enhanced feature fusion module further includes a step of compressing its channels through pointwise convolution to obtain the input feature map.
6. The real-time processing method for 3D physical IoT perception data based on edge computing according to claim 1, characterized in that It also includes feature point extraction and matching, which includes using the SIFT algorithm to extract the image features of the 3D model and the video geometry, then dividing the video geometry into grids, using the Delaunay triangulation method for triangular grid division for regular contour video geometries, and using the constrained Delaunay triangulation method for geometries with irregular contours.
7. The real-time processing method of the real-scene three-dimensional IoT perception data based on edge computing according to claim 6, characterized in that, It also includes using the Laplacian mesh deformation algorithm to adjust the vertex coordinates of the video geometry to correct texture deformation.
Citation Information
Patent Citations
Stereo vision tracking method and stereo vision tracking system based on 3D Delaunay triangulation
CN105513094A
Multi-modal 3D target detection method and device based on cloud edge collaboration
CN117496322A
Three-dimensional reconstruction method based on unmanned aerial vehicle and edge calculation
CN119478213A
Pose estimation apparatus and method for robotic arm to grasp target based on monocular infrared thermal imaging vision
US20240242377A1