Construction site monitoring system based on remote sensing image technology and equipment thereof
Through the combination of drones, low-orbit satellites and ground mobile platforms combined with multimodal remote sensing imaging technology, the accuracy of remote sensing image target recognition in complex environments is solved, and high-precision target recognition and management under different lighting, weather conditions and urban construction backgrounds are achieved, and construction safety and management efficiency are improved.
Patent Information
- Application Number
- CN202510398032.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, in complex environments, such as different lighting, weather conditions and urban building complex backgrounds, the accuracy of remote sensing image target recognition is affected, resulting in incorrect target classification, which in turn leads to incorrect interpretation of engineering monitoring information.
Remote sensing images are obtained by using drone platforms, low-orbit satellite platforms and ground mobile platforms, combined with optical + infrared images, synthetic aperture radars and lidars, real-time transmission through 5G edge computing, multimodal data fusion is used to combine multimodal data fusion, combined with self-supervised learning and transfer learning technology, an image quality standard library is established, a graph convolutional network is used to analyze spatial relationships, and an uncertainty quantization mechanism is introduced for target identification and evaluation.
It improves the accuracy of target recognition in complex environments, reduces dependence on labeled data, adapts to new task requirements, improves construction safety and management efficiency, reduces misidentification and misjudgment, and is suitable for dynamic construction site environments.
Smart Images

Figure CN120339766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of construction site monitoring, and particularly relates to a construction site monitoring system and its equipment based on remote sensing image technology. Background Art
[0002] Remote sensing image technology refers to the technology of obtaining information on the Earth's surface from a distance using sensors installed on satellites, aircraft, or other flying vehicles. These sensors can capture electromagnetic wave radiation in different bands, including visible light, infrared, ultraviolet, and microwave, etc., thus generating various types of remote sensing images. Remote sensing image technology is widely used in multiple fields such as construction site monitoring, environmental monitoring, and disaster assessment.
[0003] The patent with the publication number CN116863344A records in its specification that "the present invention discloses an engineering monitoring system based on remote sensing images, belonging to the technical field of construction monitoring, including a remote sensing satellite module and a monitoring display module; the remote sensing satellite module is used to intelligently recommend corresponding remote sensing satellite images, and the monitoring display module is used to perform engineering monitoring display, obtain remote sensing images collected by the remote sensing satellite, intercept the remote sensing images according to the target monitoring area to obtain target images, perform target classification and identification on the target images to obtain monitoring images, perform corresponding construction asset evaluation according to the obtained monitoring images to obtain corresponding asset safety values, and integrate the obtained fund safety values and monitoring images into display data for display; through the setting of the monitoring display module, combined with the current remote sensing satellite, dynamic monitoring of overseas construction projects is realized, and key monitoring and evaluation are carried out according to the points concerned by the enterprise, so as to realize the effective risk control of overseas infrastructure projects by the enterprise and ensure the safety of funds". Although the above technology uses a target recognition model to classify and identify images, although the target recognition model adopts advanced technologies such as deep learning, in complex environments such as images under different lighting and weather conditions, and images with complex urban building backgrounds, the accuracy of target recognition may be affected, resulting in incorrect classification and identification, and incorrect target classification may lead to misinterpretation of engineering monitoring information, resulting in confusion between buildings and vegetation, or misclassification of construction equipment.
[0004] In summary, developing a construction site monitoring system and its equipment based on remote sensing image technology is still a key problem urgently to be solved in the technical field of construction site monitoring. Summary of the Invention
[0005] The object of the present invention is to solve the problem that although the above-mentioned technology uses a target recognition model to classify and identify images, although the target recognition model adopts advanced technologies such as deep learning, in complex environments such as images under different lighting and weather conditions, and images in the complex background of urban buildings, the accuracy of target recognition may be affected, resulting in incorrect classification and identification, and incorrect target classification may lead to incorrect interpretation of engineering monitoring information, resulting in confusion between buildings and vegetation, and misclassification of construction equipment.
[0006] To achieve the above object, the present invention provides a construction site monitoring system based on remote sensing image technology, including: a remote sensing image acquisition module, where the drone platform, the low-earth orbit satellite platform, and the ground mobile platform all obtain remote sensing images of the construction site through sensors; A multi-modal fusion module that performs data fusion processing on the remote sensing images to obtain multi-modal data; A target recognition module that automatically recognizes construction site targets according to the multi-modal data by establishing an image quality standard library; A spatio-temporal constraint module that analyzes the spatial relationship using a graph convolutional network (GCN) according to the construction site targets and classifies the targets using temporal and spatial information; A doubt evaluation module that introduces an uncertainty quantification mechanism to evaluate the confidence of the target classification.
[0007] Further, the operation process of the remote sensing image acquisition module includes: The sensors include optical + infrared images, synthetic aperture radar, and lidar, and the remote sensing images are transmitted in real time using 5G edge computing. At time , the remote sensing images collected by each platform are expressed as: , where is the remote sensing image collected by platform through the sensor at position and time , represents the relationship between the trajectory of platform and the sampling point, is the modulation method of different sensors on the remote sensing image determined by the sensor model, represents the noise at time and spatial position . Since the amount of data of the remote sensing image is large, to achieve real-time transmission, let the size of the remote sensing image data be , the transmission bandwidth be , and the 5G transmission delay be . It is expressed as: , where is the signal propagation delay, is the distance from the construction site to the edge server, is the speed of light, The preprocessing delay of edge computing depends on the amount of computing tasks and computing resources , and the expression formula is: , where is the computational complexity of the task, is the computing resource, represents the total number of computing tasks that the system needs to process. The remote sensing image is compressed at the edge computing end, and the compression ratio affects the transmission delay: , where is the amount of data after compression, refers to the size of the remote sensing image data before compression, The purpose is to minimize the 5G transmission delay , is the constraint that the amount of data after compression cannot be lower than the minimum available data volume , and the edge computing server performs denoising, feature extraction, and target detection on the remote sensing image. The expression formula is: , where is the denoised and optimized remote sensing image, refers to the remote sensing image data collected by the sensor, refers to the noise-free remote sensing image under ideal conditions, is used to measure the credibility of different pixel points.
[0008] Furthermore, the operation process of the multimodal fusion module includes: The multimodal data is obtained by fusing the remote sensing image in the way of a deep learning algorithm with a Transformer+CNN hybrid architecture, combined with data augmentation technology, multi-layer attention mechanism, self-supervised learning technology, and transfer learning technology. The data augmentation technology is used to perform data augmentation processing on the denoised and optimized remote sensing image , and the expression formula is: , where is the transformation function, is the specific parameter that includes transformations such as rotation, scaling, and flipping, is the augmented dataset. Use a multi-layer convolutional neural network to extract features, and set the convolutional network , and the expression formula is: , where includes the weights, bias terms, and activation function parameters of the convolutional kernel, obtains local features through multi-layer convolution and pooling operations, and combines the local features extracted by convolution Input into the Transformer architecture for global context modeling, and set the Transformer model , the expression formula: , where are the parameters in the Transformer, including the weight matrix and position encoding in the attention mechanism. The multi-layer attention mechanism weights and fuses the features of multiple modalities The weighted fusion operation: , the expression formula of the multi-layer attention mechanism: , where is the weight matrix, are the parameters of the attention mechanism.
[0009] Furthermore, the operation process of the multi-modal fusion module includes: In the self-supervised learning technology stage, the structural features of unlabeled data are learned through the reconstruction loss, and the self-supervised loss is set as: , where is the loss function of self-supervised learning, is the weight hyperparameter that controls the influence degree of the feature contrast loss term, is to sum over all feature dimensions , is the -dimensional predicted feature, is the -dimensional target feature, is to calculate the square of the Euclidean distance between the predicted feature and the true feature, is the weight hyperparameter that controls the influence degree of the information entropy loss, is used to measure the probability distribution (true distribution) and (predicted distribution) between the differences. The transfer learning technology uses a pre-trained model to perform transfer learning on the task, and the transfer loss is set as: , where is the transfer learning loss function, represents the sum over all feature dimensions from 1 to , are the features of the target task and the source task respectively, is the hyperparameter of transfer learning. The multi-modal data output after training and optimization, the expression formula: , where is the finally obtained multi-modal data set, is the -th fused feature data, is the feature extracted by fusing CNN and Transformer through the attention mechanism, is the image spatial feature extracted by CNN, is the global feature extracted by Transformer, denotes is an element in the feature set after fusion through the attention mechanism.
[0010] Furthermore, the operation process of the target recognition module includes: Set the image quality standard library as , extract the image features through the deep learning network, and the expression formula: , where is the deep learning feature extraction network based on the Transformer+CNN hybrid architecture, is the extracted high-dimensional feature vector, is the finally obtained multi-modal data set, and the extracted features are matched with the target features in the standard library to calculate the feature similarity, and the expression formula: , where represents the cosine similarity.
[0011] Furthermore, the operation process of the target recognition module includes: Adopt the Transformer encoder to perform multi-scale feature learning, and the expression formula: , where is the deep feature representation after being processed by the Transformer. Subsequently, use the Softmax classifier to calculate the target category probability, and the expression formula: , where are the trainable parameters of the classifier, Given the input multi-modal data set the probability distribution of the target category after is the bias term. Finally, select the category with the highest probability as the recognition result, and the expression formula: , where represents the finally recognized construction site target.
[0012] Furthermore, the operation process of the spatio-temporal constraint module includes: Aggregate the construction site targets obtained from the target recognition module , and the expression formula: , where represents the th target's feature vector, and establish the spatial relationship graph between targets, and the expression formula: , where is the construction site target set, and the edge set Represents the spatial relationship between targets, is to establish a graph structure between targets, define edge weights according to Euclidean distance and topological structure, and establish the spatial relationship between targets After the graph is established, spatial feature extraction is performed through a graph convolutional network (GCN). The expression formula is: , where is the feature representation of the layer of GCN initially (i.e., target features), is the adjacency matrix representing the connection relationship between targets, is the degree matrix , is the weight matrix of the layer of GCN, is the non-linear activation function. Finally, the spatial relationship features of the construction site targets are obtained , the expression formula is: , where is the spatial relationship feature of the construction site targets, is the adjacency matrix representing the connection relationship between targets, is the feature of the targets on the construction site.
[0013] Furthermore, the operation process of the spatio-temporal constraint module includes: In the time dimension, a time series feature extraction method based on the Transformer structure is adopted. The expression formula for the time series information of the construction site targets is: , where are the query, key, and value obtained through linear transformation, is the spatio-temporal feature. The SAtn calculation formula is expressed as: , where is the dimension of Key. Based on the spatio-temporal feature , target classification is performed using Sofx. The expression formula is: , where is the parameter of the classification layer, is the classification bias vector, is used for normalizing the output, represents the classification probability distribution of the construction site targets, is the construction site target.
[0014] Furthermore, the operation process of the question evaluation module includes: Obtain the classification probability distribution of the construction site targets from the spatio-temporal constraint module, and then introduce Bayesian inference for quantification. The expression formula is: , where are the weight and bias parameters of the neural network, represents the weight matrix for feature transformation in the neural network, represents the bias term of the neural network, is not a fixed parameter but follows a probability distribution, denotes follows a normal distribution with a mean of and a covariance matrix of The normal distribution, denotes follows a normal distribution with a mean of and a covariance matrix of The normal distribution, and then through Monte Carlo for rounds of sampling, the expression formula: where is the predicted probability calculated through the forward propagation of multiple random perturbations, is the weight of the round of sampling. Then, the information entropy is used to measure the uncertainty of the prediction, and the expression formula is: where is the uncertainty measure when predicting the target category, denotes that under the condition of input data the probability that the predicted target belongs to category is is the logarithm of the probability. At the same time, the prediction variance is calculated to define the final confidence, and the expression formula is: where is the confidence, is the uncertainty measure when predicting the target category, is the maximum prediction variance, represents the maximum possible information entropy, is the total number of classification categories, denotes the output fluctuation under different weight samplings. The final confidence score of the target category ranges from 0 to 1. A confidence close to 1 indicates that the classification result of the construction site target is reliable, and close to 0, the target needs to be re-evaluated.
[0015] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the described construction site monitoring system based on remote sensing image technology, and the memory is a non-transitory computer-readable storage medium.
[0016] Beneficial effects Adopting the technical solution provided by the present invention, compared with the known public technology, it has the following beneficial effects: When in use, the present invention uses a drone platform, a low-earth orbit satellite platform, and a ground mobile platform to collect remote sensing image data, which is convenient for accurately identifying various targets at the construction site, reducing misidentifications, and is beneficial for identifying and positioning construction site targets in complex environments such as images under different lighting and weather conditions and images with complex backgrounds of urban buildings. In data-scarce and new construction site environments, self-supervised learning and transfer learning technologies can effectively improve the learning ability of the model, reduce the dependence on labeled data, and quickly adapt to new task requirements. When in use, the present invention combines a Transformer + CNN hybrid architecture, which is convenient for extracting local and global features of images, improving the accuracy of recognition, and is applicable to complex construction site environments such as lighting changes, angle changes, occlusion, etc. Through multi-scale feature learning, regardless of how the size and angle of the target change, and regardless of whether the camera shoots from the ground or the drone takes an aerial photo, the same target can be recognized. For unregistered and abnormal devices, it improves construction safety and management efficiency. At the same time, when encountering new targets or failed recognition cases, the recognition ability of the system can be gradually improved by updating the image standard library, which is applicable to dynamic construction site environments, ensuring the long-term adaptability of the system, and is beneficial for providing an automatic adjustment mechanism when the classification is uncertain, avoiding misjudgment, reducing the misjudgment of construction site management caused by misclassification, improving construction safety, and enhancing the adaptability to complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a system diagram of a construction site monitoring system and its equipment based on remote sensing image technology according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] It should be noted that the terms "first", "second", etc. in the description, claims and drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0020] The present invention will be further described in detail below with reference to the accompanying drawings: Embodiment: As Figure 1 shown, the present invention provides a construction site monitoring system based on remote sensing image technology, including: a remote sensing image acquisition module, and the unmanned aerial vehicle platform, the low-earth orbit satellite platform and the ground mobile platform all obtain remote sensing images of the construction site through sensors; The operation process of the remote sensing image acquisition module includes: The sensors include optical + infrared images, synthetic aperture radar and lidar, and the remote sensing images are transmitted in real time using 5G edge computing. At time , the remote sensing images collected by each platform are expressed as: , where is the remote sensing image collected by platform through the sensor at position and time , represents the relationship between the trajectory of platform and the sampling points, is the modulation method of different sensors for remote sensing images determined by the sensor model, represents the noise at time and spatial position . Since the amount of data of the remote sensing images is large, in order to achieve real-time transmission, let the size of the remote sensing image data be , the transmission bandwidth be , and the 5G transmission delay be , which is expressed as: , where is the signal propagation delay, is the distance from the construction site to the edge server, is the speed of light, is the preprocessing delay of edge computing, which depends on the amount of computing tasks and computing resources , and the expression formula: , where is the computational complexity of the task, is the computing resource, represents the total number of computing tasks that the system needs to process. The remote sensing image is compressed at the edge computing end, and the compression ratio affects the transmission delay: , where is the amount of data after compression, refers to the size of the remote sensing image data before compression, The purpose is to minimize the 5G transmission delay , is the constraint that the amount of data after compression cannot be lower than the minimum available data volume , and the edge computing server performs denoising, feature extraction, and target detection on the remote sensing image. The expression formula: , where is the denoised and optimized remote sensing image, refers to the remote sensing image data collected by the sensor, refers to the noise-free remote sensing image under ideal conditions, is used to measure the credibility of different pixel points; Specifically, in the monitoring application of construction sites, unmanned aerial vehicle platforms, low-earth orbit satellite platforms, and ground mobile platforms are used to collect remote sensing image data. Optical + infrared sensors and lidar are used to obtain remote sensing images with different resolutions. The low-earth orbit satellite provides macroscopic regional data to supplement the blank of unmanned aerial vehicle data and obtain information on large-scale construction sites. The ground mobile platform is in close contact with the construction site targets. All remote sensing image data is transmitted to the edge server through the 5G edge computing network. The image data is compressed during the transmission process to reduce the transmission delay, which is beneficial to improving the response speed of construction site monitoring, helping to grasp the construction progress in real time, facilitating the accurate identification of various targets at the construction site, and reducing misidentification.
[0021] The multi-modal fusion module performs data fusion processing according to the remote sensing image to obtain multi-modal data; The operation process of the multi-modal fusion module includes: The multi-modal data is obtained by fusing the remote sensing image through a deep learning algorithm with a Transformer + CNN hybrid architecture, combined with data augmentation techniques, multi-layer attention mechanisms, self-supervised learning techniques, and transfer learning techniques. The data augmentation technique is used to perform data augmentation processing on the denoised and optimized remote sensing image The expression formula: , where is the transformation function, is the specific parameter that includes transformations such as rotation, scaling, and flipping, It is an enhanced dataset. A multi-layer convolutional neural network is used to extract features, and the convolutional network is set . The expression formula is: , where contains the weights of the convolutional kernel, the bias term, and the parameters of the activation function. Local features are obtained through multi-layer convolution and pooling operations. The local features extracted by convolution are input into the Transformer architecture for global context modeling. The Transformer model is set . The expression formula is: , where are the parameters in the Transformer, including the weight matrix and position encoding in the attention mechanism. The multi-layer attention mechanism weights and fuses the features of multiple modalities . The weighted fusion operation is: . The expression formula of the multi-layer attention mechanism is: , where is the weight matrix, and are the parameters of the attention mechanism. The operation process of the multi-modal fusion module includes: In the self-supervised learning technology stage, the structural features of unlabeled data are learned through the reconstruction loss. The self-supervised loss is set as: , where is the loss function of self-supervised learning, is the weight hyperparameter that controls the influence degree of the feature contrast loss term, is to sum over all feature dimensions , is the -dimensional predicted feature, is the -dimensional target feature, is to calculate the square of the Euclidean distance between the predicted feature and the true feature, is the weight hyperparameter that controls the influence degree of the information entropy loss, is used to measure the probability distribution (true distribution) and (predicted distribution) difference. The transfer learning technology uses a pre-trained model to perform transfer learning on the task. The transfer loss is set as: , where is the transfer learning loss function, means to sum over all feature dimensions from 1 to , are the features of the target task and the source task respectively. is a hyperparameter for transfer learning, and the multi-modal data output after training and optimization, with the expression formula: , where is the finally obtained multi-modal data set, is the th fused feature data, is the feature extracted by fusing CNN and Transformer through the attention mechanism, is the image spatial feature extracted by CNN, is the global feature extracted by Transformer, represents is an element in the feature set after being fused through the attention mechanism; Specifically, through multi-modal data fusion, this system can integrate the advantages of optical, infrared, and radar images, overcome the limitations of a single modality, and is beneficial to providing more accurate and comprehensive information about construction site targets. Combining the feature extraction capabilities of CNN and Transformer, as well as the weighted fusion of multi-layer attention mechanisms, can effectively improve the accuracy of target recognition, and is beneficial to identifying and locating construction site targets in complex environments such as images under different lighting and weather conditions, and images with complex backgrounds of urban buildings. In data-scarce and new construction site environments, self-supervised learning and transfer learning techniques can effectively improve the learning ability of the model, reduce the dependence on labeled data, and quickly adapt to new task requirements.
[0022] The target recognition module automatically recognizes construction site targets based on the multi-modal data by establishing an image quality standard library; The operation process of the target recognition module includes: Set the image quality standard library as , extract image features through a deep learning network, with the expression formula: , where is a deep learning feature extraction network based on a Transformer+CNN hybrid architecture, is the extracted high-dimensional feature vector, is the finally obtained multi-modal data set, and the extracted feature is matched with the target feature in the standard library, and the feature similarity is calculated, with the expression formula: , where represents the cosine similarity; The operation process of the target recognition module includes: Adopt a Transformer encoder for multi-scale feature learning, with the expression formula: , where It is the deep feature representation after being processed by the Transformer. Subsequently, a Softmax classifier is used to calculate the target category probabilities, and the expression formula is: , where are the trainable parameters of the classifier, Given the input multi-modal data set the probability distribution of the target category , is the bias term. Finally, the category with the highest probability is selected as the recognition result, and the expression formula is: , where represents the finally recognized construction site target; Specifically, combined with the Transformer+CNN hybrid architecture, it is convenient to extract local and global features of the image, improve the accuracy of recognition, and is applicable to complex construction site environments, such as lighting changes, angle changes, occlusion, etc. For multi-scale feature learning, no matter how the target size and angle change, whether the camera shoots from the ground or the drone takes an aerial photo, the same target can be recognized. For unregistered devices and abnormal devices, it can improve construction safety and management efficiency. At the same time, in the case of encountering new targets or recognition failures, the recognition ability of this system can be gradually improved by updating the image standard library, which is applicable to dynamic construction site environments and ensures the long-term adaptability of the system.
[0023] The spatio-temporal constraint module, according to the construction site target, uses a graph convolutional network (GCN) to analyze the spatial relationship and uses temporal and spatial information for target classification; The operation process of the spatio-temporal constraint module includes: Collect the construction site targets obtained from the target recognition module , and the expression formula is: , where represents the th feature vector of the target, establish a spatial relationship graph between targets, and the expression formula is: , where is the set of construction site targets, and the edge set represents the spatial relationship between targets, is to establish a graph structure between targets, define the edge weights according to the Euclidean distance and topological structure, and after establishing the spatial relationship graph between targets, perform spatial feature extraction through a graph convolutional network (GCN), and the expression formula is: , where is the initial feature representation of the th layer of GCN (i.e., the target feature), is the adjacency matrix representing the connection relationship between targets, is the degree matrix , is the weight matrix of the -th layer GCN, is a non-linear activation function. Finally, the spatial relationship features of the construction site targets are obtained , and the expression formula is: , where is the spatial relationship features of the construction site targets, is the adjacency matrix representing the connection relationship between targets, is the feature of the targets on the construction site; The operation process of the spatio-temporal constraint module includes: In the time dimension, a time series feature extraction method based on the Transformer structure is adopted, and the expression formula for the time series information of the construction site targets is: , where are the query, key, and value obtained through linear transformation, is the spatio-temporal feature, and the expression formula for the SAtn calculation formula is: , where is the dimension of the Key. Based on the spatio-temporal feature , target classification is performed using Sofx, and the expression formula is: , where are the classification layer parameters, is the classification bias vector, is used for normalization output, represents the classification probability distribution of the construction site targets, is the construction site target;
[0024] The doubt evaluation module introduces an uncertainty quantification mechanism to evaluate the confidence of the target classification; The operation process of the doubt evaluation module includes: Obtain the classification probability distribution of the construction site targets from the spatio-temporal constraint module, and then introduce Bayesian inference for quantification. The expression formula is: , where are the weight and bias parameters of the neural network, Represents the weight matrix for feature transformation in the neural network, Represents the bias term of the neural network, which is not a fixed parameter but follows a probability distribution, denotes follows a normal distribution with a mean of and a covariance matrix of , denotes follows a normal distribution with a mean of and a covariance matrix of , and then performs rounds of sampling through Monte Carlo. The expression formula is: where is the predicted probability calculated through the forward propagation of multiple random perturbations, is the weight for the round of sampling. Then, the uncertainty of the prediction is measured using information entropy. The expression formula is: where is the uncertainty measure when predicting the target category, denotes the probability that the predicted target belongs to category under the condition of input data , is the logarithm of the probability. At the same time, the prediction variance is calculated to define the final confidence. The expression formula is: where is the confidence, is the uncertainty measure when predicting the target category, is the maximum prediction variance, represents the maximum possible information entropy, is the total number of classification categories, denotes the output fluctuation under different weight samplings. The final confidence score for the target category ranges from 0 to 1. A confidence close to 1 indicates that the classification result of the construction site target is reliable, and close to 0, the target needs to be re-evaluated; Specifically, during the classification process of construction site targets, the questioning and evaluation module is mainly used to evaluate the confidence of the classification results, improve the reliability of the prediction, introduce Bayesian inference and Monte Carlo sampling, which is beneficial to providing an automatic adjustment mechanism when the classification is uncertain, avoiding misjudgment, reducing the misjudgment of construction site management caused by misclassification, improving construction safety, and enhancing the adaptability to complex environments.
[0025] Furthermore, the present invention also provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The processor can call the computer program in the memory to execute the described construction site monitoring system based on remote sensing image technology.
[0026] In addition, when the computer program in the above-mentioned memory is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-transitory computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: disks, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical disks, etc., which can store program codes. The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, all the functions of the described construction site monitoring system based on remote sensing image technology of the above system are realized. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A construction site monitoring system based on remote sensing image technology, characterized in that, Including: A remote sensing image acquisition module, where the unmanned aerial vehicle platform, low-earth orbit satellite platform, and ground mobile platform all obtain remote sensing images of the construction site through sensors; A multimodal fusion module that performs data fusion processing on the remote sensing images to obtain multimodal data; A target recognition module that automatically recognizes construction site targets based on the multimodal data by establishing an image quality standard library; A spatio-temporal constraint module that analyzes spatial relationships using a graph convolutional network (GCN) based on the construction site targets and classifies the targets using temporal and spatial information; A doubt evaluation module that introduces an uncertainty quantification mechanism to evaluate the confidence level of the target classification.
2. The construction site monitoring system based on remote sensing image technology according to claim 1, characterized in that, The operation process of the remote sensing image acquisition module includes: The sensor includes an optical + infrared image, a synthetic aperture radar, and a lidar. The remote sensing image is transmitted in real time using 5G edge computing. At time , the remote sensing images collected by each platform are expressed as: , where is the platform through the sensor at the location and time collected remote sensing image, represents the relationship between the trajectory of the platform and the sampling point. is the sensor model that determines the modulation method of different sensors for the remote sensing image. represents the noise at time and spatial position . The amount of data in the remote sensing image is large. To achieve real-time transmission, let the size of the remote sensing image data be , the transmission bandwidth be , and the 5G transmission delay be , which is expressed as: , where is the signal propagation delay, is the distance from the construction site to the edge server, is the speed of light, is the preprocessing delay of edge computing, which depends on the amount of computing tasks and computing resources , and the expression formula is: , where is the computational complexity of the task, is the computing resource, represents the total number of computing tasks that the system needs to process. The remote sensing image is compressed at the edge computing end, and the compression ratio affects the transmission delay: , where is the amount of data after compression, refers to the size of the remote sensing image data before compression, The purpose is to minimize the 5G transmission delay , is the constraint that the amount of data after compression cannot be lower than the minimum available data volume . The edge computing server performs denoising, feature extraction, and target detection on the remote sensing image, and the expression formula is: , where is the denoised and optimized remote sensing image, refers to the remote sensing image data collected by the sensor, refers to the noise-free remote sensing image under ideal conditions, is used to measure the credibility of different pixel points.
3. The construction site monitoring system based on remote sensing image technology according to claim 2, characterized in that, The operation process of the multimodal fusion module includes: The multi-modal data is obtained by fusing the remote sensing images through a deep learning algorithm with a Transformer+CNN hybrid architecture, combined with data augmentation techniques, multi-layer attention mechanisms, self-supervised learning techniques, and transfer learning techniques. The data augmentation technique is used to perform data augmentation on the denoised and optimized remote sensing images for data augmentation processing, and the expression formula is: , where is the transformation function, is the specific parameter including rotation, scaling, and flipping transformations, is the augmented dataset. A multi-layer convolutional neural network is used to extract features, and the convolutional network is set, and the expression formula is: , where includes the weights of the convolutional kernels, bias terms, and activation function parameters, obtains local features through multi-layer convolution and pooling operations. The local features extracted by convolution are input into the Transformer architecture for global context modeling, and the Transformer model is set, and the expression formula is: , where are the parameters in the Transformer, including the weight matrix and position encoding in the attention mechanism. The multi-layer attention mechanism performs weighted fusion on the features of multiple modalities , and the weighted fusion operation is: , and the expression formula of the multi-layer attention mechanism is: , where is the weight matrix, are the parameters of the attention mechanism.
4. The construction site monitoring system based on remote sensing image technology according to claim 3, characterized in that, The operation process of the multimodal fusion module includes: In the self-supervised learning technology stage, the structural features of unlabeled data are learned through reconstruction loss, and the self-supervised loss is set as: , where is the loss function of self-supervised learning, is the weight hyperparameter that controls the influence degree of the feature contrast loss term, is to sum over all feature dimensions for summation, is the -th dimensional predicted feature, is the -th dimensional target feature, is to calculate the square of the Euclidean distance between the predicted feature and the true feature, is the weight hyperparameter that controls the influence degree of the information entropy loss, is used to measure the probability distribution (true distribution) and (predicted distribution) between the differences. The transfer learning technology uses a pre-trained model to perform transfer learning on the task, and the transfer loss is set as: , where is the transfer learning loss function, represents summing over all feature dimensions from 1 to for summation, are the features of the target task and the source task respectively, is the hyperparameter of transfer learning. The multi-modal data output after training and optimization is expressed by the formula: , where is the finally obtained multi-modal data set, is the -th fused feature data, is the feature extracted by fusing CNN and Transformer through the attention mechanism, is the image spatial feature extracted by CNN, is the global feature extracted by Transformer, represents is an element in the feature set fused through the attention mechanism.
5. The construction site monitoring system based on remote sensing image technology according to claim 4, characterized in that, The operation process of the target recognition module includes: Set the image quality standard library as , extract image features through a deep learning network, and the expression formula is: , where is a deep learning feature extraction network based on a Transformer+CNN hybrid architecture, is the extracted high-dimensional feature vector, is the finally obtained multi-modal data set, and the extracted features are matched with the target features in the standard library, and the feature similarity is calculated. The expression formula is: , where represents the cosine similarity.
6. The construction site monitoring system based on remote sensing image technology according to claim 5, characterized in that, The operation process of the target recognition module includes: Adopt a Transformer encoder Perform multi-scale feature learning. The expression formula is: , where is the deep feature representation after being processed by the Transformer. Subsequently, use a Softmax classifier to calculate the target class probability. The expression formula is: , where are the trainable parameters of the classifier, Given the input multi-modal data set the probability distribution of the target class , is the bias term. Finally, select the class with the highest probability as the recognition result. The expression formula is: , where represents the finally recognized construction site target.
7. The construction site monitoring system based on remote sensing image technology according to claim 6, characterized in that, The operation process of the spatio-temporal constraint module includes: The construction site targets obtained from the target recognition module are aggregated, and the expression formula is: , where represents the feature vector of the th target, and a spatial relationship graph between targets is established, and the expression formula is: , where is the set of construction site targets, and the edge set represents the spatial relationship between targets, is the graph structure established between targets. The edge weights are defined according to the Euclidean distance and topological structure. After establishing the spatial relationship graph between targets, spatial feature extraction is performed through a graph convolutional network (GCN), and the expression formula is: , where is the initial feature representation of the th layer of GCN (i.e., the target feature), is the adjacency matrix representing the connection relationship between targets, is the degree matrix , is the weight matrix of the th layer of GCN, is the non-linear activation function. Finally, the spatial relationship features of the construction site targets are obtained, and the expression formula is: , where is the spatial relationship feature of the construction site targets, is the adjacency matrix representing the connection relationship between targets, is the feature of the targets on the construction site.
8. A construction site monitoring system based on remote sensing image technology according to claim 7, characterized in that, The operation process of the spatio-temporal constraint module includes: In the time dimension, a time series feature extraction method based on the Transformer structure is adopted. The expression formula for the time series information of construction site targets is: , where are the query, key, and value obtained through linear transformation, are the spatio-temporal features. The expression formula for the SAtn calculation formula is: , where is the dimension of the Key. Based on the spatio-temporal features , target classification is performed using Sofx. The expression formula is: , where are the classification layer parameters, is the classification bias vector, is used for normalizing the output, represents the classification probability distribution of construction site targets, are the construction site targets.
9. The construction site monitoring system based on remote sensing image technology according to claim 8, characterized in that The operation process of the doubt evaluation module includes: Obtain the classification probability distribution of the construction site target from the spatio-temporal constraint module, and then introduce Bayesian inference for quantification. The expression formula is: , where are the weight and bias parameters of the neural network, represents the weight matrix for feature transformation in the neural network, represents the bias term of the neural network, is not a fixed parameter but follows a probability distribution, denotes follows a normal distribution with mean and covariance matrix , denotes follows a normal distribution with mean and covariance matrix . Then, perform rounds of sampling through Monte Carlo. The expression formula is: , where is the predicted probability calculated through the forward propagation of multiple random perturbations, is the weight of the th round of sampling. Then, use information entropy to measure the uncertainty of the prediction. The expression formula is: , where is the uncertainty measure when predicting the target category, denotes the probability that the predicted target belongs to category under the condition of input data , is the logarithm of the probability. At the same time, calculate the prediction variance and define the final confidence. The expression formula is: , where is the confidence, is the uncertainty measure when predicting the target category, is the maximum prediction variance, represents the maximum possible information entropy, is the total number of classification categories, denotes the output fluctuation under different weight samplings. The final confidence score of the target category ranges from 0 to 1. A confidence close to 1 indicates that the classification result of the construction site target is reliable, and close to 0, the target needs to be re-evaluated.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a construction site monitoring system based on remote sensing image technology according to any one of claims 1-9, and the memory is a non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Engineering monitoring system based on remote sensing image
CN116863344A
Cited By
Transmission and transformation project progress identification method and device based on space-air-ground integration technology
CN120656068A
A method and device for identifying the progress of power transmission and transformation projects based on integrated space-air-ground technology
CN120656068B
Method for automatically transparentizing white edges of multistage image slices of three-dimensional map
CN120726072A