Remote sensing big data automatic processing method based on cloud computing
By adopting a cloud-based automated processing method for remote sensing big data, combined with a transmission routing task planning algorithm for low-Earth orbit remote sensing constellations and a container cloud integration model, the problems of low efficiency and poor security in remote sensing big data processing are solved. This achieves efficient and secure data transmission and flexible data processing, and improves the visualization capabilities of the results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 周军亚
- Filing Date
- 2023-04-19
- Publication Date
- 2026-04-21
AI Technical Summary
Existing remote sensing big data processing methods suffer from low processing efficiency, poor data security, and insufficient intuitive visualization of results.
We adopt a cloud-based automated processing method for remote sensing big data, combining a transmission routing task planning algorithm for low-Earth orbit remote sensing constellations and a container cloud integration model. Through data fusion, high-speed fiber optic data transmission, MapReduce cloud computing processing, hierarchical storage, and data visualization technologies, we achieve automated processing and analysis of remote sensing big data.
It improves the transmission efficiency and security of remote sensing big data, enhances the flexibility of data processing, and improves the intuitiveness of results through visualization technology.
Smart Images

Figure CN121904572A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing big data processing, and more specifically to a cloud computing-based automated processing method for remote sensing big data. Background Technology
[0002] With the rapid upgrading of space and computer technologies, remote sensing technology has made tremendous progress. Satellite remote sensing big data, derived from land, ocean, atmosphere, and human activities, is not only a new strategic resource for the nation but also an important area of competition between nations. The processing and analysis of remote sensing big data has become a crucial issue in fields such as geographic information, environmental monitoring, and agriculture. Remote sensing big data is characterized by its macroscopic nature, massive volume, objectivity, diversity, authenticity, and real-time nature. Remote sensing images are real-time reflections of objective things on the Earth's surface. Remote sensing big data is characterized by its large data volume, diverse data types, and high processing difficulty. Traditional single-machine processing methods can no longer meet the needs of remote sensing big data processing. Therefore, using cloud computing technology to automate the processing and analysis of remote sensing big data has become a hot research topic. Currently, some cloud computing-based remote sensing big data processing methods have been proposed. Researchers at home and abroad have improved the classification and recognition accuracy of remote sensing images through deep learning, machine learning, and other methods.
[0003] However, despite some research results, existing methods still suffer from low processing efficiency, poor data security, and insufficient visualization of results in the processing and analysis of remote sensing big data. Therefore, developing an automated remote sensing big data processing method based on cloud computing is of great practical significance. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention discloses a cloud-based automated processing method for remote sensing big data. This method improves the efficiency, security, and flexibility of automated data processing and data transmission of remote sensing big data by integrating a transmission routing task planning algorithm for low-Earth orbit remote sensing constellations and a container cloud integration model while simultaneously achieving automated processing of remote sensing big data based on cloud computing.
[0005] The present invention adopts the following technical solution:
[0006] A cloud-based automated processing method for remote sensing big data, wherein the method includes:
[0007] (S1) Acquire remote sensing data, including satellite imagery, aerial photography, and radar data. The received remote sensing data is fused using a data fusion module. This data fusion module uses a data fusion function, Q, to achieve data fusion. VAC for:
[0008]
[0009] In equation (1), δ(t) represents the value of the information collision variable generated during the fusion process, m represents the fusion redundancy, t represents the data fusion delay, and ω k The Fourier transform constant represents the fusion process, K represents the fusion data type, A(n) represents the data matrix to be fused, and u k (t) represents the degree of fusion of the output data information of the fusion function. γ represents the data fusion error adjustment variable, ρ represents the data fusion margin, and ρ represents the noise generated during the fusion process.
[0010] (S2) Remote sensing data is transmitted via a high-speed fiber optic data transmission link module;
[0011] The high-speed fiber optic data transmission link module uses a transmission routing task planning algorithm model for low-Earth orbit remote sensing constellations to upload to the distributed file system of the cloud computing platform data processing center. The distributed file system forms a remote sensing file system network by connecting remote sensing data nodes from multiple locations.
[0012] The transmission routing task planning algorithm model includes an image conversion module, a data protocol identification module, a data flow routing module, a task scheduling module, an optimal path algorithm model, and a planning output module. The image conversion module improves the data conversion capability of received data information through format conversion, encoding recognition, or image resolution recognition. The data protocol identification module identifies the protocols in data communication. The data flow routing module routes the data information to the task scheduling module according to the identified protocols. The task scheduling module schedules the identified data information. The optimal path algorithm model plans the optimal data transmission route and improves data transmission efficiency. The planning output module transmits the optimal path to the data output control port of the high-speed fiber optic data transmission link module.
[0013] The output of the image conversion module is connected to the input of the data protocol recognition module, the output of the data protocol recognition module is connected to the input of the data stream diversion module, the output of the data stream diversion module is connected to the input of the task mobilization module, the output of the task mobilization module is connected to the input of the optimal path algorithm model, and the output of the optimal path algorithm model is connected to the input of the planning output module.
[0014] (S3) The MapReduce cloud computing processing model is used to perform distributed processing and analysis on remote sensing data. The analysis results are transmitted to the microservice architecture toolkit unit. The atomic service layer of the microservice architecture toolkit unit receives the remote sensing data information and configures and manages the remote sensing data information to obtain the remote sensing data compressed package.
[0015] The MapReduce cloud computing processing model includes a similarity calculation module, a classification model, and a regression algorithm model. The similarity calculation module is used to calculate the similarity between adjacent information points of each data unit. The classification model classifies remote sensing image processing scenes according to pixel size. The regression algorithm model is used to regress and predict the time series data of remote sensing images.
[0016] The output of the similarity calculation module is connected to the input of the classification model, and the output of the classification model is connected to the input of the regression algorithm model.
[0017] (S4) The remote sensing data compressed package is saved to a hierarchical storage data warehouse on the cloud computing platform. The hierarchical storage data warehouse uses open-source virtualization container technology to achieve seamless operation and storage across different machine environments and uses a container cloud integration model to uniformly debug data processing algorithms with different learning depths on the cloud platform; the data information is continuously updated by continuously updating the constraint function; the constraint function Fm is continuously updated. l (t) is:
[0018]
[0019] In equation (2), m li Indicates the initial update time coefficient, g l Indicates the constraint width, g i Indicates the constraint depth, k l D represents the factors affecting communication link delay. l The factors affecting communication stability are represented by β, where β represents the center frequency noise.
[0020] (S5) The processing results are displayed and analyzed using the Tableau data visualization platform through data visualization technology.
[0021] Furthermore, the remote sensing big data mining process is as follows:
[0022] (i) Using convolutional neural network features to acquire and integrate remote sensing image data of artificial buildings and natural geographic environment to obtain remote sensing image data packets, storing the remote sensing image data packets in cloud storage, and processing the remote sensing image data packets by denoising, sampling and filtering on the cloud platform to form a remote sensing image dataset.
[0023] (ii) Use template matching to match dataset information with target information, find image information in the image that matches the template features, and classify the image based on the snake model according to the image contour;
[0024] (iii) Extract page fingerprint features from the classified image information data using a nonlinear web page classification model, and classify the data according to the correlation of the page fingerprint features;
[0025] (iv) The classified dataset is further broken down, and the decision tree algorithm model is deployed in the same data cluster using batch processing methods to further explore the cohesion between these data.
[0026] (v) Select a business intelligence dashboard tool to visualize the classification status, relevance and cohesion of the dataset, and users can understand and analyze remote sensing data based on the visualized management interface.
[0027] Furthermore, the big data automated processing method includes the following steps:
[0028] (1) Using web crawling and batch import techniques, remote sensing data with the same feature values are collected from the remote sensing information database by using location data acquisition commands;
[0029] (2) The collected remote sensing data is cleaned, filtered and converted in format to reduce data calculation errors and generate a remote sensing data cluster. The distributed computing framework Spark is used to analyze the remote sensing data cluster through Spark SQL database technology.
[0030] (3) A clustering model is constructed using convolutional neural network machine learning. This model is then used to classify the analyzed remote sensing data clusters using a similarity matrix. The clustering model uses a similarity relation construction function for categorical data to complete the clustering process. The similarity relation construction function θ2(p m ,p n )for:
[0031]
[0032] In formula (3), r represents the classification dimension parameter, t represents the similarity cohesion coefficient, and p m p represents the width of the similarity matrix. n Let ξ1 represent the adjacent width covariance and ξ2 represent the adjacent length covariance.
[0033] (4) Store the remote sensing data after classification by similarity matrix into a MongoDB database based on a file-interactive operating system;
[0034] (5) The data visualization unit calls remote sensing data and displays the processing and analysis results to users through the Tableau visualization data platform.
[0035] Furthermore, the distributed processing and analysis process adopts the HDFS distributed file system, which uses a master-slave structure. The HDFS distributed file system cluster is the master server that manages the file namespace and regulates client access to files. The master and slave servers connect data nodes to store and manage data information. The HDFS distributed file system exposes the file namespace and allows user data to be stored in the form of files.
[0036] Furthermore, the steps of the transmission routing task planning algorithm for low-Earth orbit remote sensing constellations are as follows:
[0037] (Step 1) Obtain the router IP data packet file for the imaging task in all physically planned remote sensing missions. The data packet file set contains the wireless satellite communication node address for acquiring remote sensing data, the remote sensing data packet file size data, and the specific acquisition time node.
[0038] (Step 2) Measure the delay of adjacent routers, and update the transmit response matrix N and receive response matrix M at the start of the delay. Matrix N contains the round-trip time measurement information for data packet transmission, and matrix M contains the satellite node ID and the amount of data K to be transmitted by the satellite node. i Total set of routing links P i Defined as the set of remote sensing satellite nodes that have a ground-based time window that satisfies the constraints within the current time delay timescale, set P i Includes remote sensing satellite transmission node and air-to-ground transmission time port information K j ;
[0039] (Step 3) Construct an undirected network graph MAP = (V, E) using the shared information matrix V of the nearest router at the delay time, where V contains all remote sensing satellite location nodes and E contains the relevant information of every two adjacent nodes. When the remote sensing satellite link is enabled, the edge weight between two adjacent remote sensing satellite nodes is 1; when the remote sensing satellite link is disconnected, the edge weight between two adjacent remote sensing satellite nodes is infinite.
[0040] (Step 4) Exhaustively search for each satellite node j in the remote sensing satellite node set. First, determine whether node j exists in the remote sensing data set. If it exists, download the remote sensing data to the satellite that is transmitting the data. The satellite first downloads data to itself, and the amount of data downloaded is DSG. Then, update the M and N matrices. DSG is defined as DSG = min{K} i ,P i ,K j};
[0041] (Step 5) If satellite node j does not exist in the remote sensing data set, then iteratively search for every satellite node i in the set, using node i as the source and node j as the destination. Utilize the shortest path first algorithm to find the path with the minimum hop count. If a shortest path exists, then node j represents the amount of data (DSG) transmitted from node i to the satellite. Update matrices M and N, delete this used path from the network topology graph, update the undirected network graph MAP, and begin a new round of iterative search until node j contains the required amount of remote sensing data (DG). j The loop ends when the value is 0, and a list of downlink data volume DSSL is generated.
[0042] (Step 6) Determine the changes in the data transmission status of remote sensing satellite inter-satellite and ground-to-satellite links in each shortest path in the previous time slice during each time delay experiment time slice switching. If there are no changes, merge the two time slices, retain the transmission path, and update the M and N matrices. Delete the determined path in the network topology graph and update the undirected network graph MAP. Repeat steps 2 to 4 until all remote sensing data transmission nodes have finished transmitting, thus completing the routing task planning.
[0043] Furthermore, the MapReduce cloud computing processing model uses Map task commands to decompose complex tasks into simple tasks during the Map phase and performs a global aggregation of the decomposition results from the Map phase during the Reduce phase.
[0044] Furthermore, the training steps for the container cloud integration model are as follows:
[0045] Step (1): View the images required by the local server from the public and private repositories and use the image library to pull the image to obtain the required basic environment image. Based on the basic environment image, add the image startup container file package and uniformly set the remote connection tool login port for the image startup container file package to train the code.
[0046] Step (2): Connect the local integrated development environment to the server terminal, write code allocation instructions on the server terminal master node, allocate the container built by the image to the underlying operating system, allocate the graphics processor memory to the underlying operating system and store the node log file data to the corresponding node container, write algorithm code in the master node scheduler and set up a unified interface for the algorithm model, write the container cloud integration model development parameters into the Java object and transmit them to the container computing node.
[0047] Step (3): The container computing node trains the container cloud integration model and outputs the model output log in the form of a simplified Java object spectrum. The model output log is sent to the basic input port of the master node.
[0048] Step (4): The master node's basic input port receives the model output logs and calls the visualization service in the master node's visualization management interface to visualize the model output logs and the parameter results of the model debugging process.
[0049] Furthermore, the cloud computing processing model includes a management and automation layer and a cloud service layer. The management and automation layer adopts automated deployment, expansion and backup recovery operations to manage and schedule virtual resources of the cloud platform. The cloud service layer uses infrastructure as a service units, platform as a service units and software as a service units to provide computing services, storage services and network services.
[0050] Furthermore, the Tableau data visualization platform employs a Level of Detail (LOD) data model to update the canvas data grid and page data panes in real time, enabling a LOD LOD data visualization experience.
[0051] Furthermore, the similarity calculation module's similarity calculation model is used to calculate the similarity between adjacent information points of each data unit. The similarity calculation model employs a regression classification function to classify adjacent information points, and the regression classification function K(t) is:
[0052] K(t) = jmod(b1t) 2θ ) / b m +1(4)
[0053] In formula (4), b1 represents the synchronous regression parameter, b m denoted as asynchronous transmission noise, j represents the phase quadrature difference, mod(·) represents the phase classification function, and θ represents the error amplification parameter.
[0054] Positive and beneficial effects:
[0055] This invention employs a transmission routing task planning algorithm for low-Earth orbit remote sensing constellations during the remote sensing data transmission process, thereby achieving flexibility in algorithm development, training, and deployment, and improving the efficiency of large-scale remote sensing data transmission.
[0056] This invention employs a container cloud integration model during the uploading and storage of remote sensing data compressed packages, enabling the integration of different deep learning algorithms onto the same physical machine, thereby improving the security and flexibility of data processing.
[0057] This invention improves the remote sensing data acquisition capability through a cloud-based automated remote sensing big data processing method, enhances the data transmission capability of the link module through high-speed fiber optic data, and uploads the transmitted data information to the distributed file system of the cloud computing platform data processing center after routing the task planning algorithm model, greatly improving the data information computing capability. Attached Figure Description
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:
[0059] Figure 1 This is a schematic diagram of a cloud computing-based automated processing method for remote sensing big data according to the present invention.
[0060] Figure 2 This is a schematic diagram of the transmission routing task planning algorithm for a low-Earth orbit remote sensing constellation, which is part of the cloud computing-based automated processing method for remote sensing big data according to the present invention.
[0061] Figure 3 This is a schematic diagram of the container cloud integration model training steps of a cloud computing-based automated processing method for remote sensing big data according to the present invention.
[0062] Figure 4 This is a schematic diagram of the cloud computing processing model architecture of a cloud computing-based automated processing method for remote sensing big data according to the present invention. Detailed Implementation
[0063] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0064] like Figure 1 As shown, a cloud-based automated processing method for remote sensing big data includes: (S1) acquiring remote sensing data, including satellite image data, aerial photography data, and radar data; fusing the received remote sensing data information through a data fusion module, wherein the data fusion module achieves data information fusion through a data fusion function Q. VAC for:
[0065]
[0066] In equation (1), δ(t) represents the value of the information collision variable generated during the fusion process, m represents the fusion redundancy, t represents the data fusion delay, and ω k The Fourier transform constant represents the fusion process, K represents the fusion data type, A(n) represents the data matrix to be fused, and u k (t) represents the degree of fusion of the output data information of the fusion function. γ represents the data fusion error adjustment variable, ρ represents the data fusion margin, and ρ represents the noise generated during the fusion process.
[0067] In specific embodiments, fusing different data information can improve the processing capability of remote sensing data. Information collision variable values occur when data information of different modes or formats is fused. Fusion redundancy refers to the accuracy of data information fusion during fusion, and various information reserves may appear during fusion. Data fusion delay represents the duration during which different data information is combined into one during the fusion process. In specific applications, Fourier transform constants are used to improve the accuracy of information calculation in order to enhance data information computation capabilities. The data types to be fused are usually multiple types of data information, such as PDF, image, and video formats. During the fusion process, multiple information quantities are combined to improve the data information processing capability. The fusion function outputs the data information fusion degree, which represents the degree of data information fusion. The data matrix to be fused is a large data information matrix that aggregates the fused data information to improve the data information computation capability. Data fusion errors are prone to occur during data fusion, and the data information is adjusted using a data fusion error adjustment variable. Due to the variability of data types, multiple information is converted into microscopic function quantities during application to improve data information analysis capabilities.
[0068] (S2) Remote sensing data is transmitted via a high-speed fiber optic data transmission link module;
[0069] The high-speed fiber optic data transmission link module uses a transmission routing task planning algorithm model for low-Earth orbit remote sensing constellations to upload to the distributed file system of the cloud computing platform data processing center. The distributed file system forms a remote sensing file system network by connecting remote sensing data nodes from multiple locations.
[0070] The transmission routing task planning algorithm model includes an image conversion module, a data protocol identification module, a data flow routing module, a task scheduling module, an optimal path algorithm model, and a planning output module. The image conversion module improves the data conversion capability of received data information through format conversion, encoding recognition, or image resolution recognition. The data protocol identification module identifies the protocols in data communication. The data flow routing module routes the data information to the task scheduling module according to the identified protocols. The task scheduling module schedules the identified data information. The optimal path algorithm model plans the optimal data transmission route and improves data transmission efficiency. The planning output module transmits the optimal path to the data output control port of the high-speed fiber optic data transmission link module.
[0071] The output of the image conversion module is connected to the input of the data protocol recognition module, the output of the data protocol recognition module is connected to the input of the data stream diversion module, the output of the data stream diversion module is connected to the input of the task mobilization module, the output of the task mobilization module is connected to the input of the optimal path algorithm model, and the output of the optimal path algorithm model is connected to the input of the planning output module.
[0072] (S3) The MapReduce cloud computing processing model is used to perform distributed processing and analysis on remote sensing data. The analysis results are transmitted to the microservice architecture toolkit unit. The atomic service layer of the microservice architecture toolkit unit receives the remote sensing data information and configures and manages the remote sensing data information to obtain the remote sensing data compressed package.
[0073] The MapReduce cloud computing processing model includes a similarity calculation module, a classification model, and a regression algorithm model. The similarity calculation module is used to calculate the similarity between adjacent information points of each data unit. The classification model classifies remote sensing image processing scenes according to pixel size. The regression algorithm model is used to regress and predict the time series data of remote sensing images.
[0074] The output of the similarity calculation module is connected to the input of the classification model, and the output of the classification model is connected to the input of the regression algorithm model.
[0075] (S4) The remote sensing data compressed package is saved to a hierarchical storage data warehouse on the cloud computing platform. The hierarchical storage data warehouse uses open-source virtualization container technology to achieve seamless operation and storage across different machine environments and uses a container cloud integration model to uniformly debug data processing algorithms with different learning depths on the cloud platform; the data information is continuously updated by continuously updating the constraint function; the constraint function Fm is continuously updated. l (t) is:
[0076]
[0077] In equation (2), m li Indicates the initial update time coefficient, g l Indicates the constraint width, g i Indicates the constraint depth, k l D represents the factors affecting communication link delay. l The factors affecting communication stability are represented by β, where β represents the center frequency noise.
[0078] (S5) The processing results are displayed and analyzed using the Tableau data visualization platform through data visualization technology.
[0079] In a specific embodiment, the image conversion module may, for example, perform format conversion on different input data information, such as data information.
[0080] Data protocol identification modules, such as network features including handshake protocol information, byte distribution information, packet length information, time sequence information, protocol header information, flow header features, and communication behavior features, improve information identification capabilities by identifying the above data.
[0081] In a specific embodiment, the data stream routing module, upon receiving a data stream, can identify the application layer protocol used by the data stream based on the messages within the data stream. This protocol could be FTP (File Transfer Protocol), SMTP (Simple Mail Transfer Protocol), or HTTP (Hypertext Transfer Protocol), as well as the characteristics of each message, such as message length and the information carried in each message field. Then, based on the identified information, a log file is generated and sent to the aforementioned analysis platform. This routing is performed according to different data communication protocols to improve information exchange capabilities.
[0082] In the task scheduling module, there are two options for task storage: memory (the default configuration) and a database. Using memory is simple and efficient, but the downside is that if the program encounters a problem and restarts, previously executed tasks will be re-executed. Using a database, on the other hand, allows the program to resume normal operation from where it left off after a crash. Several options are available, such as setting a scheduler like BlockingScheduler: this is suitable when the scheduler is the only running process in the process hierarchy; calling the start function will block the current thread and cannot return immediately. In specific implementations, the appropriate scheduling instruction is selected based on specific requirements.
[0083] Optimal path algorithm models, such as harmonic search algorithm and gray wolf optimization algorithm, can be further summarized as follows in a more advanced embodiment, such as the improved gray wolf optimization algorithm:
[0084] Step 1: Set the population size, initial values of algorithm parameters, and maximum number of iterations;
[0085] Step 2: Initialize the gray wolf's position;
[0086] Step 3: Calculate the fitness of each gray wolf. The top three individuals with the highest fitness are denoted as Alpha Gray Wolf, Beta Gray Wolf, and Delta Gray Wolf.
[0087] Step 4: For each iteration, update the parameter α according to equation (9);
[0088] Step 5: Calculate the distances between other gray wolves and Alpha, Beta and Delta gray wolves according to equation (6), and update the positions of other gray wolves according to equations (7) and (8);
[0089] Step 6: Determine if the termination condition has been met. If the termination condition is not met, proceed to step 3 to continue execution; otherwise, end the process.
[0090] By optimizing the optimal path and minimizing the communication path, with the path as the constraint, this study applies the Grey Wolf Algorithm to obstacle avoidance path planning for mobile robots, which can improve data information application and planning capabilities. This application is not limited to the path-based methods described above.
[0091] The planning output module is used to output the data information of the above planning.
[0092] Furthermore, the remote sensing big data mining process is as follows:
[0093] (i) Using convolutional neural network features to acquire and integrate remote sensing image data of artificial buildings and natural geographic environment to obtain remote sensing image data packets, storing the remote sensing image data packets in cloud storage, and processing the remote sensing image data packets by denoising, sampling and filtering on the cloud platform to form a remote sensing image dataset.
[0094] (ii) Use template matching to match dataset information with target information, find image information in the image that matches the template features, and classify the image based on the snake model according to the image contour;
[0095] (iii) Extract page fingerprint features from the classified image information data using a nonlinear web page classification model, and classify the data according to the correlation of the page fingerprint features;
[0096] (iv) The classified dataset is further broken down, and the decision tree algorithm model is deployed in the same data cluster using batch processing methods to further explore the cohesion between these data.
[0097] (v) Select a business intelligence dashboard tool to visualize the classification status, relevance and cohesion of the dataset, and users can understand and analyze remote sensing data based on the visualized management interface.
[0098] Furthermore, the big data automated processing method includes the following steps:
[0099] (1) Using web crawling and batch import techniques, remote sensing data with the same feature values are collected from the remote sensing information database by using location data acquisition commands;
[0100] (2) The collected remote sensing data is cleaned, filtered and converted in format to reduce data calculation errors and generate a remote sensing data cluster. The distributed computing framework Spark is used to analyze the remote sensing data cluster through Spark SQL database technology.
[0101] (3) A clustering model is constructed using convolutional neural network machine learning. This model is then used to classify the analyzed remote sensing data clusters using a similarity matrix. The clustering model uses a similarity relation construction function for categorical data to complete the clustering process. The similarity relation construction function θ2(p m ,p n )for:
[0102]
[0103] In formula (3), r represents the classification dimension parameter, t represents the similarity cohesion coefficient, and p m p represents the width of the similarity matrix. n Let ξ1 represent the adjacent width covariance and ξ2 represent the adjacent length covariance.
[0104] (4) Store the remote sensing data after classification by similarity matrix into a MongoDB database based on a file-interactive operating system;
[0105] (5) The data visualization unit calls remote sensing data and displays the processing and analysis results to users through the Tableau visualization data platform.
[0106] Furthermore, the distributed processing and analysis process adopts the HDFS distributed file system, which uses a master-slave structure. The HDFS distributed file system cluster is the master server that manages the file namespace and regulates client access to files. The master and slave servers connect data nodes to store and manage data information. The HDFS distributed file system exposes the file namespace and allows user data to be stored in the form of files.
[0107] In a specific embodiment, after enabling the HDFS distributed file system, the user views the system shell user command interpretation, then views the file content of the specified directory rule, downloads the specified directory rule file and moves it to the specified server location, creates a new file directory on the server system to save the downloaded specified directory rule file, and finally modifies the administrator privileges of the directory rule file for their own use.
[0108] like Figure 2As shown, the transmission routing task planning algorithm for low-Earth orbit remote sensing constellations further follows: (Step 1) Obtain the router IP data packet file for the imaging task in all physically planned remote sensing tasks. The data packet file set contains the wireless satellite communication node address for acquiring remote sensing data, the remote sensing data packet file size data, and the specific acquisition time node.
[0109] (Step 2) Measure the delay of adjacent routers, and update the transmit response matrix N and receive response matrix M at the start of the delay. Matrix N contains the round-trip time measurement information for data packet transmission, and matrix M contains the satellite node ID and the amount of data K to be transmitted by the satellite node. i Total set of routing links P i Defined as the set of remote sensing satellite nodes that have a ground-based time window that satisfies the constraints within the current time delay timescale, set P i Includes remote sensing satellite transmission node and air-to-ground transmission time port information K j ;
[0110] (Step 3) Construct an undirected network graph MAP = (V, E) using the shared information matrix V of the nearest router at the delay time, where V contains all remote sensing satellite location nodes and E contains the relevant information of every two adjacent nodes. When the remote sensing satellite link is enabled, the edge weight between two adjacent remote sensing satellite nodes is 1; when the remote sensing satellite link is disconnected, the edge weight between two adjacent remote sensing satellite nodes is infinite.
[0111] (Step 4) Exhaustively search for each satellite node j in the remote sensing satellite node set. First, determine whether node j exists in the remote sensing data set. If it exists, download the remote sensing data to the satellite that is transmitting the data. The satellite first downloads data to itself, and the amount of data downloaded is DSG. Then, update the M and N matrices. DSG is defined as DSG = min{K} i ,P i ,K j};
[0112] (Step 5) If satellite node j does not exist in the remote sensing data set, then iteratively search for every satellite node i in the set, using node i as the source and node j as the destination. Utilize the shortest path first algorithm to find the path with the minimum hop count. If a shortest path exists, then node j represents the amount of data (DSG) transmitted from node i to the satellite. Update matrices M and N, delete this used path from the network topology graph, update the undirected network graph MAP, and begin a new round of iterative search until node j contains the required amount of remote sensing data (DG). j The loop ends when the value is 0, and a list of downlink data volume DSSL is generated.
[0113] (Step 6) Determine the changes in the data transmission status of remote sensing satellite inter-satellite and ground-to-satellite links in each shortest path in the previous time slice during each time delay experiment time slice switching. If there are no changes, merge the two time slices, retain the transmission path, and update the M and N matrices. Delete the determined path in the network topology graph and update the undirected network graph MAP. Repeat steps 2 to 4 until all remote sensing data transmission nodes have finished transmitting, thus completing the routing task planning.
[0114] Furthermore, the MapReduce cloud computing processing model uses Map task commands to decompose complex tasks into simple tasks during the Map phase and performs a global aggregation of the decomposition results from the Map phase during the Reduce phase.
[0115] In a specific embodiment, the MapReduce cloud computing processing model is divided into two stages. The first stage, the Map stage, can process large-scale data in parallel, dividing the data into multiple small blocks. Each small block can be processed on different computing nodes, thus significantly improving the data processing speed of the cloud computing model. The Reduce stage integrates the data processed in the previous stage, restoring the data structure to its original state. Figure 3 As shown, the training steps for the container cloud integration model are as follows:
[0116] Step (1): View the images required by the local server from the public and private repositories and use the image library to pull the image to obtain the required basic environment image. Based on the basic environment image, add the image startup container file package and uniformly set the remote connection tool login port for the image startup container file package to train the code.
[0117] Step (2): Connect the local integrated development environment to the server terminal, write code allocation instructions on the server terminal master node, allocate the container built by the image to the underlying operating system, allocate the graphics processor memory to the underlying operating system and store the node log file data to the corresponding node container, write algorithm code in the master node scheduler and set up a unified interface for the algorithm model, write the container cloud integration model development parameters into the Java object and transmit them to the container computing node.
[0118] Step (3): The container computing node trains the container cloud integration model and outputs the model output log in the form of a simplified Java object spectrum. The model output log is sent to the basic input port of the master node.
[0119] Step (4): The master node's basic input port receives the model output logs and calls the visualization service in the master node's visualization management interface to visualize the model output logs and the parameter results of the model debugging process.
[0120] Specifically, container cloud technology can integrate different deep learning algorithms onto the same physical machine. Compared with traditional virtualization technology, it has advantages such as fine granularity, scalability, flexibility and support for GPU computing, which makes it possible to isolate the environment and resources of a large number of deep learning algorithms, and improves the flexibility of algorithm development, training and deployment.
[0121] like Figure 4 As shown, the cloud computing processing model further includes a management and automation layer and a cloud service layer. The management and automation layer adopts automated deployment, expansion and backup recovery operations to manage and schedule virtual resources of the cloud platform. The cloud service layer uses infrastructure as a service units, platform as a service units and software as a service units to provide computing services, storage services and network services.
[0122] In a specific embodiment, the management and automation layer includes a physical infrastructure module, which includes data centers, servers, and network equipment to provide infrastructure support for cloud computing; the cloud service layer includes a virtualization module and a security and monitoring module. The virtualization module uses virtualization technology to virtualize physical resources into logical resources. The virtualization technology includes server virtualization, storage virtualization, and network virtualization. The security and monitoring module provides security and monitoring services to ensure the secure and stable operation of cloud computing. The security and monitoring technologies include identity authentication, access control, and log monitoring.
[0123] Furthermore, the Tableau data visualization platform employs a Level of Detail (LOD) data model to update the canvas data grid and page data panes in real time, enabling a LOD LOD data visualization experience.
[0124] In a specific embodiment, the Tableau desktop analytics tool of the Tableau data visualization platform can connect to a variety of data sources. After connecting to the data source, the desktop analytics tool can quickly create interactive, beautiful, and intelligent views and dashboards by dragging and dropping. The high-performance data engine of the Tableau data visualization platform can process data quickly, and the visualization analysis of millions of data points can be completed through mouse operations, allowing users to quickly obtain the answers they need during operation.
[0125] Furthermore, the similarity calculation module's similarity calculation model is used to calculate the similarity between adjacent information points of each data unit. The similarity calculation model employs a regression classification function to classify adjacent information points, and the regression classification function K(t) is:
[0126] K(t) = jmod(b1t) 2θ ) / b m +1(4)
[0127] In formula (4), b1 represents the synchronous regression parameter, bm denoted as asynchronous transmission noise, j represents the phase quadrature difference, mod(·) represents the phase classification function, and θ represents the error amplification parameter.
[0128] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Various omissions, substitutions, and changes can be made to the details of the methods and systems described above without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result using substantially the same method falls within the scope of the present invention. Therefore, the scope of the present invention is defined only by the appended claims.
Claims
1. A cloud computing-based automated processing method for remote sensing big data, characterized in that: The method includes: (S1) Acquire remote sensing data, including satellite imagery, aerial photography, and radar data. The received remote sensing data is fused using a data fusion module. This data fusion module uses a data fusion function, Q, to achieve data fusion. VAC for: In equation (1), δ(t) represents the value of the information collision variable generated during the fusion process, m represents the fusion redundancy, t represents the data fusion delay, and ω k The Fourier transform constant represents the fusion process, K represents the fusion data type, A(n) represents the data matrix to be fused, and u k (t) represents the degree of fusion of the output data information of the fusion function. denoted by the data fusion error adjustment variable, Y represents the data fusion margin, and ρ represents the noise generated during the fusion process; (S2) Remote sensing data is transmitted through a high-speed optical fiber data transmission link module. The high-speed fiber optic data transmission link module uses a transmission routing task planning algorithm model for low-Earth orbit remote sensing constellations to upload to the distributed file system of the cloud computing platform data processing center. The distributed file system forms a remote sensing file system network by connecting remote sensing data nodes from multiple locations. The transmission routing task planning algorithm model includes an image conversion module, a data protocol identification module, a data flow routing module, a task scheduling module, an optimal path algorithm model, and a planning output module. The image conversion module improves the data conversion capability of received data information through format conversion, encoding recognition, or image resolution recognition. The data protocol identification module identifies the protocols in data communication. The data flow routing module routes the data information to the task scheduling module according to the identified protocols. The task scheduling module schedules the identified data information. The optimal path algorithm model plans the optimal data transmission route and improves data transmission efficiency. The planning output module transmits the optimal path to the data output control port of the high-speed fiber optic data transmission link module. The output of the image conversion module is connected to the input of the data protocol recognition module, the output of the data protocol recognition module is connected to the input of the data stream diversion module, the output of the data stream diversion module is connected to the input of the task mobilization module, the output of the task mobilization module is connected to the input of the optimal path algorithm model, and the output of the optimal path algorithm model is connected to the input of the planning output module. (S3) The MapReduce cloud computing processing model is used to perform distributed processing and analysis on remote sensing data. The analysis results are transmitted to the microservice architecture toolkit unit. The atomic service layer of the microservice architecture toolkit unit receives the remote sensing data information and configures and manages the remote sensing data information to obtain the remote sensing data compressed package. The MapReduce cloud computing processing model includes a similarity calculation module, a classification model, and a regression algorithm model. The similarity calculation module is used to calculate the similarity between adjacent information points of each data unit. The classification model classifies remote sensing image processing scenes according to pixel size. The regression algorithm model is used to regress and predict the time series data of remote sensing images. The output of the similarity calculation module is connected to the input of the classification model, and the output of the classification model is connected to the input of the regression algorithm model. (S4) The remote sensing data compressed package is saved to a hierarchical storage data warehouse on the cloud computing platform. The hierarchical storage data warehouse uses open-source virtualization container technology to achieve seamless operation and storage across different machine environments and uses a container cloud integration model to uniformly debug data processing algorithms with different learning depths on the cloud platform; the data information is continuously updated by continuously updating the constraint function; the constraint function Fm is continuously updated. l (t) is: In equation (2), m li Indicates the initial update time coefficient, g l Indicates the constraint width, g i Indicates the constraint depth, k l D represents the factors affecting communication link delay. l The factors affecting communication stability are represented by β, where β represents the center frequency noise. (S5) The processing results are displayed and analyzed using the Tableau data visualization platform through data visualization technology.
2. The automated remote sensing big data processing method based on cloud computing according to claim 1, characterized in that: The remote sensing big data mining process is as follows: (i) Using convolutional neural network features to acquire and integrate remote sensing image data of artificial buildings and natural geographic environment to obtain remote sensing image data packets, storing the remote sensing image data packets in cloud storage, and processing the remote sensing image data packets by denoising, sampling and filtering on the cloud platform to form a remote sensing image dataset. (ii) Use template matching to match dataset information with target information, find image information in the image that matches the template features, and classify the image based on the snake model according to the image contour; (iii) Extract page fingerprint features from the classified image information data using a nonlinear web page classification model, and classify the data according to the correlation of the page fingerprint features; (iv) The classified dataset is further broken down, and the decision tree algorithm model is deployed in the same data cluster using batch processing methods to further explore the cohesion between these data. (v) Select a business intelligence dashboard tool to visualize the classification status, relevance and cohesion of the dataset, and users can understand and analyze remote sensing data based on the visualized management interface.
3. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The automated big data processing method includes the following steps: (1) Using web crawling and batch import techniques, remote sensing data with the same feature values are collected from the remote sensing information database by using location data acquisition commands; (2) The collected remote sensing data is cleaned, filtered and converted in format to reduce data calculation errors and generate a remote sensing data cluster. The distributed computing framework Spark is used to analyze the remote sensing data cluster through Spark SQL database technology. (3) A clustering model is constructed using convolutional neural network machine learning. This model is then used to classify the analyzed remote sensing data clusters using a similarity matrix. The clustering model uses a similarity relation construction function for categorical data to complete the clustering process. The similarity relation construction function θ2(p m ,p n )for: In formula (3), r represents the classification dimension parameter, t represents the similarity cohesion coefficient, and p m p represents the width of the similarity matrix. n Let ξ1 represent the adjacent width covariance and ξ2 represent the adjacent length covariance. (4) Store the remote sensing data after classification by similarity matrix into a MongoDB database based on a file-interactive operating system.
4. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The distributed processing and analysis process uses the HDFS distributed file system, which adopts a master-slave structure. The HDFS distributed file system cluster is the master server that manages the file namespace and regulates client access to files. The master and slave servers connect data nodes to store and manage data information. The HDFS distributed file system exposes the file namespace and allows user data to be stored in the form of files.
5. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The steps of the transmission routing task planning algorithm for low-Earth orbit remote sensing constellations are as follows: (Step 1) Obtain the router IP data packet file for the imaging task in all physically planned remote sensing missions. The data packet file set contains the wireless satellite communication node address for acquiring remote sensing data, the remote sensing data packet file size data, and the specific acquisition time node. (Step 2) Measure the delay of adjacent routers, and update the transmit response matrix N and receive response matrix M at the start of the delay. Matrix N contains the round-trip time measurement information for data packet transmission, and matrix M contains the satellite node ID and the amount of data K to be transmitted by the satellite node. i Total set of routing links P i Defined as the set of remote sensing satellite nodes that have a ground-based time window that satisfies the constraints within the current time delay timescale, set P i Includes remote sensing satellite transmission node and air-to-ground transmission time port information K j ; (Step 3) Construct an undirected network graph MAP = (V, E) using the shared information matrix V of the nearest router at the delay time, where V contains all remote sensing satellite location nodes and E contains the relevant information of every two adjacent nodes. When the remote sensing satellite link is enabled, the edge weight between two adjacent remote sensing satellite nodes is 1; when the remote sensing satellite link is disconnected, the edge weight between two adjacent remote sensing satellite nodes is infinite. (Step 4) Exhaustively search for each satellite node j in the remote sensing satellite node set. First, determine whether node j exists in the remote sensing data set. If it exists, download the remote sensing data to the satellite that is transmitting the data. The satellite first downloads data to itself, and the amount of data downloaded is DSG. Then update the M and N matrices. DSG is defined as DSG = min{K i P i K j }; (Step 5) If satellite node j does not exist in the remote sensing data set, then iteratively search for every satellite node i in the set, using node i as the source and node j as the destination. Utilize the shortest path first algorithm to find the path with the minimum hop count. If a shortest path exists, then node j represents the amount of data (DSG) transmitted from node i to the satellite. Update matrices M and N, delete this used path from the network topology graph, update the undirected network graph MAP, and begin a new round of iterative search until node j contains the required amount of remote sensing data (DG). j The loop ends when the value is 0, and a list of downlink data volume DSSL is generated. (Step 6) Determine the changes in the data transmission status of remote sensing satellite inter-satellite and ground-to-satellite links in each shortest path in the previous time slice during each time delay experiment time slice switching. If there are no changes, merge the two time slices, retain the transmission path, and update the M and N matrices. Delete the determined path in the network topology graph and update the undirected network graph MAP. Repeat steps 2 to 4 until all remote sensing data transmission nodes have finished transmitting, thus completing the routing task planning.
6. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The MapReduce cloud computing processing model uses Map task commands to decompose complex tasks into simple tasks during the Map phase and then globally summarizes the results of the Map phase decomposition during the Reduce phase.
7. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The training steps for the container cloud integration model are as follows: Step (1): View the images required by the local server from the public and private repositories and use the image library to pull the image to obtain the required basic environment image. Based on the basic environment image, add the image startup container file package and uniformly set the remote connection tool login port for the image startup container file package to train the code. Step (2): Connect the local integrated development environment to the server terminal, write code allocation instructions on the server terminal master node, allocate the container built by the image to the underlying operating system, allocate the graphics processor memory to the underlying operating system and store the node log file data to the corresponding node container, write algorithm code in the master node scheduler and set up a unified interface for the algorithm model, write the container cloud integration model development parameters into the Java object and transmit them to the container computing node. Step (3): The container computing node trains the container cloud integration model and outputs the model output log in the form of a simplified Java object spectrum. The model output log is sent to the basic input port of the master node. Step (4): The master node's basic input port receives the model output logs and calls the visualization service in the master node's visualization management interface to visualize the model output logs and the parameter results of the model debugging process.
8. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The cloud computing processing model includes a management and automation layer and a cloud service layer. The management and automation layer uses automated deployment, expansion and backup recovery operations to manage and schedule virtual resources of the cloud platform. The cloud service layer uses infrastructure as a service units, platform as a service units and software as a service units to provide computing services, storage services and network services.
9. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The Tableau data visualization platform uses a Level of Detail (LOD) data model to update the canvas data grid and page data panes in real time, enabling a LOD data visualization experience.
10. The automated processing method for remote sensing big data based on cloud computing according to claim 1, characterized in that: The similarity calculation module's similarity calculation model is used to calculate the similarity between adjacent information points for each data unit. The similarity calculation model uses a regression classification function K(t) to classify adjacent information points: K(t)=jmod(b1t 2θ ) / b m +1(4) In formula (4), b1 represents the synchronous regression parameter, b m denoted as asynchronous transmission noise, j represents the phase quadrature difference, mod(·) represents the phase classification function, and θ represents the error amplification parameter.