Computer system and method for cloud-enabled data processing platform

The cloud-enabled data processing platform addresses inefficiencies in traditional systems by leveraging machine learning and secure data transfer mechanisms, ensuring efficient and secure data processing and analysis for environmental monitoring and disaster response.

WO2026044394A1PCT designated stage Publication Date: 2026-03-05MDA SYST LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Traditional data processing platforms face challenges such as high costs, inefficiencies, and complexity in handling large datasets, limited scalability, and lack of interoperability, which hinder accessibility and delay valuable insights in environmental monitoring and disaster response.

Method used

A cloud-enabled data processing platform utilizing machine learning, sensor-agnostic data collection, secure push-pull data transfer, and local caching techniques, enabling efficient, scalable, and secure data processing and analysis of diverse data sources in near real-time.

Benefits of technology

The platform reduces data transfer and processing times, minimizes costs, and enhances data security and usability, facilitating timely and accurate insights for environmental monitoring and disaster response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050595_05032026_PF_FP_ABST
    Figure CA2025050595_05032026_PF_FP_ABST
Patent Text Reader

Abstract

A system, method, and server for cloud-enabled data processing are provided. The system includes receiving data, pre-processing the data, and a server including a data acquisition module for receiving the data, a data ingestion module for ingesting and pre-filtering the data, transferring the pre-processed input data to edge computing units, an extraction module for performing geocoding and enrichment on the pre-processed input data to extract multidimensional landscape features including georeferenced data, an ML pre-processing module for applying one or more supervised or unsupervised ML algorithms to analyze the georeferenced data to detect anomalies and potential threats, data archiving storage for storing data and the multidimensional landscape features according to lifecycle management rules, a data indexing and cataloguing module for integrating with data catalog services to create metadata for the archived data stored in the data archiving storage, a prioritized data transmission module for transmitting prioritized data to end users.
Need to check novelty before this filing date? Find Prior Art

Description

COMPUTER SYSTEM AND METHOD FOR CLOUD-ENABLED DATA PROCESSING PLATFORMTechnical Field

[0001] The following relates generally to data processing platforms, and more particularly to systems and methods for integrating, processing, and analyzing data from diverse sources in near real-time for environmental monitoring and disaster response.Introduction

[0002] Traditional data analytics processing platforms often employ a data-centric or “reactive” model, which requires users to download and transfer substantial datasets, irrespective of their actual value. This model imposes significant financial burdens due to costs associated with computation, storage, and data transfer. Furthermore, this approach may lead to delays as users must locate and process data, often without prior knowledge of its relevance. Such costs are incurred before assessing the applicability of the data to specific purposes, leading to notable delays in processing and extracting meaningful insights.

[0003] High initial costs and substantial processing power requirements present significant barriers to entry for prospective users. The complexity of data and processing abstraction requires users to independently manage intricate aspects of data storage and processing, which may be cumbersome and technically challenging. Additionally, reliance on continuous high-speed internet connectivity may limit usability in remote locations with limited internet service.

[0004] For user-friendly platforms for Big Data, including Earth Observation (EO) and other large datasets, it is highly desirable to offer high-level abstraction, reproducibility, scalability, and interoperability. Such platforms may significantly reduce entry barriers and facilitate more efficient and effective data management and analysis. Currently, users incur significant costs due to the inefficiencies of traditional platforms. The complexity of data and processing abstraction exacerbates the difficulty for users to access and manage extensive datasets, thereby increasing operational costs.

[0005] The lack of infrastructure replicability and limited interoperability with external tools impede flexibility and integration, resulting in delays in processing and extracting valuable insights. Moreover, substantial initial costs and processing power requirements in traditional data analytics pose significant barriers to entry for potential users, limiting the accessibility and practical utility of Big Data for scientific and practical applications.

[0006] Visualization tools in modem platforms leverage advanced technologies to provide immersive and interactive experiences. These tools transform raw data into meaningful visual representations, enabling users to interpret and analyze geospatial information. Key features often include near real-time rendering, dynamic layering, and customizable dashboards that allow users to tailor the display to their specific requirements. In remote locations with limited internet access, visualization tools may be designed to operate efficiently with minimal bandwidth. They employ data compression techniques and local caching to ensure smooth performance and user experience. This capability is important in field operations and near real-time monitoring scenarios where consistent connectivity cannot be guaranteed.

[0007] The foregoing difficulties lead to several key issues in traditional EO platforms:

[0008] Data Transfer: Transferring large datasets can be expensive, especially with limited bandwidth. High-throughput requirements and data transfer inefficiencies are significant concerns.

[0009] Computational Resources: Performing data encryption and processing massive datasets requires substantial computing power, leading to high costs. The computational demands for data processing are intensive.

[0010] Storage: Storing vast amounts of data locally or in the cloud can be expensive. The increasing volume of data makes scalable storage solutions highly valuable.

[0011] Human Analysis Time: Analysts spend considerable time sifting through irrelevant data, hindering efficiency. Manual data analysis is time-consuming and inefficient.

[0012] Delays in Insights: The data-centric model leads to time-consuming processing procedures, resulting in delays in extracting valuable security insights from the data.

[0013] Limited Scalability: Traditional platforms often struggle to scale efficiently when confronted with the steadily increasing volume of data where response time is considered a key value driver.

[0014] Velocity: For high-frequency data collection and data archiving requirements, advanced systems capable of handling rapid data influx are highly desirable.

[0015] Variety: For diverse types of data, flexible systems capable of handling various data formats and sources are highly desirable. Traditional platforms often lack the flexibility needed to process and integrate heterogeneous data, resulting in fragmented and inefficient workflows.

[0016] Data Governance: Data governance in platforms may advantageously ensure that data is managed in a manner that maximizes its value while maintaining compliance with end-user licence agreements (EULA's) and ethical and security standards. Traditional platforms often lack robust governance mechanisms, which may lead to compromising data integrity and compliance.

[0017] Data Security: Ensuring the protection of data from unauthorized access and maintaining its confidentiality and integrity is very important. Such protection includes the implementation of robust security measures such as encryption, access controls, and regular security audits. Addressing varying levels of data sensitivity preferably includes a comprehensive data segregation strategy, categorizing data into public, unclassified, classified, and secret tiers. Each category may preferable have specific security protocols:

[0018] Public Data: Implementing minimal access controls to prevent unauthorized modifications while ensuring general accessibility. Comprehensive audit logs are to monitor access and modifications to maintain data integrity.

[0019] Unclassified Data: Apply encryption standards to safeguard unclassified data from unauthorized access. Enforce role-based access controls to restrict data access to authorized personnel only. Periodic security audits are preferably implemented to verify adherence to access controls and encryption protocols.

[0020] Classified Data: Utilize advanced encryption techniques to secure classified data both at rest and in transit. Implement comprehensive access controls based on the principle of least privilege, ensuring that only personnel with appropriate security clearances are able to access classified data. Detailed audit logs preferably track access and modifications, with regular reviews to identify and address unauthorized activities.

[0021] Secret Data: Employ high levels of encryption, preferably the highest level of encryption available, including frequent rotation and secure management of encryption keys. Multi-factor authentication (MFA) to be implemented to add an additional layer of security, including requiring multiple methods of identity verification for access. Secret data to be stored in isolated environments with no direct external network access, minimizing the risk of breaches. Continuous monitoring tools to be deployed to detect and respond to unauthorized access attempts or suspicious activities in real time.

[0022] Compliance: Comprehensive data security policies tailored to each data category to be developed and enforced. Regular training sessions to ensure awareness of these policies. An incident response plan to be established to promptly address any breaches. Compliance with relevant regulatory standards and industry best practices to be maintained, with policies and procedures regularly updated to reflect changes in the legal landscape. Regular audits and reviews to be performed to ensure the effectiveness of security measures and adherence to established policies.

[0023] Automated Machine Learning Algorithms: Automated machine learning (ML) algorithms may advantageously be deployed for data processing to advantageously reduce manual labor and enhance system efficiency. Many features extracted from SAR,such as ship detections, vary widely in size and shape, and so robust systems are preferably used to minimize false positives and improve detection accuracy. Automation significantly enhances the precision and reliability of oceanographic parameter estimation. The output of the automated machine learning algorithms may further be used to ask future EO missions and / or identify EO archive datasets.

[0024] Scalability of SAR Imagery: Given the massive size of high-resolution SAR imagery, scalable big data platforms and short revisit times are preferably deployed to detect changes. Traditional systems struggle with the volume and velocity of data, and so advanced big data architectures are strongly desirable to handle these demands effectively.

[0025] Indexing and Query Optimization: Efficient management and querying of large-scale Earth Observation (EO) data is a significant challenge due to the high dimensionality, spatial and temporal diversity, and sheer volume of data. Traditional database systems struggle to handle the complexity and scale of EO datasets, leading to slow data retrieval times and inefficient data processing. To overcome this gap, the development and utilization of specialized database management systems and computational frameworks are preferably tailored to the unique characteristics of EO data.

[0026] Platform Display: Tools may or may not include display options that enable users to open and view datasets. Such inconsistency may limit the ability of users to effectively visualize and analyze data, impacting the overall usability and accessibility of the platform.

[0027] Low-Bandwidth Environments: Many visualization tools are not optimized for environments with limited internet connectivity. Consequently, they often perform poorly in remote or field settings where consistent high-speed internet access is unavailable, impeding real-time data analysis and decision-making.

[0028] Accordingly, there is a need for an improved system and method for earth observation data processing that overcomes at least some of the disadvantages of existing systems and methods.Summary

[0029] A system for cloud-enabled data processing is provided. The system includes one or more input devices for receiving data, one or more data processing devices for pre-processing the received data, and a server for cloud-enabled data processing. The server includes a data acquisition module for receiving the pre- processed input data, a data ingestion module for ingesting the pre-processed input data and pre-filtering the pre-processed input data to automatically perform preliminary quality and format checks, thereby ensuring compliance with data requirements, a data transferring module for transferring the pre-processed input data to edge computing units within cloud infrastructure positioned to optimize data transfer times, an extraction module for performing geocoding and enrichment on the pre-processed input data to extract multidimensional landscape features including georeferenced data, an ML preprocessing module for performing regional processing by applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near-real-time, data archiving storage for storing raw input data corresponding to the pre-processed input data and the multidimensional landscape features according to lifecycle management rules, a data indexing and cataloguing module for integrating with data catalog services to create metadata for the archived data stored in the data archiving storage, and a prioritized data transmission module for transmitting prioritized data to end users.

[0030] The server may further include a push-pull mechanism, the push-pull mechanism including a push module for providing the prioritized data to an end user, and a pull mechanism for receiving a request for further information from the end user responsive to the provided prioritized data. The server may transmit the further information responsive to the received request.

[0031] The extraction module may perform the geocoding and enrichment through synthetic aperture radar (SAR) interferometry and / or interferometric SAR (InSAR).

[0032] The extraction module may further conduct noise reduction and radiometric correction and assign geographic coordinates to each data point within the multidimensional landscape features to produce the georeferenced features.

[0033] The created metadata may include a searchable index that facilitates data retrieval for future reference or analysis.

[0034] The prioritized data may include the georeferenced data of the multidimensional landscaped features.

[0035] The input data may be SAR data captured or otherwise generated by satellites at regular intervals.

[0036] A method of cloud-enabled data processing is provided. The method includes receiving or acquiring input data, ingesting the input data and pre-filtering the input data to automatically perform preliminary quality and format checks, transferring the input data to edge computing units within a cloud infrastructure positioned to optimize data transfer times, performing geocoding and enrichment on the input data to extract multidimensional landscape features including georeferenced data, performing regional preprocessing by applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near- real-time, archiving raw input data corresponding to the pre-processed input data and the multidimensional landscape features according to lifecycle management rules, integrating with data catalog services to create metadata for the archived data stored in the data archiving storage, and transmitting prioritized data to end users.

[0037] The method may further include receiving a request for further information from the end user responsive to the provided prioritized data and transmitting the further information responsive to the received request.

[0038] The geocoding and enrichment may be performed through SAR interferometry and / or InSAR.

[0039] The geocoding and enrichment may further include noise reduction, radiometric correction, and assigning geographic coordinates to each data point within the multidimensional landscape features to produce the georeferenced features.

[0040] The created metadata may include a searchable index that facilitates data retrieval for future reference or analysis.

[0041] The prioritized data may include the georeferenced data of the multidimensional landscaped features.

[0042] The input data may be SAR data captured or otherwise generated by satellites at regular intervals.

[0043] A server for cloud-enabled data processing is provided. The server receives pre-processed input data. The server includes a data acquisition module for receiving the pre-processed input data, a data ingestion module for ingesting the pre-processed input data and pre-filtering the pre-processed input data to automatically perform preliminary quality and format checks, thereby ensuring compliance with data requirements, a data transferring module for transferring the pre-processed input data to edge computing units within cloud infrastructure positioned to optimize data transfer times, an extraction module for performing geocoding and enrichment on the pre-processed input data to extract multidimensional landscape features including georeferenced data, an ML preprocessing module for applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near- real-time, data archiving storage for storing raw input data corresponding to the pre- processed input data and the multidimensional landscape features according to lifecycle management rules, a data indexing and cataloguing module for integrating with data catalog services to create metadata for the archived data stored in the data archiving storage, and a prioritized data transmission module for transmitting prioritized data to end users.

[0044] The server may further include a push-pull mechanism, the push-pull mechanism including a push module for providing the prioritized data to an end user, and a pull mechanism for receiving a request for further information from the end userresponsive to the provided prioritized data. The server may transmit the further information responsive to the received request.

[0045] The extraction module may perform the geocoding and enrichment through SAR interferometry and / or InSAR.

[0046] The extraction module may further conduct noise reduction and radiometric correction and assign geographic coordinates to each data point within the multidimensional landscape features to produce the georeferenced features.

[0047] The created metadata may include a searchable index that facilitates data retrieval for future reference or analysis.

[0048] The prioritized data may include the georeferenced data of the multidimensional landscaped features.

[0049] The input data may be SAR data captured or otherwise generated by satellites at regular intervals.

[0050] The cloud-enabled data processing platform (the “Platform") is developed to address the limitations of traditional data processing platforms. The Platform integrates machine learning, sensor-agnostic data collection, secure push-pull data transfer mechanisms, advanced data compression, and local caching techniques. The Platform offers a secure and scalable environment for deploying applications, tailored to the needs of organizations involved in environmental monitoring and disaster response.

[0051] Other aspects and features will become apparent, to those ordinarily skilled in the art, upon review of the following description of some exemplary embodiments.Brief Description of the Drawings

[0052] The drawings included herewith are for illustrating various examples of articles, methods, and apparatuses of the present specification. In the drawings:

[0053] Figure 1 is a schematic diagram of a system for cloud-enabled data processing, according to an embodiment;

[0054] Figure 2 is a block diagram of a device for cloud-enabled data processing, according to an embodiment;

[0055] Figure 3 is a block diagram of a computer system for cloud-enabled data processing, according to an embodiment;

[0056] Figure 4 is a flowchart of a method of cloud-enabled data processing, according to an embodiment;

[0057] Figure 5 is a schematic diagram of a computer system for cloud-enabled processing of Earth observation data, according to an embodiment; and

[0058] Figure 6 is a flowchart of a method of cloud-enabled data processing of Earth observation data, according to an embodiment.Detailed Description

[0059] Various apparatuses or processes will be described below to provide an example of each claimed embodiment. No embodiment described below limits any claimed embodiment and any claimed embodiment may cover processes or apparatuses that differ from those described below. The claimed embodiments are not limited to apparatuses or processes having all of the features of any one apparatus or process described below or to features common to multiple or all of the apparatuses described below.

[0060] One or more systems described herein may be implemented in computer programs executing on programmable computers, each comprising at least one processor, a data storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. For example, and without limitation, the programmable computer may be a programmable logic unit, a mainframe computer, server, and personal computer, cloud-based program or system, laptop, personal data assistance, cellular telephone, smartphone, or tablet device.

[0061] Each program is preferably implemented in a high-level procedural or object-oriented programming and / or scripting language to communicate with a computer system. However, the programs can be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language. Each such computer program is preferably stored on a storage media or a device readable by a general or special purpose programmable computer for configuring and operating thecomputer when the storage media or device is read by the computer to perform the procedures described herein.

[0062] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the present invention.

[0063] Further, although process steps, method steps, algorithms or the like may be described (in the disclosure and I or in the claims) in a sequential order, such processes, methods and algorithms may be configured to work in alternate orders. In other words, any sequence or order of steps that may be described does not necessarily indicate a requirement that the steps be performed in that order. The steps of processes described herein may be performed in any order that is practical. Further, some steps may be performed simultaneously.

[0064] When a single device or article is described herein, it will be readily apparent that more than one device I article (whether or not they cooperate) may be used in place of a single device I article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be readily apparent that a single device I article may be used in place of the more than one device or article.

[0065] The following relates generally to data processing platforms, and more particularly to systems and methods for integrating, processing, and analyzing data from diverse sources in near real-time for environmental monitoring and disaster response.

[0066] The foregoing may include or may be applied to earth observation and geointelligence, and more particularly to systems and methods for processing earth observation data using a cloud enabled data processing platform.

[0067] The following further relates to Big Data Analytics and geospatial processing, with an emphasis on the integration of advanced computational techniques and emerging technologies. The following includes methods and systems for the near real-time processing, analysis, and visualization of large and complex datasets derived from various sensors and data sources. These sources include Earth Observation (EO)datasets, Space Domain Awareness (SDA) data, Ground-Based Automatic Identification Systems (AIS), and other geospatial data sources. The present disclosure may advantageously handle diverse and heterogeneous data types, providing robust and scalable solutions for near real-time data processing and analytics across multiple domains. The present disclosure ensures interoperability and effective data exchange through the adoption and enhancement of relevant standards, enabling comprehensive and efficient analysis of complex datasets.

[0068] The following further relates to a cloud-enabled data processing platform (the "Platform") developed to meet the demands of modem data processing platforms specifically for environmental monitoring and disaster response.

[0069] The Platform incorporates advanced technologies, including machine learning, sensor-agnostic data collection, and secure data transfer mechanisms. These technologies enable the efficient processing and analysis of large and complex datasets from various sources, such as Earth Observation (EO) datasets, Space Domain Awareness (SDA) data, and Ground-Based Automatic Identification Systems (AIS). The Platform's near real-time analysis and robust data security features facilitate informed decision-making promptly and securely.

[0070] The computer platform of the present disclosure revolutionizes earth observation (“EO”) data processing by implementing a user-centric or “proactive” approach that prioritizes efficiency and cost-effectiveness. This approach departs from the data-centric model of existing systems by focusing processing power on the origin of data acquisition location.

[0071] The platform of the present disclosure leverages a novel architecture utilizing edge computing and regional computation systems for efficient preliminary data processing at the source, identifying the relevant portions for further analysis. This significantly reduces data transfer and processing times.

[0072] Employing machine learning (ML) algorithms, the platform analyzes the data and identifies valuable information, ensuring users only process and store the data truly relevant to their needs, leading to significant cost savings. The machine learning algorithms may include supervised or unsupervised learning algorithms.

[0073] By prioritizing efficiency and cost-effectiveness, the platform of the present disclosure offers a more advantageous solution compared to traditional data centric EO data processing platforms.

[0074] The present disclosure provides systems, methods, and devices suitable for processing diverse data types through a push-pull data transfer mechanism coupled with real-time, Al-powered analysis. For example, the proactive nature of the push-pull data transfer mechanism is particularly advantageous in applications such as wildfire monitoring and management, where timely data transfer and analysis are highly desirable. This mechanism enables the Platform to deliver real-time insights and facilitate rapid decision-making in dynamic and high-stakes environments. While this example pertains to wildfire monitoring, the system is equally applicable to other domains, including but not limited to, environmental monitoring, space domain awareness, ground- based automatic identification systems (AIS), and other sensor-based data collection systems. By integrating this mechanism, the platform supports timely and informed responses across a wide range of applications, enhancing overall operational efficiency and effectiveness.

[0075] As an example, in a remote forest area, an EO satellite using machine learning (ML) algorithms conducts routine surveillance. The ML system on board the satellite detects a wildfire by identifying patterns consistent with smoke and heat signatures from the sensed data. Upon downlinking the data into the Platform, the ML system verifies the fire based on the heat signature and its geometry and activates a push-pull data transfer mechanism to manage the response. The alert includes initial data on the location, intensity, and spread of the fire.

[0076] The central command system, powered by Al, begins real-time analysis of the fire data. The Al calculates the potential spread of the fire based on parameters such as wind speed and direction, humidity levels, soil moisture, and biomass patterns. Using the push-pull system, the central cloud system pulls detailed environmental data from remote sensing systems and weather databases. The central cloud system pushes commands to deploy additional remote sensing systems to capture higher resolutionimages and data. These systems include but are not limited to high resolution SAR, infrared, and multispectral imagery onboard aerial spacecrafts in the area of interest.

[0077] A new satellite pass over the area of interest provides high-resolution imagery that is fed into the system for continuous analysis. The Al integrates this data to update the fire’s progression and refine predictions on its potential direction and impact. Based on the integrated data, the Al system predicts the wildfire's path, considering realtime changes in weather conditions and terrain features. The Al system alerts emergency response teams with detailed maps and predicted spread patterns, enabling them to strategize and distribute resources more efficiently.

[0078] While this example pertains to wildfire monitoring, the same Platform is applicable to various other scenarios where real-time data processing and decisionmaking are highly desired. These scenarios include, but are not limited to, environmental monitoring, space domain awareness, ground-based automatic identification systems (AIS), and other sensor-based data collection systems.

[0079] The integration of sensor-agnostic data collection, advanced machine learning algorithms, and a secure push-pull data transfer mechanism offers unparalleled efficiency and accuracy in processing and analyzing Synthetic Aperture Radar (SAR) imagery. The system's automated image processing pipeline eliminates the need for manual validation, significantly reducing operational costs and human resource requirements. The machine learning algorithms, trained on extensive datasets, provide real-time, Al-powered analysis, ensuring that insights are delivered with unmatched speed and precision. This capability is particularly desirable for responding to critical situations such as wildfires and military operations, where timely and accurate information is highly desirable.

[0080] The present disclosure includes prioritized data transfer functionality for advantageously mitigating or entirely eliminating the need to transfer vast datasets. Conventional data processing methods typically require centralized computing resources to handle analysis, which includes moving large volumes of data from the point of collection (e.g., sensors, devices) to centralized servers, often resulting in high latency, significant bandwidth costs, and potential data loss and security risks during transfer. Thepresent disclosure may advantageously overcome the foregoing disadvantages through mitigating or eliminating the need for large-scale data transfers by employing edge computing and localized data processing. In localized data processing, data is processed close to the point of collection using distributed computing resources such as local servers, smart devices, and even Internet of Things (loT) devices with processing capabilities, as well as leveraging cloud-based technologies and user-driven code deployment. Users may deploy their own code (custom algorithms, filtering tools, data visualization tools) directly within the infrastructure, closer to the origin of the data. This approach may not only advantageously minimize data transfer requirements but may further enhance processing speed, reduce latency, and strengthen security by keeping data within the local infrastructure.

[0081] The present disclosure may achieve seamless data integration and up-to- date information repository management through a novel system architecture that prioritizes security. This architecture leverages a combination of technologies, which may include any one or more of:• advanced middleware with transformation capabilities, which acts as an intermediary layer, facilitating communication between a platform and various data storage solutions, and translates data requests from the platform into specific commands understood by each storage solution, regardless of its underlying protocol or data format. Additionally, the middleware may be configured to perform data transformation tasks, such as cleansing or standardizing data formats, to ensure compatibility with the platform's internal repository, which eliminates the need for custom connectors for each data source, ensuring broader compatibility.• RESTful APIs with authentication, which function as a set of standardized instructions that define how the platform interacts with each data storage solution. The middleware utilizes these APIs to translate the platform's requests into the specific language understood by the target storage solution, enabling efficient data retrieval. All APIs leverage OAuth2 authentication tokens, enabling user segregation and access control based on specific roles.• Data source (Apache Kafka) connectors with fault tolerance, which act as specialized adapters that streamline the connection process for specific data sources with unique protocols (like cloud storage or databases) and efficiently bridge the gap between the platform and these sources, ensuring seamless data transfer. These connectors are also designed with fault tolerance in mind, ensuring data transfer continues even if there are temporary disruptions with a specific data source.• enhanced security through robust protocols, whereby the platform prioritizes data security throughout the data integration process. To ensure the confidentiality and integrity of information during transfer, the system utilizes RSA encryption, a robust industry-standard algorithm. This encrypts all data in transit between the platform and various data storage solutions. Additionally, a load balancer acts as a security gateway, encrypting all data flowing to and from the TKE Kafka brokers, which are the central nervous system for real-time data streaming within the platform.

[0082] Further features of the Platform may include any one or more of the following:• Automated SAR image processing: The Platform may automate the processing of Synthetic Aperture Radar (SAR) imagery using advanced machine learning algorithms, which may advantageously eliminate any need for manual validation, thereby ensuring rapid and accurate analysis of satellite imagery and facilitating proactive responses to natural disasters by detecting changes and anomalies in the environment;• Secure push-pull data transfer mechanism: Ensuring the timely and secure delivery of data and insights through a push-pull mechanism may advantageously enhance operational efficiency and maintain data integrity during transfer, which may further enable near-real-time updates and secure data exchange, which may be extremely valuable in the emergency response context;• Sensor-agnostic data collection: The Platform may support data collection from various sensors, including light detection and ranging (“LiDAR”), infrared, optical, and SAR sensors. Such sensor-agnostic data collection may advantageouslyprovide comprehensive data capture capabilities under diverse conditions and may further offer a holistic view of the monitored area by integrating a plurality of sensor types;• Near-real-time Al powered analysis: The Platform may employ machine learning (ML) algorithms to analyze data in near real time and provide immediate insights and predictive modelling to improve the effectiveness of interventions by enabling rapid response and informed decision-making;• Segregated environment for deployment: The Platform may allow users to deploy applications in secure, isolated environments, thereby ensuring data security and optimizing performance in order to protect sensitive data against breaches and unauthorized access;• Integrated billing system: Users may provision resources and may be billed via linked credit cards, which may advantageously simplify resource management and cost tracking and provide transparency and ease of use, thereby facilitating budget management;• Application gallery: The Platform may facilitate the publication and sharing of applications, machine learning models, and data processing scripts to foster a collaborative ecosystem, thereby promoting collaboration and continuous improvement within the user community;• Scalability and flexibility: The Platform may support scalable deployment and adapt to various applications to provide a versatile tool for diverse user needs, thereby ensuring the platform grows with user requirements and thus is suitable for both small-scale and large-scale operation;• Data indexing and prioritization: The Platform may employ near-real-time object detection algorithms to detect and transfer only relevant data to improve efficiency by reducing data transfer volume, thereby ensuring only the most relevant data is analyzed, reducing bandwidth usage, and accelerating decision making;• Data compression and local caching: The Platform may implement a data compression and local caching, which may enhance the performance in remote orlow bandwidth environment to ensure smooth performance even in challenging connectivity conditions, thereby improving usability in remote locations, ensuring consistent performance;• User-defined data retention policies: The Platform may allow users to define their data retention policy to provide flexibility in managing data storage, thereby offering control over storage cost and maintaining compliance with data management regulations;• Integrated development tools: The Platform may offer integrated development tools and environment for coding, testing, and deploying applications to streamline the development cost and effort, thereby simplifying the workflow for developers to execute their code close to the data origin, enhancing efficiency.• Advanced visualization and dashboarding capabilities: The Platform may provide sophisticated visualization tools and customizable dashboards to enhance data interpretation and decision-making, thereby improving the quality of insights and decisions;• Multi-cloud and hybrid cloud support: The Platform may support deployment across multiple cloud providers and hybrid cloud environments to increase flexibility and resilience, thereby enhancing versatility and reliability while reducing the vendor lock-in;• Edge and regional computing integration: The Platform may integrate edge and regional computing capabilities to process data closer to the source to reduce latency and bandwidth usage, thereby supporting near-real-time applications for which immediate insights are strongly desired; and• Compliance and governance tools: The Platform may provide built-in tools to ensure compliance with data governance policies and regulatory requirements to help organizations adhere to legal and regulatory standards, thereby reducing the risk of non-compliance.

[0083] The present disclosure may be applicable to the broader security context, which conventionally relies on vast quantities of data collected from ground-based radarstations and overhead satellites, e.g., optical imagery, geospatial information, and radar data. Processing and analyzing such massive data sets may be time-consuming and resource-intensive, often leading to delays in threat detection and hindering timely response efforts. The present disclosure relates to a paradigm shift in data processing through a distributed architecture with a focus on regional computing, including preprocessing data at the point of acquisition, leveraging machine learning (“ML”) algorithms specifically designed for object detection and change detection within the context of border security.

[0084] The present disclosure may be applicable in a wide variety of contexts and scenarios, e.g., for monitoring deforestation, disaster response, space domain awareness applications.

[0085] The present disclosure may include deploying only software tools considered necessary for pre-processing and further deploying other software tools during processing only. Only where value is identified or pre-processing complete may remaining data be transferred for further processing.

[0086] The improved platform described herein may advantageously streamline the process of managing and analyzing large datasets by offering user-friendly interfaces and advanced computational capabilities. By abstracting the complexities of data storage and processing, the following systems and methods may advantageously enable users to focus on deriving insights rather than managing infrastructure. Enhanced reproducibility may ensure that scientific analyses may be consistently replicated, while scalability may allow the system to efficiently handle growing volumes of data. Interoperability with external tools and adherence to industry standards may advantageously enhance integration and flexibility, making it easier to incorporate diverse data sources and analytical methods.

[0087] By addressing these issues through the development and implementation of a comprehensive solution, the cloud-enabled data processing platform of the present disclosure may advantageously overcome the limitations of traditional data process platforms.

[0088] The present disclosure includes a data acquisition module configured to receive raw input data concurrently from various data sources (including satellites) and integrate and standardize data from disparate sources into a united format. The present disclosure further includes a data ingestion module configured to pre-filter the pre- processed input data and automatically perform preliminary quality and format checks to ensure compliance with data requirements. The present disclosure further includes a data transferring module configured to transfer the pre-processed input data to edge computing units within the cloud infrastructure and optimize data transfer times by dynamically selecting the most efficient paths based on real-time network conditions. The present disclosure further includes a feature extraction module configured to perform geospatial encoding and enrichment on the pre-processed input data, including SAR interferometry and / or interferometric SAR (InSAR), utilizing advanced noise reduction and radiometric correction techniques to produce georeferenced features with high precision and accuracy and extract multidimensional landscape features, including georeferenced data. The present disclosure further includes an ML pre-processing module configured to apply supervised and unsupervised ML algorithms, including convolutional neural networks (CNNs) and K-Means Clustering, analyze the georeferenced data to detect anomalies, potential threats, and geospatial changes in near-real-time, and flag relevant data segments for further analysis, where for object detection the output includes identified objects with bounding boxes, classification labels, and confidence scores, and for change detection the output includes annotated bounding boxes highlighting areas of significant change with timestamps and change vectors. The present disclosure further includes a data archiving storage module configured to store raw input data, pre-processed input data, and multidimensional georeferenced features according to lifecycle management rules. The present disclosure further includes a data indexing and cataloging module configured to integrate with data catalog services to create metadata for archived data and facilitate efficient search and retrieval based on metadata. The present disclosure further includes a prioritized data transmission module configured to transmit prioritized data to end users based on relevance, urgency, and user preferences and further based on a dynamically updated hierarchy of importance, incorporating machine learning insights to refine prioritization criteria. Relevance iscontinuously assessed on how pertinent the data is to the user's current needs or context using adaptive algorithms. Urgency is automatically determined for the immediacy desired for the data to be actionable based on predefined thresholds and real-time conditions. User preferences incorporate and update the specific preferences and desires of the end user through a user-friendly interface. The present disclosure further includes data archiving storage for storing raw input data corresponding to the pre-processed input data and the multidimensional landscape features according to lifecycle management rules. The present disclosure further includes a data indexing and cataloguing module for integrating with data catalog services to create metadata for the archived data stored in the data archiving storage.

[0089] Referring now to Figure 1 , shown therein is a system for cloud-enabled data processing, according to an embodiment.

[0090] The system 10 includes a user interface device 12 for providing a user interface (“Ul”) which communicates with a logic device 14, a data stream storage device 16, a data processing device 18, and a server 22 via a network 20. Throughout the present disclosure, the terms “server”, “server platform”, and “platform” may be used interchangeably except where context indicates otherwise. The foregoing discussion of the “platform” may be applicable to the server 22 and / or to the entire system 10.

[0091] The server platform 22 may be a purpose-built machine designed specifically for cloud-enabled data processing of earth observation data. The server platform 22 may communicate directly with the Ul device 12 and / or mediate communications between the Ul device 12 and any one or more of the logic device 14, the data stream storage device 16, and the data processing device 18.

[0092] The system 10 may include a plurality of the data stream storage devices 16.

[0093] The data stream storage device 16 is configured to receive or store data (e.g., represented as data streams). The device 16 may receive or access data streams according to commands from the Ul device 12. The device 16 provides or transmits the data streams to the server 22, which further provides or transmits the data streams to the Ul device 12.

[0094] The device 16 further provides or transmits the data streams to the data processing device 18 for pre-processing. Specifically, the data processing device 18 is configured to perform pre-processing on the data streams. The data processing device transmits such preprocessed data to the server 22.

[0095] For example, the device 18 may analyze incoming data from sensors and satellites (e.g., as received at the device 14 and transmitted through the device 16 as data streams to the device 18) and use pre-processing to identify areas with rapidly rising water levels in the context of flood monitoring, which may allow for earlier and faster detection (compared to conventionally processing all incoming data). Leveraging realtime processing in the foregoing may advantageously ensure that critical information is available as soon as possible to enable timely decision-making for disaster response teams. The foregoing pre-processing functionality may advantageously minimize data transfer and processing with respect to irrelevant information to help reduce costs associated with disaster response activities.

[0096] The data processing device 18 may include a plurality of data processing devices, and such pre-processing may be performed across the plurality of data processing devices through distributed computing, specifically regional computing. Such distributed computing, specifically regional computing, may further be performed across the devices 14, 16, 18, each of which may include a plurality of such devices.

[0097] In embodiments, the pre-processing includes leveraging ML algorithms specifically designed for object detection and change detection. For example, object detection algorithms or change detection algorithms may include convolutional neural networks (“CNNs”) trained on labelled datasets (e.g., in the border security context, including known threats) to identify relevant data (e.g., potential threats such as vehicles, vessels, or unusual heat signatures) within incoming synthetic aperture radar (“SAR”) and optical data streams (e.g., as collected and provided by the data stream storage device 16).

[0098] Data prioritization may be improved by performing some or all of the foregoing pre-processing at the edge.

[0099] Through the foregoing pre-processing functionality, the most relevant portions of data (e.g., imagery including detected objects or areas with anomalous changes), are transferred to the server 22, to the III device 12, or otherwise to a central processing facility or cloud environment for further analysis, e.g., by human experts who review the prioritized data and leverage additional processing power. The foregoing may significantly reduce data transfer costs and minimize the computational burden placed on centralized processing resources (e.g., the server 22).

[0100] Each device or component performing some or all of the pre-processing functionality (e.g., the devices 14, 16, 18) and / or the server 22 includes a push mechanism (e.g., a dedicated module configured to perform push functionality) for sending highly relevant information to users (e.g., imagery including detected objects or areas with anomalous changes) for further analysis by a human expert, which may occur directly (through sending such data to the III device 12) or indirectly (through continuing to transmit such data through the system 10 as herein described, e.g., to the server 22, which may further transmit such data to the human expert at the III device 12).

[0101] Such prioritized data transfer may advantageously reduce data transfer costs and minimize the burden on centralized processing resources.

[0102] The push mechanism may be provided by or may include one or more satellites continuously monitoring for potential wildfires using sensors that detect heat signatures and other signs of fire. When a potential wildfire is identified, relevant data (e.g., location, initial size, thermal signature) is actively “pushed” from the satellite directly to the server 22, e.g., using a messaging system such as Apache Kafka topics (which are dedicated channels for specific types of messages). Because of such specificity, subscribers (such as user applications or wildfire monitoring systems, can “listen” to specific Kafka topics for incoming wildfire alerts, thereby allowing the satellite to publish the initial wildfire detection event.

[0103] The push mechanism employs Apache Kafka to facilitate instantaneous, proactive data transmissions. For example, critical alerts (e.g., social media notifications) such as wildfire notifications are instantly pushed to users when specific changes in data are detected, thereby ensuring timely updates. The present functionality represents aparticu lar improvement in the wildfire context, as wildfires in particular spread rapidly such that even slight delays in data transmission (or slight delays in such data transmissions being received or, in conventional reactive systems, first requested) may translate to catastrophic loss of life or other property damage.

[0104] The system 10 further enables on-demand retrieval (“pull” functionality provided by a “pull mechanism” that together with the push mechanism represents a “push-pull” mechanism) of more detailed data (for example, facts, parameters, or statistics about the wildfire in respect of which the wildfire notification was pushed as a critical alert), further enhancing the speed and relevance of the data provided. The pull mechanism may include a device (e.g., the III device 12) initiating a targeted request to the satellite for specific data after receiving the wildfire alert through the push mechanism (e.g., through the Kafka topic). Such specific data may include high-resolution imagery of the specific location to confirm the fire and assess size and intensity thereof.

[0105] The push-pull mechanism may advantageously accelerate response times and further optimize resource usage by eliminating the conventional need for continuous data polling, which represents a substantial advancement in both efficiency and functionality in real-time data handling and distribution.

[0106] By combining the proactive push of initial data with the targeted pull for details, the platform achieves sub-second delivery compared to conventional methods, thereby providing for faster delivery, as conventional methods require the system to constantly check with the satellite (polling) for updates, which introduces delays in such conventional systems.

[0107] The push mechanism advantageously eliminates unnecessary polling, reducing the overall workload on both the system (e.g., on the server 22) and the satellites. This translates to better resource utilization and improved efficiency.

[0108] The push-pull mechanism provides a timelier response, as near real-time delivery of initial wildfire data allows for faster response times, and emergency services may be alerted more quickly, potentially leading to faster containment and minimizing damage.

[0109] The foregoing data prioritization functionality further advantageously prioritizes real-time analysis at the edge (“real-time processing”), allowing for near- instantaneous identification of potential threats. The foregoing pre-processing functionality permits analysis to begin nearly immediately.

[0110] The system 10, applied in the border security context, may trigger alerts and notifications only for confirmed threats in order to minimize false positive and ensure that centralized processing resources and / or human expertise are focused most efficiently on critical situations. The foregoing may advantageously significantly improve the efficiency of human intervention in border security operations.

[0111] The pre-processing includes leveraging change detection algorithms, such as K-means clustering, to analyze geospatial data to identify anomalies in activity patterns or deviations from baseline readings, which may potentially indicate suspicious activity (and which may be further analyzed at or by the server 22 when the pre-processed data is provided thereto).

[0112] The logic device 14 may sense, record, or otherwise track data and provide such data to the data stream storage device 16.

[0113] The logic device 14 may receive such data from another source, e.g., as specified by or uploaded by the III device 12.

[0114] The logic device 14 organizes, pre-processes, or otherwise logically arranges such data as the data streams received by, stored by, and further provided or transmitted by the device 16.

[0115] The logic device 14 may further pre-process such data before providing the data to the data stream storage device 16, thereby reducing amounts of data transmitted (e.g., by preprocessing the data to send a relevant subset of the data or the output of calculations).

[0116] The server 22 mediates the foregoing data interactions by communicating with the III device 12 and / or autonomously (e.g., in the absence of a specific command from the III device 12, the server 22 performs autonomous mediation of the devices 14, 16, 18).

[0117] The server 22 performs further processing on the data streams received from the device 16. The server 22 performs further processing on the preprocessed data received from the device 18.

[0118] The system 10 may advantageously reduce data acquisition costs by minimizing data transfer through pre-processing data at the source (e.g., at the logic device 14). Such a reduction in data acquisition costs may further include a reduction in an amount of data subject to licensing fees, particularly for large datasets where only a portion thereof may be ultimately used.

[0119] The system 10 may further advantageously optimize computational costs by leveraging cloud-based processing and distributed architecture to improve efficiency and lead to reduced computational costs compared to local processing or the temporary use of high-power computing resources. For example, each of the devices 12, 14, 16, 18 and / or the server 22 may be an abstraction implemented on the cloud across a plurality of machines located at different physical locations (or of virtual machines that are themselves abstractions implemented on the cloud across a plurality of machines located at different physical locations, etc.).

[0120] By processing only relevant portions of the data (e.g., processing at the server 22 only the portions of data or data streams received from one or more of the devices 14, 16, 18 after such devices have performed pre-processing thereon), the system 10 minimizes storage costs through reducing the amount of information stored, both locally (e.g., at a particular user device) and on the cloud.

[0121] By pre-processing data at the source (e.g., at the logic device 14) and transferring only desired or needed data, the system 10 further minimizes data transfer requirements.

[0122] The foregoing real-time processing may, for example, enable faster response times from border security agencies. The foregoing real-time processing may allow durable and reliable data to be replicated across regions, for example in governance and security applications (leveraging established principles associated with Infrastructure as Code (laC) and adhering to recognized security standards to ensure the securedeployment, processing, and transfer of Earth Observation (EO) data for border security applications.

[0123] Data provided to the system 10 (e.g., data recorded at the logic device 14, data provided as input via the III device 12) may include EO data. Such EO data may include data captured according to SAR, which provides high-resolution images of the Earth’s surface independent of weather conditions. Such SAR data may be particularly advantageous for border security applications to see through camouflage or at night. Such EO data may further include multispectral imagery data, which captures information from multiple bands of the electromagnetic spectrum (beyond merely visible light) and may be used to identify different types of vegetation, soil moisture content, and potential camouflage materials. Such EO data may further include hyperspectral imagery, which is an advanced form of multispectral imagery with even more spectral bands, allowing for very detailed analysis of materials and their properties, and which may be used for threat detection by identifying specific signatures of explosive or other contraband materials, and may include non-EO based data such as LiDAR and / or ground-based AIS.

[0124] Any one or more of the devices 12, 14, 16, 18, and the server 22 may include machine learning functionality for real-time processing. For example, preprocessing at any one or more of the devices 14, 16, and / or 18 and / or processing at the server 22 may use a combination of supervised and unsupervised learning algorithms, e.g., to analyze Earth observation data at the source (e.g., at the device 14), to identify valuable information therein (e.g., identifying a subset of such data or pre-processing such data to generate the valuable information as output), and minimize data transfer and processing requirements (through providing only the valuable information to the server 22 for further processing, thereby transferring less data to the server 22 and processing less data at the server 22). Such machine learning may include reinforcement learning.

[0125] The III device 12, devices 14, 16, 18 and server 22 may be a server computer, desktop computer, notebook computer, tablet, PDA, smartphone, or another computing device. The III device 12, devices 14, 16, 18 and server 22 may include a connection with the network 20 such as a wired or wireless connection to the Internet. In some cases, the network 20 may include other types of computer or telecommunicationnetworks. The III device 12, devices 14, 16, 18 and server 22 may include one or more of a memory, a secondary storage device, a processor, an input device, a display device, and an output device. Memory may include random access memory (RAM) or similar types of memory. Also, memory may store one or more applications for execution by processor. Applications may correspond with software modules comprising computer executable instructions to perform processing for the functions described below. Secondary storage device may include a hard disk drive, floppy disk drive, CD drive, DVD drive, Blu-ray drive, or other types of non-volatile data storage. Processor may execute applications, computer readable instructions or programs. The applications, computer readable instructions or programs may be stored in memory or in secondary storage or may be received from the Internet or other network 20. Input device may include any device for entering information into III device 12, devices 14, 16, 18 and server 22. For example, input device may be a keyboard, keypad, cursor-control device, touchscreen, camera, or microphone. Display device may include any type of device for presenting visual information. For example, display device may be a computer monitor, a flat-screen display, a projector or a display panel. Output device may include any type of device for presenting a hard copy of information, such as a printer for example. Output device may also include other types of output devices such as speakers, for example. In some cases, the III device 12, devices 14, 16, 18 and server 22 may include multiple of any one or more of processors, applications, software modules, second storage devices, network connections, input devices, output devices, and display devices.

[0126] Although the III device 12, devices 14, 16, 18 and server 22 are described with various components, one skilled in the art will appreciate that the III 12, devices 14, 16, 18 and server 22 may in some cases contain fewer, additional or different components. In addition, although aspects of an implementation of the III 12, devices 14, 16, 18 and server 22 may be described as being stored in memory, one skilled in the art will appreciate that these aspects can also be stored on or read from other types of computer program products or computer-readable media, such as secondary storage devices, including hard disks, floppy disks, CDs, or DVDs; a carrier wave from the Internet or other network; or other forms of RAM or ROM. The computer-readable media mayinclude instructions for controlling the III device 12, devices 14, 16, 18 and server 22 and / or processor to perform a particular method.

[0127] In the description that follows, devices such as the III device 12, devices 14, 16, 18 and server 22 are described performing certain acts. It will be appreciated that any one or more of these devices may perform an act automatically or in response to an interaction by a user of that device. That is, the user of the device may manipulate one or more input devices (e.g. a touchscreen, a mouse, or a button) causing the device to perform the described act. In many cases, this aspect may not be described below, but it will be understood.

[0128] As an example, it is described herein that the III device 12 or devices 14, 16, 18 may send information to the server 22. For example, a user using the III device 12 may manipulate one or more input devices (e.g. a mouse and a keyboard) to interact with the III. Generally, the III device 12 may receive a user interface from the network 20 (e.g. in the form of a webpage). Alternatively, or in addition, a user interface may be stored locally at a device (e.g. a cache of a webpage or a mobile application).

[0129] The server 22 may be configured to receive a plurality of information, from each of the plurality of devices 12, 14, 16, 18. Generally, the information may comprise at least an identifier identifying an administrator, or user. For example, the information may comprise one or more of a username, e-mail address, password, or social media handle.

[0130] In response to receiving information, the server 12 may store the information in a storage database. The storage database may correspond with secondary storage of the device 12, 14, 16, or 18. Generally, the storage database may be any suitable storage device such as a hard disk drive, a solid state drive, a memory card, or a disk (e.g. CD, DVD, or Blu-ray etc.). Also, the storage database may be locally connected with the server 22. In some cases, the storage database may be located remotely from the server 22 and accessible to the server 22 across the network 20, for example. In some cases, the storage database may comprise one or more storage devices located at a networked cloud storage provider.

[0131] Each device 12, 14, 16, 18 may be associated with an account, e.g., a user account. Any suitable mechanism for associating a device 12, 14, 16, 18 with an account is expressly contemplated. In some cases, a device 12, 14, 16, 18 may be associated with an account by sending credentials (e.g. a cookie, login, or password etc.) to the server 22. The server 22 may verify the credentials (e.g. determine that the received password matches a password associated with the account). If a device 12, 14, 16, 18 is associated with an account, the server 22 may consider further acts by that device 12, 14, 16, 18 to be associated with that account.

[0132] A key feature of the present disclosure includes or otherwise relates to regional computing to provide a transformative approach to data processing. Conventional platforms typically transfer all data to a central processing facility and rely solely on human analysis of the entire unprocessed dataset. However, the devices 14, 16, and / or 18 utilize edge computing, whereby ML algorithms perform initial analysis (or “pre-processing”) directly at data acquisition points (e.g., SAR stations, satellites). The pre-processing advantageously significantly reduces data transfer by focusing only on relevant information identified by the ML algorithms, such as potential threats or anomalies.

[0133] In embodiments, the ML functionality of the present disclosure includes convolutional neural networks (CNNs) trained to detect objects such as vehicles, vessels, or unusual heat signatures within incoming data streams (e.g., as provided by the data stream storage device 16 to the server 22), enabling identification of potential threats at the edge.

[0134] In embodiments, the ML functionality further includes K-means clustering for analyzing geospatial data to identify anomalies in activity patterns or deviations from baseline readings, potentially indicating suspicious activity.

[0135] The system 10 may operate based on an end-to-end encryption strategy to meet security standards, such as Fed Ramp, ISO 27001 , and HIPPA.

[0136] When the system 10 or components thereof are at rest, data may be encrypted according to industry-standard algorithms, e.g., AES-256, whether the data is stored on edge devices (e.g., the devices 14, 16, 18 or the server 22 operating accordingto an edge paradigm), regional processing centers (e.g., the devices 14, 16, 18 performing pre-processing), or the III device 12. The foregoing encryption may advantageously provide protection to sensitive information even in the event of a security breach.

[0137] The system 10 enforces rigorous data segregation practices to ensure that classified and unclassified data remain separate throughout the system 10. Such segregation may be achieved through dedicated storage solutions and the implementation of access controls for classified data.

[0138] The system 10 may be deployed automatically with integrated security standards to minimize the risk of human error and further to provide consistent security configurations across deployments. The foregoing may advantageously strengthen the overall security posture of the system 10.

[0139] Compliance with established security frameworks may expedite security audits, particularly for users subject to adherence with specific regulations. The foregoing secure data transfer functionality and data segregation practices may further advantageously reduce the audit scope for user destinations.

[0140] The system 10 leverages DevOps principles to automate and streamline data transfers, thereby enhancing operational efficiency and minimizing the risk of errors during data movement. The application of Continuous Integration / Continuous Deployment (CI / CD) pipelines automates the integration of code changes, testing, and deployment to production environments. The concept of Infrastructure as Code (laC) is utilized to automate and manage infrastructure configurations, ensuring consistent and repeatable deployments. This reduces manual configuration errors and expedites the provisioning of infrastructure required for data transfers. Real-time monitoring and logging of data transfer processes are facilitated through automated monitoring systems, which ensure the prompt identification and resolution of any issues, thereby maintaining the reliability and integrity of data transfers. Alerting systems are integrated to notify the operations team of anomalies or failures in data transfer, enabling swift remediation and minimizing downtime. The integration of security practices into the CI / CD pipeline, known as DevSecOps, ensures compliance with security standards and regulations, therebyreducing the risk of data breaches. Compliance automation tools are employed to ensure that data transfers adhere to organizational and regulatory requirements. Containerization of data transfer applications ensures consistent execution across different environments, thereby enhancing scalability and flexibility. The system provides a comprehensive and robust solution for managing heterogeneous data from diverse formats and multiple sources. This approach significantly enhances the system's capability to handle varied data formats seamlessly, providing a dynamic and efficient data exploration and analysis experience. The system 10 may break down data transfer processes into smaller, manageable microservices to advantageously allow for independent scaling and updating of each service, improving the overall efficiency and reliability.

[0141] The present disclosure may achieve seamless data integration and up-to- date information repository management through a novel architecture for the system 10 that prioritizes security. This architecture leverages a combination of technologies, including:• advanced middleware with transformation capabilities, which acts as an intermediary layer, facilitating communication between the server 22 and the data stream storage device 16 and translating data requests from the server 22 into specific commands understood by each device 16, regardless of underlying protocol or data format. Additionally, the middleware may be configured to perform data transformation tasks, such as cleansing or standardizing data formats, to ensure compatibility with an internal repository of the server 22, which eliminates the need for custom connectors for each data source, ensuring broader compatibility.• RESTful APIs with authentication, which function as a set of standardized instructions that define how the server 22 interacts with each data stream storage device 16. The middleware utilizes these APIs to translate the requests of the server 22 into the specific language understood by a target device 16, enabling efficient data retrieval. All APIs leverage OAuth2 authentication tokens, enabling user segregation and access control based on specific roles.• Apache Kafka connectors with fault tolerance, which act as specialized adapters that streamline the connection process for specific data sources with unique protocols (like cloud storage or databases) and efficiently bridge the gap between the server 22 and these sources, ensuring seamless data transfer. These connectors are also designed with fault tolerance in mind, ensuring data transfer continues even if there are temporary disruptions with a specific data source.• enhanced security through robust protocols, whereby the server 22 prioritizes data security throughout the data integration process. To ensure the confidentiality and integrity of information during transfer, the system 10 utilizes RSA encryption, a robust industry-standard algorithm. This encrypts all data in transit between the server 22 and the data stream storage devices 16. Additionally, a load balancer acts as a security gateway, encrypting all data flowing to and from the TKE Kafka brokers, which are the central nervous system for real-time data streaming within the server 22.

[0142] Referring now to Figure 2, shown therein is a block diagram of a computing device 200 of the system 10 of Figure 1 , according to an embodiment. The computing device 200 may be, for example, any one of devices 12, 14, 16, 18 or the server 22 of Figure 1 .

[0143] The computing device 200 includes multiple components such as a processor 202 that controls the operations of the computing device 200. Communication functions, including data communications, voice communications, or both may be performed through a communication subsystem 204. Data received by the computing device 200 may be decompressed and decrypted by a decoder 206. The communication subsystem 204 may receive messages from and send messages to a wireless network 250.

[0144] The wireless network 250 may be any type of wireless network, including, but not limited to, data-centric wireless networks, voice-centric wireless networks, and dual-mode networks that support both voice and data communications.

[0145] The computing device 200 may be a battery-powered device and as shown includes a battery interface 242 for receiving one or more rechargeable batteries 244.

[0146] The processor 202 also interacts with additional subsystems such as a Random Access Memory (RAM) 208, a flash memory 210, a display 212 (e.g., with a touch-sensitive overlay 214 connected to an electronic controller 216 that together comprise a touch-sensitive display 218), an actuator assembly 220, one or more optional force sensors 222, an auxiliary input / output (I / O) subsystem 224, a data port 226, a speaker 228, a microphone 230, short-range communications systems 232 and other device subsystems 234.

[0147] In some embodiments, user-interaction with the graphical user interface may be performed through the touch-sensitive overlay 214. The processor 202 may interact with the touch-sensitive overlay 214 via the electronic controller 216. Information, such as text, characters, symbols, images, icons, and other items that may be displayed or rendered on a computing device generated by the processor 202 may be displayed on the touch-sensitive display 212.

[0148] The processor 202 may also interact with an accelerometer 236 as shown in Figure 2. The accelerometer 236 may be utilized for detecting direction of gravitational forces or gravity-induced reaction forces.

[0149] To identify a subscriber for network access according to the present embodiment, the computing device 200 may use a Subscriber Identity Module or a Removable User Identity Module (SIM / RUIM) card 238 inserted into a SIM / RUIM interface 240 for communication with a network (such as the wireless network 250). Alternatively, user identification information may be programmed into the flash memory 210 or performed using other techniques.

[0150] The computing device 200 also includes an operating system 246 and software components 248 that are executed by the processor 202 and which may be stored in a persistent data storage device such as the flash memory 210. Additional applications may be loaded onto the computing device 200 through the wireless network 250, the auxiliary I / O subsystem 224, the data port 226, the short-range communications subsystem 232, or any other suitable device subsystem 234.

[0151] In use, a received signal such as a text message, an e-mail message, web page download, or other data may be processed by the communication subsystem 204and input to the processor 202. The processor 202 then processes the received signal for output to the display 212 or alternatively to the auxiliary I / O subsystem 224. A subscriber may also compose data items, such as e-mail messages, for example, which may be transmitted over the wireless network 250 through the communication subsystem 204.

[0152] For voice communications, the overall operation of the computing device 200 may be similar. The speaker 228 may output audible information converted from electrical signals, and the microphone 230 may convert audible information into electrical signals for processing.

[0153] Referring now to Figure 3, shown therein is a computer system 300 for cloud-enabled data processing, according to an embodiment. The computer system 300 may be or may implement any one or more of the devices 12, 14, 16, 18 and / or the server 22 of Figure 1 .

[0154] The system 300 includes a processor 302 for executing software models and modules.

[0155] The system 300 further includes a memory 304 for storing data, including output data from the processor 302.

[0156] The system 300 further includes a communication interface 306 for communicating with other devices, such as through receiving and sending data via a network connection (e.g., the network 20 of Figure 1 ).

[0157] The system 300 further includes a display 308 for displaying various data generated by the computer system 300 in human-readable format.

[0158] The processor 302 includes a data acquisition module 310 for receiving input data 312. The input data 312 is stored in the memory 304.

[0159] The memory 304 may represent one or more memories of one or more computers arranged according to a distributed computing paradigm, that is, physically separate to one another and cooperating or sharing resources (viz., their memories) over a network (e.g., the network 20 of Figure 1 ).

[0160] Accordingly, the input data 312 may be stored in a predetermined cloudbased computational resource or a regional blob storage facility near a ground station, which is considered part of the memory 304.

[0161] The input data 312 may be sensor data captured or otherwise generated by satellites at regular intervals. The input data 312 may include SAR data.

[0162] The storage facility may be managed by cloud storage providers, e.g., Amazon Web Services, Google, or K-SAT Ground Station.

[0163] The processor 302 further includes a data ingestion module 314.

[0164] The data ingestion module 314 ingests the input data 312 and filters the input data 312 to automatically perform preliminary quality and format checks, ensuring compliance with requirements of the system 300 before further processing.

[0165] The processor 302 may represent one or more processors of one or more computers arranged according to the foregoing distributed computing paradigm such that physically separate computers cooperate or share their processors over a network. The processor 302 may be understood to represent an abstraction of such shared processing functionality. Accordingly, the ingesting and pre-filtering of the input data 312 may take place at a specific device whose processor forms part of the processor 302 (e.g., any one or more of the devices 14, 16, 18 of the system 10 of Figure 1 ).

[0166] The processor 302 further includes a data transferring module 316.

[0167] The data transferring module 316 transfers the input data 312 or otherwise relocates the input data 312 to edge computing units within the cloud infrastructure positioned to optimize data transfer times (e.g., to specific processing units within the device 16 or the device 18 of Figure 1 ). The data transferring module 316 performs the foregoing functionality according to stringent security classifications, best practices, and access controls, which are meticulously enforced.

[0168] The processor 302 further includes an extraction module 318. The extraction module 318 performs geocoding and enrichment on the input data 312.

[0169] The input data 312 is initially two-dimensional. The extraction module 318 performs enrichment techniques on the input data 312 to extract multidimensionallandscape features 320 from the input data 312 through techniques including but not limited to SAR interferometry and / or InSAR.

[0170] SAR image processing may include Surface Deformation and Change (SDC), Digital Elevation Modeling (DEM), and Change Detection (CD).

[0171] The foregoing relate to methods and systems for geocoding and enriching two-dimensional synthetic aperture radar (SAR) data to extract multidimensional landscape features. The techniques employed include SAR interferometry and interferometric synthetic aperture radar (InSAR), which enable the generation of DEMs and the extraction of temporal surface deformation and change (SDC) information.

[0172] The extraction module 318 described herein performs geocoding and enrichment on multi-dimensional input data using SAR interferometry and InSAR to extract multidimensional landscape features. Additionally, the extraction module 318 integrates data fusion and machine learning algorithms for feature extraction and multidimensional analysis.

[0173] The multidimensional landscape features 320 are stored in the memory 304.

[0174] The extraction module 318 further conducts noise reduction and radiometric correction and assigns geographic coordinates to each data point within the multidimensional landscape features 320. This enhances the multidimensional landscape features 320 with external sources, such as digital elevation models, for improved anomaly detection. The multidimensional landscape features 320 include georeferenced data produced according to the foregoing.

[0175] The processor 302 further includes an ML pre-processing module 322.

[0176] The ML pre-processing module 322 performs regional preprocessing using ML.

[0177] The ML pre-processing module 322 may be specific to one or more devices or may be an abstraction of some or all processing or pre-processing functionality of one or more devices connected by a network (e.g., one or more of the devices 14, 16, 18 connected by the network 20 of Figure 1 ).

[0178] The ML pre-processing module 322 applies one or more supervised or unsupervised ML algorithms to analyze the georeferenced data of the multidimensional landscape features 320 to detect anomalies and potential threats in near-real-time. This may significantly reduce data transfer requirements by flagging only relevant data segments for further analysis. Supervised ML algorithms may include CNNs. Unsupervised ML algorithms may include K-Means Clustering.

[0179] The regional pre-processing using machine learning (ML) is performed by applying advanced machine learning algorithms to analyze the georeferenced data to detect anomalies, potential threats, and geospatial changes in near-real-time using virtually unlimited distributed computing power. The ML pre-processing flags only relevant data segments for further analysis, which includes identifying and marking specific portions of the data that exhibit particular characteristics or meet predefined criteria, thus rendering such specific portions of the data significant or noteworthy for subsequent detailed analysis. For vessel detection on the surface of the water, the input to the system 300 includes high-resolution satellite imagery, such as Synthetic Aperture Radar (SAR) data, along with hyperspectral analysis and radar signals. The supervised ML algorithm, such as Convolutional Neural Networks (CNNs), processes this imagery to identify features corresponding to known objects, such as vessels. The regional pre-processing classifies these features and identifies vessels within the imagery, outputting bounding boxes around detected vessels, along with classification labels and confidence scores. For example, if the system receives SAR data capturing a maritime area, the CNN may detect and classify various vessels (e.g., cargo ships, fishing boats) on the water's surface. The output includes annotated images where each detected vessel is enclosed in a bounding box, labeled with the type of vessel and a confidence score indicating the certainty of the classification. The flagged data segments include the original satellite images with these annotations. The flagged data is then prioritized for further analysis, ensuring that only the most pertinent information is forwarded for more detailed examination, thereby reducing the volume of data to which further processing is applied. The system integrates SAR, hyperspectral analysis, and radar signals at the edge of the network, utilizing finely tuned software that performs this processing multiple times faster than transferring terabytes of data into a centralized location for subsequent processing.By conducting the foregoing analysis at the edge, the time from data acquisition to actionable insight is significantly reduced. Monitoring an expansive area in real time, such as an ocean, presents significant challenges, including data transfer costs, network bottlenecks, and the computational burden of processing vast amounts of data. Traditional methods that include transferring terabytes of raw data to centralized facilities for processing are inefficient and costly. By deploying machine learning logic to the edge, the system 300 may process data in real time, focusing only on the relevant portions thereof. For example, in the South China Sea, which spans approximately 3.5 million square kilometers, the traditional approach would include capturing and transferring large datasets to detect and monitor vessel movements. However, by applying ML algorithms at the edge, only data segments indicating potential vessel activity are processed and flagged for further analysis. This targeted approach significantly reduces the amount of data to be transferred, stored, and processed centrally. Accordingly, real-time edge processing may identify and flag only 1 % of the total data as relevant, which translates to a reduction in data transfer volume by 99%. Such reduction not only alleviates network congestion but further minimizes the costs associated with data storage and computational resources. Furthermore, the reduced latency in processing and analyzing relevant data enhances the responsiveness and effectiveness of maritime security operations. The intelligence embedded within the ML computing ensures that only the most critical information is processed and analyzed, thereby streamlining operations and reducing costs.

[0180] The memory 304 further includes data archiving storage 324 for storing all raw SAR-related data (e.g., the input data 312) and derived metadata (e.g., the multidimensional landscape features 320 including the georeferenced data).

[0181] The data archiving storage 324 employs a cost-effective, regulation- compliant storage solution with lifecycle management rules that automatically transition data to more economical storage options, balancing cost and accessibility effectively. This solution is optimized for data archiving and long-term backup, providing secure, durable, and low-cost storage with options for different retrieval times, thereby ensuring efficient storage of less frequently accessed data while maintaining accessibility when needed. The data archiving storage 324 is structured to adhere to cost-efficiency and regulatorycompliance through advanced lifecycle management rules. These rules facilitate the automated transition of data to economical storage options over time, ensuring a balance between cost and accessibility. The system 10 utilizes a dual-layer architecture comprising separate computing and storage layers that operate in tandem to enhance data processing capabilities. Unlike traditional single-system storage management, this architecture employs an event-driven approach with a multi-tiered storage layer and embedded compute nodes. This architecture advantageously allows for simultaneous distribution of data into active and cold-storage tiers. Accordingly, the active storage tier handles frequently accessed data, providing high performance and immediate accessibility, while the cold-storage tier is optimized for long-term data archiving, contributing to security and cost-effectiveness. This distributed architecture ensures that data transitions from active to cold storage seamlessly, based on predefined lifecycle management rules, without requiring manual intervention. These protocols automate the migration process, optimizing storage costs while maintaining compliance with regulatory standards. The data archiving storage 324 supports scalability and durability, offering a comprehensive solution for robust data management. The data archiving storage 324 integrates advanced lifecycle management protocols that automate data transitions between different storage tiers, thus ensuring that data is stored in the most appropriate and cost-effective tier throughout its lifecycle. The system 10 is designed to comply with various regulatory requirements, ensuring data protection and accessibility in a secure and efficient manner. By leveraging a distributed architecture and event-driven design, the system 10 enhances overall data management performance and cost-efficiency, providing a practical solution for modem data archiving and long-term storage needs. The multi-tiered storage 324, with distinct active and cold-storage layers, distributes storage and processing responsibilities across a distributed network, enhancing data accessibility and compliance while optimizing costs. This approach ensures efficient data management and aligns with the dynamic requirements of contemporary enterprises, providing a reliable and economical solution for long-term data retention.

[0182] The data archiving storage 324 stores data into a two-layered architecture including respective computing and storage layers, where both layers cohesively worktogether to enhance the data processing capability. Unlike traditional systems where storage is managed by a single storage system, this architecture leverages a distributed approach known as event-driven architecture. Such event-driven architecture includes a multi-tiered storage layer embedding a compute node to distribute data into active and cold storage in parallel.

[0183] The data archiving storage 324 may correspond to be or may be implemented by one or more of the devices 14, 16, 18 of Figure 1 , particularly the data storage device 16 of Figure 1 , or by multiple such devices 14, 16, 18, particularly multiple such data storage devices 16, according to the distributed computing paradigm hereinabove discussed.

[0184] The processor 302 further includes a data indexing and cataloguing module 326.

[0185] The data indexing and cataloguing module 326 integrates with data catalog services to create metadata 328 for the archived data stored in the data archiving storage 324.

[0186] The created metadata 328 includes a searchable index 330. The searchable index may facilitate easy data retrieval for future reference or analysis.

[0187] Traditional data lakes face significant challenges in scalability and flexibility when dealing with high-velocity, high-volume, and diverse data sources. Centralized storage systems, such as relational databases or file systems, often become bottlenecks, leading to inefficiencies in data processing and management. The present disclosure addresses these limitations by providing a distributed architecture that leverages event- driven mechanisms and scalable storage solutions. The present disclosure describes a distributed data lake ingestion and storage system configured to efficiently handle high- velocity data streams and ensure data durability. This system utilizes a two-layer middleware architecture that integrates computing and storage functionalities, enabling seamless data ingestion, storage, and archiving.

[0188] The data ingestion process implemented by the foregoing system is managed by a network of brokers using event-driven architecture principles, employingstreaming technologies such as but not limited to Apache Kafka. These brokers ingest millions of messages per second, ensuring high throughput and low latency. The event- driven nature of the system enables efficient real-time data processing, making the system suitable for applications requiring rapid data ingestion and analysis.

[0189] The system implements a data replication strategy to ensure data durability and availability. Data ingested by the brokers is replicated across multiple nodes in different regions, providing fault tolerance and disaster recovery capabilities. The system incorporates an adjustable data retention policy, allowing efficient management of active and archival data. Active data may typically be retained in the high-velocity storage layer for an adjustable period ranging from 7 days to 2 weeks.

[0190] After the retention period, data is automatically transitioned to long-term storage solutions such as cloud-based blob storage. This transition is managed by streaming technologies such as but not limited to Apache Kafka Connectors, which replicate data from the high-velocity layer to the archival storage according to desired frequency of access. The archival storage strategy adheres to best practices, including multi-cloud support, allowing data to be moved to different storage tiers based on cost and retrieval speed requirements.

[0191] The distributed architecture of the system allows for easy scalability. New nodes may be added to the broker network to increase ingestion capacity. Storage resources may be scaled horizontally to accommodate growing data volumes. The flexibility of the system advantageously renders the system suitable for various applications, from real-time analytics to long-term data archiving.

[0192] The architecture supports a multi-tiered storage strategy that aligns with best practices for data lifecycle management. Data initially stored in high-performance storage such as but not limited to AWS S3 may be transitioned to lower-cost storage options such as S3 Glacier after a predefined period, based on user-defined policies. This hierarchical storage model optimizes storage costs and access speeds according to data usage patterns.

[0193] The middleware includes Al-driven management tools that continuously monitor data ingestion, storage, and access patterns. These tools optimize resourceallocation, adjust retention policies dynamically, and ensure compliance with data governance requirements. Machine learning algorithms predict data usage trends and preemptively manage storage needs. For instance, the Al system may dynamically adjust data retention periods based on real-time usage patterns, ensuring that frequently accessed data may advantageously remain readily available while less critical data may be archived.

[0194] The created metadata 328 preferably includes a searchable index that facilitates easy data retrieval for future reference or analysis. This searchable index advantageously helps maintain data accessibility and usability across the data lake. By indexing metadata, the system allows users to quickly locate and retrieve specific datasets, even as the volume of stored data grows. The index is created using technologies such as Elasticsearch or Apache Sole, which support full-text search capabilities and can handle large volumes of data efficiently.

[0195] Traditional data storage techniques include the "schema-on-write" procedure, which necessitates formatting and constructing the data in accordance with a predetermined structure before storing it. This requirement may disadvantageously lead to significant delays in data processing, especially when dealing with large volumes of diverse data. Furthermore, this requirement imposes rigid constraints on data ingestion, limiting the flexibility to accommodate varying data types and formats. The "schema-on- read" approach adopted by data lakes according to the present disclosure advantageously allows for writing raw data first and determining the schema later during data retrieval and analysis. Accordingly, extra flexibility is provided by permitting different schemas for each dataset. Additionally, technologies such as AWS Glue and Apache NiFi may advantageously facilitate the automated discovery, cataloging, and schema inference of data, enabling efficient management of structured and unstructured data. This approach significantly reduces the time and complexity involved in data processing, enhancing the overall efficiency and adaptability of the data lake architecture.

[0196] The architecture described in the present disclosure provides several advantages over traditional data lake systems, including enhanced scalability, improved fault tolerance, and better management of diverse data types. The distributed, event-driven approach ensures high performance and efficiency, while advanced security measures and data lifecycle management policies maintain data integrity and reduce costs.

[0197] The processor 302 further includes a prioritized data transmission module 332.

[0198] The prioritized data transmission module 332 transmits prioritized data 334 to end users (e.g., via the display 308). End users may include, for example, border security agencies. Such transmission may further be performed to or by real-time delivery systems, such as Apache Kafka or the like, to enhance threat response capabilities.

[0199] The system 300 may be further configured to support automated critical threat alerts and promptly notify security personnel through various channels.

[0200] The prioritized data 334 may include the georeferenced data of the multidimensional landscaped features 320.

[0201] The prioritized data 334 may be used by security personnel for detailed analysis and strategic response, including mobilization of patrols to investigate identified threats.

[0202] The computer system 300 offers a vertically integrated solution by combining cloud computing, edge processing, and advanced machine learning. The foregoing may advantageously ensure adherence to stringent security compliance protocols and may further advantageously significantly reduce operational costs related to data processing and storage.

[0203] The computer system 300 may advantageously enhance data transfer efficiency by allowing users to selectively access critical information, thereby conserving resources.

[0204] The dynamic, user-centric approach of the computer system 300 may advantageously support comprehensive data queries and retrieval on demand, facilitating scalability and enabling sophisticated data fusion applications for optimal utility and cost efficiency.

[0205] In an embodiment, the computer system 300 supports users deploying their own code thereon to facilitate on-site data analysis close to the origin of the data. Such functionality may advantageously significantly reduce latency and processing times, thereby enabling tailored and efficient data handling.

[0206] In an embodiment, the computer system 300 further includes robust security measures to ensure data privacy and integrity, even with user-deployed code. Scalability of the foregoing security measures allows users with smaller datasets or limited resources to benefit from flexibility in the security measures. The computer system 300 may further include collaborative features to enable teams to work together on code and analysis, further enhancing the transformative impact of the system 300 on data processing.

[0207] Referring now to Figure 4, shown therein is a flowchart of a method 400 for cloud-enabled data processing, according to an embodiment. The method 400 may be performed by the system 10 of Figure 1 or by the system 300 of Figure 3. The method 400 may be encoded as computer-executable instructions which, when executed by one or more processors, perform the steps of method 400.

[0208] At 402, the method 400 includes receiving or acquiring input data.

[0209] The input data may be stored in cloud memory according to a distributed computing paradigm, that is, across one or more devices physically separate to one another and cooperating or sharing resources (viz., their memories) over a network (e.g., the network 20 of Figure 1 ).

[0210] Accordingly, the input data may be stored in a predetermined cloud-based computational resource or a regional blob storage facility near a ground station, which is considered part of the cloud memory.

[0211] The input data 312 may be SAR data captured or otherwise generated by satellites at regular intervals.

[0212] At 404, the method 400 further includes ingesting the input data and prefiltering the input data to automatically perform preliminary quality and format checks.

[0213] At 406, the method 400 further includes transferring the input data 312 to edge computing units within a cloud infrastructure positioned to optimize data transfer times according to stringent security classifications, best practices, and access controls.

[0214] At 408, the method 400 further includes performing geocoding and enrichment on the input data.

[0215] The input data is initially two-dimensional, and the enrichment techniques performed thereon extract multidimensional landscape features therefrom through techniques such as SAR interferometry and / or InSAR.

[0216] Performing geocoding and enrichment on the input data includes performing noise reduction and radiometric correction and assigning geographic coordinates to each data point within the multidimensional landscape features, thereby enhancing the multidimensional landscape features with external sources, such as digital elevation models, for improved anomaly detection. The multidimensional landscape features include georeferenced data produced according to the foregoing.

[0217] At 410, the method 400 further includes performing regional preprocessing using ML.

[0218] Performing the regional preprocessing includes applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near-real-time. This may significantly reduce data transfer requirements by flagging only relevant data segments for further analysis.

[0219] At 412, the method 400 further includes archiving all raw SAR-related data (e.g., the input data) and derived metadata (e.g., the multidimensional landscape features including the georeferenced data) for long-term storage in a cost-effective, regulation- compliant storage solution with lifecycle management rules automatically transitioning data to cost-efficient storage options, balancing cost and accessibility.

[0220] At 414, the method 400 further includes integrating the foregoing functionality with data catalog services to create metadata for the archived data.

[0221] The metadata includes a searchable index that facilitates easy data retrieval for future reference or analysis.

[0222] At 416, the method 400 further includes transmitting prioritized data to end users.

[0223] The systems, methods, and devices described in the present disclosure may advantageously foster seamless connection and may provide for up-to-date information repository management. The systems, methods, and devices may provide automatic data discovery by leveraging data discovery techniques to identify a wide range of data sources, including both newly created collections and archived data.

[0224] The systems, methods, and devices described in the present disclosure may provide for standardized communication by leveraging standardized APIs, such as RESTful APIs with robust authentication mechanisms (e.g., OAuth 2.0, API keys, or JSON Web Tokens), to interact with various data storage solutions. These APIs work in conjunction with a data catalog, such as Apache Hive or a service such as Amazon Athena, which acts as a centralized registry for all the connected data sources of the system (e.g., the system 10 of Figure 1 ). This data catalog provides users with a comprehensive view of available datasets, including metadata such as schema and location. Integration with such services allows users to leverage familiar SQL queries for data exploration and analysis directly within the data catalog, eliminating the need for complex data movement or manipulation.

[0225] The present disclosure encompasses systems and methods for comprehensive data ingestion, which is highly valuable for managing heterogeneous data from diverse formats and multiple sources. In particular, systems such as but not limited to Apache NiFi facilitate this by supporting scalable ingestion processes, including both batch and stream methods. The concept of schema-on-read, utilized by data catalogs such as Hive and Athena, enables flexible querying of heterogeneous data without the necessity of a predefined schema. This approach significantly enhances the system's capability to handle varied data formats seamlessly, thereby providing a dynamic and efficient data exploration and analysis experience. This data catalog provides users with a comprehensive view of available datasets, including metadata such as schema andlocation. The input into the data catalog includes various datasets from different data sources, which may include structured data (e.g., databases, spreadsheets), semistructured data (e.g., JSON, XML), and unstructured data (e.g., text, images, videos). Each dataset is ingested into the system through standardized APIs and is registered within the data catalog, where metadata such as the schema (defining the structure of the data) and location (indicating where the data is stored) are recorded. The output from the data catalog includes a unified interface that allows users to access and query the datasets. This output interface provides metadata information, enabling users to understand the structure and storage details of the data. Users are able to perform SQL queries directly within the data catalog, leveraging the provided metadata to explore and analyze the datasets without the need for complex data movement or manipulation. This output interface enhances the usability and accessibility of the data, facilitating efficient data exploration and analysis.

[0226] The systems, methods, and devices described in the present disclosure may provide for automatic synchronization and up-to-date repository. The system (such as the system 10 of Figure 1 ) may be configured for automatic synchronization using crawlers, batch operations, and Apache Kafka topics. Such configuration may advantageously ensure that the data catalog and underlying repository remain constantly updated with the latest data from connected sources.

[0227] The systems, methods, and devices described in the present disclosure may provide for an improved data access speed. By seamlessly connecting to diverse data sources, users may access a comprehensive data repository without needing to manage individual connections or navigate different protocols, which may advantageously result in faster data retrieval.

[0228] The systems, methods, and devices described in the present disclosure may provide for flexibility and choice. Users may choose the most appropriate data exploration tool based on their needs. SQL expertise allows leveraging data catalogs like Athena, while KSQL caters to real-time data analysis requirements.

[0229] The systems, methods, and devices described in the present disclosure may provide for faster time to insights, as a seamless connection through middleware andstandardized APIs may advantageously ensure efficient data retrieval, reducing the time taken to access and analyze data.

[0230] The systems, methods, and devices described in the present disclosure may provide for reduced development time, as standardized APIs and pre-built integrations with data catalogs such as Athena minimize custom development efforts for connecting to diverse data sources.

[0231] The systems, methods, and devices described in the present disclosure may provide for reduced management overhead, as seamless connection may advantageously eliminate the conventional need for custom connectors for each data source, reducing the administrative burden for managing individual data source connections.

[0232] The systems, methods, and devices described in the present disclosure provide deployment flexibility and an option for on-premises data storage. The system (such as the system 10 of Figure 1 ) may advantageously cater to diverse security requirements by offering flexible deployment options. Users may opt for a cloud-based deployment with secure facilities, which offers advantages such as scalability, cost efficiency, and ease of access. Alternatively, for organizations with stricter data security regulations, the system supports on-premises data storage, enabling them to maintain complete control over their data within their own secure facilities. This flexibility is particularly advantageous in the context of big data, where vast amounts of sensitive information are to be managed and protected. By accommodating both cloud-based and on-premises storage, the system ensures that organizations may advantageously be able to meet their specific security and compliance needs while efficiently handling large datasets. Additionally, the system is adaptable to both unclassified and classified data environments, ensuring secure handling and compliance with regulatory requirements for sensitive information. The system may advantageously segregate and manage unclassified data in the cloud while maintaining classified data within on-premises facilities or secure cloud environments. Seamless data interactions in a hybrid deployment are facilitated by technologies such as VPC peering, AWS Transit Gateway, or equivalents from other cloud providers (e.g., Azure Virtual Network Peering or GoogleCloud Interconnect). These technologies enable secure and efficient communication between different parts of the hybrid infrastructure, with VPC peering allowing direct network connections between Virtual Private Clouds for low-latency, high-bandwidth data transfer, and AWS Transit Gateway acting as a hub that connects multiple VPCs and onpremises networks, simplifying network architecture and management. This hybrid deployment approach ensures that data may advantageously flow seamlessly between cloud and on-premises environments, providing organizations with the flexibility to leverage the benefits of both. Furthermore, the system incorporates robust encryption and access control mechanisms, enhancing data protection in both deployment scenarios, thus enhancing the capability of the system to handle big data securely and efficiently, meeting the dynamic needs of modem enterprises.

[0233] In an embodiment, the present disclosure is applied in the context of both unclassified and classified systems, e.g., within the context of Amazon Web Services (AWS), e.g., AWS GovCIoud (US) and AWS Secret Region. The system 10 may handle varying security requirements, from handling controlled unclassified data to managing highly classified information. AWS GovCIoud (US) adheres to regulations such as ITAR, FedRAMP High, DoD SRG Levels 2 and 4, making it suitable for sensitive unclassified data. In contrast, AWS Secret Region is tailored for classified data, meeting DoD SRG Levels 5 and 6 and ICD 503 standards. In practical terms, the regional computing framework processes data in a step-by-step manner to ensure robust security at each stage. At a data acquisition stage, input data, such as SAR data from satellites, is stored in AWS cloud protected storage across a distributed computing environment. AWS GovCIoud (US) is used for unclassified data, ensuring compliance with ITAR and FedRAMP High. Classified data is stored in AWS Secret Region, adhering to DoD SRG Level 5 and ICD 503 standards. At an ingestion and pre-filtering stage, preliminary quality and format checks are performed upon data ingestion. AWS services such as AWS Lambda and AWS Glue enforce data integrity and compliance. AWS Key Management Service (KMS) encrypts classified data using FIPS 140-2 validated cryptographic modules. At a transfer to edge computing stage, pre-processed data is transferred to edge computing units within the AWS cloud infrastructure, optimization is performed usingAWS Direct Connect and AWS Transit Gateway, ensuring secure data paths. VPC peering with strict access controls in compliance to ITAR and FedRAMP advantageously maintains compliance for classified data. At a geocoding and enrichment stage, geocoding and enrichment techniques such as SAR interferometry are applied. AWS Elastic Kubernetes Service (EKS) provides a secure, scalable environment for these computations. Data is enriched with external sources stored in Amazon S3, with encryption at rest and in transit. At a regional preprocessing using machine learning (ML) stage, machine learning algorithms, including convolutional neural networks (CNNs) for object detection and K-Means for clustering, detect anomalies and threats. AWS Sage Maker facilitates this processing, employing security measures such as network isolation and role-based access control. Classified data processing follows DoD SRG Level 5 guidelines. At an archiving and cataloging stage, all raw SAR data and derived metadata are archived in Amazon S3 with lifecycle management rules for cost-effective long-term storage. The data is cataloged using AWS Glue Data Catalog for efficient metadata creation and searchability, with encryption and access controls according to ISO / IEC 27001 standards. At a prioritized data transmission stage, prioritized data is transmitted to end users based on relevance, urgency, and user preferences. AWS KMS provides a secure messaging layer for real-time notifications and updates. Machine learning insights dynamically adjust prioritization settings. Each stage is fortified with robust security measures aligned with AWS standards, ensuring compliance with protocols such as FedRAMP, DoD SRG, ICD 503, NIST SP 800-53, ISO / IEC 27001 , NIST SP-800 171 R3, ITSG-33 and SOC certifications.

[0234] The systems, methods, and devices described in the present disclosure may provide for automatically triggering SAR image capture alongside user alerts in specific scenarios, further streamlining the response process.

[0235] Referring now to Figure 5 shown therein is a computer system 500 for cloud-enabled processing of earth observation data, according to an embodiment. The system 500 is implemented using a cloud infrastructure.

[0236] The system 500 may be configured to implement any one or more of the methods described herein or portions thereof.

[0237] Components of the computer system 500 may be implemented at one or more devices, such as the logic device 14, the data stream device 16, the data processing device 18, or the server 22 of Figure 1 .

[0238] The system 500 includes a memory 502 and a processor 504 in communication with the memory 502.

[0239] The memory 502 stores various data received by, operated on, and generated by the system 500.

[0240] The processor 504 is configured to execute various software modules and components. In some embodiments, modules or components executed by the processor 504 may include server-side software components and client-side software components that communicate with each other in order to provide various features and functionalities of the system 500. In some cases, server-side components may be executed at a server computer and client-side components may be executed at a user device.

[0241] The system 504 includes a communication interface device 506 for transmitting and receiving data to and from other computing devices. The communication interface device 506 may include a network interface device for transmitting and receiving data via a network connection (e.g., local area network, wide area network, etc.). The network connection may be wired or wireless connection.

[0242] The system 500 includes a display device 508 for displaying data generated by the system 500. The display device 508 may be located at a user device, such as the user interface device 12 of Figure 1 .

[0243] The system 500 includes an input device 510 for receiving input data from a user interacting with the system 500. For example, a user may use input device 510 to interact with the system 500 through a graphical user interface generated by the processor 504 and displayed via the display device 508. The input device 510 may be located at the user device, such as user interface device 12 of Figure 1 . A user may input information requests about environmental monitoring and disaster responses through the input device 510.

[0244] The processor 504 includes a data acquisition module 512, a data ingestion module 514, a data transferring module 516, a feature extraction module 518, a machine learning (ML) pre-processing module 520, a data archiving and storage module 522, a data indexing and cataloging module 524, and a prioritized data transmission module 526. In other embodiments, the system 500 may include only some of the foregoing modules.

[0245] The data acquisition module 512 configures raw input data 528 concurrently from various data sources (including satellites). The data acquisition module 512 integrates and standardizes data 530 from disparate sources into a united format. The data 530 may be considered pre-processed input data.

[0246] The data ingestion module 514 pre-filters the pre-processed input data 530. The data ingestion module 514 automatically performs preliminary quality and format checks to ensure compliance with data requirements.

[0247] The data transferring module 516 transfers the pre-processed input data 530 to edge computing units within the cloud infrastructure.

[0248] The data transferring module 516 optimizes data transfer times by dynamically selecting the most efficient paths 531 based on real-time network conditions.

[0249] The data transferring module 516 may use a dynamic routing algorithm. The dynamic routing algorithm may optimize data transfer paths 531 based on real-time network conditions using machine learning to continuously improve transfer efficiency.

[0250] The data transfer module 516 may use an adaptive transfer protocol. The adaptive transfer protocol may adjust data compression and transmission parameters to enhance transfer efficiency and reliability. The adaptive transfer protocol may leverage historical data to predict optimal settings.

[0251] The feature extraction module 518 performs geospatial encoding and enrichment on the pre-processed input data 530, including SAR interferometry and / or InSAR. The feature extraction module 518 may use advanced noise reduction and radiometric correction techniques. The feature extraction module 518 generatesgeoreferenced features 532. The feature extraction module 518 may generate the feature data 532 with high precision and accuracy.

[0252] The feature extraction module 518 extracts multidimensional landscape features, including georeferenced data.

[0253] The ML pre-processing module 520 applies supervised or unsupervised learning algorithms to the georeferenced feature data 532. The ML pre-processing module 520 may use, for example, one or more convolutional neural networks (CNNs) or clustering techniques (e.g., K-means clustering).

[0254] The ML pre-processing module 520 analyzes the georeferenced data 532 to detect anomalies, potential threats, and geospatial changes in near-real-time.

[0255] The ML pre-processing module 520 flags relevant data segments 534 for further analysis. This may include object detection or change detection. For object detection, the output includes identified objects with bounding boxes, classification labels, and confidence scores. For change detection, the output includes annotated bounding boxes highlighting areas of significant change with timestamps and change vectors.

[0256] The ML pre-processing module 520 may implement a continuous learning framework. The continuous learning framework updates the machine learning models based on new data inputs and feedback from end users, incorporating reinforcement learning techniques for improved adaptability.

[0257] The ML pre-processing module 520 may include a contextual analysis engine. The contextual analysis engine enhances anomaly detection by considering temporal and spatial context, improving the accuracy and relevance of flagged data segments through advanced pattern recognition and predictive analytics.

[0258] The data archiving storage module 522 stores raw input data 528, pre- processed input data 530 and multidimensional georeferenced features 532 according to lifecycle rules.

[0259] The data indexing and cataloguing module 524 integrates with data catalog services to create metadata 540 for archived data. The created metadata may include a searchable index that facilitates data retrieval for future reference or analysis

[0260] The data indexing and cataloguing module 524 facilitates efficient search and retrieval based on the metadata 540.

[0261] The data indexing and cataloguing module 524 includes a semantic search capability. The semantic search capability allows users to query archived data using advanced natural language processing (NLP) techniques, including sentiment analysis and contextual understanding.

[0262] The data indexing and cataloguing module 524 includes an automated tagging system. The automated tagging system enriches metadata 540 with additional contextual information. This may improve data discoverability and usability through the integration of machine learning algorithms for continuous metadata 540 enhancement.

[0263] The prioritized data transmission module 526 transmits prioritized data 542 to end users based on relevance, urgency, and user preferences. The prioritized data may include the georeferenced data of the multidimensional landscape features, dynamically ranked based on machine learning-driven relevance and urgency assessments.

[0264] The input data may be SAR data captured or otherwise generated by satellites at regular intervals, ensuring consistent and reliable data collection for comprehensive analysis.

[0265] The prioritized data transmission module 526 may be designed to transmit data based on a dynamically updated hierarchy of importance. The prioritized data transmission module 526 may incorporate machine learning insights to refine prioritization criteria. This dynamic prioritization may ensure that the most critical information reaches the end users first, facilitating prompt and effective responses.

[0266] The prioritized data transmission module 526 evaluates the significance of data segments 534 to ensure timely delivery of critical information. Relevance evaluation may include continuously assessing how pertinent the data is to the user's current needs or context using adaptive algorithms. Urgency evaluation may include automatically determining the immediacy required for the data to be actionable based on predefined thresholds and real-time conditions. User preferences evaluation may includeincorporating and updating the specific preferences and requirements of the end user through a user interface.

[0267] The prioritized data transmission module 526 executes a prioritization process. The prioritization process includes evaluating flagged data segments 534 based on relevance, urgency, and user preferences to obtain evaluated data segments 544. The prioritization process includes ranking the evaluated data segments 544 (e.g., from highest to lowest priority). This ranking helps in organizing the data flow efficiently and can be adjusted in real-time based on changing conditions.

[0268] The prioritized data transmission module 526 also segments data into priority levels 546. The priority levels may include, for example, high, medium, and low priority. High priority data includes data that is critical and needs immediate attention. For example, in an emergency response system, alerts about natural disasters or critical medical updates would fall into this category.

[0269] Medium priority data includes data that is important but not urgent. This may include scheduled reports or updates that are necessary but can be delayed if needed.

[0270] Low priority data includes data that is informative but not time-sensitive, such as routine updates or non-critical notifications.

[0271] The prioritized data transmission module 526 transmits high-priority data based on priority to ensure timely response and decision-making.

[0272] High-priority data may be sent immediately using the most efficient transmission paths to ensure it reaches the user as quickly as possible. For instance, in an agricultural monitoring system, alerts about sudden weather changes affecting crops would be transmitted without delay.

[0273] Medium-priority data may be sent based on network availability and current load, ensuring efficient use of resources without compromising the delivery of more urgent data.

[0274] Low-priority data may be transmitted when the system has sufficient bandwidth and capacity, ensuring that it does not interfere with the delivery of more critical information.

[0275] In an emergency response example, high-priority data may include realtime updates about ongoing incidents, while medium-priority may cover status reports from different response teams, and low-priority may be periodic summaries.

[0276] In an agricultural monitoring example, high-priority data may be immediate weather alerts, medium-priority data may include weekly health reports of crops, and low- priority data may be historical data analysis reports.

[0277] In a wildfire monitoring system example, high-priority data may include immediate alerts of detected fires using satellite imagery and thermal sensors. Mediumpriority data may include updates on fire spread and weather conditions affecting the fire, while low-priority data may include historical fire occurrence data and trend analysis.

[0278] The prioritized data transmission module 526 may transmit high-priority data based on priority to ensure timely response and decision-making using adaptive learning and feedback integration. The system may incorporate adaptive learning algorithms that refine data prioritization based on user feedback and changing data patterns, ensuring continuous improvement in data transmission efficiency and relevance.

[0279] The prioritized data transmission module 526 may transmit high-priority data based on priority to ensure timely response and decision-making using security and compliance. This may include ensuring that all data transmissions comply with industrystandard security protocols and regulatory requirements, maintaining data integrity and confidentiality throughout the prioritization and transmission process.

[0280] The prioritized data transmission module 526 may include real-time edge processing. The module 526 may leverage real-time edge processing capabilities to perform initial data analysis at the source. This may reduce latency and bandwidth usage. This approach may ensure that only relevant and high-priority data is transmitted to central servers, enhancing overall system efficiency.

[0281] The prioritized data transmission module 526 may implement scalability and integration. The system may be scalable to handle increasing data volumes and may be flexible enough to integrate with various data sources and types. The system may supporta wide range of applications, from security to environmental monitoring, providing users with a customizable and robust solution for their specific needs.

[0282] The prioritized data transmission module 526 may implement data fusion techniques. For example, the module 526 may combine SAR imagery with other satellite data types to provide a comprehensive and enriched data set. This capability enhances the system’s ability to deliver more accurate and actionable insights for various applications such as disaster response, environmental monitoring, and broader security applications.

[0283] The prioritized data transmission module 526 may provide a user- configurable prioritization framework that allows end users to define custom prioritization criteria based on their specific needs and preferences. This may include machine learning insights to dynamically adjust prioritization settings based on historical data and user behavior (such as geo-reference, location, time / date, sensor type, etc.).

[0284] The prioritized data transmission module 526 may include a push-pull notification system. The push-pull notification system may alert users to high-priority data updates. This provides real-time insights and decision support and may include adaptive learning algorithms to refine notification relevance and timeliness.

[0285] The push-pull mechanism may include a push module and a pull mechanism. The push module may provide the prioritized data to an end user. This may leverage real-time data analytics to ensure timely and relevant data delivery. The pull mechanism may integrate additional data sources to provide comprehensive insights. For example, the pull mechanism may integrate weather data, other satellite images, other sensors, and geospatial databases.

[0286] The push-pull mechanism may enable faster and more efficient responses by combining proactive data pushes with targeted pulls for detailed information, reducing latency, improving resource utilization, and ensuring timely and actionable insights for emergency response and other critical applications.

[0287] In some embodiments, the system 500 transmits further information responsive to a received request, incorporating machine learning to predict and pre-fetchrelevant data, enhancing response efficiency and enabling on-demand tasking for immediate satellite capture and real-time Al-powered analysis.

[0288] The system 500 may ensure data security and compliance by implementing techniques such as: end-to-end encryption for data in transit and at rest using industrystandard algorithms such as AES-256 to protect sensitive information; strict access controls and authentication mechanisms, including OAuth2 tokens and role-based access controls, to ensure that only authorized users can access and modify data; secure data transfer protocols and fault-tolerant messaging systems like Apache Kafka to maintain data integrity and availability; data segregation practices to separate classified and unclassified data, ensuring compliance with regulatory requirements such as but not limited to FedRAMP, ISO 27001 , and HIPAA.

[0289] The system 500 may enhance efficiency and response times through adaptive learning by incorporating machine learning algorithms that continuously refine the push-pull mechanism based on user feedback and changing data patterns.

[0290] The system 500 may enhance efficiency and response times by using edge computing for initial data analysis at the source, reducing latency and bandwidth usage.

[0291] The system 500 may enhance efficiency and response times through automated data prioritization, which may include dynamically adjusting the prioritization of data based on real-time conditions and historical data.

[0292] The system 500 may enhance efficiency and response times through dynamic data retrieval including ensuring immediate access to supplementary information upon user request through dynamic data retrieval techniques, improving the speed and relevance of responses.

[0293] The system 500 may enhance efficiency and response times through comprehensive data integration by combining data from multiple sources, including SAR imagery, weather data, other sensors, and geospatial databases, to provide a holistic view. This integration may enhance situational awareness and decision-making by merging diverse datasets into a unified format, facilitating advanced analysis and more accurate insights.

[0294] The system 500 may enhance efficiency and response times through improved data discoverability by utilizing advanced metadata indexing and semantic search capabilities for efficient data retrieval.

[0295] The system 500 may enhance efficiency and response times through implementing cost-saving measures, such as reduced data transfer and storage costs, through efficient pre-processing and prioritization.

[0296] Referring now to Figure 6, shown therein is a method 600 of cloud enabled data processing of earth observation data, according to an embodiment. The method 600 may be implemented by the system 10 of Figure 1 . One or more steps of the method 600 may be encoded as computer-executable instructions which, when executed by one or more processors, cause the computer system 10 to perform such steps.

[0297] At 602, the method 600 includes receiving data from multiple sources via one or more input devices.

[0298] At 604, the method 600 includes pre-processing the received data to ensure quality and format compliance.

[0299] At 606, the method 600 includes encoding and decoding the pre-processed data to extract multidimensional landscape features.

[0300] Geocoding and enrichment may be performed through SAR interferometry and / or InSAR. Geocoding and enrichment may further include noise reduction, radiometric correction, and assigning geographic coordinates to each data point within the multidimensional landscape features. This may produce highly accurate georeferenced features.

[0301] At 608, the method 600 includes acquiring, ingesting, and transferring the data to regional cloud servers. This may include concurrently receiving raw data from various sources, pre-filtering and ensuring compliance with data requirements, optimizing data transfer times using dynamic routing algorithms, and prioritizing and transmitting data based on predefined criteria.

[0302] At 610, the method 600 includes performing feature extraction and machine learning processing. This may include applying ML algorithms to analyze georeferenceddata, leveraging adaptive learning algorithms to refine prioritization dynamically. This may also include flagging relevant data segments for further analysis, storing and indexing the processed data with lifecycle management rules, and prioritizing and transmitting data based on predefined criteria, leveraging adaptive learning algorithms to refine prioritization dynamically.

[0303] The method 600 may further include receiving a request for further information from the end user responsive to the provided prioritized data and transmitting the further information responsive to the received request.

[0304] The present disclosure may be applied to the paradigm of wildfire detection, which concerns urgency, cost-efficiency, and reliability of information. The input data may include high-resolution satellite imagery, thermal infrared data, and ancillary geospatial datasets such as wind patterns and vegetation indices. A system according to the present disclosure (e.g., the computer system 300) deploys advanced machine learning algorithms, particularly Convolutional Neural Networks (CNNs), to meticulously analyze this imagery and identify features indicative of fire hotspots. These algorithms classify the detected features, discerning active fire regions within the imagery, and produce outputs delineating bounding boxes around detected hotspots, accompanied by classification labels and confidence scores. For instance, upon acquiring thermal infrared data of a forested region, the CNNs detect and classify various fire hotspots, ranging from small flare-ups to extensive conflagrations, across the landscape. The resultant output comprises annotated images where each detected hotspot is encased in a bounding box, labeled with the type of fire and a confidence score indicating the classification's certainty. The flagged data segments include the original satellite images enriched with these annotations. This flagged data is subsequently prioritized for further analysis, ensuring that only the most critical information is forwarded for in-depth examination, thereby minimizing the volume of data necessitating further processing. The system integrates thermal infrared analysis and other geospatial data at the edge of the network, utilizing precisely calibrated software that processes this data tenfold faster than the conventional method of transferring terabytes of data to a centralized location for subsequent processing. Conducting this analysis at the edge significantly reduces the interval from data acquisition to actionable insight. This capability is vital in rapidly evolving situationssuch as wildfires, where immediate and reliable information can substantially influence response efforts, thus mitigating damage and preserving lives. This approach prioritizes pertinent data without necessitating extensive data transfers to central hubs. By deploying machine learning logic at the edge, the system executes real-time processing, focusing exclusively on relevant data sections. This strategy markedly curtails costs associated with data transfer, waiting periods, network bottlenecks, data duplication, storage, and computing resources. The system compresses the time from data acquisition to insight by a factor of ten, leveraging edge computing to perform localized, intelligent data processing. For instance, in a vast forested area susceptible to wildfires, processing every pixel of satellite imagery for potential fire risks would be inefficient and costly. Instead, the system employs edge computing to analyze incoming data streams in real time, flagging only sections with detected fire hotspots for further scrutiny. This method enhances the speed and reliability of fire detection and optimizes resource utilization, ensuring that only critical data is transferred and processed centrally. By focusing computational efforts on the most relevant data, the system provides expedited, more accurate insights, enabling swifter and more effective response measures. The system is advantageously able to execute real-time data processing at the edge, obviating extensive data transfers and facilitating rapid responses to dynamic situations. Traditional systems often necessitate the transfer of vast amounts of data to central locations for processing, resulting in delays and escalated costs. The edge computing capabilities of a system according to the present disclosure ensure that only the most relevant data is processed and transmitted, significantly reducing time, cost, and resource consumption. This innovative approach not only enhances efficiency but further augments the reliability of critical applications such as wildfire detection, providing for enhanced real-time geospatial data analysis. Moreover, by integrating a finely tuned software system that operates ten times faster than traditional methods, the present disclosure addresses the core challenges of modern data processing platforms to seamlessly handle the complexities of diverse and heterogeneous data types, providing robust and scalable solutions for near real-time data processing and analytics across multiple domains. This ensures interoperability and effective data exchange, enabling comprehensive and efficient analysis of complex datasets. The strategic advantage offered by the present disclosure is its unparalleledefficiency in pre-processing data at the source, thereby eliminating the high costs and delays associated with traditional data-centric models. The system’s architecture is designed to handle exponentially growing volumes of data, ensuring that the system remains a formidable tool in the ever-evolving landscape of geospatial data analysis.

[0305] Prioritization of data according to the present disclosure includes a process of evaluation and ranking of data segments based on their assessed importance, thereby facilitating an organized and dynamic data flow that adapts to real-time changes. This process includes categorizing data into high, medium, and low priority levels. High priority data includes critical information that may necessitate immediate attention, such as alerts pertaining to natural disasters or urgent medical updates. Medium priority data includes important, yet not urgent, information such as scheduled reports and necessary updates that can endure a slight delay. Low priority data includes informative, non-time-sensitive content such as routine updates and non-critical notifications. A system or platform according to the present disclosure may advantageously employ near-real-time object detection algorithms to detect and transfer only relevant data, thus enhancing efficiency by reducing the volume of data transferred, minimizing bandwidth usage, and accelerating decision-making processes. The system provides a user-configurable prioritization framework, enabling end users to define custom criteria based on their specific needs and preferences. This includes the integration of machine learning insights to dynamically adjust prioritization settings based on historical data and user behavior, including geo-referencing, location, time / date, and sensor type. The push-pull notification mechanism alerts users to high-priority updates, offering real-time insights and decision support. Adaptive learning algorithms refine the relevance and timeliness of notifications. The system's real-time edge processing capabilities enable near-instantaneous identification of potential threats. Analysis begins almost immediately upon data receipt, evaluating flagged data segments for relevance, urgency, and user preferences. This process includes ranking data segments from highest to lowest priority, which enables organizing data flow efficiently and adjusting in real-time based on changing conditions. High-priority data is transmitted immediately via the most efficient paths to ensure rapid delivery. Medium-priority data is sent based on current network conditions and load, ensuring efficient resource use. Low-priority data is transmitted when bandwidth permits,ensuring no interference with the delivery of more critical information. Detailed examples of applications include emergency response systems where high-priority data includes real-time updates on ongoing incidents, agricultural monitoring systems where immediate weather alerts are considered high-priority data, and wildfire detection systems where immediate fire alerts using satellite imagery and thermal sensors are classified as high- priority. Adaptive learning algorithms continuously refine data prioritization based on evolving patterns and user feedback, ensuring continuous improvement in transmission efficiency and relevance. All data transmissions comply with industry-standard security protocols and regulatory requirements, ensuring data integrity and confidentiality. The system is scalable to handle increasing data volumes and flexible enough to integrate with various data sources and types. The system supports a wide array of applications from security to environmental monitoring, offering a customizable and robust solution. The system supports data fusion techniques, combining Synthetic Aperture Radar (SAR) imagery with other satellite data types to create enriched datasets. This capability enhances the system’s ability to deliver precise and actionable insights for diverse applications, including disaster response and environmental monitoring. The advanced data transferring module employs a dynamic routing algorithm that optimizes data transfer paths based on real-time network conditions using machine learning. An adaptive transfer protocol adjusts data compression and transmission parameters to enhance transfer efficiency and reliability, leveraging historical data to predict optimal settings. A machine learning pre-processing module includes a continuous learning framework that updates machine learning models based on new data inputs and feedback from end users, incorporating reinforcement learning techniques for improved adaptability. Furthermore, a contextual analysis engine enhances anomaly detection by considering temporal and spatial context, improving the accuracy and relevance of flagged data segments through advanced pattern recognition and predictive analytics.

[0306] While the above description provides examples of one or more apparatus, methods, or systems, it will be appreciated that other apparatus, methods, or systems may be within the scope of the claims as interpreted by one of skill in the art.

Claims

Claims:1 . A system for cloud-enabled data processing, the system comprising: one or more input devices for receiving data; one or more data processing devices for pre-processing the received data; and a server for cloud-enabled data processing, the server comprising: a data acquisition module for receiving the pre-processed input data; a data ingestion module for ingesting the pre-processed input data and prefiltering the pre-processed input data to automatically perform preliminary quality and format checks, thereby ensuring compliance with data requirements; a data transferring module for transferring the pre-processed input data to edge computing units within cloud infrastructure positioned to optimize data transfer times; an extraction module for performing geocoding and enrichment on the pre- processed input data to extract multidimensional landscape features including georeferenced data; an ML pre-processing module for performing regional processing by applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near-real-time;data archiving storage for storing raw input data corresponding to the pre- processed input data and the multidimensional landscape features according to lifecycle management rules; a data indexing and cataloguing module for integrating with data catalog services to create metadata for the archived data stored in the data archiving storage; and a prioritized data transmission module for transmitting prioritized data to end users.

2. The system of claim 1 , wherein the server further includes a push-pull mechanism, the push-pull mechanism comprising a push module for providing the prioritized data to an end user, and a pull mechanism for receiving a request for further information from the end user responsive to the provided prioritized data; wherein the server transmits the further information responsive to the received request.

3. The server of claim 2, wherein the extraction module performs the geocoding and enrichment through SAR interferometry and / or InSAR.

4. The server of claim 3, wherein the extraction module further conducts noise reduction and radiometric correction and assigns geographic coordinates to each data point within the multidimensional landscape features to produce the georeferenced features.

5. The server of claim 1 , wherein the created metadata includes a searchable index that facilitates data retrieval for future reference or analysis.

6. The server of claim 1 , wherein the prioritized data includes the georeferenced data of the multidimensional landscaped features.

7. The server of claim 1 , wherein the input data is SAR data captured or otherwise generated by satellites at regular intervals.

8. A method of cloud-enabled data processing, the method comprising: receiving or acquiring input data; ingesting the input data and pre-filtering the input data to automatically perform preliminary quality and format checks; transferring the input data to edge computing units within a cloud infrastructure positioned to optimize data transfer times; performing geocoding and enrichment on the input data to extract multidimensional landscape features including georeferenced data; performing regional preprocessing by applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near-real-time; archiving raw input data corresponding to the pre-processed input data and the multidimensional landscape features according to lifecycle management rules; integrating with data catalog services to create metadata for the archived data stored in the data archiving storage; and transmitting prioritized data to end users.

9. The method of claim 8 further comprising:receiving a request for further information from the end user responsive to the provided prioritized data; and transmitting the further information responsive to the received request.

10. The method of claim 8, wherein the geocoding and enrichment are performed through SAR interferometry and / or InSAR.11 . The method of claim 10, wherein the geocoding and enrichment further comprise noise reduction, radiometric correction, and assigning geographic coordinates to each data point within the multidimensional landscape features to produce the georeferenced features.

12. The method of claim 8, wherein the created metadata includes a searchable index that facilitates data retrieval for future reference or analysis.

13. The method of claim 8, wherein the prioritized data includes the georeferenced data of the multidimensional landscaped features.

14. The method of claim 8, wherein the input data is SAR data captured or otherwise generated by satellites at regular intervals.

15. A server for cloud-enabled data processing, the server receiving pre-processed input data, the server comprising: a data acquisition module for receiving the pre-processed input data; a data ingestion module for ingesting the pre-processed input data and prefiltering the pre-processed input data to automatically perform preliminary quality and format checks, thereby ensuring compliance with data requirements;a data transferring module for transferring the pre-processed input data to edge computing units within cloud infrastructure positioned to optimize data transfer times; an extraction module for performing geocoding and enrichment on the pre- processed input data to extract multidimensional landscape features including georeferenced data; an ML pre-processing module for applying one or more supervised or unsupervised ML algorithms, including CNNs and K-Means Clustering, to analyze the georeferenced data of the multidimensional landscape features to detect anomalies and potential threats in near-real-time; data archiving storage for storing raw input data corresponding to the pre- processed input data and the multidimensional landscape features according to lifecycle management rules; a data indexing and cataloguing module for integrating with data catalog services to create metadata for the archived data stored in the data archiving storage; and a prioritized data transmission module for transmitting prioritized data to end users.

16. The server of claim 15, wherein the server further includes a push-pull mechanism, the push-pull mechanism comprising a push module for providing the prioritized data to an end user, and a pull mechanism for receiving a request for further information from the end user responsive to the provided prioritized data; wherein the server transmits the further information responsive to the received request.

17. The server of claim 15, wherein the extraction module performs the geocoding and enrichment through SAR interferometry and / or InSAR.

18. The server of claim 17, wherein the extraction module further conducts noise reduction and radiometric correction and assigns geographic coordinates to each data point within the multidimensional landscape features to produce the georeferenced features.

19. The server of claim 15, wherein the created metadata includes a searchable index that facilitates data retrieval for future reference or analysis.

20. The server of claim 15, wherein the prioritized data includes the georeferenced data of the multidimensional landscaped features.21 . The server of claim 15, wherein the input data is SAR data captured or otherwise generated by satellites at regular intervals.