Dynamic fusion processing method and system for multi-source multi-domain heterogeneous data

By employing dynamic acquisition, semantic fusion, and distributed parallel computing, the problems of acquisition difficulties and low fusion accuracy in processing heterogeneous data from multiple sources and domains have been solved, achieving efficient and reliable data processing, meeting real-time requirements, and improving data quality and storage efficiency.

CN122020501APending Publication Date: 2026-05-12CHINA SCIENCE SATELLITE (ANHUI) DATA TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA SCIENCE SATELLITE (ANHUI) DATA TECHNOLOGY CO LTD
Filing Date
2025-09-25
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies suffer from low data acquisition efficiency, low fusion accuracy, and poor processing efficiency when processing heterogeneous data from multiple sources and domains, and cannot meet the needs of application scenarios with high real-time requirements.

Method used

It adopts dynamic acquisition of multi-source heterogeneous data, dynamically generates acquisition strategies by identifying data source types, formats and transmission protocols in real time, and converts data into a unified intermediate format; it achieves deep semantic fusion of multi-domain data and uses knowledge graphs to display semantic relationships; it adopts distributed intelligent parallel computing to decompose processing tasks and dynamically allocate resources; and it implements hierarchical storage and lifecycle management.

Benefits of technology

It improved the coverage, accuracy, and timeliness of data collection, enhanced the quality and availability of the fused data, met real-time requirements, optimized the utilization rate of data storage resources, and ensured the security and integrity of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020501A_ABST
    Figure CN122020501A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic fusion processing method and system for multi-source and multi-domain heterogeneous data. The method comprises the following steps: S1, dynamically collecting multi-source heterogeneous data; s2, realizing multi-domain data semantic deep fusion; s3, executing distributed intelligent parallel computing; and S4, implementing hierarchical storage and life cycle management. Through the innovative data acquisition, fusion, calculation and management method, the bottleneck of multi-source multi-domain heterogeneous data processing in the prior art is broken through, efficient and accurate processing of complex data is realized, the utilization value of the data is improved, and a firm and reliable technical support is provided for data-based decision and application in each field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a dynamic fusion processing method and system for heterogeneous data from multiple sources and domains. Background Technology

[0002] In today's digital age, data across various fields is experiencing explosive growth and exhibits complex characteristics of being multi-source, multi-domain, and heterogeneous. Taking environmental monitoring as an example, data sources encompass large-area macro-environmental data acquired by satellite remote sensing equipment, localized high-resolution data collected by low-altitude drones, real-time point-based data monitored by ground sensor equipment, and various related texts, images, and statistical data from the internet. Satellite remote sensing data is characterized by its wide coverage and strong periodicity, providing global or large-area environmental information such as vegetation cover, land use, and water distribution; however, its resolution is relatively low, making it difficult to accurately capture local details. Drone data, on the other hand, boasts high mobility and high resolution, allowing for flexible observation of specific areas and the acquisition of detailed information such as topography and vegetation health; however, its observation range is limited, and data acquisition is significantly constrained by flight conditions. Ground sensor equipment can monitor environmental parameters at its location in real time and accurately, such as temperature, humidity, and air quality, but it only reflects the situation at a local point and lacks overall spatial information. Internet data comes from a wide range of sources, including environmental-related images and texts posted by users on social media, research reports and statistical data from professional websites, etc. These data are diverse in format and quality, and differ significantly from professional monitoring data in semantics and structure.

[0003] Current traditional data processing technologies exhibit numerous problems when faced with such complex, multi-source, and multi-domain heterogeneous data. In the data acquisition phase, there is a lack of unified, efficient, and adaptive acquisition mechanisms for different types of data sources. For example, satellite remote sensing data receiving equipment is typically designed for specific satellites and data formats, making it difficult to quickly adapt to new satellite data sources or changes in data formats; UAV data acquisition relies on manual operation or specific flight planning software, failing to respond in real-time to complex and ever-changing monitoring needs; ground sensor equipment data acquisition is limited by communication protocols and equipment compatibility, making it difficult to simultaneously acquire data from equipment from different manufacturers; and internet data acquisition faces challenges such as the legality of data crawling, website anti-crawling mechanisms, and the diversity of data formats, resulting in low data acquisition efficiency and difficulty in ensuring data integrity.

[0004] In the data fusion stage, most existing fusion methods operate only based on simple data structure or format conversion, failing to delve into the inherent semantic relationships between data from different domains. Different domains define, describe, and express the same concept in different ways. For example, in environmental monitoring, satellite remote sensing data describes vegetation through spectral features, UAV data analyzes vegetation conditions from image texture, ground sensor data measures vegetation using indicators such as biomass, and internet text data mentions vegetation-related information in natural language. Traditional fusion methods struggle to effectively integrate these data with different expressions, resulting in fused data failing to fully realize its potential value and providing comprehensive and accurate support for decision-making.

[0005] From the perspective of data processing efficiency, with the rapid increase in data volume, traditional single-machine or simple parallel computing models are proving inadequate for handling large-scale, multi-source, multi-domain heterogeneous data. The large-scale nature of multi-source heterogeneous data makes data storage and transmission bottlenecks. Different types of data have varying processing requirements, and traditional computing models cannot fully utilize distributed computing resources or achieve efficient parallel processing for different data types. This results in long processing times, making it difficult to meet the real-time requirements of applications such as disaster early warning systems, which require timely analysis of multi-source heterogeneous data for rapid decision-making. Therefore, developing a key technology for processing multi-source, multi-domain heterogeneous data that can effectively solve these problems is urgently needed. Summary of the Invention

[0006] To address the existing problems, this invention provides a dynamic fusion processing method and system for multi-source, multi-domain heterogeneous data, the specific solution of which is as follows: A dynamic fusion processing method for heterogeneous data from multiple sources and domains includes the following steps: S1, dynamically acquires heterogeneous data from multiple sources: Real-time identification of data source type, data format, and transmission protocol; Data acquisition strategies are dynamically generated based on the recognition results, including acquisition frequency, transmission method and parsing rules; Perform data acquisition and convert heterogeneous data into a unified intermediate format using an adaptive parser; S2 enables deep semantic fusion of multi-domain data: Extract entity and attribute information from multi-source heterogeneous data and semantically label categories; Construct knowledge graphs for different domains to graphically represent the semantic relationships between data; By comparing entities, attributes, and relationships in knowledge graphs from different domains, semantic associations are mined, and semantic associations are used to fuse multi-source heterogeneous data to generate a semantically unified dataset. S3 performs distributed intelligent parallel computing: The processing task is decomposed into independent but related subtasks according to data type and processing logic; Real-time monitoring of computing node resource status and dynamic allocation of subtasks to distributed computing nodes; Subtasks are executed using parallel strategies such as data partitioning parallelism or task pipeline parallelism. S4 implements tiered storage and lifecycle management: Choose the storage medium based on the degree of data structuring and access frequency—structured data is stored in a relational database, while semi / unstructured data is stored in a non-relational database or a distributed file system; Establish a data catalog and indexing mechanism; Data is migrated to tiered storage media based on its timeliness, and regular cleanup and backup are performed.

[0007] Preferably, the data source types mentioned in step S1 include: satellite remote sensing data, UAV data, ground sensor equipment data, and Internet data.

[0008] Preferably, the dynamic data acquisition strategy in step S1 includes: For satellite remote sensing data, the acquisition frequency is dynamically adjusted according to the satellite's orbital period, data update frequency, and monitoring needs. Data transmission protocols are used to ensure rapid data transmission, and a dedicated satellite data analyzer is used to accurately analyze data of different bands and resolutions. Regarding drone data, when a new flight mission or change in the monitoring area is detected, the intelligent sensing module automatically adjusts the acquisition strategy, adjusts the data transmission rate in real time to adapt to network conditions, and uses image recognition technology to automatically identify and filter valid image data. For ground sensor equipment data, when a new device is connected or the device parameters change, the communication parameters are automatically configured, and corresponding data parsing rules are generated according to the sensor type to ensure accurate data collection. For internet data, intelligent algorithms are used to bypass website anti-scraping mechanisms, and collection tools and parsing methods are selected according to data type to ensure the legality and integrity of the data.

[0009] Preferably, the semantic annotation categories mentioned in step S2 specifically include: Satellite remote sensing data is analyzed for spectral characteristics, and semantic relationships are established between data in different spectral bands and environmental elements. Ground sensor data is semantically labeled based on the correspondence between monitoring parameters and environmental concepts; UAV image data is used to identify ground features using image recognition technology and assign corresponding semantic labels; Internet text data is processed through natural language processing to extract key information and perform semantic classification.

[0010] Preferably, the dynamic allocation of subtasks in step S3 includes: allocating computationally intensive subtasks to high CPU performance computing nodes; and allocating data-intensive subtasks to high storage resource computing nodes.

[0011] Preferably, the parallel strategy in step S3 includes: using data partitioning parallelism for large-scale spatial data; and using task pipeline parallelism for image processing tasks.

[0012] Preferably, the dynamic fusion processing method for multi-source, multi-domain heterogeneous data further includes lifecycle management, which includes: In the raw data collection phase, ensure the integrity and accuracy of the data by performing preliminary data cleaning; During the data processing and analysis phase, data is efficiently calculated and analyzed to generate valuable information. During the data archiving phase, data that has not been used for a long time or is no longer frequently accessed is migrated to low-cost storage media for storage, while a small amount of key index information is retained in the relational database so that data can be quickly located and recovered when needed. Data cleaning and backup are performed regularly to ensure data security and availability. The data that has not been used for a long time or is no longer frequently accessed is historical order data, and the low-cost storage media is a tape library.

[0013] The present invention also discloses a system based on any of the methods described above, comprising: The multi-source heterogeneous data adaptive acquisition module is used to identify the data source type, format and transmission protocol in real time and dynamically generate data acquisition strategies; The multi-domain data semantic deep fusion module realizes data semantic annotation and association fusion through a cross-domain general semantic model library; A distributed intelligent optimization parallel computing framework that performs parallel computing based on task decomposition and resource scheduling strategies; The hybrid data storage management system uses relational databases, non-relational databases, and distributed file systems to store data in a hierarchical manner.

[0014] Preferably, the multi-source heterogeneous data adaptive acquisition module includes: an intelligent sensing unit that identifies data source attributes through protocol detection and file header analysis; a dynamic strategy library that generates acquisition frequency, transmission method, and parsing rules according to the data source type; and an adaptive parser that converts heterogeneous data into a unified intermediate format. The multi-domain data semantic deep fusion module includes: a semantic annotation unit, which extracts environmental element entities using named entity recognition technology; a knowledge graph construction unit, which establishes entity associations in different domains; and a semantic mapping unit, which realizes multi-source heterogeneous data fusion based on common entities / attributes. The distributed parallel computing framework includes: a task decomposition unit, which splits and processes tasks according to data type and logic; an intelligent scheduler, which monitors node resources in real time and dynamically allocates subtasks; and a parallel optimization unit, which adopts data partitioning parallelism and task pipeline parallelism strategies. The hybrid data storage management system includes: a storage selection unit that allocates storage media according to data structure and access frequency; an index management unit that establishes composite indexes for relational / non-relational data; and a lifecycle management unit that migrates data to tiered storage media according to data timeliness.

[0015] The beneficial effects of this invention are as follows: It greatly improves the coverage, accuracy and timeliness of data collection, can automatically adapt to various complex and ever-changing data sources and data formats, reduces manual intervention, lowers data collection costs, and provides a rich and reliable data foundation for subsequent data processing.

[0016] By using deep semantic fusion, richer and more valuable correlation information between multi-domain data is unearthed, significantly improving the quality and usability of the fused data, providing more comprehensive and accurate data support for decision-making, and enhancing the scientific nature and reliability of decision-making.

[0017] Based on a distributed intelligent optimization parallel computing framework, it makes full use of distributed computing resources, effectively improves data processing efficiency, and can meet the needs of application scenarios with high real-time requirements, such as disaster early warning and real-time environmental monitoring, providing strong support for rapid response and decision-making.

[0018] The innovative data storage and management system enables efficient data storage, rapid retrieval, and secure management. It makes reasonable plans based on data characteristics and lifecycle, improves the utilization rate of data storage resources, ensures data security and integrity, and facilitates long-term data management and maintenance. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic flowchart of the present invention; Figure 2 This is a flowchart illustrating the dynamic acquisition of multi-source heterogeneous data according to the present invention. Figure 3 This is a flowchart illustrating the process of achieving deep semantic fusion of multi-domain data in this invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] This invention provides a dynamic fusion processing method and system for multi-source, multi-domain heterogeneous data. It aims to solve the current challenges in processing multi-source, multi-domain heterogeneous data, such as difficulties in data acquisition, low fusion accuracy, and poor processing efficiency. Through innovative data acquisition, fusion, computation, and management methods, it overcomes the bottlenecks in existing multi-source, multi-domain heterogeneous data processing technologies, improving the accuracy, efficiency, and usability of data processing to meet the needs of complex data processing in various fields such as environmental monitoring, urban planning, and disaster early warning.

[0023] like Figure 1 A dynamic fusion processing method for heterogeneous data from multiple sources and domains includes the following steps: S1, dynamically collects heterogeneous data from multiple sources, such as Figure 2 : S11 identifies data source type, data format, and transmission protocol in real time.

[0024] The data sources include: satellite remote sensing data, UAV data, ground sensor equipment data, and internet data. Data formats include: specific band formats for satellite data, UAV image formats, numerical formats for sensor data, and text, image, and table formats for internet data. Transmission protocols include: specific communication protocols commonly used for satellite data, UAV data transmission protocols, wired or wireless communication protocols for sensor equipment, and protocols such as HTTP and FTP for acquiring internet data.

[0025] Specifically, the intelligent sensing module continuously scans the network environment. When a new data source is detected, it quickly identifies the data source type, data format, and transmission protocol through techniques such as protocol detection and file header analysis. For a newly connected database data source, the intelligent sensing module can attempt to connect to the database to obtain the database type, such as MySQL or Oracle, and analyze the database table structure and data storage method to determine the data format.

[0026] S12: Dynamically generate data acquisition strategies based on the recognition results. According to the recognition results, select or dynamically generate an appropriate acquisition strategy from a predefined strategy library. The strategy includes data acquisition frequency, data transmission method, and data parsing rules. For sensor data sources with high real-time requirements, generate high-frequency acquisition strategies and adopt streaming transmission; for file data sources with complex data formats, generate corresponding complex data parsing rules.

[0027] Specifically, the dynamically generated data acquisition strategy includes: For satellite remote sensing data, the acquisition frequency is dynamically adjusted according to the satellite's orbital period, data update frequency, and monitoring needs. Data transmission protocols are used to ensure rapid data transmission, and a dedicated satellite data analyzer is used to accurately analyze data of different bands and resolutions. Regarding drone data, when a new flight mission or change in the monitoring area is detected, the intelligent sensing module automatically adjusts the acquisition strategy, adjusts the data transmission rate in real time to adapt to network conditions, and uses image recognition technology to automatically identify and filter valid image data. For ground sensor equipment data, when a new device is connected or the device parameters change, the communication parameters are automatically configured, and corresponding data parsing rules are generated according to the sensor type to ensure accurate data collection. For internet data, intelligent algorithms are used to bypass website anti-scraping mechanisms, and collection tools and parsing methods are selected according to data type to ensure the legality and integrity of the data.

[0028] S13 executes data acquisition and converts heterogeneous data into a unified intermediate format using an adaptive parser. Specifically, according to the generated acquisition strategy, the data acquisition program is started, and the acquired data is converted into a unified intermediate format, such as JSON, using an adaptive data parser. When acquiring configuration file data in XML format, the data parser converts it into a format conforming to the JSON specification based on the XML tag structure and data type definitions, facilitating subsequent processing.

[0029] S2 enables deep semantic fusion of multi-domain data: such as Figure 3This project aims to create a cross-domain general semantic model library, utilizing advanced technologies such as natural language processing, knowledge graphs, and machine learning to perform in-depth semantic annotation and analysis of data from different domains. In the field of environmental monitoring, for satellite remote sensing data, spectral feature analysis and semantic association are used to establish semantic connections between different bands of data and environmental elements, such as vegetation type and water pollution levels. For UAV image data, image recognition technology is used to identify ground features and assign corresponding semantic labels. Ground sensor data is semantically annotated based on the correspondence between monitoring parameters and environmental concepts. Internet text data is extracted using natural language processing technology and semantically classified. By mining potential semantic connections between data from different domains, such as based on common environmental entities (rivers, forests, etc.), attributes (temperature, area, etc.), or events (natural disasters, environmental changes, etc.), mapping relationships are established between data, achieving deep fusion of multi-domain data. When fusing satellite remote sensing data and ground sensor data, rivers are used as a common entity. The location and shape information of rivers in satellite remote sensing images are integrated with data such as river flow and water quality monitored by ground sensors based on semantic association to form more comprehensive and accurate river environmental information.

[0030] Specifically: S21, Semantic Labeling: For data from different domains, natural language processing techniques such as named entity recognition and part-of-speech tagging are used to process the textual information in the data, extract key information such as entities and attributes, and perform semantic labeling. In medical record text data in the medical field, entities such as patient names, disease names, and symptoms are identified and their semantic categories are labeled.

[0031] S22, Knowledge Graph Construction: Based on semantic annotation results and combined with domain knowledge, construct knowledge graphs for various domains. Knowledge graphs graphically represent the semantic relationships between data, including associations between entities and relationships between attributes. In the financial domain, construct a knowledge graph containing entities such as customers, accounts, and transactions, and their relationships.

[0032] S23, Semantic Association Mining and Fusion: By comparing entities, attributes, and relationships in knowledge graphs from different domains, potential semantic associations are mined. These associations are then used to fuse data from different domains, forming a unified, semantically rich dataset. When fusing data from the financial and e-commerce domains, it was discovered that customer IDs exist in both domain knowledge graphs. Using this as a bridge, customer credit information in the financial domain and consumer behavior information in the e-commerce domain are linked and fused.

[0033] S3 performs distributed intelligent parallel computing: Employing a distributed computing architecture, it meticulously decomposes large-scale, multi-source, multi-domain heterogeneous data processing tasks into multiple independent yet related subtasks based on data type, processing logic, and other factors. Using an intelligent task scheduler, it monitors the resource usage of each computing node (CPU utilization, memory usage, network bandwidth, etc.) and the characteristics of subtasks (computational complexity, data size, real-time requirements, etc.) in real time, dynamically allocating tasks based on optimization algorithms to achieve efficient utilization of computing resources. Parallel computing optimization strategies are introduced. For example, for large-scale satellite remote sensing data, a data partitioning parallel strategy is used, distributing remote sensing data from different regions to different computing nodes for parallel processing. For complex image processing tasks involving UAV image data, a task pipeline parallel strategy is used, assembling tasks such as image preprocessing, feature extraction, and target recognition into a pipeline, with different computing nodes responsible for different stages, improving the efficiency of parallel computing. When processing comprehensive tasks involving satellite remote sensing, UAV, ground sensor, and internet data, multiple parallel strategies are flexibly combined according to data characteristics and task requirements, fully leveraging the advantages of distributed computing resources and significantly shortening data processing time.

[0034] Specifically: S31, Task Decomposition: The task of processing multi-source, multi-domain heterogeneous data is decomposed into multiple independent but related sub-tasks based on factors such as data type and processing logic. When processing multi-source heterogeneous data containing images and text, tasks such as image feature extraction and classification, and text segmentation and sentiment analysis are treated as different sub-tasks.

[0035] S32, Resource Assessment and Task Allocation: The intelligent task scheduler monitors the resource usage of each computing node in real time, including CPU utilization, remaining memory, and network bandwidth. Based on the computational complexity and data volume requirements of the subtasks, it allocates them to the most suitable computing nodes. For computationally intensive image recognition subtasks, they are allocated to computing nodes with high CPU performance; for data-intensive text data storage and retrieval subtasks, they are allocated to computing nodes with abundant storage resources.

[0036] S33, Parallel Computing Execution: Each computing node executes subtasks in parallel according to the task allocation results. During execution, optimization strategies such as data partitioning parallelism and task pipeline parallelism are adopted to improve the efficiency of parallel computing. In data partitioning parallelism, large-scale text data is divided into multiple partitions according to certain rules (such as by paragraph, by keyword, etc.), and processed on different computing nodes respectively. In task pipeline parallelism, tasks such as image preprocessing, feature extraction, and classification are composed of pipelines, with each computing node responsible for one step. Data is processed sequentially on each node, reducing the waiting time between nodes.

[0037] S4, Implementing Tiered Storage: A hybrid data storage architecture integrating the advantages of relational databases, non-relational databases, and distributed file systems is designed. The appropriate storage method is intelligently selected based on the data structure, access frequency, and storage requirements. Historical monitoring data from ground sensor equipment with high structure and frequent transaction processing needs is stored in a relational database, leveraging its powerful transaction processing and complex query capabilities. Semi-structured and unstructured satellite remote sensing image data, UAV image data, and internet text data with large volumes and relatively flexible query methods are stored in a non-relational database or distributed file system, with an indexing mechanism ensuring fast data retrieval. A detailed data catalog and index are established to classify and manage data from different sources and of different types, facilitating rapid data location and retrieval.

[0038] Specifically: S41, Storage Method Selection: Determine the data storage method based on the data structure, access frequency, and storage requirements. For structured, highly transactional business data requiring frequent complex queries, choose a relational database. For semi-structured, unstructured data with large volumes and relatively simple query methods, such as log files and multimedia files, choose a non-relational database or distributed file system. Enterprise order data, due to its high degree of structure and frequent transaction processing requirements, is stored in the relational database MySQL; while user-uploaded images, videos, and other multimedia files are stored in the distributed file system Ceph and indexed using the non-relational database MongoDB for convenient and fast querying.

[0039] S42, Data Catalog and Index Creation: While storing data, establish a detailed data catalog and indexes. For relational databases, utilize the database's built-in indexing mechanism to create indexes on frequently queried fields. For non-relational databases and distributed file systems, create indexes based on key data characteristics (such as filename, file type, creation time, etc.) to improve data retrieval speed. In MongoDB, create a composite index based on filename and file type for the collection storing image files to quickly retrieve specific types of image files.

[0040] The method of this invention also introduces a data lifecycle management strategy. Based on factors such as the frequency of data use and timeliness, the data is divided into different stages (such as raw data collection, processed and analyzed data, and long-term archived data). Different storage and management methods are adopted for data in different stages. For example, historical data that has not been used for a long time is migrated to a low-cost storage medium. At the same time, the data is cleaned and backed up regularly to ensure the security and effectiveness of the data.

[0041] Data Lifecycle Management: Develop a data lifecycle management strategy, dividing data into different stages, such as raw data acquisition, data processing and analysis, and data archiving. Different management measures are adopted for data at different stages. In the raw data acquisition stage, ensure data integrity and accuracy by performing preliminary data cleaning. In the data processing and analysis stage, perform efficient calculations and analysis to generate valuable information. In the data archiving stage, migrate long-unused or infrequently accessed data to low-cost storage media and perform regular data cleaning and backups to ensure data security and availability. For historical order data from one year ago—data that has been unused or infrequently accessed for a year—migrate it from the relational database to a tape library (a low-cost storage medium) for archiving, while retaining a small amount of key index information in the database for quick data location and recovery when needed.

[0042] This invention also discloses a dynamic fusion processing system for heterogeneous data from multiple sources and domains, comprising: The multi-source heterogeneous data adaptive acquisition module is used to identify the data source type, format, and transmission protocol in real time and dynamically generate data acquisition strategies. It includes: an intelligent sensing unit, which identifies data source attributes through protocol detection and file header analysis; a dynamic strategy library, which generates acquisition frequency, transmission method, and parsing rules according to the data source type; and an adaptive parser, which converts heterogeneous data into a unified intermediate format.

[0043] The multi-domain data semantic deep fusion module realizes data semantic annotation and association fusion through a cross-domain general semantic model library; it includes a semantic annotation unit, which uses named entity recognition technology to extract environmental element entities; a knowledge graph construction unit, which establishes the relationship between entities in different domains; and a semantic mapping unit, which realizes the fusion of multi-source heterogeneous data based on common entities / attributes.

[0044] The distributed intelligent optimization parallel computing framework performs parallel computing based on task decomposition and resource scheduling strategies. It includes a task decomposition unit that splits and processes tasks according to data type and logic; an intelligent scheduler that monitors node resources in real time and dynamically allocates subtasks; and a parallel optimization unit that adopts data partitioning parallelism and task pipeline parallelism strategies.

[0045] The hybrid data storage management system uses relational databases, non-relational databases, and distributed file systems to store data in a hierarchical manner. It includes a storage selection unit that allocates storage media according to data structure and access frequency; an index management unit that creates composite indexes for relational / non-relational data; and a lifecycle management unit that migrates data to hierarchical storage media according to its timeliness.

[0046] This invention greatly improves the coverage, accuracy, and timeliness of data collection, can automatically adapt to various complex and ever-changing data sources and formats, reduces manual intervention, lowers data collection costs, and provides a rich and reliable data foundation for subsequent data processing.

[0047] This invention, through deep semantic fusion, uncovers richer and more valuable correlations between multi-domain data, significantly improving the quality and usability of the fused data, providing more comprehensive and accurate data support for decision-making, and enhancing the scientific rigor and reliability of decision-making.

[0048] This invention is based on a distributed intelligent optimization parallel computing framework, which makes full use of distributed computing resources, effectively improves data processing efficiency, and can meet the needs of application scenarios with high real-time requirements, such as disaster early warning and real-time environmental monitoring, providing strong support for rapid response and decision-making.

[0049] This invention's innovative data storage and management system achieves efficient data storage, rapid retrieval, and secure management. It rationally plans data based on its characteristics and lifecycle, improving the utilization rate of data storage resources, ensuring data security and integrity, and facilitating long-term data management and maintenance.

[0050] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0051] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dynamic fusion processing method for heterogeneous data from multiple sources and domains, characterized in that, Includes the following steps: S1, dynamically acquires heterogeneous data from multiple sources: Real-time identification of data source type, data format, and transmission protocol; Data acquisition strategies are dynamically generated based on the recognition results, including acquisition frequency, transmission method and parsing rules; Perform data acquisition and convert heterogeneous data into a unified intermediate format using an adaptive parser; S2 enables deep semantic fusion of multi-domain data: Extract entity and attribute information from multi-source heterogeneous data and semantically label categories; Construct knowledge graphs for different domains to graphically represent the semantic relationships between data; By comparing entities, attributes, and relationships in knowledge graphs from different domains, semantic associations are mined, and semantic associations are used to fuse multi-source heterogeneous data to generate a semantically unified dataset. S3 performs distributed intelligent parallel computing: The processing task is decomposed into independent but related subtasks according to data type and processing logic; Real-time monitoring of computing node resource status and dynamic allocation of subtasks to distributed computing nodes; Subtasks are executed using parallel strategies such as data partitioning parallelism or task pipeline parallelism. S4 implements tiered storage: Choose the storage medium based on the degree of data structuring and access frequency—structured data is stored in a relational database, while semi / unstructured data is stored in a non-relational database or a distributed file system; Establish a data catalog and indexing mechanism; Data is migrated to tiered storage media based on its timeliness, and regular cleanup and backup are performed.

2. The method according to claim 1, characterized in that, The data source types mentioned in step S1 include: satellite remote sensing data, UAV data, ground sensor equipment data, and Internet data.

3. The method according to claim 2, characterized in that, The dynamic data acquisition strategy described in step S1 includes: For satellite remote sensing data, the acquisition frequency is dynamically adjusted according to the satellite's orbital period, data update frequency, and monitoring needs. Data transmission protocols are used to ensure rapid data transmission, and a dedicated satellite data analyzer is used to accurately analyze data of different bands and resolutions. Regarding drone data, when a new flight mission or change in the monitoring area is detected, the intelligent sensing module automatically adjusts the acquisition strategy, adjusts the data transmission rate in real time to adapt to network conditions, and uses image recognition technology to automatically identify and filter valid image data. For ground sensor equipment data, when a new device is connected or the device parameters change, the communication parameters are automatically configured, and corresponding data parsing rules are generated according to the sensor type to ensure accurate data collection. For internet data, intelligent algorithms are used to bypass website anti-scraping mechanisms, and collection tools and parsing methods are selected according to data type to ensure the legality and integrity of the data.

4. The method according to claim 2, characterized in that, The semantic annotation categories mentioned in step S2 specifically include: Satellite remote sensing data is analyzed for spectral characteristics, and semantic relationships are established between data in different spectral bands and environmental elements. Ground sensor data is semantically labeled based on the correspondence between monitoring parameters and environmental concepts; UAV image data is used to identify ground features using image recognition technology and assign corresponding semantic labels; Internet text data is processed through natural language processing to extract key information and perform semantic classification.

5. The method according to claim 1, characterized in that, The dynamic allocation of subtasks in step S3 includes: allocating computationally intensive subtasks to high CPU performance computing nodes; and allocating data-intensive subtasks to high storage resource computing nodes.

6. The method according to claim 1, characterized in that, The parallel strategies described in step S3 include: using data partitioning parallelism for large-scale spatial data; and using task pipeline parallelism for image processing tasks.

7. The method according to claim 1, characterized in that, It also includes lifecycle management, which includes: In the raw data collection phase, ensure the integrity and accuracy of the data by performing preliminary data cleaning; In the data processing and analysis stage, data is efficiently calculated and analyzed to generate valuable information; During the data archiving phase, data that has not been used for a long time or is no longer frequently accessed is migrated to a low-cost storage medium for storage, while a small amount of key index information is retained in the relational database so that the data can be quickly located and recovered when needed; data cleaning and backup are performed regularly to ensure data security and availability; the data that has not been used for a long time or is no longer frequently accessed is historical order data, and the low-cost storage medium is a tape library.

8. A system based on the method of any one of claims 1-7, characterized in that, include: The multi-source heterogeneous data adaptive acquisition module is used to identify the data source type, format and transmission protocol in real time and dynamically generate data acquisition strategies; The multi-domain data semantic deep fusion module realizes data semantic annotation and association fusion through a cross-domain general semantic model library; A distributed intelligent optimization parallel computing framework that performs parallel computing based on task decomposition and resource scheduling strategies; The hybrid data storage management system uses relational databases, non-relational databases, and distributed file systems to store data in a hierarchical manner.

9. The system according to claim 8, characterized in that: The multi-source heterogeneous data adaptive acquisition module includes: an intelligent sensing unit that identifies data source attributes through protocol detection and file header analysis; a dynamic strategy library that generates acquisition frequency, transmission method, and parsing rules based on the data source type; and an adaptive parser that converts heterogeneous data into a unified intermediate format. The multi-domain data semantic deep fusion module includes: a semantic annotation unit, which extracts environmental element entities using named entity recognition technology; a knowledge graph construction unit, which establishes entity associations in different domains; and a semantic mapping unit, which realizes multi-source heterogeneous data fusion based on common entities / attributes. The distributed parallel computing framework includes: a task decomposition unit, which splits and processes tasks according to data type and logic; an intelligent scheduler, which monitors node resources in real time and dynamically allocates subtasks; and a parallel optimization unit, which adopts data partitioning parallelism and task pipeline parallelism strategies. The hybrid data storage management system includes: a storage selection unit that allocates storage media according to data structure and access frequency; an index management unit that establishes composite indexes for relational / non-relational data; and a lifecycle management unit that migrates data to tiered storage media according to data timeliness.