Multi-source heterogeneous data hierarchical convergence method and device, equipment and storage medium
By hierarchically aggregating multi-source heterogeneous data from unmanned helicopters and classifying and storing the data based on data format, popularity, and time dimensions, the problem of low query efficiency in data lakes is solved, and efficient data query and management are achieved.
Patent Information
- Application Number
- CN202510794546.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, storing multi-source heterogeneous data generated during unmanned helicopter flight missions in a data lake results in low query efficiency and makes it difficult to meet specific data query requests.
A multi-source heterogeneous data hierarchical aggregation method is adopted. Based on the data format dimension, usage popularity dimension, and time dimension, the category information of multi-source heterogeneous data is determined and stored in different data pools of the data lake, including structured, semi-structured, unstructured, hot data pool and time series data pool, and stored and queried through tag information.
It significantly improves data query speed and efficiency, reduces invalid scanning of raw data in the data lake, enables rapid location and retrieval, and supports real-time monitoring and fault diagnosis of unmanned helicopters.
Smart Images

Figure CN120873008A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage technology, and in particular to a method, apparatus, device, and storage medium for hierarchical aggregation of multi-source heterogeneous data. Background Technology
[0002] With the increasingly widespread application of unmanned helicopters in key areas such as military reconnaissance, environmental monitoring, and logistics transportation, the amount of data generated during their flight missions is experiencing explosive growth. This data is not only diverse in type but also significantly heterogeneous, encompassing structured sensor data, equipment operation logs and status monitoring information, semi-structured mission scheduling instructions and configuration files, as well as unstructured images, videos, and text.
[0003] Existing methods typically store massive amounts of multi-source heterogeneous data in data lakes. However, the sheer volume of multi-source heterogeneous data results in high internal storage complexity within the data lake, which can easily lead to "data swamps." Data query requests for specific needs require traversing and querying the data stored in the data lake, resulting in low data query efficiency. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and storage medium for hierarchical aggregation of multi-source heterogeneous data, which addresses the shortcomings of low data query efficiency in existing technologies when querying data for specific needs in unmanned helicopter scenarios, thereby improving data query efficiency.
[0005] This invention provides a method for hierarchical aggregation of multi-source heterogeneous data, comprising the following steps: Acquire multi-source heterogeneous data during the flight of unmanned helicopters; Based on the dimensions of data format, usage popularity, and time, the multiple categories of information contained in the multi-source heterogeneous data are determined. The multi-source heterogeneous data is stored in the storage layer of the data lake, and based on the multiple category information, the tag information of the multi-source heterogeneous data is stored in the corresponding level of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools built in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0006] According to a method for hierarchical aggregation of multi-source heterogeneous data provided by the present invention, the process of determining the tag information of the multi-source heterogeneous data includes: Based on the data type of the multi-source heterogeneous data, an extraction method for text extraction from the multi-source heterogeneous data is determined; Based on the extraction method, text extraction is performed on the multi-source heterogeneous data to obtain the text information of the multi-source heterogeneous data; Keywords are extracted from the text information based on term frequency-inverse document frequency, and the tag information of the multi-source heterogeneous data is determined based on the extracted keywords.
[0007] According to a hierarchical aggregation method for multi-source heterogeneous data provided by the present invention, the method for determining multiple categories of information contained in the multi-source heterogeneous data based on data format, usage frequency, and time dimensions includes: When the original metadata of the multi-source heterogeneous data exists, the original metadata of the multi-source heterogeneous data is obtained, and the original metadata is identified according to the data format dimension, usage popularity dimension and time dimension to determine the multiple categories of information contained in the multi-source heterogeneous data. In the absence of original metadata for the multi-source heterogeneous data, the file extensions of the multi-source heterogeneous data are identified according to the dimensions of data format, usage frequency, and time to determine the multiple categories of information contained in the multi-source heterogeneous data.
[0008] According to the multi-source heterogeneous data hierarchical aggregation method provided by the present invention, after storing the multi-source heterogeneous data in the storage layer of the data lake, the method further includes: The consistency of the multi-source heterogeneous data is verified, and it is determined that the multi-source heterogeneous data verification is successful.
[0009] According to the present invention, a method for hierarchical aggregation of multi-source heterogeneous data includes performing consistency verification on the multi-source heterogeneous data, comprising: Based on multiple categories of information from multi-source heterogeneous data, a verification method for the multi-source heterogeneous data is determined from the data verification rules. The data verification rules include multi-source data consistency verification, cross-layer data consistency verification, and time-series data consistency verification. The multi-source data consistency verification is used to verify based on the data source, the cross-layer data consistency verification is used to verify based on the data processing flow, and the time-series data consistency verification is used to verify based on the data time attribute. Based on the aforementioned verification method, consistency verification is performed on the multi-source heterogeneous data.
[0010] The multi-source heterogeneous data hierarchical aggregation method provided by the present invention further includes: Receive a data query request for target data from the unmanned helicopter, wherein the target data is used to perform a requirements analysis process; In response to the data query request, and based on the category information and hierarchical information of the target data, the target data is queried from multiple data pools of the data lake.
[0011] The multi-source heterogeneous data hierarchical aggregation method provided by the present invention further includes: Receive a data acquisition request from the unmanned helicopter for the data to be acquired; In response to the data acquisition request, the storage location of the data to be acquired is determined based on the metadata federated directory, and the data to be acquired is acquired from the storage location. The metadata federated directory is determined based on the directory information of all data stored in the data lake. The directory information includes the data's tag information, the hierarchical information in the data pool, and the location information in the storage layer.
[0012] The present invention also provides a multi-source heterogeneous data hierarchical aggregation device, comprising the following modules: The data acquisition module is used to acquire multi-source heterogeneous data during the flight of the unmanned helicopter; The category determination module is used to determine multiple category information contained in the multi-source heterogeneous data based on the data format dimension, usage popularity dimension, and time dimension. The storage module is used to store the multi-source heterogeneous data to the storage layer of the data lake, and based on the multiple category information, store the tag information of the multi-source heterogeneous data to the corresponding level in different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools constructed in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the multi-source heterogeneous data hierarchical aggregation method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-source heterogeneous data hierarchical aggregation method as described above.
[0015] The present invention provides a method, apparatus, device, and storage medium for hierarchical aggregation of multi-source heterogeneous data. By determining multiple categories of information contained in multi-source heterogeneous data based on data format, usage frequency, and time dimensions, the multi-source heterogeneous data is stored in the storage layer of a data lake. Based on these multiple categories, the tag information of the multi-source heterogeneous data is stored in corresponding levels within different data pools. This tag-based storage method simplifies the process by requiring only the location of the corresponding level for a specific device within the time-series data pool, and then filtering the tag information for a specific time period at that level. This avoids invalid scanning of other raw data in the data lake, significantly improving query speed and efficiency. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the multi-source heterogeneous data hierarchical aggregation method provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the data storage architecture provided by the present invention.
[0019] Figure 3 This is a schematic diagram of the structure of the multi-source heterogeneous data hierarchical aggregation device provided by the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0022] Figure 1 This is a flowchart illustrating the multi-source heterogeneous data hierarchical aggregation method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 110: Acquire multi-source heterogeneous data during the flight of the unmanned helicopter; Step 120: Based on the data format dimension, usage popularity dimension, and time dimension, determine the multiple categories of information contained in the multi-source heterogeneous data; Step 130: Store the multi-source heterogeneous data in the storage layer of the data lake, and based on the multiple category information, store the tag information of the multi-source heterogeneous data in the corresponding levels of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools constructed in the data lake. The levels in the data pools are determined based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0023] The execution subject of the multi-source heterogeneous data hierarchical aggregation method provided by this invention can be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, a mobile electronic device can be a tablet computer, laptop computer, handheld computer, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., while a non-mobile electronic device can be a server, network attached storage (NAS), or personal computer (PC), etc. This invention does not impose specific limitations.
[0024] The following example, using a computer executing the multi-source heterogeneous data hierarchical aggregation method provided by this invention, illustrates the technical solution of this invention in detail.
[0025] In step 110, multi-source heterogeneous data during the flight of the unmanned helicopter are acquired.
[0026] Unmanned helicopters generate various types of data during flight, and this data typically has the following characteristics: Multi-source: from different sensors and systems, such as Global Positioning System (GPS), Inertial Measurement Unit (IMU), cameras, radar, weather sensors, engine condition monitoring, etc.
[0027] Heterogeneity: The data formats are diverse, including structured data (such as CSV and JSON formats), semi-structured data (such as Extensible Markup Language (XML) and log files), and unstructured data (such as images, videos, and audio).
[0028] Dynamic: Data is generated at a fast rate, and some data has time-series characteristics (such as sensor readings changing over time).
[0029] Hierarchical structure: The data may come from different subsystems or devices of the unmanned helicopter and have a certain hierarchical relationship (such as fuselage level, sensor level, mission level).
[0030] Because the data acquired during unmanned helicopter missions or reconnaissance requires extensive analysis, it is necessary to store the acquired data in a data lake to facilitate subsequent analysis.
[0031] It should be noted that a data lake is a storage architecture that can store large amounts of structured, semi-structured, and unstructured data. It breaks the limitations of traditional data warehouses on data format and structure, enabling the storage of massive amounts of data at a lower cost and supporting a variety of data processing and analysis tools.
[0032] In step 120, based on the data format dimension, usage popularity dimension, and time dimension, the multiple categories of information contained in the multi-source heterogeneous data are determined.
[0033] After acquiring multi-source heterogeneous data during the flight of the unmanned helicopter, the acquired multi-source heterogeneous data is analyzed in three dimensions: data format, usage popularity, and time, to determine the multiple categories of information contained in the multi-source heterogeneous data.
[0034] Specifically, after analysis based on three dimensions, the identified category information can include structured data categories, semi-structured data categories, unstructured data categories, hot data categories, cold data categories, and time-series data categories.
[0035] It is understandable that image data acquired by cameras on unmanned helicopters can include unstructured data categories, hot data categories, and time categories. Therefore, acquired multi-source heterogeneous data typically includes information from multiple categories.
[0036] Structured data is data with a defined structure and format, such as records in a flight parameter database, which can be labeled as "structured data".
[0037] Semi-structured data can be categorized as similar to XML configuration files or JSON-formatted communication protocol messages. They possess a certain structure, but it is not as strictly defined as structured data.
[0038] Unstructured data can include images, videos, audio, LiDAR point clouds, and other data categorized as "unstructured data." The processing and analysis of this data is relatively complex, requiring specialized tools and techniques, such as image recognition algorithms and point cloud processing software, to extract valuable information.
[0039] Hot data is data that is frequently accessed and used during the flight of an unmanned helicopter, such as real-time flight attitude data and control command feedback data.
[0040] The "cold data" category refers to data that is rarely accessed during the flight of an unmanned helicopter, such as backup data.
[0041] Time series data can be categorized based on the time information at which the data was generated.
[0042] In step 130, the multi-source heterogeneous data is stored in the storage layer of the data lake, and the category tags are stored in the corresponding levels of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools constructed in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0043] Multi-source heterogeneous data is stored in the storage layer of the data lake. The storage layer is the physical storage location for the data, providing a unified storage space for different types of data while ensuring data reliability and scalability. The storage layer can select appropriate storage media based on the characteristics of the data, such as high-speed disk arrays, distributed file systems, or object storage.
[0044] Multiple data pools are pre-built in the data lake to store the tag information determined in step 120. Specifically, the tag information can be determined based on important keywords in multi-source heterogeneous data. Specifically, it can be based on term frequency-inverse document frequency analysis of multi-source heterogeneous data to determine the corresponding keywords.
[0045] The data pools built in a data lake specifically include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools.
[0046] Each constructed data pool is layered, resulting in multiple different levels within the data pool. Specifically, the data acquisition equipment of the unmanned helicopter has a hierarchical relationship, which reflects the organizational structure, functional dependencies, and direction of data flow between the equipment.
[0047] For example, data acquisition for unmanned helicopters can include a sensor layer, a control data layer, and a processing data layer. The next layer after the sensor layer can specifically include optical sensors, lidar, infrared thermal imagers, and weather sensors. The next layer after optical sensors can include high-resolution cameras, video cameras, and low-light / night vision cameras.
[0048] Based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter, each data pool is layered. After layering, the determined tag information is stored in the corresponding layer within the data pool. For example, in the case of acquiring image data collected by video cameras during the unmanned helicopter's flight, the acquired image data is analyzed based on data format, usage frequency, and time dimensions to determine whether the image data falls into the category of unstructured data, hot data, and data acquisition time.
[0049] The tag information of image data can be stored in the video camera layer in the unstructured data pool, the tag information in the video camera layer in the hot data pool, and the tag information in the video camera layer in the time-series data pool.
[0050] It's important to note that data analysis of unmanned helicopters often requires acquiring large amounts of data of a specific type. Since the data is already categorized and stored according to type and hierarchy, queries only need to scan the tags within a specific data pool hierarchy, rather than performing a full scan of the entire data lake. For example, to analyze data from a specific device within a certain time period, we only need to find the corresponding hierarchy for that device in the time-series data pool and filter the tag information for that specific time period within that hierarchy. This avoids invalid scanning of other raw data in the data lake, significantly improving query speed and efficiency.
[0051] Based on the storage method of various categorized tag information, data can be quickly located and retrieved using the tag information. For example, when it is necessary to find real-time data acquired by sensors related to flight safety, one only needs to search for the corresponding level in the corresponding data pool according to the relevant tag. This efficient retrieval mechanism can quickly respond to data query requests, providing timely data support for applications such as real-time monitoring and fault diagnosis of unmanned helicopters.
[0052] Storage based on categories allows for finer-grained access control. For example, different access permissions can be assigned to different users and systems based on data popularity, data format, or other categories. Only users with the appropriate permissions can access data in a specific data pool, thereby preventing unauthorized access and data misuse, and further enhancing data security.
[0053] The multi-source heterogeneous data hierarchical aggregation method provided by this invention determines multiple categories of information contained in multi-source heterogeneous data based on data format, usage frequency, and time dimensions. It then stores the multi-source heterogeneous data in the storage layer of a data lake and, based on these categories, stores the tag information of the multi-source heterogeneous data in corresponding levels within different data pools. This tag-based storage method only requires finding the level corresponding to a specific device in the time-series data pool and filtering the tag information for a specific time period within that level. This avoids invalid scanning of other raw data in the data lake, significantly improving query speed and efficiency.
[0054] In one embodiment, the process of determining the tag information of multi-source heterogeneous data includes: determining an extraction method for text extraction of the multi-source heterogeneous data based on the data type of the multi-source heterogeneous data; performing text extraction on the multi-source heterogeneous data based on the extraction method to obtain text information of the multi-source heterogeneous data; extracting keywords from the text information based on word frequency-inverse document frequency, and determining the tag information of the multi-source heterogeneous data based on the extracted keywords.
[0055] When dealing with multi-source heterogeneous data, direct analysis and application are challenging due to the wide range of data sources and diverse types, including structured, semi-structured, and unstructured data. Therefore, the primary task is to accurately determine the appropriate text extraction method based on the characteristics of these data types.
[0056] For semi-structured and unstructured data, such as text documents, PDF documents, and images, there are typically technical metadata and content metadata. Technical metadata includes information from the original metadata, such as file size, creation time, camera model, geographic location, and creation software. Content metadata for text documents and PDF documents requires keyword extraction and named entity recognition using natural language processing. For image data, content metadata uses object detection and scene classification algorithms to extract the contained text information. For audio and video data, content metadata uses a combination of speech-to-text and natural language processing to extract the contained text information.
[0057] To extract key information from these texts and facilitate data classification, retrieval, and understanding, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is used for keyword extraction. The TF-IDF algorithm effectively measures the importance of words to the document's topic by calculating the frequency (TF) of a word in a document and its inverse document frequency (IDF) in the entire document set, thereby selecting the most representative keywords.
[0058] Based on these keywords, we can further determine the label information of multi-source heterogeneous data. These labels not only reflect the main content or characteristics of the data, but also provide strong support for the automated classification of data, the construction of recommendation systems and knowledge graphs, greatly improving the management efficiency and application value of data.
[0059] In one embodiment, determining multiple categories of information contained in the multi-source heterogeneous data based on data format, usage popularity, and time dimensions includes: when the multi-source heterogeneous data has original metadata, acquiring the original metadata of the multi-source heterogeneous data, and identifying the original metadata according to the data format, usage popularity, and time dimensions to determine multiple categories of information contained in the multi-source heterogeneous data; when the multi-source heterogeneous data does not have original metadata, identifying the file extension, content, attributes, and associated information of the multi-source heterogeneous data according to the data format, usage popularity, and time dimensions to determine multiple categories of information contained in the multi-source heterogeneous data.
[0060] When processing multi-source heterogeneous data, if original metadata exists, the first step is to actively acquire the metadata information. Metadata typically includes core attributes such as the data's source, format, creation time, and modification history. Subsequently, based on three key dimensions—data format (e.g., structured, semi-structured, unstructured), usage frequency (measured by metrics such as access frequency and citation count), and time (e.g., data generation time and latest update time)—the original metadata is identified and analyzed. This process determines the multiple categories of information covered by the multi-source heterogeneous data, achieving preliminary structured classification of the data.
[0061] When multi-source heterogeneous data lacks original metadata, the file extension, which is readily available from the data itself, can be used as a starting point for classification. Similarly, multi-source heterogeneous data can be categorized based on data format (inferred directly from file extensions), usage frequency (estimated indirectly through log analysis, user behavior statistics, etc.), and time (such as file modification timestamps).
[0062] In one embodiment, after storing the multi-source heterogeneous data in the storage layer of the data lake, the method further includes: performing a consistency check on the multi-source heterogeneous data and determining that the multi-source heterogeneous data has passed the check.
[0063] The consistency verification of the multi-source heterogeneous data includes: Based on multiple categories of information from multi-source heterogeneous data, a verification method for the multi-source heterogeneous data is determined from the data verification rules. The data verification rules include multi-source data consistency verification, cross-layer data consistency verification, and time-series data consistency verification. The multi-source data consistency verification is used to verify based on the data source, the cross-layer data consistency verification is used to verify based on the data processing flow, and the time-series data consistency verification is used to verify based on the data time attribute. Based on the aforementioned verification method, consistency verification is performed on the multi-source heterogeneous data.
[0064] The consistency verification module includes multi-source data consistency verification, cross-layer data consistency verification, and time-series data consistency verification. Multi-source data consistency verification addresses data source issues, cross-layer data consistency verification addresses data processing flow issues, and time-series data consistency verification addresses data time attribute issues. One or more combinations of these methods can be selected based on business needs to form a comprehensive data verification mechanism, improving verification efficiency and accuracy.
[0065] Optionally, multi-source data consistency verification can ensure that multi-source data has unified semantics and consistency, avoiding conflicts or redundancy. Multi-source data consistency verification includes the following verification rules: Format and type validation: During the data loading process, you can use the built-in type conversion and validation functions of the ETL (Extract, Transform, Load) tool to check the data type, use regular expressions to validate strings with specific formats, validate the valid range of predefined field values, use enumeration values to validate the list of predefined allowed values, use NOT NULL validation to check whether required fields are empty, and use length validation to check the length of strings or numbers.
[0066] Field mapping validation: Use a field mapping table to unify the differences in field names from different sources into standard fields; Primary key / uniqueness verification: Based on the database's PRIMARY KEY or UNIQUE constraint, uniqueness verification is performed on fields with unique identifiers to avoid data redundancy caused by duplicate insertions.
[0067] Conflict detection rules: Records from different sources that point to the same entity are grouped based on their unique identifiers. Within each group, all common fields with non-unique identifiers that exist in multiple sources are traversed, and their values are compared. If different values are detected for the same field of the same entity in different sources, a record conflict is identified. Conflicts are resolved using subsequent priority rules and data merging rules.
[0068] Priority rule: The value of the data source with higher credibility among conflicting data is selected based on factors such as the weight of the data source and the historical accuracy of the data source.
[0069] Data merging rules: When priority rules are insufficient to resolve conflicts, or when it is desirable to utilize information from multiple sources, a voting method is used to select the value that appears most frequently, or the value of the most recently timestamped record is selected. Conflicts that cannot be resolved automatically are stored in a special structure for manual review.
[0070] The above validation rules are typically applied sequentially or in combination within an ETL or data integration process. First, the basic format and uniqueness are identified and validated. Then, conflicts are detected. Finally, conflicts are resolved according to preset priorities and merging strategies, with the goal of generating a clean, consistent, and reliable dataset.
[0071] Cross-layer data consistency verification integrates data from different system layers to avoid information silos and synchronization delays. Optionally, cross-layer data consistency verification includes the following verification rules: Hierarchical association verification: Foreign key / association key checks, record count comparisons, sampling traceability, and business rule verification are used to ensure that the mapping relationship between data at each level is correct.
[0072] Synchronization verification: By calculating the checksum or hash value of the data batches sent at the source end, and recalculating and comparing it after receiving it at the target end, data loss during cross-layer transmission is checked; by recording and comparing the timestamp of data leaving the source layer with the timestamp of data arriving at the target layer, the delay during cross-layer data transmission is checked.
[0073] Data conversion verification: Write unit tests for functions or modules that perform data conversion, using predefined input data and expected output to verify the correctness of the conversion logic.
[0074] Data compensation rules: In the event of data loss or delay, a compensation mechanism is adopted to repair the missing data. The compensation mechanism is implemented by re-requesting or data interpolation. Idempotency verification: Idempotency is ensured in cross-layer data fusion operations through unique identifiers or deduplication tables. Idempotency means that identical data will not cause errors due to repeated processing. The unique identifier is a globally unique identifier (such as a UUID or business primary key) generated for each data operation. Before processing, it is checked whether the identifier has been recorded. If it already exists, the operation is skipped or an existing result is returned. The deduplication table maintains a table of processed requests in the database or cache. Before each operation, the table is checked to see if a corresponding record exists. If it exists, it is considered a duplicate request.
[0075] Time-series data consistency verification aligns data from different data sources along the time dimension to avoid inconsistencies caused by latency and time zone differences. Optionally, time-series data consistency verification can include the following verification rules: Timestamp alignment rules: NTP (Network Time Protocol) is used to ensure system time synchronization across data sources. To address the drift issue of IoT devices or mobile devices, logical clocks or vector clocks are used to resolve timestamp inconsistencies at cross-boundary points.
[0076] Time sequence verification: Check whether the time sequence of multi-source data with a unified time base according to timestamp alignment rules conforms to the order in which events occurred.
[0077] Delay tolerance and expired data verification: Set delay tolerance for different application scenarios, and discard or mark data that exceeds the delay threshold as expired.
[0078] Lost data compensation: If time series data is missing for a certain period of time, compensation is made through interpolation and historical data backfilling.
[0079] In one embodiment, the method further includes: receiving a data query request for target data in the unmanned helicopter, the target data being used to perform a requirements analysis process; and, in response to the data query request, querying the target data from multiple data pools of the data lake based on the category information and hierarchical information of the target data.
[0080] In unmanned helicopter applications, data querying is a crucial step in subsequent data analysis. Upon receiving a query request for target data, the system responds by locating and querying the target data within multiple partitioned data pools in the data lake, based on the category and hierarchical information obtained from prior analysis of the target data (specifically, the hierarchy of the data within the data pool, which can be determined by the hierarchy of the data acquisition equipment within the unmanned helicopter).
[0081] In one embodiment, the method further includes: receiving a data acquisition request for data to be acquired from the unmanned helicopter; responding to the data acquisition request, determining the storage location of the data to be acquired based on a metadata federated directory, and acquiring the data to be acquired from the storage location, wherein the metadata federated directory is determined based on the directory information of all data stored in the data lake, and the directory information includes the data's tag information, the hierarchical information in the data pool, and the location information in the storage layer.
[0082] Based on the multi-source heterogeneous data hierarchical aggregation architecture provided in this embodiment of the invention, the data to be acquired can be obtained as follows: Figure 2 The data storage architecture provided by this invention is shown in the schematic diagram.
[0083] First, data is collected, including multi-source heterogeneous data and its raw metadata.
[0084] The category labels of the collected multi-source heterogeneous data are hierarchically aggregated and then aggregated into the corresponding data pool.
[0085] A pre-built metadata federated catalog is established. A unified metadata catalog is constructed for data in structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools. The raw metadata is summarized, including data source, creation time, filename, and the data pool category to which the second step belongs. For data that cannot be parsed in semi-structured and unstructured data pools, such as text documents, PDF documents, and images, technical metadata and content metadata are typically included. Technical metadata includes file size, creation time, camera model, geographic location, and creation software from the raw metadata. Content metadata for text documents and PDF documents requires keyword extraction and named entity recognition using natural language processing. Image data content metadata uses object detection and scene classification algorithms to extract the contained content entities. Audio and video data content metadata uses a combination of speech-to-text and natural language processing to extract the contained content entities. Data in both hot and cold data pools can be structured, semi-structured, and unstructured; therefore, the metadata analysis and extraction methods are the same.
[0086] After registration in the metadata federated directory, consistency checks are performed on multi-source heterogeneous data. These consistency checks specifically include multi-source data consistency checks, cross-layer data consistency checks, and time-series data consistency checks.
[0087] The metadata federated directory stores the location index of the data to be retrieved, so the storage location of the data can be determined by reading the metadata federated directory. After determining the storage location, the data to be retrieved is obtained from the storage location, and data analysis or applications are performed based on the retrieved data.
[0088] The multi-source heterogeneous data hierarchical aggregation device provided by the present invention is described below. The multi-source heterogeneous data hierarchical aggregation device described below can be referred to in correspondence with the multi-source heterogeneous data hierarchical aggregation method described above.
[0089] like Figure 3 As shown, the device includes: Data acquisition module 310 is used to acquire multi-source heterogeneous data during the flight of the unmanned helicopter; The category determination module 320 is used to determine multiple category information contained in the multi-source heterogeneous data based on the data format dimension, usage popularity dimension, and time dimension. Storage module 330 is used to store the multi-source heterogeneous data to the storage layer of the data lake, and based on the multiple category information, store the tag information of the multi-source heterogeneous data to the corresponding level in different data pools. The different data pools include structured data pool, semi-structured data pool, unstructured data pool, hot data pool, cold data pool and time-series data pool constructed in the data lake. The level in the data pool is obtained based on the hierarchical relationship of the data acquisition equipment in the unmanned helicopter.
[0090] In one embodiment, the storage module 330 is specifically used for: The process of determining the label information of the multi-source heterogeneous data includes: Based on the data type of the multi-source heterogeneous data, an extraction method for text extraction from the multi-source heterogeneous data is determined; Based on the extraction method, text extraction is performed on the multi-source heterogeneous data to obtain the text information of the multi-source heterogeneous data; Keywords are extracted from the text information based on term frequency-inverse document frequency, and the tag information of the multi-source heterogeneous data is determined based on the extracted keywords.
[0091] In one embodiment, the category determination module 320 is specifically used for: The determination of multiple categories of information contained in the multi-source heterogeneous data based on data format, usage popularity, and time dimensions includes: When the original metadata of the multi-source heterogeneous data exists, the original metadata of the multi-source heterogeneous data is obtained, and the original metadata is identified according to the data format dimension, usage popularity dimension and time dimension to determine the multiple categories of information contained in the multi-source heterogeneous data. In the absence of original metadata for the multi-source heterogeneous data, the file extensions, content, attributes, and associated information of the multi-source heterogeneous data are identified according to the dimensions of data format, usage popularity, and time, thereby determining the multiple categories of information contained in the multi-source heterogeneous data.
[0092] In one embodiment, the storage module 330 is further configured to: After storing the multi-source heterogeneous data in the storage layer of the data lake, the method further includes: The consistency of the multi-source heterogeneous data is verified, and it is determined that the multi-source heterogeneous data verification is successful.
[0093] In one embodiment, the storage module 330 is further configured to: The consistency verification of the multi-source heterogeneous data includes: Based on multiple categories of information from multi-source heterogeneous data, a verification method for the multi-source heterogeneous data is determined from the data verification rules. The data verification rules include multi-source data consistency verification, cross-layer data consistency verification, and time-series data consistency verification. The multi-source data consistency verification is used to verify based on the data source, the cross-layer data consistency verification is used to verify based on the data processing flow, and the time-series data consistency verification is used to verify based on the data time attribute. Based on the aforementioned verification method, consistency verification is performed on the multi-source heterogeneous data.
[0094] In one embodiment, the storage module 330 is further configured to: Receive a data query request for target data from the unmanned helicopter, wherein the target data is used to perform a requirements analysis process; In response to the data query request, and based on the category information and hierarchical information of the target data, the target data is queried from multiple data pools of the data lake.
[0095] In one embodiment, the storage module 330 is further configured to: Receive a data acquisition request from the unmanned helicopter for the data to be acquired; In response to the data acquisition request, the storage location of the data to be acquired is determined based on the metadata federated directory, and the data to be acquired is acquired from the storage location. The metadata federated directory is determined based on the directory information of all data stored in the data lake. The directory information includes the data's tag information, the hierarchical information in the data pool, and the location information in the storage layer.
[0096] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions from the memory 430 to execute a multi-source heterogeneous data hierarchical aggregation method, which includes acquiring multi-source heterogeneous data during the flight of an unmanned helicopter. Based on the dimensions of data format, usage popularity, and time, the multiple categories of information contained in the multi-source heterogeneous data are determined. The multi-source heterogeneous data is stored in the storage layer of the data lake, and based on the multiple category information, the tag information of the multi-source heterogeneous data is stored in the corresponding level of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools built in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0097] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the multi-source heterogeneous data hierarchical aggregation method provided by the above methods, the method including: acquiring multi-source heterogeneous data during the flight of an unmanned helicopter; Based on the dimensions of data format, usage popularity, and time, the multiple categories of information contained in the multi-source heterogeneous data are determined. The multi-source heterogeneous data is stored in the storage layer of the data lake, and based on the multiple category information, the tag information of the multi-source heterogeneous data is stored in the corresponding level of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools built in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for hierarchical aggregation of multi-source heterogeneous data provided by the above methods, the method comprising: acquiring multi-source heterogeneous data during the flight of an unmanned helicopter; Based on the dimensions of data format, usage popularity, and time, the multiple categories of information contained in the multi-source heterogeneous data are determined. The multi-source heterogeneous data is stored in the storage layer of the data lake, and based on the multiple category information, the tag information of the multi-source heterogeneous data is stored in the corresponding level of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools built in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hierarchical aggregation method for multi-source heterogeneous data, characterized in that, include: Acquire multi-source heterogeneous data during the flight of unmanned helicopters; Based on the dimensions of data format, usage popularity, and time, the multiple categories of information contained in the multi-source heterogeneous data are determined. The multi-source heterogeneous data is stored in the storage layer of the data lake, and based on the multiple category information, the tag information of the multi-source heterogeneous data is stored in the corresponding level of different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools built in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
2. The multi-source heterogeneous data hierarchical aggregation method according to claim 1, characterized in that, The process of determining the label information of the multi-source heterogeneous data includes: Based on the data type of the multi-source heterogeneous data, an extraction method for text extraction from the multi-source heterogeneous data is determined; Based on the extraction method, text extraction is performed on the multi-source heterogeneous data to obtain the text information of the multi-source heterogeneous data; Keywords are extracted from the text information based on term frequency-inverse document frequency, and the tag information of the multi-source heterogeneous data is determined based on the extracted keywords.
3. The multi-source heterogeneous data hierarchical aggregation method according to claim 1, characterized in that, The determination of multiple categories of information contained in the multi-source heterogeneous data based on data format, usage popularity, and time dimensions includes: When the original metadata of the multi-source heterogeneous data exists, the original metadata of the multi-source heterogeneous data is obtained, and the original metadata is identified according to the data format dimension, usage popularity dimension and time dimension to determine the multiple categories of information contained in the multi-source heterogeneous data. In the absence of original metadata for the multi-source heterogeneous data, the file extensions of the multi-source heterogeneous data are identified according to the dimensions of data format, usage frequency, and time to determine the multiple categories of information contained in the multi-source heterogeneous data.
4. The multi-source heterogeneous data hierarchical aggregation method according to claim 1, characterized in that, After storing the multi-source heterogeneous data in the storage layer of the data lake, the method further includes: The consistency of the multi-source heterogeneous data is verified, and it is determined that the multi-source heterogeneous data verification is successful.
5. The multi-source heterogeneous data hierarchical aggregation method according to claim 4, characterized in that, The consistency verification of the multi-source heterogeneous data includes: Based on multiple categories of information from multi-source heterogeneous data, a verification method for the multi-source heterogeneous data is determined from the data verification rules. The data verification rules include multi-source data consistency verification, cross-layer data consistency verification, and time-series data consistency verification. The multi-source data consistency verification is used to verify based on the data source, the cross-layer data consistency verification is used to verify based on the data processing flow, and the time-series data consistency verification is used to verify based on the data time attribute. Based on the aforementioned verification method, consistency verification is performed on the multi-source heterogeneous data.
6. The multi-source heterogeneous data hierarchical aggregation method according to claim 1, characterized in that, Also includes: Receive a data query request for target data from the unmanned helicopter, wherein the target data is used to perform a requirements analysis process; In response to the data query request, and based on the category information and hierarchical information of the target data, the target data is queried from multiple data pools of the data lake.
7. The multi-source heterogeneous data hierarchical aggregation method according to claim 1, characterized in that, Also includes: Receive a data acquisition request from the unmanned helicopter for the data to be acquired; In response to the data acquisition request, the storage location of the data to be acquired is determined based on the metadata federated directory, and the data to be acquired is acquired from the storage location. The metadata federated directory is determined based on the directory information of all data stored in the data lake. The directory information includes the data's tag information, the hierarchical information in the data pool, and the location information in the storage layer.
8. A multi-source heterogeneous data hierarchical aggregation device, characterized in that, include: The data acquisition module is used to acquire multi-source heterogeneous data during the flight of the unmanned helicopter; The category determination module is used to determine multiple category information contained in the multi-source heterogeneous data based on the data format dimension, usage popularity dimension, and time dimension. The storage module is used to store the multi-source heterogeneous data to the storage layer of the data lake, and based on the multiple category information, store the tag information of the multi-source heterogeneous data to the corresponding level in different data pools. The different data pools include structured data pools, semi-structured data pools, unstructured data pools, hot data pools, cold data pools, and time-series data pools constructed in the data lake. The levels in the data pools are divided based on the hierarchical relationship of the data acquisition devices in the unmanned helicopter.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the multi-source heterogeneous data hierarchical aggregation method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-source heterogeneous data hierarchical aggregation method as described in any one of claims 1 to 7.