Intelligent terminal and method for supporting distributed storage and retrieval of heterogeneous data
Through modular design and intelligent processing, the storage incompatibility and slow retrieval of traditional smart terminals when processing heterogeneous data is solved, and efficient and secure heterogeneous data storage and retrieval is achieved, suitable for big data analysis and the Internet of Things.
Patent Information
- Application Number
- CN202510413535.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional smart terminals are difficult to effectively process heterogeneous data, resulting in storage incompatibility, data loss, slow retrieval speed and low accuracy, and cannot meet the reliability and scalability needs of large data volumes.
It adopts a modular design, including data acquisition, preprocessing, distributed storage, multi-level indexing and security modules. Through cleaning, format conversion, load balancing and encryption methods, the work of each module is coordinated to achieve efficient storage and rapid retrieval.
It improves the processing efficiency and accuracy of heterogeneous data, optimizes storage performance and retrieval speed, ensures data security and privacy, and is suitable for big data analysis and Internet of Things applications.
Smart Images

Figure CN120353845A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage and retrieval, and more specifically, to an intelligent terminal and method for supporting distributed storage and retrieval of heterogeneous data. Background Art
[0002] With the rapid development of information technology, the amount of data generated by various industries has increased exponentially, and the types of data have become extremely rich, covering structured, semi-structured, and unstructured data, that is, heterogeneous data. Traditional intelligent terminals have many problems in processing these heterogeneous data.
[0003] In the data processing stage, due to the lack of a professional heterogeneous data processing module, it is difficult to effectively receive, parse, and convert data in different formats, resulting in format incompatibility during storage and retrieval, leading to data errors or losses, and seriously affecting the accuracy and integrity of data processing. In terms of storage, the traditional centralized storage method has poor reliability and scalability in the face of large amounts of data. Once the storage device fails, the data is at risk of loss, and the expansion cost is high and the efficiency is low. In the retrieval link, traditional retrieval algorithms do not fully consider data diversity and complex user query requirements, with slow retrieval speed and low accuracy, and cannot quickly locate the required information from a large amount of heterogeneous data, making it difficult to meet the rapidly developing digital application scenarios.
[0004] Therefore, how to provide an intelligent terminal and method that can efficiently process various types of data, achieve distributed storage and fast retrieval, and improve data processing efficiency and accuracy is an urgent problem for those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides an intelligent terminal and method for supporting distributed storage and retrieval of heterogeneous data. Through modular design, it realizes the cleaning, format conversion, and standardization processing of structured, semi-structured, and unstructured data, and uses technologies such as encryption, load balancing, multi-level indexing, and intelligent scheduling to optimize storage performance, improve retrieval efficiency, ensure data security, meet the needs of diverse application scenarios, and promote the intelligent development of data management.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] On the one hand, the present invention provides an intelligent terminal for supporting distributed storage and retrieval of heterogeneous data, characterized by including:
[0008] A data acquisition module for collecting heterogeneous data from multiple data sources, where the heterogeneous data includes structured data, semi-structured data, and unstructured data;
[0009] A data preprocessing module for cleaning, format conversion, and standardization of the collected heterogeneous data;
[0010] A distributed storage module for data sharding and load balancing of the preprocessed heterogeneous data, and distributed storage on multiple storage nodes;
[0011] A data indexing module for establishing multi-level indexes for the stored heterogeneous data;
[0012] A data retrieval module for performing fast data retrieval using the multi-level indexes according to the user query request;
[0013] A data security module for ensuring the security of data during storage and transmission by using encryption methods and access control mechanisms;
[0014] An intelligent scheduling module for coordinating the work between various modules.
[0015] Preferably, the data preprocessing module includes: a structured processing module, a semi-structured processing module, and an unstructured processing module;
[0016] The structured processing module includes:
[0017] A first data cleaning unit for processing missing values, outliers, and duplicate values in the structured data;
[0018] A first format conversion unit for converting the data type and encoding of the cleaned structured data;
[0019] A first standardization unit for standardizing the format-converted structured data using normalization;
[0020] The semi-structured processing module includes:
[0021] A second data cleaning unit for parsing the semi-structured data, extracting key information and data fields, and removing irrelevant information;
[0022] A second format conversion unit for converting the cleaned semi-structured data into a structured format;
[0023] A second standardization unit for standardizing the value range and data format of the format-converted data, and performing standardized classification and encoding on the classification information in the format-converted data, and converting the text-based classification data into digital codes;
[0024] The unstructured processing module includes:
[0025] A third data cleaning unit for removing and cleaning the noise in the unstructured data and filtering out irrelevant information;
[0026] A third format conversion unit, configured to perform text extraction on unstructured data and convert the extracted text into feature vectors;
[0027] A third normalization unit, which normalizes the extraction of the feature vectors.
[0028] Preferably, the distributed storage module includes:
[0029] A data processing module, configured to perform encryption and segmentation processing on the preprocessed heterogeneous data to obtain data shards;
[0030] A data location module, configured to predict the future access trend of the data shards based on the analysis of historical access data, and determine the storage location of the data shards according to the access trend;
[0031] A data storage module, configured to store the data shards into storage nodes according to the storage location;
[0032] A data balancing module, configured to perform data migration on the data shards stored in the storage nodes when the storage nodes are overloaded.
[0033] Preferably, the data processing module includes:
[0034] An encryption processing module, configured to perform encryption processing on the preprocessed heterogeneous data according to an encryption algorithm;
[0035] A data sharding processing unit, configured to perform sharding processing on the encrypted heterogeneous data to obtain N data shards, where each of the N data shards includes a shard identifier, a shard offset, and a shard flag bit.
[0036] Preferably, the data location module includes:
[0037] A relationship extraction unit, configured to construct a relationship of access volume change based on the historical access data; the relationship of access volume change is used to represent the change of access volume over time;
[0038] A feature extraction unit, configured to extract at least one access volume feature from the relationship of access volume change; the access volume feature is used to describe the change trend of the access volume; the access volume feature includes a feature point clustering coefficient and a change graph clustering coefficient; the data category corresponding to the historical access data is determined by using the access volume feature;
[0039] An access trend prediction unit, configured to predict the access trend of the data shards through the data category and the access volume feature;
[0040] A positioning unit, which is used to determine the storage location of the data shard according to the access trend.
[0041] Preferably, the data retrieval module includes:
[0042] A receiving module, which is used to receive a query request input by a user, and perform a preliminary parsing on the query request to identify a query type and key information;
[0043] An index level determination module, which is used to determine an index level according to the query type and the structure of the multi-level index;
[0044] A data positioning module, which is used to start from the determined index level and search for an index item matching the key information in the corresponding index structure;
[0045] A data feedback module, which is used to locate a data pointer satisfying a query condition according to the index item, determine a data storage location according to the pointer, and obtain data at the data storage location to form a complete query result and feedback it to the user.
[0046] On the other hand, the present invention provides a method for supporting distributed storage and retrieval of heterogeneous data, including the following steps:
[0047] Collect heterogeneous data from multiple data sources, where the heterogeneous data includes structured data, semi-structured data, and unstructured data;
[0048] Clean, format-convert, and standardize the collected heterogeneous data;
[0049] Perform data sharding and load balancing on the preprocessed heterogeneous data, and distribute and store it in multiple storage nodes;
[0050] Establish a multi-level index for the stored heterogeneous data;
[0051] According to a user query request, perform fast data retrieval by using the multi-level index;
[0052] Adopt an encryption method and an access control mechanism to ensure the security of data during storage and transmission;
[0053] Coordinate the work among various modules according to the workload and resource requirements of each module.
[0054] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an intelligent terminal and method for supporting distributed storage and retrieval of heterogeneous data. By adopting modular design and intelligent processing, it realizes efficient acquisition, preprocessing, distributed storage, rapid retrieval, and security protection of heterogeneous data. In data preprocessing, structured, semi-structured, and unstructured data are processed separately to improve data consistency and availability, laying a solid foundation for subsequent processes. In terms of storage, data sharding and load balancing strategies are used, combined with intelligent data location and migration mechanisms, to reasonably distribute data, improve scalability and performance, and reduce storage costs. When retrieving data, with the help of a multi-level index structure, the retrieved data can be quickly located to meet the large-scale data query requirements and improve the user experience. At the security level, encryption methods and access control mechanisms are adopted to ensure data confidentiality, integrity, and availability, preventing data leakage and unauthorized access. Through the intelligent scheduling module, the work of each module is dynamically coordinated, and the resource configuration is automatically adjusted according to the system load and resource requirements to achieve intelligent management and optimization, improving the system operation efficiency and stability. Moreover, the invention can adapt to diverse application scenarios such as big data analysis and Internet of Things data processing, and has significant advantages in improving the efficiency of heterogeneous data processing and optimizing storage performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0056] Figure 1 Structural schematic diagram provided by the present invention;
[0057] Figure 2 Specific structural schematic diagrams of the data processing module and the data location module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0059] An embodiment of the present invention discloses an intelligent terminal for supporting distributed storage and retrieval of heterogeneous data, which is characterized in that, as Figure 1 shown, it includes:
[0060] A data acquisition module for collecting heterogeneous data from multiple data sources, where the heterogeneous data includes structured data, semi-structured data, and unstructured data;
[0061] A data preprocessing module for cleaning, format conversion, and standardization of the collected heterogeneous data to facilitate subsequent storage and retrieval.
[0062] A distributed storage module for data sharding and load balancing of the preprocessed heterogeneous data, and distributed storage on multiple storage nodes;
[0063] A data indexing module for establishing multi-level indexes for the stored heterogeneous data, including full-text indexes, inverted indexes, and spatial indexes to support efficient data retrieval.
[0064] A data retrieval module for quickly retrieving data using the established multi-level indexes according to user query requests, and supporting multiple query methods, such as keyword query, range query, and spatial query.
[0065] A data security module for ensuring the security of data during storage and transmission by using encryption methods and access control mechanisms;
[0066] An intelligent scheduling module for coordinating the work between various modules.
[0067] Furthermore, the data preprocessing module includes: a structured processing module, a semi-structured processing module, and an unstructured processing module;
[0068] The structured processing module includes:
[0069] A first data cleaning unit for handling missing values, outliers, and duplicate values in structured data. Specifically:
[0070] First, identify the missing values in the dataset. You can quickly locate which columns have missing values by counting the number of non-null values in each column. For a small number of missing values, if it is numerical data, statistical measures such as mean, median, and mode can be used for filling; if it is categorical data, the most frequently occurring category can be used for filling. For a large number of missing values, it is necessary to evaluate the importance of the data column. If it is not very important, the column can be directly deleted. If it is important, machine learning algorithms such as multiple imputation methods based on decision trees can be considered for filling.
[0071] For the handling of outliers, methods such as box plots and the 3σ principle are used to identify outliers. For obvious incorrect outliers, such as extremely large or small values caused by data entry errors, they can be directly deleted or corrected. For data that may be real but deviate greatly, according to business logic and data characteristics, the winsorization method can be used to adjust it to a reasonable boundary value.
[0072] For the processing of duplicate values, use the query statements or data processing tools of the database to find the duplicate records in the dataset, and then keep one of them according to the business requirements and delete the remaining duplicate records.
[0073] The first format conversion unit is used to perform data type conversion and encoding conversion on the cleaned structured data.
[0074] For data type conversion, convert the data type according to the actual meaning of the data and the subsequent analysis requirements. For example, convert the date field from string type to date and time type for date-related calculations and analysis; convert the classification field from numeric type to character type to enhance the readability of the data.
[0075] For encoding conversion, if the data has different encoding formats, such as UTF-8, GBK, etc., it is necessary to convert them to a unified encoding format to prevent garbled problems and affect data processing and analysis.
[0076] The first standardization unit is used to standardize the structured data after format conversion by normalization; specifically, map the data to the interval [0,1] or [-1,1], and the commonly used method is min-max normalization.
[0077] The semi-structured processing module includes:
[0078] The second data cleaning unit is used to parse the semi-structured data, extract key information and data fields, and remove irrelevant information.
[0079] For semi-structured data, such as data in XML and JSON formats, first parse the data, extract the key information and data fields in it, and remove irrelevant tags, metadata, etc.
[0080] For the processing of inconsistencies, check for problems such as format inconsistencies and semantic inconsistencies in the data. For example, in JSON data, there may be cases where fields with the same meaning have different names in different records, which need to be unified.
[0081] The second format conversion unit is used to convert the cleaned semi-structured data into a structured format; convert the semi-structured data into a structured format suitable for analysis, such as converting JSON data into a relational database table structure, or converting XML data into CSV format, which is convenient for subsequent processing using database tools or data analysis tools. If there are semi-structured data from multiple sources, it is necessary to integrate them into a unified format and merge and align the same or similar data fields.
[0082] A second standardization unit, which is used to standardize the value range and data format of the data after format conversion, and to standardize the classification and encoding of the classification information in the data after format conversion, converting the text-type classification data into digital codes;
[0083] The unstructured processing module includes:
[0084] A third data cleaning unit, which is used to remove and clean the noise in the unstructured data and filter out irrelevant information; for the noise in the text data, such as garbled characters, special characters, HTML tags, etc., use tools such as regular expressions to remove and clean, and retain the clean text content. Filter out irrelevant and low-quality data according to certain rules and conditions. For example, for a large amount of web page text data, filter out irrelevant information such as advertising content, navigation bars, copyright statements, etc.
[0085] A third format conversion unit, which is used to extract text from the unstructured data and convert the extracted text into feature vectors; for unstructured document data, such as PDF, Word documents, etc., use corresponding text extraction tools to extract the text content therein and convert it into plain text format for subsequent processing and analysis. Further, convert the unstructured data into the form of feature vectors so that the computer can process and understand. For example, for text data, methods such as bag-of-words model, TF-IDF, etc. can be used to extract the feature vectors of the text.
[0086] A third standardization unit, which standardizes the feature vector extraction by normalization.
[0087] Furthermore, the distributed storage module includes:
[0088] A data processing module, which is used to encrypt and split the preprocessed heterogeneous data to obtain data shards;
[0089] A data location module, which is used to predict the future access trend of the data shards according to the analysis of historical access data and determine the storage location of the data shards according to the access trend;
[0090] A data storage module, which is used to store the data shards to the storage nodes according to the storage location;
[0091] A data balancing module, which is used to perform data migration on the data shards stored in the storage nodes when the storage nodes are overloaded.
[0092] As Figure 2 shown, the data processing module includes:
[0093] An encryption processing module, which is used to encrypt the preprocessed heterogeneous data according to the encryption algorithm; the specific encryption process:
[0094] The system randomly generates a digital obfuscation rule a3 within any range from (0 - 9), numbers each specified position of each specified item according to the digital obfuscation rule a3, and simultaneously selects the digital obfuscation rule a3 of the corresponding specified position according to the algorithm factor d to perform digital obfuscation on the specified position of the shuffled card number, obtaining a set of obfuscated numbers;
[0095] Ten encryption algorithm factors d are preset for the rule a4 of inserting into the obfuscated card number, and the ten rules of inserting the encryption algorithm factors d into the obfuscated card number are respectively defined as 0, 1......9 as the decoding keys;
[0096] Randomly select a decoding key from the preset decoding keys, and according to the rule a4 corresponding to the decoding key, combine the encryption algorithm factor d with at least three identical numbers in the set of obfuscated numbers.
[0097] The data sharding processing unit is used to perform sharding processing on the encrypted heterogeneous data to obtain N data shards, where each of the N data shards includes a shard identifier, a shard offset, and a shard flag bit.
[0098] Furthermore, as Figure 2 shown, the data positioning module includes:
[0099] The relationship extraction unit is used to construct a relationship of access volume change based on historical access data; the relationship of access volume change is used to represent the change of access volume over time;
[0100] The feature extraction unit is used to extract at least one access volume feature from the relationship of access volume change; the access volume feature is used to describe the change trend of the access volume; the access volume feature includes a feature point clustering coefficient and a change graph clustering coefficient; the data category corresponding to the historical access data is determined using the access volume feature;
[0101] The access trend prediction unit is used to predict the access trend of the data shard through the data category and the access volume feature;
[0102] The positioning unit is used to determine the storage location of the data shard according to the access trend.
[0103] Furthermore, the data retrieval module includes:
[0104] The receiving module is used to receive the query request input by the user, and perform preliminary parsing on the query request to identify the query type and key information;
[0105] The determining index level module is used to determine the index level according to the query type and the structure of the multi-level index;
[0106] A data location module, configured to start from the determined index level and search for index entries matching the key information in the corresponding index structure;
[0107] A data feedback module, configured to locate the data pointer satisfying the query condition according to the index entry, determine the data storage location according to the pointer, and obtain the data at the data storage location to form a complete query result and feedback it to the user.
[0108] On the other hand, the present invention provides a method for supporting distributed storage and retrieval of heterogeneous data, including the following steps:
[0109] Collect heterogeneous data from multiple data sources, where the heterogeneous data includes structured data, semi-structured data, and unstructured data;
[0110] Clean, format-convert, and standardize the collected heterogeneous data;
[0111] Perform data sharding and load balancing on the preprocessed heterogeneous data, and distribute it for storage in multiple storage nodes;
[0112] Establish a multi-level index for the stored heterogeneous data;
[0113] According to the user's query request, perform fast data retrieval using the multi-level index;
[0114] Adopt encryption methods and access control mechanisms to ensure the security of data during storage and transmission;
[0115] Coordinate the work between each module according to the workload and resource requirements of each module.
[0116] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0117] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent terminal supporting distributed storage and retrieval of heterogeneous data, characterized in that, Including: A data acquisition module for acquiring heterogeneous data from multiple data sources, where the heterogeneous data includes structured data, semi-structured data, and unstructured data; A data preprocessing module for cleaning, format conversion, and standardization processing of the acquired heterogeneous data; A distributed storage module for data sharding and load balancing of the preprocessed heterogeneous data and distributed storage on multiple storage nodes; A data indexing module for establishing multi-level indexes for the stored heterogeneous data; A data retrieval module for performing fast data retrieval using the multi-level indexes according to user query requests; A data security module for ensuring the security of data during storage and transmission by using encryption methods and access control mechanisms; An intelligent scheduling module for coordinating the work between various modules.
2. The intelligent terminal for supporting distributed storage and retrieval of heterogeneous data according to claim 1, characterized in that The data preprocessing module includes: a structured processing module, a semi-structured processing module, and an unstructured processing module; The structured processing module includes: A first data cleaning unit for performing missing value processing, outlier processing, and duplicate value processing on the structured data; A first format conversion unit for performing data type conversion and encoding conversion on the cleaned structured data; A first standardization unit for standardizing the format-converted structured data using normalization; The semi-structured processing module includes: A second data cleaning unit for parsing the semi-structured data, extracting key information and data fields, and removing irrelevant information; A second format conversion unit for converting the cleaned semi-structured data into a structured format; A second standardization unit for standardizing the value range and data format of the format-converted data, and performing standardized classification and encoding on the classification information in the format-converted data, converting the text-based classification data into digital codes; The unstructured processing module includes: A third data cleaning unit for removing and cleaning the noise in the unstructured data and filtering out irrelevant information; A third format conversion unit for extracting text from the unstructured data and converting the extracted text into feature vectors; A third standardization unit for standardizing the extraction of the feature vectors using normalization.
3. An intelligent terminal for supporting distributed storage and retrieval of heterogeneous data according to claim 1, characterized in that, The distributed storage module includes: A data processing module for encrypting and splitting the preprocessed heterogeneous data to obtain data shards; A data location module for predicting the future access trend of the data shards based on the analysis of historical access data and determining the storage location of the data shards according to the access trend; A data storage module for storing the data shards to the storage nodes according to the storage location; A data balancing module for performing data migration on the data shards stored on the storage nodes when the storage nodes are overloaded.
4. The intelligent terminal for supporting distributed storage and retrieval of heterogeneous data according to claim 3, characterized in that, The data processing module includes: An encryption processing module for encrypting the preprocessed heterogeneous data according to an encryption algorithm; A data sharding processing unit, which is used to perform sharding processing on the encrypted heterogeneous data to obtain N data shards. Each of the N data shards includes a shard identifier, a shard offset, and a shard flag bit.
5. The intelligent terminal for supporting distributed storage and retrieval of heterogeneous data according to claim 3, characterized in that, The data location module includes: A relationship extraction unit, which is used to construct a change relationship of access volume based on the historical access data; the change relationship of access volume is used to represent the change of access volume over time; A feature extraction unit, which is used to extract at least one access volume feature from the change relationship of access volume; the access volume feature is used to describe the change trend of access volume; the access volume feature includes a feature point clustering coefficient and a change graph clustering coefficient; the data category corresponding to the historical access data is determined by using the access volume feature; An access trend prediction unit, which is used to predict the access trend of the data shard through the data category and the access volume feature; A location unit, which is used to determine the storage location of the data shard according to the access trend.
6. The intelligent terminal for supporting distributed storage and retrieval of heterogeneous data according to claim 1, wherein The data retrieval module includes: A receiving module, which is used to receive a query request input by a user, and perform preliminary parsing on the query request to identify a query type and key information; An index level determination module, which is used to determine an index level according to the query type and the structure of a multi-level index; A data location module, which is used to start from the determined index level and search for an index entry matching the key information in the corresponding index structure; A data feedback module, which is used to locate a data pointer that meets the query condition according to the index entry, determine a data storage location according to the pointer, and obtain the data at the data storage location to form a complete query result and feedback it to the user.
7. A method for supporting distributed storage and retrieval of heterogeneous data, characterized in that, It includes the following steps: Collect heterogeneous data from multiple data sources, where the heterogeneous data includes structured data, semi-structured data, and unstructured data; Perform cleaning, format conversion, and standardization processing on the collected heterogeneous data; Perform data sharding and load balancing on the preprocessed heterogeneous data, and distribute and store them in multiple storage nodes; Establish a multi-level index for the stored heterogeneous data; Perform fast data retrieval by using the multi-level index according to the user query request; Adopt an encryption method and an access control mechanism to ensure the security of data during storage and transmission; Coordinate the work between each module according to the workload and resource requirements of each module.
Citation Information
Cited By
Intelligent fragmentation storage and access method and system for secondary operation and maintenance multi-modal data
CN121657932A