Big data processing and analyzing device based on artificial intelligence
Through data folding and template multiplexing technology, combined with distributed storage and parallel computing, and dynamic resource scheduling, the flexibility, efficiency and security issues in big data processing are solved, and efficient and secure data analysis is achieved.
Patent Information
- Application Number
- CN202510350760.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing big data processing and analysis technologies lack flexibility and accuracy in processing complex data types, improper resource scheduling leads to inefficiency and insufficient data security.
It adopts a big data processing and analysis device based on artificial intelligence, and through data folding and template multiplexing technology, combining distributed storage and parallel computing, dynamic resource scheduling, supports multi-grained data processing, and strengthens data security.
Significantly reduce duplicate data processing and storage overhead, improve data processing efficiency, improve processing speed and resource utilization, and ensure data security.
Smart Images

Figure CN120276846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and more specifically, to a big data processing and analysis device based on artificial intelligence. Background Art
[0002] With the rapid development of information technology, big data has become an important resource in all walks of life. Big data processing and analysis technology, as a key means to mine data value and improve business decision-making efficiency, is receiving unprecedented attention. Existing big data processing and analysis technologies mainly rely on distributed computing frameworks (such as Hadoop, Spark, etc.) and artificial intelligence technologies (such as machine learning, deep learning, etc.). These technologies can process data volumes of PB level and provide efficient data storage, query, and analysis capabilities.
[0003] In terms of data collection, existing technologies usually extract data from multiple data sources (such as relational databases, non-relational databases, log files, sensor data, etc.) and provide high-quality inputs for subsequent analysis through preprocessing steps such as data cleaning, formatting, and feature extraction. Distributed storage technologies, such as the Hadoop Distributed File System (HDFS), achieve high availability and fault tolerance of data by dispersing data storage on multiple nodes. Parallel computing frameworks, such as MapReduce and Spark, can make full use of the computing resources of the cluster to efficiently process large-scale data.
[0004] Although existing technologies have achieved remarkable results in big data processing and analysis, there are still some obvious disadvantages. First, existing technologies often lack sufficient flexibility and accuracy when processing complex data types. For example, for unstructured or semi-structured data such as log files and sensor data, existing technologies may not be able to effectively extract key information, resulting in inaccurate analysis results or missing important information.
[0005] Secondly, existing technologies have limitations in resource scheduling. As the data volume continues to increase, the demand for computing resources also continues to grow. However, existing technologies often cannot dynamically allocate computing resources according to task requirements, resulting in low resource utilization or slow processing speed. This not only increases costs but also affects the user experience.
[0006] In addition, there are also certain risks in data security for existing technologies. A large amount of sensitive information, such as user privacy and business secrets, is involved in the big data processing and analysis process. Without effective data encryption and integrity verification mechanisms, this information may be leaked or tampered with, causing serious losses to users. Summary of the Invention
[0007] 1. Technical Problems to be Solved
[0008] In view of the problems existing in the prior art, the purpose of the present invention is to provide a big data processing and analysis device based on artificial intelligence. It significantly reduces the processing and storage overhead of duplicate data through data folding and template reuse technologies, improves data processing efficiency, and adopts distributed storage and parallel computing technologies to make full use of hardware resources and further enhance the data processing speed.
[0009] 2. Technical Solution
[0010] To solve the above problems, the present invention adopts the following technical solutions.
[0011] A big data processing and analysis device based on artificial intelligence includes:
[0012] A data acquisition module, which is used to collect data from multiple data sources and perform adaptive preprocessing on the data based on deep learning, including data cleaning, formatting, and feature extraction;
[0013] A distributed storage module, which is used to store the preprocessed data distributedly on multiple nodes and manage it using a distributed file system;
[0014] A parallel computing module, which is used to perform parallel computing on the data stored distributedly and analyze the data using artificial intelligence algorithms, including deep learning, machine learning, and graph neural networks;
[0015] A data folding module, which is used to compare new data with historical data, identify duplicate or similar parts, and reuse the processing results of historical data. The data folding module adopts a comparison method based on hash algorithm or similarity calculation and supports dynamic adjustment of the folding granularity;
[0016] A data template module, which is used to generate data templates, extract key features as template identifiers, and support hierarchical comparison and template reuse. The data template module can dynamically generate and update templates according to the coincidence degree between new data and templates, and preferentially compare template identifiers with high frequencies;
[0017] A resource optimization and allocation module, which is used to dynamically allocate computing resources according to the requirements of processing tasks and adopts a dynamic resource scheduling algorithm based on reinforcement learning;
[0018] A visualization display module, which is used to display the analysis results to users in a visual way and support various visualization charts and interactive operations;
[0019] A data security module, which is used to ensure the security of data processing and storage, including data encryption and integrity verification;
[0020] A cross-platform adaptation module, which is used to support cross-platform data processing and analysis, provide a unified API interface and multi-format protocol support.
[0021] As a further improvement of the present invention, the data folding module further includes:
[0022] A data comparison unit for identifying duplicate or similar parts of data through a hash algorithm, similarity calculation, or machine learning model;
[0023] A folding processing unit for directly invoking the processing results of historical data for duplicate or similar parts and performing normal processing on non-duplicate parts;
[0024] A result fusion unit for fusing the processing results of the folded part and the non-folded part and using a data verification mechanism to ensure the accuracy of the fusion result.
[0025] As a further improvement of the present invention, the data template module further includes:
[0026] A template generation unit for forming a data template from the folded data set and extracting key features as template identifiers;
[0027] A template comparison unit for preferentially comparing the features of new data with the template identifiers and determining whether to reuse the processing results in the template according to the coincidence degree threshold;
[0028] A template update unit for generating a new template when the coincidence degree is lower than the threshold but higher than the set value, and updating the template identifier and suffix, where the suffix includes a folding strategy and a folding frequency.
[0029] As a further improvement of the present invention, the data acquisition module further includes:
[0030] A multi-granularity data processing unit for supporting high-granularity mode and low-granularity mode. The high-granularity mode is used for high-precision data analysis, adopting deep data cleaning, feature extraction, and complex model analysis. The low-granularity mode is used for fast data processing, adopting data sampling, simplification, and lightweight analysis;
[0031] A dynamic switching subunit for automatically switching the processing mode according to task requirements and identifying the task type through a machine learning model.
[0032] As a further improvement of the present invention, the data folding module further includes a distributed folding processing unit for supporting distributed data folding processing:
[0033] A distributed data comparison subunit for comparing data in parallel on multiple nodes;
[0034] A distributed folding processing subunit for invoking the processing results of historical data in parallel on multiple nodes;
[0035] A distributed result fusion subunit for parallelly fusing processing results on multiple nodes.
[0036] As a further improvement of the present invention, the data security module further includes:
[0037] A data encryption unit for encrypting and storing the historical data processing results;
[0038] An integrity verification unit for verifying the integrity of the historical data processing results through digital signature or hash verification algorithms;
[0039] An access control unit for dynamically controlling data access and operations according to user permissions.
[0040] As a further improvement of the present invention, the data template module further includes a multi-dimensional optimization unit for supporting multi-dimensional data folding optimization:
[0041] A time dimension optimization subunit for folding historical data according to time windows;
[0042] A space dimension optimization subunit for folding data according to geographical locations or data partitions;
[0043] A semantic dimension optimization subunit for folding data according to data categories or themes;
[0044] A dynamic optimization subunit for selecting the optimal folding dimension according to specific requirements.
[0045] As a further improvement of the present invention, the visualization display module further includes a data folding visualization unit for supporting the visualization and debugging of data folding:
[0046] A visualization subunit for displaying the results of data comparison and folding processing;
[0047] A debugging interface subunit for allowing users to manually adjust folding parameters, including folding granularity, comparison thresholds, and resource allocation strategies;
[0048] A real-time monitoring subunit for real-time monitoring of data processing status and performance metrics.
[0049] As a further improvement of the present invention, the cross-platform adaptation module further includes:
[0050] A unified API interface unit for facilitating integration into existing systems;
[0051] A multi-format protocol support unit for supporting multiple data formats and protocols;
[0052] A cross-platform deployment unit for implementing efficient template management and data folding processing in a distributed environment;
[0053] A containerization support unit for achieving rapid deployment and expansion through container technology.
[0054] As a further improvement of the present invention, the data template module further includes an intelligent folding unit for supporting intelligent data folding based on machine learning:
[0055] A feature extraction subunit for extracting key features of the data template through a deep learning model or a clustering algorithm;
[0056] A dynamic threshold adjustment subunit for automatically optimizing the overlap threshold according to data characteristics and processing requirements;
[0057] A template compression subunit for compressing and optimizing the data template to reduce storage space and transmission overhead.
[0058] 3. Beneficial effects
[0059] Compared with the prior art, the advantages of the present invention are as follows:
[0060] (1) Through data folding and template reuse technologies, the present invention significantly reduces the processing and storage overhead of duplicate data, improves data processing efficiency, and adopts distributed storage and parallel computing technologies to make full use of hardware resources, further enhancing the data processing speed.
[0061] (2) The dynamic resource scheduling algorithm based on reinforcement learning in the present invention can dynamically adjust resource allocation according to real-time workloads, optimize the utilization of computing resources, and effectively avoid resource waste and improve the overall performance of the system in multi-task concurrent scenarios.
[0062] (3) The present invention supports high-granularity and low-granularity data processing modes, can automatically switch according to task requirements, adapts to data processing requirements in different scenarios, the high-granularity mode is suitable for high-precision data analysis, and the low-granularity mode is suitable for fast data processing, enhancing the flexibility and applicability of the device.
[0063] (4) Through data encryption, integrity verification, and access control mechanisms in the present invention, the security of data processing and storage is ensured, data leakage and tampering are prevented, and the requirements of high-security scenarios are met. Description of the drawings
[0064] Figure 1 It is a structural schematic diagram of the present invention. Detailed implementation manners
[0065] A detailed description of an implementation manner of the present application will be given below with reference to the drawings.
[0066] Example:
[0067] This embodiment provides a big data processing and analysis device based on artificial intelligence, and its structure is as Figure 1 shown, including the following modules:
[0068] Data acquisition module: It is used to collect data from multiple data sources (such as relational databases, non-relational databases, log files, sensor data, etc.), and perform adaptive preprocessing on the data based on deep learning.
[0069] Distributed storage module: It is used to distribute and store the preprocessed data on multiple nodes, and manage it using a distributed file system (such as HDFS).
[0070] Parallel computing module: It is used to perform parallel computing on the data stored distributively, and analyze the data using artificial intelligence algorithms (such as deep learning, machine learning, graph neural networks).
[0071] Data folding module: It is used to compare new data with historical data, identify duplicate or similar parts, and reuse the processing results of historical data.
[0072] Data template module: It is used to generate data templates, extract key features as template identifiers, and support hierarchical comparison and template reuse.
[0073] Resource optimization and allocation module: It is used to dynamically allocate computing resources according to the requirements of processing tasks, and adopt a dynamic resource scheduling algorithm based on reinforcement learning.
[0074] Visualization display module: It is used to display the analysis results to users in a visual way, and supports a variety of visualization charts and interactive operations.
[0075] Data security module: It is used to ensure the security of data processing and storage, including data encryption and integrity verification.
[0076] Cross-platform adaptation module: It is used to support cross-platform data processing and analysis, and provides a unified API interface and multi-format protocol support.
[0077] The working principle of this device is as follows:
[0078] I. Data acquisition and preprocessing
[0079] Data acquisition: The data acquisition module collects data from multiple data sources, including structured data (such as relational databases), semi-structured data (such as JSON, XML), and unstructured data (such as log files, sensor data).
[0080] The data acquisition module supports real-time data acquisition and batch data acquisition, and can dynamically adjust the acquisition frequency and scale according to task requirements.
[0081] Data Preprocessing:
[0082] The data acquisition module uses an adaptive preprocessing method based on deep learning to clean, format, and extract features from the data.
[0083] For example, in the scenario of log data analysis, the data acquisition module can automatically identify the log format and extract key fields (such as timestamp, event type, error code).
[0084] The data acquisition module also supports multi-granularity data processing modes, including high-granularity mode and low-granularity mode:
[0085] High-granularity mode: Used for high-precision data analysis, adopting deep data cleaning, feature extraction, and complex model analysis. For example, in financial transaction data analysis, the high-granularity mode can extract detailed features of each transaction (such as transaction amount, transaction time, transaction location).
[0086] Low-granularity mode: Used for fast data processing, adopting data sampling, simplification, and lightweight analysis. For example, in real-time log monitoring, the low-granularity mode can quickly extract key error information.
[0087] Dynamic switching subunit: Automatically switches the processing mode according to the task requirements and identifies the task type through a machine learning model. For example, in tasks with high real-time requirements, it automatically switches to the low-granularity mode; in tasks with high precision requirements, it automatically switches to the high-granularity mode.
[0088] II. Distributed Storage and Parallel Computing
[0089] Distributed storage: The preprocessed data is distributedly stored on multiple nodes and managed using a distributed file system (such as HDFS).
[0090] The distributed storage module supports data sharding and replication mechanisms to ensure high availability and fault tolerance of the data.
[0091] Parallel computing: The parallel computing module performs parallel computing on the data stored distributively and analyzes the data using artificial intelligence algorithms (such as deep learning, graph neural networks).
[0092] For example, in the scenario of financial transaction data analysis, the parallel computing module can calculate the risk score of each transaction in parallel.
[0093] The parallel computing module supports multiple computing frameworks (such as MapReduce, Spark) and can dynamically select the optimal computing framework according to the task requirements.
[0094] III. Data Folding and Template Reuse
[0095] Data folding: The data folding module compares new data with historical data, identifies duplicate or similar parts, and reuses the processing results of historical data.
[0096] The data folding module includes the following subunits:
[0097] Data comparison unit: Identifies duplicate or similar parts of data through hash algorithms, similarity calculations, or machine learning models. For example, in log data analysis, the data comparison unit can quickly identify duplicate error logs through hash algorithms.
[0098] Folding processing unit: Directly invokes the processing results of historical data for duplicate or similar parts and processes non-duplicate parts normally. For example, in sensor data analysis, the folding processing unit can directly invoke the analysis results of historical temperature data.
[0099] Result fusion unit: Fuses the processing results of the folded part and the non-folded part and adopts a data verification mechanism to ensure the accuracy of the fusion result. For example, in financial transaction data analysis, the result fusion unit can fuse the risk scores of the folded part and the non-folded part to generate a final risk report.
[0100] Dynamic folding optimization unit: Dynamically adjusts the folding strategy according to data characteristics and processing requirements, including folding granularity and comparison threshold. For example, in real-time data processing, the dynamic folding optimization unit can automatically reduce the folding granularity to improve processing speed.
[0101] Data template: The data template module generates data templates, extracts key features as template identifiers, and supports hierarchical comparison and template reuse.
[0102] The data template module includes the following subunits:
[0103] Template generation unit: Forms a data template from the folded data set and extracts key features as template identifiers. For example, in log data analysis, the template generation unit can extract error codes and occurrence times as template identifiers.
[0104] Template comparison unit: Prioritizes comparing the features of new data with the template identifiers and decides whether to reuse the processing results in the template based on the coincidence degree threshold. For example, in financial transaction data analysis, the template comparison unit can prioritize comparing high-frequency trading templates.
[0105] Template update unit: Generates a new template when the coincidence degree is lower than the threshold but higher than the set value, and updates the template identifier and suffix, where the suffix includes the folding strategy and folding frequency. For example, in sensor data analysis, the template update unit can generate a new temperature data template.
[0106] Template Compression Unit: Compresses and optimizes data templates to reduce storage space and transmission overhead. For example, in log data analysis, the template compression unit can compress and store historical log templates.
[0107] The generation of the suffix includes:
[0108] 1. Definition of the suffix
[0109] The suffix is a part of the data template identifier and is used to record the folding strategy and folding frequency of the data template. The generation of the suffix is an important part of data template management, which can help the system quickly identify and reuse historical data processing results.
[0110] The format of the suffix is usually:
[0111] Template identifier _ Folding strategy _ Folding frequency
[0112] For example: Template123_StrategyA_5, which means the template identifier is Template123, the folding strategy is StrategyA, and the folding frequency is 5.
[0113] 2. Generation of the folding strategy
[0114] The folding strategy refers to the specific method or rule adopted when folding data. The generation of the folding strategy is based on data characteristics, processing tasks, and system requirements, and usually includes the following steps:
[0115] 2.1 Types of folding strategies
[0116] Folding strategies can be divided into the following types:
[0117] Hash-based folding strategy: Uses hash algorithms (such as MD5, SHA-256) to compare data and identify duplicate parts, suitable for structured data (such as database records).
[0118] Similarity-based folding strategy: Uses similarity calculations (such as cosine similarity, Jaccard similarity) to compare data and identify similar parts, suitable for unstructured data (such as text, logs).
[0119] Machine learning-based folding strategy: Uses machine learning models (such as clustering algorithms, deep learning models) to compare data and identify duplicate or similar parts, suitable for complex data (such as images, time series).
[0120] 2.2 Generation logic of the folding strategy
[0121] The generation logic of the folding strategy is as follows:
[0122] Data Feature Analysis: Extract features from the data and analyze the data type, structure, and distribution. For example, in log data analysis, extract features such as timestamps, event types, and error codes.
[0123] Task Requirement Matching: Select the optimal folding strategy according to the requirements of the processing task. For example, in real-time log monitoring, select a hash-based folding strategy to improve processing speed.
[0124] System Performance Optimization: Dynamically adjust the folding strategy according to system performance (such as computing resources and storage space). For example, in high-load scenarios, select a lightweight folding strategy to reduce resource consumption.
[0125] 2.3 Dynamic Adjustment of Folding Strategy
[0126] The folding strategy supports dynamic adjustment and can be automatically optimized according to data characteristics and processing requirements. For example, in scenarios with a high data duplication rate, automatically switch to a hash-based folding strategy.
[0127] In scenarios with a high data similarity, automatically switch to a similarity-based folding strategy.
[0128] 3. Generation of Folding Frequency
[0129] The folding frequency refers to the number of times a data template is reused. The generation of folding frequency is based on the usage of data templates and usually includes the following steps:
[0130] 3.1 Recording of Folding Frequency
[0131] The folding frequency records the number of times a data template is reused through a counter. For example: each time a data template is reused, the counter is incremented by 1, and the initial value of the counter is 0, indicating that the data template has not been reused yet.
[0132] 3.2 Application of Folding Frequency
[0133] The folding frequency is used to optimize the management and reuse of data templates. Specific applications include:
[0134] Preferentially Reuse High-Frequency Templates: When comparing data, preferentially reuse data templates with a higher folding frequency. For example, in log data analysis, preferentially reuse high-frequency error log templates.
[0135] Dynamically Eliminate Low-Frequency Templates: When the storage space is insufficient, dynamically eliminate data templates with a lower folding frequency. For example, in sensor data analysis, eliminate low-frequency temperature data templates.
[0136] 3.3 Dynamic Update of Folding Frequency
[0137] The folding frequency supports dynamic updates and can be automatically adjusted according to the usage of the data template. For example, when the data template is reused, the folding frequency is automatically updated, and when the data template is phased out, the folding frequency is automatically reset.
[0138] 4. Generation logic of the suffix
[0139] The generation logic of the suffix is as follows:
[0140] Template identifier generation: Extract the key features of the data template to generate a template identifier. For example, in log data analysis, extract the error code and occurrence time as the template identifier.
[0141] Folding strategy generation:
[0142] Generate a folding strategy based on data characteristics, processing tasks, and system requirements. For example, in real-time log monitoring, generate a hash-based folding strategy.
[0143] Folding frequency generation: Record the number of times the data template is reused to generate the folding frequency. For example, in log data analysis, record the folding frequency of high-frequency error log templates.
[0144] Suffix concatenation: Concatenate the template identifier, folding strategy, and folding frequency into a suffix. For example, generate the suffix Template123_StrategyA_5.
[0145] 5. Application scenarios of the suffix
[0146] The suffix plays an important role in data template management and reuse. The specific application scenarios include:
[0147] Data template retrieval: Quickly retrieve data templates through the suffix to improve data comparison efficiency. For example, in log data analysis, quickly retrieve high-frequency error log templates through the suffix Template123_StrategyA_5.
[0148] Data template optimization: Optimize the management and reuse of data templates through the suffix, reducing storage space and computing resources. For example, in sensor data analysis, dynamically phase out low-frequency temperature data templates through the suffix.
[0149] Data template debugging: Debug the generation and reuse process of data templates through the suffix to improve system stability. For example, in financial transaction data analysis, debug the generation process of high-frequency transaction templates through the suffix
[0150] IV. Resource Optimization and Dynamic Scheduling
[0151] Resource optimization allocation module: Dynamically allocate computing resources according to the requirements of processing tasks, and adopt a dynamic resource scheduling algorithm based on reinforcement learning.
[0152] For example, in a high-load scenario, the resource optimization allocation module can dynamically increase computing nodes to improve processing efficiency.
[0153] V. Visualization Display and Debugging
[0154] Visualization display module: Presents the analysis results to users in a visual manner, supporting multiple visualization charts (such as line charts, bar charts, heat maps) and interactive operations.
[0155] The visualization display module includes the following subunits:
[0156] Visualization subunit: Displays the results of data comparison and folding processing. For example, in medical data analysis, the visualization subunit can generate a health trend chart of patients.
[0157] Debugging interface subunit: Allows users to manually adjust folding parameters, including folding granularity, comparison threshold, and resource allocation strategy.
[0158] Real-time monitoring subunit: Monitors the data processing status and performance metrics in real time.
[0159] VI. Data Security and Cross-Platform Adaptation
[0160] Data security module: Encrypts and stores the historical data processing results, and verifies the integrity of the data through digital signature or hash verification algorithm.
[0161] The data security module includes the following subunits:
[0162] Data encryption unit: Encrypts and stores the historical data processing results.
[0163] Integrity verification unit: Verifies the integrity of the historical data processing results through digital signature or hash verification algorithm.
[0164] Access control unit: Dynamically controls data access and operations according to user permissions.
[0165] Cross-platform adaptation module:
[0166] Provides a unified API interface and multi-format protocol support, facilitating integration into existing systems.
[0167] The cross-platform adaptation module includes the following subunits:
[0168] Unified API interface unit: Facilitates integration into existing systems.
[0169] Multi-format protocol support unit: Supports multiple data formats and protocols.
[0170] Cross-platform deployment unit: Implements efficient template management and data folding processing in a distributed environment.
[0171] Containerization Support Unit: Achieves rapid deployment and scalability through container technologies (such as Docker).
[0172] This device can be applied to the following scenarios:
[0173] Financial Risk Control: Performs high-precision analysis on transaction data to identify abnormal trading behaviors, and improves processing efficiency by folding and reusing the analysis results of historical data.
[0174] Medical Diagnosis: Performs in-depth analysis on patients' medical data to assist doctors in diagnosis, and reduces repeated calculations by reusing historical diagnosis results through data templates.
[0175] Log Monitoring: Performs real-time analysis on massive log data to identify system anomalies, and improves real-time performance by quickly processing log data in low-granularity mode.
[0176] Intelligent Recommendation: Performs real-time analysis on user behavior data to generate personalized recommendations, and improves recommendation efficiency by folding and reusing historical recommendation results.
[0177] Internet of Things Data Analysis: Performs real-time analysis on sensor data to monitor device status, and reduces storage space by reusing historical data analysis results through data templates.
[0178] The following is a specific implementation example, taking log data analysis as an example:
[0179] Data Collection and Preprocessing: The data collection module collects data from log files and extracts key fields (such as timestamps, event types, error codes) using an adaptive preprocessing method based on deep learning. The data collection module automatically switches between high-granularity mode and low-granularity mode according to task requirements.
[0180] Distributed Storage and Parallel Computing: The preprocessed data is distributed and stored on multiple nodes, and the parallel computing module performs parallel analysis on the log data to identify error types and occurrence frequencies.
[0181] Data Folding and Template Reuse: The data folding module compares new log data with historical log data to identify duplicate error logs and directly invokes the analysis results of historical data. The data template module generates error log templates and extracts key features (such as error codes, occurrence times) as template identifiers.
[0182] Resource Optimization and Dynamic Scheduling: The resource optimization allocation module dynamically allocates computing resources according to the volume of log data to ensure efficient processing.
[0183] Visualization Display and Debugging: The visualization display module generates a trend chart of error logs to help operation and maintenance personnel quickly locate problems.
[0184] Data security and cross-platform adaptation: The data security module encrypts and stores the log data and verifies the integrity of the data through a hash check algorithm. The cross-platform adaptation module provides a unified API interface for easy integration into existing operation and maintenance systems.
[0185] This embodiment details the structure, working principle, and specific application scenarios of the big data processing and analysis device based on artificial intelligence. Through technologies such as data folding, template reuse, and intelligent resource scheduling, this device can significantly improve data processing efficiency, reduce storage space, and ensure data security. The device has broad application prospects and is applicable to multiple fields such as finance, healthcare, log monitoring, intelligent recommendation, and the Internet of Things.
[0186] The above is only a preferred specific embodiment of the present invention; however, the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A big data processing and analysis device based on artificial intelligence, characterized in that: Including: A data acquisition module, which is used to collect data from multiple data sources and perform adaptive preprocessing on the data based on deep learning, including data cleaning, formatting, and feature extraction; A distributed storage module, which is used to store the preprocessed data distributedly on multiple nodes and manage it using a distributed file system; A parallel computing module, which is used to perform parallel computing on the distributedly stored data and analyze the data using artificial intelligence algorithms, including deep learning, machine learning, and graph neural networks; A data folding module, which is used to compare new data with historical data, identify duplicate or similar parts, and reuse the processing results of historical data. The data folding module adopts a comparison method based on a hash algorithm or similarity calculation and supports dynamic adjustment of the folding granularity; A data template module, which is used to generate data templates, extract key features as template identifiers, and support hierarchical comparison and template reuse. The data template module can dynamically generate and update templates according to the coincidence degree between new data and templates, and preferentially compare high-frequency template identifiers; A resource optimization and allocation module, which is used to dynamically allocate computing resources according to the requirements of processing tasks and adopts a dynamic resource scheduling algorithm based on reinforcement learning; A visualization display module, which is used to display the analysis results to users in a visual way and support various visualization charts and interactive operations; A data security module, which is used to ensure the security of data processing and storage, including data encryption and integrity verification; A cross-platform adaptation module, which is used to support cross-platform data processing and analysis, provide a unified API interface, and support multi-format protocols.
2. The big data processing and analysis device based on artificial intelligence according to claim 1, wherein: The data folding module further includes: A data comparison unit, which is used to identify duplicate or similar parts of data through a hash algorithm, similarity calculation, or machine learning model; A folding processing unit, which is used to directly call the processing results of historical data for duplicate or similar parts and perform normal processing on non-duplicate parts; A result fusion unit, which is used to fuse the processing results of the folded part and the non-folded part and adopt a data verification mechanism to ensure the accuracy of the fusion result.
3. The big data processing and analysis device based on artificial intelligence according to claim 1, wherein: The data template module further includes: A template generation unit, which is used to form a data template from the folded data set and extract key features as template identifiers; A template comparison unit, which is used to preferentially compare the features of new data with template identifiers and decide whether to reuse the processing results in the template according to a coincidence degree threshold; A template update unit, which is used to generate a new template when the coincidence degree is lower than the threshold but higher than a set value, and update the template identifier and suffix. The suffix includes a folding strategy and a folding frequency.
4. The big data processing and analysis device based on artificial intelligence according to claim 1, characterized in that: The data acquisition module further includes: A multi-granularity data processing unit, which is used to support a high-granularity mode and a low-granularity mode. The high-granularity mode is used for high-precision data analysis, adopting in-depth data cleaning, feature extraction, and complex model analysis. The low-granularity mode is used for fast data processing, adopting data sampling, simplification, and lightweight analysis; A dynamic switching subunit, which is used to automatically switch the processing mode according to task requirements and identify the task type through a machine learning model.
5. The big data processing and analysis device based on artificial intelligence according to claim 1, wherein: The data folding module further includes a distributed folding processing unit for supporting distributed data folding processing: A distributed data comparison subunit for parallelly comparing data on multiple nodes; A distributed folding processing subunit for parallelly invoking historical data processing results on multiple nodes; A distributed result fusion subunit for parallelly fusing processing results on multiple nodes.
6. The big data processing and analysis device based on artificial intelligence according to claim 1, characterized in that: The data security module further includes: A data encryption unit for encrypting and storing historical data processing results; An integrity verification unit for verifying the integrity of historical data processing results through digital signature or hash verification algorithms; An access control unit for dynamically controlling data access and operations according to user permissions.
7. The big data processing and analysis device based on artificial intelligence according to claim 1, characterized in that: The data template module further includes a multi-dimensional optimization unit for supporting multi-dimensional data folding optimization: A time dimension optimization subunit for folding historical data by time window; A space dimension optimization subunit for folding data by geographical location or data partition; A semantic dimension optimization subunit for folding data by data category or theme; A dynamic optimization subunit for selecting the optimal folding dimension according to specific requirements.
8. The big data processing and analysis device based on artificial intelligence according to claim 1, characterized in that: The visualization display module further includes a data folding visualization unit for supporting the visualization and debugging of data folding: A visualization subunit for displaying the results of data comparison and folding processing; A debugging interface subunit for allowing users to manually adjust folding parameters, including folding granularity, comparison threshold, and resource allocation strategy; A real-time monitoring subunit for real-time monitoring of data processing status and performance metrics.
9. The big data processing and analysis device based on artificial intelligence according to claim 1, characterized in that: The cross-platform adaptation module further includes: A unified API interface unit for facilitating integration into existing systems; A multi-format protocol support unit for supporting multiple data formats and protocols; A cross-platform deployment unit for achieving efficient template management and data folding processing in a distributed environment; A containerization support unit for achieving rapid deployment and expansion through container technology.
10. The big data processing and analysis device based on artificial intelligence according to claim 1, characterized in that: The data template module further includes an intelligent folding unit for supporting machine learning-based intelligent data folding: A feature extraction subunit for extracting key features of data templates through deep learning models or clustering algorithms; A dynamic threshold adjustment subunit for automatically optimizing the coincidence threshold according to data features and processing requirements; A template compression subunit for compressing and optimizing data templates to reduce storage space and transmission overhead.
Citation Information
Patent Citations
Internet big data processing system and method based on artificial intelligence
CN117076810A
Data classification cleaning method
CN117290315A
Efficient data processing system based on cloud network fusion and implementation method thereof
CN118964470A
Data analysis and governance integrated platform based on multi-dimensional data
CN119025582A
System for real-time data aggregation and analysis in IoT networks
DE202024107370U1