Artificial intelligence big data processing system

By designing an artificial intelligence big data processing system, using technical means such as data integration, cleaning, compression and deduplication, the existing system has solved the problems of low processing accuracy and high computing resource consumption when processing data with multiple sources, different formats and uneven quality, and achieved efficient and accurate data processing and analysis.

CN120123329AInactive Publication Date: 2025-06-10MAISIEVO (BEIJING) TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510206494.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When existing big data processing systems process data with multiple sources, different formats and uneven quality, they have low processing accuracy and high computing resources, and cannot integrate and analyze data in a timely manner.

Method used

An artificial intelligence big data processing system was designed, including a demand analysis planning module, a data comprehensive acquisition and integration module, a database module, a data comprehensive processing and analysis module and a security monitoring and assurance module. The system improves data quality and processing efficiency through technical means such as data integration, cleaning, compression and deduplication.

Benefits of technology

Effectively remove noise, redundancy and errors in the data, improve the accuracy and efficiency of data analysis, realize the standardization of different data sources, and reduce computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123329A_ABST
    Figure CN120123329A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence big data processing system, and belongs to the technical field of data processing, and the system comprises a demand analysis planning module which is used for determining business demands and system targets; the data comprehensive acquisition and integration module is used for acquiring multi-party data and integrating the acquired multi-party data; the data comprehensive acquisition and integration module and the data comprehensive processing and analysis module are used for performing comprehensive integration, processing and analysis on data, so that noise, redundancy and errors in data information can be effectively removed, and the accuracy and effectiveness of data analysis are effectively improved; and meanwhile, the data is integrated and analyzed through the data comprehensive acquisition and integration module and the data comprehensive processing and analysis module, so that the data of different data sources can be effectively standardized, the data quality is improved, the possibility of data errors is reduced, the accuracy and processing efficiency of data processing are further improved, and the data processing efficiency is improved. And consumption of computing resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and more specifically, relates to an artificial intelligence big data processing system. Background Art

[0002] In the current context of the digital wave sweeping across the globe, human society is generating and accumulating data at an unprecedented speed, and the big data era has already arrived. From business operations to scientific research, from healthcare to public management, massive amounts of data are being generated in various fields. These data contain huge value and have become the key driving force for social development and innovation.

[0003] However, the performance of existing big data processing systems is relatively single. During the actual operation and calculation process, due to the complexity of big data content, multi-source data information is likely to have different formats, and the quality of some data is uneven, with problems such as missing values, error values, and duplicate values. Existing data processing systems cannot promptly judge and integrate data information, and the processing accuracy may not be high during subsequent processing and analysis. At the same time, when processing and calculating different formats of data information, more computing resources need to be invested, and the consumption of computing resources may be relatively large. Summary of the Invention

[0004] Problems to be Solved

[0005] In view of the problems raised in the existing background art, the present invention provides an artificial intelligence big data processing system.

[0006] Technical Solution

[0007] To solve the above problems, the present invention adopts the following technical solutions.

[0008] An artificial intelligence big data processing system, comprising:

[0009] A requirements analysis and planning module: used to clarify business requirements and system goals;

[0010] A data comprehensive collection and integration module: used to collect multi-party data and integrate the collected multi-party data;

[0011] A database module: used to store and collect the integrated multi-party data;

[0012] A data comprehensive processing and analysis module: used to process and analyze the stored multi-party data to improve data quality;

[0013] A security monitoring and guarantee module: used to monitor and manage data information and encrypt data information.

[0014] Preferably, the data comprehensive acquisition and integration module includes a data identification module, a data acquisition module, and a data integration module:

[0015] The data identification module is used to determine the data source;

[0016] The data acquisition module is used to select a suitable data acquisition tool according to the data source and characteristics, and use the data acquisition tool to extract and collect data information;

[0017] The data integration module is used to integrate the extracted and collected data, and the data integration module will also perform preliminary cleaning and conversion on the integrated data.

[0018] Further, the database module includes a data storage module, a data management module, a data security module, a data operation module, and a data backup and recovery module:

[0019] The data storage module is used for data storage and management, and the data storage module includes two parts: physical storage and logical storage. Physical storage is responsible for the physical storage of data, and logical storage is responsible for the logical organization and management of data;

[0020] The data operation module is used to perform various operations on the data in the database, such as insertion, deletion, update, and query;

[0021] The data management is responsible for managing and controlling the data in the database module;

[0022] The data security module is used to protect the data in the database module from being illegally accessed and modified;

[0023] The data backup and recovery module is responsible for regularly backing up the data in the database module and performing data recovery when needed. The data backup and recovery module includes the formulation of data backup strategies, the execution of data backup, the testing of data recovery, and the implementation of data recovery.

[0024] Preferably, the data comprehensive processing and analysis module includes a data cleaning module, a data integration and optimization module, a data compression and feature selection module, and a data deduplication module;

[0025] The data cleaning module is used to clean the data in the database, identify and correct errors and inconsistencies in the data;

[0026] The data integration and optimization module is used to integrate the cleaned data, eliminate data redundancy, and the data integration and optimization module can also establish data lineage for tracking data sources and historical changes;

[0027] The data compression and feature selection module is used to compress the integrated data, reduce the storage space requirement, while maintaining the data's feature representation ability. The data compression and feature selection module also introduces a sparse coding method to reduce the feature dimension, improve the data processing efficiency, and reduce the consumption of computing resources;

[0028] The data deduplication module introduces a hash function for data deduplication, quickly identifying and merging duplicate records. At the same time, the data deduplication module uses a clustering algorithm to identify duplicate data and deduplicate the data through classification labels.

[0029] Furthermore, the security monitoring and guarantee module includes a monitoring and diagnosis module and an encryption module;

[0030] The monitoring and diagnosis module monitors the data information in real time, timely discovers potential problems and bottlenecks in the data, and provides data support for optimization; the monitoring and diagnosis module also uses machine learning and statistical analysis methods to identify abnormal behaviors and fault patterns in the database module, improving the accuracy and timeliness of fault detection;

[0031] The encryption module is used to encrypt sensitive data, including static data encryption and transmission data encryption.

[0032] Even further, the formula for establishing the data lineage relationship in the data integration and optimization module is:

[0033] G = (V, E)

[0034] V represents the set of nodes, E represents the set of edges, V represents data entities such as data sources, data tables, and data fields, and E represents the flow and relationship between data.

[0035] Even further, the formula for the sparse coding method is:

[0036]

[0037] x is the input signal, D is the dictionary matrix, and s is the coefficient vector.

[0038] Even further, when the data deduplication module introduces a hash function for data deduplication, the formula for the hash function is:

[0039] h(k) = k mod m

[0040] k is the keyword for hash processing, and m is the size of the hash table.

[0041] Even further, the formula for the machine learning and statistical analysis method is:

[0042]

[0043] Cov(X,Y) is the covariance of X and Y, and σX and σY are the standard deviations of X and Y respectively;

[0044] The calculation formula of the said Cov(X,Y) is:

[0045]

[0046] and are the sample means of X and Y respectively.

[0047] Advantageous Effects

[0048] Compared with the prior art, the advantageous effects of the present invention are as follows:

[0049] (1) By using the data comprehensive acquisition and integration module and the data comprehensive processing and analysis module to comprehensively integrate, process and analyze the data, the present invention can effectively remove the noise, redundancy and errors in the data information, effectively improve the accuracy and effectiveness of data analysis. At the same time, by integrating and analyzing the data through the data comprehensive acquisition and integration module and the data comprehensive processing and analysis module, it can effectively standardize the data from different data sources, improve the data quality, reduce the possibility of data errors, further improve the accuracy and processing efficiency of data processing, and reduce the consumption of computing resources. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments or exemplifications of the present application, the following will briefly introduce the drawings required for use in the description of the embodiments or exemplifications. Obviously, the drawings in the following description are only some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the drawings shown.

[0051] Figure 1 It is a schematic diagram of the system structure of the present invention. Detailed Embodiments

[0052] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0053] Accordingly, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0054] Embodiment 1

[0055] As Figure 1 shown, in specific implementation, the present system can be applied in an e-commerce user behavior data analysis system, an artificial intelligence big data processing system, including:

[0056] Requirement analysis and planning module: When an e-commerce enterprise hopes to optimize the product recommendation strategy and improve the user purchase conversion rate by analyzing user behavior data, after communicating with the enterprise's market, operation and other departments, the requirement analysis and planning module clarifies that the business requirement is to accurately analyze user browsing, collection, purchase and other behavior data. The system goal is to build a platform that can process and analyze big data in real time, providing high-quality data support for the product recommendation algorithm, for clarifying business requirements and system goals;

[0057] Data comprehensive collection and integration module: used to collect multi-party data and integrate the collected multi-party data;

[0058] Database module: used to store and collect the integrated multi-party data;

[0059] Data comprehensive processing and analysis module: used to process and analyze the stored multi-party data to improve data quality;

[0060] Security monitoring and guarantee module: used to monitor and manage data information and encrypt data information.

[0061] The data comprehensive collection and integration module includes a data identification module, a data collection module and a data integration module:

[0062] The data identification module determines that the data sources include user log records of e-commerce websites, operation data of mobile applications, and third-party marketing data;

[0063] The data collection module selects appropriate collection tools according to the characteristics of different data sources. For website logs, a log collection tool (such as Fluentd) is used for real-time collection; for third-party marketing data, data extraction is performed through an API interface;

[0064] The data integration module integrates the collected multi-party data, and performs preliminary cleaning and conversion on key information such as user ID and timestamp to ensure consistent data formats and prepare for subsequent analysis.

[0065] The database module includes a data storage module, a data management module, a data security module, a data operation module, and a data backup and recovery module:

[0066] The data storage module is used for data storage and management. The data storage module includes two parts: physical storage and logical storage. Physical storage is responsible for the physical storage of data, while logical storage is responsible for the logical organization and management of data;

[0067] The data operation module is used to perform various operations on the data in the database, such as insertion, deletion, update, and query;

[0068] Data management is responsible for managing and controlling the data in the database module;

[0069] The data security module is used to protect the data in the database module from being illegally accessed and modified;

[0070] The data backup and recovery module is responsible for regularly backing up the data in the database module and performing data recovery when needed. The data backup and recovery module includes the formulation of data backup strategies, the execution of data backup, the testing of data recovery, and the implementation of data recovery.

[0071] The data comprehensive processing and analysis module includes a data cleaning module, a data integration and optimization module, a data compression and feature selection module, and a data deduplication module;

[0072] The data cleaning module is used to clean the data in the database, identify and correct errors and inconsistencies in the data;

[0073] The data integration and optimization module is used to integrate the cleaned data, eliminate data redundancy, and the data integration and optimization module can also establish data lineage relationships for tracking data sources and historical changes;

[0074] The data compression and feature selection module is used to compress the integrated data, reduce the storage space requirement, while maintaining the data's feature representation ability, and the data compression and feature selection module also introduces a sparse coding method to reduce the feature dimension, improve data processing efficiency, and reduce computing resource consumption;

[0075] The data deduplication module introduces a hash function for data deduplication, quickly identifies and merges duplicate records, and at the same time, the data deduplication module uses a clustering algorithm to identify duplicate data and deduplicate the data through user behavior classification labels.

[0076] The security monitoring and guarantee module includes a monitoring and diagnosis module and an encryption module;

[0077] The monitoring and diagnosis module monitors data information in real time, discovers potential problems and bottlenecks in the data in a timely manner, and provides data support for optimization; the monitoring and diagnosis module also uses machine learning and statistical analysis methods to identify abnormal behaviors and fault patterns in the database module, improving the accuracy and timeliness of fault detection;

[0078] The encryption module performs static data encryption on users' sensitive information (such as ID numbers and bank card numbers), and uses the SSL / TLS protocol to encrypt the data transmission process.

[0079] The formula for the data integration and optimization module to establish data lineage is:

[0080] G = (V, E)

[0081] V represents the set of nodes, and E represents the set of edges.

[0082] The formula for the sparse coding method is:

[0083]

[0084] x is the input signal, D is the dictionary matrix, and s is the coefficient vector.

[0085] When the data deduplication module introduces a hash function for data deduplication, the formula for the hash function is:

[0086] h(k) = k mod m

[0087] k is the keyword for hash processing, and m is the size of the hash table.

[0088] The formula for machine learning and statistical analysis methods is:

[0089]

[0090] Cov(X, Y) is the covariance of X and Y, and σX and σY are the standard deviations of X and Y respectively;

[0091] The calculation formula for Cov(X, Y) is:

[0092]

[0093] and are the sample means of X and Y respectively.

[0094] Example 2:

[0095] In specific implementation, this processing system can be applied in a medical and health data analysis system, which is basically the same as Example 1, except that:

[0096] Requirement analysis and planning module

[0097] When a medical institution hopes to improve the accuracy of disease diagnosis and treatment effectiveness by analyzing patients' medical data. After communicating with medical experts and management personnel, the requirements analysis and planning module clarifies that the business requirement is to integrate data such as patients' medical records, examination reports, and medication records. The system goal is to build a medical big data analysis platform to provide support for clinical decision-making.

[0098] Data Comprehensive Collection and Integration Module

[0099] Data Identification Module: Determine that the data sources include the hospital's electronic medical record system, inspection and examination equipment, pharmacy management system, etc.

[0100] Data Collection Module: Use different collection tools for different data sources. For the electronic medical record system, extract data through the database interface; for inspection and examination equipment, use a data collector to collect data.

[0101] Data Integration Module: Integrate the multi-party data collected, perform preliminary cleaning and conversion on patients' basic information, diagnosis results, etc., and unify the data format.

[0102] Database Module

[0103] Data Storage Module: Use the relational database MySQL for physical storage to ensure data consistency and integrity; use database views for logical storage to facilitate data query and analysis.

[0104] Data Operation Module: Develop stored procedures to perform insert, delete, update, and query operations on the data in the database, such as updating medical record data according to patients' hospitalization information.

[0105] Data Management Module: Establish a data quality management system, responsible for managing and controlling the data in the database, including data quality assessment, data standardization, etc.

[0106] Data Security Module: Protect the data in the database from unauthorized access and modification through an access control list (ACL), and encrypt the storage of patients' privacy information.

[0107] Data Backup and Recovery Module: Develop a strategy of full backup every month and incremental backup every day, and use cloud storage for data backup. Conduct regular data recovery drills to ensure quick recovery in case of data loss or damage.

[0108] Data Comprehensive Processing and Analysis Module

[0109] Data Cleaning Module: Clean the medical data in the database, identify and correct errors and inconsistencies in the data, such as correcting data entry errors in inspection reports.

[0110] Data Integration Optimization Module: Integrates the cleaned data to eliminate data redundancy. Establishes data lineage, such as recording the transfer process of a patient's test results from the test equipment to the electronic medical record system, facilitating the tracing of data sources and historical changes.

[0111] Data Compression and Feature Selection Module: Uses the Gzip algorithm to compress the integrated data, reducing the storage space requirement. Introduces sparse coding methods to process the symptom features of patients, reducing the feature dimension and improving data processing efficiency.

[0112] Data Deduplication Module: Introduces a multiplicative hash function for data deduplication to quickly identify and merge duplicate patient records. Also uses the DBSCAN clustering algorithm to identify duplicate data and deduplicate data through the disease classification labels of patients.

[0113] Security Monitoring and Assurance Module

[0114] Monitoring and Diagnosis Module: Monitors data information in real-time, uses classification algorithms in machine learning (such as decision trees) and statistical analysis methods (such as correlation coefficient analysis) to identify abnormal behaviors and fault patterns in the database, and timely discovers potential problems and bottlenecks in the data processing process.

[0115] Encryption Module: Performs static data encryption on patients' sensitive information (such as medical records, diagnosis results), and uses VPN to encrypt the data transmission process.

[0116] In summary, through the use of the Data Comprehensive Acquisition and Integration Module and the Data Comprehensive Processing and Analysis Module to comprehensively integrate and process analyze data, this system can effectively remove noise, redundancy, and errors in data information, effectively improve the accuracy and effectiveness of data analysis. At the same time, through the integration and analysis of data by the Data Comprehensive Acquisition and Integration Module and the Data Comprehensive Processing and Analysis Module, it can effectively standardize the data from different data sources, improve data quality, reduce the possibility of data errors, further improve the accuracy and processing efficiency of data processing, and reduce the consumption of computing resources.

[0117] The above-described embodiments only represent the preferred embodiments of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations, improvements, and substitutions can be made, and these all belong to the protection scope of the present invention.

Claims

1. An artificial intelligence big data processing system, characterized in that: include: Requirements analysis and planning module: used to clarify business requirements and system goals; Data comprehensive collection and integration module: used to collect data from multiple parties and integrate the collected data from multiple parties; Database module: used to store and collect integrated multi-party data; Data comprehensive processing and analysis module: used to process and analyze the stored multi-party data to improve data quality; Security monitoring and assurance module: used to monitor and manage data information and encrypt data information.

2. The artificial intelligence big data processing system according to claim 1, characterized in that: The data comprehensive acquisition and integration module includes a data identification module, a data acquisition module and a data integration module: The data identification module is used to determine the source of data; The data collection module is used to select appropriate data collection tools according to the data source and characteristics, and use the data collection tools to extract and collect data information; The data integration module is used to integrate the extracted and collected data, and the data integration module will also perform preliminary cleaning and conversion on the integrated data.

3. The artificial intelligence big data processing system according to claim 2, characterized in that: The database module includes a data storage module, a data management module, a data security module, a data operation module and a data backup and recovery module: The data storage module is used for data storage and management, and the data storage module includes two parts: physical storage and logical storage. The physical storage is responsible for the physical storage of data, and the logical storage is responsible for the logical organization and management of data. The data operation module is used to perform various operations on the data in the database, such as inserting, deleting, updating and querying; The data management is responsible for managing and controlling the data in the database module; The data security module is used to protect the data in the database module from being illegally accessed and modified; The data backup and recovery module is responsible for regularly backing up the data in the database module and restoring the data when necessary. The data backup and recovery module includes the formulation of data backup strategy, the execution of data backup, the testing of data recovery and the implementation of data recovery.

4. The artificial intelligence big data processing system according to claim 1, characterized in that: The data comprehensive processing and analysis module includes a data cleaning module, a data integration optimization module, a data compression and feature selection module, and a data deduplication module; The data cleaning module is used to clean the data in the database, identify and correct errors and inconsistencies in the data; The data integration and optimization module is used to integrate the cleaned data and eliminate data redundancy. The data integration and optimization module can also establish data lineage relationships to track data sources and historical changes. The data compression and feature selection module is used to compress the integrated data, reduce the storage space requirement, and maintain the feature representation capability of the data. The data compression and feature selection module also introduces a sparse coding method to reduce feature dimensions, improve data processing efficiency, and reduce computing resource consumption; The data deduplication module introduces a hash function to perform data deduplication, quickly identify and merge duplicate records, and at the same time, the data deduplication module uses a clustering algorithm to identify duplicate data and deduplicates data through classification labels.

5. The artificial intelligence big data processing system according to claim 1, characterized in that: The safety monitoring and assurance module includes a monitoring and diagnosis module and an encryption module; The monitoring and diagnosis module monitors data information in real time, promptly discovers potential problems and bottlenecks in the data, and provides data support for optimization; the monitoring and diagnosis module also uses machine learning and statistical analysis methods to identify abnormal behaviors and failure modes in the database module, thereby improving the accuracy and timeliness of fault detection; The encryption module is used to encrypt sensitive data, including static data encryption and transmission data encryption.

6. The artificial intelligence big data processing system according to claim 4, characterized in that: The formula for establishing data kinship relationship in the data integration optimization module is: G=(V,E) V represents the node set and E represents the edge set.

7. The artificial intelligence big data processing system according to claim 4, characterized in that: The formula of the sparse coding method is: x is the input signal, D is the dictionary matrix, and s is the coefficient vector.

8. The artificial intelligence big data processing system according to claim 4, characterized in that: When the data deduplication module introduces a hash function to perform data deduplication, the formula of the hash function is: h(k(=k mod m k is the keyword for hash processing, and m is the size of the hash table.

9. The artificial intelligence big data processing system according to claim 4, characterized in that: The formula for the machine learning and statistical analysis method is: Cov(X,Y) is the covariance of X and Y, σX and σY are the standard deviations of X and Y respectively; The calculation formula of Cov(X,Y) is: and are the sample means of X and Y respectively.

Citation Information

Patent Citations

  • Face data set construction method and system

    CN114863525A

  • Data deduplication method and system of DBSCAN algorithm based on tolerable clustering deviation

    CN115994133A

  • Big data acquisition and analysis system

    CN117033501A

  • Data analysis processing system based on big data

    CN117851490A

  • Intelligent power grid data processing method based on NLP and dynamic consanguinity

    CN118585516A