Data processing method based on artificial intelligence

By introducing modules such as data preprocessing, feature selection and dimensionality reduction, data security and privacy protection, distributed processing and intelligent model training and evaluation in data processing, the challenges in the existing technology such as data quality, data imbalance, feature selection and dimensionality reduction, data security and privacy protection, data processing and storage efficiency, as well as model training and evaluation, are solved, and more efficient and more accurate data processing and analysis are achieved.

CN119988834APending Publication Date: 2025-05-13ENSHI VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510089965.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing AI-based data processing methods face many challenges in data quality, data imbalance, feature selection and dimensionality reduction, data security and privacy protection, data processing and storage efficiency, and model training and evaluation.

Method used

A data processing method based on artificial intelligence is proposed, including data preprocessing module, feature selection and dimensionality reduction module, data security and privacy protection module, distributed data processing and storage module and intelligent model training and evaluation module. This method solves multiple problems in data processing by filling vacant values, handling outliers, feature selection, data encryption, anonymization processing, differential privacy protection, distributed processing and model integration.

Benefits of technology

It significantly improves data quality and consistency, optimizes feature selection and dimensionality reduction effects, enhances data security and privacy protection capabilities, improves distributed data processing and storage efficiency, and improves intelligent model training and evaluation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988834A_ABST
    Figure CN119988834A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a data processing method based on artificial intelligence. Comprising the following steps: a) a data preprocessing module: performing primary processing on input data, including but not limited to filling vacant values, identifying and processing abnormal values, removing repeated data and unifying data formats, so as to ensure the quality and consistency of the data; b) a feature selection and dimension reduction module: screening out key features from original data and reducing data dimensions by using an automatic feature selection method; c) a data security and privacy protection module; d) a distributed data processing and storage module; e) an intelligent model training and evaluation module; and f) a feedback and optimization module. The invention provides the artificial intelligence-based data processing method which can improve the data preprocessing capability, optimize the feature selection and dimension reduction effect, enhance the data security and privacy protection capability, improve the distributed data processing and storage efficiency, improve the intelligent model training and evaluation performance and perfect the feedback and optimization mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data processing method based on artificial intelligence. Background Art

[0002] With the rapid development of information technology, data has become an indispensable and important resource in modern society. As a powerful tool for data processing and analysis, artificial intelligence has been widely used in many industries such as medicine, finance, manufacturing, and retail. However, existing data processing methods based on artificial intelligence still face a series of challenges and problems in practical applications, including but not limited to the following aspects:

[0003] 1. Data quality issues

[0004] Missing values, outliers, and duplicate values: There are often problems such as missing values, outliers, and duplicate values ​​in the raw data, which can seriously affect the accuracy and reliability of artificial intelligence algorithms.

[0005] Data inconsistency: Due to the diversity of data sources, data formats and attributes may be inconsistent, making data integration and processing more difficult.

[0006] 2. Data imbalance problem

[0007] Class imbalance: In classification tasks, the number of samples in some categories is significantly less than that in other categories, which makes traditional machine learning algorithms easily affected by the categories with more samples and reduces the prediction performance of minority categories.

[0008] Data scarcity: In some professional fields (such as medicine, law, finance, etc.), obtaining labeled data is time-consuming and expensive, and data scarcity has become a key factor limiting the application of AI.

[0009] 3. Feature selection and dimensionality reduction issues

[0010] Feature redundancy: The original data usually contains a lot of redundant information, which increases the complexity of the model and reduces the performance of the algorithm.

[0011] High-dimensional data: High-dimensional data not only increases computational complexity, but may also lead to model overfitting. Therefore, feature selection and dimensionality reduction are required to improve the efficiency and accuracy of the model.

[0012] 4. Data security and privacy protection issues

[0013] Data privacy: When processing sensitive data (such as user data, medical data, financial data, etc.), how to ensure data privacy and security has become an issue that cannot be ignored in AI development.

[0014] Compliance requirements: Data privacy laws (such as GDPR, CCPA, etc.) impose strict compliance requirements on the storage, processing, and sharing of data, increasing the difficulty and cost of data processing.

[0015] 5. Data processing and storage efficiency issues

[0016] Big data processing: As the amount of data continues to increase, how to efficiently process and store data has become a pressing issue to be solved.

[0017] Distributed storage and computing: In order to improve the efficiency of data processing and storage, distributed storage and computing technologies need to be adopted, but this increases the complexity of data management and maintenance.

[0018] 6. Model training and evaluation issues

[0019] Model training: Selecting an appropriate machine learning model for training is an important step in data processing. However, model training often faces challenges due to limitations in data quality and quantity.

[0020] Model evaluation: Model evaluation is a key step in judging model performance. However, due to problems such as data imbalance and data scarcity, the accuracy of model evaluation is often affected.

[0021] In summary, existing data processing methods based on artificial intelligence still face many challenges in terms of data quality, data imbalance, feature selection and dimensionality reduction, data security and privacy protection, data processing and storage efficiency, model training and evaluation, etc. In order to solve these problems, the present invention proposes a data processing method based on artificial intelligence. Summary of the invention

[0022] The problems to be solved by the present invention are missing values, abnormal values ​​and duplicate values ​​in the data, data inconsistency, data category imbalance, data scarcity, feature redundancy, high-dimensional data, poor data security and privacy protection, inability to efficiently process and store data, limitations on data quality and quantity, challenges faced by model training, which in turn affect evaluation.

[0023] The technical solution adopted by the present invention is: a data processing method based on artificial intelligence, comprising the following steps:

[0024] a) Data preprocessing module: First, the input data is preliminarily processed, including but not limited to filling in missing values, identifying and processing outliers, removing duplicate data, and unifying data formats to ensure data quality and consistency; at the same time, in order to address the problem of data imbalance, methods such as oversampling, undersampling, or synthetic minority oversampling techniques (such as SMOTE) are used to balance the number of samples in each category, and transfer learning techniques are used to transfer knowledge from related fields to alleviate data scarcity;

[0025] b) Feature selection and dimensionality reduction module: Use correlation analysis, recursive feature elimination (RFE), principal component analysis (PCA) or automatic feature selection methods based on deep learning to screen out key features from the original data and reduce data dimensions, reduce redundant information, and improve model training efficiency and prediction accuracy;

[0026] c) Data security and privacy protection module: In the process of data processing, data encryption, anonymization, differential privacy protection and other technologies are implemented to ensure the security and privacy of sensitive data, while complying with relevant laws, regulations and industry standards;

[0027] d) Distributed data processing and storage module: Use distributed file systems (such as HDFS) and distributed computing frameworks (such as Spark) to achieve efficient processing and storage of big data, and improve data processing speed and scalability through data sharding and parallel computing technology;

[0028] e) Intelligent model training and evaluation module: According to the data type and task requirements, select or design appropriate machine learning or deep learning models, and use preprocessed data for training; in the model evaluation stage, use cross-validation, confusion matrix, AUC-ROC curve and other evaluation indicators, combined with data balancing strategy, to comprehensively and accurately evaluate model performance; at the same time, use model integration technology (such as bagging, boosting) to improve model robustness and generalization ability;

[0029] f) Feedback and optimization module: Based on the model evaluation results, the model is optimized through methods such as gradient descent, hyperparameter tuning, grid search or random search, and new data is continuously collected. The model is updated through incremental learning or online learning mechanisms to adapt to data changes and maintain model performance.

[0030] As a further solution of the present invention: the specific steps of the data preprocessing module are:

[0031] a1) Data cleaning: preliminary processing of input data, including but not limited to:

[0032] Filling in missing values: Use mean filling, median filling, mode filling, forward / backward filling, interpolation, or prediction filling based on machine learning models to reasonably estimate and fill in missing values ​​in the data;

[0033] Identify and handle outliers: Use statistical methods (such as the 3σ principle, box plots), machine learning algorithms (such as isolation forests), or domain knowledge to identify outliers, and choose to delete, modify, or mark outliers based on actual conditions;

[0034] Remove duplicate data: Identify and remove duplicate records in a data set through methods such as hash functions, unique key detection, or similarity calculation;

[0035] Unified data format: Convert data from different sources and formats into a unified format, including data type conversion, date format unification, text encoding standardization, etc., to facilitate subsequent processing and analysis;

[0036] a2) Data balancing: To address the data imbalance problem, the following method is used to balance the number of samples in each category:

[0037] Oversampling: Duplicate or synthesize minority class samples to increase their number to match the number of majority class samples;

[0038] Undersampling: Randomly select some samples from the majority class samples and delete them to reduce their number to make it close to the number of minority class samples;

[0039] Synthetic minority oversampling techniques (such as SMOTE): Generate new synthetic samples based on minority samples. These synthetic samples are located at the linear interpolation points between minority samples and their nearest neighbor samples to increase the diversity of minority samples.

[0040] a3) Transfer learning technology: In the case of high data scarcity, transfer learning technology is used to transfer knowledge from related fields, including but not limited to:

[0041] Feature-based transfer: The feature representation learned from the source domain is transferred to the target domain to enhance the feature expression ability of the target domain data;

[0042] Model-based migration: Migrate the pre-trained model parameters or structure to the target domain as the starting point for training the target domain model, accelerate model convergence and improve performance;

[0043] Instance-based transfer: According to the similarity between the source domain and target domain data, part of the source domain data is selected as auxiliary data for target domain training to alleviate the data scarcity problem.

[0044] As a further solution of the present invention: the specific steps of the feature selection and dimensionality reduction module are:

[0045] b1) Data preprocessing stage:

[0046] Receive raw data sets as input; preprocess the raw data sets, including but not limited to missing value filling, outlier processing, data standardization or normalization, to ensure data quality and consistency;

[0047] b2) Feature selection stage:

[0048] Correlation analysis: Calculate the correlation index between each feature in the original data set and the target variable (such as Pearson correlation coefficient, Spearman rank correlation coefficient, etc.), and filter out the feature subset that is significantly correlated with the target variable according to the preset correlation threshold;

[0049] Recursive feature elimination (RFE): Using a basic machine learning model (such as support vector machine, random forest, etc.), iteratively removes the features that contribute the least to model performance until the predetermined number of features is reached or the model performance is no longer significantly improved;

[0050] Automatic feature selection based on deep learning: Utilize the automatic feature learning capability of deep learning models (such as neural networks, convolutional neural networks, recurrent neural networks, etc.) to automatically extract and select the most critical features for prediction tasks through the training process; this method may need to be combined with regularization techniques (such as L1 regularization) to induce sparse feature selection;

[0051] b3) Dimensionality reduction stage:

[0052] Principal Component Analysis (PCA): Perform principal component analysis on the data set after feature selection, project the data into a new low-dimensional space through linear transformation, and each dimension (principal component) of the new space retains the variance information of the original data as much as possible, while achieving data dimensionality reduction;

[0053] Other dimensionality reduction techniques (optional): Based on the data characteristics and requirements, other dimensionality reduction techniques such as linear discriminant analysis (LDA), t-distributed neighbor embedding (t-SNE), U-Map, etc. can be selected to further reduce the data dimension and retain key information;

[0054] b4) Output stage:

[0055] Output a dataset after feature selection and dimensionality reduction. This dataset has lower dimensions and less redundant information, which helps improve the training efficiency and prediction accuracy of subsequent machine learning models.

[0056] b5) Verification and optimization:

[0057] In the process of feature selection and dimensionality reduction, strategies such as cross-validation are used to evaluate the impact of different feature subsets and dimensionality reduction schemes on model performance to determine the optimal feature selection and dimensionality reduction strategy;

[0058] Based on the verification results, necessary adjustments and optimizations are made to the feature selection and dimensionality reduction modules to ensure that the final output data set can maximize the performance of the machine learning model.

[0059] As a further solution of the present invention: the specific steps of the data security and privacy protection module are:

[0060] c1) Data encryption

[0061] Key and certificate management: Strict permission management should be set for certificates, symmetric keys and asymmetric keys in the encryption mechanism to ensure that only authorized personnel can access and operate them; important keys (such as SMK and DMK) and certificates should be backed up off-site to prevent data loss; the validity period of keys and certificates should be checked and updated regularly to ensure that they are in a safe state;

[0062] Application of encryption technology: Advanced encryption technologies, such as AES and RSA, should be used during data transmission and storage to ensure the confidentiality and integrity of data; sensitive data, such as personal identity information and financial information, should be protected by higher-level encryption standards;

[0063] c2) Anonymization

[0064] Implementation of anonymization technology: Use anonymization methods such as generalization, compression, decomposition, substitution and interference to process personal data to reduce personal privacy risks; the anonymized data should meet the anonymization standards under legal semantics, that is, the data itself cannot point to a specific individual, and even if combined with other data, it cannot point to a specific individual;

[0065] Use of anonymized data: Anonymized data can be used in fields such as data analysis and scientific research, but personal privacy should not be leaked during use; institutions should separately and securely store specific information that can restore identity data from anonymized data to prevent data from being restored;

[0066] c3) Differential Privacy Protection

[0067] Differential privacy technology principle: Differential privacy protects personal privacy by adding random noise to the original data, making it impossible for the data analysis results to accurately point to a certain individual. Differential privacy technology is suitable for scenarios where large-scale data sets are analyzed, such as training machine learning models.

[0068] Differential privacy implementation strategy: Before data is released, differential privacy processing should be performed on the original data to ensure the privacy of the data; during the data analysis process, differential privacy algorithms should be used to protect the query results to prevent personal privacy leakage;

[0069] c4) Comply with relevant laws, regulations and industry standards

[0070] Compliance with laws and regulations: Relevant laws and regulations should be strictly observed to ensure the legality of data processing activities; for cross-border data transfer, the laws and regulations of relevant countries and regions, such as the EU's GDPR, should be observed;

[0071] Implementation of industry standards: Information security and privacy information management system standards such as ISO 27001 and ISO 29151 should be followed to ensure the standardization and effectiveness of data security and privacy protection measures; in specific industries (such as automobiles, finance, etc.), the data protection and privacy requirements of the industry, such as ISO 21434, should also be followed;

[0072] c5) Comprehensive measures and continuous improvement

[0073] Implementation of comprehensive measures: Data encryption, anonymization, differential privacy protection and other technologies should be combined to form a comprehensive data security and privacy protection framework; a risk management plan for data security and privacy protection should be established to identify risks through risk assessment and define the organization's risk tolerance and risk preferences;

[0074] Continuous improvement and optimization: Data security and privacy protection measures should be audited and evaluated regularly to ensure their effectiveness and compliance; data security and privacy protection measures should be continuously optimized and improved based on technological development and business needs to adapt to the ever-changing security environment.

[0075] As a further solution of the present invention: the specific steps of the distributed data processing and storage module are:

[0076] d1) Distributed file system: Hadoop Distributed File System (HDFS) or equivalent distributed storage technology is used to achieve efficient and reliable storage of big data; the distributed file system has automatic data sharding, replication and fault tolerance mechanisms to ensure high availability and persistence of data; data is divided into multiple data blocks and stored on multiple physical nodes, and each data block has multiple copies distributed on different nodes to improve data reliability and reading efficiency;

[0077] d2) Distributed computing framework: Apache Spark or equivalent distributed computing technology is used to achieve efficient processing of big data; the distributed computing framework supports data parallel processing and memory computing, which can significantly improve data processing speed and efficiency; through data parallel processing technology, large-scale data sets are divided into multiple subsets and processed in parallel on multiple nodes; through memory computing technology, data that is frequently accessed during the processing process is stored in memory to reduce disk I / O operations and improve processing speed;

[0078] d3) Data sharding and load balancing: The module has a data sharding mechanism that can reasonably divide data into multiple data blocks according to the data volume and processing requirements, and distribute them to different nodes for storage and processing; at the same time, the module also has a load balancing function that can dynamically adjust the distribution of data blocks according to the processing capacity of each node to ensure balanced distribution and efficient execution of data processing tasks;

[0079] d4) Scalability and fault tolerance: The module supports dynamic expansion and can add new nodes as needed to expand storage and processing capabilities. At the same time, the module has a fault-tolerant mechanism that can automatically migrate data to other normal nodes when a node fails, ensuring the continuity and reliability of data processing.

[0080] d5) Data processing interfaces and tools: The module provides a wealth of data processing interfaces and tools, supporting a variety of data processing operations such as data cleaning, conversion, aggregation, and analysis; users can use these interfaces and tools to easily write data processing programs to meet complex data processing requirements.

[0081] As a further solution of the present invention: the specific steps of the intelligent model training and evaluation module are:

[0082] e1) Model selection and design: Based on the type of input data (such as images, text, time series, etc.) and specific task requirements (such as classification, regression, clustering, prediction, etc.), the module can automatically select or design appropriate machine learning or deep learning models; models include but are not limited to support vector machines, decision trees, random forests, neural networks, convolutional neural networks, recurrent neural networks, etc.;

[0083] e2) Data preprocessing and enhancement: Before model training, the module uses preprocessed data for training; preprocessing steps may include data cleaning, normalization, feature selection, feature scaling, data enhancement (such as image flipping, rotation, cropping, etc.) to improve data quality and model training effect;

[0084] e3) Model training and optimization: The module uses appropriate training algorithms and optimization strategies to train the model, such as gradient descent, stochastic gradient descent, Adam optimizer, etc. At the same time, the module supports hyperparameter tuning, and finds the optimal hyperparameter combination through grid search, random search or Bayesian optimization to improve the performance and generalization ability of the model;

[0085] e4) Model evaluation and validation: During the model evaluation phase, the module uses a variety of evaluation indicators to conduct a comprehensive and accurate evaluation of the model, including but not limited to accuracy, recall, F1 score, cross-validation, confusion matrix, AUC-ROC curve, etc. At the same time, the module considers data balancing strategies, such as processing unbalanced data through resampling, weighted loss function, etc., to avoid overfitting of the model to the majority class samples;

[0086] e5) Model integration and improvement: The module uses model integration techniques (such as bagging, boosting, stacking, etc.) to improve the robustness and generalization ability of the model; by training multiple base models and combining their prediction results, the stability and accuracy of the overall model can be improved;

[0087] e6) Performance monitoring and tuning: The module monitors the performance indicators of the model in real time during the training process, such as loss function value, accuracy, etc., and performs necessary tuning operations based on the monitoring results. At the same time, the module supports model visualization, and helps users better understand the behavior and performance of the model by drawing charts such as training curves and feature importance graphs.

[0088] As a further solution of the present invention: the specific steps of the feedback and optimization module are:

[0089] f1) Feedback on model evaluation results: Based on the model evaluation results output by the intelligent model training and evaluation module, the module can identify the deficiencies of the model, such as overfitting, underfitting, large deviation or large variance;

[0090] f2) Model optimization strategy: The module can automatically select or recommend appropriate optimization strategies for the feedback of model evaluation results; these strategies include but are not limited to:

[0091] Gradient descent and its variants: such as standard gradient descent, stochastic gradient descent, mini-batch gradient descent, momentum gradient descent, Adam optimizer, etc., are used to adjust model parameters to minimize the loss function;

[0092] Hyperparameter tuning: Use grid search, random search, Bayesian optimization and other methods to find the optimal hyperparameter combination in the preset hyperparameter space to improve model performance;

[0093] Model structure adjustment: such as increasing or decreasing the number of network layers, changing the number of neurons, adjusting the activation function, etc., to improve the model's expressiveness and generalization capabilities;

[0094] f3) Continuous data collection: The module can continuously collect new data, which may come from real-time data streams, user feedback, newly collected samples, etc. The collection of new data helps the model better adapt to data changes and improve the practicality and accuracy of the model;

[0095] f4) Incremental or online learning mechanisms: Using newly collected data, the module can use incremental learning or online learning mechanisms to update the model; incremental learning allows the model to gradually learn new data while retaining the original knowledge; online learning allows the model to be continuously updated in real-time data streams to adapt to data changes; both mechanisms help maintain the performance and accuracy of the model;

[0096] f5) Performance monitoring and alarm: During the model optimization and update process, the module can monitor the model's performance indicators in real time, such as accuracy, recall rate, F1 score, etc.; when the performance indicators are lower than the preset threshold, the module can automatically trigger the alarm mechanism to remind the user to make further optimization or adjustments.

[0097] As a further solution of the present invention: the intelligent model training and evaluation module supports the automated machine learning (AutoML) framework, which can automatically complete the model selection, parameter tuning and evaluation process, reduce manual intervention and improve data processing efficiency.

[0098] As a further solution of the present invention: the method is applied to but not limited to data processing and analysis tasks in multiple fields such as finance, medical care, education, e-commerce, and the Internet of Things.

[0099] Beneficial effects of the present invention:

[0100] 1. Improvement of data preprocessing capabilities:

[0101] By filling missing values, identifying and processing outliers, removing duplicate data, and unifying data formats, the quality and consistency of the data are significantly improved, laying a solid foundation for subsequent data processing and analysis.

[0102] To address the data imbalance problem, methods such as oversampling, undersampling or synthetic minority class oversampling technology are used to effectively balance the number of samples in each category and improve the prediction performance of the model, especially when dealing with minority category samples.

[0103] By using transfer learning technology to transfer knowledge from related fields, the challenges brought by data scarcity are effectively alleviated and the generalization ability of the model is improved.

[0104] 2. Optimization of feature selection and dimensionality reduction effect:

[0105] Through correlation analysis, recursive feature elimination (RFE), principal component analysis (PCA) or automatic feature selection methods based on deep learning, key features are successfully screened out and data dimensions are reduced, redundant information is reduced, and model training efficiency and prediction accuracy are improved.

[0106] During the feature selection and dimensionality reduction process, strategies such as cross-validation were used to evaluate the impact of different feature subsets and dimensionality reduction schemes on model performance, ensuring the formulation of the optimal feature selection and dimensionality reduction strategy.

[0107] 3. Enhanced data security and privacy protection capabilities:

[0108] During the data processing process, technologies such as data encryption, anonymization, and differential privacy protection are implemented to effectively ensure the security and privacy of sensitive data, comply with relevant laws, regulations, and industry standards, and reduce the risk of data leakage and abuse.

[0109] A risk management plan for data security and privacy protection has been established, and the standardization and effectiveness of data security and privacy protection measures have been ensured through risk assessment, continuous improvement and optimization.

[0110] 4. Improvement of distributed data processing and storage efficiency:

[0111] The use of distributed file systems (such as HDFS) and distributed computing frameworks (such as Spark) has enabled efficient processing and storage of big data, and the data processing speed and scalability have been significantly improved through data sharding and parallel computing technology.

[0112] The distributed data processing and storage module supports dynamic expansion and fault-tolerant mechanisms, ensuring the continuity and reliability of data processing and reducing the risk of data loss due to node failures.

[0113] 5. Improvements in intelligent model training and evaluation performance:

[0114] Automatically select or design appropriate machine learning or deep learning models based on data type and task requirements, improving the adaptability and accuracy of the models.

[0115] During the model evaluation phase, a variety of evaluation indicators and data balancing strategies are used to comprehensively and accurately evaluate the model performance, avoiding overfitting of the model to majority class samples.

[0116] Model integration technology is used to improve the robustness and generalization ability of the model. By training multiple base models and combining their prediction results, the stability and accuracy of the overall model are improved.

[0117] 6. Improvement of feedback and optimization mechanism:

[0118] According to the model evaluation results, the appropriate optimization strategy is automatically selected or recommended to optimize the model, thereby improving the performance and accuracy of the model.

[0119] Continuously collect new data and update the model through incremental learning or online learning mechanisms, so that the model can better adapt to data changes and maintain the performance and accuracy of the model.

[0120] The performance monitoring and alarm mechanism can monitor the performance indicators of the model in real time, and automatically trigger an alarm when the performance indicators are lower than the preset threshold, reminding the user to make further optimization or adjustments.

[0121] 7. Support for Automated Machine Learning (AutoML) framework:

[0122] The intelligent model training and evaluation module supports the automated machine learning (AutoML) framework, which can automatically complete the model selection, parameter tuning and evaluation processes, reducing manual intervention and improving data processing efficiency.

[0123] 8. Wide range of applications:

[0124] The method of the present invention is not only applicable to data processing and analysis tasks in multiple fields such as finance, medical treatment, education, e-commerce, and the Internet of Things, but can also be flexibly adjusted and expanded according to the characteristics and needs of different fields, and has broad application prospects and market demands. BRIEF DESCRIPTION OF THE DRAWINGS

[0125] Figure 1 The figure is a schematic diagram of the overall structure of an artificial intelligence-based data processing method of the present invention. DETAILED DESCRIPTION

[0126] The present invention will be further described below.

[0127] Example 1

[0128] 1. Implementation of Data Preprocessing Module

[0129] 1. Data cleaning

[0130] Filling missing values: Use mean filling, median filling, or mode filling to fill missing values. For time series data, you can consider using forward / backward filling or interpolation. In certain cases, you can also use machine learning models (such as regression models) for prediction filling.

[0131] Identify and handle outliers: Use statistical methods (such as the 3σ principle and box plots) to identify outliers. For outliers that are significantly deviated from the normal range, delete, modify, or mark them according to the actual situation. At the same time, machine learning algorithms (such as isolation forests) can also be used for outlier detection.

[0132] Remove duplicate data: Identify and remove duplicate records in a data set through hash functions, unique key detection, or similarity calculation to ensure the uniqueness of the data.

[0133] Unified data format: Convert data from different sources and formats into a unified format, including data type conversion, date format unification, text encoding standardization, etc., to facilitate subsequent processing and analysis.

[0134] 2. Data balancing

[0135] Oversampling: Duplicate or synthesize minority class samples (such as the SMOTE algorithm) to increase their number to match the number of majority class samples.

[0136] Undersampling: Randomly select some samples from the majority class and delete them to reduce their number to make it close to the number of minority class samples. However, care should be taken to avoid losing important information.

[0137] Transfer learning: In the case of high data scarcity, transfer learning techniques are used to transfer knowledge from related fields. For example, feature-based transfer can transfer feature representations learned in the source domain to the target domain; model-based transfer can transfer pre-trained model parameters or structures to the target domain as a starting point; instance-based transfer can transfer based on the similarity between the source and target domain data.

[0138] 2. Implementation of Feature Selection and Dimensionality Reduction Module

[0139] 1. Data preprocessing stage: Perform preprocessing operations such as missing value filling, outlier processing, data standardization or normalization on the original data set to ensure the quality and consistency of the data.

[0140] 2. Feature selection stage

[0141] Correlation analysis: Calculate the correlation index between each feature and the target variable (such as Pearson correlation coefficient, Spearman rank correlation coefficient, etc.), and filter out the feature subset that is significantly correlated with the target variable based on the preset correlation threshold.

[0142] Recursive Feature Elimination (RFE): It uses basic machine learning models such as support vector machines and random forests to iteratively remove the features that contribute the least to model performance.

[0143] Automatic feature selection based on deep learning: Utilize the automatic feature learning capabilities of deep learning models such as neural networks and convolutional neural networks to automatically extract and select the most critical features for prediction tasks through the training process.

[0144] 3. Dimensionality reduction stage

[0145] Principal Component Analysis (PCA): Perform principal component analysis on the dataset after feature selection, project the data into a new low-dimensional space, and retain as much variance information of the original data as possible.

[0146] Other dimensionality reduction techniques: Depending on the data characteristics and requirements, you can choose linear discriminant analysis (LDA), t-distributed neighbor embedding (t-SNE), U-Map and other dimensionality reduction techniques.

[0147] 4. Validation and optimization: Use strategies such as cross-validation to evaluate the impact of different feature subsets and dimensionality reduction schemes on model performance to determine the optimal feature selection and dimensionality reduction strategy.

[0148] 3. Implementation of Data Security and Privacy Protection Module

[0149] 1. Data encryption

[0150] Key and certificate management: Strict permission management should be implemented for certificates, symmetric keys, and asymmetric keys in encryption mechanisms to ensure that only authorized personnel can access and operate them. Important keys and certificates should be backed up off-site, and their validity periods should be checked and updated regularly.

[0151] Application of encryption technology: During data transmission and storage, advanced encryption technologies such as AES and RSA are used to ensure the confidentiality and integrity of data. For sensitive data, higher-level encryption standards are used for protection.

[0152] 2. Anonymization

[0153] Implementation of anonymization technology: Use anonymization methods such as generalization, compression, decomposition, substitution and interference to process personal data to reduce personal privacy risks. The anonymized data should meet the anonymization standards under legal semantics.

[0154] Use of anonymized data: Anonymized data can be used in fields such as data analysis and scientific research, but it should be ensured that personal privacy will not be leaked during use.

[0155] 3. Differential Privacy Protection

[0156] The principle of differential privacy technology: By adding random noise to the original data, the data analysis results cannot be accurately pointed to a certain individual.

[0157] Differential privacy implementation strategy: Before data is released and during data analysis, differential privacy algorithms are used to process data to protect personal privacy.

[0158] 4. Comply with relevant laws, regulations and industry standards: Strictly comply with relevant laws, regulations and information security and privacy information management system standards (such as ISO 27001, ISO 29151, etc.) to ensure the legality and standardization of data processing activities.

[0159] 4. Implementation of Distributed Data Processing and Storage Module

[0160] 1. Distributed file system: Use Hadoop Distributed File System (HDFS) or equivalent technology to achieve efficient and reliable storage of big data. Data is divided into multiple data blocks and stored on multiple physical nodes. Each data block has multiple copies distributed on different nodes to improve reliability and reading efficiency.

[0161] 2. Distributed computing framework: Use Apache Spark or equivalent technology to achieve efficient processing of big data. Support data parallel processing and memory computing technology to significantly improve data processing speed and efficiency.

[0162] 3. Data sharding and load balancing: Data is reasonably divided into multiple data blocks according to data volume and processing requirements and distributed to different nodes for storage and processing. At the same time, the distribution of data blocks is dynamically adjusted according to the processing capacity of each node to ensure balanced distribution and efficient execution of tasks.

[0163] 4. Scalability and fault tolerance: Supports dynamic expansion of new nodes to expand storage and processing capabilities. It has a fault-tolerant mechanism that automatically migrates data to other normal nodes when a node fails to ensure continuity and reliability.

[0164] 5. Data processing interfaces and tools: Provides a variety of data processing interfaces and tools to support data cleaning, conversion, aggregation, analysis and other operations. Users can use these interfaces and tools to easily write data processing programs to meet complex needs.

[0165] 5. Implementation of Intelligent Model Training and Evaluation Module

[0166] 1. Model selection and design: Automatically select or design appropriate machine learning or deep learning models based on data types and task requirements. Models include but are not limited to support vector machines, decision trees, random forests, neural networks, etc.

[0167] 2. Data preprocessing and enhancement: Use preprocessed data for training before model training. Preprocessing steps may include data cleaning, normalization, feature selection, etc. For image data, data enhancement operations such as flipping and rotation can be considered to improve model training results.

[0168] 3. Model training and optimization: Use appropriate training algorithms and optimization strategies to train the model, such as gradient descent, Adam optimizer, etc. Support hyperparameter tuning, and use grid search, random search and other methods to find the optimal hyperparameter combination to improve model performance.

[0169] 4. Model evaluation and verification: Use a variety of evaluation indicators to comprehensively and accurately evaluate the model, including accuracy, recall, F1 score, etc. At the same time, consider data balancing strategies to avoid overfitting of the model to majority class samples.

[0170] 5. Model integration and improvement: Use model integration techniques (such as bagging, boosting, etc.) to improve the robustness and generalization ability of the model. Improve the stability and accuracy of the overall model by training multiple base models and combining their prediction results.

[0171] 6. Performance monitoring and tuning: Monitor the performance indicators of the model in real time and perform necessary tuning operations based on the monitoring results. Support model visualization to help users better understand model behavior and performance.

[0172] 6. Implementation of the Feedback and Optimization Module

[0173] 1. Model evaluation result feedback: Identify the deficiencies of the model, such as overfitting, underfitting, etc., based on the model evaluation results output by the intelligent model training and evaluation module.

[0174] 2. Model optimization strategy: Automatically select or recommend appropriate optimization strategies based on the feedback from model evaluation results, such as gradient descent variants, hyperparameter tuning, model structure adjustment, etc.

[0175] 3. Continuous data collection: Continuously collect new data, which may come from real-time data streams, user feedback, etc. The collection of new data helps the model better adapt to data changes and improve practicality and accuracy.

[0176] 4. Incremental learning or online learning mechanism: Use incremental learning or online learning mechanism to update the model using newly collected data. Incremental learning allows the model to gradually learn new data while retaining the original knowledge; online learning allows the model to be continuously updated in real-time data streams to adapt to data changes.

[0177] 5. Performance monitoring and alarm: Monitor the performance indicators of the model in real time and automatically trigger the alarm mechanism to remind users to make further optimization or adjustments when the performance indicators are lower than the preset threshold.

[0178] Through the above specific implementation methods, the data processing method based on artificial intelligence of the present invention can significantly improve data quality, optimize feature selection and dimensionality reduction effects, enhance data security and privacy protection capabilities, improve distributed data processing and storage efficiency, improve intelligent model training and evaluation performance, and improve feedback and optimization mechanisms. This method is applicable to data processing and analysis tasks in multiple fields such as finance, medical care, education, e-commerce, and the Internet of Things, and has broad application prospects and market demand.

[0179] Example 2

[0180] 1. Implementation of Data Preprocessing Module

[0181] 1. Data cleaning and formatting

[0182] Missing value processing: First, identify the missing values ​​in the input data set. For numerical data, use mean filling, median filling, or mode filling methods to ensure data integrity. For categorical data, choose the most likely category to fill based on context or domain knowledge, or use a predictive filling method based on a machine learning model.

[0183] Outlier processing: Statistical methods (such as the 3σ principle and box plots) and machine learning algorithms (such as isolation forests) are combined to identify outliers in the data set. For outliers that are significantly deviated from the normal range, they are deleted, corrected to normal values, or marked as special values ​​according to the actual situation to avoid interference with subsequent analysis.

[0184] Duplicate data removal: Identify and remove duplicate records in a data set through methods such as hash functions or unique key detection to ensure the uniqueness and accuracy of the data.

[0185] Data format unification: Convert data from different sources and formats into a unified format, including data type conversion (such as converting strings to numeric types), date format unification (such as converting dates in different formats to a unified YYYY-MM-DD format), text encoding standardization, etc., to facilitate subsequent processing and analysis.

[0186] 2. Data balancing

[0187] Oversampling and undersampling: To address the data imbalance problem, oversampling technology is used to increase the number of minority class samples, or undersampling technology is used to reduce the number of majority class samples to achieve class balance. Specific methods include randomly copying minority class samples, synthesizing new minority class samples based on the SMOTE algorithm, or randomly selecting some samples from the majority class samples for deletion.

[0188] Transfer learning: When data is scarce, transfer learning techniques are used to transfer knowledge from related fields. The knowledge learned in the source domain is transferred to the target domain through methods such as feature-based transfer, model-based transfer, or instance-based transfer to enhance the feature expression ability of the target domain data or improve the model performance.

[0189] 2. Implementation of Feature Selection and Dimensionality Reduction Module

[0190] 1. Data preprocessing stage: Receive the raw data set that has been initially cleaned and formatted as input, and further fill in missing values, handle outliers, and standardize or normalize the data set to ensure data quality and consistency.

[0191] 2. Feature selection stage:

[0192] Correlation analysis: Calculate the correlation index between each feature in the original data set and the target variable (such as Pearson correlation coefficient, Spearman rank correlation coefficient, etc.), and filter out the feature subset that is significantly correlated with the target variable based on the preset correlation threshold.

[0193] Recursive feature elimination (RFE): Using machine learning models such as support vector machines and random forests as base models, the features that contribute the least to model performance are gradually removed in a recursive manner until the predetermined number of features is reached or the model performance is no longer significantly improved.

[0194] Automatic feature selection based on deep learning: Utilize the automatic feature learning capabilities of deep learning models such as neural networks and convolutional neural networks to automatically extract and select the most critical features for prediction tasks through the training process.

[0195] 3. Dimensionality reduction stage: Perform principal component analysis (PCA) or other dimensionality reduction techniques (such as LDA, t-SNE, etc.) on the data set after feature selection to project the data into a new low-dimensional space while retaining the variance information and key information of the original data as much as possible.

[0196] 4. Output stage: Output the data set after feature selection and dimensionality reduction. This data set has lower dimensions and less redundant information, which helps to improve the training efficiency and prediction accuracy of subsequent machine learning models.

[0197] 3. Implementation of Data Security and Privacy Protection Module

[0198] 1. Data encryption: During data transmission and storage, advanced encryption technologies such as AES and RSA are used to ensure data confidentiality and integrity. Sensitive data (such as personal identity information, financial information, etc.) is protected by a higher level of encryption standards.

[0199] 2. Anonymization: Use anonymization methods such as generalization, compression, decomposition, substitution, and interference to process personal data to reduce personal privacy risks. Ensure that the anonymized data meets the anonymization standards under legal semantics.

[0200] 3. Differential privacy protection: Before data is released and during data analysis, differential privacy technology is used to add random noise to the original data, so that the data analysis results cannot be accurately pointed to a certain individual, thereby protecting personal privacy.

[0201] 4. Compliance with laws and regulations and implementation of industry standards: Strictly comply with relevant laws and regulations and industry standards to ensure the legality of data processing activities. For cross-border data transfer, comply with the laws and regulations of relevant countries and regions.

[0202] 4. Implementation of Distributed Data Processing and Storage Module

[0203] 1. Distributed file system: Use Hadoop Distributed File System (HDFS) or equivalent distributed storage technology to achieve efficient and reliable storage of big data. Use automatic data sharding, replication and fault tolerance mechanisms to ensure high availability and persistence of data.

[0204] 2. Distributed computing framework: Use Apache Spark or equivalent distributed computing technology to achieve efficient processing of big data. Through data parallel processing and memory computing technology, the data processing speed and efficiency are significantly improved.

[0205] 3. Data sharding and load balancing: According to the data volume and processing requirements, the data is reasonably divided into multiple data blocks and allocated to different nodes for storage and processing. At the same time, the allocation of data blocks is dynamically adjusted according to the processing capacity of each node to ensure balanced distribution and efficient execution of data processing tasks.

[0206] 4. Scalability and fault tolerance: It supports dynamic expansion and can add new nodes as needed to expand storage and processing capabilities. It has a fault-tolerant mechanism that can automatically migrate data to other normal nodes when a node fails, ensuring the continuity and reliability of data processing.

[0207] 5. Implementation of Intelligent Model Training and Evaluation Module

[0208] 1. Model selection and design: Automatically select or design appropriate machine learning or deep learning models based on the type of input data and specific task requirements. Models include but are not limited to support vector machines, decision trees, random forests, neural networks, etc.

[0209] 2. Data preprocessing and enhancement: Before model training, use preprocessed data for training. Preprocessing steps may include data cleaning, normalization, feature selection, feature scaling, data enhancement, etc.

[0210] 3. Model training and optimization: Use appropriate training algorithms and optimization strategies to train the model, such as gradient descent, stochastic gradient descent, etc. At the same time, support hyperparameter tuning, and find the optimal hyperparameter combination through grid search, random search and other methods.

[0211] 4. Model evaluation and verification: Use a variety of evaluation indicators to conduct a comprehensive and accurate evaluation of the model, including but not limited to accuracy, recall, F1 score, etc. At the same time, consider data balancing strategies to avoid overfitting of the model to the majority class samples.

[0212] 5. Model integration and improvement: Use model integration techniques (such as bagging, boosting, etc.) to improve the robustness and generalization ability of the model.

[0213] 6. Implementation of the Feedback and Optimization Module

[0214] 1. Feedback on model evaluation results: Identify the deficiencies of the model based on the model evaluation results output by the intelligent model training and evaluation module.

[0215] 2. Model optimization strategy: Automatically select or recommend appropriate optimization strategies, such as gradient descent variants, hyperparameter tuning, model structure adjustment, etc., based on the feedback from model evaluation results.

[0216] 3. Continuous data collection: Continuously collect new data, including real-time data streams, user feedback, etc., to update the model and adapt to data changes.

[0217] 4. Incremental learning or online learning mechanism: Using newly collected data, incremental learning or online learning mechanism is used to update the model to maintain the performance and accuracy of the model.

[0218] 5. Performance monitoring and alarm: Monitor the performance indicators of the model in real time and automatically trigger the alarm mechanism when the performance indicators are lower than the preset threshold.

[0219] Through the above specific implementation methods, the present invention can significantly improve the efficiency and accuracy of data processing and analysis, while ensuring the security and privacy protection of data.

[0220] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data processing method based on artificial intelligence, characterized in that: The following steps are involved: a) Data preprocessing module: First, the input data is preliminarily processed, including but not limited to filling in missing values, identifying and processing outliers, removing duplicate data, and unifying data formats to ensure data quality and consistency; at the same time, in order to address the problem of data imbalance, oversampling, undersampling, or synthetic minority oversampling techniques are used to balance the number of samples in each category, and transfer learning techniques are used to transfer knowledge from related fields to alleviate data scarcity; b) Feature selection and dimensionality reduction module: Use correlation analysis, recursive feature elimination, principal component analysis or automatic feature selection methods based on deep learning to screen out key features from the original data and reduce data dimensions, reduce redundant information, and improve model training efficiency and prediction accuracy; c) Data security and privacy protection module: In the process of data processing, data encryption, anonymization, and differential privacy protection technologies are implemented to ensure the security and privacy of sensitive data, while complying with relevant laws, regulations, and industry standards; d) Distributed data processing and storage module: Use distributed file systems and distributed computing frameworks to achieve efficient processing and storage of big data, and improve data processing speed and scalability through data sharding and parallel computing technologies; e) Intelligent model training and evaluation module: According to the data type and task requirements, select or design a suitable machine learning or deep learning model, and use the preprocessed data for training; in the model evaluation stage, use cross-validation, confusion matrix, AUC-ROC curve and other evaluation indicators, combined with data balancing strategy, to comprehensively and accurately evaluate the model performance; at the same time, use model integration technology to improve the model robustness and generalization ability; f) Feedback and optimization module: Based on the model evaluation results, the model is optimized through gradient descent, hyperparameter tuning, grid search or random search methods, and new data is continuously collected. The model is updated through incremental learning or online learning mechanisms to adapt to data changes and maintain model performance.

2. The data processing method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the data preprocessing module are: a1) Data cleaning: preliminary processing of input data, including but not limited to: Filling missing values: Use mean filling, median filling, mode filling, forward / backward filling, interpolation or prediction filling methods based on machine learning models to reasonably estimate and fill missing values ​​in the data; Identify and handle outliers: Use statistical methods, machine learning algorithms, or domain knowledge to identify outliers, and choose to delete, modify, or mark outliers based on actual conditions; Remove duplicate data: Identify and remove duplicate records in a data set through hash functions, unique key detection, or similarity calculation methods; Unified data format: Convert data from different sources and formats into a unified format, including data type conversion, date format unification, and text encoding standardization, to facilitate subsequent processing and analysis; a2) Data balancing: To address the data imbalance problem, the following methods are used to balance the number of samples in each category: Oversampling: Duplicate or synthesize minority class samples to increase their number to make it equal to the number of majority class samples; Undersampling: Randomly select some samples from the majority class samples and delete them to reduce their number to make it close to the number of minority class samples; Synthetic minority oversampling technology: Generate new synthetic samples based on minority samples. These synthetic samples are located at the linear interpolation points between minority samples and their nearest neighbor samples to increase the diversity of minority samples. a3) Transfer learning technology: In the case of high data scarcity, transfer learning technology is used to transfer knowledge from related fields, including but not limited to: Feature-based transfer: The feature representation learned from the source domain is transferred to the target domain to enhance the feature expression ability of the target domain data; Model-based migration: Migrate the pre-trained model parameters or structure to the target domain as the starting point for training the target domain model, accelerate model convergence and improve performance; Instance-based transfer: According to the similarity between the source domain and target domain data, part of the source domain data is selected as auxiliary data for target domain training to alleviate the data scarcity problem.

3. The data processing method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the feature selection and dimensionality reduction module are as follows: b1) Data preprocessing stage: Receive raw data sets as input; preprocess the raw data sets, including but not limited to missing value filling, outlier processing, data standardization or normalization, to ensure data quality and consistency; b2) Feature selection stage: Correlation analysis: Calculate the correlation index between each feature in the original data set and the target variable, and filter out the feature subset that is significantly correlated with the target variable based on the preset correlation threshold; Recursive feature elimination: Take a basic machine learning model and iteratively remove the features that contribute the least to the model performance until a predetermined number of features is reached or the model performance is no longer significantly improved. Automatic feature selection based on deep learning: Utilize the automatic feature learning capability of deep learning models to automatically extract and select the most critical features for prediction tasks through the training process; this method needs to be combined with regularization techniques to induce sparse feature selection; b3) Dimensionality reduction stage: Principal component analysis: Perform principal component analysis on the data set after feature selection, project the data into a new low-dimensional space through linear transformation, and each dimension of the new space retains the variance information of the original data while achieving data dimensionality reduction; Other dimensionality reduction techniques: Based on the data characteristics and requirements, other dimensionality reduction techniques such as linear discriminant analysis, t-distribution neighborhood embedding, and U-Map are selected to further reduce the data dimension and retain key information; b4) Output stage: Output a dataset after feature selection and dimensionality reduction. This dataset has lower dimensions and less redundant information, which helps improve the training efficiency and prediction accuracy of subsequent machine learning models. b5) Verification and optimization: In the process of feature selection and dimensionality reduction, a cross-validation strategy is used to evaluate the impact of different feature subsets and dimensionality reduction schemes on model performance to determine the optimal feature selection and dimensionality reduction strategy; Based on the verification results, necessary adjustments and optimizations are made to the feature selection and dimensionality reduction modules to ensure that the final output data set can maximize the performance of the machine learning model.

4. The data processing method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the data security and privacy protection module are as follows: c1) Data encryption Key and certificate management: Strict permission management should be set for certificates, symmetric keys and asymmetric keys in encryption mechanisms to ensure that only authorized personnel can access and operate them; important keys and certificates should be backed up off-site to prevent data loss; The validity period of keys and certificates should be checked and updated regularly to ensure they are in a secure state; Application of encryption technology: Advanced encryption technology should be used during data transmission and storage to ensure the confidentiality and integrity of data; For sensitive data, a higher level of encryption standards should be used for protection; c2) Anonymization Implementation of anonymization technology: Use generalization, compression, decomposition, substitution and interference anonymization methods to process personal data to reduce personal privacy risks; the anonymized data should meet the anonymization standards under legal semantics, that is, the data itself cannot point to a specific individual, and even if combined with other data, it cannot point to a specific individual; Use of anonymized data: Anonymized data is used in data analysis and scientific research, but personal privacy should not be leaked during use; institutions should separately and securely store specific information that can restore identity data from anonymized data to prevent data from being restored; c3) Differential Privacy Protection Differential privacy technology principle: Differential privacy protects personal privacy by adding random noise to the original data, making it impossible for the data analysis results to accurately point to a certain individual; Differential privacy technology is suitable for scenarios where large-scale data sets are analyzed; Differential privacy implementation strategy: Before data is released, differential privacy processing should be performed on the original data to ensure the privacy of the data; During the data analysis process, differential privacy algorithms should be used to protect query results and prevent personal privacy leaks; c4) Comply with relevant laws, regulations and industry standards Compliance with laws and regulations: Relevant laws and regulations should be strictly observed to ensure the legality of data processing activities; for cross-border data transfer, the laws and regulations of relevant countries and regions should be observed; Implementation of industry standards: ISO 27001 and ISO 29151 information security and privacy information management system standards should be followed to ensure the standardization and effectiveness of data security and privacy protection measures; in specific industries, the data protection and privacy requirements of the industry should also be followed; c5) Comprehensive measures and continuous improvement Implementation of comprehensive measures: Data encryption, anonymization, and differential privacy protection technologies should be combined to form a comprehensive data security and privacy protection framework; A risk management plan for data security and privacy protection should be established to identify risks through risk assessment and define the organization's risk tolerance and risk appetite; Continuous improvement and optimization: Data security and privacy protection measures should be audited and evaluated regularly to ensure their effectiveness and compliance; data security and privacy protection measures should be continuously optimized and improved based on technological development and business needs to adapt to the ever-changing security environment.

5. The data processing method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the distributed data processing and storage module are as follows: d1) Distributed file system: Hadoop distributed file system or equivalent distributed storage technology is used to achieve efficient and reliable storage of big data; the distributed file system has automatic data sharding, replication and fault tolerance mechanisms to ensure high availability and persistence of data; data is divided into multiple data blocks and stored on multiple physical nodes, and each data block has multiple copies distributed on different nodes to improve data reliability and reading efficiency; d2) Distributed computing framework: Apache Spark or equivalent distributed computing technology is used to achieve efficient processing of big data; the distributed computing framework supports data parallel processing and in-memory computing, which can significantly improve data processing speed and efficiency; through data parallel processing technology, large-scale data sets are divided into multiple subsets and processed in parallel on multiple nodes; Through in-memory computing technology, frequently accessed data during processing is stored in memory to reduce disk I / O operations and increase processing speed; d3) Data sharding and load balancing: The module has a data sharding mechanism that can reasonably divide data into multiple data blocks according to the data volume and processing requirements, and distribute them to different nodes for storage and processing; at the same time, the module also has a load balancing function that can dynamically adjust the distribution of data blocks according to the processing capacity of each node to ensure balanced distribution and efficient execution of data processing tasks; d4) Scalability and fault tolerance: The module supports dynamic expansion and can add new nodes as needed to expand storage and processing capabilities. At the same time, the module has a fault-tolerant mechanism that can automatically migrate data to other normal nodes when a node fails, ensuring the continuity and reliability of data processing. d5) Data processing interfaces and tools: The module provides a wealth of data processing interfaces and tools, supporting data cleaning, conversion, aggregation, and analysis of various data processing operations; users can easily write data processing programs through these interfaces and tools to meet complex data processing requirements.

6. The data processing method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the intelligent model training and evaluation module are as follows: e1) Model selection and design: Based on the type of input data and specific task requirements, the module can automatically select or design appropriate machine learning or deep learning models; models include but are not limited to support vector machines, decision trees, random forests, neural networks, convolutional neural networks, and recurrent neural networks; e2) Data preprocessing and enhancement: Before model training, the module uses preprocessed data for training; the preprocessing steps include data cleaning, normalization, feature selection, feature scaling, and data enhancement to improve data quality and model training results; e3) Model training and optimization: The module uses appropriate training algorithms and optimization strategies to train the model. At the same time, the module supports hyperparameter tuning, and finds the optimal hyperparameter combination through grid search, random search or Bayesian optimization method to improve the performance and generalization ability of the model. e4) Model evaluation and validation: During the model evaluation phase, the module uses a variety of evaluation indicators to conduct a comprehensive and accurate evaluation of the model, including but not limited to accuracy, recall, F1 score, cross-validation, confusion matrix, and AUC-ROC curve; at the same time, the module considers data balancing strategies to avoid overfitting of the model to the majority class samples; e5) Model integration and improvement: The module uses model integration technology to improve the robustness and generalization ability of the model; by training multiple base models and combining their prediction results, the stability and accuracy of the overall model are improved; e6) Performance monitoring and tuning: The module monitors the performance indicators of the model in real time during the training process and performs necessary tuning operations based on the monitoring results. At the same time, the module supports model visualization and helps users better understand the behavior and performance of the model by drawing training curves and feature importance graphs.

7. The data processing method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the feedback and optimization module are: f1) Feedback on model evaluation results: Based on the model evaluation results output by the intelligent model training and evaluation module, the module can identify the deficiencies of the model; f2) Model optimization strategy: The module can automatically select or recommend appropriate optimization strategies for the feedback of model evaluation results; these strategies include but are not limited to: Gradient descent and its variants: standard gradient descent, stochastic gradient descent, mini-batch gradient descent, momentum gradient descent, Adam optimizer, which is used to adjust model parameters to minimize the loss function; Hyperparameter tuning: Use grid search, random search, and Bayesian optimization methods to find the optimal hyperparameter combination in the preset hyperparameter space to improve model performance; Model structure adjustment: increase or decrease the number of network layers, change the number of neurons, and adjust the activation function to improve the model's expressiveness and generalization capabilities; f3) Continuous data collection: The module can continuously collect new data, which comes from real-time data streams, user feedback, and newly collected samples; The collection of new data helps the model better adapt to data changes and improves the practicality and accuracy of the model; f4) Incremental learning or online learning mechanism: Using the newly collected data, the module can adopt incremental learning or online learning mechanism to update the model; Incremental learning allows the model to gradually learn new data while retaining the original knowledge; Online learning allows the model to be continuously updated in real-time data streams to adapt to changes in the data; both mechanisms help maintain the performance and accuracy of the model; f5) Performance monitoring and alarm: During the model optimization and update process, the module can monitor the performance indicators of the model in real time; when the performance indicators are lower than the preset threshold, the module can automatically trigger the alarm mechanism to remind the user to make further optimization or adjustments.

8. The data processing method based on artificial intelligence according to claim 6, characterized in that: The intelligent model training and evaluation module supports an automated machine learning framework and can automatically complete model selection, parameter tuning and evaluation processes, reducing manual intervention and improving data processing efficiency.

9. The data processing method based on artificial intelligence according to claim 1, characterized in that: The method is applied to but not limited to data processing and analysis tasks in multiple fields such as finance, medical care, education, e-commerce, and the Internet of Things.

Citation Information

Cited By

  • Data processing method and device, computer equipment and storage medium

    CN120561567A

  • AI quality inspection system based on cloud upgrade

    CN120723779A