Intelligent data analysis system and method based on artificial intelligence
Through data preprocessing, feature engineering, and model optimization, the intelligent data analysis system solves the problems of low efficiency and inaccurate results in traditional data analysis, and achieves efficient and reliable data analysis and interactive display, which is applicable to fields such as e-commerce, finance, healthcare, and transportation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JINGKUN COMPUTER TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional data analysis methods are inefficient, data quality and analysis results are greatly affected by human factors, algorithm selection is unscientific, computing resource utilization is inefficient, result presentation is monotonous, and user interactivity is poor.
An AI-based intelligent data analysis system is adopted, including modules for data acquisition, preprocessing, feature engineering, model training and optimization, result evaluation and deployment, and visualization interaction. Through data cleaning, feature extraction, model selection and optimization, cross-validation, and various display formats, combined with big data processing technology, data quality and computational efficiency are improved.
It improves the accuracy and reliability of data analysis, enhances user interactivity, optimizes the utilization of computing resources, ensures the reliability and speed of analysis results, and meets the personalized needs of different fields.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and data processing technology, and specifically to an intelligent data analysis system and method based on artificial intelligence. Background Technology
[0002] With the rapid development of information technology, the amount of data generated by various industries is exploding. How to extract valuable information from massive amounts of data to support decision-making has become an urgent problem to be solved. Traditional data analysis methods often rely on manual operation, which is inefficient, difficult to handle large-scale data, and the accuracy and reliability of the analysis results are greatly affected by human factors.
[0003] Although some data analysis techniques have incorporated computer technology, numerous problems remain in practical applications. Regarding data quality, collected data often contains a large amount of duplicate, invalid, and erroneous information; without effective processing, this can lead to biased analysis results. In terms of algorithm selection, existing technologies lack scientific analysis of the compatibility between data characteristics and algorithms, making it difficult to select the most suitable algorithm for data analysis and affecting the analysis results. Regarding computing resources, large-scale data processing requires substantial computing resources, and existing technologies are inefficient in utilizing these resources, often resulting in slow analysis speeds or even failure to complete the analysis. Furthermore, the presentation of existing data analysis results is relatively limited, making it difficult for users to intuitively and conveniently obtain key information or engage in interactive exploration, thus hindering the full realization of the value of the data analysis results. Therefore, those skilled in the art have provided intelligent data analysis systems and methods based on artificial intelligence to address the problems mentioned in the background. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides an artificial intelligence-based intelligent data analysis system, comprising a data acquisition module, a data preprocessing module, a feature engineering module, a model training and optimization module, a result evaluation and deployment module, and a visualization and interaction module. The data acquisition module acquires data from different sources; the data preprocessing module processes the acquired data to improve its quality; the feature engineering module performs feature-related operations on the preprocessed data; the model training and optimization module selects a suitable model and performs training and parameter adjustments; the result evaluation and deployment module evaluates model performance, formulates deployment plans, and performs monitoring and optimization; and the visualization and interaction module displays data analysis results in various formats and supports user interaction.
[0005] Preferably, the data acquisition module acquires data from at least one of the following sources: database, API interface, social media platform, sensor device, and web data acquired by web crawler.
[0006] Preferably, the data preprocessing module performs the following operations: data cleaning, data transformation, and data normalization. Data cleaning is used to remove duplicate data, invalid data, and erroneous data. Data transformation is used to convert data from one format or structure to another. Data normalization is used to scale the data to a uniform numerical range.
[0007] Preferably, the feature engineering module performs feature-related operations including feature extraction, feature selection, feature transformation, and feature dimensionality reduction; the feature extraction is used to extract features related to the analysis target from the original data.
[0008] Preferably, the model training and optimization module selects models including machine learning models and deep learning models; the machine learning models include classification algorithm models, clustering algorithm models, and regression algorithm models.
[0009] Preferably, the model performance evaluation of the result evaluation deployment module is performed using a test dataset, and the deployment scheme includes a hardware configuration scheme and a software environment configuration scheme; the monitoring and optimization is used to monitor the running status of the model after deployment in real time, and adjust the model parameters according to the monitoring results.
[0010] Preferably, the display formats of the visualization interaction module include charts, visualization dashboards, interactive visualization interfaces, and visualization large screens.
[0011] An artificial intelligence-based intelligent data analysis method includes the following steps: S1: Data source selection and data collection and organization, determine the data source and collect the data, and perform preliminary organization of the collected data; S2: Data preprocessing, which involves cleaning, transforming, and normalizing the processed data. S3: Data feature engineering, which performs feature extraction, feature selection, feature transformation, and feature dimensionality reduction operations on preprocessed data; S4: Model selection, training and optimization. Select a suitable model based on data characteristics and business needs, train the model using the training dataset, and adjust the model parameters using cross-validation techniques. S5: Results evaluation and deployment. Use test datasets to evaluate model performance, develop and deploy the model, and monitor and optimize the deployed model in real time. S6: Visualization of data analysis results. The data analysis results are displayed in various forms through the visualization and interactive module, and user interaction is supported.
[0012] Preferably, in step S1, data collection and organization also includes classifying and storing the collected data.
[0013] Preferably, in step S5, the monitoring and optimization after model deployment includes periodically collecting model running data, analyzing the trend of model performance changes, and retraining or adjusting the parameters when the model performance declines beyond a preset threshold.
[0014] The technical effects and advantages of this invention are as follows: 1. In this invention, the collected data is comprehensively processed through a data preprocessing module, effectively removing duplicate, invalid, and erroneous data, improving data quality, providing a reliable data foundation for subsequent data analysis and model training, and avoiding deviations in analysis results due to data quality issues.
[0015] 2. This invention, through in-depth analysis of data characteristics and scientific selection of appropriate models based on business needs, and through optimization of model parameters using techniques such as cross-validation, improves the adaptability of algorithms to data, ensures the accuracy and reliability of data analysis results, and solves the problem of blind algorithm selection in existing technologies.
[0016] 3. This invention employs big data processing technologies such as data sharding, parallel processing, data compression, and data stream processing, which effectively improves the utilization efficiency of computing resources, accelerates the processing of large-scale data and model training, and solves the problem of low analysis efficiency caused by insufficient computing resources. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an artificial intelligence-based intelligent data analysis system provided in an embodiment of this application; Figure 2 This is a schematic diagram of the data acquisition module in the artificial intelligence-based intelligent data analysis system provided in this application embodiment; Figure 3 This is a schematic diagram of the feature engineering module in the artificial intelligence-based intelligent data analysis system provided in the embodiments of this application; Figure 4 This is a schematic diagram of the model training and optimization module in the artificial intelligence-based intelligent data analysis system provided in this application embodiment; Figure 5 This is a schematic diagram of the result evaluation deployment module in the artificial intelligence-based intelligent data analysis system provided in this application embodiment; Figure 6 This is a schematic diagram of the visualization interaction module in the artificial intelligence-based intelligent data analysis system provided in this application embodiment; Figure 7 This is a schematic diagram of an artificial intelligence-based intelligent data analysis method provided in an embodiment of this application. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose. Please see Figures 1-7 As shown in the figure, this embodiment provides an intelligent data analysis system and method based on artificial intelligence, including a data acquisition module, a data preprocessing module, a feature engineering module, a model training and optimization module, a result evaluation and deployment module, and a visualization and interaction module. The data acquisition module is used to acquire data from different sources; the data preprocessing module is used to process the acquired data to improve data quality; the feature engineering module is used to perform feature-related operations on the preprocessed data; the model training and optimization module is used to select a suitable model and perform training and parameter adjustment; the result evaluation and deployment module is used to evaluate model performance, formulate deployment plans, and perform monitoring and optimization; the visualization and interaction module is used to display data analysis results in various forms and support user interaction.
[0019] Specifically, in this invention, the collected data is comprehensively processed through a data preprocessing module to effectively remove duplicate, invalid, and erroneous data, improve data quality, provide a reliable data foundation for subsequent data analysis and model training, and avoid deviations in analysis results due to data quality issues.
[0020] By conducting in-depth analysis of data characteristics and scientifically selecting appropriate models based on business needs, and by optimizing model parameters through techniques such as cross-validation, the adaptability of algorithms to data has been improved, ensuring the accuracy and reliability of data analysis results and solving the problem of blind algorithm selection in existing technologies.
[0021] Furthermore, the data acquisition module obtains data from at least one of the following sources: databases, API interfaces, social media platforms, sensor devices, and web data obtained by web crawlers.
[0022] Furthermore, the data preprocessing module includes data cleaning, data transformation, and data normalization. Data cleaning is used to remove duplicate, invalid, and erroneous data. Data transformation is used to convert data from one format or structure to another. Data normalization is used to scale data to a uniform numerical range.
[0023] Furthermore, the feature engineering module's feature-related operations include feature extraction, feature selection, feature transformation, and feature dimensionality reduction; feature extraction is used to extract features relevant to the analysis target from the raw data.
[0024] Specifically, the feature engineering module's feature-related operations include feature extraction, feature selection, feature transformation, and feature dimensionality reduction. Feature extraction is used to extract features relevant to the analysis objective from the raw data. Feature selection is used to screen features that contribute to the model's predictive performance. Feature transformation is used to normalize, standardize, or uniquely encode features. Feature dimensionality reduction is used to reduce the feature dimension through principal component analysis and factor analysis.
[0025] Furthermore: the model training and optimization module selects machine learning models and deep learning models; the machine learning models include classification algorithm models, clustering algorithm models, and regression algorithm models.
[0026] Specifically, the model training and optimization module selects machine learning models and deep learning models; machine learning models include classification algorithm models, clustering algorithm models, and regression algorithm models; classification algorithm models include at least one of logistic regression model, support vector machine model, and decision tree model; clustering algorithm models include at least one of K-means model and hierarchical clustering model; regression algorithm models include at least one of linear regression model and ridge regression model; and deep learning models are neural network-based models used to process large-scale, high-dimensional data.
[0027] Furthermore: The performance evaluation of the model in the result evaluation deployment module is conducted using a test dataset, and the deployment scheme includes hardware configuration schemes and software environment configuration schemes; monitoring and optimization are used to monitor the running status of the model after deployment in real time, and adjust the model parameters based on the monitoring results.
[0028] Specifically, the performance evaluation of the model in the result evaluation deployment module is conducted using a test dataset, and the evaluation metrics include at least one of accuracy, precision, recall, F1 score, and mean squared error; the deployment scheme includes hardware configuration scheme and software environment configuration scheme; monitoring and optimization are used to monitor the running status of the model after deployment in real time and adjust the model parameters according to the monitoring results.
[0029] Furthermore, the visualization and interactive modules can be displayed in various formats, including charts, visualization dashboards, interactive visualization interfaces, and large visualization screens.
[0030] Specifically, the display formats of the visualization interaction module include charts, visualization dashboards, interactive visualization interfaces, and visualization screens; charts include at least one of bar charts, line charts, pie charts, and scatter plots; interactive visualization interfaces allow users to explore data through clicking, dragging, and filtering operations.
[0031] Please see Figure 7As shown, an intelligent data analysis method based on artificial intelligence includes the following steps: S1: Data source selection and data collection and organization, determining data sources and collecting data, and performing preliminary organization of the collected data; S2: Data preprocessing, performing data cleaning, data transformation, and data normalization on the organized data; S3: Data feature engineering, performing feature extraction, feature selection, feature transformation, and feature dimensionality reduction operations on the preprocessed data; S4: Model selection, training, and optimization, selecting a suitable model based on data characteristics and business needs, training the model using a training dataset, and adjusting model parameters using cross-validation techniques; S5: Result evaluation and deployment, evaluating model performance using a test dataset, developing a model deployment plan and deploying it, and monitoring and optimizing the deployed model in real time; S6: Visualization of data analysis results, displaying data analysis results in various forms through a visualization interaction module and supporting user interaction.
[0032] Further: In step S1, data collection and organization also includes classifying and storing the collected data.
[0033] Specifically, in step S1, data collection and organization also includes classifying and storing the collected data, and the storage method includes at least one of relational database storage, NoSQL database storage, and data warehouse storage.
[0034] Furthermore: In step S4, big data processing technology is used to improve training efficiency during model training. Big data processing technology includes at least one of data sharding technology, parallel processing technology, data compression technology, and data stream processing technology. Data sharding technology divides the large dataset into small blocks for distributed processing. Parallel processing technology utilizes multiple processors to process data simultaneously. Data compression technology uses Huffman coding and the LZ77 algorithm to reduce data storage space. Data stream processing technology uses the Apache Kafka and Storm frameworks to process real-time incoming data streams.
[0035] Furthermore, in step S5, the post-deployment monitoring and optimization includes periodically collecting model running data, analyzing the trend of model performance changes, and retraining or adjusting the parameters when the model performance declines beyond a preset threshold.
[0036] Specifically, in step S5, the post-deployment monitoring and optimization includes periodically collecting model running data, analyzing the trend of model performance changes, and retraining the model or adjusting the parameters when the model performance declines beyond a preset threshold.
[0037] The present invention will be further described in detail below with reference to specific embodiments. Example 1: E-commerce User Behavior Analysis This example applies the intelligent data analysis system and implementation method of the present invention to e-commerce user behavior analysis to achieve personalized recommendations and precision marketing.
[0038] The intelligent data analysis system of this invention operates as follows: Data acquisition module: Acquires database data (including basic user information, product information, and transaction records) from e-commerce platforms, API interface data (such as payment data from third-party payment platforms), user browsing history, search keywords, favorites, and reviews on e-commerce platforms.
[0039] Data preprocessing module: Cleans the collected data, deleting duplicate user browsing records, invalid transaction records (such as transactions with an amount of 0), and incorrect user information (such as incorrectly formatted phone numbers); converts data in different formats (such as JSON format browsing record data and CSV format transaction data) into unified tabular data; normalizes data such as user spending amount and browsing time, scaling it to the [0,1] range.
[0040] Feature engineering module: Extracts user behavior features from preprocessed data, such as browsing frequency, purchase frequency, average purchase amount, preferred product categories, and search keyword frequency; uses correlation coefficient method to screen features with high relevance to user purchase intention, such as purchase frequency, preferred product categories, and average purchase amount; performs independent encoding and transformation on category features such as preferred product categories; and performs dimensionality reduction on multiple extracted features through principal component analysis to reduce the number of features and simplify the model.
[0041] Model Training and Optimization Module: Based on the business needs of e-commerce user behavior analysis (predicting whether users will purchase specific products, a classification task), a logistic regression model was selected as the initial model. The dataset was divided into training and testing datasets in an 8:2 ratio, and the logistic regression model was trained using the training dataset. During training, data sharding technology was used to divide the large-scale user behavior data into smaller chunks for distributed processing; the model's regularization parameters were adjusted using 5-fold cross-validation to improve its generalization ability. After training and optimization, a high-performance user purchase prediction model was obtained.
[0042] The results evaluation and deployment module evaluates the user purchase prediction model using a test dataset. The model's accuracy is calculated to be 89%, precision 85%, recall 87%, and F1 score 86%, meeting the personalized recommendation needs of e-commerce platforms. A deployment plan is developed, selecting a server with high CPU and memory configuration as the model's runtime server, and installing a Linux operating system, MySQL database, Python language environment, and Scikit-learn machine learning library. After deploying the model to the server, its running status is monitored in real time, collecting data such as prediction accuracy and response time. When the model's accuracy drops below 85%, it is retrained and optimized using the latest user behavior data.
[0043] The interactive visualization module displays user behavior analysis results in multiple formats. It uses bar charts to show the number of purchases by users across different product categories, line charts to show monthly changes in user purchase frequency, and visual dashboards to display key metrics such as daily active users, order conversion rate, and average order value. An interactive visual interface allows operations staff to filter users by age group and region to view their purchasing behavior characteristics. Based on the results of the user purchase prediction model, it recommends products users are likely to purchase, achieving personalized recommendations and precise marketing, thereby improving the e-commerce platform's order conversion rate and user satisfaction.
[0044] Example 2: Financial Risk Prediction This example applies the intelligent data analysis system and implementation method of the present invention to financial risk prediction, in order to predict financial market trends and investment risks, and provide a basis for investment decisions.
[0045] The intelligent data analysis system of this invention operates as follows: Data Acquisition Module: Acquires historical data (such as daily closing prices, trading volume, and price changes over the past 10 years), real-time data (such as current market prices and trading volume), macroeconomic data (such as GDP growth rate, interest rates, and inflation rates), and corporate financial data (such as listed companies' net profit, debt-to-equity ratio, and revenue growth rate) from financial markets such as stocks, futures, and foreign exchange. Data sources include API interfaces provided by financial data service providers, public databases of stock exchanges, and macroeconomic reports published by economic statistics departments.
[0046] Data preprocessing module: Cleans the collected financial data, removing outliers caused by data transmission errors (such as sudden and unexplained large fluctuations in stock prices) and duplicate transaction records; converts financial data in different formats (such as macroeconomic data in XML format and historical stock data in TXT format) into unified structured data; normalizes data such as stock prices and trading volumes to eliminate the impact of price magnitude differences between different financial products. For example, it converts stock price data into yield data, calculated using the following formula: in Let be the rate of return on day t. Let be the closing price on day t. The closing price on day t-1.
[0047] Feature Engineering Module: Extracts features related to market trends and risks from preprocessed financial data, such as moving averages (MA), relative strength index (RSI), stochastic oscillator (KDJ), trading volume volatility, interest rate change rate, and GDP growth rate trends. A recursive feature elimination method is used to screen features with a significant impact on financial risk prediction, such as moving averages, RSI, trading volume volatility, and interest rate change rate. The screened features are standardized to ensure they conform to a normal distribution. Factor analysis is used to reduce the dimensionality of multiple related features, extracting key factors, reducing feature dimensionality, and lowering model complexity.
[0048] Model Training and Optimization Module: Based on the business requirements of financial risk prediction (predicting the price trends of financial products, a regression task), a gradient boosting regression model (such as the XGBoost model) is selected. Financial data is divided into training and testing datasets in a 7:3 ratio, and the XGBoost model is trained using the training dataset. During training, parallel processing technology is employed, utilizing multiple CPU cores to process data simultaneously, accelerating model training. Parameters such as the learning rate, tree depth, and number of trees are adjusted using a grid search method to improve the model's prediction accuracy. Simultaneously, data stream processing technology is used to process real-time acquired financial data, enabling real-time model updates and optimization, ensuring the model can adapt to market changes promptly.
[0049] The results evaluation and deployment module evaluates the optimized XGBoost model using a test dataset. The model's mean squared error (MSE) is calculated to be 0.02, and its mean absolute error (MAE) is 0.15, indicating good predictive performance that meets the needs of financial risk prediction. A deployment plan is developed, selecting a server with high-performance GPUs and ample memory as the model's runtime server. A Windows Server operating system, Oracle database, Python language environment, and XGBoost library are installed. After deploying the model to the server, its running status is monitored in real time, collecting information such as the deviation between the model's predictions and actual market data, and the model's response time. When the model's prediction deviation exceeds a preset threshold (e.g., MSE exceeding 0.03), the model is retrained using the latest financial data, and the model parameters are adjusted accordingly.
[0050] The visualization and interactive module displays real-time trends of major financial market indices (such as the Shanghai Composite Index, Shenzhen Component Index, and Nasdaq Index), price fluctuations of various financial products, and risk warning indicators in the form of a large visual screen; it uses line charts to compare the price trends of financial products predicted by the model with the actual price trends; it uses heatmaps to show the risk level distribution of different industry sectors; and it allows investors to filter data for specific financial products and specific time periods through an interactive interface, view detailed analysis results and risk assessment reports, and provide intuitive and accurate reference for investment decisions.
[0051] Example 3: Smart City Traffic Flow Prediction This example applies the intelligent data analysis system and implementation method of the present invention to smart city traffic flow prediction in order to predict traffic flow and congestion and optimize traffic management.
[0052] The intelligent data analysis system of this invention operates as follows: Data acquisition module: Acquires real-time data such as traffic flow, vehicle speed, and vehicle type through traffic sensors (such as geomagnetic sensors and video detectors) deployed on urban roads; retrieves historical traffic flow data, traffic accident records, and road construction information from the database of traffic management departments; acquires weather data (such as rainfall, visibility, and wind force) through the API interface of meteorological departments; and extracts holiday information, weekday and weekend information from calendar data.
[0053] Data preprocessing module: Cleans the collected traffic data, removing abnormal data caused by sensor malfunctions (such as traffic flow suddenly dropping to 0 or data far exceeding the normal range) and duplicate traffic accident records; converts data from different sources (such as real-time data acquired by sensors and historical data in the database) into a unified timestamp format and data structure; normalizes traffic flow, vehicle speed, and other data, converting traffic flow data into the number of vehicles per unit time, vehicle speed data into kilometers per hour, and scaling it to an appropriate numerical range.
[0054] Feature Engineering Module: Extracts features related to traffic flow and congestion from preprocessed traffic data, such as peak hours (7:00-9:00 AM, 5:00-7:00 PM), weather conditions (sunny, rainy, snowy), holidays, road types (main roads, secondary roads, local roads), traffic accident frequency, and construction section length. Chi-square tests are used to screen features that significantly impact traffic flow prediction, such as peak hours, weather conditions, holidays, and road types. Weather conditions and road types are independently coded. Principal component analysis is used to reduce the dimensionality of the extracted features, decreasing the number of features and improving model training efficiency.
[0055] Model Training and Optimization Module: Based on the business requirements of traffic flow prediction (predicting traffic flow over a future period, a regression task), a Long Short-Term Memory (LSTM) deep learning model is selected. This model can effectively handle time-series data and is suitable for time-series traffic flow prediction. Traffic data is divided into training and testing datasets in a 7:3 ratio, and the LSTM model is trained using the training dataset. During training, data compression techniques are used to compress and store large-scale historical traffic data, reducing storage space usage; parallel processing techniques are employed to accelerate model training; and the prediction accuracy of the model is improved by adjusting parameters such as the number of hidden layers, the number of neurons, and the learning rate. Simultaneously, real-time traffic data streams are used, and data stream processing techniques are employed to enable online learning and updates of the model, allowing it to adapt to dynamic changes in traffic conditions.
[0056] The results evaluation and deployment module evaluates the optimized LSTM model using a test dataset. The mean absolute percentage error (MASE) is calculated to be 8%, indicating that the prediction performance meets the needs of traffic management departments. A deployment plan is developed, selecting a server with high-performance GPUs as the model's runtime server. A Linux operating system, MongoDB database, Python language environment, and TensorFlow deep learning framework are installed. After deploying the model to the server, its running status is monitored in real time, collecting data such as the deviation between the model's predicted traffic flow and actual traffic flow, and the model's response time. When the MASE exceeds 10%, the model is retrained using the latest traffic data to optimize the model parameters.
[0057] The visualization and interactive module displays real-time traffic flow, vehicle speed, and congestion rate information for various urban areas in the form of a visual dashboard, using different colors to indicate the degree of road congestion (green for smooth traffic, yellow for slow traffic, and red for congestion); it uses line graphs to show the predicted traffic flow trends for the next 1, 3, and 6 hours; an interactive visualization interface allows traffic management personnel to filter data for specific roads and time periods, and view detailed traffic flow analysis results and congestion cause analysis reports; and it displays real-time traffic congestion warning information on a large visualization screen, issuing timely warnings when severe congestion is predicted on a certain road, providing decision support for traffic management departments to formulate traffic diversion plans and optimize traffic signal timing, thereby improving urban traffic management efficiency and alleviating traffic congestion.
[0058] Example 4: Medical Image Diagnosis This example applies the intelligent data analysis system and implementation method of the present invention to medical image diagnosis to automatically identify lesions and abnormalities and assist doctors in making diagnoses.
[0059] The intelligent data analysis system of this invention operates as follows: Data acquisition module: Collects medical image data from the hospital's PACS system (image storage and transmission system), including X-ray films, CT scan images, MRI (magnetic resonance imaging) images, etc.; and obtains basic patient information, medical history, diagnosis results, and other data from the hospital's electronic medical record system as auxiliary information for image diagnosis.
[0060] The data preprocessing module preprocesses the acquired medical image data. First, it performs noise reduction using methods such as Gaussian filtering and median filtering to remove noise and improve image quality. Then, it performs image enhancement by adjusting contrast and brightness, and employing techniques such as histogram equalization to highlight lesion areas in the images. It also performs unified format and resolution conversion for image data of different formats and resolutions to facilitate subsequent feature extraction and model training. Simultaneously, it cleans patient information data, removing duplicate and erroneous information to ensure data accuracy.
[0061] Feature Engineering Module: Extracts features related to lesions and abnormalities from preprocessed medical image data. Image segmentation techniques (such as thresholding, edge segmentation, and region growing segmentation) are used to separate organs, tissues, and background in the images. Shape features (such as area, perimeter, and roundness), texture features (such as gray-level co-occurrence matrix and texture entropy), and gray-level features (such as average gray value and gray-level standard deviation) of the lesion region are then extracted. A comprehensive feature set is constructed by combining patient medical history, age, and gender information. L1 regularization is used to select features important for lesion diagnosis and remove redundant features. Due to the high dimensionality of medical image features, principal component analysis is used to reduce the dimensionality of the features, retaining key feature information and reducing the computational load of the model.
[0062] Model Training and Optimization Module: Based on the business needs of medical image diagnosis (determining the presence and type of lesions in images, a classification task), a Convolutional Neural Network (CNN) deep learning model was selected. CNN models have excellent performance in image recognition and can automatically extract deep features from images. A large amount of medical image data annotated by doctors was collected as training samples. The dataset was divided into training and testing datasets in an 8:2 ratio, and the CNN model was trained using the training dataset. During training, data partitioning technology was used to divide the large-scale image data into smaller blocks for distributed training; data augmentation techniques (such as image rotation, flipping, scaling, and adding noise) were used to expand the number of training samples and improve the model's generalization ability; by adjusting parameters such as the number of convolutional layers, pooling layers, fully connected layer neurons, and learning rate of the CNN model, cross-validation technology was used to optimize model performance, enabling the model to accurately identify different types of lesions.
[0063] The results evaluation and deployment module: The optimized CNN model was evaluated using a test dataset, achieving an accuracy of 92%, precision of 90%, recall of 91%, and an F1 score of 90.5%. The model performance met the requirements for medical image-assisted diagnosis. A deployment plan was developed, selecting a server with a high-performance GPU as the model's runtime server, and installing a Linux operating system, MySQL database, Python language environment, and PyTorch deep learning framework. Before deployment, the model underwent security and reliability testing to ensure it did not leak patient privacy information and that the diagnostic results were stable and reliable. After deploying the model to the hospital's medical image diagnostic system, the model's running status was monitored in real time, collecting data such as the consistency between the model's diagnostic results and the doctor's diagnoses, and the model's processing speed. When the model's diagnostic accuracy dropped below 90%, the model was retrained using the latest labeled medical image data, and the model parameters were updated.
[0064] The visualization and interaction module displays medical image diagnostic results in an interactive visual interface. It shows the original medical images and marks the lesion areas identified by the model with different colored markers, while also displaying the lesion area's characteristic information (such as lesion type, lesion size, and confidence level). It allows doctors to zoom in, zoom out, and rotate the image to examine the details of the lesion area in detail. It provides a comparison function between the model's diagnostic results and historical diagnostic results for doctors' reference. The module generates a diagnostic report from the model's diagnostic results and image feature data, which doctors can edit and modify to form a complete medical diagnostic report. This helps doctors improve diagnostic efficiency and accuracy, reducing misdiagnosis and missed diagnosis.
[0065] 1. The intelligent data analysis system and implementation method of the present invention comprehensively processes the collected data through a data preprocessing module, effectively removing duplicate, invalid, and erroneous data, improving data quality, providing a reliable data foundation for subsequent data analysis and model training, and avoiding deviations in analysis results due to data quality issues.
[0066] 2. Regarding algorithm selection, this invention scientifically selects appropriate models by conducting in-depth analysis of data characteristics and combining them with business needs. It also optimizes model parameters through techniques such as cross-validation, thereby improving the adaptability of the algorithm to the data, ensuring the accuracy and reliability of the data analysis results, and solving the problem of blind algorithm selection in existing technologies.
[0067] 3. This invention employs big data processing technologies such as data sharding, parallel processing, data compression, and data stream processing, which effectively improves the utilization efficiency of computing resources, accelerates the processing of large-scale data and model training, and solves the problem of low analysis efficiency caused by insufficient computing resources.
[0068] 4. The visualization and interaction module presents data analysis results in a variety of intuitive and vivid forms and supports user interaction, enabling users to quickly obtain key information, explore data in depth, fully leverage the value of data analysis results, provide strong support for decision-making, and improve the situation of single result display and poor user experience in existing technologies.
[0069] 5. This invention has wide applicability and can be flexibly applied to multiple fields such as e-commerce, finance, healthcare, and transportation. It can be customized according to the business needs of different fields, providing technical support for the intelligent development of various industries. It has important practical significance and great application value.
[0070] Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art and related fields based on the embodiments of the present invention without inventive effort should fall within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described and explained in the present invention, unless otherwise specified or limited, shall be implemented according to conventional means in the art.
Claims
1. An intelligent data analysis system based on artificial intelligence, characterized in that: It includes a data acquisition module, a data preprocessing module, a feature engineering module, a model training and optimization module, a result evaluation and deployment module, and a visualization and interaction module; The data acquisition module is used to acquire data from different sources; the data preprocessing module is used to process the acquired data to improve data quality. The feature engineering module is used to perform feature-related operations on the preprocessed data; The model training and optimization module is used to select a suitable model and perform training and parameter adjustment; the result evaluation and deployment module is used to evaluate model performance, formulate deployment plans, and perform monitoring and optimization; the visualization and interaction module is used to display data analysis results in various forms and support user interaction.
2. The intelligent data analysis system based on artificial intelligence according to claim 1, characterized in that, The data acquisition module acquires data from at least one of the following sources: databases, API interfaces, social media platforms, sensor devices, and web data obtained by web crawlers.
3. The intelligent data analysis system based on artificial intelligence according to claim 1, characterized in that, The data preprocessing module performs the following operations: data cleaning, data transformation, and data normalization. Data cleaning is used to remove duplicate, invalid, and erroneous data. Data transformation is used to convert data from one format or structure to another. Data normalization is used to scale the data to a uniform numerical range.
4. The intelligent data analysis system based on artificial intelligence according to claim 1, characterized in that, The feature engineering module's feature-related operations include feature extraction, feature selection, feature transformation, and feature dimensionality reduction; the feature extraction is used to extract features related to the analysis target from the original data.
5. The intelligent data analysis system based on artificial intelligence according to claim 1, characterized in that, The model training and optimization module selects models including machine learning models and deep learning models; the machine learning models include classification algorithm models, clustering algorithm models, and regression algorithm models.
6. The intelligent data analysis system based on artificial intelligence according to claim 1, characterized in that, The performance evaluation of the model in the result evaluation deployment module is performed using a test dataset. The deployment scheme includes a hardware configuration scheme and a software environment configuration scheme. The monitoring and optimization are used to monitor the running status of the model after deployment in real time and adjust the model parameters according to the monitoring results.
7. The intelligent data analysis system based on artificial intelligence according to claim 1, characterized in that, The visualization interaction module can be displayed in various formats, including charts, visualization dashboards, interactive visualization interfaces, and large visualization screens.
8. The method for intelligent data analysis based on artificial intelligence according to any one of claims 1-7, characterized in that, Includes the following steps: S1: Data source selection and data collection and organization, determine the data source and collect the data, and perform preliminary organization of the collected data; S2: Data preprocessing, which involves cleaning, transforming, and normalizing the processed data. S3: Data feature engineering, which performs feature extraction, feature selection, feature transformation, and feature dimensionality reduction operations on preprocessed data; S4: Model selection, training and optimization. Select a suitable model based on data characteristics and business needs, train the model using the training dataset, and adjust the model parameters using cross-validation techniques. S5: Results evaluation and deployment. Use test datasets to evaluate model performance, develop and deploy the model, and monitor and optimize the deployed model in real time. S6: Visualization of data analysis results. The data analysis results are displayed in various forms through the visualization and interactive module, and user interaction is supported.
9. The method for intelligent data analysis based on artificial intelligence according to claim 8, characterized in that, In step S1, data collection and organization also includes classifying and storing the collected data.
10. The method for intelligent data analysis based on artificial intelligence according to claim 9, characterized in that, In step S5, post-deployment monitoring and optimization includes periodically collecting model running data, analyzing the trend of model performance changes, and retraining or adjusting the parameters when the model performance declines beyond a preset threshold.