Computer data analysis method based on artificial intelligence

Through the computer data analysis method based on artificial intelligence, the data analysis module and data update module are used to solve the problem that models cannot be cross-compared in the existing technology, and the accuracy of data analysis and the optimization of model performance are achieved.

CN120011738AInactive Publication Date: 2025-05-16天津仁爱学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411891558.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In computer data analysis, the existing technology models cannot be cross-compared after analyzing the data, resulting in analysis redundancy and deviation, which requires manual review.

Method used

A computer data analysis method based on artificial intelligence is adopted, including data collection module, data analysis module and data update module. The data analysis module ensures the accuracy and generalization ability of the model through algorithm selection and model training comparison. The data update module optimizes model performance through manual data adjustment and model hyperparameter adjustment.

Benefits of technology

Through artificial intelligence analysis and cross-comparison, reduce the redundancy and deviation of analysis, improve the accuracy of data, and ensure the effectiveness and performance optimization of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011738A_ABST
    Figure CN120011738A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer data analysis, in particular to a computer data analysis method based on artificial intelligence, which comprises a host, the host is connected with a data collection module, a data analysis module, a data updating module and a display, specifically, the data collection module comprises data preprocessing and data collection, the data analysis module comprises algorithm selection and model training comparison, the data updating module comprises manual data adjustment and data model updating, the data preprocessing comprises the steps that a user inputs data needing to be analyzed and collected into the host to form comparison data, the host converts the comparison data into a format suitable for analysis, and then the data is updated; according to the method and the system, data of different sources can be processed and compared in a unified manner in a TXT format, a DOCX format, a PDF format and a CSV format, the data is analyzed, meanwhile, the data is collected and compared with network data through artificial intelligence, the accuracy rate of the data is increased, and analysis redundancy and deviation are reduced through cross comparison.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer data analysis, and in particular to a computer data analysis method based on artificial intelligence. Background Art

[0002] Data analysis refers to the process of analyzing a large amount of collected first-hand and second-hand data using appropriate statistical methods in order to maximize the development of the data and give full play to the role of data. It is a process of detailed research and summary of data in order to extract useful information and form conclusions. In the computer field, computer data analysis is the process of processing and analyzing data with the help of computer technology and tools.

[0003] When analyzing data, some computers use their own internal models to analyze data. However, the models they carry directly derive conclusions after analyzing the data and cannot be cross-compared, which increases the redundancy and bias of the analysis and requires multiple manual reviews. Therefore, it is necessary to propose a computer data analysis method based on artificial intelligence. Summary of the invention

[0004] In view of the problems in the prior art, the present invention provides a computer data analysis method based on artificial intelligence.

[0005] The technical solution adopted by the present invention to solve its technical problem is: a computer data analysis method based on artificial intelligence, including a host, and the host is connected to a data collection module, a data analysis module, a data update module and a display.

[0006] Specifically, the data collection module includes data preprocessing and data collection; The data analysis module includes algorithm selection and model training comparison; The data updating module includes manual data adjustment and data model updating.

[0007] Specifically, the data preprocessing is that the user enters the data to be analyzed and collected into the host to form comparative data. The host first converts the comparative data into a format suitable for analysis, such as TXT format, DOCX format, PDF format, CSV format, to ensure that data from different sources can be uniformly processed and compared, and then removes duplicate data, missing values, abnormal values, etc. in the converted data, extracts nouns, verbs or adjectives in the data, and scales the data to a standard range to re-form streamlined data. The standard range can be the number of words in the streamlined data; The data collection is that the host transmits the simplified data to the crawler software, the crawler software compares the data through the network, crawls the data similar to the simplified data, and displays the crawled data through the display.

[0008] Specifically, the algorithm selection is to analyze the characteristics of the collected data and the task requirements as needed, select a suitable artificial intelligence algorithm for use in a dynamic environment where decision-making is required, and is suitable for solving multi-step decision-making problems; The model training comparison is to train the data set according to the selected algorithm. During the data training process, the test data set is first used to evaluate the generalization ability of the model. The evaluation indicators include accuracy, precision, recall rate and F1 value. After the data to be analyzed is input into the model, the parameters of the model are adjusted through the training set, the performance of the model is evaluated through cross-validation and other methods, and the model is optimized by adjusting the learning rate, regularization, etc., and the model is prevented from overfitting. At the same time, it is ensured that the model has sufficient generalization ability. During training, the data crawled by the crawler can be compared with the data after model analysis, and the data output by the model can be compared and analyzed to ensure the accuracy of the model.

[0009] Specifically, the manual data adjustment is to perform secondary deletion of the concentrated repeated records in the streamlined data when the host processes the data, so as to avoid redundancy and deviation in the model training process, and adopt a filling or deletion strategy for the missing data, and modify the data by a certain amount. The modification amount depends on the number and distribution of the missing values. During the model training analysis, the streamlined data can be interfered by adding text data and other methods to generate new samples to enhance data analysis; The data model update is the host selecting and adjusting the model hyperparameters during the model training process to achieve the purpose of optimizing the model performance. The hyperparameters are parameters set before model training, such as learning rate, batch size, regularization coefficient, etc., which are adjusted through online learning on the network. The model learns and updates its parameters from the new data extracted by the crawler. After the model is updated, A / B testing can be used to evaluate whether the performance of the new model is better than the old model. By comparing the performance of the new and old models on the same data set, a model with better performance can be selected for deployment.

[0010] Specifically, the artificial intelligence algorithm includes supervised learning and unsupervised learning. The supervised learning, such as regression analysis, decision tree, support vector machine, neural network, etc., is suitable for labeled data sets. The unsupervised learning, such as hierarchical clustering, etc., is suitable for unlabeled scenarios and is usually used for data exploration, pattern recognition, etc.

[0011] Specifically, during the manual data adjustment process, it is necessary to ensure the privacy and security of the data to avoid leakage of sensitive information. When adjusting the data, data from different sources and backgrounds should be included as much as possible.

[0012] Specifically, when the host updates the model, it can continuously optimize the model through incremental learning for new data without directly replacing the old model. When the model is manually adjusted, the model can be monitored regularly to ensure the effectiveness of the model in practical applications.

[0013] Beneficial effects of the present invention: (1) The computer data analysis method based on artificial intelligence described in the present invention analyzes data through artificial intelligence, simplifies the data during analysis, unifies the data format during data simplification, and screens the unified data. The data that has been simplified multiple times is placed in a model for analysis. The data that has been simplified multiple times can also be compared and collected with network data, and the compared and collected data is imported into the model. After the model performs self-analysis on the data, it is cross-compared with the compared and collected data. The data is displayed on a display. While analyzing the data, it is collected and compared with network data through artificial intelligence to increase the accuracy of the data. Through cross-comparison, the redundancy and deviation of the analysis are reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0015] Figure 1 A schematic diagram of the data analysis process structure of a computer data analysis method based on artificial intelligence provided by the present invention; Figure 2 A schematic diagram of the model update process structure of a computer data analysis method based on artificial intelligence provided by the present invention; Figure 3 A schematic diagram of the data collection process structure of an artificial intelligence-based computer data analysis method provided by the present invention. DETAILED DESCRIPTION

[0016] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.

[0017] like Figure 1-Figure 3 As shown, the computer data analysis method based on artificial intelligence described in the present invention includes a host, and the host is connected to a data collection module, a data analysis module, a data update module and a display.

[0018] Among them, the data collection module includes data preprocessing and data collection; The data analysis module includes algorithm selection and model training comparison; The data update module includes manual data adjustment and data model update.

[0019] Among them, data preprocessing is that the user enters the data to be analyzed and collected into the host to form comparative data. The host first converts the comparative data into a format suitable for analysis, such as TXT format, DOCX format, PDF format, CSV format, to ensure that data from different sources can be uniformly processed and compared, and then removes duplicate data, missing values, and abnormal values ​​in the converted data, extracts nouns, verbs, or adjectives in the data, and scales the data to a standard range to re-form streamlined data. The standard range can be the number of words in the streamlined data; Data collection is when the host transmits streamlined data to the crawler software. The crawler software compares data on the network, crawls data similar to the streamlined data, and displays the crawled data on the monitor.

[0020] Among them, algorithm selection is to analyze the characteristics of collected data and task requirements as needed, select a suitable artificial intelligence algorithm, which is used in a dynamic environment where decision-making is required and is suitable for solving multi-step decision-making problems; Model training comparison is to train the data set according to the selected algorithm. During the data training process, the test data set is first used to evaluate the generalization ability of the model. The evaluation indicators include accuracy, precision, recall and F1 value. After the data to be analyzed is input into the model, the parameters of the model are adjusted through the training set. The performance of the model is evaluated through methods such as cross-validation, and the model is optimized by adjusting the learning rate and regularization. The model is prevented from overfitting and ensures that the model has sufficient generalization ability. During training, the data crawled by the crawler can be compared with the data after model analysis, and the data output by the model can be compared and analyzed to ensure the accuracy of the model.

[0021] Among them, manual data adjustment is to perform secondary deletion of concentrated and repeated records in the streamlined data when the host processes the data, so as to avoid redundancy and deviation in the model training process, and adopt a filling or deletion strategy for the missing data to modify the data to a certain extent. The specific amount of modification depends on the number and distribution of missing values. During the model training and analysis, the streamlined data can be interfered by adding text data and other methods to generate new samples to enhance data analysis; Data model update is the host's selection and adjustment of model hyperparameters during model training to optimize model performance. Hyperparameters are parameters set before model training, such as learning rate, batch size, regularization coefficient, etc. They are adjusted through online learning on the network. The model learns and updates its parameters from new data extracted by the crawler. After the model is updated, A / B testing can be used to evaluate whether the performance of the new model is better than the old model. By comparing the performance of the new and old models on the same data set, the model with better performance can be selected for deployment.

[0022] Among them, artificial intelligence algorithms include supervised learning and unsupervised learning. Supervised learning, such as regression analysis, decision trees, support vector machines, neural networks, etc., is suitable for labeled data sets. Unsupervised learning, such as hierarchical clustering, is suitable for unlabeled scenarios and is usually used for data exploration, pattern recognition, etc.

[0023] During the manual data adjustment process, it is necessary to ensure the privacy and security of the data to avoid leaking sensitive information. When adjusting the data, data from different sources and backgrounds should be included as much as possible.

[0024] When the host updates the model, it can continuously optimize the model through incremental learning for new data without directly replacing the old model. When the model is manually adjusted, the model can be monitored regularly to ensure the effectiveness of the model in practical applications.

[0025] When in use, first input the data to be analyzed into the host, the host converts the data format, changes the image or text into TXT format, DOCX format, PDF format, CSV format, etc., uniformly processes data from different sources, cross-compares the formatted data, compares duplicate data, processes missing values, abnormal values, etc. in the data, removes duplicate data, processes missing values, abnormal values, etc. in the data after comparison, extracts nouns, verbs or adjectives in the data, scales the data to a standard range to re-form streamlined data, crawls the streamlined data through the network through artificial intelligence, and downloads and collects the crawled data. The collected data can be displayed on the display, the host selects artificial intelligence algorithms according to the type of streamlined data, such as the financial industry, medical field, retail industry and manufacturing industry, imports the streamlined data into the artificial intelligence algorithm for analysis, and the data initially analyzed by the artificial intelligence algorithm is compared with the network data The artificial intelligence algorithm is used to perform cross-comparison, and the self-analyzed data is judged according to the network data to determine whether there are major discrepancies in the data. After the judgment is completed, the artificial intelligence algorithm will perform analysis and display the analysis on the monitor. When updating the artificial intelligence algorithm, the staff can add a certain amount of redundancy to the streamlined data when the artificial intelligence algorithm analyzes the streamlined data, and interfere with the artificial intelligence algorithm by generating new samples to enhance the data analysis of the artificial intelligence algorithm. The staff can also perform secondary deletion of concentrated and repeated records in the streamlined data, or fill or delete missing data, and add the deleted or filled data to the artificial intelligence algorithm to increase the data analysis of the artificial intelligence algorithm. The artificial intelligence algorithm can learn from the new crawled data and update its parameters. After the model is updated, A / B testing can be used to evaluate whether the performance of the new model is better than the old model. By comparing the performance of the new and old models on the same data set, the model with better performance can be selected for deployment.

[0026] Convert image to TXT code example; Use Python's OCR library (such as Tesseract) to recognize the text in the image and then save it as a TXT file.

[0027] Install required libraries: pip install pytesseract pillow from PIL import Image import pytesseract def image_to_txt(image_path, output_path): Open the image img = Image.open(image_path) OCR recognition using Tesseract text = pytesseract.image_to_string(img) Save the recognized text to a TXT file with open(output_path, 'w', encoding='utf-8') as f: f.write(text) The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of the present invention is defined by the attached claims and their equivalents.

Claims

1. A computer data analysis method based on artificial intelligence, characterized in that: The system comprises a host computer, wherein the host computer is connected with a data collection module, a data analysis module, a data update module and a display.

2. The computer data analysis method based on artificial intelligence according to claim 1, characterized in that: The data collection module includes data preprocessing and data collection; The data analysis module includes algorithm selection and model training comparison; The data updating module includes manual data adjustment and data model updating.

3. The computer data analysis method based on artificial intelligence according to claim 2, characterized in that: The data preprocessing is that the user enters the data to be analyzed and collected into the host to form comparative data. The host first converts the comparative data into a format suitable for analysis, such as TXT format, DOCX format, PDF format, CSV format, to ensure that data from different sources can be uniformly processed and compared, and then removes duplicate data, missing values, abnormal values, etc. in the converted data, extracts nouns, verbs or adjectives in the data, and scales the data to a standard range to re-form streamlined data. The standard range can be the number of words in the streamlined data; The data collection is that the host transmits the simplified data to the crawler software, the crawler software compares the data through the network, crawls the data similar to the simplified data, and displays the crawled data through the display.

4. The computer data analysis method based on artificial intelligence according to claim 2, characterized in that: The algorithm selection is to select a suitable artificial intelligence algorithm based on the characteristics of the collected data and the task requirements, which is used in a dynamic environment where decision-making is required and is suitable for solving multi-step decision-making problems; The model training comparison is to train the data set according to the selected algorithm. During the data training process, the test data set is first used to evaluate the generalization ability of the model. The evaluation indicators include accuracy, precision, recall rate and F1 value. After the data to be analyzed is input into the model, the parameters of the model are adjusted through the training set, the performance of the model is evaluated through cross-validation and other methods, and the model is optimized by adjusting the learning rate, regularization, etc., and the model is prevented from overfitting. At the same time, it is ensured that the model has sufficient generalization ability. During training, the data crawled by the crawler can be compared with the data after model analysis, and the data output by the model can be compared and analyzed to ensure the accuracy of the model.

5. The computer data analysis method based on artificial intelligence according to claim 2, characterized in that: The manual data adjustment is to perform secondary deletion of the concentrated repeated records in the streamlined data when the host processes the data, so as to avoid redundancy and deviation in the model training process, and adopt a filling or deletion strategy for the missing data, and modify the data by a certain amount. The modification amount depends on the number and distribution of the missing values. During the model training analysis, the streamlined data can be interfered by adding text data and other methods to generate new samples to enhance data analysis; The data model update is the host selecting and adjusting the model hyperparameters during the model training process to achieve the purpose of optimizing the model performance. The hyperparameters are parameters set before model training, such as learning rate, batch size, regularization coefficient, etc., which are adjusted through online learning on the network. The model learns and updates its parameters from the new data extracted by the crawler. After the model is updated, A / B testing can be used to evaluate whether the performance of the new model is better than the old model. By comparing the performance of the new and old models on the same data set, a model with better performance can be selected for deployment.

6. The computer data analysis method based on artificial intelligence according to claim 4, characterized in that: The artificial intelligence algorithms include supervised learning and unsupervised learning. The supervised learning algorithms, such as regression analysis, decision trees, support vector machines, neural networks, etc., are applicable to labeled data sets. The unsupervised learning algorithms, such as hierarchical clustering, etc., are applicable to unlabeled scenarios and are usually used for data exploration, pattern recognition, etc.

7. The computer data analysis method based on artificial intelligence according to claim 5, characterized in that: During the manual data adjustment process, it is necessary to ensure the privacy and security of the data to avoid leakage of sensitive information. When adjusting the data, data from different sources and backgrounds should be included as much as possible.

8. The computer data analysis method based on artificial intelligence according to claim 5, characterized in that: When the host updates the model, it can continuously optimize the model through incremental learning for new data without directly replacing the old model. When the model is manually adjusted, the model can be monitored regularly to ensure the effectiveness of the model in practical applications.