Operation redisk data mining method and system
By formulating unified data specifications and standards, establishing diversified tables, and implementing fine data preprocessing processes, the problems of data inconsistency and diverse formats are solved, and efficient and unified data sorting and management are achieved, and the accuracy and efficiency of business review are improved.
Patent Information
- Application Number
- CN202510337439.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, data standardization and collation lack uniform specifications, resulting in data inconsistency and diverse formats that are difficult to manage in a unified manner, increasing the complexity of data processing and reducing efficiency, affecting the accuracy and efficiency of business review.
Formulate unified data specifications and standards, establish diversified standards and format tables, accurately fill in the corresponding tables through the standardized data processed, and implement a fine data preprocessing process, including data integration, cleaning, conversion and quality inspection to ensure data consistency and comparability.
It significantly enhances the readability and comparability of data, simplifies the complexity of data management, provides convenient and reliable data support for subsequent data analysis and applications, and improves the accuracy and efficiency of data processing.
Smart Images

Figure CN120336748A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for mining business review data. Background Art
[0002] In the current business environment, data has become an important basis for enterprise decision-making. As an important part of enterprise management, the quality and efficiency of business review directly affect the strategic planning and business adjustment of the enterprise. However, in actual operation, enterprises often face many challenges in data standardization and collation.
[0003] The data standardization and collation process in the prior art often lacks unified data specifications and standards, resulting in inconsistencies in data at the initial generation stage. Such inconsistencies not only increase the difficulty of subsequent data processing, but may also lead to data misinterpretation and misuse, thus affecting the accuracy of business review. In addition, due to the wide variety of data sources and formats, it is difficult for the prior art to achieve unified and effective management of these data. Different formats of data require different processing methods, which not only increases the complexity of data processing, but also reduces the efficiency of data processing.
[0004] In view of the above problems, the present invention proposes a new method and system for mining business review data, aiming to achieve efficient and unified collation of data by formulating unified data specifications and standards, and establishing diversified standard and format tables, so as to provide more convenient and reliable data support for subsequent data analysis and application. Summary of the Invention
[0005] In order to overcome the problems raised in the above background art, the present invention proposes a method and system for mining business review data.
[0006] The technical solution of the present invention is as follows: A method for mining business review data includes the following steps: S11: Determine the goals and scope, define the goals of the review, and define the time range, business areas involved, and key indicators of the review; S12: Data preparation, collect data from the set data sources, and preprocess the data to remove useless data; S13: Data exploration, explore the patterns and anomalies in the data through data exploration techniques, where the data exploration techniques used include data visualization and data clustering; S14: Model construction, select an analysis model according to the research problem and data characteristics, and train and validate the model; S15: Model application, comprehensively evaluate the performance of the model, and apply the evaluated model to the actual business scenario.
[0007] Preferably, when defining the objectives of the review, as well as the time scope, business areas involved, and key metrics of the review, the following steps are included: S21: Clearly define the task background and understand the origin, purpose, and importance of the task; S22: Set specific goals and formulate quantifiable and achievable goals based on the task background; S23: Determine the key metrics, identify the key factors affecting the achievement of the goals, and set corresponding measurement metrics.
[0008] Preferably, when collecting data from the set data sources and preprocessing the data to remove useless data, the following steps are included: S31: Determine the data sources, and based on the review objectives and business areas, determine the data sources, where the data sources include internal databases, external databases, online data, and sensor data; S32: Design a data collection plan and formulate a detailed data collection plan, where the data collection plan includes collection methods, collection tools, and collection frequencies; S33: Execute data collection and carry out data collection work according to the set data collection plan; S34: Data preprocessing, perform preprocessing operations on the data, where the preprocessing operations adopted include data cleaning and data transformation.
[0009] Preferably, when performing preprocessing operations on the data, the following steps are included: S41: Data integration, integrate the collected data, where the integration operations adopted include data merging, data deduplication, and data association; S42: Data cleaning and preprocessing, clean the integrated data, including handling missing values, handling outliers, and handling duplicate values, and perform data transformation, data classification, and data encoding on the data; S43: Data standardization and formatting, first perform standardization processing on the data, and then uniformly organize the data; S44: Data quality inspection, perform quality inspection on the integrated, cleaned, preprocessed, and standardized data through data quality tools; S45: Data storage, store the processed data and establish a data update backup mechanism.
[0010] Preferably, when performing standardization processing on the data and then uniformly organizing the data, the processing rules adopted are: A11: First formulate unified data specifications and standards, and when generating data, enter the data according to the specifications and standards; A12: For data in different formats, establish tables with different standards and formats, and achieve unified collation of data by filling the standardized data into the tables.
[0011] Preferably, when exploring patterns and anomalies in data through data exploration techniques, the following steps are included: S51: Data visualization, presenting the data in a visual form using visualization tools, where the visual forms adopted include bar charts, pie charts, scatter plots, histograms, and box plots; S52: Identify patterns and trends, identify patterns and trends in the data through visual analysis, including seasonal fluctuations in sales and the distribution of customer groups; S53: Data clustering, grouping similar data points using clustering algorithms to discover potential categories and groups in the data, where the clustering algorithms adopted include K-means clustering, hierarchical clustering, and DBSCA; S54: Outlier detection, detecting outliers in the data through statistical methods to identify possible problems.
[0012] Preferably, when identifying patterns and trends in the data, correlation analysis is also included, where the principle formula of correlation analysis is: ; where x i and y i are the observed values of two variables respectively, and are the means of x i and y i respectively.
[0013] Preferably, when selecting an analysis model according to the research question and data characteristics and training and validating the model, the following steps are included: S61: Select the model type, select a suitable analysis model type according to the review objective and data characteristics; S62: Determine the model parameters, determine the parameters of the model according to the data type and analysis requirements, where the parameters of the model include regression coefficients and classification thresholds; S63: Train the model, train the model using the training dataset and adjust the model parameters by minimizing the error function; S64: Validate the model, evaluate the performance of the model using the validation dataset; S65: Optimize the model, optimize the model according to the validation results.
[0014] Preferably, when selecting a suitable analysis model type according to the review objective and data characteristics, the analysis model types include: A21: Regression model, the principle formula is: ; Among them, y i represents the actual value of the i-th sample, represents the intercept term of the model, represents the weight of the j-th feature, x ij represents the value of the j-th feature in the i-th sample; A22: Classification model, its principle formula is: ; Among them, is the probability that the dependent variable Y takes the value of 1 under the condition of the given independent variable X A23: Clustering model, its principle formula is: ; Among them, J represents the total sum of squared errors of the clustering effect, represents the coordinate vector of the i-th sample point, represents the coordinate vector of the j-th cluster center.
[0015] An operation review data mining system includes: A data collection module for collecting raw data from different sources; A data storage module for effectively storing and managing the collected data; A data preprocessing module for cleaning, transforming, and integrating the collected raw data; A data analysis module for extracting valuable information and knowledge from the data using statistical analysis, machine learning, and deep learning techniques; A report generation module for organizing the data analysis results into the form of reports and charts; A data security module for protecting the data using encryption means and security protection means; A system integration and maintenance module for ensuring the normal operation and continuous optimization of the data mining system, including regularly checking, updating, and optimizing the system, as well as handling system failures and performance issues.
[0016] Advantages of the present invention: 1. Compared with the possible disadvantages in the prior art such as the lack of unified specifications, diverse formats that are difficult to manage uniformly in data standardization and collation, this solution first formulates a detailed set of data specifications and standards to ensure that data is entered in accordance with established rules at the initial stage of data generation, enhancing the consistency of data from the source. Furthermore, for data in different formats, this solution innovatively establishes diverse standard and format tables. By accurately filling the standardized data into the corresponding tables, efficient and unified data collation is achieved. This solution not only significantly enhances the readability and comparability of data but also greatly simplifies the complexity of data management, providing more convenient and reliable data support for subsequent data analysis and applications; 2. Compared with the possible disadvantages in the prior art such as the absence, lack of systematicness, or lax quality control in data preprocessing steps, this solution adopts a complete and refined data preprocessing process, including data integration, in-depth cleaning and preprocessing, standardization and formatting, strict quality inspection, and secure and efficient storage management. This solution ensures the integrity of data through data merging, deduplication, and correlation, improves the accuracy and usability of data by handling missing values, outliers, and duplicate values, as well as performing data conversion, classification, and encoding. The data standardization and formatting steps further ensure the consistency and comparability of data, while strict data quality inspection effectively avoids potential data errors. Finally, through a secure data storage and update backup mechanism, the persistence and recoverability of data are guaranteed. This solution significantly improves the effect of data preprocessing, providing a more reliable and high-quality data foundation for subsequent data analysis and business decision-making; 3. This invention integrates multiple modules such as data collection, storage, preprocessing, analysis, report generation, data security, and system integration and maintenance, forming an efficient, secure, and comprehensive data mining solution. By automatically collecting and storing multi-source data, implementing fine-grained data preprocessing, and applying advanced statistical analysis, machine learning, and deep learning technologies for in-depth mining, this system can quickly extract valuable information and knowledge and present it in the form of intuitive reports and charts. At the same time, a powerful data security module ensures the security and privacy protection of data, while the system integration and maintenance module guarantees the stable operation and continuous optimization of the system, overall improving the efficiency and accuracy of enterprise operation review. Brief Description of the Drawings
[0017] Figure 1 Shown is a schematic flow chart of the operation review data mining method of this invention; Figure 2 Shown is a schematic structural diagram of the operation review data mining system of this invention. Detailed Embodiments
[0018] The present invention will be further described below in conjunction with the drawings and embodiments.
[0019] Please refer to Figure 1 , the present invention provides an embodiment: a method for mining business review data, including the following steps: S11: Determine the goals and scope, define the goals of the review, and define the time range, business areas involved, and key indicators of the review; S12: Data preparation, collect data from the set data sources, and preprocess the data to remove useless data; S13: Data exploration, explore the patterns and anomalies in the data through data exploration techniques, where the data exploration techniques used include data visualization and data clustering; S14: Model construction, select an analysis model according to the research question and data characteristics, and train and validate the model; S15: Model application, comprehensively evaluate the performance of the model, and apply the evaluated model to the actual business scenario.
[0020] As described above, the present invention effectively improves the accuracy and efficiency of business review by clearly defining the review goals and scope, carefully preparing and preprocessing data, deeply exploring the patterns and anomalies in the data, scientifically constructing and validating the analysis model, and comprehensively evaluating the model performance and applying it to the actual business scenario, provides strong data support for enterprise decision-making, and promotes business optimization and growth.
[0021] Preferably, when defining the goals of the review and defining the time range, business areas involved, and key indicators of the review, it includes the following steps: S21: Clarify the task background, understand the origin, purpose, and importance of the task; S22: Set specific goals, and formulate quantifiable and achievable goals according to the task background; S23: Determine the key indicators, identify the key factors affecting the achievement of the goals, and set corresponding measurement indicators.
[0022] As described above, the present invention grasps the origin, purpose, and importance by clarifying the task background, then sets specific, quantifiable, and achievable goals, and at the same time accurately identifies and sets the key indicators affecting the achievement of the goals. This systematic process greatly improves the pertinence and effectiveness of the review, ensures that the review activities can focus on the core issues, and lays a solid foundation for subsequent data analysis and business optimization.
[0023] Preferably, when collecting data from the set data sources, preprocessing the data, and removing useless data, it includes the following steps: S31: Determine the data sources. Based on the review objectives and business areas, determine the data sources, where the data sources include internal databases, external databases, online data, and sensor data; S32: Design a data collection plan. Develop a detailed data collection plan, where the data collection plan includes collection methods, collection tools, and collection frequencies; S33: Execute data collection. Conduct data collection work according to the set data collection plan; S34: Data preprocessing. Perform preprocessing operations on the data, where the preprocessing operations adopted include data cleaning and data transformation.
[0024] As described above, the present invention accurately determines the data sources, carefully designs the data collection plan, efficiently executes the data collection tasks, and uses preprocessing means such as data cleaning and data transformation to remove useless data, effectively ensuring the comprehensiveness, accuracy, and availability of the data, laying a solid data foundation for subsequent data exploration and analysis, and improving the quality and efficiency of data review.
[0025] Preferably, when performing preprocessing operations on the data, the following steps are included: S41: Data integration. Integrate the collected data, where the integration operations adopted include data merging, data deduplication, and data association; S42: Data cleaning and preprocessing. Clean the integrated data, including handling missing values, handling outliers, and handling duplicate values, and perform data transformation, data classification, and data encoding on the data; S43: Data standardization and formatting. First, perform standardization processing on the data, and then uniformly organize the data; S44: Data quality inspection. Conduct quality inspection on the integrated, cleaned, preprocessed, and standardized data through data quality tools; S45: Data storage. Store the processed data and establish a data update and backup mechanism.
[0026] As described above, in view of the possible deficiencies, lack of systematicness or lax quality control in the data preprocessing steps in the prior art, the present solution adopts a complete and refined data preprocessing process, including data integration, in-depth cleaning and preprocessing, standardization and formatting, strict quality inspection, and safe and efficient storage management. This solution ensures data integrity through data merging, deduplication, and correlation, and improves data accuracy and usability by handling missing values, outliers, and duplicate values, as well as performing data conversion, classification, and encoding. The data standardization and formatting steps further ensure data consistency and comparability, while strict data quality inspection effectively avoids potential data errors. Finally, through a secure data storage and update backup mechanism, the persistence and recoverability of data are guaranteed. This solution significantly improves the effect of data preprocessing and provides a more reliable and high-quality data foundation for subsequent data analysis and business decision-making.
[0027] Preferably, when standardizing the data and then uniformly organizing the data, the processing rules adopted are as follows: A11: First, formulate unified data specifications and standards, and when generating data, enter the data according to the specifications and standards. A12: For data in different formats, establish tables with different standards and formats, and achieve the unified organization of data by filling the standardized data into the tables.
[0028] As described above, in view of the possible deficiencies such as lack of unified specifications and difficulty in unified management due to diverse formats in data standardization and organization in the prior art, the present solution first formulates a detailed set of data specifications and standards to ensure that data is entered according to the established rules at the initial stage of data generation, thus enhancing data consistency from the source. Furthermore, for data in different formats, the present solution innovatively establishes diverse standard and format tables, and achieves the efficient unified organization of data by accurately filling the standardized data into the corresponding tables. This solution not only significantly enhances the readability and comparability of data, but also greatly simplifies the complexity of data management, providing more convenient and reliable data support for subsequent data analysis and applications.
[0029] Preferably, when exploring patterns and anomalies in data through data exploration techniques, the following steps are included: S51: Data visualization, presenting the data in a visual form using visualization tools, where the visual forms include bar charts, pie charts, scatter plots, histograms, and box plots. S52: Identify patterns and trends, and identify patterns and trends in the data through visual analysis, including seasonal fluctuations in sales and the distribution of customer groups. S53: Data clustering, using clustering algorithms to group similar data points and discover potential categories and groups in the data. Among them, the clustering algorithms adopted include K-means clustering, hierarchical clustering, and DBSCAN.
[0030] S54: Outlier detection, detecting outliers in the data through statistical methods to identify possible problems.
[0031] As described above, the present invention intuitively presents data characteristics through various visualization forms such as bar charts, pie charts, scatter plots, histograms, and box plots, making the patterns and trends in the data clear at a glance, such as the seasonal fluctuations of sales and the distribution of customer groups. At the same time, by using clustering algorithms such as K-means clustering, hierarchical clustering, and DBSCAN, similar data points are accurately grouped, effectively discovering potential categories and groups in the data. This series of measures not only greatly enhances the depth and breadth of data exploration, but also provides strong support for subsequent data analysis and business decision-making, enabling enterprises to more accurately grasp market dynamics and optimize business strategies.
[0032] As a preference, when identifying patterns and trends in the data, it also includes performing correlation analysis. Among them, the principle formula of correlation analysis is: ; where x i and y i are respectively the observed values of two variables, and are respectively the means of x i and y i .
[0033] As described above, the present invention uses statistical principles to calculate the ratio of the covariance between the observed values of two variables and their respective means to the standard deviation, that is, the correlation coefficient, to accurately quantify the linear correlation degree between variables. The implementation of this technical solution enables enterprises to more accurately grasp the internal relationship between data variables, reveal hidden laws and trends, and provide a more scientific and rigorous basis for data-driven decision-making.
[0034] As a preference, when selecting an analysis model according to the research question and data characteristics and training and validating the model, it includes the following steps: S61: Select the model type, select a suitable analysis model type according to the review objective and data characteristics; S62: Determine the model parameters, determine the parameters of the model according to the data type and analysis requirements. Among them, the parameters of the model include regression coefficients and classification thresholds; S63: Train the model, train the model using the training data set and adjust the model parameters by minimizing the error function; S64: Verify the model and evaluate the performance of the model using the validation dataset; S65: Optimize the model and optimize the model according to the verification results.
[0035] As described above, the present invention first carefully selects the model type according to the review objective and data characteristics, and then accurately determines the model parameters according to the data type and analysis requirements. The model is fully trained using the training dataset, and the error function is minimized to optimize the parameter configuration. Then, the performance of the model is comprehensively evaluated using the validation dataset to ensure the accuracy and reliability of the model. Finally, the model is optimally adjusted according to the verification results. The implementation of this series of steps not only improves the pertinence and effectiveness of model construction, but also ensures the high-performance performance of the model in actual applications, providing more accurate and reliable data analysis support for enterprises.
[0036] Preferably, when selecting a suitable analysis model type according to the review objective and data characteristics, the analysis model types include: A21: Regression model, the principle formula is: ; where y i represents the actual value of the i-th sample, represents the intercept term of the model, represents the weight of the j-th feature, x ij represents the value of the j-th feature in the i-th sample; A22: Classification model, the principle formula is: ; where is the probability that the dependent variable Y takes the value of 1 under the condition of the given independent variable X A23: Clustering model, the principle formula is: ; where J represents the total error sum of squares of the clustering effect, represents the coordinate vector of the i-th sample point, represents the coordinate vector of the j-th cluster center.
[0037] As described above, the present invention flexibly employs various analysis models such as regression models, classification models, and clustering models. The regression model provides a powerful tool for prediction and interpretation by precisely fitting the relationship between features and dependent variables; the classification model can accurately determine the class membership of samples, providing a scientific basis for classification decisions; the clustering model effectively discovers potential groups and structures in the data by minimizing the sum of squared clustering errors. This technical solution that comprehensively uses multiple analysis models not only enhances the flexibility and adaptability of data analysis but also significantly improves the accuracy and depth of analysis, providing enterprises with more comprehensive and in-depth data insights.
[0038] Please refer to Figure 2 , the present invention provides an embodiment: an operation review data mining system, including: A data collection module for collecting raw data from different sources; A data storage module for effectively storing and managing the collected data; A data preprocessing module for cleaning, transforming, and integrating the collected raw data; A data analysis module for extracting valuable information and knowledge from the data using statistical analysis, machine learning, and deep learning techniques; A report generation module for organizing the data analysis results into the form of reports and charts; A data security module for protecting the data using encryption means and security protection means; A system integration and maintenance module for ensuring the normal operation and continuous optimization of the data mining system, including regularly checking, updating, and optimizing the system, as well as handling system failures and performance issues.
[0039] As described above, the present invention integrates multiple modules such as data collection, storage, preprocessing, analysis, report generation, data security, and system integration and maintenance, forming an efficient, secure, and comprehensive data mining solution. By automatically collecting and storing multi-source data, implementing fine-grained data preprocessing, and using advanced statistical analysis, machine learning, and deep learning techniques for in-depth mining, the system can quickly extract valuable information and knowledge and present them in the form of intuitive reports and charts. At the same time, the powerful data security module ensures the security and privacy protection of the data, while the system integration and maintenance module guarantees the stable operation and continuous optimization of the system, overall improving the efficiency and accuracy of enterprise operation review.
[0040] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the purpose of the present invention.
Claims
1. A method for mining business review data, characterized in that: It includes the following steps: S11: Determine the objectives and scope, define the objectives of the review, and define the time scope, business areas involved, and key indicators of the review; S12: Data preparation, collect data from the set data sources, and preprocess the data to remove useless data; S13: Data exploration, explore the patterns and anomalies in the data through data exploration techniques. Among them, the data exploration techniques used include data visualization and data clustering; S14: Model construction, select an analysis model according to the research question and data characteristics, and train and validate the model; S15: Model application, comprehensively evaluate the performance of the model, and apply the evaluated model to the actual business scenario.
2. The operation review data mining method according to claim 1, wherein: When defining the objectives of the review, and defining the time scope, business areas involved, and key indicators of the review, it includes the following steps: S21: Clarify the task background, understand the origin, purpose, and importance of the task; S22: Set specific objectives, and formulate quantifiable and achievable objectives according to the task background; S23: Determine the key indicators, identify the key factors affecting the achievement of the objectives, and set corresponding measurement indicators.
3. The method for mining operation review data according to claim 2, wherein: When collecting data from the set data sources, and preprocessing the data to remove useless data, it includes the following steps: S31: Determine the data sources, determine the data sources according to the review objectives and business areas. Among them, the data sources include internal databases, external databases, online data, and sensor data; S32: Design a data collection plan, formulate a detailed data collection plan. Among them, the data collection plan includes collection methods, collection tools, and collection frequencies; S33: Execute data collection, and carry out data collection work according to the set data collection plan; S34: Data preprocessing, perform preprocessing operations on the data. Among them, the preprocessing operations used include data cleaning and data transformation.
4. A method for mining operation review data according to claim 3, characterized in that: When performing preprocessing operations on the data, it includes the following steps: S41: Data integration, integrate the collected data. Among them, the integration operations used include data merging, data deduplication, and data association; S42: Data cleaning and preprocessing, clean the integrated data, including handling missing values, handling outliers, and handling duplicate values, and perform data transformation, data classification, and data encoding on the data; S43: Data standardization and formatting, first perform standardization processing on the data, and then uniformly organize the data; S44: Data quality inspection, perform quality inspection on the integrated, cleaned, preprocessed, and standardized data through data quality tools; S45: Data storage, store the processed data, and establish a data update backup mechanism.
5. A business review data mining method according to claim 4, characterized in that: When performing standardization processing on the data, and then uniformly organizing the data, the processing rules used are: A11: First, formulate unified data specifications and standards, and when generating data, enter the data according to the specifications and standards; A12: For data in different formats, establish tables with different standards and formats, and realize the unified organization of the data by filling the standardized processed data into the tables.
6. The operation review data mining method according to claim 5, wherein: When exploring patterns and anomalies in data through data exploration techniques, the following steps are included: S51: Data visualization, presenting data in a visual form using visualization tools, where the visual forms include bar charts, pie charts, scatter plots, histograms, and box plots; S52: Identifying patterns and trends, identifying patterns and trends in data through visual analysis, including seasonal fluctuations in sales and the distribution of customer groups; S53: Data clustering, grouping similar data points using clustering algorithms to discover potential classes and groups in the data, where the clustering algorithms used include K-means clustering, hierarchical clustering, and DBSCA; S54: Outlier detection, detecting outliers in data through statistical methods to identify potential problems.
7. A method for mining operation review data according to claim 6, characterized in that: When identifying patterns and trends in data, correlation analysis is also included, where the principle formula for correlation analysis is: ; where x i and y i are the observed values of two variables, and are the means of x i and y i respectively.
8. A method for mining operation review data according to claim 7, characterized in that: When selecting an analysis model based on the review objective and data characteristics, and training and validating the model, the following steps are included: S61: Selecting the model type, selecting a suitable analysis model type based on the review objective and data characteristics; S62: Determining model parameters, determining the parameters of the model according to the data type and analysis requirements, where the parameters of the model include regression coefficients and classification thresholds; S63: Training the model, training the model using the training dataset and adjusting the model parameters by minimizing the error function; S64: Validating the model, evaluating the performance of the model using the validation dataset; S65: Optimizing the model, optimizing the model based on the validation results.
9. A method for mining business review data according to claim 8, characterized in that: When selecting a suitable analysis model type according to the review objective and data characteristics, the analysis model types include: A21: Regression model, the principle formula is: ; Among them, y i represents the actual value of the i-th sample, represents the intercept term of the model, represents the weight of the j-th feature, x ij represents the value of the j-th feature in the i-th sample; A22: Classification model, the principle formula is: ; wherein, is the probability that the dependent variable Y takes the value of 1 under the condition of a given independent variable X A23: Clustering model, the principle formula is: ; Among them, J represents the total sum of squared errors of the clustering effect, represents the coordinate vector of the i-th sample point, represents the coordinate vector of the j-th cluster center.
10. A business review data mining system, characterized in that: Including: Data acquisition module, used to collect raw data from different sources; Data storage module, used to effectively store and manage the collected data; Data preprocessing module, used to clean, transform, and integrate the collected raw data; Data analysis module, used to extract valuable information and knowledge from data using statistical analysis, machine learning, and deep learning techniques; Report generation module, used to organize the data analysis results into the form of reports and charts; Data security module, used to protect data using encryption means and security protection means; System integration and maintenance module, used to ensure the normal operation and continuous optimization of the data mining system, including regularly checking, updating, and optimizing the system, as well as handling system failures and performance issues.