Big data-based comprehensive evaluation data analysis method and system
By constructing data evaluation dimensions of big data and determining weights using the global entropy method, and monitoring indicator changes in real time, the problems of low data processing efficiency and insufficient evaluation accuracy in traditional methods are solved, and efficient and accurate comprehensive evaluation is achieved.
Patent Information
- Application Number
- CN202510386347.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the big data era, traditional comprehensive evaluation data analysis methods have the problems of low data processing efficiency, single evaluation dimension, and highly subjective determination of indicator weights. They are difficult to meet the accuracy and timeliness requirements of data analysis and evaluation. Existing data are difficult to update in real time, resulting in waste of resources or low timeliness.
A comprehensive evaluation data analysis method based on big data is adopted. By obtaining preset category information, data evaluation dimensions are constructed, and the indicator weights are determined using the global entropy method. Weighted summation is performed through the comprehensive evaluation method. Indicator change information is monitored in real time, indicators are updated, and a comprehensive evaluation index is obtained.
It improves the efficiency and automation of data processing, ensures the timeliness and accuracy of evaluation, supports multi-level and multi-angle data analysis, provides a scientific basis for decision-making, and reduces labor costs.
Smart Images

Figure CN119884103B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application provides a big data-based comprehensive evaluation data analysis method and system, and relates to the technical field of data analysis, in particular to the technical field of comprehensive evaluation data analysis. BACKGROUND
[0002] As an important data analysis method, comprehensive evaluation data analysis aims to comprehensively and objectively evaluate complex systems by constructing reasonable evaluation dimensions and index systems. However, traditional comprehensive evaluation data analysis methods often have low data processing efficiency, single evaluation dimensions, strong subjectivity in determining index weights, and other problems, making it difficult to meet the accuracy and timeliness requirements of data analysis and evaluation in the big data era. Existing data cannot be updated according to data changes for various types of indicators, and real-time updating wastes data processing resources or updating too late leads to low data timeliness. SUMMARY
[0003] The application provides a big data-based comprehensive evaluation data analysis method and system to solve the above problems.
[0004] The application provides a big data-based comprehensive evaluation data analysis method and system, which comprises the following steps:
[0005] S1, obtaining preset category information and preset type information, and then obtaining first category information, second category information and third category information, constructing data evaluation dimensions, updating and matching category information, and obtaining dimension update information;
[0006] S2, collecting preset analysis demand data by a big data method to obtain analysis collection data, and classifying the analysis collection data by data evaluation dimensions to obtain a data analysis library;
[0007] S3, determining the weight data of each index of the third category by a global entropy method, and using a comprehensive evaluation method to weight and sum each index to obtain a comprehensive evaluation index;
[0008] S4, obtaining data change information of the indexes of the third category and the second category, calculating the third category index update value and the second category index update value, determining whether to update the indexes of the second category and the first category, and obtaining an updated comprehensive evaluation update index.
[0009] Further, the S1 comprises the following steps:
[0010] obtaining preset first category information;
[0011] classifying the first category information according to preset second type information to obtain second category information;
[0012] Classify the second category information according to the preset third category information to obtain third category information;
[0013] The first category information includes the second category information, and the second category information includes the third category information;
[0014] Combining the first category information, the second category information, and the third category information to obtain data evaluation dimensions;
[0015] Acquire category update information, and perform information matching on the category update information with the first category information, the second category information, and the third category information, respectively, to obtain a matching category level corresponding to the category update information;
[0016] Performing a preset category update on the matching category level using the category update information to obtain category update information of the corresponding matching category level;
[0017] performing a preset category update on the lower category information of the matching category level to obtain category update information of the lower category information;
[0018] After the data is updated, the first category information, the second category information, and the third category information are obtained and combined to obtain dimension update information.
[0019] Furthermore, the S2 includes:
[0020] Acquire preset analysis requirement information, collect the preset analysis requirement information through big data methods, and obtain analysis and collection data;
[0021] Preprocessing the analysis and collection data to obtain preprocessed analysis and collection data;
[0022] Classify the analyzed collected data according to the first category information, the second category information, and the third category information of the data evaluation dimension to obtain classified first category data, second category data, and third category data;
[0023] The classified first category data, second category data and third category data are combined to obtain a data analysis library.
[0024] Furthermore, the S3 includes:
[0025] Establish a three-dimensional time series data table based on the data analysis library and form an initial global evaluation matrix;
[0026] The extreme value method is used to standardize each indicator, and the standardized data is distributed between 0.1-0.9. The specific standardization formula includes positive indicators and negative indicators;
[0027] The indicators are normalized to obtain normalized indicators;
[0028] The entropy value and redundancy of each indicator are calculated;
[0029] The weight of each indicator is determined;
[0030] The comprehensive evaluation index is obtained by comprehensive weighting.
[0031] Further, the S4 comprises:
[0032] The data variation information of the third category of indicators is obtained in real time, and the third category indicator update value is calculated according to the data variation information of the third category of indicators;
[0033] The third category indicator update value is compared with the preset three-type update threshold to obtain a three-level comparison result;
[0034] The second category of indicators is updated according to the three-level comparison result, the data variation information of the second category of indicators is obtained, and the second category indicator update value is calculated according to the data variation information of the second category of indicators;
[0035] The second category indicator update value is compared with the preset two-type update threshold to obtain a two-level comparison result;
[0036] The first category of indicators is updated according to the two-level comparison result, and then the indicator update database is obtained;
[0037] The comprehensive evaluation update index of the indicator update database is obtained.
[0038] Further, the system comprises:
[0039] The category dimension construction module is used to obtain preset category information and preset kind information, and then obtain first category information, second category information and third category information, construct data evaluation dimensions, update and match category information, and obtain dimension update information;
[0040] The data acquisition classification module is used to collect preset analysis demand data by a big data method to obtain analysis collection data, and divide and classify the analysis collection data by the data evaluation dimensions to obtain a data analysis library;
[0041] The comprehensive evaluation analysis module is used to determine the weight data of each indicator of the third category by a global entropy value method, and to obtain a comprehensive evaluation index by weighted summation of each indicator by a comprehensive evaluation method;
[0042] The index updating and evaluating module is configured to acquire data variation information of indexes of the third category and the second category, calculate third-category index updating values and second-category index updating values, determine whether to update indexes of the second category and the first category, and acquire an updated comprehensive evaluation updating index.
[0043] Further, the category dimension construction module comprises:
[0044] The preset category division module is configured to acquire preset first-category information.
[0045] The first-category information is classified according to preset second-category information to obtain second-category information.
[0046] The second-category information is classified according to preset third-category information to obtain third-category information.
[0047] The first-category information comprises the second-category information, and the second-category information comprises the third-category information.
[0048] The first-category information, the second-category information and the third-category information are combined to obtain data evaluation dimensions.
[0049] The preset category updating module is configured to acquire category updating information, match the category updating information with the first-category information, the second-category information and the third-category information respectively to obtain matching category levels corresponding to the category updating information.
[0050] The matching category levels are updated according to the category updating information to obtain category updating information of the corresponding matching category levels.
[0051] Lower-category information of the matching category levels is updated according to the preset category to obtain category updating information of the lower-category information.
[0052] After data updating, the first-category information, the second-category information and the third-category information are combined to obtain dimension updating information.
[0053] Further, the data collection and classification module comprises:
[0054] The data collection module is configured to acquire preset analysis requirement information, collect the preset analysis requirement information by using a big data method, and obtain analysis collection data.
[0055] The analysis collection data is preprocessed to obtain preprocessed analysis collection data.
[0056] The collection data classification module is configured to classify the analysis collection data according to the first category information, the second category information and the third category information of the data evaluation dimension, and obtain first category data, second category data and third category data after classification.
[0057] The analysis library establishment module is configured to combine the first category data, the second category data and the third category data after classification, and obtain a data analysis library.
[0058] Further, the comprehensive evaluation analysis module comprises:
[0059] The standardization processing module is configured to establish a three-dimensional time sequence data table according to the data analysis library, and form an initial global evaluation matrix.
[0060] The extreme value method is used to standardize each index, and the standardized data is distributed between 0.1 and 0.9. The specific standardization formula includes positive indexes and negative indexes.
[0061] The index data processing module is configured to normalize each index, and obtain index normalization.
[0062] The entropy value and redundancy of each index are calculated.
[0063] The weight of each index is determined.
[0064] The comprehensive evaluation calculation module is configured to perform comprehensive weighting, and obtain a comprehensive evaluation index.
[0065] Further, the index update evaluation module comprises:
[0066] The three-type update analysis module is configured to obtain data change information of the third category of indexes in real time, calculate a third category index update value according to the data change information of the third category of indexes.
[0067] The third category index update value is compared with a preset three-type update threshold value, and a three-level comparison result is obtained.
[0068] The two-type update analysis module is configured to update the second category of indexes according to the three-level comparison result, obtain data change information of the second category of indexes, and calculate a second category index update value according to the data change information of the second category of indexes.
[0069] The second category index update value is compared with a preset two-type update threshold value, and a two-level comparison result is obtained.
[0070] The first category of indexes is updated according to the two-level comparison result, and an index update database is further obtained.
[0071] The comprehensive update evaluation module is configured to acquire a comprehensive evaluation update index of the index update database.
[0072] The present application has the following advantages: by constructing data evaluation dimensions and using a global entropy value method to determine weights, the system can more accurately reflect the actual state and value of data, avoiding interference from human factors. The system can continuously monitor data change information and update indexes and comprehensive evaluation updates as needed, ensuring the timeliness and accuracy of evaluations. By constructing different levels of category information and data evaluation dimensions, the system can support multi-level, multi-angle data analysis, meeting the needs of different users and business scenarios. Using big data methods and automated processes for data collection, classification, and evaluation calculation greatly improves the efficiency and automation of data processing, reducing labor costs. The comprehensive evaluation index and update index provided by the system can provide scientific basis and data support for decision-making, improving the scientificity and accuracy of decision-making. BRIEF DESCRIPTION OF DRAWINGS
[0073] Figure 1 FIG. 1 is a schematic diagram of a comprehensive evaluation data analysis method based on big data. DETAILED DESCRIPTION
[0074] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0075] In one embodiment of the present application, the present application proposes a comprehensive evaluation data analysis method and system based on big data, and the method comprises:
[0076] S1, acquire preset category information and preset kind information, and then acquire first category information, second category information, and third category information, construct data evaluation dimensions, update and match the category information, and acquire dimension update information;
[0077] S2, collect preset analysis demand data by a big data method to obtain analysis collection data, and divide and classify the analysis collection data by the data evaluation dimensions to obtain a data analysis library;
[0078] S3, determine the weight data of each index of the third category by a global entropy value method, and use a comprehensive evaluation method to weight and sum each index to obtain a comprehensive evaluation index;
[0079] S4, acquire data change information of the indexes of the third category and the second category, calculate the third category index update value and the second category index update value, determine whether to update the indexes of the second category and the first category, and acquire an updated comprehensive evaluation update index, as shown in Figure 1
[0080] The working principle of the above technical solution is as follows: in the initial stage, the system obtains preset category information and type information. Based on these preset information, the system further refines and constructs first category information, second category information and third category information, which represent different data levels or analysis angles. The system constructs a data evaluation dimension, and also continuously updates and matches the preset category information to ensure the timeliness and accuracy of the preset categories of the data evaluation dimension. The preset analysis requirement data is collected by a big data method. These data may come from various data sources such as databases, log files, sensors, etc. The collected data is classified by the previously constructed data evaluation dimension to form a structured data analysis library. This step ensures the orderliness and analyzability of the data. The system uses a global entropy value method to determine the weight data of each index in the third category. The global entropy value method is a method based on information entropy, which can objectively reflect the information amount and importance of the index data. After determining the weight, the system uses a comprehensive evaluation method to weight and sum each index to calculate a comprehensive evaluation index. This index reflects the overall performance or state of the third category data. The system continuously monitors the data variation information of the third category and the second category index, in order to timely discover the abnormality or trend change in the data. Based on this variation information, the system calculates the third category index update value and the second category index update value, and judges whether the second category and the first category index need to be updated. If it needs to be updated, the system recalculates the comprehensive evaluation update index to reflect the latest data state and evaluation result.
[0081] The technical effect of the above technical solution is as follows: by constructing the data evaluation dimension and using the global entropy value method to determine the weight, the system can more accurately reflect the actual state and value of the data, avoiding the interference of human factors. The system can continuously monitor the data variation information, and update the index and the comprehensive evaluation according to the need, ensuring the timeliness and accuracy of the evaluation. By constructing different levels of category information and data evaluation dimensions, the system can support multi-level and multi-angle data analysis, meeting the needs of different users and business scenarios. Using big data methods and automated processes for data collection, classification and evaluation calculation greatly improves the efficiency and automation level of data processing, reducing labor costs. The comprehensive evaluation index and update index provided by the system can provide scientific basis and data support for decision-making, which can improve the scientificity and accuracy of decision-making.
[0082] In an embodiment of the present application, the S1 comprises:
[0083] obtaining preset first category information;
[0084] classifying the first category information according to preset second type information to obtain second category information;
[0085] According to the preset third kind of information, the second kind of information is classified to obtain third kind of information;
[0086] The first kind of information includes second kind of information, and the second kind of information includes third kind of information;
[0087] The first kind of information, the second kind of information and the third kind of information are combined to obtain data evaluation dimension;
[0088] The category update information is obtained, and the category update information is matched with the first kind of information, the second kind of information and the third kind of information respectively to obtain the matching category level corresponding to the category update information; the matching category level includes the first kind of information, the second kind of information and the third kind of information; when the category update information is matched with the first kind of information, the matching category level corresponding to the category update information is the category level of the first kind of information; the matching is passed through the calculation of information similarity, and the calculation result is compared with a preset matching threshold; according to the comparison result, the category matching information is matched; the similarity calculation and the threshold setting can be set by the person skilled in the art according to the prior art.
[0089] The matching category level is updated by the category update information to obtain the category update information of the corresponding matching category level;
[0090] The lower category information of the matching category level is updated by preset category to obtain the category update information of the lower category information; for example, after the first kind of information is updated, the second and third kind of information are adaptively updated.
[0091] After the data is updated, the first kind of information, the second kind of information and the third kind of information are combined to obtain the dimension update information.
[0092] The working principle of the above technical solution is as follows: the first type of information is obtained, which constitutes the basic framework of data analysis. The first type of information is classified according to the second type of information, forming the second type of information. This step is a further refinement and organization of the first type of information. The second type of information is classified according to the third type of information, obtaining the third type of information. At this point, the data has been divided into three levels, each level representing a different data perspective or analysis dimension. The first type of information, the second type of information and the third type of information are combined to form the data evaluation dimension. Obtain category update information, which comes from external data sources or internal data updates. Match the first type of information, the second type of information and the third type of information with the category update information to determine the category level to which the update information belongs. The matching process involves information similarity calculation, and by comparing the calculation result with the preset matching threshold, the system can determine which category level the update information is most matched with. Once the matching is successful, the system will update the matching category level with the preset category update, and at the same time, in order to ensure the consistency and integrity of the data, the system will also adaptively update the lower category information of the matching category level. For example, if the first type of information is updated, the system will adjust the second and third type of information accordingly. After completing the category update, the system re-obtains the first type of information, the second type of information and the third type of information, and combines them to obtain the latest dimension update information. These information reflect the latest state of the data evaluation dimension.
[0093] The technical effect of the above technical solution is that by dividing the data into different levels of category information, the system can more flexibly organize and manage the data, and at the same time, this hierarchical structure also makes the system easy to extend and adapt to new data requirements. By matching the category update information with the preset category level, the update information can be accurately applied to the correct data level. At the same time, the adaptive update of the lower category information ensures the consistency and integrity of the data, and improves the efficiency of data update. The constructed data evaluation dimension provides a basis for multi-level and multi-angle data analysis. The system can analyze and evaluate the data from different dimensions and perspectives according to user needs or business scenarios. By introducing intelligent methods such as information similarity calculation and matching threshold setting, the system can automatically complete the matching and processing of category update information, reducing the degree of manual intervention and improving the intelligent level of data processing.
[0094] In an embodiment of the present application, the S2 comprises:
[0095] Obtain preset analysis requirement information, collect the preset analysis requirement information by big data method, and obtain analysis collection data;
[0096] Preprocess the analysis collection data to obtain preprocessed analysis collection data;
[0097] Classify the analysis collection data according to the first category information, the second category information and the third category information of the data evaluation dimension to obtain classified first category data, second category data and third category data;
[0098] Combine the classified first category data, the second category data and the third category data to obtain a data analysis library.
[0099] The working principle of the above technical solution is as follows: the starting point of the whole process involves determining the specific content or target that needs to be analyzed. The preset analysis requirement information comes from business requirements, market analysis, data research and other aspects. According to the preset analysis requirement, relevant data is collected from various data sources (such as databases, log files, social media, Internet of Things devices, etc.) using big data methods (such as distributed storage, parallel computing, data mining, etc.). The purpose of this step is to collect as much information as possible related to the analysis requirement. The data collection method used is an existing data collection method in the prior art. The raw data collected often has problems such as noise, missing values, outliers, etc., so preprocessing is needed. The preprocessing steps may include data cleaning (removing invalid or incorrect data), data transformation (such as normalization, standardization), data integration (merging data from multiple data sources), etc. to obtain clean, consistent and usable analysis collection data. The preprocessed data is classified according to the first category information, the second category information and the third category information of the data evaluation dimension. These category information may be defined based on the nature, source, importance or other relevant characteristics of the data. The purpose of classification is to organize the data into a format that is easier to analyze and understand. The classified data (first category data, second category data and third category data) is combined to form a comprehensive data analysis library.
[0100] The technical effect of the above technical solution is as follows: through big data collection and preprocessing methods, the quality of the data used for analysis can be ensured, and the analysis errors caused by data problems can be reduced. Data classification and combination make the data more orderly and easy to access, thereby improving the efficiency of data analysis. Classification according to different data evaluation dimensions enables the data analysis library to support diverse analysis requirements. Whether it is trend analysis, correlation analysis or prediction analysis, appropriate data sets can be found in the classified data. A high-quality, structured data analysis library provides a solid foundation for decision-making. The analysis results based on these data can more accurately and reliably guide business decisions. By establishing a data analysis library, enterprises can better manage and utilize their data assets and improve data governance. This can ensure the compliance, security and accessibility of data, laying a foundation for the long-term development of the enterprise.
[0101] In one embodiment of the present application, the S3 comprises:
[0102] A stereoscopic time series data table is established according to the data analysis library, and an initial global evaluation matrix is formed;
[0103] An extreme value method is used to perform standardization processing on each index, and the standardized data is distributed between 0.1 and 0.9, and the specific standardization formula includes positive indexes and negative indexes;
[0104] Normalization is performed on each index to obtain index normalization;
[0105] Entropy and redundancy of each index are calculated;
[0106] The weight of each index is determined;
[0107] Comprehensive weighting is performed to obtain a comprehensive evaluation index.
[0108] The working principle of the above technical solution is that the comprehensive evaluation module comprises:
[0109] A matrix construction module is configured to establish a stereoscopic time series data table and form an initial global evaluation matrix;
[0110] The initial global evaluation matrix is:
[0111]
[0112] wherein, is the jth index of the ith county in the tth year, m is the number of evaluated counties, n is the number of evaluation indexes, and T is the number of evaluation years;
[0113] A standardization module is configured to perform standardization processing on each index, and the standardized data is distributed between 0.1 and 0.9, and the specific standardization formula includes positive indexes and negative indexes;
[0114] The calculation formula of the positive index is:
[0115]
[0116] The calculation formula of the negative index is:
[0117]
[0118] wherein, is the standardized value of , and are the maximum value and the minimum value of the jth index of all counties;
[0119] a normalization module configured to normalize each index to obtain index normalization;
[0120] The index normalization calculation formula is:
[0121]
[0122] wherein, is the normalized value of the jth index;
[0123] an entropy value calculation module configured to calculate the entropy value of each index, and the entropy value calculation formula is:
[0124]
[0125]
[0126] wherein, is the entropy value of the jth index;
[0127] a redundancy calculation module configured to calculate the redundancy of each index, and the redundancy calculation formula is:
[0128]
[0129] wherein, is the redundancy of the jth index;
[0130] a weight calculation module configured to calculate the weight of each index, and the weight calculation formula of each index is:
[0131]
[0132] wherein, W j is the weight of each index;
[0133] a comprehensive evaluation module configured to comprehensively weight each index, and the comprehensive weighting calculation formula is:
[0134]
[0135] wherein, is the comprehensive evaluation index.
[0136] Check the integrity, accuracy and consistency of the data, handle missing values, outliers, etc.
[0137] Use the pandas library of Python or other data processing tools to construct a three-dimensional data table containing time, county and index according to the collected data.
[0138] The three-dimensional data table is expanded into a two-dimensional matrix, where rows represent different county areas in different years and columns represent different indicators.
[0139] The positive indicators are standardized using the formula.
[0140] The negative indicators are standardized using the formula.
[0141] Vectorized operations are performed using numpy or pandas libraries to improve computational efficiency.
[0142] The normalized data is normalized using the formula to obtain the normalized value.
[0143] Vectorized operations are also performed using numpy or pandas libraries.
[0144] The entropy of each indicator is calculated using the formula.
[0145] The redundancy of each indicator is calculated.
[0146] Mathematical operations are performed using numpy libraries, including logarithmic operations and summation operations.
[0147] The weight of each indicator is calculated using the formula.
[0148] Vectorized operations are performed using numpy libraries to calculate the weight of each indicator.
[0149] The comprehensive evaluation index of each county is calculated using the formula.
[0150] The comprehensive evaluation index is sorted or visualized by county and time for further analysis and comparison.
[0151] Data processing and result output are performed using pandas libraries, and visualization can be performed using matplotlib or seaborn libraries.
[0152] Through the implementation of the above technical means, each step in the above technical solution can be completed, thereby obtaining the comprehensive evaluation index of each county at a specific time.
[0153] The technical effects of the above technical solutions are: through data preprocessing, standardization and normalization and other steps, the accuracy and consistency of the input data are ensured, and the distortion of the evaluation result caused by data errors or inconsistencies is avoided. Using the efficient data processing libraries such as pandas and numpy of Python, the data is quickly cleaned, converted and calculated, which greatly improves the efficiency of data processing. Through the calculation of entropy and redundancy, and the weight distribution based on redundancy, the objectivity and importance of each index in the evaluation process are fully considered, and the influence of subjective factors on the evaluation result is avoided. The method of comprehensive weighted evaluation is adopted, the normalized values of each index are combined with the weights, and the comprehensive evaluation index of each county in different years is obtained. By constructing a three-dimensional time series data table and an initial global evaluation matrix, the index change of each county in different years can be clearly seen, and the advantages and disadvantages of county development can be accurately positioned. Based on the comprehensive evaluation index, targeted development suggestions and measures can be provided for each county. Using visualization tools such as matplotlib or seaborn, the comprehensive evaluation index can be visualized. By constructing a perfect county evaluation index system and data analysis model, the county governance system can be continuously improved and optimized. With the help of data analysis results, the accuracy and scientificity of county governance can be improved, and the efficiency and level of county governance can be improved.
[0154] In one embodiment of the present application, the S4 comprises:
[0155] Real-time acquisition of data variation information of the third category of indicators, calculation of third category indicator update values according to the data variation information of the third category of indicators;
[0156] The calculation formula of the third category indicator update value is:
[0157]
[0158] Wherein, is the third category indicator update value of each second category, P is the number of third category indicators of the second category, is the data variation amount of the a-th third category indicator of the second category, is the initial data amount of the a-th third category indicator of the second category.
[0159] The third category indicator update value is compared with a preset three-type update threshold to obtain a three-level comparison result;
[0160] According to the three-level comparison result, the second category of indicators is updated to obtain data variation information of the second category of indicators, and second category indicator update values are calculated according to the data variation information of the second category of indicators;
[0161] The calculation formula of the second category index update value is:
[0162]
[0163] wherein, is the second category index update value of each first category, r is the number of second category indexes of the first category, is the data variation of the e-th second category index of the first category, is the initial data amount of the e-th second category index of the first category.
[0164] The second category index update value is compared with a preset second category update threshold to obtain a second level comparison result;
[0165] The index of the first category is updated according to the second level comparison result, and then an index update database is obtained;
[0166] A comprehensive evaluation update index of the index update database is obtained.
[0167] The working principle of the above technical solution is that: the data variation information of the third category index is collected in real time through a specific data source (such as a sensor, a database, an API, etc.). These indexes may involve business operation, production process, environmental monitoring, etc. According to the obtained data variation information, the update value of the third category index is calculated. The calculated third category index update value is compared with the preset third category update threshold. These thresholds are usually set by ordinary technical personnel according to business logic and actual needs, and are used to judge the amplitude and importance of index variation. According to the third level comparison result, the second category index is updated. If the variation of the third category index exceeds a certain threshold, it may trigger the adjustment or recalculation of the second category index. Similarly, according to the data variation information of the updated second category index, its update value is calculated. The second category index update value is compared with the preset second category update threshold to further judge the amplitude and importance of its variation. According to the second level comparison result, the index of the first category is updated. The updated first, second and third category indexes are stored in the index update database. The data in the index update database is used to calculate the comprehensive evaluation update index. This index reflects the health status, efficiency or performance of the overall business or system.
[0168] The technical effects of the above technical solutions are: by acquiring and updating the index data in real time, the system can quickly reflect the latest state of the business or system, improve the accuracy and timeliness of decision-making. By classifying the indicators into different categories and updating and comparing them in turn, the system can provide strong support for decision-making at different levels. The system design is flexible and can adjust the categories, thresholds and update algorithms of the indicators according to actual needs. At the same time, the system is easy to extend and can accommodate more indicators and data sources. By calculating the comprehensive evaluation update index, the system can provide data support for decision-making and reduce the risk of subjective speculation and blind decision-making.
[0169] The system can automatically complete the acquisition, updating and comparison process of the indicators, reducing manual intervention and improving work efficiency.
[0170] In an embodiment of the present application, the system comprises:
[0171] The category dimension construction module is used to acquire preset category information and preset kind information, and then acquire first category information, second category information and third category information, construct data evaluation dimensions, update and match the category information, and acquire dimension update information.
[0172] The data acquisition and classification module is used to acquire preset analysis demand data through a big data method, obtain analysis acquisition data, perform data division and classification on the analysis acquisition data through the data evaluation dimensions, and obtain a data analysis library.
[0173] The comprehensive evaluation analysis module is used to determine the weight data of each indicator of the third category through a global entropy value method, and perform weighted summation on each indicator using a comprehensive evaluation method to obtain a comprehensive evaluation index.
[0174] The index update evaluation module is used to acquire data variation information of the indicators of the third category and the second category, calculate the third category indicator update value and the second category indicator update value, determine whether to update the indicators of the second category and the first category, and acquire an updated comprehensive evaluation update index.
[0175] The working principle of the above technical solution is as follows: in the initial stage, the system obtains preset category information and type information. Based on these preset information, the system further refines and constructs first category information, second category information and third category information, which represent different data levels or analysis angles. The system constructs a data evaluation dimension, and also continuously updates and matches the preset category information to ensure the timeliness and accuracy of the preset categories of the data evaluation dimension. The preset analysis requirement data is collected by a big data method. These data may come from various data sources such as databases, log files, sensors, etc. The collected data is classified by the previously constructed data evaluation dimension to form a structured data analysis library. This step ensures the orderliness and analyzability of the data. The system uses a global entropy value method to determine the weight data of each index in the third category. The global entropy value method is a method based on information entropy, which can objectively reflect the information amount and importance of the index data. After determining the weight, the system uses a comprehensive evaluation method to weight and sum each index to calculate a comprehensive evaluation index. This index reflects the overall performance or state of the third category data. The system continuously monitors the data variation information of the third category and the second category index, in order to timely discover the abnormality or trend change in the data. Based on this variation information, the system calculates the third category index update value and the second category index update value, and judges whether the second category and the first category index need to be updated. If it needs to be updated, the system recalculates the comprehensive evaluation update index to reflect the latest data state and evaluation result.
[0176] The technical effects of the above technical solution are as follows: by constructing the data evaluation dimension and using the global entropy value method to determine the weight, the system can more accurately reflect the actual state and value of the data, avoiding the interference of human factors. The system can continuously monitor the data variation information and update the index and the comprehensive evaluation according to the need, ensuring the timeliness and accuracy of the evaluation. By constructing different levels of category information and data evaluation dimensions, the system can support multi-level and multi-angle data analysis, meeting the needs of different users and business scenarios. Using the big data method and automatic process for data collection, classification and evaluation calculation greatly improves the efficiency and automation degree of data processing, and reduces the labor cost. The comprehensive evaluation index and update index provided by the system can provide scientific basis and data support for decision-making, which can improve the scientificity and accuracy of decision-making.
[0177] In an embodiment of the present application, the category dimension construction module comprises:
[0178] The preset category division module is configured to obtain preset first category information.
[0179] The first category information is classified according to the preset second type information to obtain second category information.
[0180] Classify the second category information according to the preset third category information, and obtain third category information;
[0181] The first category information includes second category information, and the second category information includes third category information;
[0182] Combine the first category information, the second category information and the third category information to obtain data evaluation dimensions;
[0183] A preset category updating module is configured to obtain category updating information, match the category updating information with the first category information, the second category information and the third category information respectively, and obtain a matching category level corresponding to the category updating information; the matching category level includes the first category information, the second category information and the third category information; when the category updating information matches the first category information, the matching category level corresponding to the category updating information is the category level of the first category information; the matching is performed by calculating information similarity, comparing the calculation result with a preset matching threshold, matching the category matching information according to the comparison result, and similarity calculation and threshold setting are settings that can be made by those skilled in the art according to the prior art.
[0184] Update the matching category level by the category updating information to obtain category updating information corresponding to the matching category level;
[0185] Update the lower category information of the matching category level by the preset category updating to obtain category updating information of the lower category information; for example, after updating the first category information, the second and third category information are updated adaptively.
[0186] After the data is updated, combine the first category information, the second category information and the third category information to obtain dimension updating information.
[0187] The working principle of the above technical solution is as follows: the first type of information is obtained, which constitutes the basic framework of data analysis. The first type of information is classified according to the second type of information, forming the second type of information. This step is a further refinement and organization of the first type of information. The second type of information is classified according to the third type of information, obtaining the third type of information. At this point, the data has been divided into three levels, each level representing a different data perspective or analysis dimension. The first type of information, the second type of information and the third type of information are combined to form the data evaluation dimension. Obtain category update information, which comes from external data sources or internal data updates. Match the first type of information, the second type of information and the third type of information with the category update information to determine the category level to which the update information belongs. The matching process involves information similarity calculation, and by comparing the calculation result with the preset matching threshold, the system can determine which category level the update information is most matched with. Once the matching is successful, the system will update the matching category level with the preset category, and at the same time, in order to ensure the consistency and integrity of the data, the system will also adaptively update the lower category information of the matching category level. For example, if the first type of information is updated, the system will adjust the second and third type of information accordingly. After completing the category update, the system re-obtains the first type of information, the second type of information and the third type of information, and combines them to obtain the latest dimension update information. These information reflect the latest state of the data evaluation dimension.
[0188] The technical effect of the above technical solution is: by dividing the data into different levels of category information, the system can more flexibly organize and manage the data, and at the same time, this hierarchical structure also makes the system easy to extend and adapt to new data requirements. By matching the category update information with the preset category level, the update information can be accurately applied to the correct data level. At the same time, the adaptive update of the lower category information ensures the consistency and integrity of the data, and improves the efficiency of data update. The constructed data evaluation dimension provides a basis for multi-level and multi-angle data analysis. The system can analyze and evaluate the data from different dimensions and perspectives according to user needs or business scenarios. By introducing intelligent methods such as information similarity calculation and matching threshold setting, the system can automatically complete the matching and processing of category update information, reducing the degree of manual intervention and improving the intelligent level of data processing.
[0189] In an embodiment of the present application, the data collection and classification module comprises:
[0190] The data collection module is used to obtain preset analysis requirement information, collect the preset analysis requirement information through big data method, and obtain analysis collection data.
[0191] Preprocess the analysis collection data to obtain preprocessed analysis collection data;
[0192] A collection data classification module is configured to classify the analysis collection data according to first, second, and third category information of the data evaluation dimension, to obtain first, second, and third category data after classification;
[0193] An analysis library establishment module is configured to combine the first, second, and third category data after classification to obtain a data analysis library.
[0194] The working principle of the above technical solution is as follows: the starting point of the entire process involves determining the specific content or target that needs to be analyzed. The preset analysis requirement information comes from business requirements, market analysis, data research, and other aspects. According to the preset analysis requirement, relevant data is collected from various data sources (such as databases, log files, social media, and Internet of Things devices) using big data methods (such as distributed storage, parallel computing, and data mining). The purpose of this step is to collect as much information as possible related to the analysis requirement. The data collection method used is an existing data collection method in the prior art. The raw data collected often has problems such as noise, missing values, and outliers, so preprocessing is needed. The preprocessing steps may include data cleaning (removing invalid or erroneous data), data transformation (such as normalization, standardization), data integration (merging data from multiple data sources), and the like, to obtain clean, consistent, and usable analysis collection data. The preprocessed data is classified according to first, second, and third category information of the data evaluation dimension. These category information may be defined based on the nature, source, importance, or other relevant characteristics of the data. The purpose of classification is to organize the data into a format that is easier to analyze and understand. The classified data (first, second, and third category data) is combined to form a comprehensive data analysis library.
[0195] The technical effects of the above technical solution are: through the big data collection and preprocessing method, the quality of the data used for analysis can be ensured, and the analysis error caused by data problems can be reduced. Data classification and combination make the data more orderly and easy to access, thereby improving the efficiency of data analysis. Classification according to different data evaluation dimensions enables the data analysis library to support diversified analysis needs. Whether it is trend analysis, correlation analysis or prediction analysis, appropriate data sets can be found from the classified data. A high-quality, structured data analysis library provides a solid foundation for decision-making. The analysis results based on these data can more accurately and reliably guide business decisions. By establishing a data analysis library, enterprises can better manage and utilize their data assets and improve data governance. This can ensure data compliance, security and accessibility, laying a foundation for long-term development of the enterprise.
[0196] In one embodiment of the present application, the comprehensive evaluation analysis module comprises:
[0197] A standardization processing module is configured to establish a three-dimensional time series data table according to the data analysis library and form an initial global evaluation matrix.
[0198] An extreme value method is used to standardize each index, and the standardized data is distributed between 0.1 and 0.9. The specific standardization formula includes positive and negative indicators.
[0199] An index data processing module is configured to normalize each index to obtain index normalization.
[0200] The entropy and redundancy of each index are calculated.
[0201] The weight of each index is determined.
[0202] A comprehensive evaluation calculation module is configured to perform comprehensive weighting to obtain a comprehensive evaluation index.
[0203] The working principle of the above technical solution is:
[0204] Check the completeness, accuracy and consistency of the data, and handle missing values, outliers, etc.
[0205] Use the pandas library of Python or other data processing tools to construct a three-dimensional data table containing time, county and indicators according to the collected data.
[0206] Expand the three-dimensional data table into a two-dimensional matrix, where the rows represent the index values of different counties in different years, and the columns represent different indicators.
[0207] Use the formula to standardize the positive indicators.
[0208] Use the formula to standardize the negative indicators.
[0209] Vectorize operations using numpy or pandas libraries for improved computational efficiency.
[0210] Normalize the standardized data using the formula to obtain normalized values.
[0211] Similarly, vectorize operations using numpy or pandas libraries.
[0212] Calculate the entropy value of each indicator using the formula.
[0213] Calculate the redundancy of each indicator.
[0214] Perform mathematical operations using numpy libraries, including logarithmic operations and summation operations.
[0215] Calculate the weight of each indicator using the formula.
[0216] Vectorize operations using numpy libraries to calculate the weight of each indicator.
[0217] Calculate the comprehensive evaluation index of each county using the formula.
[0218] Sort or visualize the comprehensive evaluation index by county and time for further analysis and comparison.
[0219] Perform data processing and result output using pandas libraries, and use matplotlib or seaborn libraries for visualization.
[0220] Through the implementation of the above technical means, each step in the above technical solution can be completed, thereby obtaining the comprehensive evaluation index of each county at a specific time.
[0221] The technical effects of the above technical solutions are: through data preprocessing, standardization and normalization and other steps, the accuracy and consistency of the input data are ensured, and the distortion of the evaluation result caused by data errors or inconsistencies is avoided. Using efficient data processing libraries such as pandas and numpy of Python, the data is quickly cleaned, converted and calculated, greatly improving the efficiency of data processing. Through the calculation of entropy and redundancy, and the weight distribution based on redundancy, the objectivity and importance of each index in the evaluation process are fully considered, and the influence of subjective factors on the evaluation result is avoided. The method of comprehensive weighted evaluation is adopted, the normalized value of each index is combined with the weight, and the comprehensive evaluation index of each county in different years is obtained. By constructing a three-dimensional time series data table and an initial global evaluation matrix, the index change of each county in different years can be clearly seen, and the advantages and disadvantages of county development can be accurately positioned. Based on the comprehensive evaluation index, targeted development suggestions and measures can be provided for each county. Using visualization tools such as matplotlib or seaborn, the comprehensive evaluation index can be visualized. By constructing a perfect county evaluation index system and data analysis model, the county governance system can be continuously improved and optimized. With the help of data analysis results, the accuracy and scientificity of county governance can be improved, and the efficiency and level of county governance can be improved.
[0222] In an embodiment of the present application, the index updating and evaluation module comprises:
[0223] The third type of updating analysis module is used to obtain the data variation information of the third type of index in real time, and calculate the third type of index updating value according to the data variation information of the third type of index.
[0224] The calculation formula of the third type of index updating value is:
[0225]
[0226] Wherein, is the third type of index updating value of each second type, and P is the number of third type of indexes of the second type, is the data variation of the a-th third type of index of the second type, is the initial data amount of the a-th third type of index of the second type.
[0227] The third type of index updating value is compared with the preset third type of updating threshold value to obtain a three-level comparison result.
[0228] The second-type updating analysis module is configured to update the second-type indicators according to the third-level comparison result, obtain data variation information of the second-type indicators, and calculate second-type indicator updating values according to the data variation information of the second-type indicators;
[0229] The calculation formula of the second-type indicator updating value is:
[0230]
[0231] wherein, is each second-type indicator updating value of the first type, r is the number of second-type indicators of the first type, is the data variation amount of the e-th second-type indicator of the first type, is the initial data amount of the e-th second-type indicator of the first type.
[0232] The second-type indicator updating value is compared with a preset second-type updating threshold to obtain a second-level comparison result;
[0233] The first-type indicators are updated according to the second-level comparison result, and then an indicator updating database is obtained;
[0234] The comprehensive updating evaluation module is configured to obtain a comprehensive evaluation updating index of the indicator updating database.
[0235] All thresholds of the present application are set by ordinary technical personnel in the art according to historical knowledge and actual conditions.
[0236] The working principle of the above technical solution is as follows: the data variation information of the third-type indicators is collected in real time through specific data sources (such as sensors, databases, APIs, etc.). These indicators may involve business operation, production process, environmental monitoring, etc. According to the obtained data variation information, the updating values of the third-type indicators are calculated. The calculated third-type indicator updating values are compared with preset third-type updating thresholds. These thresholds are usually set by ordinary technical personnel according to business logic and actual needs, and are used to judge the amplitude and importance of the indicator variation. According to the third-level comparison result, the second-type indicators are updated. If the variation of the third-type indicators exceeds a certain threshold, it may trigger the adjustment or recalculation of the second-type indicators. Similarly, according to the data variation information of the updated second-type indicators, their updating values are calculated. The second-type indicator updating values are compared with preset second-type updating thresholds to further judge the amplitude and importance of their variation. According to the second-level comparison result, the first-type indicators are updated. The updated first-type, second-type and third-type indicators are stored in the indicator updating database. The comprehensive evaluation updating index is calculated using the data in the indicator updating database. This index reflects the health status, efficiency or performance of the overall business or system.
[0237] The technical effects of the above technical solutions are: by acquiring and updating the index data in real time, the system can quickly reflect the latest state of the business or system, improve the accuracy and timeliness of decision-making. By classifying the indicators into different categories and updating and comparing them in turn, the system can provide strong support for decision-making at different levels. The system design is flexible and can adjust the categories, thresholds and update algorithms of the indicators according to actual needs. At the same time, the system is easy to extend and can accommodate more indicators and data sources. By calculating the comprehensive evaluation update index, the system can provide data support for decision-making, reducing the risk of subjective speculation and blind decision-making.
[0238] The system can automatically complete the acquisition, updating and comparison process of the indicators, reducing manual intervention and improving work efficiency.
[0239] Obviously, those skilled in the art can make various modifications and changes to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and changes of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and changes.
Claims
1. A comprehensive evaluation data analysis method based on big data, characterized in that: The method comprises: S1. Obtain preset category information and preset type information, and then obtain first category information, second category information, and third category information, construct a data evaluation dimension, update and match the category information, and obtain dimension update information; Wherein, the S1 includes: Obtaining preset first category information; Classify the first category information according to preset second category information to obtain second category information; Classify the second category information according to the preset third category information to obtain third category information; The first category information includes the second category information, and the second category information includes the third category information; Combining the first category information, the second category information, and the third category information to obtain data evaluation dimensions; Acquire category update information, and perform information matching on the category update information with the first category information, the second category information, and the third category information, respectively, to obtain a matching category level corresponding to the category update information; Performing a preset category update on the matching category level using the category update information to obtain category update information of the corresponding matching category level; performing a preset category update on the lower category information of the matching category level to obtain category update information of the lower category information; After the data is updated, the first category information, the second category information, and the third category information are obtained and combined to obtain dimension update information; S2. Collecting preset analysis requirement data using a big data method to obtain analysis and collection data, and classifying the analysis and collection data using data evaluation dimensions to obtain a structured data analysis library; S3. Determine the weight data of each indicator of the third category by using the global entropy method, and use the comprehensive evaluation method to perform weighted summation on each indicator to obtain a comprehensive evaluation index; Among them, the weight data of each indicator in the third category is determined by the global entropy method, and the comprehensive evaluation method is used to perform weighted summation of each indicator to obtain a comprehensive evaluation index, including: The global entropy method and comprehensive evaluation method were used to analyze multiple evaluation indicators of multiple counties in multiple years to obtain a comprehensive indicator index; S4. Obtain data change information of the third category and the second category indicators, calculate the updated value of the third category indicator and the updated value of the second category indicator, determine whether to update the second category and the first category indicators, and obtain the updated comprehensive evaluation update index; Wherein, the S4 includes: Acquire data change information of the third category indicator in real time, and calculate the updated value of the third category indicator according to the data change information of the third category indicator; Comparing the updated value of the third category indicator with the preset three categories of update thresholds to obtain a three-level comparison result; updating the second category indicator according to the three-level comparison result, obtaining data change information of the second category indicator, and calculating the second category indicator update value according to the data change information of the second category indicator; Comparing the updated value of the second category indicator with a preset second category update threshold to obtain a secondary comparison result; updating the first category of indicators according to the secondary comparison results, thereby obtaining an indicator update database; Obtaining a comprehensive evaluation update index of the indicator update database.
2. The comprehensive evaluation data analysis method based on big data according to claim 1, characterized in that: The S2 includes: Acquire preset analysis requirement information, collect the preset analysis requirement information through big data methods, and obtain analysis and collection data; Preprocessing the analysis and collection data to obtain preprocessed analysis and collection data; Classify the analyzed collected data according to the first category information, the second category information, and the third category information of the data evaluation dimension to obtain classified first category data, second category data, and third category data; The classified first category data, second category data and third category data are combined to obtain a structured data analysis library.
3. The comprehensive evaluation data analysis method based on big data according to claim 1 is characterized in that: The S3 includes: Establish a three-dimensional time series data table based on the data analysis library and form an initial global evaluation matrix; The extreme value method is used to standardize each indicator, and the standardized data is distributed between 0.1-0.
9. The specific standardization formula includes positive indicators and negative indicators; Normalize each indicator to obtain indicator normalization; Calculate the entropy and redundancy of each indicator; Determine the weight of each indicator; Perform comprehensive weighting to obtain a comprehensive evaluation index.
4. A comprehensive evaluation data analysis system based on big data, characterized in that: The system comprises: A category dimension construction module is used to obtain preset category information and preset type information, and then obtain first category information, second category information, and third category information, construct a data evaluation dimension, update and match category information, and obtain dimension update information; The category dimension building module includes: A preset category classification module is used to obtain preset first category information; Classify the first category information according to preset second category information to obtain second category information; Classify the second category information according to the preset third category information to obtain third category information; The first category information includes the second category information, and the second category information includes the third category information; Combining the first category information, the second category information, and the third category information to obtain data evaluation dimensions; A preset category update module is used to obtain category update information, match the category update information with the first category information, the second category information, and the third category information, and obtain a matching category level corresponding to the category update information; Performing a preset category update on the matching category level using the category update information to obtain category update information of the corresponding matching category level; performing a preset category update on the lower category information of the matching category level to obtain category update information of the lower category information; After the data is updated, the first category information, the second category information, and the third category information are obtained and combined to obtain dimension update information; The data collection and classification module is used to collect the preset analysis requirement data through big data methods to obtain the analysis and collection data, and to divide and classify the analysis and collection data according to the data evaluation dimension to obtain a structured data analysis library; The comprehensive evaluation analysis module is used to determine the weight data of each indicator of the third category through the global entropy method, and use the comprehensive evaluation method to perform weighted summation on each indicator to obtain a comprehensive evaluation index; Among them, the weight data of each indicator in the third category is determined by the global entropy method, and the comprehensive evaluation method is used to perform weighted summation of each indicator to obtain a comprehensive evaluation index, including: The global entropy method and comprehensive evaluation method were used to analyze multiple evaluation indicators of multiple counties in multiple years to obtain a comprehensive indicator index; An indicator update evaluation module is used to obtain data change information of indicators of the third category and the second category, calculate the updated value of the third category indicator and the updated value of the second category indicator, determine whether to update the indicators of the second category and the first category, and obtain the updated comprehensive evaluation update index; Wherein, the indicator update evaluation module includes: A third-category update analysis module, configured to obtain data change information of the third-category indicator in real time, and calculate an updated value of the third-category indicator based on the data change information of the third-category indicator; Comparing the updated value of the third category indicator with the preset three categories of update thresholds to obtain a three-level comparison result; a second-category update analysis module, configured to update the second-category indicator according to the third-level comparison result, obtain data change information of the second-category indicator, and calculate the second-category indicator update value according to the data change information of the second-category indicator; Comparing the updated value of the second category indicator with a preset second category update threshold to obtain a secondary comparison result; updating the first category of indicators according to the secondary comparison results, thereby obtaining an indicator update database; The comprehensive update evaluation module is used to obtain the comprehensive evaluation update index of the indicator update database.
5. The comprehensive evaluation data analysis system based on big data according to claim 4 is characterized in that: The data collection and classification module includes: A data collection module is used to obtain preset analysis requirement information, collect the preset analysis requirement information through a big data method, and obtain analysis and collection data; Preprocessing the analysis and collection data to obtain preprocessed analysis and collection data; a collected data classification module, configured to classify the analyzed collected data according to the first category information, the second category information, and the third category information of the data evaluation dimension, and obtain classified first category data, second category data, and third category data; The analysis library establishment module is used to combine the classified first category data, second category data and third category data to obtain a structured data analysis library.
6. The comprehensive evaluation data analysis system based on big data according to claim 4 is characterized in that: The comprehensive evaluation and analysis module includes: The standardization processing module is used to establish a three-dimensional time series data table based on the data analysis library and form an initial global evaluation matrix; The extreme value method is used to standardize each indicator, and the standardized data is distributed between 0.1-0.
9. The specific standardization formula includes positive indicators and negative indicators; The indicator data processing module is used to normalize various indicators and obtain indicator normalization; Calculate the entropy and redundancy of each indicator; Determine the weight of each indicator; The comprehensive evaluation calculation module is used to perform comprehensive weighting to obtain a comprehensive evaluation index.
Citation Information
Patent Citations
Power customer classification method and device
CN113111924A
The method and kit of the selection of Molecule-Binding Nucleic Acids and the identification of the targets, and their use
KR1020180041331A