A simulation data evaluation system based on database

By designing a database-based simulation data evaluation system, including data preprocessing, feature analysis, clustering division and historical evaluation modules, the existing system is solved to solve the problem that existing systems are difficult to cope with complex data changes and evaluation delays, and efficient data management and accurate evaluation are achieved.

CN119761227BActive Publication Date: 2025-05-09FUJIAN KEDE ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510275236.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-09
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing simulation data evaluation systems are difficult to cope with complex simulation data changes and cannot adjust the evaluation strategy in real time. The evaluation process is delayed due to the large amount of data and complex types, which affects the timeliness of decision-making.

Method used

Design a database-based simulation data evaluation system, including data collection, preprocessing, characteristic analysis, clustering division, data analysis, historical analysis, evaluation standards, data evaluation and data storage modules. By analyzing the simulation data set, the distribution characteristics of the input parameters, the change trend of the output results, and the influence of environmental information are determined, and normalized to obtain the characteristic data set. Then, clustering division and feature extraction of data fragments are performed based on the characteristic data set, matching degree is calculated based on historical evaluation data, and evaluation scope and standards are dynamically adjusted.

Benefits of technology

It significantly improves the interpretability and consistency of data, improves data management and retrieval efficiency, can capture potential patterns and laws in the data, achieve more accurate evaluation and decision support, and adapt to complex and changeable simulation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761227B_ABST
    Figure CN119761227B_ABST
Patent Text Reader

Abstract

The present invention provides a simulation data evaluation system based on a database, and relates to the technical field of data processing. The method comprises: collecting simulation data from a simulation environment, and preprocessing the data to obtain a simulation data set, analyzing the simulation data set to obtain a characteristic data set, clustering the simulation data set according to the characteristic data set, and adding a clustering identifier to obtain a data segment set, identifying the simulation type of the data segment according to the clustering identifier, extracting characteristic data of the data segment to obtain a characteristic data set, and determining a historical evaluation rule to obtain an evaluation standard range, calculating the score values ​​of data segments of different simulation types according to the characteristic data set, and screening the data segments according to the evaluation standard range to obtain an evaluation result. The present invention can evaluate simulation data based on a database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of data processing, and in particular to a simulation data evaluation system based on a database. Background Art

[0002] In the prior art, the simulation data evaluation process mainly relies on preset rules to evaluate the results. Although this method can meet the needs in some cases, due to the diversity of data, it may cause the simulation data evaluation system to be unable to cope with complex simulation data changes and unable to adjust the evaluation strategy in real time.

[0003] The simulation data evaluation system may cause delays in the evaluation process due to the large amount of data and complex data types. For example, in mechanical design simulation, when the amount of data increases or there are multiple simulation scenarios, the database's efficient data indexing and retrieval functions can be used to locate and extract data, but the applicability and efficiency of the preset rules will be greatly reduced. The simulation data evaluation system may be difficult to adjust, resulting in delayed evaluation results, which in turn affects the timeliness of decision-making. Summary of the invention

[0004] The purpose of the present invention is to provide a simulation data evaluation system based on a database, aiming to solve the problems mentioned in the background technology.

[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0006] A simulation data evaluation system based on a database, the system comprising:

[0007] A data collection module is used to collect simulation data from the simulation environment, where the simulation data includes input parameters and output results of multiple simulation environments;

[0008] The data processing module is used to pre-process the simulation data, remove redundant data, and convert it into a unified format to obtain a simulation data set;

[0009] The characteristic analysis module is used to analyze the simulation data set to determine the distribution characteristics of the input parameters, the change trend of the output results and the influence of the environmental information, and perform normalization processing on them to obtain the characteristic data set;

[0010] A clustering module is used to cluster the simulation data set according to the characteristic data set, and add a cluster identifier to it to obtain a data segment set;

[0011] A data analysis module, used to identify the simulation type of the data segment according to the cluster identifier, and extract feature data of the data segment according to the simulation type to obtain a feature data set;

[0012] A historical analysis module is used to extract historical evaluation data sets of the same simulation type from the database according to the simulation type, and calculate the matching degree between the data segment set and the historical evaluation data set according to the feature data set to determine the historical evaluation rules;

[0013] An evaluation standard module is used to extract the historical evaluation range according to the historical evaluation rules, and adjust it through the feature data set to obtain the evaluation standard range;

[0014] The data evaluation module is used to calculate the score values ​​of data segments of different simulation types according to the characteristic data set, and screen them according to the evaluation standard range to obtain the evaluation results;

[0015] The data storage module is used to store the data segments that meet the evaluation criteria in the database according to the evaluation results.

[0016] Furthermore, by analyzing the simulation data set, the distribution characteristics of the input parameters, the change trend of the output results and the influence of the environmental information are determined, and normalized, a characteristic data set is obtained, including:

[0017] Extract input parameters according to the simulation data set, and determine the distribution characteristics of the input parameters by calculating the mean, variance and square deviation of the input parameters;

[0018] By analyzing the trend line fitting of the output results, the growth trend, decline trend and periodic fluctuation of the output results can be identified, and the changing trend of the output results can be determined;

[0019] By analyzing the correlation between the input parameters and their corresponding output results, the degree of mapping between the input parameters and the output results is determined;

[0020] According to the mapping degree, the interference of the output results due to the environmental information is analyzed to obtain the sensitivity of the input parameters;

[0021] According to the sensitivity, calculate the impact of environmental information on the simulation process;

[0022] The characteristic data set is obtained by normalizing the distribution characteristics, change trends and influence and unifying the format.

[0023] Furthermore, the calculation formula of influence is:

[0024] ,

[0025] in, is the influence of environmental information on the simulation process, is the index of the input parameter, is the total number of input parameters, For the input parameters, For the The sensitivity of the input parameters to the environmental information, For the The degree of mapping between input parameters and their corresponding output results, is the environmental disturbance function,

[0026] , are the input parameters and output results respectively. is the mean of the input parameters, is the variance of the input parameter,

[0027] is an exponential function, , is the coefficient.

[0028] Furthermore, according to the characteristic data set, the simulation data set is clustered and a cluster identifier is added thereto to obtain a data segment set, including;

[0029] According to the characteristic data set, the simulation data set is divided into clusters and randomly select The simulation data are used as the initial cluster centers;

[0030] The initial cluster is formed by calculating the distance from each simulation data to the initial cluster center and assigning it to the nearest initial cluster center;

[0031] By calculating the average value of all simulation data in each initial cluster, it is updated as the cluster center;

[0032] Repeat the steps of randomly selecting simulation data as the initial cluster center and updating the cluster center. When the change range of the cluster center is less than the preset change value, the cluster formed by the simulation data is the final cluster.

[0033] Further, according to the cluster identifier, the simulation type of the data segment is identified, including:

[0034] By extracting data segments with cluster identifiers and classifying them according to the cluster identifiers, multiple data subsets are formed;

[0035] By extracting the characteristic vectors of all data segments in the data subset, a characteristic vector set of the data subset is obtained;

[0036] According to the feature vector set, the feature mean and feature range of each dimension are calculated one by one to obtain the feature dimension set;

[0037] Compare the characteristic dimension set of the data subset with the statistical characteristic results of all simulation types in turn, analyze the matching degree between the characteristic dimension set and the statistical characteristic results of each simulation type, and obtain the comparison result;

[0038] According to the comparison results, the data subset is classified into the simulation type with the highest matching degree, and the simulation type of the data segment is determined.

[0039] Furthermore, feature data of the data segments are extracted according to the simulation type to obtain feature data sets, including:

[0040] According to the characteristic dimension set, the correlation between different characteristic dimensions of the data fragments is analyzed to obtain the dimensional relationship:

[0041] According to the dimension relationship, the completely linearly correlated dimensions are eliminated to obtain the significant dimensions;

[0042] Number the significant dimensions according to the dimension order to obtain a list of characteristic dimensions;

[0043] The feature dimension list and feature dimension set of each simulation type are merged to obtain a feature data set.

[0044] Furthermore, the matching degree between the data segment set and the historical evaluation data set is calculated based on the feature data set, and the historical evaluation rules are determined, including:

[0045] According to the feature data set, the initial feature values ​​of different feature dimensions are extracted and normalized to eliminate the dimension difference and obtain the feature value;

[0046] According to the eigenvalues, the variance of the eigenvalues ​​of different characteristic dimensions is calculated to obtain the classification ability value;

[0047] According to the classification capability value, determine the information increment provided by different characteristic dimensions;

[0048] According to the information increment, the contribution of different characteristic dimensions to the data fragment is determined to obtain the dimension weight;

[0049] According to the dimension weights, the matching degree between the data segment set and the historical evaluation data set is calculated.

[0050] Furthermore, the calculation formula of matching degree is:

[0051] ,

[0052] in, For the The matching degree between the data segments of the simulation type and the historical evaluation data set, is the index of the simulation type, is the index of the feature dimension, is the total number of feature dimensions, For the The first simulation type The weight of feature dimensions, The first The first simulation type The standard deviation of the feature dimension, The first The first simulation type The feature value of feature dimension, The historical evaluation dataset The first simulation type The feature value of feature dimension, is the coefficient,

[0053] is an exponential function, .

[0054] Furthermore, according to the historical evaluation rules, the historical evaluation range is extracted and adjusted through the feature data set to obtain the evaluation standard range, including:

[0055] According to the historical evaluation rules, the means and standard deviations of different characteristic dimensions are extracted and grouped according to the simulation type to obtain the historical evaluation range;

[0056] According to the feature data set, determine the mean and standard deviation of different feature dimensions in the data segment, and obtain the center position and distribution range;

[0057] Analyze the differences between the data segments and historical evaluation data based on the center position and distribution range, identify the shift in the mean and the change in the standard deviation, and obtain the difference information;

[0058] The historical evaluation range is adjusted based on the difference information. When the difference information is greater than the preset difference value, the historical evaluation range is adjusted proportionally. When the difference information is not greater than the preset difference value, the deviation amplitude of the standard deviation is calculated and the historical evaluation range is adjusted accordingly.

[0059] Furthermore, the calculation formula of the score value is:

[0060] ,

[0061] in, For the The score of the data segment of the simulation type, is the index of the simulation type, is the index of the feature dimension, is the total number of feature dimensions, For the The first simulation type The weight of feature dimensions, For the The first simulation type The feature value of feature dimension, For the The average eigenvalue of the simulation type,

[0062] is an exponential function, , is the coefficient.

[0063] The above solution of the present invention includes at least the following beneficial effects:

[0064] The present invention can significantly improve the interpretability and consistency of data by analyzing and normalizing the distribution characteristics, output trends and environmental impact of simulation data sets. The core of this function is to convert multi-dimensional and complex simulation data into a unified characteristic data set, making the subsequent clustering and analysis processes more efficient and accurate. Normalization not only reduces the scale differences between different variables, but also effectively reduces the errors caused by environmental noise. This data processing method is particularly important for dynamic simulation environments, enabling the system to adapt to complex and changeable simulation scenarios, while providing high-quality input data for the functional implementation of subsequent modules, thereby ensuring the reliability and stability of the overall evaluation system.

[0065] The present invention can decompose complex simulation data into multiple data segment sets with similar characteristics by performing cluster analysis on the characteristic data set. This function classifies the data to make the data structure clearer. By adding a clustering identifier to each data segment, the system can quickly locate and classify simulation data, thereby improving data management and retrieval efficiency. It can also capture potential patterns and regularities in the data, providing data support for further simulation type identification and feature extraction. This efficient classification capability is particularly significant when processing large-scale, multi-dimensional simulation data, which helps the system achieve more accurate evaluation.

[0066] The present invention can identify the simulation type of data and extract feature data sets by extracting and analyzing the characteristic vectors of data segments. Through the calculation and comparative analysis of the characteristic vectors, the system can accurately classify the simulation types of data segments, providing a reliable basis for subsequent historical analysis. The characteristic vector analysis process can filter out irrelevant or redundant information, thereby optimizing data storage and processing efficiency. By comparing the characteristic ranges of data subsets, new or unknown simulation types can also be identified, expanding the adaptability and application scope of the system. This function is particularly suitable for simulation tasks that require rapid response to changing scenarios.

[0067] The present invention can accurately match historical evaluation rules by utilizing historical evaluation data in the database in combination with the feature data set of the current data segment. Through the effective use of historical data, the system can quickly generate evaluation rules for current data without additional human intervention. The weight calculation and matching degree evaluation in the matching process can ensure the applicability and scientificity of historical rules. By dynamically adjusting the matching strategy, it can cope with the challenges of data changes or new types of simulation tasks, thereby greatly improving the accuracy and efficiency of the evaluation. This analysis method based on the combination of historical experience and real-time data significantly improves the decision-making ability of the system.

[0068] The present invention can ensure the applicability and consistency of the system to different data segments by adjusting the evaluation scope based on historical evaluation rules and feature data sets. It solves the problem of evaluation error caused by differences in simulation data characteristics by dynamically adjusting the evaluation scope. It can identify the offset of the evaluation scope through difference analysis and make adjustments according to actual conditions to ensure the scientificity and rationality of the evaluation criteria. It can automatically adapt to the data characteristics of different simulation types and provide a unified basis for the scoring and screening of data segments. This flexibility and adaptability enable the evaluation system to be more widely used in different fields and scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a flowchart of a simulation data evaluation system based on a database provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0070] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0071] like Figure 1 As shown, an embodiment of the present invention provides a simulation data evaluation system based on a database, the system comprising:

[0072] A data collection module is used to collect simulation data from the simulation environment, where the simulation data includes input parameters and output results of multiple simulation environments;

[0073] The data processing module is used to pre-process the simulation data, remove redundant data, and convert it into a unified format to obtain a simulation data set;

[0074] The characteristic analysis module is used to analyze the simulation data set to determine the distribution characteristics of the input parameters, the change trend of the output results and the influence of the environmental information, and perform normalization processing on them to obtain the characteristic data set;

[0075] A clustering module is used to cluster the simulation data set according to the characteristic data set, and add a cluster identifier to it to obtain a data segment set;

[0076] A data analysis module, used to identify the simulation type of the data segment according to the cluster identifier, and extract feature data of the data segment according to the simulation type to obtain a feature data set;

[0077] A historical analysis module is used to extract historical evaluation data sets of the same simulation type from the database according to the simulation type, and calculate the matching degree between the data segment set and the historical evaluation data set according to the feature data set to determine the historical evaluation rules;

[0078] An evaluation standard module is used to extract the historical evaluation range according to the historical evaluation rules, and adjust it through the feature data set to obtain the evaluation standard range;

[0079] The data evaluation module is used to calculate the score values ​​of data segments of different simulation types according to the characteristic data set, and screen them according to the evaluation standard range to obtain the evaluation results;

[0080] The data storage module is used to store the data segments that meet the evaluation criteria in the database according to the evaluation results.

[0081] In an embodiment of the present invention, simulation data is collected from a simulation environment, and the simulation data includes input parameters and output results of multiple simulation environments, providing a reliable data basis for subsequent modules; the simulation data is preprocessed to remove redundant data and converted into a unified format to obtain a simulation data set, thereby ensuring data quality during the analysis process and improving the consistency and accuracy of the overall workflow; the simulation data set is analyzed to determine the distribution characteristics of the input parameters, the changing trend of the output results, and the influence of the environmental information, and normalize them to obtain a characteristic data set, thereby improving the ability to parse complex simulation data, thereby effectively responding to the diversity and changing trends of simulation data.

[0082] According to the characteristic data set, the simulation data set is clustered and divided, and cluster identifiers are added to it to obtain a data fragment set, which optimizes the data storage and management methods, reduces the processing redundancy of data fragments, and provides more accurate basic data for subsequent analysis and evaluation; according to the cluster identifier, the simulation type of the data fragment is identified, and the characteristic data of the data fragment is extracted according to the simulation type to obtain a characteristic data set, avoiding the misevaluation phenomenon caused by type confusion in traditional methods, thereby improving the reliability of the evaluation results; according to the simulation type, the historical evaluation data set of the same simulation type is extracted from the database, and the matching degree between the data fragment set and the historical evaluation data set is calculated based on the characteristic data set, and the historical evaluation rules are determined. According to the accumulation of historical data, the evaluation process is continuously optimized to improve the accuracy and intelligence of the evaluation.

[0083] According to the historical evaluation rules, the historical evaluation range is extracted and adjusted through the feature data set to obtain the evaluation standard range, ensuring a high degree of match between the evaluation results and the actual situation, and improving the system's adaptability to dynamic data environments; according to the feature data set, the scoring values ​​of data fragments of different simulation types are calculated, and they are screened according to the evaluation standard range to obtain evaluation results, providing accurate quantitative evaluation for simulation data and improving evaluation efficiency and decision-making support capabilities; according to the evaluation results, the data fragments that meet the evaluation standards are stored in the database, which can support rapid retrieval of data and provide an important reference basis for evaluation, improving the scalability of the system and the level of data management.

[0084] The simulation data is collected from the simulation environment, and the simulation data includes input parameters and output results of multiple simulation environments, specifically including:

[0085] Design a dedicated data interface for each simulation environment, including a real-time data stream interface and a file import interface, and ensure that the interface supports multiple data transmission protocols, such as HTTP, Web Socket, MQTT, etc. to be compatible with multiple types of simulation environments; initialize the data interface, connect to the target simulation environment, periodically sample the simulation data, collect input parameters and corresponding output results according to the preset frequency, support incremental collection function, only extract new or changed data, and reduce the transmission and storage burden of duplicate data; perform real-time verification of data format, integrity and accuracy during the collection stage, mark abnormal data for subsequent processing, and attach timestamps and environment identifiers to the collected raw data to ensure that the timing and source of the data are clear.

[0086] Among them, the simulation data is preprocessed to remove redundant data and converted into a unified format to obtain a simulation data set, which specifically includes:

[0087] Apply hash algorithm or eigenvalue analysis method to detect and remove duplicate entries in simulation data. For nearly duplicate data, use the set similarity threshold to filter and retain representative data. For abnormal data marked in the acquisition phase, perform parameter range test. For slightly abnormal data, correct it through interpolation or averaging method, and directly eliminate seriously abnormal data. Convert the encoding format of data, for example, from XML to JSON, unify the unit system for numerical data, ensure that all parameters use the same dimension, and store the processed data as structured data for subsequent use.

[0088] According to the evaluation results, the data segments that meet the evaluation criteria are stored in the database, including:

[0089] Create a relational database and define the table structure to contain relevant information of the data fragments, such as simulation type, feature vector, evaluation score, etc. Set the partition storage strategy according to usage requirements, such as classified storage by simulation type, time or evaluation result; receive data fragments that meet the evaluation criteria from the data evaluation module, assign a unique identifier to each data, and write it to the corresponding table or partition in the database; set version control for historical data sets, record the time point and operation content of each data update, and when the update involves the adjustment of the evaluation criteria, retain a copy of the original data for reference to avoid data loss or overwriting; create indexes for stored data, optimize data retrieval efficiency, and provide multi-dimensional retrieval functions, such as by time range, simulation type or score value, to facilitate users to quickly locate the required data.

[0090] In a preferred embodiment of the present invention, the distribution characteristics of the input parameters, the change trend of the output results and the influence of the environmental information are determined by analyzing the simulation data set, and normalized, thereby obtaining a characteristic data set, including:

[0091] Extract input parameters according to the simulation data set, and determine the distribution characteristics of the input parameters by calculating the mean, variance and square deviation of the input parameters;

[0092] By analyzing the trend line fitting of the output results, the growth trend, decline trend and periodic fluctuation of the output results can be identified, and the changing trend of the output results can be determined;

[0093] By analyzing the correlation between the input parameters and their corresponding output results, the degree of mapping between the input parameters and the output results is determined;

[0094] According to the mapping degree, the interference of the output results due to the environmental information is analyzed to obtain the sensitivity of the input parameters;

[0095] According to the sensitivity, calculate the impact of environmental information on the simulation process;

[0096] The characteristic data set is obtained by normalizing the distribution characteristics, change trends and influence and unifying the format.

[0097] In an embodiment of the present invention, input parameters are extracted according to a simulation data set, and the distribution characteristics of the input parameters are determined by calculating the mean, variance and square difference of the input parameters, thereby reducing the randomness of the data and the interference of outliers on the results; by analyzing the trend line fitting of the output results, the growth trend, recession trend and periodic fluctuation of the output results are identified, the change trend of the output results is determined, and the change law of the output results in the simulation process is revealed; by analyzing the correlation between the input parameters and their corresponding output results, the mapping degree between the input parameters and the output results is determined, and it is determined which parameters contribute more to the output results; according to the mapping degree, the interference of the output results due to environmental information is analyzed, the sensitivity of the input parameters is obtained, and the interference of the external environment on the simulation accuracy is identified; by normalizing the distribution characteristics, change trends and influence degrees, and unifying the format, a characteristic data set is obtained, ensuring that each characteristic data has the same influence weight in subsequent data analysis and modeling, thereby improving the accuracy and consistency of the analysis results.

[0098] Among them, by analyzing the trend line fitting of the output results, identifying the growth trend, recession trend and cyclical fluctuation of the output results, and determining the change trend of the output results, specifically including:

[0099] Analyze the time series data of the output results in the simulation data set, and use the fitting algorithm to generate a trend line. The fitting algorithm can use linear regression, polynomial fitting or moving average method to identify the growth trend, recession trend and periodic fluctuation of the output results based on the fitted trend line. For example, perform polynomial fitting on the temperature change data of the output results to obtain its growth or recession law in different time periods.

[0100] The mapping degree between the input parameters and the output results is determined by analyzing the correlation between the input parameters and the corresponding output results, which specifically includes:

[0101] The correlation coefficient between the input parameters and the corresponding output results is calculated through correlation analysis. The correlation coefficient can be the Pearson correlation coefficient, which reflects the linear or nonlinear relationship between the input parameters and the output results. For example, by analyzing the correlation between temperature and output, it is found that the temperature change is proportional to the output, and the proportional value is the correlation coefficient.

[0102] Among them, according to the mapping degree, the interference of the output result due to the environmental information is analyzed to obtain the sensitivity of the input parameters, which specifically includes:

[0103] For each environmental condition, record the corresponding input parameters and output results, and determine the extent of change in the output results under different environments. Use statistical analysis methods to measure the interference of environmental information on the output results. For example, analyze the impact of temperature changes on yield under different humidity levels, and evaluate whether changes in humidity significantly affect the output results. Based on the degree of mapping, analyze the sensitivity of input parameters in different environments by taking partial derivatives. For example, if the input parameters are temperature and humidity, their impact on the output results under different environmental conditions can be evaluated by taking the partial derivatives of temperature and humidity.

[0104] In a preferred embodiment of the present invention, the calculation formula of the influence degree is:

[0105] ,

[0106] in, is the influence of environmental information on the simulation process, is the index of the input parameter, is the total number of input parameters, For the input parameters, For the The sensitivity of the input parameters to the environmental information, For the The degree of mapping between input parameters and their corresponding output results, is the environmental disturbance function,

[0107] , are the input parameters and output results respectively. is the mean of the input parameters, is the variance of the input parameter,

[0108] is an exponential function, , is the coefficient.

[0109] In an embodiment of the present invention, by integrating data of multiple dimensions such as input parameters, sensitivity, weights and environmental disturbance functions, the comprehensive impact of the environment on the simulation process can be effectively evaluated. The calculation formula not only improves the accuracy of the simulation evaluation, but also provides a reliable input basis for subsequent modules. The data in the formula are flexible and diverse, and can adapt to different types of simulation scenarios.

[0110] in, For the The sensitivity of an input parameter to environmental information. The sensitivity quantifies the impact of environmental disturbance on a certain input parameter. The greater the sensitivity, the more significant the impact of the change in the input parameter on the environmental adaptability of the system output. The sensitivity of each input parameter is obtained by calculating the partial derivative of the output result to the environmental disturbance.

[0111] In a preferred embodiment of the present invention, the simulation data set is clustered and divided according to the characteristic data set, and a cluster identifier is added thereto to obtain a data segment set, including:

[0112] According to the characteristic data set, the simulation data set is divided into clusters and randomly select The simulation data are used as the initial cluster centers;

[0113] The initial cluster is formed by calculating the distance from each simulation data to the initial cluster center and assigning it to the nearest initial cluster center;

[0114] By calculating the average value of all simulation data in each initial cluster, it is updated as the cluster center;

[0115] Repeat the steps of randomly selecting simulation data as the initial cluster center and updating the cluster center. When the change range of the cluster center is less than the preset change value, the cluster formed by the simulation data is the final cluster.

[0116] In the embodiment of the present invention, the simulation data set is divided into clusters and randomly select The simulation data is used as the initial cluster center to provide an initial reference point for the subsequent clustering process, ensuring the diversity and flexibility of the clustering process; the distance from each simulation data to the initial cluster center is calculated and assigned to the nearest initial cluster center to form an initial cluster, and the accuracy and effectiveness of clustering are improved by the distance measurement method; the average value of all simulation data in each initial cluster is calculated and updated as the cluster center, and the clustering process is gradually converged by updating the cluster center; the steps of randomly selecting simulation data as the initial cluster center and updating the cluster center are repeated. When the change range of the cluster center is less than the preset change value, the cluster formed by the simulation data is the final cluster. The convergence judgment mechanism effectively avoids infinite loops in the clustering process, ensures that the clustering operation can be completed within a limited time, and the clustering result has high stability.

[0117] Among them, according to the characteristic data set, the simulation data set is divided into clusters and randomly select The simulation data is used as the initial cluster center, including:

[0118] The simulation data set is divided according to the characteristics of the characteristic data set. By using the K-means algorithm, the simulation data set is preliminarily grouped according to the attributes of the characteristic data set. clusters and randomly selected from the simulation dataset data points as the initial cluster centers.

[0119] In a preferred embodiment of the present invention, identifying the simulation type of the data segment according to the cluster identifier includes:

[0120] By extracting data segments with cluster identifiers and classifying them according to the cluster identifiers, multiple data subsets are formed;

[0121] By extracting the characteristic vectors of all data segments in the data subset, a characteristic vector set of the data subset is obtained;

[0122] According to the feature vector set, the feature mean and feature range of each dimension are calculated one by one to obtain the feature dimension set;

[0123] Compare the characteristic dimension set of the data subset with the statistical characteristic results of all simulation types in turn, analyze the matching degree between the characteristic dimension set and the statistical characteristic results of each simulation type, and obtain the comparison result;

[0124] According to the comparison results, the data subset is classified into the simulation type with the highest matching degree, and the simulation type of the data segment is determined.

[0125] In an embodiment of the present invention, data segments with cluster identifiers are extracted and classified according to the cluster identifiers to form multiple data subsets, thereby effectively reducing the complexity of data analysis; the feature vector set of the data subset is obtained by extracting the feature vectors of all data segments in the data subset, thereby converting the original complex data into representative numerical data; according to the feature vector set, the feature mean and feature range of each dimension are calculated one by one to obtain a feature dimension set, thereby simplifying and summarizing the multi-dimensional characteristics of the data; the feature dimension set of the data subset is compared with the statistical feature results of all simulation types in turn, and the degree of matching between the data subset and the statistical feature results of each simulation type is analyzed to obtain a comparison result, thereby accurately evaluating the similarity between the data subset and each simulation type; according to the comparison result, the data subset is classified into the simulation type with the highest matching degree, the simulation type of the data segment is determined, and it is accurately assigned to the corresponding simulation type according to the characteristics, thereby realizing automatic classification of the simulation type.

[0126] In a preferred embodiment of the present invention, feature data of the data segment is extracted according to the simulation type to obtain a feature data set, including:

[0127] According to the characteristic dimension set, the correlation between different characteristic dimensions of the data fragments is analyzed to obtain the dimensional relationship:

[0128] According to the dimension relationship, the completely linearly correlated dimensions are eliminated to obtain the significant dimensions;

[0129] Number the significant dimensions according to the dimension order to obtain a list of characteristic dimensions;

[0130] The feature dimension list and feature dimension set of each simulation type are merged to obtain a feature data set.

[0131] In an embodiment of the present invention, according to the characteristic dimension set, the correlation between different characteristic dimensions of the data segment is analyzed to obtain the dimensional relationship, and the dimension pairs with strong correlation are screened out, so as to improve the effectiveness and accuracy of subsequent analysis: according to the dimensional relationship, the dimensions with complete linear correlation are eliminated to obtain the significant dimensions, and focusing on the significant dimensions can reduce the computational burden; the significant dimensions are numbered according to the dimensional order to obtain a characteristic dimension list, which is helpful to systematically manage the data and ensure the repeatability and transparency of the subsequent analysis; the characteristic dimension list and the characteristic dimension set of each simulation type are merged to obtain a feature data set, which ensures the integrity and diversity of the feature data.

[0132] Among them, according to the characteristic dimension set, the correlation between different characteristic dimensions of the data fragment is analyzed to obtain the dimension relationship, which specifically includes:

[0133] Each data segment is composed of multiple dimensions, which can be physical quantities or system parameters. The value of each dimension constitutes a vector, which represents the observed value of the data segment on that dimension. The Pearson correlation coefficient method is used to measure the linear relationship strength between two dimensions, and its value is between -1 and 1. For all characteristic dimensions, the Pearson correlation coefficients between them are calculated to obtain a correlation matrix. For example, if there are n characteristic dimensions, the correlation matrix is ​​an n×n symmetric matrix, in which each element represents the correlation value between different characteristic dimensions. Based on the calculated correlation matrix, analyze which characteristic dimensions have a higher correlation. Usually, a correlation threshold is preset. If the absolute value of the correlation coefficient is greater than the threshold, the two dimensions are highly correlated and have a significant linear correlation.

[0134] Among them, according to the dimension relationship, the completely linearly related dimensions are eliminated to obtain the significant dimensions, including:

[0135] According to the preset correlation threshold, those dimension pairs with correlation higher than the threshold can be screened out. These dimension pairs are defined as completely linearly correlated and provide duplicate information in data analysis, so they need to be eliminated. Once two or more highly correlated dimensions are identified, one of the dimensions can be eliminated, and the principal component analysis dimensionality reduction technique can be used to merge these redundant dimensions into a new dimension, retaining the maximum variance information, and map the original data to a new coordinate system through linear transformation, and select the new dimension with the largest variance. For example, first calculate the covariance matrix of the data, and then obtain the principal components through eigenvalue decomposition. The principal component with the largest eigenvalue is selected as the significant dimension, and the principal component with smaller variance is eliminated. After eliminating redundant dimensions, the correlation matrix in the characteristic dimension set can be recalculated to verify whether the correlation between dimensions has reached the ideal degree of separation. If there are still highly correlated dimensions, they can continue to be eliminated until the remaining dimensions are relatively independent.

[0136] In a preferred embodiment of the present invention, the matching degree between the data segment set and the historical evaluation data set is calculated based on the feature data set to determine the historical evaluation rule, including:

[0137] According to the feature data set, the initial feature values ​​of different feature dimensions are extracted and normalized to eliminate the dimension difference and obtain the feature value;

[0138] According to the eigenvalues, the variance of the eigenvalues ​​of different characteristic dimensions is calculated to obtain the classification ability value;

[0139] According to the classification capability value, determine the information increment provided by different characteristic dimensions;

[0140] According to the information increment, the contribution of different characteristic dimensions to the data fragment is determined to obtain the dimension weight;

[0141] According to the dimension weights, the matching degree between the data segment set and the historical evaluation data set is calculated.

[0142] In an embodiment of the present invention, based on a feature data set, initial eigenvalues ​​of different feature dimensions are extracted, and normalized to eliminate dimensional differences to obtain eigenvalues, which can effectively extract multidimensional feature information; based on the eigenvalues, the variance of the eigenvalues ​​of different feature dimensions is calculated to obtain classification capability values, which can determine the classification capability of each feature dimension; based on the classification capability values, the information increment provided by different feature dimensions is determined, and the contribution of each dimension to the decision-making process is quantified through the information increment; based on the information increment, the contribution of different feature dimensions to the data fragment is determined to obtain dimension weights, and through the determination of the weights, the relative importance of different feature dimensions is clarified; based on the dimension weights, the matching degree of the data fragment set and the historical evaluation data set is calculated, and by comprehensively considering the weights of the feature dimensions and the actual content of the data, the similarity between the data fragment and the historical evaluation data set can be accurately measured.

[0143] Among them, according to the classification capability value, the information increment provided by different characteristic dimensions is determined, including:

[0144] The classification ability value is usually expressed by measuring the variance or standard deviation of the feature. The larger the variance, the larger the range of variation of the feature and the greater its contribution to classification. The information increment measures the new information provided by a specific feature in the classification task. It is measured by using information theory indicators such as mutual information. For each feature dimension, the information increment of each feature dimension is calculated through mutual information. This can clarify the amount of additional information provided by each feature dimension in the classification process. Dimensions with larger information increments contribute more to the classification task.

[0145] Among them, according to the information increment, the contribution of different characteristic dimensions to the data fragment is determined, and the dimension weight is obtained, which specifically includes:

[0146] Contribution refers to the degree of influence of each feature dimension on the overall data segment, and is calculated by the ratio of its information increment to the total information increment. Dimension weight is an indicator that reflects the importance of each feature dimension in the calculation of data segment matching. Dimension weight is proportional to contribution, that is, the dimension with larger information increment has larger weight. Dimension weight is further optimized according to actual business needs. For example, higher weight can be given to certain feature dimensions according to the needs of specific tasks, while other dimensions are ignored. After multiple iterative calculations, dimension weight will be more stable and reflect the actual importance of each feature in the overall matching calculation.

[0147] In a preferred embodiment of the present invention, the calculation formula of the matching degree is:

[0148] ,

[0149] in, For the The matching degree between the data segments of the simulation type and the historical evaluation data set, is the index of the simulation type, is the index of the feature dimension, is the total number of feature dimensions, For the The first simulation type The weight of feature dimensions, The first The first simulation type The standard deviation of the feature dimension, The first The first simulation type The feature value of feature dimension, The historical evaluation dataset The first simulation type The feature value of feature dimension, is the coefficient,

[0150] is an exponential function, .

[0151] In an embodiment of the present invention, the difference between the feature dimension of the simulation type in the data segment and the feature dimension in the historical evaluation data set is quantified to evaluate the similarity between the two. The formula comprehensively evaluates the similarity between multi-dimensional features, avoids the influence of single-dimensional differences, and considers various information such as the mean, standard deviation and weight of the characteristic dimension, so that the calculation result more accurately reflects the true matching degree of the data segment. The exponential function term and the regularization term are combined to quantify the matching degree between the data segment and the historical evaluation data, fully considering the importance of the feature dimensions, the numerical differences, and the scale of the data, and is suitable for complex and diverse simulation scenarios.

[0152] In a preferred embodiment of the present invention, according to the historical evaluation rules, the historical evaluation range is extracted and adjusted by the feature data set to obtain the evaluation standard range, including:

[0153] According to the historical evaluation rules, the means and standard deviations of different characteristic dimensions are extracted and grouped according to the simulation type to obtain the historical evaluation range;

[0154] According to the feature data set, determine the mean and standard deviation of different feature dimensions in the data segment, and obtain the center position and distribution range;

[0155] Analyze the differences between the data segments and historical evaluation data based on the center position and distribution range, identify the shift in the mean and the change in the standard deviation, and obtain the difference information;

[0156] The historical evaluation range is adjusted based on the difference information. When the difference information is greater than the preset difference value, the historical evaluation range is adjusted proportionally. When the difference information is not greater than the preset difference value, the deviation amplitude of the standard deviation is calculated and the historical evaluation range is adjusted accordingly.

[0157] In an embodiment of the present invention, according to historical evaluation rules, the means and standard deviations of different characteristic dimensions are extracted, and they are grouped according to the simulation type to obtain the historical evaluation range, and preliminarily define the central trend and dispersion of the data; according to the feature data set, the means and standard deviations of different characteristic dimensions in the data segment are determined, the center position and distribution range are obtained, and the characteristic distribution of the current data is clarified; according to the center position and distribution range, the difference between the data segment and the historical evaluation data is analyzed, the offset of the mean and the change of the standard deviation are identified, the difference information is obtained, and the deviation point of the current data from the historical evaluation range is accurately located; according to the difference information, the historical evaluation range is adjusted, the robustness and practicality of the evaluation are improved, and the reliability of the results is ensured.

[0158] Among them, the historical evaluation range is adjusted according to the difference information. When the difference information is greater than the preset difference value, the historical evaluation range is adjusted according to the proportion; when the difference information is not greater than the preset difference value, the deviation range of the standard deviation is calculated and the historical evaluation range is adjusted accordingly, including:

[0159] When the difference information is greater than the preset difference value, the historical evaluation range is adjusted proportionally according to the difference amount. For dimensions with large mean shifts, the center value of the historical range is moved to a position closer to the current data. For dimensions with significant changes in standard deviation, the discreteness of the evaluation range is appropriately expanded or reduced.

[0160] When the difference information is not greater than the preset difference value, the center value is not adjusted, only the deviation amplitude of the standard deviation is calculated, and the historical evaluation range is fine-tuned according to the deviation amplitude, for example, by increasing or decreasing the upper and lower limits of the evaluation range.

[0161] The adjusted historical evaluation range is matched and verified with the feature data set again to ensure that the adjusted range can effectively include the current data. When the verification result shows that the adjustment is insufficient or excessive, the adjustment process is repeated until the evaluation range is stable.

[0162] In a preferred embodiment of the present invention, the calculation formula of the score value is:

[0163] ,

[0164] in, For the The score of the data segment of the simulation type, is the index of the simulation type, is the index of the feature dimension, is the total number of feature dimensions, For the The first simulation type The weight of feature dimensions, For the The first simulation type The feature value of feature dimension, For the The average eigenvalue of the simulation type,

[0165] is an exponential function, , is the coefficient.

[0166] In an embodiment of the present invention, by comprehensively considering the weight of the eigenvalue, the square of the eigenvalue, and the nonlinear change of the eigenvalue, the performance of the data segment is evaluated, which can effectively capture the importance and change characteristics of the feature dimension and improve the accuracy and applicability of the model score; the eigenvalue is logarithmically processed to slow down the growth rate of large eigenvalues, effectively reduce the impact of outliers or abnormal values, and improve the robustness of the model; the volatility of the eigenvalue is captured, and the deviation from the mean is amplified by an exponential function, so that the model can respond more sensitively to feature dimensions with a larger degree of discreteness, which is suitable for measuring the variability of features.

[0167] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A simulation data evaluation system based on a database, characterized in that: The system comprises: A data collection module is used to collect simulation data from the simulation environment, where the simulation data includes input parameters and output results of multiple simulation environments; The data processing module is used to pre-process the simulation data, remove redundant data, and convert it into a unified format to obtain a simulation data set; The characteristic analysis module is used to analyze the simulation data set to determine the distribution characteristics of the input parameters, the change trend of the output results and the influence of the environmental information, and perform normalization processing on them to obtain the characteristic data set; A clustering module is used to cluster the simulation data set according to the characteristic data set, and add a cluster identifier to it to obtain a data segment set; A data analysis module, used to identify the simulation type of the data segment according to the cluster identifier, and extract feature data of the data segment according to the simulation type to obtain a feature data set; A historical analysis module is used to extract historical evaluation data sets of the same simulation type from the database according to the simulation type, and calculate the matching degree between the data segment set and the historical evaluation data set according to the feature data set to determine the historical evaluation rules; An evaluation standard module is used to extract the historical evaluation range according to the historical evaluation rules, and adjust it through the feature data set to obtain the evaluation standard range; The data evaluation module is used to calculate the score values ​​of data segments of different simulation types according to the characteristic data set, and screen them according to the evaluation standard range to obtain the evaluation results; A data storage module, used for storing data segments that meet the evaluation criteria in a database according to the evaluation results; The calculation formula of the influence degree is: ; in, is the influence of environmental information on the simulation process, is the index of the input parameter, is the total number of input parameters, For the input parameters, For the The sensitivity of the input parameters to the environmental information, For the The degree of mapping between input parameters and their corresponding output results, is the environmental disturbance function, , are the input parameters and output results respectively. is the mean of the input parameters, is the variance of the input parameter, is an exponential function, , is the coefficient; The calculation formula of the matching degree is: ; in, For the The matching degree between the data segments of the simulation type and the historical evaluation data set, is the index of the simulation type, is the index of the feature dimension, is the total number of feature dimensions, For the The first simulation type The weight of feature dimensions, The first The first simulation type The standard deviation of the feature dimension, The first The first simulation type The feature value of feature dimension, The historical evaluation dataset The first simulation type The feature value of feature dimension, is the coefficient, is an exponential function, ; The calculation formula of the score value is: ; in, For the The score of the data segment of the simulation type, For the The average eigenvalue of the simulation type, is an exponential function, , is the coefficient.

2. A simulation data evaluation system based on a database according to claim 1, characterized in that: By analyzing the simulation data set, the distribution characteristics of the input parameters, the change trend of the output results and the influence of the environmental information are determined, and normalized, a characteristic data set is obtained, including: Extract input parameters according to the simulation data set, and determine the distribution characteristics of the input parameters by calculating the mean, variance and square deviation of the input parameters; By analyzing the trend line fitting of the output results, the growth trend, decline trend and periodic fluctuation of the output results can be identified, and the changing trend of the output results can be determined; By analyzing the correlation between the input parameters and their corresponding output results, the degree of mapping between the input parameters and the output results is determined; According to the mapping degree, the interference of the output results due to the environmental information is analyzed to obtain the sensitivity of the input parameters; According to the sensitivity, calculate the impact of environmental information on the simulation process; The characteristic data set is obtained by normalizing the distribution characteristics, change trends and influence and unifying the format.

3. A simulation data evaluation system based on a database according to claim 2, characterized in that: According to the characteristic data set, the simulation data set is clustered and divided, and a cluster identifier is added to it to obtain a data segment set, including; According to the characteristic data set, the simulation data set is divided into clusters and randomly select The simulation data are used as the initial cluster centers; The initial cluster is formed by calculating the distance from each simulation data to the initial cluster center and assigning it to the nearest initial cluster center; By calculating the average value of all simulation data in each initial cluster, it is updated as the cluster center; Repeat the steps of randomly selecting simulation data as the initial cluster center and updating the cluster center. When the change range of the cluster center is less than the preset change value, the cluster formed by the simulation data is the final cluster.

4. A database-based simulation data evaluation system according to claim 3, characterized in that: According to the cluster identification, the simulation type of the data segment is identified, including: By extracting data segments with cluster identifiers and classifying them according to the cluster identifiers, multiple data subsets are formed; By extracting the characteristic vectors of all data segments in the data subset, a characteristic vector set of the data subset is obtained; According to the feature vector set, the feature mean and feature range of each dimension are calculated one by one to obtain the feature dimension set; Compare the characteristic dimension set of the data subset with the statistical characteristic results of all simulation types in turn, analyze the matching degree between the characteristic dimension set and the statistical characteristic results of each simulation type, and obtain the comparison result; According to the comparison results, the data subset is classified into the simulation type with the highest matching degree, and the simulation type of the data segment is determined.

5. A database-based simulation data evaluation system according to claim 4, characterized in that: Extract feature data of the data segment according to the simulation type to obtain a feature data set, including: According to the characteristic dimension set, the correlation between different characteristic dimensions of the data fragments is analyzed to obtain the dimensional relationship: According to the dimension relationship, the completely linearly correlated dimensions are eliminated to obtain the significant dimensions; Number the significant dimensions according to the dimension order to obtain a list of characteristic dimensions; The feature dimension list and feature dimension set of each simulation type are merged to obtain a feature data set.

6. A database-based simulation data evaluation system according to claim 5, characterized in that: The matching degree between the data segment set and the historical evaluation data set is calculated based on the feature data set, and the historical evaluation rules are determined, including: According to the feature data set, the initial feature values ​​of different feature dimensions are extracted and normalized to eliminate the dimension difference and obtain the feature value; According to the eigenvalues, the variance of the eigenvalues ​​of different characteristic dimensions is calculated to obtain the classification ability value; According to the classification capability value, determine the information increment provided by different characteristic dimensions; According to the information increment, the contribution of different characteristic dimensions to the data fragment is determined to obtain the dimension weight; According to the dimension weights, the matching degree between the data segment set and the historical evaluation data set is calculated.

7. A database-based simulation data evaluation system according to claim 6, characterized in that: According to the historical evaluation rules, the historical evaluation range is extracted and adjusted through the feature data set to obtain the evaluation standard range, including: According to the historical evaluation rules, the means and standard deviations of different characteristic dimensions are extracted and grouped according to the simulation type to obtain the historical evaluation range; According to the feature data set, determine the mean and standard deviation of different feature dimensions in the data segment, and obtain the center position and distribution range; Analyze the differences between the data segments and historical evaluation data based on the center position and distribution range, identify the shift in the mean and the change in the standard deviation, and obtain the difference information; The historical evaluation range is adjusted based on the difference information. When the difference information is greater than the preset difference value, the historical evaluation range is adjusted proportionally. When the difference information is not greater than the preset difference value, the deviation amplitude of the standard deviation is calculated and the historical evaluation range is adjusted accordingly.

Citation Information

Patent Citations

  • Evaluation index weight data processing method based on clustering analysis

    CN117332287A

  • Statistical method and device based on test results of simulation model and simulation equipment

    CN118228018A