Analysis system, analysis program, and analysis method
The analytical system addresses the challenge of inaccurate data evaluation by identifying a fitted distribution model and calculating deviations, enabling precise data analysis and improved decision-making.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- S I LAB CO LTD
- Filing Date
- 2025-10-08
- Publication Date
- 2026-06-03
AI Technical Summary
Existing data analysis methods fail to accurately evaluate target data due to their inability to account for underlying data trends, leading to inaccurate decision-making and strategy development.
An analytical system comprising an acquisition unit, parameter estimation unit, goodness-of-fit calculation unit, distribution identification unit, and evaluation unit, which analyzes source data to identify a fitted distribution model and calculate deviations, enabling a comprehensive evaluation of the target data.
This system allows for accurate evaluation of target data by specifying a distribution model that fits the data distribution, grasping the underlying structure, and quantifying structural characteristics, thereby supporting better decision-making.
Smart Images

Figure 0007869608000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an analysis system, an analysis program, and an analysis method. [Background technology]
[0002] Conventionally, technologies for automatically classifying diverse data, such as business data and medical data, into multiple classification classes have been widely used. One example of such technology is proposed in Patent Document 1.
[0003] Patent Document 1 discloses the construction of a model that infers classification classes from target data. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2024-162867 [Overview of the Initiative] [Problems that the invention aims to solve]
[0005] However, the invention described in Patent Document 1 simply learns the trends of the target data using an inference model and classifies the target data, but it has the problem that it cannot take into account the data trends underlying the target data and therefore cannot accurately evaluate the target data.
[0006] In view of the above circumstances, the problem to be solved by the present invention is to provide a novel technology that enables more accurate evaluation of target data. [Means for solving the problem]
[0007] To solve the above problems, the present invention provides an analytical system for analyzing an object to be analyzed, The analysis system comprises an acquisition unit, a parameter estimation unit, a goodness-of-fit calculation unit, a distribution identification unit, a deviation calculation unit, and an evaluation unit. The acquisition unit acquires source data including an analysis index that serves as an indicator in the analysis of the target of analysis, and actual values of the target of analysis for the analysis index. The parameter estimation unit estimates parameters for each candidate distribution model that is a candidate distribution model for the data distribution of a plurality of targets of analysis based on the source data. The goodness-of-fit calculation unit calculates the distribution fit for each candidate distribution model whose parameters have been estimated, for the data distribution of the target of analysis. The distribution identification unit identifies a fitted distribution model that fits the data distribution of the target of analysis from among the plurality of candidate distribution models based on the distribution fit. The deviation calculation unit calculates a theoretical value for the analysis index of the target of analysis based on the parameters of the identified fitted distribution model, calculates the deviation between the theoretical value and the actual value of the analysis index of the target of analysis included in the source data. The evaluation unit evaluates the target of analysis based on the deviation.
[0008] Furthermore, in order to solve the above problems, an analysis program for analyzing an object to be analyzed is provided, wherein the analysis program functions as an acquisition unit, a parameter estimation unit, a goodness-of-fit calculation unit, a distribution identification unit, a deviation calculation unit, and an evaluation unit, the acquisition unit acquires source data for each object to be analyzed, including an analysis index that serves as an indicator in the analysis of the object to be analyzed, and the actual values of the object to be analyzed for the analysis index, the parameter estimation unit estimates the parameters for each candidate distribution model that is a candidate distribution model for the data distribution of a plurality of objects to be analyzed based on the source data, the goodness-of-fit calculation unit calculates the distribution goodness for each candidate distribution model whose parameters have been estimated for the data distribution of the object to be analyzed, the distribution identification unit identifies a fitted distribution model that fits the data distribution of the object to be analyzed from among the plurality of candidate distribution models based on the distribution goodness-of-fit, the deviation calculation unit calculates the theoretical value of the analysis index of the object to be analyzed based on the parameters of the identified fitted distribution model, calculates the deviation between the theoretical value and the value of the analysis index of the object to be analyzed included in the source data, and the evaluation unit evaluates the object to be analyzed based on the deviation.
[0009] Also, in order to solve the above problems, the present invention is an analysis method using an analysis system for analyzing an analysis target, wherein the analysis system acquires analysis source data including an analysis index serving as an index in the analysis of the analysis target and an actual value of the analysis target with respect to the analysis index; based on the analysis source data, performs an estimation process on parameters for each candidate distribution model that is a candidate for a distribution model of data distributions of a plurality of the analysis targets, and calculates a distribution fitness for the data distribution of the analysis target for each candidate distribution model for which the parameters have been estimated; based on the distribution fitness, performs a process of specifying a suitable distribution model that matches (fits) the data distribution of the analysis target among a plurality of candidate distribution models; based on the parameters of the specified suitable distribution model, performs a process of calculating a theoretical value in the analysis index of the analysis target; performs a process of calculating a deviation between the theoretical value and the actual value of the analysis index of the analysis target included in the analysis source data; and performs a process of evaluating the analysis target based on the deviation.
[0010] [[ID=*]] [[ID=*]]
[0011] [[ID=*]] With such a configuration, it is possible to specify a distribution model that matches (fits) the data distributions of a plurality of analysis targets using the indexes included in the analysis source data, and to grasp the structure underlying the analysis source data. Thereby, it is possible to more accurately evaluate the analysis target.
[0011] In a preferred form of the present invention, the evaluation unit calculates a deviation vector based on the deviation for each analysis target, calculates a relative position index indicating an overall tendency of the analysis target based on a plurality of the analysis indexes based on the deviation vector, calculates a structural distance index indicating a degree of divergence in which the distance between the actual value and the theoretical value is corrected by a correlation relationship between the plurality of analysis indexes based on the deviation vector, calculates a consistency index indicating a degree of variation of the actual values of the plurality of analysis indexes in the analysis target based on the deviation vector, and calculates an integrated evaluation index for evaluating the analysis target based on the relative position index, the structural distance index, and the consistency index.
[0012] With such a configuration, it becomes possible to comprehensively and quantitatively evaluate the structural characteristics of a certain analysis target.
[0013] In a preferred form of the present invention, the distribution specifying unit specifies one or more of the fitting distribution models based on the first distribution fitness, and when a plurality of the fitting distribution models are specified, the fitness calculation unit calculates the second distribution fitness based on the number of parameters for each of the plurality of specified fitting distribution models and the degree of fit to the data distribution, and the distribution specifying unit specifies one fitting distribution model based on the second distribution fitness.
[0014] With such a configuration, it is possible to specify the optimal distribution model that best fits the data distribution from among a plurality of distribution models.
[0015] In a preferred form of the present invention, the evaluation unit transmits an instruction to the generation system AI to specify the analysis indicators that have a significant impact on the calculation process of the integrated evaluation index for the analysis source data and the analysis target, and specifies the analysis indicators that have a significant impact on the calculation process of the integrated evaluation index among the plurality of analysis indicators included in the analysis source data.
[0016] With such a configuration, by specifying the analysis indicators that have a significant positive or negative impact on the score calculation among a plurality of analysis indicators, it is possible to enable efficient differentiation from other analysis targets existing in the same environment.
[0017] In a preferred form of the present invention, the analysis system further includes an analysis means specifying unit. The analysis means specifying unit specifies the analysis field of the analysis source data based on a plurality of the analysis indicators included in the acquired analysis source data, and determines whether the analysis source data is an analysis target by comparing the analysis field and the analysis target reference information.
[0018] In this invention, appropriate analytical fields are defined, and by adopting this configuration, it is possible to prevent obtaining inappropriate analytical results by not performing analysis on analytical fields that are outside the scope of analysis in this invention from the outset.
[0019] In a preferred embodiment of the present invention, the analysis system further comprises an environmental suitability index calculation unit, which processes to identify the degree of concentration of the analysis environment, which indicates the degree of weekly variation of the analysis environment corresponding to the identified suitability distribution model; processes to calculate the degree of change of the analysis environment, which indicates the degree of temporal change of the analysis environment composed of the plurality of analysis targets; and processes to calculate an environmental suitability index, which indicates the degree of suitability of the analysis target to the analysis environment, based on the degree of concentration of the analysis environment, the degree of change of the analysis environment, and the sales deviation.
[0020] In a preferred embodiment of the present invention, the environmental compatibility index calculation unit calculates the environmental compatibility index using a calculation formula that has weightings for each variable according to the identified compatibility distribution model, with the analysis environment concentration, the analysis environment change rate, and the sales deviation as variables.
[0021] This configuration allows us to quantify, for example, the likelihood of a company (the subject of analysis) surviving in the market (the analysis environment), thereby supporting users in deciding on future strategies. [Effects of the Invention]
[0022] This invention can provide a novel technique that enables more accurate evaluation of target data. [Brief explanation of the drawing]
[0023] [Figure 1] This is a block diagram showing the system configuration in a case of aggregated data. [Figure 2] A block diagram showing the hardware configuration of a system in one embodiment of the present invention. [Figure 3] This is an example of a processing flowchart in one embodiment of the present invention. [Figure 4] This is an example of a flowchart relating to the details of calculating the integrated evaluation index in one embodiment of the present invention. [Modes for carrying out the invention]
[0024] Further details will be provided below with reference to the attached screenshots. The drawings illustrate preferred embodiments. However, many different forms are possible and the embodiments are not limited to those described herein.
[0025] For example, in this embodiment, the configuration and operation of the analysis system will be described, but similar effects can be achieved by the method of execution, the apparatus, the computer program, etc. Furthermore, the program may be stored on a recording medium. Using this recording medium, for example, the program can be installed on a computer, thereby configuring an analysis device or analysis system. Here, the recording medium on which the program is stored may be a non-transient recording medium such as a CD-ROM.
[0026] <1. Overview of the Invention> This invention relates to a system for analyzing objects of analysis. Traditionally, scientific inquiry has long centered on "comparison" between groups, and its statistical basis has been rooted in the worldview of the normal distribution, which uses the mean and variance as criteria. While this paradigm has been extremely successful in understanding the majority and discovering general laws, individual peculiarities, treated as measurement errors and considered "outliers," tend to be overlooked as noise in the analysis. As a result, problems have emerged where distributional characteristics such as oligopolistic structures and long-tail structures underlying the data are overlooked, preventing accurate decision-making and strategy development.
[0027] In contrast, in today's world, where personalized medicine and respect for diversity have become indispensable in all aspects of society, these "outliers" are precisely the crucial sources of information that hold the key to breakthroughs. In the era of personalized medicine pioneered by genomics, the discovery of drugs that show dramatic effects in patients with specific gene mutations, even if they have no effect on the majority, demonstrates that what was once considered an "outlier" is actually a "super-responder" that holds the key to therapeutic breakthroughs. This trend is not limited to medicine, but extends to all fields such as business, education, and finance. A uniform approach to the majority has reached its limits, and "diversity" and "individualization" that respond to individual talents, risks, and needs are becoming the source of new value creation.
[0028] Against this backdrop, this new era demands a new role from data analysis. Specifically, a paradigm shift was needed, moving from analysis that compares and generalizes groups to analysis that understands structure and gives meaning to individuals.
[0029] To address this paradigm shift, this invention provides a new analytical method called Distributional Structure Analysis (DSA), which respects the inherent distributional structure of the data used for analysis (hereinafter referred to as the source data) and quantifies the relative position of each data point within that structure. This allows for the redefinition of conventional outliers as "structural singularities" that represent the essence of the distribution, and enables the elucidation of their unique meaning and value. Furthermore, it avoids uniform normalization, which can distort the data structure, and is extremely effective in accurately capturing the non-normality exhibited by much of the data in the real world (e.g., tail risk in financial markets, extreme variability in corporate sales size).
[0030] Specifically, the present invention performs a distributional analysis on each of the actual indicator values based on source data that includes indicators used for analysis (for example, data items, hereinafter referred to as analysis indicators) and the values of the analysis indicators possessed by each target of analysis (hereinafter referred to as actual indicator values), and then performs an analysis of the target of analysis. Here, in the present invention, all targets included in the source data of analysis are targets of analysis.
[0031] In this embodiment, the analysis of the target of analysis is performed using the following procedure: (1) Based on the actual values of the indicators included in the source data, identify a single distribution model (hereinafter referred to as the fitted distribution model) that fits the distribution of multiple targets for analysis within the source data. (2) The deviation between the theoretical value based on the identified fitted distribution model and the actual value of the indicator included in the source data is calculated for each indicator included in the source data. (3) Based on the deviations for each calculated indicator, a relative position indicator is calculated that shows the overall trend of the subject of analysis based on multiple analytical indicators, a structural distance indicator is calculated that shows how far it deviates from a state in which the actual values of the indicators and the theoretical values are in agreement (ideal structure), and a consistency indicator is calculated that shows how consistent the multiple indicators are. An integrated evaluation indicator is then calculated to evaluate the subject of analysis using the relative position indicator, structural distance indicator, and consistency indicator. (4) Based on the calculated integrated evaluation indicators, evaluation comments are generated for the subject of analysis.
[0032] In this embodiment, the target of analysis is a production entity such as a company, but it may also be a target of analysis in any field, such as medical device development, epidemiology and public health, healthcare management analysis, or clinical research.
[0033] <2. System Configuration> Figure 1 is a block diagram showing the configuration of an analysis system according to one embodiment. As shown in Figure 1, the analysis system 0 includes an analysis device 1. A general-purpose personal computer or the like can be used as the analysis device 1. In this embodiment, the analysis system 0 is configured with a client computer on which the analysis program is installed as the analysis device 1 (a so-called standalone type), but the client computer may be configured to communicate with a server via wired or wireless connection, and some or all of the functional components (means) of the analysis device 1 may be provided on the server (a so-called server-client type).
[0034] <2.1. Overview of Functional Configuration> As shown in Figure 1, the analysis device 1 comprises, functionally, an acquisition unit 101, an analysis means identification unit 102, a distribution candidate identification unit 103, a parameter estimation unit 104, a fitness calculation unit 105, a distribution identification unit 106, an evaluation unit 107, an environmental suitability index calculation unit 108, a display processing unit 109, and a database 2. This represents the concrete realization of information processing by software (stored in the storage unit 12 described later) by hardware (processing unit 11, etc.).
[0035] <2.2. Hardware Configuration> Figure 2 shows an example of the hardware configuration of the analysis device 1 that constitutes the analysis system of this embodiment. The analysis device 1 comprises a processing unit 11, a storage unit 12, a communication unit 13, an input unit 14, and an output unit 15 as its hardware configuration. The processing unit 11 includes one or more processors such as a CPU, and controls the entire operation process of the analysis device 1 by executing the analysis program, OS, and other applications according to the present invention. The memory unit 12 includes a volatile memory such as RAM capable of storing instruction sets, and a non-volatile storage device such as an HDD or SSD capable of storing the OS and analysis programs. The communication unit 13 has a communication interface device for connecting to a network, and performs communication control with the communication network NW to input and output information. The input unit 14 has an input device capable of input processing, such as a keyboard or a touch panel. The output unit 95 has a display device capable of display processing, such as a display.
[0036] <2.3. Details of the System Functional Configuration> Next, we will explain the processing details of each part of the above-mentioned functional configuration.
[0037] <2.3.1. Acquisition part 101> The acquisition unit 101 acquires the source data for analysis. The acquisition unit 101 acquires source data (e.g., spreadsheet files, CSV files, etc.) for each analysis target, which includes the same analysis indicator and contains the actual values for each analysis indicator.
[0038] The source data for analysis includes, for example, if the subject of analysis is a company, indicators such as productivity indicators showing the company's productivity (sales, number of employees, value added, total assets, number of orders, sales personnel, etc.), profitability indicators showing sales (operating profit margin, return on equity, gross profit margin, average customer spending, etc.), and customer indicators showing market acquisition status and customer attraction power (number of customers, customer composition ratio, number of repeat customers, etc.), along with their actual values.
[0039] Furthermore, if the subject of analysis is a medical device, for example, it would include indicators such as quality indicators that evaluate the safety, effectiveness, and reliability of the medical device (clinical performance data, product quality data, process quality data, etc.), profitability indicators that show how efficiently profits are being generated from the medical device (market share, gross profit margin, return on investment, R&D cost to sales ratio, etc.), and productivity indicators that measure the efficiency of the development and manufacturing processes (data related to the development and manufacturing processes (development cycle, equipment utilization rate, etc.), data related to the supply chain (raw material shortage rate, etc.)), as well as their actual values.
[0040] Furthermore, if the subject of analysis is an epidemic, for example, it may include indicators such as morbidity-related indicators (prevalence, mortality rate, etc.) that show how often the disease occurs in a particular region or population, risk indicators for evaluating the causes and related factors of the disease (smoking rate, hypertension prevalence, water pollution level, etc.), and medical resource indicators (number of doctors, hospitalization rate, bed occupancy rate, etc.) that show the healthcare delivery system and access status for each region or population, along with their actual values.
[0041] Furthermore, if the subject of analysis is a medical institution, for example, it will include indicators such as profitability indicators that show the stability and profitability of management (e.g., medical fee unit price, equity ratio, cost ratio by medical department), efficiency indicators to visualize the operational efficiency of medical institutions (e.g., bed occupancy rate, outpatient waiting time, number of consultations per doctor), and human resource indicators to visualize the adequacy and skill development status of medical personnel (e.g., number of doctors, turnover rate, training participation rate), along with their actual values.
[0042] Furthermore, if the subject of analysis is a clinical trial, for example, it will include indicators such as efficacy indicators showing effectiveness and safety (survival time, symptom improvement rate, mortality rate, etc.), patient background indicators that evaluate the interpretation of clinical trial results and external validity (age distribution, disease severity, smoking / drinking rates, etc.), and trial efficiency indicators that evaluate the efficiency, speed, and quality of the clinical trial (investigational drug supply lead time, screening failure rate, number of audit findings, etc.), as well as their actual values.
[0043] <2.3.2. Analysis means identification section 102> The analysis means identification unit 102 determines whether predetermined conditions are met and identifies an analysis means according to the determination result. The analysis means identification unit 102 determines whether the source data is suitable as data to be analyzed and identifies an analysis means based on the determination result. In this embodiment, the analysis means identification unit 102 determines whether the source data is suitable based on the number of data points (number of samples) included in the acquired source data and the data type of the indicator performance value.
[0044] <2.3.3. Distribution candidate identification unit 103> The distribution candidate identification unit 103 identifies distribution candidates. Based on one or more predetermined distribution models, the distribution candidate identification unit 103 identifies candidate distribution models that are suitable for multiple data distributions to be analyzed.
[0045] In this embodiment, various distribution models can be used, but the following are preferred. (1) Normal distribution (parameters: μ (mean), σ (standard deviation)):
number
number
number
number
number
number
number
number
number
number
[0046] <2.3.4. Parameter Estimation Unit 104> The parameter estimation unit 104 estimates the parameters of the candidate distribution model. Based on the actual values of each of the multiple analysis indicators included in the source data, the parameter estimation unit 104 estimates the parameters of the distribution model identified as a candidate distribution model for each analysis indicator.
[0047] <2.3.5. Suitability Calculation Unit 105> The goodness-of-fit calculation unit 105 calculates the distribution fit of the distribution model fitted to multiple analysis targets. Based on the source data, the goodness-of-fit calculation unit 105 calculates the distribution fit for each distribution model fitted to multiple analysis targets with respect to the values of the analysis indicators.
[0048] <2.3.6.Distribution identification unit 106> The distribution identification unit 106 identifies a distribution model that fits the distributions of multiple analysis targets. Based on the calculated distribution fit, the distribution identification unit 106 identifies a suitable distribution model from among the multiple distribution models that appropriately represents the distributions of the multiple analysis targets.
[0049] <2.3.7. Evaluation Section 107> The evaluation unit 107 evaluates the object of analysis based on the deviation between the theoretical value based on the parameters of the fitted distribution model and the actual value. The evaluation unit 107 includes a deviation calculation unit 1071, a relative position calculation unit 1072, a deviation degree calculation unit 1073, a consistency calculation unit 1074, an integrated evaluation calculation unit 1075, and a comment generation unit 1076.
[0050] <2.3.7.1. Deviation Calculation Unit 1071> The deviation value calculation unit 1071 calculates the theoretical values based on the distribution model for a plurality of analysis targets included in the analysis source data. The deviation value calculation unit 1071 calculates the theoretical values of each analysis target using the calculation formula corresponding to each specified candidate distribution model and the fitting distribution model.
[0051] Here, examples of the calculation formula for the theoretical value corresponding to each fitting distribution model are as follows. Note that the subscript i in the following calculation formula is a symbol indicating the i-th analysis index. (1) Normal distribution (parameters: μ, σ):
Equation
Equation
Equation
number
number
number
[0052] Furthermore, the deviation calculation unit 1071 calculates the deviation between the theoretical value of each indicator based on the fitted distribution model and the actual value of the object being analyzed. The deviation calculation unit 1071 calculates the deviation between the theoretical value of the analysis indicator based on the fitted distribution model and the value of the analysis indicator of the object being analyzed included in the source data (actual indicator value) for each fitted distribution model, and calculates a deviation vector.
[0053] <2.3.7.2. Relative Position Calculation Unit 1072> The relative position calculation unit 1072 calculates a relative position index that shows the overall trend of the subject being analyzed based on multiple analysis indicators. The relative position calculation unit 1072 calculates the relative position index based on the deviation vector.
[0054] <2.3.7.3. Deviation Calculation Unit 1073> The deviation calculation unit 1073 calculates a structural distance index that shows the structural deviation of each actual value of an indicator from its theoretical value, taking into account the correlation between the analytical indicators in the subject of analysis. Based on the deviation vector, the deviation calculation unit 1073 calculates a structural distance index that shows the deviation of correcting the distance between the actual value and the theoretical value due to the correlation between multiple analytical indicators.
[0055] <2.3.7.4. Consistency Calculation Unit 1074> The consistency calculation unit 1074 calculates a consistency index that shows the degree of variability in the actual values of multiple analytical indicators in the subject of analysis. The consistency calculation unit 1074 calculates the consistency index based on the deviation vector.
[0056] <2.3.7.5. Integrated Evaluation Calculation Unit 1075> The integrated evaluation calculation unit 1075 calculates integrated evaluation indicators. Based on relative position indicators, structural distance indicators, and consistency indicators, the integrated evaluation calculation unit 1075 calculates integrated evaluation indicators to evaluate the subject of analysis.
[0057] <2.3.7.6. Comment Generation Section 1076> The comment generation unit 1076 processes evaluation comments for the analysis target based on the integrated evaluation indicators. The comment generation unit 1076 sends instructions to the generation system AI to identify factors that significantly influenced the integrated evaluation indicators and generate evaluation comments based on the source data, the relative position indicator, structural distance indicator, and consistency indicator used to calculate the integrated evaluation indicators, and the type of fitted distribution model, causing the AI to identify these factors as influencing factors for the analysis target and generate and obtain evaluation comments that include these influencing factors.
[0058] <2.3.8.Environmental compatibility index calculation unit 108> The environmental suitability index calculation unit 108 calculates an environmental suitability index that indicates the degree of suitability of the analysis target to the analysis environment. The environmental suitability index calculation unit 108 identifies the degree of concentration that corresponds to the identified suitability distribution model of the analysis environment, calculates the degree of change in the analysis environment that indicates the degree of change in the analysis environment composed of multiple analysis targets, and calculates an environmental suitability index that indicates the degree of suitability of the analysis target to the analysis environment based on the analysis environment concentration, the degree of change in the analysis environment, and the deviation related to the analysis index.
[0059] Specifically, the environmental suitability index calculation unit 108 calculates the standard deviation and power exponent calculated to identify the optimal suitability distribution model as the concentration of the analysis environment. The environmental suitability index calculation unit 108 also compares the latest source data with source data from past points in time that include the same analysis targets as the latest source data, and calculates the degree of change in the analysis environment, which indicates how much the analysis environment composed of the multiple analysis targets has changed over time, based on the difference between the actual values of multiple indicators included in the source data from past points in time and the actual values of multiple indicators included in the latest source data.
[0060] The environmental suitability index calculation unit 108 then calculates an environmental suitability index that indicates the degree of suitability of the analysis target to the analysis environment, based on the concentration of the analysis environment, the degree of change in the analysis environment, and the relative position index. Specifically, among the analysis indicators, important analysis indicators are set, and the environmental suitability index calculation unit 108 uses the concentration of the analysis environment, the degree of change in the analysis environment, and the relative position index as variables, and calculates the environmental suitability index using a weighted sum of the weights of each variable according to the optimal suitability distribution model identified in the important analysis indicator.
[0061] <2.3.9. Display Processing Unit 109> The display processing unit 109 processes various screens operated by the user and displays the display processing results to the user terminal device (analysis device 1).
[0062] <3. Processing Flowchart> The analysis method using the analysis system of the present invention will be described below with reference to Figures 3 and 4. Figure 3 is a flowchart showing the process from acquiring the source data to identifying a fitted distribution model that fits the distribution of multiple analysis targets and evaluating the analysis targets according to the fitted distribution model.
[0063] <3.1. Obtaining the source data for analysis> First, in step S1 (hereinafter referred to as "step SX"), the acquisition unit 101 acquires the source data for analysis. In this embodiment, the acquisition unit 101 receives a specification from the user of the source data to be analyzed, performs data cleansing on the specified source data, and acquires the source data for analysis.
[0064] Specifically, data cleansing processes involve removing duplicate data, correcting invalid values, and standardizing formats. Additionally, machine learning-based anomaly detection can be used to automatically identify invalid data caused by human error or system errors.
[0065] The acquisition unit 101 may acquire the source data by receiving input of the source data from the user, or it may acquire the source data stored in database 2 or an external database.
[0066] <3.2. Assessment of suitability of the source data for analysis> In S2, the analysis means identification unit 102 determines whether the source data is suitable for analysis. In this embodiment, the analysis means identification unit 102 determines that the source data is not suitable for analysis if the number of data points included in the acquired source data is less than a predetermined number. Furthermore, the analysis means identification unit 102 determines that the source data is not suitable for analysis if the data type of the indicator performance values included in the acquired source data is a nominal scale (gender, blood type, place of origin, etc.).
[0067] Then, if it is determined that the source data is not suitable for analysis (NO in S2), in S3 the analysis means identification unit 102 sends instructions to the generation system AI to identify the source data and the alternative analysis means, causing the AI to identify an alternative analysis means to analyze the source data. For example, if the data type of the indicator performance value included in the source data is a category, the analysis means identification unit 102 will determine χ 2 The test is identified as an alternative analytical means. Then, in S14, the display processing unit 109 processes and displays the analysis results obtained by the alternative analytical means.
[0068] In a preferred embodiment of the present invention, the analysis means identification unit 102 identifies an analysis field based on a plurality of analysis indicators contained in the source data, and determines whether the source data is suitable for analysis based on the analysis field. Specifically, the analysis means identification unit 102 transmits instructions to the generation system AI to identify the source data and the analysis field of the source data, causes the AI to perform language analysis processing on the plurality of analysis indicators contained in the source data, and identifies the analysis field by combining the plurality of analysis indicators. The analysis means identification unit 102 then compares the analysis field with the analysis target reference information indicating the field to be analyzed, and determines whether the analysis field falls under a predetermined non-analysis target field (a field where the removal of outliers is a prerequisite; for example, an SPC for quality control), and determines that the source data is not suitable for analysis if the analysis field falls under a non-analysis target field.
[0069] <3.3. Identifying Candidate Distribution Models> On the other hand, if the source data is determined to be suitable for analysis (YES in S2), the distribution candidate identification unit 103 identifies a candidate distribution model in S3. In this embodiment, the distribution candidate identification unit 103 identifies the above-mentioned distribution model from well-known distribution models based on the number of data points (number of data to be analyzed) in the source data acquired in S1.
[0070] Specifically, the distribution candidate identification unit 103 refers to the number of data points (number of data to be analyzed) in the source data acquired in S1 and identifies as a candidate distribution model the distribution model in which the minimum number of data points required for fitting the distribution model is smaller than the number of data points.
[0071] <3.4. Parameter estimation process for each candidate distribution model> In S4, the parameter estimation unit 104 processes the parameters of the distribution model based on each candidate distribution model. In this embodiment, the parameter estimation unit 104 sends the source data acquired in S1 and instructions to the generative AI to estimate the parameters of each candidate distribution model to be applied to the distribution of the actual values of each analysis indicator (e.g., sales, symptom improvement rate, etc.) of multiple analysis targets (e.g., companies, clinical trials) included in the source data. Based on likelihood maximization (a method that maximizes the probability that the actual values of the source data will appear with the estimated candidate distribution parameters), the AI estimates and acquires the parameters of each candidate distribution model for each analysis indicator.
[0072] <3.5. Identifying a Fitted Distribution Model Based on First Goodness-of-Fit> In S5, the distribution identification unit 106 identifies a suitable distribution model based on the first goodness of fit. In this embodiment, first, the goodness of fit calculation unit 105 sends an instruction to the generation system AI to identify a candidate distribution model that exceeds a predetermined first goodness of fit threshold from the candidate distribution models whose parameters were obtained in S4. Using a first goodness of fit calculation means different for each candidate distribution model and the candidate distribution model for which parameters have been calculated, the first goodness of fit is calculated for each candidate distribution model of each analysis index.
[0073] Specifically, the goodness-of-fit calculation unit 105 uses theoretical values based on the parameters of each candidate distribution model obtained in S4 and empirical cumulative distribution functions based on the actual values of multiple indicators being analyzed.
number
[0074] More specifically, the goodness-of-fit calculation unit 105 uses, for example, the Kolmogorov-Smirnov test, the Anderson-Darling test, or the Shapiro-Wilk test as the first goodness-of-fit calculation means. For example, if the candidate distribution model is a power-law distribution, the Kolmogorov-Smirnov test is used as the first goodness-of-fit calculation means; if the candidate distribution model is a log-product distribution, the Anderson-Darling test is used as the first goodness-of-fit calculation means; and if the candidate distribution model is a normal distribution, the Shapiro-Wilk test or the like is used as the first goodness-of-fit calculation means.
[0075] The distribution identification unit 106 then identifies one or more candidate distribution models whose first goodness of fit exceeds the first goodness of fit threshold (for example, 0.05) based on the first goodness of fit of each candidate distribution model calculated and a predetermined first goodness of fit threshold (for example, 0.05) as the fitted distribution model.
[0076] The goodness-of-fit calculation unit 105 may also calculate the first distribution goodness-of-fit by comparing theoretical values based on the parameters of the candidate distribution model with actual values of multiple analysis indicators. Specifically, the goodness-of-fit calculation unit 105 compares the ranks of multiple analysis targets arranged in ascending order based on the actual values of multiple indicators with quantiles based on the candidate distribution model whose parameters have been estimated.
number
number
[0077] <3.6. Determination of whether a fitted distribution model can be identified based on the first distribution's goodness of fit> In S7, the distribution identification unit 106 determines whether a suitable distribution model was identified by the identification process in S6. If it is determined that a suitable distribution model cannot be identified (there are no candidate distribution models that exceed the first goodness-of-fit threshold) (NO in S7), then in S3, the analysis means identification unit 102 causes the generative AI to identify an alternative analysis means.
[0078] <3.7. Identifying the optimally fitted distribution model based on the second distribution's goodness of fit> On the other hand, if the distribution identification unit 106 determines in S7 that a suitable distribution model has been identified (YES in S7), then in S8 the distribution identification unit 106 determines whether there are multiple suitable distribution models identified in S6. If there is only one suitable distribution model (NO in S8), the process proceeds to S10.
[0079] On the other hand, if there are multiple fitted distribution models (YES in S8), in S9 the distribution identification unit 106 identifies a fitted distribution model based on the second goodness of fit. In this embodiment, first the goodness of fit calculation unit 105 sends the multiple fitted distribution models identified in S6, and an instruction to the generation system AI to identify one fitted distribution model from the multiple fitted distribution models, and causes the second goodness of fit calculation means to calculate the second goodness of fit based on the parameters of the multiple fitted distribution models.
[0080] Specifically, the goodness-of-fit calculation unit 105 calculates the second distribution goodness-of-fit by inputting the parameters of multiple fitted distribution models into a formula that evaluates the balance between the degree of model fit (maximum likelihood) and the number of parameters, using an information criterion (e.g., Akaike information criterion, Bayesian information criterion, etc.) as a second goodness-of-fit calculation means.
[0081] The distribution identification unit 106 then identifies one fitted distribution model from among multiple fitted distribution models for each analysis index based on the second distribution fit. Specifically, the distribution identification unit 106 identifies the fitted distribution model with the smallest information criterion for each of the multiple fitted distribution models and acquires this fitted distribution model as the optimal fitted distribution model.
[0082] <3.8. Calculation Process for Integrated Evaluation Indicators> In S10, the evaluation unit 107 calculates an integrated evaluation index. In this embodiment, the evaluation unit 107 sends instructions to the generation system AI to calculate the optimally fitted distribution model identified in S9, the multiple analysis targets included in the source data acquired in S1, and the integrated evaluation index, and calculates an integrated evaluation index for the analysis targets included in the source data.
[0083] Figure 4 is a flowchart showing the detailed process for calculating the integrated evaluation index.
[0084] <3.8.1. Calculation process for theoretical values> In S101, the deviation calculation unit 1071 calculates theoretical values based on the optimally fitted distribution model. In this embodiment, for each analysis indicator included in the source data acquired in S1, the deviation calculation unit 1071 sends an instruction to the generation system AI to calculate the theoretical value along with the actual indicator value corresponding to the optimally fitted distribution model and analysis indicator identified in S9. The AI then inputs the actual indicator value for each analysis target into the theoretical value calculation formula corresponding to the optimally fitted distribution model, thereby calculating theoretical values for each analysis target and each analysis indicator. The deviation calculation unit 1071 then acquires the calculated theoretical values for each analysis target and each analysis indicator.
[0085] <3.8.2. Calculation of the deviation vector between theoretical and actual values> In S102, the deviation calculation unit 1071 calculates a deviation vector based on the theoretical value. In this embodiment, the deviation calculation unit 1071 sends an instruction to the generation system AI to calculate the deviation between the theoretical value of the first analysis indicator calculated in S101 and the actual value of the second analysis indicator, which is the same as or different from the first analysis indicator, for a plurality of analysis indicators included in the source data acquired in S1, and the deviation between the theoretical value and the actual value of the indicator between the analysis indicators.
number
number
[0086] Then, the deviation calculation unit 1071 uses the deviation or deviation rate in the i-th analysis target to calculate the deviation vector R i =(r i1 ,r i2 ,···,r im The following is calculated: Here, m represents the number of items to be analyzed included in the source data.
[0087] <3.8.3. Calculation process of relative position index based on deviation vectors> In S103, the relative position calculation unit 1072 calculates a relative position index based on the deviation vector. In this embodiment, the relative position calculation unit 1072 sends an instruction to the generation system AI to calculate the deviation vector and relative position index calculated in S102, and the average value of the components of the deviation vector in the i-th analysis target.
number
[0088] Here, the interpretation of the relative position index is RPI i When RPI > 0, the combined performance of multiple indicators is generally higher than theoretically expected (upward trend). i If <0, the overall performance is below the theoretical value (downward trend). RPI i When the result is approximately 0, it indicates that the overall value is close to the theoretical value, or that the positive and negative deviations cancel each other out. This allows us to evaluate the relative position of each analysis target within the data structure.
[0089] <3.8.4. Calculation process of structural distance index based on deviation vectors> In S104, the deviation calculation unit 1073 calculates a structural distance index based on the deviation vector. In this embodiment, the deviation calculation unit 1073 sends an instruction to the generation system AI to calculate the deviation vector and structural distance index calculated in S102, and corrects the Euclidean distance between the analysis indicators in the i-th analysis target based on the correlation between these analysis indicators (the so-called Mahalanobis distance).
number
[0090] Furthermore, the structural distance index is interpreted as a quantification of how much the i-th analysis target deviates from the "ideal structure" as a multidimensional distance. The ideal structure is defined as a state where the measured values and theoretical values coincide for all indicators, i.e., the deviation vector is the zero vector R. i This is the state where (0,0,···,0). In other words, SD i A larger value indicates that the i-th analysis target deviates significantly from the ideal state, taking into account the correlation structure between variables, suggesting the possibility of some structural peculiarity or instability. i The smaller the value, the closer it is to the ideal structure and the more stable it is considered to be.
[0091] <3.8.5. Calculation process of consistency index based on deviation vectors> In S105, the consistency calculation unit 1074 calculates a consistency index based on the deviation vector. In this embodiment, the consistency calculation unit 1074 transmits the deviation calculated in S102, the relative position index calculated in S103, and an instruction to calculate the consistency index to the generation system AI, and calculates the standard deviation of each deviation of the deviation vector Ri based on the actual values of multiple indicators in the i-th analysis target.
number
[0092] Here, the interpretation of the consistency metric is CI i The smaller the value, the more aligned the direction and magnitude of the deviations of each indicator are, indicating a state of "high consistency" or "balance." On the other hand, CI i A higher value indicates that performance differs significantly across indicators, suggesting a state of "low consistency" or "imbalance," and representing a distorted structure where specific strengths and weaknesses coexist. This allows us to assess the quality of the data and the sustainability of the subject of analysis.
[0093] <3.8.6. Calculation Process for Integrated Evaluation Indicators> In S106, the integrated evaluation calculation unit 1075 calculates the integrated evaluation index. In this embodiment, the integrated evaluation calculation unit 1075 transmits the acquired relative position index, structural distance index, and consistency index, along with an instruction to calculate the integrated evaluation index, to the generative AI to normalize the relative position index, structural distance index, and consistency index (e.g., to a 0-1 scale), and then calculates a weighted sum using the respective scores and the weights w of each score according to the analysis objective, taking into consideration that smaller structural distance index (SD) and consistency index (CI) are preferable.
number
[0094] The integrated performance indicator is an evaluation score that comprehensively considers performance level (RPI), stability (SD), and balance (CI). By setting the weights, it is possible to design diverse evaluation axes, such as "evaluation that emphasizes growth" or "evaluation that emphasizes stability."
[0095] Furthermore, these weights are set by accepting the desired weight settings from the user who specified the source data in S1, and the integrated evaluation index is calculated using the weights corresponding to the weight settings. In addition, the formula for calculating the integrated evaluation index is not limited to the above formula, as long as it takes into account that smaller structural distance (SD) and consistency (CI) indices are preferable. This allows for automatic evaluation based on the context of the data structure, enabling efficient analysis of the target data.
[0096] As described above, the integrated evaluation index is calculated by executing processes S101 to S106, and the following evaluation comment generation process is then executed.
[0097] <3.9. Generation of Evaluation Comments> In S11, the comment generation unit 1076 processes an evaluation comment based on the integrated evaluation index. In this embodiment, the comment generation unit 1076 transmits the relative position index, structural distance index, and consistency index calculated in S10, along with instructions to identify factors influencing the integrated evaluation index, to the generation system AI to generate an evaluation comment including the influencing factors. Specifically, the comment generation unit 1076 identifies the scores with relatively large scores among the normalized relative position index, structural distance index, and consistency index as influencing factors.
[0098] Furthermore, the comment generation unit 1076 identifies a differentiation indicator that is an analytical indicator that contributes significantly to the integrated evaluation indicator and that significantly contributes to differentiation more than other indicators in the subject of analysis. Specifically, the comment generation unit 1076 sends the source data and instructions to the generation system AI to identify the analytical indicator that contributes significantly to the integrated evaluation indicator, thereby identifying the analytical indicator as a differentiation indicator. More specifically, for the scores identified as influencing factors, the comment generation unit 1076 sends instructions to the generation system AI to identify the source data and analytical indicators that contribute significantly to the scores of the influencing factors, thereby identifying the analytical indicator as a differentiation indicator. For example, in the calculation process of the structural distance indicator, an analytical indicator that has a significant influence on increasing the structural distance indicator (having a significant influence on reducing stability) is identified as a differentiation indicator. Also, for example, in the calculation process of the consistency indicator, a differentiation indicator can be identified in the same way as the structural distance indicator.
[0099] The comment generation unit 1076 then sends instructions to the generation system AI to generate an evaluation comment, along with the integrated evaluation indicator, influencing factors, and differentiation indicator, causing the AI to generate a comment that includes at least the integrated evaluation indicator, influencing factors, and differentiation indicator. The comment generation unit 1076 then acquires this comment as an evaluation comment.
[0100] <3.10. Displaying Analysis Results> In S12, the display processing unit 109 processes the analysis results for display. In this embodiment, the display processing unit 109 processes the analysis results (the analysis results from S3 and from S11) for display and displays the display processing results on the user terminal device (analysis device 1).
[0101] By executing the processes S1 to S12 described above, the distribution model that best fits the data distribution is identified for each analytical indicator, and by evaluating the target of analysis, the underlying structure of the source data can be reflected, enabling more accurate analysis. Conventional statistical methods have lost important information inherent in the data, especially the existence of tail risks and high performers, by removing "outliers" or performing uniform normalization. In contrast, DSA introduces the concept of "structural singularities" and solves conventional problems by performing processing adapted to each distribution. As a result, in fields such as finance, human resources, and healthcare, it becomes possible to obtain more realistic and practical insights that could not be obtained with conventional methods, such as improved accuracy in risk assessment and identification of true high performers.
[0102] In a preferred embodiment of the present invention, the integrated evaluation calculation unit 1075 may update the integrated evaluation index based on the environmental compatibility index. Specifically, the integrated evaluation calculation unit 1075 updates the integrated evaluation index by adding the environmental compatibility index, which has been multiplied by a predetermined weight, to the integrated evaluation index. This allows for the inclusion of environmental compatibility indicators, which take into account the degree of environmental change, in the evaluation criteria, enabling a more accurate assessment of the subject of analysis.
[0103] Furthermore, in a preferred embodiment of the present invention, key analytical indicators, which are important indicators in the analysis, are identified from among the analytical indicators, and predetermined weights used in calculating the integrated evaluation indicator and the environmental compatibility indicator are set in accordance with the optimally fitted distribution model, by referring to the optimally fitted distribution model identified in the key analytical indicators. The process of setting the key analytical indicators is performed, for example, depending on which growth period the analysis is performed in, such as short-term, medium-term, or long-term.
[0104] Specifically, this system includes an indicator setting unit (not shown), in which key analytical indicators that are important for short-term growth and key analytical indicators that are important for long-term growth are registered in database 2 for each field of analysis. The indicator setting unit accepts a selection from the user for short-term or long-term growth and sets key analytical indicators based on the field of analysis included in the acquired source data and the selected growth period. The field of analysis included in the source data can be identified by applying known text analysis techniques to the source data or by inputting or specifying data from the user.
[0105] In addition, key analytical indicators may be set using source data from past points in time. Specifically, a change index correspondence table showing key analytical indicators according to the degree of change in the analysis environment is registered in database 2, and the indicator setting unit determines whether or not source data from past points in time that includes the same multiple analysis targets can be obtained in addition to the acquired source data (for example, whether it is stored in database 2 or entered by the user). If source data from past points in time can be obtained, the indicator setting unit identifies the degree of change in the analysis environment from the fluctuations in the numerical values contained in the latest acquired source data and source data from past points in time. The indicator setting unit then compares the identified degree of change with the change index correspondence table and sets the indicator corresponding to the identified degree of change as a key analytical indicator.
[0106] Alternatively, a generative AI may be used to set key analytical indicators. Specifically, the indicator setting unit sends instructions to the generative AI to generate questions to elicit key analytical indicators, along with the source data, and receives questions from the generative AI to extract key analytical indicators from among multiple indicators included in the source data. The indicator setting unit also sends the user's answers to the received questions, instructions to extract key analytical indicators based on the answers, and instructions to generate new questions if key analytical indicators cannot be extracted, and receives the indicators extracted as key analytical indicators from the generative AI and sets them as key analytical indicators.
[0107] In this embodiment, the calculation process, identification process, and generation process are performed within the system by the execution of the program of the present invention by hardware. However, it may also be a process in which the system sends prompts containing the source data and instructions to perform each process on the source data to an external generation AI and obtains the processing results. Alternatively, the system may have a generation AI, and the generation AI may perform each process. Furthermore, different generation AIs may be used to execute each process, or multiple processes may be executed on a single generation AI.
[0108] In this embodiment, the generative AI is a generative neural network pre-trained for natural language tasks. A representative example is the GPT (Generative Pre-trained Transformer) model, but its parameter size and network configuration are not limited. For example, a hybrid model may be used, with a Transformer at its core, supplemented by a convolutional neural network (CNN) or a recurrent neural network (LSTM, GRU, etc.).
[0109] In this embodiment, the display process refers to the process in which the display processing unit 109 executes a process to generate the necessary information, transmits the generated information to the output unit 15 of the analysis device 1, and the output unit 15 displays the generated information. In the case where the analysis system 0 is composed of a server and an analysis device 1 (client), and the server is equipped with a display processing unit 109, the display processing unit 109 may execute a process to generate the information necessary for display, transmit the generated information to the analysis device 1, and the analysis device 1 may display the generated information. Alternatively, the display processing unit 109 may transmit a processing command to the analysis device 1 to generate the information necessary for display, causing the analysis device 1 to generate the information necessary for display and display the generated information. [Explanation of Symbols]
[0110] 0: Analysis System 1:Analyzer 2: Database 11: Processing Section 12: Storage section 13: Communications Department 14: Input section 15: Output section NW: Communication Network 101: Acquisition Department 102:Analysis means identification part 103:Distribution candidate identification part 104: Parameter estimation unit 105: Fit Calculation Unit 106:Distribution identification part 107: Evaluation Department 1071: Deviation calculation unit 1072: Relative position calculation unit 1073: Deviation calculation unit 1074: Consistency Calculation Unit 1075: Integrated Evaluation Calculation Unit 1076: Comment generation unit 108: Environmental compatibility index calculation department 109: Display Processing Unit
Claims
1. An analytical system for analyzing the object of analysis, The analysis system comprises an acquisition unit, a distribution candidate identification unit, a parameter estimation unit, a goodness-of-fit calculation unit, a distribution identification unit, a deviation calculation unit, and an evaluation unit. The acquisition unit acquires analysis source data including analysis indicators that serve as indicators in the analysis of the subject to be analyzed, and actual values of the subject to be analyzed for the analysis indicators. The distribution candidate identification unit identifies a predetermined distribution model to be applied to the multiple actual values of the analysis targets as a candidate distribution model. The parameter estimation unit estimates the parameters for each candidate distribution model to be applied to the multiple actual values of the targets to be analyzed included in the source data, such that the likelihood of the source data with respect to the actual values is maximized. The goodness-of-fit calculation unit calculates the distribution fit for each candidate distribution model whose parameters have been estimated, for the data distribution under analysis. The distribution identification unit, based on the distribution fit, identifies a fitted distribution model from among a plurality of candidate distribution models that is suitable for the data distribution to be analyzed. The deviation calculation unit substitutes the parameters of the specified fitted distribution model into a theoretical value calculation formula that has parameters corresponding to each fitted distribution model as variables, calculates the theoretical value of the analysis indicator under analysis, and calculates the deviation between the theoretical value and the actual value of the analysis indicator under analysis included in the source data. The evaluation unit evaluates the object of analysis based on the deviation. Analysis system.
2. The deviation calculation unit calculates the deviation between the theoretical value and the actual value of the analysis indicator for each analysis indicator, and calculates a plurality of deviation vectors based on the deviations for each analysis target. The evaluation unit calculates the average value of the deviation for each analysis indicator, which is a component of the deviation vector for each analysis target, as a relative position indicator showing the overall trend of the analysis target based on the multiple analysis indicators. As a structural distance index that shows the degree of deviation obtained by correcting the distance between the actual value and the theoretical value based on the correlation between the multiple analytical indicators, the Mahalanobis distance is calculated using the deviation vector and the covariance matrix based on the deviation vector for each of the analyzed targets as parameters. As a consistency index indicating the degree of variability of the actual values of the multiple analytical indicators in the subject of analysis, the degree of variability of the deviation vector is calculated based on the deviation for each analytical indicator, which is a component of the deviation vector for each subject of analysis, and the mean value of said deviation. Based on the relative position index, the structural distance index, and the consistency index, an integrated evaluation index for evaluating the object of analysis is calculated. The analysis system according to claim 1.
3. The goodness-of-fit calculation unit calculates the first goodness-of-fit distribution using the Kolmogorov-Smirnov test or the Anderson-Darling test, with the result of comparing the value of the theoretical cumulative distribution function based on the parameters of the fitted distribution model and the value of the empirical cumulative distribution function based on the actual values of the multiple subjects of analysis as variables. The distribution identification unit identifies one or more fitted distribution models based on the first distribution fit, The aforementioned predetermined distribution model has a predetermined number of parameters for each distribution model, When multiple fitted distribution models are identified, the goodness-of-fit calculation unit calculates a second distribution goodness-of-fit by substituting the likelihood, which indicates the degree of fit to the data distribution, and the number of parameters of the identified multiple fitted distribution models into the information criterion formula. The distribution identification unit identifies one fitted distribution model based on the second distribution fit. The analysis system according to claim 1 or claim 2.
4. The evaluation unit transmits instructions to the generation AI to identify the analysis indicators that have a significant influence on the calculation process of the integrated evaluation indicators for the source data and the target of analysis, and to identify the analysis indicators that have a significant influence on the calculation process of the integrated evaluation indicators from among the multiple analysis indicators included in the source data. The analysis system according to claim 2.
5. The analysis system further comprises an analysis means identification unit, Reference information for the fields to be analyzed is pre-configured. The analysis means identification unit identifies the analysis field of the source data based on a plurality of analysis indicators included in the acquired source data, and compares the analysis field with the analysis target reference information to determine whether or not the source data is the target of analysis. The analysis system according to claim 1.
6. The aforementioned analysis system further comprises an environmental compatibility index calculation unit, The environmental suitability index calculation unit identifies the standard deviation or power exponent, which is a parameter of the suitability distribution model, as an analysis environment concentration score indicating the degree of concentration of the analysis environment corresponding to the identified suitability distribution model. The latest source data for analysis is compared with the source data for analysis from a past point in time that includes the same analysis targets as the latest source data. The change in the value of multiple indicators from the past source data to the latest source data is calculated as the degree of change in the analysis environment, which indicates the degree of change in the analysis environment composed of the multiple analysis targets over time. Using the aforementioned concentration of the analysis environment, the degree of change in the analysis environment, and the deviation as variables, an environmental suitability index indicating the degree of suitability to the analysis environment is calculated using a calculation formula that has weighting for each variable according to the identified fit distribution model. The aforementioned analysis environment is a population in which multiple subjects of analysis exist. The analysis system according to claim 1.
7. An analytical program for analyzing the target of analysis, The aforementioned analysis program causes the computer to function as an acquisition unit, a distribution candidate identification unit, a parameter estimation unit, a goodness-of-fit calculation unit, a distribution identification unit, a deviation calculation unit, and an evaluation unit. The acquisition unit acquires, for each of the analysis targets, raw data for analysis, including analytical indicators that serve as benchmarks in the analysis of the analysis targets, and actual values of the analysis targets for the analytical indicators. The distribution candidate identification unit identifies a predetermined distribution model to be applied to the multiple actual values of the analysis targets as a candidate distribution model. The parameter estimation unit estimates the parameters for each candidate distribution model to be applied to the multiple actual values of the targets to be analyzed included in the source data, such that the likelihood of the source data with respect to the actual values is maximized. The goodness-of-fit calculation unit calculates the distribution fit for each candidate distribution model whose parameters have been estimated, for the data distribution under analysis. The distribution identification unit, based on the distribution fit, identifies a fitted distribution model from among a plurality of candidate distribution models that is suitable for the data distribution to be analyzed. The deviation calculation unit substitutes the parameters of the specified fitted distribution model into a theoretical value calculation formula that has parameters corresponding to each fitted distribution model as variables, calculates the theoretical value of the analysis index of the subject of analysis, and calculates the deviation between the theoretical value and the value of the analysis index of the subject of analysis included in the source data. The evaluation unit evaluates the object of analysis based on the deviation. Analysis program.
8. An analytical method using an analytical system for analyzing the object of analysis, The aforementioned analysis system The process of obtaining source data for analysis, including analytical indicators that serve as indicators in the analysis of the subject of analysis, and the actual values of the subject of analysis for the analytical indicators, A process to identify a predetermined distribution model as a candidate distribution model to be applied to the actual values of multiple targets for analysis, The parameters for each candidate distribution model, which are applied to the actual values of the multiple data distributions of the data to be analyzed included in the source data, are estimated so as to maximize the likelihood of the source data with respect to the actual values. A process for calculating the distribution fit for the data distribution under analysis for each candidate distribution model whose parameters have been estimated, Based on the distribution fit described above, a process is performed to identify a fitted distribution model from among several candidate distribution models that is suitable for the data distribution to be analyzed. The process involves substituting the parameters of the identified fitted distribution model into a theoretical value calculation formula that has parameters corresponding to each fitted distribution model as variables to calculate the theoretical value of the analysis indicator under analysis, and then calculating the deviation between the theoretical value and the actual value of the analysis indicator under analysis included in the source data. A process for evaluating the object of analysis based on the aforementioned deviation, An analysis method to perform this task.