Information processing system, information processing method, and information processing program
The information processing system addresses the challenge of missing data in raw material composition prediction by calculating regression vectors to fill in missing values, ensuring accurate composition prediction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RESONAC CORP
- Filing Date
- 2024-11-06
- Publication Date
- 2026-05-19
AI Technical Summary
Existing methods for predicting compositions using regression analysis or classification with raw materials face challenges when characteristic quantity data is missing, limiting their applicability.
An information processing system that calculates first and second regression coefficient vectors to fill in missing feature data values, using formulation and characteristic data to determine relationships and supplement the feature data, enabling predictions even with incomplete data.
Enables accurate prediction of composition characteristics by complementing feature data, allowing for effective composition prediction even when some raw material features are missing.
Smart Images

Figure 2026082083000001_ABST
Abstract
Description
Technical Field
[0001] One aspect of the present disclosure relates to an information processing system, an information processing method, and an information processing program.
Background Art
[0002] In regression analysis or classification using composition data, when setting the composition of raw materials and the blending ratio between raw materials as explanatory variables, there is a problem that prediction of a composition containing new raw materials cannot be made. Non-Patent Documents 1 and 2 propose methods for solving this problem.
[0003] Non-Patent Document 1 describes an optimization approach that simultaneously handles three degrees of freedom: selection of a plurality of raw materials, selection of the ratio for mixing the plurality of raw materials, and selection of processing conditions used for manufacturing the plurality of raw materials. In this approach, a partial least squares (PLS) model combining a database of previously created mixtures and a database of data on the properties of constituent materials used in these mixtures is constructed. Next, the resulting model is used in an optimization framework to select raw materials from a larger database containing raw materials that have not been used before, and to select a mixing ratio for producing a blended product having specified final properties at minimum cost.
[0004] Non-Patent Document 2 describes a new mixture property PLS model that represents the relationship between raw material properties and the properties of the final mixed product, taking into account the influence of the mixing ratio and processing conditions of each material. This document presents a new approach to mixture design of experiments (mixture DOE) based on the form of the mixture property PLS model, and shows that it provides a more efficient approach for selecting materials and conditions.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Non - Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in reality, there are raw materials with some of the characteristic quantity data indicating the characteristics of the raw materials missing, so there are few situations where the methods described in Non - Patent Documents 1 and 2 can be directly applied. Therefore, a mechanism for appropriately complementing the characteristic quantity data indicating the characteristics of individual raw materials is desired.
Means for Solving the Problems
[0007] The information processing system relating to one aspect of this disclosure comprises at least one processor. The at least one processor acquires formulation data, which includes a plurality of data records showing the correspondence between an identifier for a formulation process and the element values in the formulation process for each of a plurality of raw materials; acquires feature data, which includes one or more missing values, showing one or more features for each of the plurality of raw materials; acquires characteristic data, which includes characteristic values for each of a plurality of compositions corresponding to the plurality of formulation processes shown by the formulation data; calculates a first regression coefficient vector showing the relationship between the formulation data and the characteristic data by first regression analysis based on the formulation data and the characteristic data; calculates a second regression coefficient vector showing the relationship between the formulation data, provisional feature data obtained by replacing one or more missing values in the feature data with one or more provisional values, and the characteristic data by second regression analysis based on the formulation data, the provisional feature data, and the characteristic data, while repeatedly changing the one or more provisional values, to identify one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, and to supplement the feature data by replacing one or more missing values in the feature data with the identified one or more provisional values.
[0008] In this respect, the calculation of the second regression coefficient vector, which shows the relationship between the blending data, feature data, and characteristic data, is repeated while changing the provisional values used to fill in missing values in the feature data. Then, the missing values in the feature data are filled in by provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector, which is calculated without being affected by missing features, satisfies predetermined conditions. The relationship between the blending of multiple raw materials and the characteristic values of the composition can be determined by both a first regression analysis that does not consider features and a second regression analysis that does consider features, and by filling in missing values based on the difference between these two types of regression analyses, the feature data showing the characteristics of individual raw materials can be appropriately supplemented. [Effects of the Invention]
[0009] According to one aspect of this disclosure, feature data representing the characteristics of individual raw materials can be appropriately complemented. [Brief explanation of the drawing]
[0010] [Figure 1] This figure shows an example of the functional configuration of an information processing system. [Figure 2] This figure shows an example of data used in an information processing system. [Figure 3] This flowchart shows an example of a process for imputing feature data. [Figure 4] This flowchart shows an example of the process for calculating new characteristic data. [Modes for carrying out the invention]
[0011] The following describes various examples in this disclosure in detail with reference to the attached drawings. In the description of the drawings, identical or equivalent elements are denoted by the same reference numeral, and redundant descriptions are omitted.
[0012] [System Overview] The information processing system described herein is a computer system that complements feature data representing the characteristics of multiple raw materials. Since the feature data is used to make predictions about compositions obtained by blending multiple raw materials, predictions cannot be made if any part of the feature data is missing. However, in reality, the raw materials that can be considered in the research or development of compositions such as resin compositions are diverse, and it is difficult to obtain complete feature data for all raw materials. By using this information processing system, it becomes possible to make predictions about compositions even when there are missing features in the feature data. Feature data complementation refers to the process of replacing (filling in) one or more missing values in the feature data with meaningful values.
[0013] The information processing system uses formulation data, which shows the element values of individual raw materials in each of multiple formulation processes, and characteristic data, which shows the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes, in order to complement the feature data. The information processing system calculates a first regression coefficient vector showing the relationship between the formulation data and the characteristic data by first regression analysis. The information processing system also calculates a second regression coefficient vector showing the relationship between the formulation data, provisional feature data (feature data with one or more missing values provisionally filled), and characteristic data by second regression analysis, while changing one or more provisional values to fill one or more missing values in the feature data. The information processing system complements the feature data with one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions. In other words, the information processing system determines the relationship between the formulation data and characteristic data by first regression analysis without considering features and second regression analysis considering features, and complements the feature data based on the difference between these two types of regression analysis.
[0014] In one example, the information processing system calculates the characteristic values of individual compositions when a new raw material is added to each individual formulation process. In this example, the information processing system recalculates a second regression coefficient vector based on formulation data, characteristic data, and interpolated feature data using second regression analysis. Based on the new formulation data with added element values for the new raw material, the new feature data with added features for the new raw material, and the recalculated second regression coefficient vector, the information processing system calculates new characteristic data showing the characteristic values of individual compositions when a new raw material is added to each individual formulation process. By using the information processing system, it becomes possible to predict the characteristic values of compositions generated by a formulation process with added new raw materials, even if some of the original feature data is missing.
[0015] As mentioned above, the composition may be a resin composition. In this case, each raw material may be a monomer or a polymer.
[0016] [System Configuration] An information processing system consists of one or more computers. When multiple computers are used, these computers are connected via a communication network such as the internet or an intranet, thereby logically constructing a single information processing system.
[0017] A computer comprising an information processing system generally includes a processor, storage device (memory), and communication interface as hardware components. The processor is, for example, a CPU, and the storage device consists of flash memory, hard disks, etc. The communication interface consists of network cards, wireless communication modules, etc. Each functional module of the information processing system is realized when the processor executes programs stored in the storage device.
[0018] An information processing program for enabling a computer to function as an information processing system includes program code for realizing each functional module of the information processing system. This information processing program may be provided on a non-temporary recording medium such as a CD-ROM, DVD-ROM, or semiconductor memory. Alternatively, the information processing program may be provided via a communication network as a data signal superimposed on a carrier wave. The provided information processing program is recorded, for example, on a storage device.
[0019] The configuration of an example information processing system 10 will be described with reference to Figures 1 and 2. Figure 1 is a diagram showing the functional configuration of the information processing system 10. Figure 2 is a diagram showing an example of data used in the information processing system 10.
[0020] In one example, the information processing system 10 connects to a database 20 and a user terminal 30 via a communication network. The communication network is typically constructed using the internet, an intranet, or a combination thereof. The communication network can also be constructed using a wired network, a wireless network, or a combination thereof.
[0021] The database 20 is a storage device that stores formulation data 21, feature data 22, and characteristic data 23. The database 20 may be a component of the information processing system 10, or it may be located outside the information processing system 10.
[0022] The formulation data 21 is electronic data that shows the element values of individual raw materials in each of multiple formulation processes. In one example, the formulation data 21 includes multiple data records that show the correspondence between the identifier of the formulation process and the element values for each of the multiple raw materials in that formulation process. The element values may be quantities related to the formulation of raw materials in the formulation process, for example, they may be formulation amounts or formulation ratios. In the example in Figure 2, the identifier of the formulation process is the process ID, and multiple raw materials such as raw materials Ma, Mb, etc. are represented as a raw material list, and the individual element values are r 11 ,r 12 ,r 21 It is expressed as follows:
[0023] Feature data 22 is electronic data representing one or more features for each of several raw materials. In one example, feature data 22 includes multiple data records showing the correspondence between a raw material identifier and one or more features of that raw material. A feature is a numerical value that quantitatively represents the characteristics of a raw material. Individual features may represent characteristics related to the structure or physical properties of the raw material, for example, characteristics related to viscosity, glass transition temperature (Tg), curing time, etc. Feature data contains one or more missing values. If the feature Fy of raw material Mx is a missing value, it means that no specific value has been set for the feature Fy of raw material Mx. In Figure 2, missing values in feature data 22 are represented by blank spaces. For example, for raw material Mb, features Fa and Fd are missing values, while features Fb and Fc are meaningful specific values. Note that the value "0" is a meaningful value and not a missing value. For a given raw material, meaningful values may be set for all features, or meaningful values may be set for any raw material with respect to a given feature.
[0024] The characteristic data 23 is electronic data that shows the characteristic values of multiple compositions corresponding to multiple compounding processes indicated by the compounding data 21. In one example, the compounding data 21 includes multiple data records that show the correspondence between an identifier for a compounding process, an identifier for a composition obtained by the compounding process, and the characteristic values of the composition. The characteristic values of a composition refer to values that indicate the physical properties unique to the composition. In the example in Figure 2, the identifier for a compounding process is the process ID, and the individual characteristic values are represented as c1, c2, etc.
[0025] The user terminal 30 is a computer used by a user of the information processing system 10. The user terminal 30 can be any type of computer, such as a personal computer, workstation, tablet, smartphone, or wearable device.
[0026] The information processing system 10 includes a processor 101 that functions as a first data acquisition unit 11, a first regression analysis unit 12, a second regression analysis unit 13, a search unit 14, a completion unit 15, a second data acquisition unit 16, and a prediction unit 17.
[0027] The first data acquisition unit 11 is a functional module that acquires formulation data 21, feature data 22, and characteristic data 23 from the database 20 based on instructions from the user terminal 30.
[0028] The first regression analysis unit 12 is a functional module that calculates a first regression coefficient vector showing the relationship between the blending data 21 and the characteristic data 23 through a first regression analysis based on the blending data 21 and the characteristic data 23.
[0029] The second regression analysis unit 13 is a functional module that calculates a second regression coefficient vector showing the relationship between the blending data 21, the provisional feature data, and the characteristic data 23 through a second regression analysis based on the blending data 21, the provisional feature data, and the characteristic data 23.
[0030] The search unit 14 is a functional module that identifies one or more provisional values corresponding to a second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions. The search unit 14 repeatedly causes the second regression analysis unit 13 to calculate the second regression coefficient vector while changing one or more provisional values used to fill in one or more missing values in the feature data 22. The search unit 14 compares each of the multiple second regression coefficient vectors obtained through this iterative process with the first regression coefficient vector and selects one second regression coefficient vector that satisfies predetermined conditions regarding the difference. Then, the search unit 14 identifies one or more provisional values corresponding to the selected second regression coefficient vector.
[0031] The interpolation unit 15 is a functional module that interpolates the feature data 22 by replacing one or more missing values in the feature data 22 with one or more identified provisional values.
[0032] The second data acquisition unit 16 is a functional module that acquires new formulation data and new feature data based on instructions from the user terminal 30. The new formulation data is electronic data obtained by adding the element values of new raw materials to the formulation data 21. The new feature data is electronic data obtained by adding the feature quantities of the new raw materials to the complementary feature data.
[0033] The prediction unit 17 is a functional module that calculates new characteristic data representing the characteristic values of multiple compositions corresponding to multiple blending processes indicated by the new blending data, based on new blending data and new feature data. In other words, the prediction unit 17 calculates the characteristic values of compositions generated by blending processes to which new raw materials have been added.
[0034] [System operation] The following describes the operation of the information processing system 10 and the information processing method related to this disclosure. Since the following explanation uses calculation formulas, the blending data 21, feature data 22, and characteristic data 23 will also be referred to as "blending data R," "feature data S," and "characteristic data Y," respectively.
[0035] (Completion of Feature Data) Referring to FIG. 3, the process of completing the feature data will be described. FIG. 3 is a flowchart showing an example of the process as processing flow S1.
[0036] In step S11, the first data acquisition unit 11 acquires the formulation data R, the feature data S including one or more missing values, and the characteristic data Y based on an instruction from the user terminal 30. The user performs a predetermined user operation to cause the information processing system 10 to complete the feature data, and the user terminal 30 transmits an instruction signal to the information processing system 10 in response to the user operation. The first data acquisition unit 11 reads the formulation data R, the feature data S, and the characteristic data Y from the database 20 in response to receiving the instruction signal.
[0037] In step S12, the first regression analysis unit 12 calculates a first regression coefficient vector β c representing the relationship between the formulation data R and the characteristic data Y by performing a first regression analysis based on the formulation data R and the characteristic data Y.
[0038] Let the number of formulation processes (i.e., the number of data records) and the number of raw materials indicated by the formulation data R be N and d, respectively. The formulation data R is represented by an N-row × d-column matrix. The component r ij in the i-th row and j-th column of this matrix indicates the element value of the j-th raw material in the i-th formulation process. On the other hand, the characteristic data Y is represented by an N-dimensional vector. The i-th component c i of this vector indicates the characteristic value of the composition obtained by the i-th formulation process.
[0039] The first regression analysis unit 12 performs a first regression analysis using the linear regression model shown in Equation (1) as the first regression model to calculate the first regression coefficient vector β c . The first regression coefficient vector β c is a d-dimensional vector. ε is an error vector. For example, the first regression analysis unit 12 solves Equation (1) by the least squares method to calculate the first regression coefficient vector β c . Y=Rβ c +ε …(1)
[0040] In step S13, the search unit 14 replaces one or more missing values in the feature data S with one or more provisional values to generate provisional feature data S'. Each provisional value is a meaningful value. The feature data S is represented as a d x p matrix, where p is the number of features. The i-th row and j-th column of this matrix are elements s ij This value is either meaningful or missing. For example, the search unit 14 randomly determines an initial provisional value for each of the one or more missing values and replaces the missing value with this provisional value.
[0041] In step S14, the second regression analysis unit 13 generates a second regression coefficient vector β that shows the relationship between the blending data R, the provisional feature data S', and the characteristic data Y. m This is calculated using a second regression analysis based on the formulation data R, provisional feature data S', and characteristic data Y.
[0042] The second regression analysis unit 13 performs a second regression analysis using the linear regression model shown in equation (2) as the second regression model, and obtains the second regression coefficient vector β. m Calculate the first regression coefficient vector β. c Unlike the second regression coefficient vector β m is a p-dimensional vector. ε is the error vector. For example, the second regression analysis unit 13 solves equation (2) using the least squares method to obtain the second regression coefficient vector β m Calculate. Y=RS'β m +ε …(2)
[0043] In step S15, the search unit 14 generates the first regression coefficient vector β c and the second regression coefficient vector β m The difference between this and is calculated. The search unit 14 calculates this difference by S'β m Calculate the second regression coefficient vector β m This is converted into a d-dimensional vector. Then, the search unit 14 uses equation (3) to obtain the first regression coefficient vector β cand the second regression coefficient vector β m The difference is calculated using the squared error, and this difference is identified as the loss L. L=||Sβ m -β c || 2 …(3)
[0044] In step S16, the search unit 14 determines whether or not to search for one or more values to fill one or more missing values in the feature data S. The search may terminate when the number of iterations reaches a predetermined value. Alternatively, the termination condition may be when the change in loss L falls below a predetermined threshold, i.e., when loss L converges. If the search is to continue (NO in step S16), the process proceeds to step S17. On the other hand, if the termination condition is met (YES in step S16), the process proceeds to step S18.
[0045] In step S17, the search unit 14 changes one or more provisional values. The search unit 14 may determine each provisional value randomly, or it may determine each provisional value using a method aimed at efficient search, such as Bayesian optimization.
[0046] After step S17, the process returns to step S13. In the repeated step S13, the search unit 14 replaces one or more missing values in the feature data S with one or more modified provisional values to generate the next provisional feature data S'. In the repeated step S14, the second regression analysis unit 13 performs a second regression analysis using the provisional feature data S' to generate the next second regression coefficient vector β m The first regression coefficient vector β is calculated. In the repeated step S15, the search unit 14 calculates the first regression coefficient vector β. c and its second regression coefficient vector β m Calculate the difference (loss L).
[0047] As shown by the repetition of steps S13 to S17, the second regression analysis unit 13 and the search unit 14 work together to repeatedly perform the process of calculating the second regression coefficient vector by the second regression analysis, changing one or more provisional values.
[0048] In step S18, the interpolation unit 15 interpolates the feature data S using provisional values where the difference (loss L) satisfies predetermined conditions. The interpolation unit 15 then interpolates the first regression coefficient vector β c The difference between this and the second regression coefficient vector β satisfies the specified conditions. m It identifies one or more provisional values corresponding to this. That is, the interpolation unit 15 identifies the first regression coefficient vector β c The difference between this and the second regression coefficient vector β satisfies the specified conditions. m The derived provisional value of 1 or greater is identified. The condition for this is that the difference is locally minimized. That is, the interpolation unit 15 identifies the first regression coefficient vector β c The second regression coefficient vector β whose difference from is locally minimized. m A provisional value of 1 or more corresponding to this may be identified. The "locally minimum difference" refers to the smallest difference obtained in the search by repeating steps S13 to S17. As another example, the condition may be that the difference is less than a predetermined threshold. That is, the interpolation unit 15 determines the first regression coefficient vector β c The second regression coefficient vector β is 1 or greater and the difference from is less than a predetermined threshold. m Select one of the following, and the selected second regression coefficient vector β m One or more provisional values corresponding to each may be identified. The interpolation unit 15 interpolates the feature data S by replacing one or more missing values in the feature data S with the identified one or more provisional values. In the following, the interpolated feature data is S comp This is how it is expressed.
[0049] The interpolation unit 15 contains the interpolated feature data S. comp The output is, for example, the interpolation unit 15 outputs the interpolated feature data S. comp This data may be stored in a storage device such as the database 20, or it may be sent to the user terminal 30.
[0050] (Calculation of new characteristic data) Referring to Figure 4, we will now explain the process for calculating new characteristic data. Figure 4 is a flowchart showing an example of this process as process flow S2.
[0051] In step S21, the second regression analysis unit 13 analyzes the combination data R, characteristic data Y, and the complementary feature data S. comp Perform a second regression analysis based on the following to obtain the second regression coefficient vector β m Recalculate the equation. This recalculation is equivalent to solving equation (4), which corresponds to equation (2) above. Y=RS comp β m +ε …(4)
[0052] In step S22, the second data acquisition unit 16 acquires new formulation data R based on instructions from the user terminal 30. new And new feature data S new The system acquires the following. For example, the user performs a predetermined user operation to pass additional information about at least one new raw material to the information processing system 10, and the user terminal 30 transmits an instruction signal containing the additional information to the information processing system 10. The additional information shows the element values and each feature quantity for each blending process for each of the at least one new raw material. Meaningful values are set for all feature quantities represented by the additional information. The second data acquisition unit 16 responds to receiving the instruction signal by acquiring the blending data R new and characteristic data Y new The second data acquisition unit 16 adds the respective element values of at least one new raw material to each of the multiple data records of the formulation data R, thereby generating the new formulation data R. new This generates new formulation data R. new This is obtained by adding at least one new column corresponding to at least one new raw material to the formulation data R. The second data acquisition unit 16 also obtains one or more features for each of the at least one new raw material from the complemented feature data S. comp By adding this, new feature data S new This generates new feature data S. new This involves at least one data record corresponding to at least one new raw material being used as feature data S. comp These can be obtained by adding them to the new formulation data R.new and new feature data S new This is an example of obtaining [something].
[0053] In step S23, the prediction unit 17 generates new formulation data R new And new feature data S new And the recalculated second regression coefficient vector β m Based on this, new characteristic data Y new This is calculated. This calculation is expressed by equation (5), which corresponds to equation (2) above. Y new =R new S new β m …(5)
[0054] New characteristic data Y new This is new formulation data R new This shows the characteristic values of each of the multiple compositions corresponding to the multiple compounding processes shown, that is, the characteristic values of each composition when a new raw material is added to each compounding process.
[0055] The prediction unit 17 generates new characteristic data Y new It outputs the following. For example, the prediction unit 17 outputs new characteristic data Y new This data may be stored in a storage device such as the database 20, or it may be sent to the user terminal 30.
[0056] [Differentiation] The technology relating to this disclosure has been described in detail above based on various examples. However, this disclosure is not limited to the examples given above. The technology relating to this disclosure can be modified in various ways without departing from its essence.
[0057] The information processing system does not need to include functional modules corresponding to the second data acquisition unit 16 and prediction unit 17 described above. That is, the information processing system may execute step S21 in the processing flow S2 described above without executing steps S22 and S23, or it may not execute the entire processing flow S2. New characteristic data may be calculated by another computer system, in which case the information processing system may provide the recalculated second regression coefficient vector to the other computer system.
[0058] In the example above, the information processing system 10 reads the blending data R, feature data S, and characteristic data Y from the database 20. The information processing system may acquire at least one of these three types of data by another method, for example, by receiving the data from a user terminal. In a similar variation, the information processing system may read new blending data R new and new feature data S new At least one of these may be received from the user terminal.
[0059] In the example above, the information processing system 10 plays the role of a server in a client-server system. Alternatively, the functions of the information processing system 10 and the database 20 may be implemented on a standalone computer. Or, the information processing system may be implemented on a user terminal that can access the database 20 via a communication network.
[0060] The processing steps for a method executed by at least one processor are not limited to the examples above. For example, some of the steps described above may be omitted, or each step may be performed in a different order. Also, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be performed in addition to each of the steps described above.
[0061] In comparing the relative magnitudes of two numerical values in this disclosure, either of the two criteria, "greater than or equal to" and "greater than," may be used, or either of the two criteria, "less than or equal to" and "less than," may be used.
[0062] In this disclosure, the expression "at least one processor executes a first process, a second process, ... and the nth process," or a corresponding expression, refers to a concept that includes cases where the entity executing the n processes from the first process to the nth process, i.e., the processor, changes along the way. In other words, this expression refers to a concept that includes both cases where all n processes are executed by the same processor and cases where the processor changes at an arbitrary rate for the n processes.
[0063] [Note] As can be seen from the various examples above, this disclosure includes the following aspects:
[0064] (Note 1) Equipped with at least one processor, The at least one processor, The system retrieves formulation data that includes multiple data records showing the correspondence between the identifier of the formulation process and the element value in that formulation process for each of the multiple raw materials. The feature data obtained shows one or more feature quantities for each of the aforementioned multiple raw materials and includes one or more missing values. Characteristic data is obtained that shows the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the formulation data. A first regression coefficient vector showing the relationship between the blending data and the characteristic data is calculated by first regression analysis based on the blending data and the characteristic data. The process of calculating a second regression coefficient vector showing the relationship between the blending data, the provisional feature data obtained by replacing the one or more missing values in the feature data with one or more provisional values, and the characteristic data, by performing a second regression analysis based on the blending data, the provisional feature data, and the characteristic data, is repeatedly performed while changing the one or more provisional values, thereby identifying the one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, The feature data is supplemented by replacing the one or more missing values in the feature data with the one or more identified provisional values. Information processing system. (Note 2) The at least one processor identifies one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector is locally minimized. The information processing system described in Appendix 1. (Note 3) The at least one processor performs the second regression analysis based on the formulation data, the characteristic data, and the complementary feature data to recalculate the second regression coefficient vector. The information processing system described in Appendix 1 or 2. (Note 4) The at least one processor, New formulation data is obtained by adding the respective element values of at least one new raw material to each of the plurality of data records of the formulation data. New feature data is obtained by adding the one or more feature quantities for each of the at least one new raw materials to the complementary feature data. Based on the new formulation data, the new feature data, and the recalculated second regression coefficient vector, new characteristic data is calculated that shows the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the new formulation data. The information processing system described in Appendix 3. (Note 5) An information processing method performed by an information processing system comprising at least one processor, A step of obtaining formulation data which includes an identifier for a formulation process and multiple data records that show the correspondence between the identifier for the formulation process and the element value for each of the multiple raw materials in that formulation process, The steps include: obtaining feature data that shows one or more features of each of the aforementioned multiple raw materials and includes one or more missing values; A step of obtaining characteristic data showing the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the formulation data, A step of calculating a first regression coefficient vector showing the relationship between the blending data and the characteristic data by first regression analysis based on the blending data and the characteristic data, The process of calculating a second regression coefficient vector showing the relationship between the blending data, the provisional feature data obtained by replacing the one or more missing values in the feature data with one or more provisional values, and the characteristic data, by performing a second regression analysis based on the blending data, the provisional feature data, and the characteristic data, is repeated while changing the one or more provisional values, to identify the one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, The steps include: replacing the one or more missing values in the feature data with the one or more identified provisional values to complete the feature data; Information processing methods including (Note 6) A step of obtaining formulation data which includes an identifier for a formulation process and multiple data records that show the correspondence between the identifier for the formulation process and the element value for each of the multiple raw materials in that formulation process, The steps include: obtaining feature data that shows one or more features of each of the aforementioned multiple raw materials and includes one or more missing values; A step of obtaining characteristic data showing the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the formulation data, A step of calculating a first regression coefficient vector showing the relationship between the blending data and the characteristic data by first regression analysis based on the blending data and the characteristic data, The process of calculating a second regression coefficient vector showing the relationship between the blending data, the provisional feature data obtained by replacing the one or more missing values in the feature data with one or more provisional values, and the characteristic data, by performing a second regression analysis based on the blending data, the provisional feature data, and the characteristic data, is repeated while changing the one or more provisional values, to identify the one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, The steps include: replacing the one or more missing values in the feature data with the one or more identified provisional values to complete the feature data; An information processing program that causes a computer to execute something.
[0065] According to appendices 1, 5, and 6, the calculation of the second regression coefficient vector, which shows the relationship between the blending data, feature data, and characteristic data, is repeated while changing the provisional values used to fill in missing values in the feature data. Then, the missing values in the feature data are filled in by provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector, which is calculated without being affected by missing features, satisfies predetermined conditions. The relationship between the blending of multiple raw materials and the characteristic values of the composition can be determined by both a first regression analysis that does not consider features and a second regression analysis that does consider features, and by filling in missing values based on the difference between these two types of regression analyses, the feature data showing the characteristics of individual raw materials can be appropriately supplemented.
[0066] The first regression analysis, which can be expressed by equation (1) above, can predict the characteristic values of the composition with higher accuracy than the second regression analysis, which can be expressed by equation (2) above. However, the first regression analysis cannot make predictions considering raw materials that are not present in the blending ratio data. On the other hand, since the characteristics of raw materials can be calculated based on the Ideal Mixing Rule, the second regression analysis can predict the characteristics of the composition when new raw materials are used. However, the second regression analysis cannot be performed unless all the characteristics of the raw materials are obtained. According to appendices 1, 5, and 6, the feature data is complemented by comparing the two regression coefficient vectors corresponding to these two types of regression analyses, each of which has its own advantages and disadvantages, so that the feature data representing the characteristics of individual raw materials can be appropriately complemented.
[0067] According to Appendix 2, missing values in the feature data are filled in by provisional values corresponding to the second regression coefficient vector, which has the smallest local difference from the first regression coefficient vector. Therefore, it becomes possible to more accurately interpolate the data representing the features of individual raw materials.
[0068] According to Appendix 3, the second regression coefficient vector is recalculated based on the interpolated feature data, so a more accurate second regression coefficient vector can be obtained.
[0069] According to Appendix 4, new characteristic data reflecting the new raw material is calculated based on new formulation data with added element values for the new raw material, new feature data obtained by adding features for the new raw material to the interpolated feature data, and a recalculated second regression coefficient vector. This process makes it possible to predict the characteristic values of the composition generated by the formulation process with the new raw material added, even if there are missing values in the original feature data. [Explanation of Symbols]
[0070] 10... Information processing system, 11... First data acquisition unit, 12... First regression analysis unit, 13... Second regression analysis unit, 14... Search unit, 15... Complementary unit, 16... Second data acquisition unit, 17... Prediction unit, 20... Database, 21... Blending data, 22... Feature data, 23... Characteristic data, 30... User terminal.
Claims
1. Equipped with at least one processor, The at least one processor, The system retrieves formulation data that includes multiple data records showing the correspondence between the identifier of the formulation process and the element value in that formulation process for each of the multiple raw materials. The feature data obtained shows one or more feature quantities for each of the aforementioned multiple raw materials and includes one or more missing values. Characteristic data is obtained that shows the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the formulation data. A first regression coefficient vector showing the relationship between the blending data and the characteristic data is calculated by first regression analysis based on the blending data and the characteristic data. The process of calculating a second regression coefficient vector showing the relationship between the blending data, the provisional feature data obtained by replacing the one or more missing values in the feature data with one or more provisional values, and the characteristic data by performing a second regression analysis based on the blending data, the provisional feature data, and the characteristic data is repeatedly performed while changing the one or more provisional values, thereby identifying the one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, The feature data is supplemented by replacing the one or more missing values in the feature data with the one or more identified provisional values. Information processing system.
2. The at least one processor identifies one or more provisional values corresponding to the second regression coefficient vector whose difference from the first regression coefficient vector is locally minimized. The information processing system according to claim 1.
3. The at least one processor performs the second regression analysis based on the formulation data, the characteristic data, and the complementary feature data to recalculate the second regression coefficient vector. The information processing system according to claim 1 or 2.
4. The at least one processor, New formulation data is obtained by adding the respective element values of at least one new raw material to each of the plurality of data records of the formulation data. New feature data is obtained by adding the one or more feature quantities for each of the at least one new raw materials to the complementary feature data. Based on the new formulation data, the new feature data, and the recalculated second regression coefficient vector, new characteristic data is calculated that shows the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the new formulation data. The information processing system according to claim 3.
5. An information processing method performed by an information processing system comprising at least one processor, A step of obtaining formulation data which includes an identifier for a formulation process and multiple data records that show the correspondence between the identifier for the formulation process and the element value for each of the multiple raw materials in that formulation process, The steps include obtaining feature data that shows one or more feature quantities for each of the aforementioned multiple raw materials and includes one or more missing values, A step of obtaining characteristic data showing the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the formulation data, A step of calculating a first regression coefficient vector showing the relationship between the blending data and the characteristic data by first regression analysis based on the blending data and the characteristic data, The process of calculating a second regression coefficient vector showing the relationship between the blending data, the provisional feature data obtained by replacing the one or more missing values in the feature data with one or more provisional values, and the characteristic data, by performing a second regression analysis based on the blending data, the provisional feature data, and the characteristic data, is repeatedly performed while changing the one or more provisional values, thereby identifying the one or more provisional values that correspond to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, The steps include: replacing the one or more missing values in the feature data with the one or more identified provisional values to complete the feature data; Information processing methods including
6. A step of obtaining formulation data which includes an identifier for a formulation process and multiple data records that show the correspondence between the identifier for the formulation process and the element value for each of the multiple raw materials in that formulation process, The steps include obtaining feature data that shows one or more feature quantities for each of the aforementioned multiple raw materials and includes one or more missing values, A step of obtaining characteristic data showing the characteristic values of each of the multiple compositions corresponding to the multiple formulation processes indicated by the formulation data, A step of calculating a first regression coefficient vector showing the relationship between the blending data and the characteristic data by first regression analysis based on the blending data and the characteristic data, The process of calculating a second regression coefficient vector showing the relationship between the blending data, the provisional feature data obtained by replacing the one or more missing values in the feature data with one or more provisional values, and the characteristic data, by performing a second regression analysis based on the blending data, the provisional feature data, and the characteristic data, is repeatedly performed while changing the one or more provisional values, thereby identifying the one or more provisional values that correspond to the second regression coefficient vector whose difference from the first regression coefficient vector satisfies predetermined conditions, The steps include: replacing the one or more missing values in the feature data with the one or more identified provisional values to complete the feature data; An information processing program that causes a computer to execute something.