Information Processing Apparatus, Information Processing Method, and Program

The vectorized prediction formula and data thinning techniques enhance the speed and efficiency of calculating relative distances between data, addressing the computational inefficiencies of existing methods.

JP7708031B2Active Publication Date: 2025-07-15TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022129482
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2025-07-15
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

Existing prediction formulas for calculating relative distances between data require significant computational resources due to complex weight calculations, leading to slow processing times.

Method used

Implement a vectorized prediction formula and data thinning techniques to calculate relative distances, reducing the number of loop operations and eliminating data points with minimal contribution to the prediction, thereby accelerating the calculation process.

Benefits of technology

Achieves high-speed prediction processing by reducing computational load and memory consumption while maintaining non-linear model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708031000020
    Figure 0007708031000020
  • Figure 0007708031000021
    Figure 0007708031000021
  • Figure 0007708031000022
    Figure 0007708031000022
Patent Text Reader

Abstract

To realize high-speed processing in prediction processing using a prediction formula including a calculation of a relative distance among data.SOLUTION: An information processing device 10 includes: a data acquisition unit 101 configured to acquire data of an explanatory variable expressed as a T by m matrix and data of an objective variable expressed as a vector having T-pieces of components; and a prediction processing unit 102 configured to calculate a prediction value with regard to the objective variable, using a prescribed prediction formula including a calculation of a relative distance using data on a focused row and data of the other rows from among the data of the explanatory variable prediction formula, and the data acquired at the data acquisition unit 101. The prediction processing unit 102 calculates the prediction value by calculating using the prediction formula where the relative distance of each of the other rows is expressed by a vector, or through thinning-out the data of a row where the relative distance exceeds a prescribed threshold value from among the other rows.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program, and particularly relates to an information processing apparatus, an information processing method, and a program for calculating a predicted value.

Background Art

[0002] In the technical field of causal analysis for analyzing the causal relationship between data, a process of calculating a predicted value (estimated value) based on the relative distance between data may be performed. In relation to this, the causal relationship learning apparatus disclosed in Patent Document 1 includes a correlation determination unit that determines the correlation between measurement values measured by two sensors, and an estimation unit that determines the causal relationship between the two sensors by estimating the cause measurement value from the result measurement value when the correlation is lower than a predetermined standard. Further, as a related technique, Non-Patent Document 1 discloses causal analysis using an NPMR (Non-Parametric Multiplicative Regression) model.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Since the calculation formula of the weight based on the relative distance between data in the prediction formula used in Patent Document 1 is relatively simple, the expression power as a non-linear model is poor. In order to solve this problem, for example, it is conceivable to use a prediction formula with an increased number of weights or to use a product operation in the calculation of the weights in the prediction formula. However, when such a prediction formula is used, the amount of calculation related to the weights increases, and a large amount of calculation time is required.

[0006] The present disclosure has been made against the background of the above circumstances, and an object thereof is to provide an information processing apparatus, an information processing method, and a program capable of realizing high-speed processing in prediction processing using a prediction formula including calculation of the relative distance between data.

Means for Solving the Problems

[0007] One aspect of the present disclosure for achieving the above object includes a data acquisition unit that acquires data of explanatory variables represented as a matrix of T (where T is an integer of 2 or more) rows and m (where m is an integer of 1 or more) columns, and data of an objective variable represented as a vector having T components, a predetermined prediction formula including calculation of the relative distance between the data of the row of interest and the data of each other row among the data of the explanatory variables, and a prediction processing unit that calculates a predicted value for the objective variable using the data acquired by the data acquisition unit. The prediction processing unit calculates the predicted value by using the prediction formula in which the relative distance from each of the other rows is represented by a vector, or by performing thinning-out calculation on the data of the rows in which the relative distance from each of the other rows is equal to or greater than a predetermined threshold value.

[0008] Also, another aspect of the present disclosure for achieving the above object is that an information processing apparatus acquires data of explanatory variables represented as a matrix of T (where T is an integer of 2 or more) rows and m (where m is an integer of 1 or more) columns, and data of objective variables represented as a vector having T components, and uses a predetermined prediction formula including calculation of relative distances between data of a target row and data of each other row among the data of the explanatory variables, and the acquired data to calculate a predicted value for the objective variable. In the step of calculating the predicted value, the predicted value is calculated by using the prediction formula in which the relative distances from each of the other rows are represented by vectors, or by performing calculation with downsampling data of rows in which the relative distances from each of the other rows are equal to or greater than a predetermined threshold value. This is an information processing method.

[0009] Also, another aspect of the present disclosure for achieving the above object is to cause a computer to execute a data acquisition step of acquiring data of explanatory variables represented as a matrix of T (where T is an integer of 2 or more) rows and m (where m is an integer of 1 or more) columns, and data of objective variables represented as a vector having T components, a prediction processing step of calculating a predicted value for the objective variable by using a predetermined prediction formula including calculation of relative distances between data of a target row and data of each other row among the data of the explanatory variables, and the data acquired in the data acquisition step. In the prediction processing step, the predicted value is calculated by using the prediction formula in which the relative distances from each of the other rows are represented by vectors, or by performing calculation with downsampling data of rows in which the relative distances from each of the other rows are equal to or greater than a predetermined threshold value. This is a program.

Advantages of the Invention

[0010] According to the present disclosure, it is possible to provide an information processing apparatus, an information processing method, and a program capable of realizing high-speed processing in prediction processing using a prediction formula including calculation of relative distances between data.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing an example of the configuration of an information processing apparatus 10 according to an embodiment. The information processing apparatus 10 performs a process of predicting data of an objective variable using data of an explanatory variable and data of an objective variable in order to perform a causal analysis for analyzing a causal relationship between the data of the explanatory variable and the data of the objective variable, for example. However, the prediction process described later may be performed for any other purpose. The information processing apparatus 10 includes, as an example, a data acquisition unit 101 and a prediction processing unit 102 as shown in FIG. 1.

[0013] The data acquisition unit 101 receives the input of data used for prediction processing. The data acquisition unit 101 receives the input of the data series of the explanatory variables and the data series of the target variables. In the present embodiment, both the data of the explanatory variables and the data of the target variables are time series data, but they do not necessarily have to be time series data. In particular, the data acquisition unit 101 acquires the data of the explanatory variables represented as a matrix of T rows and m columns, and the data of the target variables represented as a vector having T components.

[0014] Specifically, the data acquisition unit 101 acquires the data of the explanatory variables represented by the following formula (1).

[0015] <Formula (1)>

Number

[0016] In addition, the data acquisition unit 101 acquires the data of the target variables represented by the following formula (2).

[0017] <Formula (2)>

Number

[0018] Here, T is an integer of 2 or more, and m is an integer of 1 or more. Note that m represents the number of types of data used as explanatory variables. That is, m predictors are used for prediction processing. Also, T is the number of data in each data series (each time series data). Taking a specific example, for example, the data series of the first column of X is the time series data of temperature, the data series of the second column is the time series data of precipitation, and the data series indicated by Y is the time series data of electricity consumption. In this case, the prediction processing unit 102 described later calculates the predicted value of electricity consumption based on the time series data of temperature, the time series data of precipitation, and the time series data of electricity consumption. Note that these are only examples, and the specific data of each data series is not limited to the above example.

[0019] Note that the data acquisition unit 101 may set the data of the explanatory variables represented as a matrix of T rows and m columns and the data of the objective variable represented as a vector having T components from the input data series. That is, the data acquisition unit 101 is the component x of the matrix X i,j (where i is an integer from 1 to T, and j is an integer from 1 to m) and the component y of the vector Y t (where t is an integer from 1 to T), and may receive an input for specifying each value, or may set these values from the input data according to the user's instructions. Further, the data acquisition unit 101 may normalize the data in advance so that the standard deviation σ j becomes a predetermined value (for example, 1).

[0020] The prediction processing unit 102 calculates a predicted value for the objective variable using a predetermined prediction formula including calculation of the relative distance (difference) between the data of the row of interest (the data of row t) and the data of each other row (the data of each row other than row t) among the data of the explanatory variables, and the data (X and Y) acquired by the data acquisition unit 101. More specifically, the predetermined prediction formula is a prediction formula including weight calculation based on the relative distance described above. In this predetermined prediction formula, the weight is set such that the greater the relative distance, the smaller the weight for the value of the objective variable.

[0021] In the present embodiment, the prediction processing unit 102 calculates a predicted value using, as the above-described predetermined prediction formula, specifically, for example, a prediction formula based on the prediction formula of the NPMR model disclosed in Non-Patent Document 1 (a prediction formula obtained by modifying the prediction formula of the NPMR model). Here, the prediction formula of the NPMR model disclosed in Non-Patent Document 1 is represented as the following formula (3). Here, in the prediction formula of formula (3), the weight w i,j applied to the value of the objective variable is represented by formula (4). Note that σ j is the standard deviation of the data in the j-th column of the explanatory variable X. Also, in the present disclosure, a variable representing a predicted value that appears on the left side of formula (3) and the like may be denoted as y t ^.

[0022] <Formula (3)>

number

[0023] <Formula (4)>

number

[0024] In this way, in NPMR, the t-th component of the data series of the response variable Y, y t The predicted value of y t ^ is calculated by weighting each component of the data series of the response variable Y other than the t-th component. As is clear from the above prediction formula, in the NPMR model, all components of the response variable Y, that is, y1 to y T If all of the data are real data (i.e., none of them are null values), then by setting the value of t from 1 to T, T Therefore, all of the real data y1 to y T and the predicted values y1^ to y T ^ The accuracy of the prediction using the explanatory variable X can be seen by comparing with each other. On the other hand, the predicted value y t To calculate ^, we use y t is not used. y1^ to y T ^ Any one of the values (e.g., unknown y t ), all of the T components of the objective variable Y do not necessarily have to be real data. That is, the data acquiring unit 101 may acquire a vector Y including one component that is a null value (missing value) as a data string of the objective variable Y. In other words, a predetermined t-th component of a vector having T components acquired by the data acquiring unit 101 may be a null value.

[0025] As can be seen from Equation (4), the weights are calculated by computing the relative distance (difference) between the data of the target row (the data of row t) and the data of each other row (the data of each row other than row t) among the data of the explanatory variables. Although the NPMR model can express sufficient non-linearity compared to the models shown in Patent Document 1 and the like, the amount of calculation related to the weights is large, and the processing time required to calculate the predicted value becomes long. Therefore, in the present embodiment, in order to achieve high-speed processing, the prediction processing unit 102 uses a prediction formula in which the relative distance is represented by a vector, and calculates the predicted value by thinning out the data of the rows where the relative distance is equal to or greater than a predetermined threshold value. That is, the calculation of the predicted value by the prediction processing unit 102 has the features of vectorization of the relative distance and thinning out of the data. Note that although the present embodiment has both of these features, the prediction processing unit 102 may calculate the predicted value by adopting only one of these features.

[0026] First, the vectorization of the relative distance will be described. The inventors have found that the prediction formula of the NPMR model can be transformed into the following prediction formula (Equation (5)) in which the relative distance is represented by a vector. In the present disclosure, the prediction formula subjected to such transformation will be referred to as a vectorized prediction formula.

[0027] <Equation (5)>

Number

[0028] Here, the subscript -t means that the component of the t-th (t-th row) has been removed. For example, for the vector y shown in the following Equation (6), y -t is defined as in the following Equation (6).

[0029] <Equation (6)>

Number

[0030] Therefore, in Equation (5), the vector y-t is a vector having T - 1 components obtained by removing the t-th component from the data sequence of the target variable represented as a vector having T components.

[0031] Also, in Equation (5), the vector Δx -t is a vector having T - 1 components whose components are the relative distances between the data set of the rows other than the t-th row and the data set of the t-th row in the data of the explanatory variables. That is, the vector Δx -t is a vector having T - 1 components obtained by removing the t-th component from T vectors whose i-th component is the relative distance between the data set of the i-th row and the data set of the t-th row in the data of the explanatory variables. The relative distance between the data set of the rows other than the t-th row and the data set of the t-th row in the data of the explanatory variables can also be said to be the relative distance between the vector represented by the m components (x i,1 , x i,2 , ···, x i,m ) of the i-th row which is a row other than the t-th row in the data of the explanatory variables and the vector represented by the m components (x t,1 , x t,2 , ···, x t,m ) of the t-th row in the data of the explanatory variables.

[0032] As a simple example, when m = 1, the vector Δx -t is shown as follows, for example.

[0033] <Equation (7)>

Number

[0034] Also, in the equations shown in the present disclosure such as Equation (5), the operator "·" is an operator indicating the inner product (inner product operator), and the operator "o" is an operator indicating the Hadamard product (Hadamard product operator). Also, in the equations shown in the present disclosure such as Equation (5), the function "exp" is an exponential function, and the function "Exp" is an exponential function acting on each component of the vector. That is, the function "Exp" means applying the function "exp" to each component respectively.

[0035] FIG. 2 is a diagram showing an example of the source code of a program when implementing the operation represented by Equation (3), which is the original NPMR prediction equation, by a computer program. Further, FIG. 3 is a diagram showing an example of the source code of a program when implementing the operation represented by Equation (5), which is the vectorized prediction equation, by a computer program. The source codes shown in FIGS. 2 and 3 show an example when using Julia as the programming language, but the operations of the prediction equation may be implemented by other programming languages. As shown in FIG. 3, when using the vectorized prediction equation, since it is possible to implement it with a code that performs vector operations, the number of executions of the calculation by loop processing is reduced compared to the code that implements the operation of the prediction equation before vectorization (the original NPMR prediction equation), and thus it is possible to calculate the predicted value with a code. Therefore, by using the vectorized prediction equation, high-speed processing can be realized in the prediction process of the predicted value. Note that the speedup of the processing when using the vectorized prediction equation can also be experimentally understood from the results shown in FIG. 7 and the like described later.

[0036] Next, the data thinning will be described. As can be seen from Equations (3) and (4), in a prediction equation in which weights are calculated based on relative distances, such as the NPMR prediction equation, the component of the target variable Y corresponding to the row with a large relative distance from the data in the t-th row of the explanatory variable X has a low contribution degree in the calculation of the predicted value y t ^. Therefore, in the present embodiment, in order to further speed up the calculation process of the predicted value, the prediction processing unit 102 calculates the predicted value by thinning out the data of the rows whose relative distance is equal to or greater than a predetermined threshold. Specifically, the prediction processing unit 102 calculates the predicted value y t ^ using the prediction equation shown in the following Equation (8). Equation (8) is a prediction equation obtained by further transforming the vectorized prediction equation shown in Equation (5). Note that the predicted value y t ^ calculated by Equation (8) is an approximate value of the predicted value calculated by the prediction equations shown in Equations (3) and (5).

[0037] <Equation (8)>

Number

[0038] Δx in the above formula -t rec is defined by the following formula (9), and y -t rec is defined by the following formula (10).

[0039] <Formula (9)>

Number

[0040] <Formula (10)>

Number

[0041] Here, R(t) is a vector having T-1 components generated based on the relative distance between the data set of the rows other than the t-th row and the data set of the t-th row in the data of the explanatory variables. R(t) is used as a filter for extracting components whose relative distance is less than a predetermined threshold value, that is, components whose contribution degree in the calculation of the predicted value is equal to or higher than a predetermined standard. The prediction processing unit 102 generates R(t), which is a filter for extracting data of rows whose relative distance is less than the threshold value, in other words, a filter for extracting components whose contribution degree in the calculation of the predicted value is equal to or higher than a predetermined standard, based on the relative distance, and thins out the data of rows whose relative distance is equal to or higher than the threshold value using the filter. R(t) corresponds to the points plotted in a recurrence plot (see FIG. 4) representing the relative distance between the data of the explanatory variables, and is defined by the following formulas (11) and (12).

[0042] <Formula (11)>

Number

[0043] <Formula (12)>

Number

[0044] Note that in Formula (12), k = 1, 2, ···, t - 1, t + 1, ···, T. That is, k is an integer from 1 to T excluding t. Also, in Formula (12), d(x k , x t ) is the relative distance between the data set of the k-th row and the data set of the t-th row in the data of the explanatory variable X. Also, δ is the predetermined threshold value described above. As can be seen from the relationship between R(t) and the recurrence plot shown in FIG. 4, R(t) shown in Formula (11) is a vector representing the distribution of the recurrence plot at time t with 0 and 1.

[0045] As described above, the component of R(t) becomes 0 when the relative distance from the data set of the t-th row in the data of the explanatory variable X is greater than or equal to the threshold value δ, and becomes 1 when the relative distance is less than the threshold value δ. And, as shown in Formula (9), Δx -t is defined by the Hadamard product of R(t) having T - 1 components taking values of 0 or 1 and Δx -t rec . That is, Δx rec -t is a vector obtained by changing the values of the components of Δx -t , and is a vector in which the values of the components corresponding to the rows where the relative distance from the data of the t-th row in the data of the explanatory variable is greater than or equal to the threshold value are changed to 0. In other words, Δx rec -t is a vector obtained by changing the values of the components of Δx -t , and can also be said to be a vector in which the values of the components having a relative distance greater than or equal to the threshold value among the components of Δx -t are changed to 0. Similarly, as shown in Formula (10), y -t is defined by the Hadamard product of R(t) having T - 1 components taking values of 0 or 1 and y -t rec . That is, y -t rec is y -tA vector obtained by changing the values of the components, and for each component corresponding to a row whose relative distance from the data of the t-th row in the explanatory variable data is equal to or greater than a threshold value, the value is changed to 0. In other words, y -t rec is a vector obtained by changing the values of the components of y -t and can also be said to be a vector obtained by changing the values of the components having the same index (element number) as the components whose values are changed to 0 in Δx -t among the components of y rec -t Therefore, in contrast to the prediction formula shown in Equation (5), in the prediction formula shown in Equation (8), among the components of Δx -t a vector obtained by changing the values of the components whose relative distance is equal to or greater than the threshold value, that is, the components that do not substantially contribute to the prediction value calculation, to 0, and y -t a vector obtained by changing the values of the components that do not substantially contribute to the prediction value calculation among the components of y

[0046] Next, an example of the hardware configuration of the information processing apparatus 10 will be described. FIG. 5 is a block diagram showing an example of the hardware configuration of the information processing apparatus 10. As shown in FIG. 5, the information processing apparatus 10 includes an input / output interface 151, a memory 152, and a processor 153.

[0047] The input / output interface 151 is an interface for communicably connecting to other devices (for example, an input device or an output device, etc.) as necessary. For example, the input / output interface 151 may be used for the data acquisition unit 101 to acquire data, or may be used for the prediction processing unit 102 to output the prediction result.

[0048] The memory 152 is constituted by, for example, a combination of a volatile memory and a non-volatile memory. The memory 152 is used to store software (computer program) including one or more instructions executed by the processor 153, and data used for various processes.

[0049] The processor 153 reads and executes software (computer program) from the memory 152 to perform the processing of each component shown in FIG. 1. The processor 153 may be, for example, a microprocessor, an MPU (Micro Processor Unit), or a CPU (Central Processing Unit). The processor 153 may include a plurality of processors. Thus, the information processing apparatus 10 has functions as a computer.

[0050] The program includes a group of instructions (or software code) for causing a computer to perform one or more functions described in the embodiments when loaded into the computer. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, the computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD), or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray (registered trademark) disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may be transmitted on a transitory computer-readable medium or a communication medium. By way of example and not limitation, the transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.

[0051] Next, the processing of the information processing apparatus 10 will be described with reference to a flowchart. FIG. 6 is a flowchart showing an example of the operation of the information processing apparatus 10 according to the present embodiment.

[0052] In step S100, the data acquisition unit 101 sets data of explanatory variables and target variables. In the present embodiment, specifically, the data acquisition unit 101 sets data of explanatory variables represented as a matrix of T rows and m columns used for calculating predicted values, and data of target variables represented as a vector having T components. Next, in step S101, the prediction processing unit 102 calculates a predicted value using a prediction formula for speedup. In the present embodiment, the prediction processing unit 102 calculates a predicted value using a vectorized prediction formula transformed so as to be able to calculate by thinning out data, like the prediction formula shown in Equation (8). However, the calculation of the predicted value by the prediction processing unit 102 may calculate the predicted value by adopting only one of the features of vectorization of relative distances and the feature of thinning out data. Next, in step S102, the prediction processing unit 102 outputs the predicted value calculated in step S101, that is, the prediction result. The prediction processing unit 102 may output to an output device such as a display, or may output to a storage device to store the prediction result.

[0053] The above describes the embodiments. According to the information processing apparatus 10, the prediction processing unit 102 uses a prediction formula in which the relative distance between the data of the row of interest and the data of each other row among the data of the explanatory variables is represented by a vector, or thins out the data of the rows in which the relative distance is equal to or greater than a predetermined threshold among each of the other rows, and calculates a predicted value. In the former case, compared with the case where a prediction formula in which the relative distance is not represented by a vector is calculated by a computer program, the loop processing to be executed is reduced, so that high-speed processing can be realized. In the latter case, since the processing is performed by thinning out the data that does not substantially contribute to the calculation of the predicted value, high-speed processing can be realized. Also, the consumption of memory can be reduced. In particular, since the prediction processing unit 102 uses, as a prediction formula used for calculating a predicted value, a prediction formula based on the prediction formula of the NPMR model, prediction processing that sufficiently reflects non-linearity can be realized at high speed.

[0054] Next, experimental results regarding the high-speed processing by the prediction processing unit 102 are shown. The measurement environment used for this experiment is as follows. Operating system: macOS Catalina Processor: 2.8 GHz, quad-core, Intel Core i7 Memory: 16 GB, 2133Mhz, LPDDR3

[0055] Also, in the first experiment, two time series data defined as follows were used as test data, and the explanatory variable X and the target variable Y were set from these time series data. In Equation (13), e1(t) and e2(t) are white noise. In the test data shown in Equation (13), there is a non-linear causality from x2 to x1.

[0056] <Equation (13)>

Number

[0057] Also, in the second experiment, three time series data defined as follows were used as test data, and explanatory variable X and objective variable Y were set from these time series data. In Equation (14), e1(t), e2(t), and e3(t) are white noise. In the test data shown in Equation (14), there are non-linear causality from x1 to x2, non-linear causality from x1 to x3, and linear causality from x2 to x3.

[0058] <Equation (14)>

Number

[0059] Figure 7 is a graph showing the experimental results in the first experiment. Also, Figure 8 is a graph showing the experimental results in the second experiment. In Figures 7 and 8, the graph of the required time when analyzing using the original NPMR (the graph labeled "NPMR"), the graph of the required time when analyzing using the NPMR with vectorized relative distance (the graph labeled "Fast NPMR-1"), and the graph of the required time when analyzing using the NPMR combined with vectorized relative distance and data thinning (the graph labeled "Fast NPMR-2") are shown. As shown in Figures 7 and 8, it can be seen that the processing is speeded up by using the technology shown in this embodiment. In particular, the higher the number of data and the more complex the analysis target, the more remarkable the speedup.

[0060] Also, in the third experiment, the effect of data thinning was confirmed using a prediction formula different from the above-mentioned prediction formula. Here, a prediction formula using the kernel method was examined. Specifically, the prediction processing unit 102 performs processing using the prediction formula shown in the following Equation (16) in order to process the prediction formula shown in the following Equation (15) at high speed.

[0061] <Equation (15)>

Number

[0062] <Formula (16)>

number

[0063] In equation (16), the data used in the calculation is thinned out using the above-mentioned filter, that is, a filter for extracting components whose relative distance is less than a predetermined threshold (components whose contribution to the calculation of the predicted value is equal to or greater than a predetermined standard). Here, R(t) in equation (16) is defined as follows, and represents a set of indexes (row numbers) indicating components whose relative distance is less than a threshold δ. Note that in equation (17), D(t i ) indicates the data of the i-th row of the explanatory variable X set by the data acquisition unit 101, and D(t i ) is calculated as the relative distance between the explanatory variable X and the data D(t) of row t and compared with a threshold.

[0064] <Formula (17)>

number

[0065] The prediction processing unit 102 generates R(t) and performs the calculation of the prediction formula shown in Equation (16). FIG. 9 is a graph showing the experimental results of the third experiment. The test data used in the third experiment is the same as that in the first experiment. FIG. 9 shows a graph of the time required for analysis without thinning out data based on relative distance (graph labeled "normal version") and a graph of the time required for analysis after thinning out data based on relative distance (graph labeled "fast version"). As shown in FIG. 9, it can be seen that the processing speed is increased by thinning out data using a filter.

[0066] Note that the present invention is not limited to the above-described embodiments and can be appropriately modified without departing from the spirit thereof. For example, in the above-described embodiments, the information processing apparatus 10 has the functions of the data acquisition unit 101 and the prediction processing unit 102, but some or all of these functions may be implemented by other devices (e.g., a server, etc.). That is, the processing described in the above-described embodiments may be realized by a system composed of one or more devices.

Explanation of Reference Numerals

[0067] 10 Information processing apparatus 101 Data acquisition unit 102 Prediction processing unit 151 Input / output interface 152 Memory 153 Processor

Claims

1. A data acquisition unit that acquires data of explanatory variables represented as a matrix of T (where T is an integer of 2 or more) rows and m (where m is an integer of 1 or more) columns, and data of a target variable represented as a vector having T components; A prediction processing unit that calculates a predicted value for the target variable using a predetermined prediction formula including calculation of a relative distance between data of a row of interest and data of each other row among the data of the explanatory variables, and the data acquired by the data acquisition unit having The prediction processing unit calculates the predicted value by using the prediction formula in which the relative distance from each of the other rows is represented by a vector, or by performing calculation after thinning out data of rows in which the relative distance is equal to or greater than a predetermined threshold among each of the other rows An information processing apparatus.

2. The prediction processing unit generates a filter for extracting data of the rows in which the relative distance is less than the threshold based on the relative distance, and thins out data of the rows in which the relative distance is equal to or greater than the threshold using the filter The information processing apparatus according to claim 1.

3. The predetermined prediction formula is a prediction formula based on a prediction formula of an NPMR (Non-Parametric Multiplicative Regression) model The information processing apparatus according to claim 1 or 2.

4. The prediction processing unit calculates the predicted value by using the following formula as the prediction formula The information processing apparatus according to claim 3. However, In the following formula, 【Number 18】 is the predicted value, t corresponds to the row number of the row of interest, Δx rec -t is a vector in which the values of the components of a vector having T-1 components whose components are the relative distances between the data sets of the rows other than the t-th row and the data set of the t-th row in the data of the explanatory variable are changed, and the values of the respective components corresponding to the rows in which the relative distance is greater than or equal to the threshold value are changed to 0. y rec -t is a vector obtained by changing the values of the components of a vector having T - 1 components obtained by removing the t-th component from the target variable represented as a vector having T components, and is a vector in which the values of the respective components corresponding to the rows where the relative distance is greater than or equal to the threshold value are changed to 0, exp is an exponential function, Exp is an exponential function acting on each component of the vector. 【Number 19】

5. The data of the explanatory variables and the data of the target variable are time series data The information processing apparatus according to claim 1.

6. An information processing apparatus acquires data of explanatory variables represented as a matrix of T (where T is an integer of 2 or more) rows and m (where m is an integer of 1 or more) columns, and data of a target variable represented as a vector having T components, calculates a predicted value for the target variable using a predetermined prediction formula including calculation of a relative distance between data of a row of interest and data of each other row among the data of the explanatory variables, and the acquired data In the step of calculating the predicted value, the predicted value is calculated by using the prediction formula in which the relative distance from each of the other rows is represented by a vector, or by thinning out data of rows in which the relative distance from each of the other rows is equal to or greater than a predetermined threshold value and then performing the calculation. Information processing method. **Claim 7** A data acquisition step of acquiring data of explanatory variables represented as a matrix of T (where T is an integer of 2 or more) rows and m (where m is an integer of 1 or more) columns, and data of an objective variable represented as a vector having T components; A prediction processing step of calculating a predicted value for the objective variable by using a predetermined prediction formula including calculation of a relative distance between data of a row of interest and data of each of the other rows among the data of the explanatory variables, and the data acquired in the data acquisition step; to be executed by a computer; In the prediction processing step, the predicted value is calculated by using the prediction formula in which the relative distance from each of the other rows is represented by a vector, or by thinning out data of rows in which the relative distance from each of the other rows is equal to or greater than a predetermined threshold value and then performing the calculation. Program.

Citation Information

Patent Citations

  • High performance midsole and method for manufacturing same

    JP2021525637A

  • Three-dimensional objects and their formation

    US20180093418A1

  • High performance footbed and method of manufacturing same

    US20190365024A1

  • Causal relationship learning method, program, device, and abnormality analysis system

    WO2018198267A1