Estimation device and estimation method
The variational Bayesian method with ELBO approximation and steepest descent addresses computational inefficiencies in Gaussian process intensity function estimation, ensuring accurate and efficient results for high-dimensional covariates.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NT T INC
- Filing Date
- 2024-11-22
- Publication Date
- 2026-05-28
AI Technical Summary
Existing methods for estimating the intensity function of a Gaussian process, such as those described in Non-Patent Document 1, face significant computational challenges with high-dimensional covariates, leading to either excessive computational costs or reduced accuracy when attempting to efficiently calculate hyperparameters.
The proposed method employs a variational Bayesian approach that approximates the marginal likelihood as a theoretical lower bound (ELBO) and uses the steepest descent method to efficiently search for optimal hyperparameters, allowing for accurate and efficient estimation of the intensity function.
This approach enables high-accuracy and efficient estimation of the intensity function by analytically calculating derivatives, reducing computational time and maintaining estimation precision, even with high-dimensional covariates.
Smart Images

Figure JP2024041439_28052026_PF_FP_ABST
Abstract
Description
Estimation device and estimation method
[0001] This disclosure relates to an estimation device and an estimation method.
[0002] Consider a situation where, within a space where covariates are defined at any given point, a point event (hereinafter also referred to as "event") occurs probabilistically, and observational data of that event is obtained. For example, (space, covariates, observational data) = (latitude and longitude, crowd density, location of accident event), etc.
[0003] In this context, a method for estimating the probability of an event occurring for a given covariate from observed data is known, as described in Non-Patent Document 1. The method described in Non-Patent Document 1 estimates the posterior predictive distribution of the event probability from observed data based on a Bayesian method using a Gaussian process as the prior distribution. The probability of an event occurring is also called the intensity, and a function that takes covariates as input and outputs the intensity is also called the intensity function.
[0004] Kim et al., "Fast Bayesian estimation of point process intensity as function of covariates", Advances in Neural Information Processing Systems 35, 2022.
[0005] However, the method described in Non-Patent Document 1 can involve enormous computational costs when calculating the hyperparameters of a Gaussian process, which can result in a very long time required to estimate the intensity function. On the other hand, if a simplified method is used to efficiently calculate the hyperparameters of a Gaussian process, the accuracy of the intensity function estimation may decrease.
[0006] This disclosure is made in view of the above points and aims to efficiently estimate an accurate intensity function.
[0007] An estimation device according to one aspect of the present disclosure includes an input unit that receives observation data including covariates in a predetermined space and the location of events observed in the space, and kernel function data relating to a kernel function used in a Gaussian process, and an estimation unit that estimates the posterior predictive distribution followed by an intensity function that takes the covariates as input and outputs the probability of the event occurring, based on the observation data and the kernel function data, wherein the estimation unit estimates the posterior predictive distribution followed by the intensity function by approximating the theoretical lower bound of the marginal likelihood for the observation data as a function of the hyperparameters of the Gaussian process and maximizing the theoretical lower bound using the steepest descent method.
[0008] It is possible to efficiently estimate an accurate intensity function.
[0009] This figure shows an example of the hardware configuration of the estimation device according to this embodiment. This figure shows an example of the functional configuration of the estimation device according to this embodiment. This flowchart shows an example of the estimation process according to this embodiment.
[0010] One embodiment of the present invention will be described in detail below with reference to the drawings.
[0011] <Conventional Methods and Their Challenges> The method described in Non-Patent Document 1 estimates the posterior predictive distribution (i.e., intensity function) of the probability of an event occurring from observed data based on a Bayesian method that uses a Gaussian process as the prior distribution.
[0012] Here, the posterior predictive distribution of the intensity function estimated by the method described in Non-Patent Document 1 is highly dependent on the hyperparameters of the Gaussian process, which is the prior distribution. Therefore, in order to accurately estimate the posterior predictive distribution, it is necessary to set the hyperparameters appropriately.
[0013] A common method for setting hyperparameters in Bayesian methods is to search for hyperparameter values that maximize the marginal likelihood calculated from the observed data, and then adopt those values. In this case, a particularly efficient search method is the steepest descent method, which uses the derivative of the marginal likelihood with respect to the hyperparameter.
[0014] However, the method described in Non-Patent Document 1 uses the Laplace approximation when calculating the marginal likelihood, making it impossible to analytically calculate the differential value of the marginal likelihood with respect to the hyperparameters. For this reason, it is necessary to use a grid search, which is less efficient than the steepest descent method (i.e., a method that calculates the marginal likelihood for pre-specified candidate hyperparameter values and searches for the hyperparameter value that maximizes that marginal likelihood).
[0015] Generally, when the dimensionality of the covariates is large, the dimensionality of the hyperparameter space also increases. Furthermore, the computational cost of grid search increases exponentially with respect to the dimensionality of the search space. For this reason, when the dimensionality of the covariates is large, the method described in Non-Patent Document 1 requires an enormous amount of time to estimate the intensity function.
[0016] Furthermore, Non-Patent Document 1 describes employing a simplified grid search to reduce the time required to estimate the intensity function; however, in that case, the estimation accuracy may decrease for high-dimensional covariates.
[0017] <Proposed Method> To address the above-mentioned problems of conventional methods, we propose a method that can efficiently estimate an accurate intensity function.
[0018] The proposed method employs a variational Bayesian method that allows for the approximate derivation of the marginal likelihood as a function of hyperparameters. Specifically, the proposed method approximates the posterior predictive distribution of the intensity function using a Gaussian process, derives an analytical solution of the theoretical lower bound of the marginal likelihood (hereinafter also referred to as "ELBO (evidence lower bound)"), and adopts this analytical solution as an approximation of the marginal likelihood. Since ELBO is an explicit function of the hyperparameters of the Gaussian process, it is possible to analytically calculate the derivative with respect to the hyperparameters. Therefore, it becomes possible to explore the hyperparameters using an efficient steepest descent method, and as a result, the intensity function for covariates can be estimated with high accuracy and efficiency, thus solving the above-mentioned problems of conventional methods.
[0019] The following describes the estimation device 10 that efficiently estimates an accurate intensity function using the proposed method described above.
[0020] <Example of Hardware Configuration of Estimation Device 10> An example of the hardware configuration of the estimation device 10 according to this embodiment will be described with reference to Figure 1. Figure 1 is a diagram showing an example of the hardware configuration of the estimation device 10 according to this embodiment.
[0021] As shown in Figure 1, the estimation device 10 according to this embodiment includes an input device 101, a display device 102, an external interface 103, a communication interface 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these hardware components is connected to each other via a bus 109 for communication.
[0022] The input device 101 is, for example, a keyboard, mouse, touch panel, or physical button. The display device 102 is, for example, a display or display panel. The estimation device 10 does not necessarily have to have at least one of the input device 101 and the display device 102.
[0023] The external I / F 103 is an interface with external devices such as the recording medium 103a. Examples of recording media 103a include CDs (Compact Discs), DVDs (Digital Versatile Disks), SD memory cards (Secure Digital memory cards), and USB (Universal Serial Bus) memory cards.
[0024] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily holds programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can hold programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as, for example, an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory. The processor 108 is an arithmetic device such as, for example, a CPU (Central Processing Unit) or GPU (Graphics Processing Unit).
[0025] Note that the hardware configuration shown in FIG. 1 is an example, and the hardware configuration of the estimation device 10 is not limited to this. For example, the estimation device 10 may have a plurality of auxiliary storage devices 107 and a plurality of processors 108, may not have a part of the illustrated hardware, or may have various hardware other than the illustrated hardware.
[0026] <Functional configuration example of the estimation device 10> A functional configuration example of the estimation device 10 according to the present embodiment will be described while referring to FIG. 2. FIG. 2 is a diagram showing an example of the functional configuration of the estimation device 10 according to the present embodiment.
[0027] As shown in FIG. 2, the estimation device 10 according to the present embodiment includes an input unit 201, an intensity function estimation unit 202, and an output unit 203. Each of these units is realized, for example, by processing that one or more programs installed in the estimation device 10 cause the processor 108 or the like to execute. Here, event occurrence data 301, covariance data 302, and kernel function data 303 are given to the estimation device 10.
[0028] <<Input unit 201>> The input unit 201 inputs event occurrence data 301, covariance data 302, and kernel function data 303. Note that the event occurrence data 301 and the covariance data 302 are observed data.
[0029] - Event occurrence data 301 The event occurrence data 301 is data regarding the occurrence position of an event to be analyzed (that is, data representing the occurrence position of the observed event). The event occurrence data 301 includes the number of observation trials U, and the number of times N of the event observed in each trial u = 1,..., U ,
[0032] and the sequence {s n u |n = 1,..., N u} of the positions of the events observed in each trial u = 1,..., U, and the observation region T u of each trial u = 1,..., U. However, the number of dimensions of the space in which the event occurs is arbitrary. For example, if it is one-dimensional, it can be time, if it is two-dimensional, it can be geographical space, if it is three-dimensional, it can be spacetime, etc. Note that a trial may also be called a "test" or the like.
[0030] - Covariate data 302 The covariate data 302 is data representing the covariates observed within the observation region T u of each trial u = 1,..., U. The covariate data 302 includes a function y u that outputs the covariate with an arbitrary point t ∈ T u in the observation region T u as the input. Note that the number of dimensions of the covariate is arbitrary.
[0031] Here, in many application examples, covariates can only be obtained for a finite number of points within the observation region T u . In this case, for example, interpolation techniques such as regression models and kriging are used to construct the function y u (t).
[0032] ・Kernel function data 303 Kernel function data 303 is data relating to the kernel function used in the Gaussian process. The kernel function data 303 includes the specification of the kernel function and the specification of the values of the parameters included in that kernel function (hereinafter also referred to as "hyperparameters"). Here, the kernel function is a function that determines the smoothness of the function being modeled in the Gaussian process. In this embodiment, the function being modeled is the intensity function for covariates (that is, a function that takes covariates as input and outputs the probability of event occurrence).
[0033] Let k(y', y) be the kernel function for any two points (y', y) in the covariate space. In the following, we assume that the kernel function k(y', y) is a function that can be expressed as the inner product of finite-dimensional feature vectors. That is, the feature vectors are represented by the following equation (1).
[0034]
[0035] In this case, the kernel function k(y', y) is represented by the following equation (2).
[0036]
[0037] However, L is the number of dimensions of the feature vector, η is the parameter of the feature vector, and the feature vector φ(y|η) is represented as a column vector. Furthermore, the following represents the transpose operation.
[0038]
[0039] The feature vector φ(y|η) is a vector that represents the characteristics of the covariate y. The feature vector φ(y|η) may also be called the feature quantity of the covariate y.
[0040] <<Intensity Function Estimation Unit 202>> The intensity function estimation unit 202 estimates the intensity function for the covariates based on the event occurrence data 301, the covariate data 302, and the kernel function data 303. Hereinafter, the intensity function for the covariates is given by λ(y) = x 2 Let's represent it as (y).
[0041] - Intensity function x for covariates 2 Method for estimating (y): First, the intensity function x 2 The square root of (y) (hereinafter referred to as "x(y)") is an element of the feature vector φ(y|η). l Let's assume it can be expressed as a linear combination of (y | η). That is, the intensity function x 2 Assume that the square root x(y) of (y) is expressed by the following equation (3).
[0042]
[0043] However, α is a weight coefficient vector represented as follows:
[0044]
[0045] Furthermore, the prior distribution p(α) of the weight coefficient vector α is a multivariate normal distribution N(α | 0, κ) with a mean of zero and a diagonal covariance matrix. -1 I L We assume that it follows the formula (4) below.
[0046]
[0047] However, κ is a positive scalar value, I L This is the L×L identity matrix.
[0048] Next, we assume that the posterior distribution q(α) of the weight coefficient vector α follows a multivariate normal distribution N(α|m, S) with mean m and covariance matrix S. That is, we assume the following equation (5).
[0049]
[0050] However, m is an L-dimensional vector and S is an L×L positive definite matrix.
[0051] In this case, the theoretical lower bound of the marginal likelihood ELBO for the observed data (i.e., event occurrence data 301 and covariate data 302) is given by the following equation (6).
[0052]
[0053] However, tr(•) is the trace of the matrix, and θ is the one-dimensional vector formed from all parameters (i.e., η, κ, m, S). Also, C is a constant, and μ(•) and σ 2 (•) is represented by the following equations (7) and (8).
[0054]
[0055] Furthermore, G(z) can be expressed using a function F(a, b, z) called a confluent hypergeometric function by the following equation (9).
[0056]
[0057] The converging hypergeometric function F(a, b, z) is a function represented by the following equation (10).
[0058]
[0059] However, (a) 0 = 1, (a) k = a(a+1)(a+2)...(a+k-1).
[0060] Based on the above facts, the intensity function estimation unit 202 calculates the intensity function x according to the following procedure 1 to 2. 2 (y) can be estimated.
[0061] Procedure 1: The intensity function estimation unit 202 searches for the parameter θ that maximizes the ELBO shown in equation (6) above. At this time, since the analytical derivative of ELBO can be obtained with respect to the parameter θ, the intensity function estimation unit 202 searches for the parameter θ that maximizes ELBO using an efficient search algorithm classified as the steepest descent method, such as gradient descent or mini-batch gradient descent. Hereinafter, the value of parameter θ obtained by this search will be called the "estimated value" and will be expressed as follows.
[0062]
[0063] In the text of this specification, the hat symbol "^" representing an estimated value shall be placed immediately before the value, not directly above it. For example, in the text of this specification, the estimated value of the parameter θ shown in equation 12 above shall be written as "^θ".
[0064] Step 2: Under the estimated value θ, the posterior distribution q(α) of the weight coefficient vector α follows a multivariate normal distribution N(α|^m,^S) with mean ^m and covariance matrix ^S. That is, the following equation (11) is obtained.
[0065]
[0066] At this time, the intensity function x for each covariate y 2 The posterior probability distribution that (y) follows is given by the scale parameter ξ 1 (y) and shape parameter ξ 2 (y) is calculated as a gamma distribution given by the following equations (12) and (13).
[0067]
[0068] However, ^μ(y) and ^σ 2 (y) is expressed by the following equations (14) and (15), respectively.
[0069]
[0070] Therefore, the intensity function estimation unit 202 uses the scale parameter ξ shown in equation (12) above. 1 Shape parameter ξ shown in (y) and (13) 2 A gamma distribution with (y) and the intensity function x 2 The posterior probability distribution of (y) (i.e., the intensity function x) 2 It can be calculated as the posterior predictive distribution of (y). This allows the scale parameter ξ shown in equation (12) above to be calculated for any covariate y. 1 Shape parameter ξ shown in (y) and (13) 2 The function that outputs the values of the gamma distribution with (y) is the intensity function x 2 (y) is obtained.
[0071] <<Output Unit 203>> The output unit 203 outputs the intensity function x estimated by the intensity function estimation unit 202. 2The intensity function data 304 containing (y) is output to a predetermined output destination. Note that the predetermined output destination is not limited to a specific destination, and any output destination can be set in advance. For example, a display device 102 such as a display, a storage area such as an auxiliary storage device 107, other devices or equipment, programs, etc. can be set as the output destination.
[0072] <Example of Estimation Process> An example of the estimation process according to this embodiment will be explained with reference to Figure 3. Figure 3 is a flowchart of an example of the estimation process according to this embodiment.
[0073] The input unit 201 receives event occurrence data 301, covariate data 302, and kernel function data 303 (step S101).
[0074] The intensity function estimation unit 202 estimates the intensity function for the covariates based on the event occurrence data 301, the covariate data 302, and the kernel function data 303 (step S102). That is, the intensity function estimation unit 202 uses the event occurrence data 301, the covariate data 302, and the kernel function data 303 to formulate the ELBO shown in equation (6) above, and then performs the intensity function x according to steps 1 and 2 above. 2 Estimate (y).
[0075] The output unit 203 outputs the intensity function x estimated in step S102 above. 2 The intensity function data 304 containing (y) is output to a predetermined output destination (step S103).
[0076] <Summary> As described above, the estimation device 10 according to this embodiment approximates the posterior predictive distribution of the intensity function with a Gaussian process, and then adopts the analytical solution of the theoretical lower bound of the marginal likelihood, ELBO, as an approximation of the marginal likelihood. Furthermore, the estimation device 10 according to this embodiment utilizes the steepest descent method when calculating the analytical solution of ELBO. This makes it possible to efficiently search for the hyperparameters of the Gaussian process, and as a result, the intensity function for covariates can be estimated with high accuracy and efficiency.
[0077] Furthermore, the estimation device 10 according to this embodiment can be applied to various fields where the estimation of intensity functions is required. For example, it can be applied to analyzing the probability of events such as accidents occurring in a geographical space, using crowd density as a covariate.
[0078] The present invention is not limited to the embodiments specifically disclosed above, and various modifications, changes, and combinations with known technologies are possible without departing from the spirit of the claims.
[0079] 10 Estimation device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Input unit 202 Intensity function estimation unit 203 Output unit 301 Event occurrence data 302 Covariate data 303 Kernel function data 304 Intensity function data
Claims
1. An estimation device comprising: an input unit that takes as input the covariates in a predetermined space, the location of events observed in the space, and kernel function data relating to a kernel function used in a Gaussian process; and an estimation unit that estimates the posterior predictive distribution followed by an intensity function that takes the covariates as input and outputs the probability of the event occurring, based on the observation data and the kernel function data, wherein the estimation unit estimates the posterior predictive distribution followed by the intensity function by approximating the theoretical lower bound of the marginal likelihood for the observation data as a function of the hyperparameters of the Gaussian process and maximizing the theoretical lower bound using the steepest descent method.
2. The estimation device according to claim 1, wherein the estimation unit calculates estimated values of hyperparameters that maximize the theoretical lower limit using the steepest descent method, and estimates the gamma distribution calculated under the estimated values as the posterior predictive distribution followed by the intensity function.
3. The estimation device according to claim 2, wherein the estimation unit calculates a scale parameter and a shape parameter from the estimated value, and estimates the gamma distribution having the scale parameter and the shape parameter as the posterior predictive distribution to which the intensity function follows.
4. An estimation method comprising: an input procedure in which observation data including covariates in a predetermined space and the location of events observed in the space, and kernel function data relating to a kernel function used for a Gaussian process are input; and an estimation procedure in which a computer performs the following steps to estimate the posterior predictive distribution followed by an intensity function that outputs the probability of the event occurring with the covariates as input, based on the observation data and the kernel function data, wherein the estimation procedure approximates the theoretical lower bound of the marginal likelihood for the observation data as a function of the hyperparameters of the Gaussian process, and estimates the posterior predictive distribution followed by the intensity function by maximizing the theoretical lower bound using the steepest descent method.