Information processing device, inference device, information processing method, and program
Patent Information
- Application Number
- PCT/JP2025/012895
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012895_01102026_PF_FP_ABST
Abstract
Description
Information processing device, inference device, information processing method, and program
[0001] This disclosure pertains to the Ministry of Internal Affairs and Communications' "Research and Development of Elemental Technologies for User-Optimized Data Provision of Remote Sensing Technology" for fiscal year 2024, and relates to information processing devices, inference devices, information processing methods, and programs.
[0002] Various mathematical models are applied in a wide range of fields. In implementing the processing in these mathematical models, it is sometimes necessary to represent relationships between variables, particularly adjacent elements in the target data, such as spatial and temporal correlations, in order to more accurately represent the underlying data. However, as the number of variables increases, the size of the covariance matrix grows, increasing the cost of operations such as inverse matrices and determinants, often making it computationally difficult.
[0003] Reflecting spatial or temporal correlations in probabilistic models is difficult, and independence has often been assumed without considering correlations between variables; however, this approach is not always applicable.
[0004] DG Tzikas, et. al., “The Variational Approximation for Bayesian Inference,” IEEE Signal Processing Magazine 25(6), 131-146, 2008
[0005] One of the non-limiting problems that embodiments of this disclosure aim to solve is to improve the accuracy of operations that reduce the computational complexity of the covariance matrix.
[0006] According to one embodiment, the information processing device comprises at least one memory and at least one processor. The at least one processor estimates the data of interest based on a probability distribution defined by a probability model expressed using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest.
[0007] A flowchart showing an example of processing in an information processing device according to one embodiment. A block diagram showing an example of hardware implementation in an information processing device according to one embodiment.
[0008] The problems to be solved by the embodiments of this disclosure are not limited to those described above, but may also include, as examples of several other problems, problems corresponding to the effects described in the embodiments. That is, any problem corresponding to at least one of the effects described in the description of the embodiments of this disclosure may be the problem to be solved in this disclosure. For example, efficiently calculating inverse matrices, determinants, operations between covariance matrices and quadratic forms, matrix and vector products between vectors, and efficiently calculating expected values using probability distributions of discrete variables that have mutual inverses may be included as at least part of the problems to be solved.
[0009] Embodiments of the present invention will be described below with reference to the drawings. The drawings and descriptions of the embodiments are provided as examples only and do not limit the present invention.
[0010] Figure 1 is a flowchart showing an example of the training process for a model used in a segmentation device, as an example of mathematical processing according to one embodiment. In this embodiment, segmentation is described as an example, but the same form can be applied to adjacent elements, or variables, in the data to be processed in various mathematical models. For example, "image" in the following form can be replaced with the data to be processed, and "segmentation" can be replaced with processing of data in a mathematical model that uses a probabilistic model.
[0011] The information processing device, acting as a training device, first initializes its parameters (S100). Parameter initialization may include setting hyperparameters and model settings. By initializing the above parameters, the training device defines a model for use in segmentation.
[0012] Note that the initialization of the various parameters mentioned above can be omitted as appropriate depending on the purpose of the training.
[0013] For example, when creating a new model, the information processing device can initialize all of the above parameters.
[0014] For example, when performing retraining such as fine-tuning, the information processing device may set some or all of the model's parameters to already learned parameters.
[0015] For example, when performing retraining such as transfer learning, the information processing device can set some of the model's parameters to already learned parameters and initialize the necessary intermediate layers and / or output layers with newly learned parameters.
[0016] The information processing device can also handle model compression processes such as pruning and distillation, and in these cases, it can set and / or initialize parameters for constructing an appropriate model.
[0017] More specifically, the information processing device may, for example, initialize all elements of the model parameters θ depending on the purpose, or it may set some elements to some value (for example, a value related to an already trained model) and then initialize the other elements.
[0018] The training device acquires training data to input into the model (S102). This training data may include, for example, an input image (an image to be segmented) and data to which segmentation labels have been assigned to the image. The image to be segmented may be, for example, an image acquired by synthetic aperture radar (SAR).
[0019] The training apparatus then optimizes the model (S104) and outputs the optimized parameters and the like (S106), thereby completing the training. An information processing apparatus as an inference apparatus (segmentation apparatus) can construct a model based on the parameters output in S106, and obtain segmented data by inputting a target image into this model. Inference can be performed using, for example, a probability distribution (approximate posterior distribution) constructed using these parameters.
[0020] The above example of processing will be described in detail later. Hereinafter, a method that considers a probability distribution for performing higher-accuracy training even when there is noise in data provided as teacher data will be described.
[0021] Here, it is desirable that all of the labels acquired in S102 indicate correct labels, but they may be labels including errors. In the case of labels including errors, it is difficult for the training apparatus to realize training by appropriate machine learning. Therefore, the training apparatus optimizes the model that can achieve higher-accuracy inference by considering the likelihood of the parameters of the model that performs segmentation.
[0022] The training of this model will be described in more detail. In the following description, H represents the height of an image, W represents the width of an image, x ∈ R H × W × C is an image to be segmented, y ∈ {0, 1} HW × L is a segmentation mask image, y* ∈ {0, 1} HW × Lrepresents the true segmentation mask image (unknown), θ denotes the parameters of the segmentation model, p(y* | x, θ) represents the likelihood of the parameters of the segmentation model, and p(y | y*) represents the probability of label error, respectively. For example, letting L be the number of label types, as will be described in detail in Equations (20) to (22) below, y, y* ∈ {0, 1} HW × L can be expressed using a one-hot vector.
[0023] The training apparatus optimizes the model using the above y as training data. When attempting optimization under the above circumstances, since it is necessary to set probabilities for the total number of elements of y* in {0, 1} HW × L , the summation calculation is difficult to complete within real time, and it is not practical to perform calculation of the posterior distribution.
[0024] Therefore, it is assumed that the prior distribution and the likelihood are independent for each pixel as follows. Equation (1) shows the prior distribution of parameters for the true label independent for each pixel i, and Equation (2) shows the likelihood of label error independent for each pixel i. In the case of this model, the posterior distribution can be calculated independently for each pixel as follows.
[0025] The log-likelihood is expressed as follows using a hidden variable. KL[・] indicates the Kullback-Leibler (KL) divergence.
[0026] Here, q(y*; y) is optimized for each image. If a common value independent of images is used for the likelihood p(y | y*) and the parameter θ, the likelihood p(y | y*) can be set such that when a label error actually occurs, adjacent pixels are also likely to have label errors similarly. When setting in this manner, it is necessary to use a model that is not independent for pixel i but has spatial or temporal correlation.
[0027] For the prior distribution of parameters p(y* | x, θ) with respect to the true label, let the label error η (∈ R H × W ) where η is known, the likelihood is defined as p(η) = N(η|μ, Σ). Σ is the covariance matrix of p(η). Here, to represent the correlation shown above, Σ can be configured to have correlation only between adjacent pixels. In particular, Σ may be defined as a covariance matrix using a Kac-Murdock-Szegö matrix (KMS matrix), which is a type of Toeplitz matrix. r(η) is a monotonically increasing function that takes values between 0 and 1. For example, r(η) may have a shape similar to a sigmoid function or the like, provided that the following condition is satisfied.
[0028] The log likelihood shown in equation (4) can be transformed as follows. As a non-limiting example, q(η) = Π i q(ηi), and further transformation yields the following inequality. Then, q(η) is set as a normal distribution, that is, configured as shown in equation (10). Equation (10) means that q(η) can be set to a correlated normal distribution without limiting q(η) to an independent normal distribution between the elements of η. Here, when this normal distribution is defined as a normal distribution with correlation in q(η), its covariance matrix shall be described using a KMS matrix. In the generalized case, Γ in equation (10) represents a covariance matrix that can be expressed using a KMS matrix.
[0029] Further, considering equations (5) and (6), by calculating the lower bound of each term, the log likelihood can be bounded from below. To calculate the first term of equation (8), it is necessary to compute the following equation.
[0030] Equation (11) calculates the sum of y*i using the calculated distribution, performing L sum calculations, and then calculating the results in parallel for each pixel i. The distribution q(y*i, yi) used to calculate this expectation can be approximated by equation (12) when r(ηi) is a sigmoid function.
[0031] Furthermore, the calculation of equation (11) can also be approximated using the Monte Carlo method with the Reparametrization Trick, based on the following equation. I represents the identity matrix. Here, By performing this reparameterization, N(ε i Alternatively, a Monte Carlo method can be used to approximate the integral using a sample εi from |0, I).
[0032] The second term of equation (8) can be transformed as follows, assuming that both p and q follow normal distributions p(η) = N(η|μ, Σ) and q(η) = N(η|m, Γ).
[0033] While this equation is generally impossible to calculate analytically, if the covariance matrix Σ of p(η) is represented by a KMS matrix, which is one of the Teplitz matrices, then the inverse and determinant can be calculated analytically. Therefore, in this disclosure, by using a Teplitz matrix (particularly a KMS matrix) for Σ, analytical calculation becomes possible.
[0034] The third term of equation (8) can be calculated in L ways for each pixel because q(y*; yi) is independent for each pixel. Furthermore, the integral calculation involving q(η) in this third term can be approximated by the Monte Carlo method using the Reparametrization Trick.
[0035] The fourth term of equation (8), presented as an unrestricted example, can be approximated by, for example, the following transformation. This fourth term can be considered as the calculation of the entropy of a normal distribution, if q(η) is a normal distribution. This entropy is described using the determinant of the normal distribution, and this determinant can be calculated analytically if the covariance matrix of q(η) is expressed using the KMS matrix. In this equation, for example, if we Taylor expand f(ηi) around ηi = mi, we get the following, and we can use this to find an approximate value.
[0036] In addition to Taylor expansion as shown in equation (16), approximation using the Monte Carlo method with the Reparametrization Trick is also possible.
[0037] The fifth term of equation (8) can be calculated by performing L sum calculations using equation (17), similar to the third term described above.
[0038] The information processing device can suppress the log-likelihood from below based on the above equation. Here, the number of pixels is represented as HW, and the state of each pixel is represented by a one-hot vector that takes 1 only for the corresponding class out of L possible classes, and 0 for all other classes. In this case, both y and y* can be represented as HW × L dimension vectors. For each i ∈ HW, it can be expressed as follows. First, In, This is how it is defined. however, That is the case.
[0039] According to the definition above, when yk = 1, it can be expressed as follows:
[0040] Similarly, when yk = 0, it can be expressed as follows.
[0041] Using this, the likelihood can be calculated as follows. This also demonstrates the usefulness of making assumptions in equations (5) and (6). Furthermore, by using this result, if r(•) is a sigmoid function and each element μi of the mean μ of p(η|μ, Σ) is μi < 0, then the following relationship can be said.
[0042] By using the above formulas, the information processing device can analytically calculate the probability distribution by analytically performing the inverse of the summation matrix. As a result, even when there are errors in the labels, an appropriate prior distribution can be set, and even when there are errors in the training data, it becomes possible to achieve highly accurate optimization based on correlations from neighboring pixels.
[0043] Here, highly accurate optimization may refer to the optimization of the posterior distribution q, or it may refer to the optimization of both the posterior distribution q and the model parameters θ. In other words, the information processing device can, for example, perform only the optimization of the posterior distribution q if the parameters θ are known, or it can perform the optimization of both the posterior distribution q and the model parameters θ if the parameters θ are unknown.
[0044] Below, Σ -1 The analytical calculation of p will be explained in detail using an example that is not limited to this. For example, p can be set such that each pixel is correlated with its four neighboring pixels. In particular, when described as Σ by a KMS matrix that affects neighboring pixels, it can be expressed as follows. If it can be written as, This can be done. In the Kronecker product, the following equation (31) holds.
[0045] Furthermore, since it can be written as equation (30), if p(η) is a normal distribution, It can be written as follows: Here, Furthermore, It can be written as follows, and a block diagonal matrix can be written as follows: At this time, the following holds true: From these, using equation (31), we get Tr(Γ Σ -1 Since it is possible to calculate Σ -1 This can be calculated analytically as follows.
[0046] Based on the above formula, we can calculate the lower bound of the log-likelihood, and thus analytically achieve optimization that maximizes this likelihood. That is, the data that can be used as training data D = {(x n , y n By optimizing the model using {|n = 1, ..., N}, it becomes possible to create a model that can perform more accurate inferences even when trained with training labels that contain label errors.
[0047] This optimization generally requires the operation of the inverse matrix of a matrix with HW × HW elements as the covariance matrix, and is therefore difficult to solve analytically. However, according to the form of this disclosure, as described above, an analytical solution can be calculated in real time by using the KMS matrix as the covariance matrix Σ.
[0048] (First application example)
[0049] As a training device, an information processing device can, in a limited number of cases, train an inference model when the training data contains noise, as shown below. However, the method using the above formula is not limited to the following series of processes. For example, if the parameters θ of the model to be trained are unknown, the information processing device approximates the posterior distribution of data errors based on a prior distribution that assumes that data errors are correlated between elements in the data, and the likelihood of a normal model, while learning the model parameters θ.
[0050] First, the information processing device initializes the parameters to be trained (S100). These parameters may include the parameters θ of the model to be trained as described above, as well as the parameters of the prior distribution and the parameters of the posterior distribution q related to the data used for training.
[0051] The information processing device extracts one or more samples from the training data acquired in S102 for each optimization step. Next, the information processing device performs optimization of each parameter using the above formula (S104).
[0052] The information processing device calculates the Evidence Lower Bound (ELBO) for the extracted samples and the gradient of this ELBO with respect to the parameter θ of the model under training, related to the parameter θ of the model under training, using the series of equations described above. Based on the calculated gradient, the information processing device updates the parameter θ of the model under training. This update of the parameter θ of the model under training may be repeated, for example, a first predetermined number of times within a step.
[0053] Next, the information processing device uses the updated parameters θ of the training model to calculate the ELBO with respect to the parameters of the posterior distribution q corresponding to the extracted samples, and the gradient of the ELBO with respect to the parameters of the posterior distribution q, using the series of equations described above. Based on the calculated gradient, the information processing device updates the parameters of the posterior distribution q of interest. This updating of the parameters of the posterior distribution q may be repeated, for example, a second predetermined number of times within the step.
[0054] The optimization of the parameters of the posterior distribution q is performed, for example, to maximize the ELBO with respect to the parameters of the posterior distribution q. By maximizing the ELBO, the KL divergence in equation (4) can be reduced, for example. The information processing device can then approximately determine the posterior distribution q through this process.
[0055] Through the above series of processes, both the parameters of the posterior distribution q and the model parameters θ can be determined by maximizing the ELBO. As a result, even when the training data contains noise, it is possible to train the inference model using data with reduced noise by considering the correlation with neighboring data of the data of interest.
[0056] An information processing device acting as an inference device can perform inference using an inference model formed by parameters θ optimized by a training device.
[0057] (Second application example)
[0058] On the other hand, as an example without limitation, the information processing device as a training device can also be applied when the model parameters θ are known or partially known.
[0059] The information processing device can use known parameters during initialization (S100). For example, the information processing device can use known model parameters θ as the parameters of the model. For example, in the process of S100, the information processing device sets the model parameters to known parameters θ and performs the initialization of the parameters of the posterior distribution q as a separate process. It is also possible to use at least some of the known parameters θ, rather than all of them.
[0060] The information processing device uses the model parameter θ to calculate the gradients of the ELBO and the parameters of the posterior distribution q corresponding to the extracted sample, using the series of equations described above. Based on the calculated gradients, the information processing device can approximately determine the posterior distribution q by updating the parameters of the posterior distribution q.
[0061] In other words, the information processing device can use the known model parameters θ to approximate the parameters of the posterior distribution q based on the likelihood of this model by maximizing the ELBO.
[0062] As described above, by introducing the likelihood of the label error distribution using a normal distribution with the KMS matrix as its covariance matrix, it becomes possible to analytically calculate the likelihood and estimate spatially correlated distributions. Using this distribution, for example, it becomes possible to optimize the segmentation model using analytical calculations, thereby improving the accuracy of inference.
[0063] As an example of the application of this disclosure, a segmentation model for images is described, but this does not preclude its application to other areas. The contents of this disclosure can be similarly applied not only to cases where the spatial correlation is two-dimensional, as described above, but also to spatial correlations in one dimension, three dimensions, or higher dimensions.
[0064] Furthermore, dimensions can, of course, include time, in which case the disclosure can also deal with temporal correlations or correlations related to time. Moreover, the disclosure can also be applied to correlations of modalities other than space and time.
[0065] In other words, the contents of this disclosure can be applied to any problem that solves computational problems related to covariance matrices, particularly covariance matrices with a large number of elements.
[0066] For example, this method can be applied to correlations that are basically represented by a normal distribution. When the size of the covariance matrix is n × n, the covariance matrix is not a diagonal matrix because of the correlation. Therefore, when n is large, the computational cost of the inverse matrix and determinant becomes high, and furthermore, depending on the amount of memory that can be allocated, the calculation itself may become difficult.
[0067] In a naive approach, the time complexity is O(n 3 ) and at least O(n 2) will be larger than . In this case, estimating the posterior distribution using a prior distribution that indicates correlation, or estimating the posterior distribution that indicates correlation, can be computationally problematic. This computational problem can be solved by utilizing the fact that the inverse and determinant of the KMS matrix can be explicitly calculated. When representing multidimensional correlation, this can be solved by using the Kronecker product as described above.
[0068] Furthermore, the representation of correlations between discrete variables can also be devised in the manner described in the series of equations above. For example, while the correlation of continuous variables can be represented by a normal distribution, in cases where discrete variables, such as segmentation labels (given as an example without limitation), are spatially correlated, the correlation between discrete variables can be indirectly represented by using a representation via a normal distribution, such as the ELBO Computable Discrete Distribution. This means that ELBO optimization is possible. This is expressed by the following equation (40).
[0069] The series of equations shown above are for the random variable x ∈ R n The fact that (n ∈ N) has a uniform correlation is that the i-th element x of the random variable x i and the j-th element x j This refers to the following correlations: This indicates a correlation with ρ, where i, j ∈ {1, ..., n} and -1 < ρ < 1.
[0070] In equation (41), if i = j, Therefore, σ1, σ j These are x i , x j This is the standard deviation. A random variable x has a uniform correlation if the correlation coefficient depends not on the absolute values of the indices i and j of x, but on the difference between the indices i and j, |i - j|. This means that the covariance matrix is determined. This covariance matrix can be represented using the KMS matrix, as mentioned above. Of course, as stated above, multidimensional correlations can be represented using the KMS matrix.
[0071] As shown in equation (42), a probability distribution with spatial correlation between discrete variables can be represented through a correlated normal distribution. By using this representation, the information processing device can perform numerical evaluation of the ELBO and / or optimization of its parameters.
[0072] The contents of this disclosure can also be summarized as follows, as an example without limitation:
[0073] (1) An information processing device comprising at least one memory and at least one processor, wherein the at least one processor estimates the data of interest based on a probability distribution defined by a probability model expressed using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest.
[0074] (2) The information processing apparatus according to (1), wherein the covariance matrix is represented using a matrix whose values are set such that neighboring data have a stronger correlation than non-neighboring data.
[0075] (3) The information processing device described in (2), wherein the matrix is described by a Teplitz matrix.
[0076] (4) The information processing apparatus according to (3), wherein the matrix is described by a matrix whose inverse can be analytically calculated.
[0077] (5) The information processing apparatus described in (4), wherein the matrix is described by a KMS matrix.
[0078] (6) The information processing apparatus according to any one of (1) to (5), wherein at least one processor trains an inference model by supervised learning using the estimated labels.
[0079] (7) An inference device that performs inference on a target using an inference model trained by the information processing device described in (6).
[0080] (8) An information processing method that estimates the data of interest based on a probability distribution defined by a probability model represented by a probability model expressed using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest, using at least one processor.
[0081] (9) A program that causes at least one processor to execute a method for estimating the data of interest based on a probability distribution defined by a probability model expressed using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest.
[0082] In the embodiments described above, some or all of the devices (information processing devices) may be composed of hardware, or they may be composed of information processing by software (programs) executed by a CPU (Central Processing Unit) or GPU (Graphics Processing Unit), etc. If the information processing is composed of software, the software that realizes at least some of the functions of each device in the embodiments described above may be stored on a non-temporary storage medium (non-temporary computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or USB (Universal Serial Bus) memory, and the software information processing may be executed by loading it into a computer. Alternatively, the software may be downloaded via a communication network. Furthermore, all or part of the software processing may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), so that the information processing by the software is executed by hardware.
[0083] The storage medium for the software may be a removable medium such as an optical disc, or a fixed storage medium such as a hard disk or memory. Furthermore, the storage medium may be located inside the computer (main memory or auxiliary storage, etc.) or outside the computer.
[0084] Figure X is a block diagram showing an example of the hardware configuration of each device (information processing device) in the embodiment described above. Each device may be implemented as a computer 7, for example, comprising a processor 71, a main memory 72 (memory), an auxiliary memory 73 (memory), a network interface 74, and a device interface 75, all connected via a bus 76.
[0085] The computer 7 in Figure X has one of each component, but it may have multiple identical components. Also, although Figure X shows one computer 7, the software may be installed on multiple computers, and each of these computers may execute the same or different parts of the software's processing. In this case, it may be a distributed computing configuration in which each computer communicates via a network interface 74 or the like to execute processing. In other words, each device (information processing device) in the above-described embodiment may be configured as a system that realizes its function by having one or more computers execute instructions stored in one or more storage devices. Alternatively, it may be configured so that information transmitted from a terminal is processed by one or more computers located on the cloud, and the processing results are transmitted to the terminal.
[0086] The various calculations performed by each device (information processing device) in the embodiments described above may be executed in parallel using one or more processors, or using multiple computers connected via a network. Alternatively, the various calculations may be distributed to multiple processing cores within a processor and executed in parallel. Furthermore, some or all of the processing and means of this disclosure may be implemented by at least one of a processor and a storage device located on a cloud that can communicate with computer 7 via a network. Thus, each device in the embodiments described above may be in the form of parallel computing using one or more computers.
[0087] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that performs at least one of the following: control of a computer or calculations. The processor 71 may also be a general-purpose processor, a dedicated processing circuit designed to perform specific calculations, or a semiconductor device including both a general-purpose processor and a dedicated processing circuit. Furthermore, the processor 71 may include optical circuits or quantum computing-based calculation functions.
[0088] The processor 71 may perform calculations based on data and software input from various devices within the computer 7, and may output calculation results and control signals to these devices. The processor 71 may also control the various components of the computer 7 by executing the computer 7's OS (Operating System) or applications.
[0089] Each of the devices (information processing devices) in the embodiments described above may be implemented by one or more processors 71. Here, processor 71 may refer to one or more electronic circuits arranged on one chip, or one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, each electronic circuit may communicate by wire or wireless.
[0090] The main memory 72 may store instructions executed by the processor 71 and various data, and the information stored in the main memory 72 may be read by the processor 71. The auxiliary storage device 73 is a storage device other than the main memory 72. These storage devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. In each of the devices (information processing devices) in the embodiments described above, the storage device for storing various data may be implemented by the main memory 72 or the auxiliary storage device 73, or by the built-in memory of the processor 71. For example, the storage unit 102 in the embodiments described above may be implemented by the main memory 72 or the auxiliary storage device 73.
[0091] In the embodiments described above, if each device (information processing device) consists of at least one storage device (memory) and at least one processor connected to (coupled with) this at least one storage device, then at least one processor may be connected to one storage device. Also, at least one storage device may be connected to one processor. Furthermore, the configuration may include at least one processor among a plurality of processors being connected to at least one storage device among a plurality of storage devices. This configuration may also be realized by storage devices and processors included in a plurality of computers. Moreover, the configuration may include a storage device integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache).
[0092] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wired connection. The network interface 74 can be any appropriate interface, such as one conforming to existing communication standards. Information may be exchanged between the computer 7 and an external device 9A connected via the communication network 8 through the network interface 74. The communication network 8 may be a WAN (Wide Area Network), LAN (Local Area Network), PAN (Personal Area Network), or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet; an example of a LAN is IEEE 802.11 or Ethernet (registered trademark); and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication).
[0093] Device interface 75 is an interface such as USB that connects directly to the external device 9B.
[0094] External device 9A is a device connected to computer 7 via a network. External device 9B is a device directly connected to computer 7.
[0095] External device 9A or external device 9B may, for example, be an input device. The input device may be, for example, a camera, microphone, motion capture device, various sensors, keyboard, mouse, or touch panel, and will provide the acquired information to computer 7. Alternatively, it may be a device equipped with an input unit, memory, and processor, such as a personal computer, tablet terminal, or smartphone.
[0096] Furthermore, external device 9A or external device 9B may, for example, be an output device. The output device may be a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or a speaker that outputs sound, etc. It may also be a device equipped with an output unit, memory, and a processor, such as a personal computer, tablet terminal, or smartphone.
[0097] Furthermore, external device 9A or external device 9B may be a storage device (memory). For example, external device 9A may be network storage, and external device 9B may be storage such as an HDD.
[0098] Furthermore, the external device 9A or external device 9B may be a device having some of the functions of the components of each device (information processing device) in the embodiments described above. In other words, the computer 7 may transmit some or all of the processing results to the external device 9A or external device 9B, or may receive some or all of the processing results from the external device 9A or external device 9B.
[0099] Where the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used herein (including the claims), it includes any of a, b, c, a-b, a-c, b-c, or a-b-c. It also includes multiple instances of any element, such as a-a, a-b-b, a-a-b-b-c-c, etc. Furthermore, it includes adding other elements other than the enumerated elements (a, b, and c), such as a-b-c-d which has d.
[0100] In this specification (including the claims), when expressions such as "using data as input / based on data / according to / in accordance with data" (including similar expressions) are used, unless otherwise specified, this includes using the data itself or using data that has been processed in some way (e.g., data with added noise, normalized data, features extracted from the data, an intermediate representation of the data, etc.). Furthermore, when it is stated that some result is obtained "using data as input / based on data / according to / in accordance with data" (including similar expressions), unless otherwise specified, this includes cases where the result is obtained based solely on the data in question or where the result is influenced by other data, factors, conditions, and / or states other than the data in question. Furthermore, when it is stated that "data is output" (including similar expressions), unless otherwise specified, this includes cases where the data itself is used as output or where data that has been processed in some way (e.g., data with added noise, normalized data, features extracted from the data, an intermediate representation of the data, etc.) is used as output.
[0101] Where the terms “connected” and “coupled” are used herein (including in the claims), they are intended to be non-restrictive terms that include any direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operational connection / coupling, or physical connection / coupling. The terms should be interpreted as appropriate in the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted non-restrictively as being included in the terms.
[0102] In this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that the permanent or temporary setting / configuration of element A is configured to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and that it is configured to actually perform operation B by the setting of a permanent or temporary program (instruction). Furthermore, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.
[0103] Wherever terms meaning "comprising" or "having" are used in this specification (including the claims), they are intended to be open-ended terms, including cases where the subject matter of such terms is not the object of the term. Where the object of such terms meaning "comprising" or "having" is an expression that does not specify a quantity or suggests a singular number (an expression with the article a or an), such expression should be interpreted as not being limited to a specific number.
[0104] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in one place, and expressions that do not specify a quantity or suggest singularity (expressions using the articles a or an) are used in another place, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest singularity (expressions using the articles a or an) should be interpreted as not necessarily limited to a specific number.
[0105] In this specification, if a particular configuration of an embodiment is described as yielding a specific advantage or result, it should be understood, unless otherwise stated, that the same advantage or result can also be obtained from one or more other embodiments having that configuration. However, it should be understood that the presence or absence of such advantage or result generally depends on various factors, conditions, and / or states, and that the configuration does not necessarily guarantee that the advantage or result can be obtained. The advantage or result can only be obtained from the configuration described in the embodiment when various factors, conditions, and / or states are met, and the advantage or result cannot necessarily be obtained in the invention claimed to define that configuration or a similar configuration.
[0106] In this specification (including the claims), when terms such as "maximize" are used, they include finding the global maximum value, finding an approximation of the global maximum value, finding the local maximum value, and finding an approximation of the local maximum value, and should be interpreted appropriately depending on the context in which the terms are used. They also include finding approximations of these maximum values probabilistically or heuristically. Similarly, when terms such as "minimize" are used, they include finding the global minimum value, finding an approximation of the global minimum value, finding the local minimum value, and finding an approximation of the local minimum value, and should be interpreted appropriately depending on the context in which the terms are used. They also include finding approximations of these minimum values probabilistically or heuristically. Similarly, when terms such as "optimize" are used, they include finding the global optimal value, finding an approximation of the global optimal value, finding the local optimal value, and finding an approximation of the local optimal value, and should be interpreted appropriately depending on the context in which the terms are used. This also includes finding approximate values of these optimal values probabilistically or heuristically.
[0107] In this specification (including the claims), when multiple hardware components perform a predetermined process, each component may cooperate to perform the predetermined process, or some components may perform all of the predetermined process. Alternatively, some components may perform part of the predetermined process, while other components perform the remainder. In this specification (including the claims), when expressions such as "one or more hardware components perform a first process, and the one or more hardware components perform a second process" (including similar expressions) are used, the hardware component performing the first process and the hardware component performing the second process may be the same or different. In other words, it is sufficient that the hardware component performing the first process and the hardware component performing the second process are included in the one or more hardware components. Hardware may include electronic circuits or devices containing electronic circuits.
[0108] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices may store only a portion of the data or the entire data. Furthermore, a configuration in which some of the multiple storage devices store data is also included.
[0109] While embodiments of this disclosure have been described in detail above, this disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, and partial deletions are possible, provided that they do not deviate from the conceptual idea and spirit of this disclosure derived from the claims and their equivalents. For example, where numerical values or mathematical formulas are used in the description of the embodiments described above, these are provided for illustrative purposes only and do not limit the scope of this disclosure. Similarly, the sequence of operations shown in the embodiments is also illustrative and does not limit the scope of this disclosure.
Claims
1. An information processing device comprising at least one memory and at least one processor, wherein the at least one processor estimates the data of interest based on a probability distribution defined by a probability model expressed using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest.
2. The information processing apparatus according to claim 1, wherein the covariance matrix is represented using a matrix whose values are set such that neighboring data have a stronger correlation than non-neighboring data.
3. The information processing apparatus according to claim 2, wherein the matrix is described by a Teplitz matrix.
4. The information processing apparatus according to claim 3, wherein the matrix is described by a matrix whose inverse can be analytically calculated.
5. The information processing apparatus according to claim 4, wherein the matrix is described by a KMS matrix.
6. The information processing apparatus according to any one of claims 1 to 5, wherein the at least one processor trains an inference model by supervised learning using the estimated labels.
7. An inference device that performs inference on a target using an inference model trained by the information processing device described in claim 6.
8. An information processing method that estimates the data of interest based on a probability distribution defined by a probability model represented by a probability model expressed using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest, using at least one processor.
9. A program that causes at least one processor to execute a method for estimating the data of interest based on a probability distribution defined by a probability model represented using a prior distribution with a covariance matrix set to correlate with neighboring data of the data of interest.