Information processing device and method

The information processing apparatus and method address the issue of accumulating prediction errors in event forecasting by projecting event data into a latent space, removing noise, and combining data, resulting in accurate time-series predictions.

WO2025109694A1PCT designated stage expired Publication Date: 2025-05-30NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/041859
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for predicting events over a certain period, such as equipment failures or infectious diseases, suffer from accumulating prediction errors due to the reuse of previous prediction results, leading to a deterioration in prediction accuracy.

Method used

An information processing apparatus and method that projects event data into a latent space to generate a latent representation, estimates and removes noise from subsequent event data, and combines the original and noise-removed event data to enhance prediction accuracy.

Benefits of technology

This approach enables accurate prediction of time-series data over a specific period by reducing noise and improving the reliability of event predictions, thereby maintaining high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023041859_30052025_PF_FP_ABST
    Figure JP2023041859_30052025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to one embodiment includes: a first output unit that projects, into a latent space, first event data indicating events that occurred between a first timing and a second timing and the timings at which the events occurred, and thereby outputs a latent representation of the first event data; and a second output unit that estimates, on the basis of the latent representation and second event data which is random event data from between the second timing and a third timing, a noise component added to the second event data, and thereby outputs event data from which the noise component has been removed.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and method

[0001] FIELD Embodiments of the present invention relate to an information processing apparatus and method.

[0002] One method known for predicting the occurrence of various events, such as equipment failures, human behavior, crimes, earthquakes, and infectious diseases, is to use point processes. A point process is a probabilistic model that describes, for example, the instantaneous value of an event that has occurred, and the timing of the event's occurrence.

[0003] To predict events using point processes, the following methods (A-1) to (A-3) are usually used: (A-1) The next event to occur is predicted based on the history of events up to the present. (A-2) The occurrence probability (intensity function) is recalculated assuming that the predicted event is an event that has occurred. (A-3) The next event is further predicted.

[0004] Furthermore, for example, Non-Patent Document 1 discloses that a point process and a diffusion process are combined to correspond to a point process in space-time, and the time and position of the next event to occur are predicted.

[0005] Spatio-temporal Diffusion Point Processes<https: / / arxiv.org / abs / 2305.12403>

[0006] However, when predicting all events occurring within a certain period of time, such as all events occurring within the next week, previous prediction results are repeatedly used to make further predictions, which results in accumulated errors during prediction and reduces prediction accuracy.

[0007] When considering practical applications, it is necessary to predict events that will occur over a certain period of time with high accuracy.

[0008] The present invention has been made in light of the above circumstances, and its object is to provide an information processing device and method that are capable of predicting time series data generated over a certain period of time with high accuracy.

[0009] An information processing device according to one aspect of the present invention includes a first output unit that outputs a latent representation of first event data by projecting first event data indicating events that occurred between a first timing and a second timing and the timings at which the events occurred into a latent space, and a second output unit that estimates a noise component added to the second event data based on second event data, which is random event data from the second timing to a third timing, and the latent representation, and outputs event data from which the noise component has been removed. An information processing device according to one aspect of the present invention includes a first output unit that outputs a latent representation of first event data by projecting first event data indicating events that occurred between a first timing and a second timing and the timings at which the events occurred into a latent space; a second output unit that estimates noise components added to the second event data based on the second event data, which is random time series data, and the latent representation, and outputs event data from which the noise components have been removed; and a combination unit that combines the first event data and the event data output by the second output unit to generate event data in which missing parts of the first event data have been complemented.

[0010] An information processing method according to one aspect of the present invention is a method performed by an information processing device, comprising: a first output unit of the information processing device projects first event data indicating events that occurred from a first timing to a second timing and the timings of the events into a latent space, thereby outputting a latent representation of the first event data; and a second output unit of the information processing device estimating a noise component added to the second event data based on second event data, which is random event data from the second timing to a third timing, and the latent representation, and outputting event data from which the noise component has been removed.

[0011] According to the present invention, it is possible to predict time series data occurring over a certain period of time with high accuracy.

[0012] FIG. 1 is a diagram showing an application example of an event prediction device according to a first embodiment of the present invention. FIG. 2 is a flowchart showing an example of processing related to learning model parameters. FIG. 3 is a flowchart showing an example of processing related to predicting event data. FIG. 4 is a diagram showing an application example of an event prediction device according to a second embodiment of the present invention. FIG. 5 is a flowchart showing an example of processing related to learning model parameters. FIG. 6 is a flowchart showing an example of processing related to predicting event data. FIG. 7 is a block diagram showing an example of the hardware configuration of an event prediction device according to an embodiment of the present invention.

[0013] An embodiment of the present invention will be described below with reference to the drawings. (First Embodiment) First, the first embodiment will be described. FIG. 1 is a diagram showing an application example of an event prediction device according to a first embodiment of the present invention. As shown in FIG. 1, an event prediction device 100 according to the first embodiment of the present invention includes an embedding unit 10, a diffusion unit 20, a dediffusion unit 30, and an update unit 40.

[0014] The event data handled in this embodiment is expressed as follows (1): This event data is point process data indicating, for example, the instantaneous value of an event that has occurred and the time at which the event occurred.

[0015]

[0016] n is the number of events. s is the timing at which the observation of the event data begins. e is the timing at which observation of the event data ends, and k and K are the number of steps in the diffusion process by the diffusion unit 20.

[0017] The event data when noise is removed by the despreading process for the steps by the despreading unit 30 is expressed as shown in (2) below.

[0018]

[0019] When k = 0, the event data in (2) above is random event data. When k = K, the event data in (2) above is real event data or data that is estimated from real event data. h is a vector of latent representations. n, t s , and t e varies depending on the individual event data.

[0020] In this embodiment, the learning dataset for the prediction model of event data is expressed as follows (3).

[0021]

[0022] In this embodiment, the prediction-time occurrence history of event data is expressed as shown in (4) below.

[0023]

[0024] In this (4), the following relationship (5) holds: s <t c <t e ...(5) The prediction target of the event data in this embodiment is expressed as follows (6).

[0025]

[0026] The prediction target shown in (6) above is a time (t c , t e ] are all events that occur in (6) above. s , t e may differ from the data used during training. m is also an object of estimation.

[0027] The embedding unit 10 projects a sequence into a latent space using an embedding layer, which is an arbitrary differentiable model that can handle sequences. For example, the embedding unit 10 may use a recurrent neural network (RNN), attention, or the configuration disclosed in the above-mentioned non-patent document 1 with the spatial portion removed.

[0028] The input to the embedding unit 10 is expressed as follows (7): This input corresponds, in the order of notation, to the event data and the start and end times of the observation period of the event data.

[0029]

[0030] The output h from the embedding unit 10 is a latent representation of the event data. The diffusion unit 20 adds noise to the event data expressed as in (8) below.

[0031]

[0032] The diffusion unit 20 can add noise to the event data by performing the following operations (B-1) to (B-3) probabilistically. The definition of probability is arbitrary, and the following operations may be performed simultaneously, the same operation may be performed multiple times, or some operations may not be performed.

[0033] Operation (B-1): The diffusion unit 20 randomly removes the data expressed by the following (10) from the event data expressed by the following (9).

[0034]

[0035] Operation (B-2): The diffusion unit 20 adds the event t i (t≦t i The added value is random and can have any distribution such as a uniform distribution.

[0036] Operation (B-3): The diffusion unit 20 adds Gaussian noise expressed by the following (11) to the data expressed by the above (10). The mean of this Gaussian noise is 0, and the variance of the Gaussian noise is predetermined depending on k.

[0037]

[0038] When the data interval deviates from the interval [t, t'] as a result of adding Gaussian noise, the diffusion unit 20 sets the data interval to either t or t', whichever is closer, or deletes the data that deviates from the interval.

[0039] Instead of performing each of the above operations by the diffusion unit 20 multiple times, it is possible to perform the calculations in one go using the following method. Regarding the above operation (B-3), the Gaussian noise that is added multiple times can be represented by a single Gaussian noise, and the diffusion unit 20 can add multiple Gaussian noises at once. The average of this Gaussian noise is 0, and the variance of the Gaussian noise varies depending on a predetermined value.

[0040] For example, when the diffusion unit 20 applies Gaussian noise addition twice, i.e., adding two Gaussian noises represented by the following (12), is the same as adding Gaussian noise represented by the following (13) once.

[0041]

[0042] Regarding the above operation (B-1), the operation of removing the data expressed by the above (10) by the diffusion unit 20 has a probability p 1 , ..., p K The probability of being removed by the number of times it is applied is 1-p K ×p K-1 It can be calculated as follows:

[0043] It is possible to calculate the distribution of the number of events added when the above operation (B-2) is applied once, and the distribution of the number of events added when the above operation (B-1) is applied multiple times. For example, if 10 events are added per application of the above operation (B-2) and the probability that no events are removed is p, then when the above operation (B-1) is applied k times, the distribution of the number of events added is 10p(p k The times of the events to be added are approximated by the same distribution as when the above operation (B-1) is applied once.

[0044] The input to the spreading unit 20 is expressed as in (14) below, and the output from the spreading unit 20 is expressed as in (15) below.

[0045]

[0046] In the above (14), k is the number of steps in the diffusion model.

[0047] The despreading unit 30 estimates and removes noise contained in the event data expressed by the following (16).

[0048]

[0049] The despreading unit 30 is realized using any differentiable model that can handle sequences. For example, the despreading unit 30 is configured as a transformer having an encoder and a decoder, and inputs a value expressed by the following (17) including the embedding of k (positional encoding) to the encoder and embeds it in a latent space, and outputs data expressed by the following (18) from the decoder.

[0050]

[0051] That is, the input to the despreading unit 30 is expressed by the following (19), and the output from the despreading unit 30 is expressed by the following (20).

[0052]

[0053] The update unit 40 updates the model parameters of the despreading unit 30 and the embedding unit 10 based on a loss function. The loss function is based on, for example, "Denoising Diffusion Probabilistic Models" (https: / / arxiv.org / abs / 2006.11239) or an extension thereof. For example, the loss function is expressed as an expected value such as the following (21).

[0054]

[0055] The spreading unit 20 prepares a value expressed by the following (22) from a value expressed by the following (23), and inputs a value expressed by the following (24) from among the values ​​expressed by the above (22) to the despreading unit 30, which estimates a value expressed by the following (25) from among the values ​​expressed by the above (22). The estimation result is expressed as the following (26).

[0056]

[0057] In the loss function above, D is the dataset. d(·,·) is a differentiable function that measures the difference between two pieces of data, such as the Van Rossum distance described in the following document: "A novel spike distance <https: / / era.ed.ac.uk / bitstream / handle / 1842 / 226 / distance.pdf> (cf. <http: / / proceedings.mlr.press / v139 / yang21n / yang21n.pdf>)" The function in question is primarily a distance function, but it does not need to satisfy the distance axioms.

[0058] Next, the process of learning the model parameters of the despreading unit 30 and the embedding unit 10 will be described as follows (C-1) to (C-10). FIG. 2 is a flowchart showing an example of the process of learning the model parameters. (C-1) First, event data is randomly selected from the learning data set D (S11). This selected data is expressed as shown in (27) below.

[0059]

[0060] (C-2) After S11, at time [t s , t e ], an arbitrary time represented by the following (28) is randomly determined (S12). In this determination, the condition represented by the following (29) is allowed. In the following, for simplicity of notation, the arbitrary time is referred to as t s , t c , t e It is assumed that it is decided that:

[0061]

[0062] (C-3) After S12, the data expressed by the following (30) is input to the embedding unit 10, and the embedding unit 10 outputs a latent expression h (S13).

[0063]

[0064] (C-4) After S13, kε{1, ..., K} is determined (S14), where K is the number of steps in the diffusion model and is a hyperparameter.

[0065] (C-5) After S14, the spreading unit 20 receives the data expressed by the following (31), k and t determined in S14. c and t e is applied, and the noise-added data expressed by the following (32) is output (S15).

[0066]

[0067] Each t in the data expressed by (31) above i t i -t c Alternatively, the data expressed by the following (33) may be used:

[0068]

[0069] In this case, the same process is performed when predicting, and when the prediction result is output, i t c is added and returned to the original value.

[0070] Furthermore, if the condition expressed by the following (34) is met and noise is added once by the spreading unit 20, the value expressed by the following (35) is obtained.

[0071]

[0072] The data expressed by (31) above is repeatedly applied to the spreading unit 20, whereby the data expressed by (32) above is obtained as the data expressed by (36) below. The data expressed by (31) above is applied to the spreading unit 20 once more, whereby the data expressed by (32) above is obtained as the data expressed by (37) below.

[0073]

[0074] (C-6) After S15, the despreading unit 30 receives the data represented by the following (38), and outputs the data represented by the following (39) from which the estimated noise has been removed (S16).

[0075]

[0076] (C-7) After S16, the update unit 40 calculates the loss function (S17). Here, the loss function is expressed as, for example, the following (40).

[0077]

[0078] (C-8) After S17, if the arbitrary number of iterations related to the calculation of the loss function has not been reached (No in S18), return to (C-4) above.

[0079] (C-9) After S17, when an arbitrary number of iterations related to the calculation of the loss function is reached (Yes in S18), the update unit 40 updates the parameters of the models of the de-diffusion unit 30 and the embedding unit 10 by the backpropagation method or the like (S19).

[0080] (C-10) After S19, if the number of iterations related to the parameter update has not been reached (No in S20), return to (C-1) above. If the number of iterations has been reached (Yes in S20), the series of processes related to learning the model parameters ends.

[0081] Next, the process for predicting event data after the learning will be described as follows (D-1) to (D-3). FIG. 3 is a flowchart showing an example of the process for predicting event data. (D-1) First, data t to be predicted, which is represented by the following (41), s and t c is input to the embedding unit 10, and the embedding unit 10 outputs a latent representation h (S31).

[0082]

[0083] (D-2) After S31, the despreading unit 30 randomly generates data expressed by the following (42) (S32). Here, for example, the number of events n is determined by a binomial distribution, and the number of events n is determined by a uniform distribution (t c , t e ] are sampled and sorted.

[0084]

[0085] (D-3) After S32, the data generated in S32 and the latent expressions h, k, and t output in S31 are sent to the despreading unit 30. c and t e is applied K times, and the prediction result expressed by the following (43) is output (S33).

[0086]

[0087] In this embodiment, the data represented by the following (45) is used to predict the data represented by the following (44), making it possible to make predictions that capture the interrelationships between events during the prediction period.

[0088]

[0089] In other words, compared to common methods that assume independence or predict each event sequentially, this method captures the mutual relationships between event data, enabling highly accurate predictions.

[0090] In this embodiment, an example of a time point process has been described. However, an extension of the point process, for example, a space-time point process (e.g., X={(t i ,lat i ,lon i )}), marked point processes (e.g., X = {(t i , M i)}), or a marked space-time point process (e.g., X = {(t i ,lat i ,lon i , M i) Each formula related to the extension of the point process is an example, and for example, the formula of the above space-time point process or the marked space-time point process may be a formula in which the concept of height is taken into consideration.

[0091] In the diffusion process, Gaussian noise is added to lat and lon. M is replaced by another mark.

[0092] d: lat,lon → The distance is calculated separately from t, and the Van Rossum distance can be calculated for each mark in M.

[0093] In this embodiment, external information a may be used as an additional input when calculating h in the embedding layer or in the despreading unit 30 .

[0094] For example, when a taxi ride is an event, the weather and temperature are used as a.

[0095] Second Embodiment Next, a second embodiment will be described. Fig. 4 is a diagram showing an application example of an event prediction device according to a second embodiment of the present invention. As shown in Fig. 4, an event prediction device 100a according to the second embodiment of the present invention includes an embedding unit 10a, a diffusion unit 20a, a dediffusion unit 30a, an update unit 40a, a division unit 50, and a combination unit 60. This second embodiment is used to fill in gaps in observations.

[0096] Capturing and analyzing the dynamics of event data is important in a variety of applications. The following methods (E-1) and (E-2) are typically used to analyze point processes. (E-1) A model is trained based on events that have occurred. (E-2) The degree of fit to the data is evaluated using likelihood.

[0097] However, in reality, it is not always possible to observe all events that occur. For example, equipment failure can prevent observation for a certain period of time. Also, probabilistic observations may be impossible due to some causes. When probabilistic observations are not possible, for example, infection with an infectious disease is considered an event, but asymptomatic people cannot be observed. Therefore, it is necessary to supplement the parts that could not be observed during analysis.

[0098] The following document discloses a method for analyzing event data: (Reference) "Intensity-Free Learning of Temporal Point Processes" <https: / / arxiv.org / abs / 1909.12127> However, analysis using the disclosed method requires that missing intervals be known in advance. This method also sequentially predicts events that occur in missing intervals, and trains model parameters to increase the likelihood of events before and after the missing interval. However, this method does not train the model parameters to correctly complement missing intervals, so the accuracy of predictions is unstable.

[0099] The embedding unit 10a projects a sequence into a latent space using an embedding layer, which is any differentiable model that can handle sequences. For example, the embedding unit 10a may use an RNN, attention, or a configuration disclosed in the above-mentioned non-patent document 1 with the spatial portion removed. The input to the embedding unit 10a is expressed as shown in the following (46). This input corresponds, in the order of notation, to event data including missing data and the start and end times of the observation period for that event data.

[0100]

[0101] The output h from the embedding unit 10a is a latent representation of the event data.

[0102] The diffusion unit 20a adds noise to the event data expressed as in (47) below.

[0103]

[0104] The diffusion unit 20a can add noise to the event data by performing the following operations (F-1) to (F-3) probabilistically. The definition of probability is arbitrary, and the following operations may be performed simultaneously, the same operation may be performed multiple times, or some operations may not be performed.

[0105] Operation (F-1): The diffusion unit 20a randomly removes the data expressed by the following (49) from the event data expressed by the following (48).

[0106]

[0107] Operation (F-2): The diffusion unit 20a adds the event t i (t≦t i The added value is random and can have any distribution such as a uniform distribution.

[0108] Operation (F-3): The diffusion unit 20a adds Gaussian noise expressed by the following (50) to the data expressed by the above (49). The mean of this Gaussian noise is 0, and the variance of the Gaussian noise is predetermined depending on k.

[0109]

[0110] When the data interval deviates from the interval [t, t'] as a result of adding Gaussian noise, the diffusion unit 20a sets the data interval to either t or t', whichever is closer, or deletes the data that deviates from the interval.

[0111] Instead of performing each of the above operations by the diffusion unit 20a multiple times, it is possible to perform the calculations in one go using the following method. For the above operation (F-3), the Gaussian noise added multiple times can be represented by a single Gaussian noise, and the diffusion unit 20a can add multiple Gaussian noises at once. The average of this Gaussian noise is 0, and the variance of the Gaussian noise varies depending on a predetermined value.

[0112] For example, when the diffusion unit 20a applies Gaussian noise addition twice, i.e., adding two Gaussian noises represented by the following (51), is the same as adding Gaussian noise represented by the following (52) once.

[0113]

[0114] Regarding the above operation (F-1), the operation of removing the data expressed by the above (49) by the diffusion unit 20a has a probability p 1, ..., p K The probability of being removed by the number of times it is applied is 1-p K ×p K-1 It can be calculated as follows:

[0115] It is possible to calculate the distribution of the number of events added when the above operation (F-2) is applied once, and the distribution of the number of events added when the above operation (F-1) is applied multiple times. For example, if 10 events are added per application and the probability that no events are removed is p, then when the above operation (F-1) is applied k times, the distribution of the number of events added is 10p(p k -1) / (p-1) events are added.

[0116] The time of the added event is approximated by the same distribution as when the above operation (F-1) is applied once.

[0117] The input to the spreading unit 20a is expressed as in (53) below, and the output from the spreading unit 20a is expressed as in (54) below.

[0118]

[0119] In the above (53), k is the number of steps in the diffusion model.

[0120] The despreading unit 30a estimates and removes noise contained in the event data expressed by the following (55).

[0121]

[0122] This despreading unit 30a is realized using any differentiable model that can handle two sequences.

[0123] For example, the model of the despreading unit 30a is configured as a transformer having two encoders and one decoder, and inputs a value expressed by the following (56) including positional encoding, which is the embedding of k, to the first encoder, inputs a value expressed by the following (57) including the embedding of k to the second encoder, and outputs data expressed by the following (58) from the decoder.

[0124]

[0125] That is, the input to the despreading unit 30a is expressed by the following (59), and the output from the despreading unit 30a is expressed by the following (60).

[0126]

[0127] The update unit 40a updates the model parameters of the despreading unit 30a and the embedding unit 10a based on a loss function.

[0128] As described in the first embodiment, the loss function is based on, for example, the "Denoising Diffusion Probabilistic Models" or an extension thereof.

[0129] In the second embodiment, for example, the loss function is expressed as an expected value as shown in the following (61).

[0130]

[0131] The spreading unit 20a prepares a value expressed by the following (62) from the value expressed by the following (63), and inputs the value expressed by the following (64) from the values ​​expressed by the above (62) to the despreading unit 30a, which estimates the value expressed by the following (65) from the values ​​expressed by the above (62). The estimation result is expressed as the following (66).

[0132]

[0133] D in the loss function above is a data set. d(·,·) in the loss function above is differentiable and is a function that measures the difference between two pieces of data, as described in the first embodiment.

[0134] The dividing unit 50 divides the event data included in the learning data set into observed event data and complemented event data.

[0135] When simulating a case where event data included in a learning dataset has missing data in a section ([t_a, t_b]) and no missing data in other sections, the dividing unit 50 defines the event data included in the corresponding section as Z and the event data not included in the corresponding section as Y. That is, the event data is divided as shown in (67) below.

[0136]

[0137] On the other hand, when simulating the presence of probabilistic missing data in the event data included in the training data set, the dividing unit 50 allocates the event data included in the training data set to missing data Y and missing data Z based on a preset probability for each event. That is, the event data is divided as shown in the following (68).

[0138]

[0139] The combining unit 60 combines the observed event with the complementary event as shown in (69) below. The combining unit 60 combines and sorts the sets of event data.

[0140]

[0141] Next, the process of learning the model parameters of the despreading unit 30a and the embedding unit 10a will be described as follows (G-1) to (G-10). FIG. 5 is a flowchart showing an example of the process of learning the model parameters. (G-1) First, event data is randomly selected from the learning data set D (S41). This selected data is expressed as (70) below.

[0142]

[0143] (G-2) After S41, the dividing unit 50 divides the selected event data into data represented by the following (71) (S42).

[0144]

[0145] Y corresponds to the observed data, and Z corresponds to the missing data. When it is assumed that the missing section of the event data is known and there is no missing data in other sections, the divided data is used as shown in (73) below instead of the data shown in (72) below. a , t b is the missing section [t a , t b]corresponds to

[0146]

[0147] (G-3) After S42, the data expressed by the following (74) is input to the embedding unit 10a, and the latent expression h is output (S43).

[0148]

[0149] (G-4) After S43, kε{1, ..., K} is determined (S44). K is the number of steps in the diffusion model and is a hyperparameter.

[0150] (G-5) After S44, the data expressed by the following (75) is applied to the spreading unit 20a, and the data to which noise has been added, expressed by the following (76), is output (S45).

[0151]

[0152] Furthermore, if the condition expressed by the following (77) is met and noise is added once by the spreading unit 20a, the value expressed by the following (78) is obtained.

[0153]

[0154] By repeatedly applying the data expressed by (75) above to the spreading unit 20a, the data expressed by (76) above is obtained as the data expressed by (79) below. By applying the data expressed by (75) above to the spreading unit 20a once more, the data expressed by (76) above is obtained as the data expressed by (80) below.

[0155]

[0156] (G-6) After S45, the despreading unit 30a receives the data represented by the following (81), and outputs the data represented by the following (82) from which the estimated noise has been removed (S46).

[0157]

[0158] In the above (81), t s , t eThere are two sets of Y and Z because they represent the periods Y and Z, respectively.

[0159] (G-7) After S46, the update unit 40 calculates the loss function (S47). Here, the loss function is expressed as, for example, the following (83).

[0160]

[0161] (G-8) After S47, if the arbitrary number of iterations related to the calculation of the loss function has not been reached (No in S48), return to (G-4) above.

[0162] (G-9) After S47, when an arbitrary number of iterations related to the calculation of the loss function is reached (Yes in S48), the update unit 40a updates the model parameters of the despreading unit 30a and the embedding unit 10a by the backpropagation method or the like (S49).

[0163] (G-10) After S49, if the number of iterations related to the parameter update has not been reached (No in S50), return to (G-1) above. If the number of iterations has been reached (Yes in S50), the series of processes related to learning the model parameters ends.

[0164] Next, the process of predicting (estimating) event data after the learning will be described as follows (H-1) to (H-4). Fig. 6 is a flowchart showing an example of the process of predicting event data.

[0165] (H-1) First, the observation data and t s , t e is input to the embedding unit 10a, and the embedding unit 10a outputs a latent expression h (S61).

[0166]

[0167] (H-2) After S61, the following data expressed by (85) is randomly generated (S62).

[0168] Here, as explained in the first embodiment, the number of events n is determined by, for example, a binomial distribution, and the number of events n is determined by a uniform distribution (t s , t e] are sampled and sorted.

[0169]

[0170] When piecewise missingness in the observation data is assumed, t s , t e is set according to the expected classification.

[0171] (H-3) After S62, the data generated in S62 and the latent expressions h, k, and t output in S61 are sent to the despreading unit 30a. s and t e is applied K times, and the prediction result expressed by the following (86) is output (S63).

[0172]

[0173] (H-4) After S63, the combining unit 60 combines the observation data described in S61 with the prediction result in S62 (S64).

[0174] In this embodiment, the data represented by the following (88) is used to predict the data represented by the following (87), making it possible to interpolate the relationship between events in both observations and missing events.

[0175]

[0176] In this embodiment, model parameters are learned in order to fill in missing event data, so high accuracy can be expected. Furthermore, this embodiment can also handle probabilistic missing event data.

[0177] Fig. 7 is a block diagram showing an example of the hardware configuration of an event prediction device 100 according to an embodiment of the present invention. In the example shown in Fig. 7, the event prediction device 100 according to the above embodiment is configured, for example, by a server computer or a personal computer, and has a hardware processor 111A such as a CPU. A program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to this hardware processor 111A via a bus 115. The same applies to the event prediction device 100a shown in Fig. 4.

[0178] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information to and from a communication network NW. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.

[0179] An input device 300 and an output device 400 that are attached to the event prediction device 100 and used by a user or the like are connected to the input / output interface 113. The input / output interface 113 takes in operation data input by a user or the like via the input device 300, such as a keyboard, a touch panel, a touchpad, or a mouse, and outputs output data to an output device 400, which includes a display device using a liquid crystal or an organic electroluminescence (EL) display, for display. Note that the input device 300 and the output device 400 may be devices built into the event prediction device 100, or may be input devices and output devices of other information terminals that can communicate with the event prediction device 100 via a network NW.

[0180] The program memory 111B is a non-transitory tangible storage medium that is a combination of a non-volatile memory that can be written to and read from at any time, such as a hard disk drive (HDD) or a solid state drive (SSD), and a non-volatile memory such as a read only memory (ROM), and stores programs necessary to execute various control processes, etc., according to one embodiment.

[0181] The data memory 112 is a tangible storage medium that is, for example, a combination of the above-mentioned nonvolatile memory and a volatile memory such as RAM (Random Access Memory), and is used to store various data acquired and created during various processing steps.

[0182] The event prediction device 100 according to an embodiment of the present invention can be configured as a data processing system or information processing device having a software-based processing function unit.

[0183] The storage system used as a work memory or the like by the event prediction device 100 can be configured by using the data memory 112 shown in Fig. 7. However, these configured storage areas are not essential components within the event prediction device 100, and may be areas provided in an external storage medium such as a USB (Universal Serial Bus) memory, or in a storage system such as a database server located in the cloud.

[0184] The processing function unit can be realized by having the hardware processor 111A read and execute a program stored in the program memory 111B, but the processing function unit may also be realized in various other forms, including an integrated circuit such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0185] The methods described in each embodiment may be stored as a program (software means) that can be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), or may be transmitted and distributed via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only executable programs but also tables and data structures) that the computer executes. The computer that realizes this device reads the program stored on the recording medium and, in some cases, configures the software means using the configuration program, and executes the above-described processing by having the operation controlled by this software means. The term "recording medium" as used herein is not limited to a storage medium for distribution, but also includes a storage medium such as a magnetic disk or semiconductor memory installed inside the computer or in a device connected via a network.

[0186] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.

[0187] REFERENCE SIGNS LIST 100, 100a... event prediction device 10, 10a... embedding unit 20, 20a... diffusion unit 30, 30a... dediffusion unit 40, 40a... update unit 50... division unit 60... combination unit

Claims

1. A first output unit that outputs a latent representation of the first event data by projecting the first event data indicating the events that occurred from the first timing to the second timing and the occurrence timings of the events into a latent space; and a second output unit that estimates a noise component added to the second event data, which is random event data from the second timing to the third timing, based on the latent representation and outputs event data with the noise component removed. An information processing apparatus comprising:

2. The first output unit projects, into the latent space, fourth event data, which is event data from the first timing to the second timing, in third event data, which is random event data from the first timing to the third timing, and outputs a latent representation of the fourth event data. Based on fifth event data, which is event data from the second timing to the third timing in the third event data, the first output unit further includes a third output unit that outputs sixth event data in which the noise component is added to the fifth event data multiple times and seventh event data in which the noise component is added to the sixth event data once. The second output unit outputs eighth event data with the noise component removed from the seventh event data based on the latent representation of the seventh event data and the fourth event data, and further includes an update unit that updates the parameters of the models used in the first and second output units based on a loss function related to the sixth event data and the eighth event data. The information processing apparatus according to claim 1.

3. The update unit updates the parameters of the models used in the first and second output units based on a loss function including a differentiable function that measures the difference between the sixth event data and the eighth event data. The information processing apparatus according to claim 2.

4. A first output unit that outputs a latent representation of the first event data by projecting first event data indicating events that occurred from a first timing to a second timing and the occurrence timings of the events into a latent space; a second output unit that estimates a noise component added to the second event data based on the second event data, which is random time-series data, and the latent representation, and outputs event data from which the noise component has been removed; and a combining unit that generates, as event data in which the loss of the first event data is complemented, combined event data obtained by combining the first event data and the event data output by the second output unit. An information processing apparatus comprising:

5. A method performed by an information processing apparatus, the method comprising: outputting, by a first output unit of the information processing apparatus, a latent representation of first event data by projecting first event data indicating events that occurred from a first timing to a second timing and the occurrence timings of the events into a latent space; and estimating, by a second output unit of the information processing apparatus, a noise component added to second event data, which is random event data from a second timing to a third timing, based on the second event data and the latent representation, and outputting event data from which the noise component has been removed. An information processing method comprising: