Information processing apparatus, information processing method, and program

By using weighted learning and region-specific meteorological factors, the information processing device enhances visibility estimation accuracy for both low and high values, addressing the biases in conventional models.

JP2025160657APending Publication Date: 2025-10-23KK TOSHIBA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024063340
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Conventional visibility estimation models struggle to accurately estimate low visibility values due to their infrequent occurrence and data truncation, leading to biased estimation results.

Method used

The information processing device employs multiple estimation models trained with weighted learning data, using different determination methods to assign higher weights to less frequent low visibility values, and incorporates region-specific meteorological factors to enhance accuracy.

Benefits of technology

This approach improves the estimation accuracy of both low and high visibility values by training models that focus on underrepresented data, resulting in more precise visibility predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160657000001_ABST
    Figure 2025160657000001_ABST
Patent Text Reader

Abstract

To enable highly accurate visibility estimation.SOLUTION: An information processing apparatus includes a processing unit. The processing unit determines a weight using different m (m is an integer equal to or larger than 2) determination methods including a first determination method that determines a larger weight for learning data including a first visibility value of which a degree of appearance is lower than those of other visibility values, than other visibility values, the method for determining the weight according to visibility values included in multiple pieces of learning data using the multiple pieces of learning data each including one or more explanatory variables and a visibility value which is an objective variable. The processing unit trains m estimation models configured to receive, as input, one or more explanatory variables to estimate visibility values, using multiple pieces of learning data with the weights which are determined by the m determination methods, respectively.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] For example, traffic control systems use meteorological data to estimate visibility for road traffic or aviation safety. In conventional visibility estimation, an estimation model is trained by machine learning using meteorological data, including data representing measured visibility, as training data, to minimize the overall estimation error. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6901647 [Patent Document 2] Patent No. 6714697 [Patent Document 3] Patent No. 4404220 [Non-patent literature]

[0004] [Non-Patent Document 1] Steininger, M., Kobs, K., Davidson, P. et al., “Density-based weighting for imbalanced regression.”, Mach Learn 110, 2187-2211 (2021). Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the prior art, there are cases where it is not possible to learn an estimation model that can estimate visibility with high accuracy. An object of the present invention is to provide an information processing device, an information processing method, and a program that enable visibility estimation with higher accuracy. [Means for solving the problem]

[0006] An information processing device according to an embodiment includes a processing unit. The processing unit determines weights using m (m is an integer equal to or greater than 2) mutually different determination methods, each of which uses a plurality of training data sets, each of which includes one or more explanatory variables and a visibility value that is a target variable, to determine weights according to the visibility values ​​included in the plurality of training data sets, including a first determination method that determines a weight greater for training data sets that include a first visibility value whose occurrence frequency is lower than other visibility values, than for other visibility values. The processing unit uses the plurality of training data sets to which weights determined by the m determination methods have been assigned, to train m estimation models that estimate visibility values ​​by inputting one or more explanatory variables. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 10 is a diagram showing an example of visibility data. [Figure 2] FIG. 1 is a block diagram of an information processing apparatus according to a first embodiment. [Figure 3] FIG. 10 is a diagram showing an example of learning input data. [Figure 4] FIG. 10 is a diagram for explaining a specific example of selection processing. [Figure 5] FIG. 10 is a diagram for explaining a specific example of selection processing. [Figure 6] FIG. 10 is a diagram for explaining a specific example of a learning process. [Figure 7] FIG. 10 is a diagram for explaining a specific example of an estimation process. [Figure 8] FIG. 10 is a diagram for explaining a specific example of a learning process of a probability model. [Figure 9] 4 is a flowchart of a learning process according to the first embodiment. [Figure 10] 4 is a flowchart of an estimation process according to the first embodiment. [Figure 11] 4 is a flowchart of a probability learning process in the first embodiment. [Figure 12] FIG. 10 is a diagram showing an example of learning data used in the model learning process. [Figure 13] FIG. 10 is a diagram showing an example of data used in a model-based estimation process. [Figure 14] FIG. 10 is a block diagram showing the configuration of an information processing apparatus according to a second embodiment. [Figure 15] FIG. 2 is a hardware configuration diagram of an information processing apparatus according to the first or second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of an information processing apparatus according to the present invention will be described in detail below with reference to the accompanying drawings.

[0009] As described above, conventional technologies have sometimes been unable to train an estimation model that can estimate visibility with high accuracy. One of the reasons for this is that the data representing visibility has the following characteristics: (F1) There is a bias in the frequency of occurrence of visibility values ​​(hereinafter referred to as visibility values). For example, small visibility values ​​(low visibility values) appear less frequently than large visibility values ​​(high visibility values). (F2) This is data for which an upper limit is set. For example, visibility data used by aircraft may be truncated if the value exceeds a certain value (e.g., 10 km) due to restrictions on flight altitude, etc. Because this is data for which an upper limit is set, there is a high possibility that high visibility values ​​close to the upper limit will occur more frequently, i.e., the situation described in (F1) above will occur.

[0010] Visibility value is a measure of how far ahead you can see, and is expressed, for example, as the maximum distance at which an object can be seen with the naked eye. Hereinafter, data that expresses visibility value may be referred to as visibility data. The scale of visibility data may change depending on the application. For example, the visibility that a person can see at eye height is several hundred meters, but for an aircraft it is several kilometers.

[0011] In the prior art, an estimation model is trained using visibility data with the above characteristics to minimize the overall estimation error. As a result, for example, high visibility values, which occur frequently, can be estimated accurately, but low visibility values, which occur less frequently, may not be estimated accurately.

[0012] Figure 1 shows an example of visibility data. Figure 1 shows the change in visibility over time, divided into actual values ​​and estimated values ​​(bold lines). The estimated values ​​show an example of visibility values ​​estimated using conventional technology.

[0013] In the example of Figure 1, the visibility value is truncated at an upper limit of 10. In addition, there are fewer data points for low visibility values, for example, values ​​of 3 or less, than for high visibility values. In other words, there is a bias in the frequency of occurrence. For this reason, while the estimation values ​​using conventional technology estimate high visibility values ​​with relatively high accuracy, it fails to estimate low visibility values.

[0014] Furthermore, the phenomena (factors) that affect visibility can differ depending on the location (region). For example, visibility is often reduced (visibility values ​​are low) due to PM2.5 in China, dense fog in India, and snowstorms in Japan. For this reason, when estimating visibility using meteorological data, it is desirable to select factors that are more appropriate for the region being estimated and use the selected factors as explanatory variables to estimate visibility (the target variable).

[0015] (First embodiment) The information processing device of the first embodiment learns (constructs) two estimation models, including an estimation model for more accurately estimating visibility values ​​that occur less frequently (e.g., low visibility values). The information processing device of this embodiment also estimates visibility values ​​for a target estimation period using data for estimation (estimation data) in the learned estimation model. The information processing device of this embodiment also selects one or more factors (visibility factors) that are effective for estimating visibility from multiple factors (multiple types of meteorological data), and uses the selected visibility factors to learn an estimation model and perform estimation using the estimation model.

[0016] In this embodiment, two estimation models are learned as described above, and estimation processing is performed using the two estimation models. The number m of estimation models is not limited to 2, and may be 3 or more. A configuration including the case where m is 3 or more will be described in the second embodiment.

[0017] 2 is a block diagram showing an example of the configuration of the information processing device 100 according to the first embodiment. As shown in FIG. 2, the information processing device 100 includes a learning control unit 10 and an estimation control unit 20.

[0018] The learning control unit 10 controls the learning process of the estimation model, and the estimation control unit 20 controls the estimation process using the learned estimation model.

[0019] At least a part of each of the above units (the learning control unit 10 and the estimation control unit 20) may be realized by one or more processing units. Each of the above units is realized, for example, by one or more processors. For example, each of the above units may be realized by having a processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) execute a program, that is, by software. Each of the above units may be realized by a processor such as a dedicated IC (Integrated Circuit), that is, by hardware. Each of the above units may be realized by a combination of software and hardware. When multiple processors are used, each processor may realize one of the units, or may realize two or more of the units.

[0020] Furthermore, the information processing device 100 may be physically configured as one device, or may be physically configured as multiple devices. For example, the information processing device 100 may be constructed in a cloud environment. Furthermore, each unit within the information processing device 100 may be distributed across multiple devices. For example, the information processing device 100 (information processing system) may be configured to include a device (e.g., a learning device) that includes the learning control unit 10 and a device (e.g., an estimation device) that includes the estimation control unit 20.

[0021] At least some of the components of the information processing device 100 may be formed as a chip. Furthermore, at least some of the components of the information processing device 100 may be incorporated into an SoC (System on Chip) such as an edge device. In this case, a storage unit (storage unit 121 described later) that stores data used for learning and a storage unit (storage unit 221 described later) that stores data used for estimation may be provided outside the SoC and may be accessible via an interface device.

[0022] The learning control unit 10 includes a storage unit 121, a selection unit 101, a weight determination unit 102_1, a weight determination unit 102_2, a model learning unit 103_1, a model learning unit 103_2, and a probability learning unit 104.

[0023] The storage unit 121 stores various types of information used by the learning control unit 10. For example, the storage unit 121 stores learning input data TD and visibility threshold TH used in the learning process, as well as output data output by the learning process. The output data includes, for example, a visibility factor VF, an estimation model EM_1, an estimation model EM_2, and a probabilistic model PRM. Note that, for convenience of explanation, only the learning input data TD is shown in the storage unit 121 in FIG. 2, but the storage unit 121 may also store other data (such as output data).

[0024] The training input data TD is data input for training the estimation models (estimation model EM_1, estimation model EM_2). For example, the training input data TD includes multiple types of meteorological data for multiple dates and times (timestamps). The meteorological data may be any type of data, such as atmospheric pressure, dew point temperature, wind speed, relative humidity, precipitation, temperature, and air pollution index (e.g., PM2.5, PM10, NOx, CO, NMHC, etc.).

[0025] Fig. 3 is a diagram showing an example of learning input data TD. In the example of Fig. 3, the learning input data TD includes a timestamp, atmospheric pressure, dew point temperature, wind speed, relative humidity, precipitation amount, PM10, temperature, PM2.5, and visibility.

[0026] The multiple meteorological data (factors) included in the learning input data TD correspond to multiple candidates for the visibility factor VF, which serves as an explanatory variable. That is, one or more explanatory variables determined by the visibility factor VF and a visibility value (visibility data), which serves as a response variable, are selected from the learning input data TD, and multiple learning data including the selected explanatory variables and response variables are generated. Hereinafter, the data selected as the visibility factor VF may be referred to as learning visibility factor data.

[0027] The one or more explanatory variables can be any data, but for example, some or all of the following data: -Visibility data observed in the past Past observed weather data other than visibility Weather data forecast

[0028] The visibility value, which is the objective variable, corresponds to the visibility value used as the correct answer data during learning. For example, when estimating the visibility value for the estimation target period after a certain time (for example, every hour from 1 hour to 6 hours) from previously observed weather data, the visibility value observed after a certain time from the time when the weather data corresponding to the explanatory variable was observed is used as the correct answer data.

[0029] In the example of Fig. 3, data 301 indicating relative humidity, data 302 indicating PM2.5, and data 303 indicating visibility are selected as visibility factors (explanatory variables). That is, in the example of Fig. 3, the learning visibility factor data includes relative humidity, PM2.5, and visibility.

[0030] The visibility threshold TH is a threshold for dividing the learning data including the learning visibility factor data into a plurality of groups. Since the visibility range changes depending on the application, the value of the visibility threshold TH can also be changed depending on the application.

[0031] The storage unit 121 can be configured with any commonly used storage medium, such as a flash memory, a memory card, a RAM (Random Access Memory), an HDD (Hard Disk Drive), an optical disk, etc. Some or all of the data stored in the storage unit 121 (learning input data TD, visibility threshold TH, visibility factor VF, estimation model EM_1, estimation model EM_2, and probabilistic model PRM) may be stored in physically different storage media, or may be stored in different storage areas of the same physically stored medium.

[0032] Returning to the description of Fig. 2, the selection unit 101 selects a visibility factor VF from the learning input data TD. For example, the selection unit 101 selects one or more candidates C1 (first candidates) that affect the estimation of the visibility value as the visibility factor VF from a plurality of candidates (weather data) of explanatory variables included in the learning input data TD. The selection unit 101 stores the selected visibility factor VF in, for example, the storage unit 121. The selection unit 101 also generates a plurality of learning data that includes the one or more selected visibility factors VF (candidates C1) as explanatory variables.

[0033] The visibility factor VF can be selected, for example, by the following method. (M1) Select the visibility factor VF according to knowledge from experts, etc. (M2) Select a candidate whose cross-correlation with the visibility value is greater than other candidates as the visibility factor VF. (M3) A candidate whose estimation accuracy by the model M_VF, which estimates visibility values ​​using the candidate as an explanatory variable, is greater than that of other candidates is selected as the visibility factor VF.

[0034] (M2) will be described. For example, assume that the learning input data TD includes air pressure, dew point temperature, wind speed, relative humidity, precipitation, PM10, temperature, and PM2.5 as candidate factors. The selection unit 101 calculates the correlation coefficient between visibility and each candidate, and selects the candidate for which the calculated correlation coefficient indicates a high correlation as the visibility factor VF. For example, in the case of a correlation coefficient in which a higher value indicates a higher correlation, the selection unit 101 selects, as the visibility factor VF, a candidate for which the correlation coefficient value is greater than a threshold, a certain number of candidates in descending order of the correlation coefficient value, or candidates with the highest integration ratio in descending order of the correlation coefficient value.

[0035] For example, if the threshold is 0.5 and the correlation coefficient between each candidate and visibility is the following value, the selection unit 101 selects the relative humidity and PM2.5 whose correlation coefficient value is greater than the threshold 0.5 as the visibility factor VF. Atmospheric pressure: 0.40 ·Dew point temperature: 0.45 ·Wind speed: 0.20 Relative humidity: 0.82 ·Precipitation: 0.34 PM10: 0.28 Temperature: 0.48 PM2.5:0.90

[0036] (M3) will be described. For example, the selection unit 101 divides the training input data TD into training data and verification data. The selection unit 101 generates various combinations of candidates included in the training data, and trains the model M_VF using training data that includes each combination as an explanatory variable. The model M_VF may be any model that estimates visibility, but for example, it is a model that has the same structure as the estimation model EM_1 or the estimation model EM_2 and is constructed by the same training method.

[0037] The selection unit 101 evaluates the trained model M_VF using the validation data and calculates an evaluation score. The evaluation score may be any index, such as the root mean squared error (RMSE), the mean absolute error (MAE), or the coefficient of determination (R2).

[0038] The selection unit 101 selects the candidate included in the combination with the best evaluation score as the visibility factor VF. As in the above, when the learning input data TD includes eight candidates, namely, atmospheric pressure, dew point temperature, wind speed, relative humidity, precipitation amount, PM10, temperature, and PM2.5, the number of candidate combinations is 255 (=2 8 -1) As shown below, some examples of combinations are listed. {Pressure}, {Dew point temperature}, {Wind speed}, {Relative humidity}, {Precipitation}, {PM10}, {Temperature}, {PM2.5}, ···, {Relative humidity, PM2.5}, {Dew point temperature, PM10}, ···, {Pressure, Dew point temperature, Wind speed, Relative humidity}, {Precipitation, PM10, Temperature, PM2.5}, ···, {Pressure, Dew point temperature, Wind speed, Relative humidity, Precipitation, PM10, Temperature, PM2.5}

[0039] In the above example, visibility is excluded from the candidates for factors to be included in the combination. In this case, the selection unit 101 may further include visibility in the visibility factor VF. The selection unit 101 may also include visibility as a candidate for factors to be included in the combination and perform the above process.

[0040] When generating candidate combinations, the selection unit 101 may use sequential forward selection (SFS), sequential backward selection (SBS), a genetic algorithm, or the like.

[0041] Returning to the explanation of Fig. 2, the weight determination unit 102_1 and the weight determination unit 102_2 determine weights according to visibility values ​​included in a plurality of learning data using mutually different determination methods. In this embodiment, the number m of determination methods is 2.

[0042] The determination method is a method for determining weights according to visibility values ​​included in a plurality of learning data by using a plurality of learning data. A weight is determined for each of the plurality of learning data (each sample), and the determined weight is assigned to the corresponding learning data. The two determination methods include a determination method DT_A (first determination method) used by the weight determination unit 102_1 and a determination method DT_B used by the weight determination unit 102_2.

[0043] The decision method DT_A assigns a weight W to the training data including the visibility value VA (first visibility value) whose frequency of occurrence is smaller than the other visibility values. A The visibility value VA is, for example, a low visibility value. The determination method DT_B may be any method different from the determination method DT_A. For example, the determination method DT_B may be a method using kernel density estimation to determine the weight W B This is a method for determining the above (second determination method).

[0044] The weight determination unit 102_1 determines weights to be assigned to the training data for training the estimation model EM_1 in accordance with a determination method DT_A. First, the weight determination unit 102_1 divides the training data into m groups using a visibility threshold TH.

[0045] When m=2 as in this embodiment, the weight determination unit 102_1 classifies the learning data including visibility values ​​(low visibility values) equal to or less than the visibility threshold TH into a group corresponding to low visibility values ​​(hereinafter referred to as a low visibility group). Visibility values ​​equal to or less than the visibility threshold TH are examples of visibility values ​​VA that appear less frequently than other visibility values. Furthermore, the weight determination unit 102_1 classifies the learning data including visibility values ​​(high visibility values) greater than the visibility threshold TH into a group corresponding to high visibility values ​​(hereinafter referred to as a high visibility group).

[0046] The weight determining unit 102_1 determines a predetermined value for each group as a weight WA The value β1 for the group corresponding to the visibility value of interest is set to a value greater than the value β2 for the other groups. For example, when focusing on low visibility, the weight determiner 102_1 assigns a weight W A The weight determining unit 102_1 determines β1=9999 as the value of W A Determine the value of β2=1.

[0047] The weight determination unit 102_2 determines the weight W to be assigned to the training data for training the estimation model EM_2. B is determined according to the determination method DT_B. Hereafter, the weight W is determined using kernel density estimation. B An example of the determination method DT_B for determining the value of the variable DT_B will be described.

[0048] As a method using kernel density estimation, a weight determination method for rare values ​​based on kernel density estimation (for example, Non-Patent Document 1) can be used. For example, the weight determination unit 102_2 determines the weight W of the visibility y using the following equation (1): B Determine (y).

number

[0049] Note that determining the weight for visibility y is equivalent to determining the weight for training data that includes visibility y. Visibility y refers to the visibility value included in a certain training data among multiple training data. Hereinafter, Y represents the set of all visibility y included in multiple training data.

[0050] In equation (1), α and ε are predetermined parameters. For example, 0.7 is used as the value of α, and 0.4 is used as the value of ε. p'(y) is expressed by the following equation (2). p(y) in equation (2) is expressed by the following equation (3).

number

number

[0051] N is the number of samples of multiple learning data, h is the bandwidth (for example, 10), and K is the kernel function. The kernel function is, for example, a standard Gaussian function, and is expressed by the following equation (4). Note that x is expressed as x=(yy i ) / h.

number

[0052] The weight determination unit 102_1 determines the weight W B Considering the weight W A For example, the weight determination unit 102_1 may determine the weight W using the following equation (5): A may be determined.

number

[0053] The threshold is the weight W B This corresponds to the visibility threshold TH for dividing the value of (y). When focusing on estimating low visibility accurately, the threshold is, for example, the weight W of the low visibility group. B (y) is the minimum value. For example, the threshold may be set to 0.5, β1 to 9999, and β2 to 1.

[0054] The model learning unit 103_1 and the model learning unit 103_2 correspond to model learning units that learn m estimation models using a plurality of learning data to which weights determined by m (two in this embodiment) determination methods are assigned. The model learning unit 103_1 and the model learning unit 103_2 correspond to the weight determination unit 102_1 (determination method DT_A) and the weight determination unit 102_2 (determination method DT_B), respectively.

[0055] For example, the model learning unit 103_1 determines the weight W determined by the weight determination unit 102_1 according to the determination method DT_A. AThe model learning unit 103_2 learns the estimation model EM_1 using the learning data to which the weights W are assigned. B The estimation model EM_2 is trained using the training data with

[0056] Any method may be used to learn the models (estimated model EM_1, estimated model EM_2) using weighted learning data, but for example, a learning method using the following technology may be applied. Gradient boosting (e.g., Light Gradient Boosting Machine: LGBM) Deep Learning ·Long Short Term Memory (LSTM) Transformer Ridge regression Linear regression DLinear

[0057] Model learning unit 103_1 and model learning unit 103_2 may use the same learning method, or may use different learning methods.

[0058] The probability learning unit 104 learns the probability model PRM using multiple learning data. The probability model PRM is a model that can estimate the probability that a visibility value is a visibility value VA (a visibility value with a low occurrence frequency) by inputting one or more explanatory variables. The probability model PRM can also be interpreted as a model that estimates the probability that a visibility value belongs to two groups (a low visibility group or a high visibility value group).

[0059] In this embodiment, the probabilistic model PRM is a model that outputs the probability that a visibility value is the visibility value VA and the probability that the visibility value is not VA. In other words, the probabilistic model PRM corresponds to a model that estimates the probability that a visibility value corresponds to the low visibility group and the probability that a visibility value corresponds to the high visibility group. The probabilistic model PRM can also be interpreted as a classifier that classifies visibility into a low visibility group and a high visibility group.

[0060] Any method may be used for learning the probabilistic model PRM, but for example, a learning method using the following technology can be applied. Logistic regression ·K-nearest neighbor method Support Vector Machine Decision Tree Deep Learning Autoencoder

[0061] The probability learning unit 104 may use weighted learning data to learn the probability model PRM. For example, the probability learning unit 104 may assign a large weight to learning data (low visibility group) including visibility values ​​VA with a low occurrence frequency, as in the determination method DT_A. The weight value assigned to each group may be the same as or different from the value used by the weight determination unit 102_1.

[0062] The estimation model EM_1 learned by the model learning unit 103_1, the estimation model EM_2 learned by the model learning unit 103_2, and the probability model PRM learned by the probability learning unit 104 are stored in the storage unit 121, for example.

[0063] Next, a description will be given of a specific example of the estimation control unit 20. As shown in FIG.

[0064] The storage unit 221 stores various types of information used by the estimation control unit 20. For example, the storage unit 221 stores input data for estimation ED used in the estimation process, and an estimated value EV that is the calculation result by the estimated value calculation unit 203. For ease of explanation, only the input data for estimation ED is shown in the storage unit 221 in FIG. 2, but the storage unit 221 may also store other data (such as the estimated value EV).

[0065] The estimation input data ED is data input for estimation processing using estimation models (estimation model EM_1, estimation model EM_2). The estimation input data ED can have the same data format as the training input data TD. That is, the estimation input data ED includes multiple types of weather data for multiple dates and times (timestamps).

[0066] The storage unit 221 can be configured from any commonly used storage medium such as a flash memory, a memory card, a RAM (Random Access Memory), an HDD (Hard Disk Drive), and an optical disk.

[0067] The estimation units 201_1 and 201_2 correspond to estimation units that estimate m (two) visibility values ​​for one or more explanatory variables for estimation using m (two in this embodiment) estimation models (estimation model EM_1, estimation model EM_2), respectively.

[0068] For example, the estimation unit 201_1 first extracts data on the visibility factor VF from the input data for estimation ED. Hereinafter, the extracted data will be referred to as visibility factor data for estimation. The visibility factor data for estimation corresponds to input data including explanatory variables to be input to the estimation model EM_1. The estimation unit 201_1 inputs the visibility factor data for estimation to the estimation model EM_1 and calculates an estimated value EV_1 of the visibility value.

[0069] The estimation unit 201_2 inputs the visibility factor data for estimation into the estimation model EM_2 and calculates an estimated value EV_2 of the visibility value. The estimation unit 201_2 may use the visibility factor data for estimation extracted by the estimation unit 201_1, or may use the visibility factor data for estimation extracted from the estimation input data ED as data of the visibility factor VF.

[0070] The probability calculation unit 202 inputs the visibility estimation factor data into the probabilistic model PRM and calculates the probability. As described above, the calculated probability includes the probability of the low visibility group and the probability of the high visibility group. This probability can also be interpreted as including the probability of the estimated value EV_1 (hereinafter, probability PR_1) and the probability of the estimated value EV_2 (hereinafter, probability PR_2).

[0071] The estimated value calculation unit 203 calculates the final estimated value EV of the visibility value using m (two) estimated values ​​of the visibility value (estimated value EV_1, estimated value EV_2). The estimated value EV is stored in, for example, the storage unit 221. For example, the estimated value calculation unit 203 calculates the estimated value EV using the estimated value EV_1, estimated value EV_2, probabilities PR_1, and probabilities PR_2.

[0072] As a method for calculating the estimated value EV using probability, for example, the following method can be applied. Between the estimated values ​​EV_1 and EV_2, the estimated value with the larger corresponding probability is calculated as the estimated value EV. · The weighted average of the estimated values ​​EV_1 and EV_2, using the probabilities PR_1 and PR_2 as weights, is calculated as the estimated value EV.

[0073] The estimation control unit 20 may have a function to output the estimated value EV. For example, the estimation control unit 20 may display the estimated value EV on a display device (such as a liquid crystal display). The estimation control unit 20 may also transmit the estimated value EV to an external device connected via a network.

[0074] Next, a description will be given of an example of processing by the information processing device 100. Figures 4 and 5 are diagrams for explaining a specific example of selection processing by the selection unit 101.

[0075] FIG. 4 shows an example in which an inappropriate visibility factor is selected. FIG. 4 shows an example in which relative humidity is selected as an inappropriate visibility factor. Data 401 represents an example of visibility included in training data. Data 411 represents an example of an estimation result by model M_VF trained using training data that includes only relative humidity as a visibility factor. When an inappropriate visibility factor is used, a peak shift may occur in the estimation result, as shown by data 411 in FIG. 4. When a peak shift occurs, the value of evaluation scores such as RMSE deteriorates.

[0076] FIG. 5 shows an example in which an appropriate visibility factor is selected. FIG. 5 shows an example in which a combination of relative humidity and PM2.5 is selected as an appropriate visibility factor. Data 412 represents an example of an estimation result by model M_VF trained using training data including a combination of relative humidity and PM2.5 as visibility factors. When an appropriate visibility factor is used, as shown in data 412 in FIG. 5, no peak shift occurs in the estimation result. As a result, a value indicating a higher evaluation is calculated as an evaluation score such as RMSE. If the evaluation score is the best value, the selection unit 101 selects relative humidity and PM2.5 as the visibility factor VF.

[0077] Next, a specific example of the learning process for each model (estimation model EM_1, estimation model EM_2, probabilistic model PRM) will be described. FIG. 6 is a diagram for explaining a specific example of the model learning process. In the example of FIG. 6, visibility, relative humidity, and PM2.5 are used as visibility factors. FIG. 6 also shows an example in which the estimation target period is every hour from one hour to six hours later.

[0078] The weight determination unit 102_1 focuses on low visibility and determines a weight so that the weight is larger for a group corresponding to a low visibility value. The model learning unit 103_1 uses the learning data to which the weight is assigned to learn an estimation model EM_1. As a result, the estimation model EM_1 becomes a model that can estimate low visibility values ​​with higher accuracy, as shown in an estimation result 601.

[0079] The weight determination unit 102_2 determines weights using, for example, kernel density estimation. The model learning unit 103_2 uses the training data to which the weights have been assigned to learn an estimation model EM_2. As a result, as shown in an estimation result 602, the estimation model EM_2 becomes a model that can estimate high visibility values ​​with higher accuracy than the estimation model EM_1.

[0080] The probability learning unit 104 learns a probability model that estimates the probability of low visibility (pl1 to pl6) and the probability of high visibility (ph1 to ph6) for up to six hours ahead.

[0081] Next, a specific example of estimation processing using the trained model will be described. FIG. 7 is a diagram for explaining a specific example of estimation processing using a model. In the example of FIG. 7, visibility, relative humidity, and PM2.5 are used as visibility factors. Also, FIG. 6 shows an example in which the estimation target period is every hour from one hour to six hours later.

[0082] The estimation unit 201_2 extracts visibility estimation factor data from the estimation input data ED, inputs the extracted visibility estimation factor data to the estimation model EM_2, and calculates an estimated value EV_2 of visibility up to six hours ahead. Data 702 corresponds to the estimated value EV_2 up to six hours ahead.

[0083] The estimation unit 201_1 inputs the visibility factor data for estimation into the estimation model EM_1 to calculate the estimated value EV_1 of the visibility up to six hours ahead. Data 701 corresponds to the estimated value EV_1 up to six hours ahead.

[0084] The probability calculation unit 202 inputs the visibility factor data for estimation into the probabilistic model PRM and estimates the probability of low visibility and the probability of high visibility up to six hours ahead.

[0085] The estimated value calculation unit 203 calculates the weighted average of the estimated values ​​EV_1 and EV_2, with the probability of low visibility and the probability of high visibility used as weights for the estimated values ​​EV_1 and EV_2, respectively. Data 711 corresponds to the estimated value EV.

[0086] Next, a specific example of the learning process of the probabilistic model PRM will be described. Fig. 8 is a diagram for explaining a specific example of the learning process of the probabilistic model PRM. In the example of Fig. 8, low visibility is defined as a distance of 3 km or less, and the learning visibility factor data is divided into a low visibility group and a high visibility group.

[0087] For example, if the visibility included in the training visibility factor data is 3 km or less, the probability learning unit 104 assigns the label "low," indicating low visibility, to the training visibility factor data, and if the visibility is greater than 3 km, the probability learning unit 104 assigns the label "high," indicating high visibility, to the training visibility factor data. The probability learning unit 104 uses the labeled training visibility factor data to learn a probabilistic model PRM that estimates the probability of low visibility and the probability of high visibility.

[0088] Next, a description will be given of the flow of the learning process by the information processing apparatus 100 of the first embodiment. Fig. 9 is a flowchart showing an example of the learning process in the first embodiment.

[0089] The selection unit 101 reads out the learning input data TD from the memory unit 121, selects a visibility factor VF from the learning input data TD, and generates multiple learning data including learning visibility factor data including the selected visibility factor VF and visibility data corresponding to the objective variable (step S101).

[0090] The weight determining unit 102_2 determines a weight for each of the plurality of learning data (step S102). The model learning unit 103_2 learns an estimation model EM_2 using the plurality of learning data to which the weights have been assigned, and outputs the learned model (step S103).

[0091] The weight determination unit 102_1 divides the training data into two visibility groups using the visibility threshold TH in consideration of the weights determined by the weight determination unit 102_2, and determines a weight for each of the training data (step S104). The model training unit 103_1 trains and outputs an estimation model EM_1 using the weighted training data (step S105).

[0092] The probability learning unit 104 divides the learning data into two visibility groups, and learns and outputs a probability model PRM that estimates the probability of each group (step S106).

[0093] Next, a description will be given of the flow of the estimation process performed by the information processing device 100 of the first embodiment. Fig. 10 is a flowchart showing an example of the estimation process in the first embodiment.

[0094] The estimation unit 201_1 reads out the input data for estimation ED from the storage unit 221, selects a visibility factor VF from the input data for estimation ED, and generates visibility factor data for estimation including the selected visibility factor VF (step S201). The visibility factor data for estimation is output to the estimation unit 201_2 and the probability calculation unit 202.

[0095] The estimation unit 201_1 inputs the visibility factor data for estimation into the estimation model EM_1 and calculates an estimated value EV_1 (step S202). The estimation unit 201_2 inputs the visibility factor data for estimation into the estimation model EM_2 and calculates an estimated value EV_2 (step S203). The probability calculation unit 202 inputs the visibility factor data for estimation into the probability model PRM and calculates a probability PR_1 of the estimated value EV_1 and a probability PR_2 of the estimated value EV_2 (step S204).

[0096] The estimated value calculation unit 203 calculates the final estimated value EV using the estimated value EV_1, the estimated value EV_2, the probabilities PR_1 and PR_2 (step S205), and outputs the calculated estimated value EV (step S206).

[0097] Next, the flow of probability learning processing by the information processing device 100 of the first embodiment will be described. The probability learning processing is processing for learning the probability model PRM, and corresponds to the processing of step S106 in Fig. 9. Fig. 11 is a flowchart showing an example of the probability learning processing in the first embodiment.

[0098] The probability learning unit 104 acquires learning data generated by, for example, the selection unit 101 (step S301). The probability learning unit 104 classifies the learning data into a plurality of groups (low visibility group, high visibility group) using a visibility threshold TH (step S302). The probability learning unit 104 generates learning data in which a label is assigned to each group (step S303).

[0099] The probability learning unit 104 uses the labeled learning data to learn a probability model PRM that estimates the probability of each group (step S304), and outputs the learned probability model PRM (step S305).

[0100] Next, an example of data used in the learning process and estimation process will be described. Fig. 12 is a diagram showing an example of learning data used in the learning process of the models (estimation model EM_1, estimation model EM_2, and probabilistic model PRM).

[0101] In the example of Figure 12, the visibility factors are relative humidity, PM2.5, and visibility. Therefore, the learning visibility factor data includes relative humidity (t), PM2.5 (t), and visibility (t), which represent the relative humidity, PM2.5, and visibility at the time (current time) t of each timestamp. Relative humidity (t-1), PM2.5 (t-1), and visibility (t-1) are the relative humidity, PM2.5, and visibility one hour before the current time t of each timestamp. Relative humidity (t-2), PM2.5 (t-2), and visibility (t-2) are the relative humidity, PM2.5, and visibility two hours before the current time t of each timestamp.

[0102] The objective variables include visibility (t+1), visibility (t+2), visibility (t+3), visibility (t+4), visibility (t+5), and visibility (t+6), which are the visibility 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, and 6 hours ahead of the current time t for each timestamp.

[0103] In the learning process, relative humidity (t-2), PM2.5 (t-2), visibility (t-2), relative humidity (t-1), PM2.5 (t-1), visibility (t-1), relative humidity (t), PM2.5 (t), and visibility (t) are the explanatory variables, and visibility (t+1), visibility (t+2), visibility (t+3), visibility (t+4), visibility (t+5), and visibility (t+6) are the objective variables.

[0104] The model learning units 103_1 and 103_2 learn estimation models (estimation model EM_1 and estimation model EM_2) that represent the relationship between explanatory variables and target variables. The probability learning unit 104 assigns labels to the values ​​of visibility (t+1), visibility (t+2), visibility (t+3), visibility (t+4), visibility (t+5), and visibility (t+6) at each timestamp using a visibility threshold TH, and learns a probability model PRM using data on the labeled explanatory variables.

[0105] FIG. 13 is a diagram showing an example of data (estimation visibility factor data) used in the estimation process using the models (estimation model EM_1, estimation model EM_2, and probabilistic model PRM).

[0106] 13, the visibility factors are relative humidity, PM2.5, and visibility. The visibility factor data for estimation includes the relative humidity (t), PM2.5 (t), and visibility (t) at the current time t for each timestamp, the relative humidity (t-1), PM2.5 (t-1), and visibility (t-1) one hour before the current time t, and the relative humidity (t-2), PM2.5 (t-2), and visibility (t-2) two hours before the current time t.

[0107] The estimation units 201_2 and 201_1 input the visibility factor data for estimation as explanatory variables into the estimation models EM_1 and EM_2, respectively, and calculate the visibility from 1 hour to 6 hours ahead, namely visibility (t+1), visibility (t+2), visibility (t+3), visibility (t+4), visibility (t+5), and visibility (t+6).

[0108] The probability calculation unit 202 inputs the visibility estimation factor data as explanatory variables into the probabilistic model PRM and calculates the probability of the estimated value EV_1 and the probability of the estimated value EV_2 for each estimation target period, i.e., visibility (t+1), visibility (t+2), visibility (t+3), visibility (t+4), visibility (t+5), and visibility (t+6), which are the visibility from 1 hour to 6 hours ahead.

[0109] Note that, up to this point, an example has been described in which the data that appear less frequently is data of low visibility values, but the data that appear less frequently is not limited to data of low visibility values ​​and may be data of any visibility value. For example, depending on the region, season, etc., data of high visibility values ​​may be data that appears less frequently. Even in such cases, the same procedure as above can be applied by treating the data that appear less frequently as data of high visibility values ​​and the data that appear more frequently as data of low visibility values.

[0110] By using a method for determining weights for rare values ​​based on kernel density estimation (for example, Non-Patent Document 1), it is possible to construct an estimation model with higher estimation accuracy compared to a method that does not use weights. In this embodiment, in addition to a method using kernel density estimation (Determination Method DT_B), a method for determining a larger weight for visibility values ​​that appear less frequently (Determination Method DT_A) is also used. This makes it possible to estimate visibility with higher accuracy, even when the frequency of appearance may be biased.

[0111] (Second embodiment) In the first embodiment, an example was described in which the number m of weight determination units, model learning units (estimation models), and estimation units is 2. As described above, m is not limited to 2, and may be 3 or more. In the second embodiment, an example of a configuration in which the number of each of these elements is generalized to m will be described.

[0112] Fig. 14 is a block diagram showing an example of the configuration of an information processing device 100-2 according to the second embodiment. As shown in Fig. 14, the information processing device 100-2 includes a learning control unit 10-2 and an estimation control unit 20-2.

[0113] Learning control unit 10-2 includes storage unit 121, selection unit 101, weight determination units 102_1 to 102_m, model learning units 103_1 to 103_m, and probability learning unit 104-2. Estimation control unit 20-2 includes storage unit 221, estimation units 201_1 to 201_m, probability calculation unit 202-2, and estimated value calculation unit 203-2.

[0114] The same components as those in FIG. 1, which is a block diagram of the information processing apparatus 100 according to the first embodiment, are denoted by the same reference numerals, and the description thereof will be omitted here.

[0115] The weight determining units 102_1 to 102_m determine weights using m different determining methods, respectively. Examples of the m determining methods will be described below.

[0116] For example, the m determination methods may be methods for determining weights for m visibility groups corresponding to m visibility ranges, in which case the visibility threshold TH may include (m-1) values ​​for dividing the visibility range into m.

[0117] The weight determination unit 102_j (j is an integer satisfying 1≦j≦m) uses the jth determination method, which is a method of determining a weight for the jth group among m groups divided using the visibility threshold TH to be a value greater than the weights of the other groups.

[0118] An example will be described in which the visibility value ranges from 0 to 10, m=4, and the visibility threshold TH includes three values: 2.5, 5.0, and 7.5. In this case, the visibility values ​​are divided into the following four groups: Visibility group G1: A group with visibility values ​​between 0.0 and 2.5 (equivalent to the low visibility group) Visibility group G2: A group with visibility values ​​greater than 2.5 and less than or equal to 5.0 Visibility group G3: A group with visibility values ​​greater than 5.0 and less than or equal to 7.5 Visibility group G4: A group with visibility values ​​greater than 7.5 and less than or equal to 10.0

[0119] In this case, for example, the weight determination unit 102_1 focuses on the visibility group G1 and determines a large weight for the visibility group G1, and determines small weights for the other three visibility groups. The determination method used by the weight determination unit 102_1 is similar to the determination method DT_A described above, that is, a weight W for training data including a visibility value VA whose occurrence frequency is smaller than other visibility values. A This corresponds to a method of determining visibility to be greater than other visibility values.

[0120] The weight determination unit 102_2 focuses on the visibility group G2 and determines a large weight for the visibility group G2, and determines small weights for the other three visibility groups. The weight determination unit 102_3 and the weight determination unit 102_4 focus on the visibility groups G3 and G4, respectively, and perform the same processing as the weight determination unit 102_2.

[0121] The method for determining the m weights is not limited to the above, and any other methods may be used as long as they are mutually different. For example, a weight determination method (second determination method) based on kernel density estimation using the training data included in the groups other than the low visibility group (visibility group G1) may be applied to some or all of the groups (visibility groups G2 to G4).

[0122] The model learning units 103_1 to 103_m learn corresponding estimation models among the estimation models EM_1 to EM_m using a plurality of learning data to which weights determined by corresponding determination methods among the m determination methods are assigned. For example, the model learning unit 103_j learns an estimation model EM_j using a plurality of learning data to which weights determined by the weight determination unit 102_j are assigned.

[0123] The probability learning unit 104-2 learns a probability model PRM that estimates the probability of each of the m groups. In this embodiment, the probability model PRM can be interpreted as a classifier that classifies training data (visibility values) into one of the m groups.

[0124] The estimation units 201_1 to 201_m estimate the visibility value using a corresponding estimation model among m estimation models (estimation models EM_1 to EM_m). For example, the estimation unit 201_j inputs visibility factor data for estimation to the estimation model EM_j and calculates an estimated value EV_j of the visibility value.

[0125] The probability calculation unit 202-2 inputs the visibility factor data for estimation into the probabilistic model PRM and calculates the probability PR_j of each of the m estimated values ​​EV_j.

[0126] The estimated value calculation unit 203-2 calculates the estimated value EV using m estimated values ​​EV_j and the probabilities PR_j of each of the m estimated values ​​EV_j. For example, when m=4, the estimated value calculation unit 203-2 calculates the estimated value EV as follows: Estimated value EV = Estimated value EV_1 × Probability PR_1 + Estimated value EV_2 × Probability PR_2 + Estimated value EV_3 × Probability PR_3 + Estimated value EV_4 × Probability PR_4

[0127] In this way, the information processing device of the second embodiment can achieve the same functions as those of the first embodiment even when m is 3 or more.

[0128] As described above, according to the first and second embodiments, it is possible to estimate visibility with higher accuracy.

[0129] Next, the hardware configuration of the information processing apparatus according to the first or second embodiment will be described with reference to Fig. 15. Fig. 15 is an explanatory diagram showing an example of the hardware configuration of the information processing apparatus according to the first or second embodiment.

[0130] The information processing device of the first or second embodiment includes a control device such as a CPU (Central Processing Unit) 51, a storage device such as a ROM (Read Only Memory) 52 or a RAM (Random Access Memory) 53, a communication I / F 54 that connects to a network and communicates, and a bus 61 that connects each part.

[0131] The programs executed by the information processing device of the first or second embodiment are provided in advance in the ROM 52 or the like.

[0132] The program executed by the information processing device of the first or second embodiment may be configured to be provided as a computer program product by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk).

[0133] Furthermore, the program executed by the information processing device of the first or second embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the program executed by the information processing device of the first or second embodiment may be provided or distributed via a network such as the Internet.

[0134] The program executed by the information processing device of the first or second embodiment can cause the computer to function as each part of the information processing device described above. In this computer, the CPU 51 can read the program from a computer-readable storage medium onto the main storage device and execute it.

[0135] A configuration example of the embodiment will be described below. (Configuration example 1) a method for determining weights according to visibility values ​​included in a plurality of learning data sets, each of which includes one or more explanatory variables and a visibility value that is a response variable, using a plurality of learning data sets, the weights being determined by m (m is an integer of 2 or more) different determination methods, including a first determination method for determining a weight for a piece of learning data set including a first visibility value whose occurrence frequency is lower than other visibility values, to be greater than the weight for other visibility values; training m estimation models that estimate the visibility value by inputting one or more of the explanatory variables using a plurality of the learning data to which weights determined by the m determination methods are assigned; Processing section An information processing device comprising: (Configuration example 2) The processing unit using a plurality of the learning data sets and inputting one or more of the explanatory variables, to learn a probabilistic model that estimates the probability that the visibility value is the first visibility value; The information processing device according to configuration example 1. (Configuration example 3) The processing unit selecting one or more first candidates that will affect the estimation of the visibility value from a plurality of candidates for the explanatory variable, and generating a plurality of the learning data including the one or more selected first candidates as one or more of the explanatory variables; The information processing device according to configuration example 1 or 2. (Configuration example 4) The processing unit selecting one or more of the first candidates having a higher accuracy of estimation by a model that estimates the visibility value using the candidates as explanatory variables than the other candidates; The information processing device according to configuration example 3. (Configuration Example 5) The processing unit selecting one or more of the first candidates having a cross-correlation with the visibility value that is greater than the other candidates; The information processing device according to configuration example 3. (Configuration Example 6) The one or more explanatory variables include at least one of relative humidity, temperature, and PM2.5. 6. The information processing device according to any one of configuration examples 1 to 5. (Configuration Example 7) The m determination methods include a second determination method that determines the weights using kernel density estimation. The information processing device according to any one of configuration examples 1 to 6. (Configuration Example 8) The processing unit Estimating the m visibility values ​​for one or more explanatory variables for estimation using the m estimation models, respectively; Calculating an estimate of the visibility value using m visibility values. The information processing device according to any one of configuration examples 1 to 7. (Configuration Example 9) The processing unit inputting one or more of the explanatory variables and using a probabilistic model that estimates the probability that the visibility value is the first visibility value, calculating the probability for the explanatory variables for estimation; calculating an estimate of the visibility value using the calculated probability and the m visibility values; The information processing device according to configuration example 8. (Configuration Example 10) using m (m is an integer of 2 or more) estimation models each for estimating a visibility value, which is a target variable, by inputting one or more explanatory variables, to estimate m visibility values ​​for the one or more explanatory variables; Calculating an estimate of the visibility value using m visibility values. a processing unit; the m estimation models are trained using a plurality of training data each including one or more of the explanatory variables and the visibility value, and the plurality of training data are assigned weights determined by m determination methods that are different from each other; The m determination methods are methods for determining weights according to the visibility values ​​included in the plurality of learning data, and include a first determination method for determining a weight for the learning data including a first visibility value whose appearance frequency is lower than other visibility values ​​so that the weight is greater than that for other visibility values. Information processing device. (Configuration Example 11) An information processing method executed by an information processing device, a step of determining weights by m (m is an integer of 2 or more) mutually different determination methods, each of which uses a plurality of learning data sets, each of which includes one or more explanatory variables and a visibility value as a response variable, to determine weights according to the visibility values ​​included in the plurality of learning data sets, the weights including a first determination method that determines a weight for the learning data set including a first visibility value whose appearance frequency is lower than other visibility values ​​to be greater than the weight for other visibility values; training m estimation models that estimate the visibility value by inputting one or more of the explanatory variables using a plurality of the training data to which weights determined by the m determination methods are assigned; An information processing method including: (Configuration Example 12) On the computer, a step of determining weights by m (m is an integer of 2 or more) mutually different determination methods, each of which uses a plurality of learning data sets, each of which includes one or more explanatory variables and a visibility value as a response variable, to determine weights according to the visibility values ​​included in the plurality of learning data sets, the weights including a first determination method that determines a weight for the learning data set including a first visibility value whose appearance frequency is lower than other visibility values ​​to be greater than the weight for other visibility values; training m estimation models that estimate the visibility value by inputting one or more of the explanatory variables using a plurality of the training data to which weights determined by the m determination methods are assigned; A program to execute.

[0136] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]

[0137] 100, 100-2 Information processing device 10, 10-2 Learning control section 20, 20-2 Estimation control section 101 Selection section 102_1~102_m Weight determination section 103_1~103_m Model learning section 104, 104-2 Probability Learning Section 121 Storage section 201_1~201_m Estimation part 202, 202-2 Probability calculation section 203, 203-2 Estimated value calculation section 221 Storage section

Claims

1. a method for determining weights according to visibility values ​​included in a plurality of learning data sets, each of which includes one or more explanatory variables and a visibility value that is a target variable, using m (m is an integer of 2 or more) different determination methods, the method including a first determination method for determining a weight for the learning data set including a first visibility value whose occurrence frequency is lower than other visibility values ​​so that the weight is greater than that for other visibility values; training m estimation models that estimate the visibility value by inputting one or more of the explanatory variables using a plurality of the learning data to which weights determined by the m determination methods are assigned; Processing section An information processing device comprising:

2. The processing unit using a plurality of the learning data and inputting one or more of the explanatory variables, to learn a probabilistic model that estimates the probability that the visibility value is the first visibility value; The information processing device according to claim 1 .

3. The processing unit selecting one or more first candidates that affect the estimation of the visibility value from a plurality of candidates for the explanatory variable, and generating a plurality of pieces of the learning data including the one or more selected first candidates as one or more of the explanatory variables; The information processing device according to claim 1 .

4. The processing unit selecting one or more of the first candidates having a higher accuracy of estimation by a model that estimates the visibility value using the candidates as explanatory variables than the other candidates; The information processing device according to claim 3 .

5. The processing unit selecting one or more of the first candidates having a cross-correlation with the visibility value that is greater than the other candidates; The information processing device according to claim 3 .

6. The one or more explanatory variables include at least one of relative humidity, temperature, and PM2.

5. The information processing device according to claim 1 .

7. The m determination methods include a second determination method that determines the weights using kernel density estimation. The information processing device according to claim 1 .

8. The processing unit using the m estimation models, respectively, to estimate the m visibility values ​​for one or more explanatory variables for estimation; calculating an estimate of the visibility value using the m visibility values; The information processing device according to claim 1 .

9. The processing unit inputting one or more of the explanatory variables and using a probabilistic model that estimates the probability that the visibility value is the first visibility value, calculating the probability for the explanatory variables for estimation; calculating an estimated value of the visibility value using the calculated probability and the m visibility values; The information processing device according to claim 8 .

10. using m (m is an integer of 2 or more) estimation models each for estimating a visibility value as a response variable by inputting one or more explanatory variables, estimating m visibility values ​​for the one or more explanatory variables; calculating an estimate of the visibility value using the m visibility values; a processing unit; the m estimation models are trained using a plurality of training data each including one or more of the explanatory variables and the visibility value, the plurality of training data being assigned weights determined by m determination methods that are different from one another; the m determination methods are methods for determining weights according to the visibility values ​​included in the plurality of learning data, and include a first determination method for determining a weight for the learning data including a first visibility value whose appearance frequency is smaller than other visibility values ​​so that the weight is larger than that for the other visibility values; Information processing device.

11. An information processing method executed by an information processing device, a step of determining weights by m (m is an integer of 2 or more) mutually different determination methods, each of which uses a plurality of learning data sets, each of which includes one or more explanatory variables and a visibility value as a response variable, to determine weights according to the visibility values ​​included in the plurality of learning data sets, the weights including a first determination method that determines a weight for the learning data set including a first visibility value whose appearance frequency is lower than other visibility values ​​to be greater than the weights for other visibility values; training m estimation models that estimate the visibility value by inputting one or more of the explanatory variables using a plurality of the training data to which weights determined by the m determination methods are assigned; An information processing method including:

12. On the computer, a step of determining weights by m (m is an integer of 2 or more) mutually different determination methods, each of which uses a plurality of learning data sets, each of which includes one or more explanatory variables and a visibility value as a response variable, to determine weights according to the visibility values ​​included in the plurality of learning data sets, the weights including a first determination method that determines a weight for the learning data set including a first visibility value whose appearance frequency is lower than other visibility values ​​to be greater than the weights for other visibility values; training m estimation models that estimate the visibility value by inputting one or more of the explanatory variables using a plurality of the training data to which weights determined by the m determination methods are assigned; A program to execute.

Citation Information

Patent Citations

  • Gas condition prediction device, method, program, and diffusion condition prediction system

    JP4404220B2

  • Method, program, and device for predicting air pollution (ultra-short-term air pollution prediction)

    JP6714697B2

  • Visibility estimation device, visibility estimation method, and recording medium

    JP6901647B1