INFORMATION PROCESSING DEVICE, PROGRAM AND INFORMATION PROCESSING METHOD

DE112022007047T5Pending Publication Date: 2025-07-24MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE112022007047
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-07-24

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An information processing device (100) comprises: an attention mechanism unit (113) that calculates a context variable by weighting and adding a plurality of time series variables using an attention mechanism learning model that is a learning model of an attention mechanism; a decision unit (114) that determines a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a last variable included in the plurality of variables; a storage unit (101) that stores result information correlating the context variable and the one decision; and an evaluation unit (115) that evaluates a training state of at least the attention mechanism learning model based on the result information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to an information processing device, a program and an information processing method. TECHNICAL BACKGROUND

[0002] Attention mechanisms exist as a technique for improving the detection accuracy of a learning model. For example, NPL 1 describes how using an attention mechanism when translating natural language with a neural network can improve translation accuracy. PRIOR ART REFERENCES NON-PATENT REFERENCES

[0003] Non-Patent Literature 1: Minh-Thang Luong et al., “Effective Approaches to Attention-based Neural Machine Translation,” arXiv preprint arXiv: 1508. 04025, Aug. 18, 2015. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION

[0004] However, the internal processing of a deep reinforcement learning model is a black box and therefore invisible. Therefore, the user cannot easily determine whether the model has been effectively trained.

[0005] Accordingly, it is a goal of one or more aspects of the disclosure to provide the ability to easily capture the training state of a learning model using an attention mechanism. MEANS TO SOLVED THE TASK

[0006] An information processing device according to one aspect of the disclosure comprises: an attention mechanism unit configured to calculate a context variable by weighting and summing a plurality of time series variables using an attention mechanism learning model, wherein the attention mechanism learning model is a learning model of an attention mechanism; a decision unit configured to determine a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a last variable included in the plurality of variables; a storage unit configured to store result information correlating the context variable and the decision;and an evaluation unit configured to evaluate a training state of at least the attention mechanism learning model based on the result information.;

[0007] A program according to one aspect of the disclosure causes a computer to operate as: an attention mechanism unit configured to calculate a context variable by weighting and summing a plurality of time series variables using an attention mechanism learning model, wherein the attention mechanism learning model is a learning model of an attention mechanism; a decision unit configured to determine a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a last variable included in the plurality of variables; a storage unit configured to store result information correlating the context variable and a decision;and an evaluation unit configured to evaluate a training state of at least the attention mechanism learning model based on the result information.;

[0008] An information processing method according to one aspect of the disclosure includes: calculating a context variable by weighting and summing a plurality of time series variables using an attention mechanism learning model, wherein the attention mechanism learning model is a learning model of an attention mechanism; determining a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a last variable included in the plurality of variables; storing result information correlating the context variable and the one decision; and evaluating a training state of at least the attention mechanism learning model based on the result information. EFFECTS OF THE INVENTION

[0009] According to one or more aspects of the disclosure, the training state of a learning model can be easily detected using an attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram schematically illustrating a configuration of an information processing device according to a first embodiment. Fig. 2A and Fig. 2B are block diagrams showing examples of a hardware configuration. Fig. 3 is a schematic diagram for explaining a process executed in the information processing device according to the first embodiment. Fig. 4 is a block diagram schematically illustrating a configuration of an information processing device according to a second embodiment. Fig. 5 is a schematic diagram for explaining processing executed in the information processing device according to the second embodiment. Fig. 6 is a block diagram schematically illustrating a configuration of an information processing device according to a third embodiment. Fig. 7 is a schematic diagram for explaining processing executed in the information processing device according to the third embodiment. Fig. 8 is a block diagram schematically illustrating a configuration of an information processing device according to a fourth embodiment. Fig. 9 is a schematic diagram for explaining processing executed in the information processing device according to the fourth embodiment. MODE FOR CARRYING OUT THE INVENTION FIRST EMBODIMENT

[0010] Fig. 1 is a block diagram schematically illustrating a configuration of an information processing device 100 according to a first embodiment.

[0011] The information processing device 100 comprises a storage unit 101, a communication unit 102, an input unit 103, a display unit 104 and a control unit 110.

[0012] The storage unit 101 stores programs and data required for the processing performed in the information processing device 100.

[0013] For example, the storage unit 101 stores at least one attention mechanism learning model, which is a learning model used in an attention mechanism executed by the control unit 110. In the first embodiment, the storage unit 101 also stores an extractive learning model and a decision learning model, as described later.

[0014] The storage unit 101 further stores result information correlating decision results of the decisions made by the control unit 110 using determination results of the attention mechanism and the determination results.

[0015] The communication unit 102 communicates with other devices. The communication unit 102 communicates with other devices, for example, via a network such as the Internet.

[0016] The input unit 103 accepts an input from a user of the information processing device 100.

[0017] The display unit 104 displays information to a user of the information processing device 100. For example, the display unit 104 displays various screen images.

[0018] The control unit 110 controls processing performed in the information processing device 100. For example, the control unit 110 calculates a context state variable by using the attention mechanism to weight and add state variables required for a decision, and determines a specific decision from the context state variable. The control unit 110 then correlates the context state variable and the decision determined from the context state variable and stores this as result information in the storage unit 101.

[0019] In the following, a state variable is also simply referred to as a variable and a context state variable is also simply referred to as a context variable.

[0020] The control unit 110 uses the result information stored in the storage unit 101 to evaluate the training state of at least the learning model used in the attention mechanism. In the first embodiment, the control unit 110 evaluates the training states of an extractive learning model, an attention mechanism learning model, and a decision learning model, as described later.

[0021] The control unit 110 includes a data acquisition unit 111, a variable extraction unit 112, an attention mechanism unit 113, a decision unit 114, and an evaluation unit 115.

[0022] The data acquisition unit 111 acquires input data. The data acquisition unit 111 can acquire input data, for example, via the communication unit 102. If input data is stored in the storage unit 101, the data acquisition unit 111 can acquire input data from the storage unit 101.

[0023] The variable extraction unit 112 extracts state variables, which are variables that can be used for decision making, from the input data acquired by the data acquisition unit 111.

[0024] Here, the variable extraction unit 112 extracts state variables using an extractive learning model, which is a learning model for extracting state variables from input data.

[0025] The attention mechanism unit 113 calculates a context state variable by causing a known attention mechanism to determine a weighted sum of the state variables extracted by the variable extraction unit 112. For example, the attention mechanism unit 113 uses a learning model stored in the storage unit 101 to weight the state variables extracted by the variable extraction unit 112 and add the weighted variables to calculate a context state variable as the determination result.

[0026] Based on the confidence levels of multiple decisions calculated from the context state variable determined by the attention mechanism unit 113 and a final state variable contained in multiple state variables, the decision unit 114 determines a decision contained in the multiple decisions from the one decision contained in the multiple decisions. The decision unit 114 then correlates the one decision and the context state variable and stores them as result information in the storage unit 101.

[0027] Here, the decision unit 114 performs a determination by using a decision learning model, which is a learning model for determining a decision from a context variable.

[0028] The evaluation unit 115 evaluates the training state of at least one attention mechanism learning model, which is a learning model used by the attention mechanism unit 113, based on the result information stored in the storage unit 101.

[0029] In the first embodiment, the evaluation unit 115 evaluates the training states of the extractive learning model, the attention mechanism learning model, and the decision learning model. However, if the state variables are not extracted from the input data, the evaluation unit 115 evaluates the training states of the attention mechanism learning model and the decision learning model.

[0030] For example, the evaluation unit 115 assigns each of the multiple decisions to a cluster to specify multiple clusters and evaluates the clusters based on the distance or similarity between the clusters. Here, the shorter the distance or the higher the similarity, the lower the evaluation.

[0031] Part or all of the control unit 110 described above may be implemented, for example, by a memory 10 and a processor 11, such as a central processing unit (CPU), which executes the programs stored in the memory 10, as shown in Fig. 2A. In other words, the information processing device 100 can be implemented by a well-known computer. Such programs can be provided via a network or can be recorded and provided on a recording medium. That is, such programs can be provided, for example, as a program product.

[0032] A part or the whole of the control unit 110 may also be implemented, for example, by a single circuit, a composite circuit, a program-operated processor, a program-operated parallel processor, a processing circuit 12 such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA), as in Fig. 2B shown, be implemented.

[0033] As described above, the control unit 110 may be implemented by a processing circuit.

[0034] The storage unit 101 may be implemented by a memory such as a hard disk drive (HDD) or a solid state drive (SSD).

[0035] The communication unit 102 may be implemented by a communication interface, such as a network interface card (NIC).

[0036] The input unit 103 may be implemented by an input interface such as a keyboard or a mouse.

[0037] The display unit 104 may be implemented by a display.

[0038] Fig. 3 is a schematic diagram for explaining processing executed in the information processing device 100 according to the first embodiment.

[0039] First, the data acquisition unit 111 acquires the input data X t-n , X t-n+1 , X t-1 , Xt (step S10). For the input data X t-n , Xt-n+1 , X t-1 , Xt are sensor values, which are observation values, and the time series data tn, t-n+1, t-1, t (where t and n are positive integers). For example, image data can be used as input data.

[0040] The data acquisition unit 111 outputs the acquired input data X t-n , X t-n+1 , X t-1 , Xt to the variable extraction unit 112.

[0041] The variable extraction unit 112 extracts from the input data X t-n , X t-n+1 , X t-1 , Xt state variables S t-n , S t-n+1 , S t-1 and St, which are advantageous variables for the decision unit 114 to make decisions (step S11).

[0042] The variable extraction unit 112 uses the extractive learning model, which is a neural network model stored in the storage unit 101, to extract the state variables S t-n , St-n+1 , S t-1 and St from the input data X t-n , X t-n+1 , X t-1 , to extract Xt.

[0043] The variable extraction unit 112 outputs the extracted state variables S t-n , S t-n+1 , S t-1 and St to the attention mechanism unit 113.

[0044] Here, the variable extraction unit 112 uses the extractive learning model; however, the first embodiment is not limited to such an example as long as the state variables S t-n , S t-n+1 , S t-1 and St can be extracted using a function.

[0045] The attention mechanism unit 113 uses a learning model to calculate weighted values of the state variables S t-n , S t-n+1 , S t-1 and St, and calculates the weighted sum to calculate a context state variable (step S12).

[0046] The attention mechanism unit 113 passes the calculated context state variable to the decision unit 114.

[0047] The decision unit 114 makes a decision based on the context state variables and the last state variable St (step S13).

[0048] The decision unit 114 uses the decision learning model, which is a neural network model stored in the storage unit 101, to determine a decision from the context state variable and the last state variable.

[0049] The decision unit 114 then correlates the decision with the context state variable and stores it as result information in the storage unit 101, thereby accumulating result information (step S14).

[0050] The evaluation unit 115 uses the result information stored in the storage unit 101 to evaluate the training state of at least the learning model used by the attention mechanism unit 113.

[0051] To facilitate evaluation, the evaluation unit 115 converts, for example, the N-dimensional data obtained by assigning the result information of the respective decisions to clusters into lower-dimensional data (step S15). Specifically, the evaluation unit 115 converts the N-dimensional data into two-dimensional data by using t-SNE (t-distributed stochastic neighbor embedding) to visualize the clusters of the respective decision.

[0052] The evaluation unit 115 then calculates, for example, the distance or similarity between the clusters as evaluation values and thereby evaluates the training state (step S16).

[0053] For example, the evaluation unit 115 performs an evaluation by comparing the evaluation values between clusters with a threshold. If the distance between clusters is less than a predetermined threshold or if the similarity between clusters is greater than a predetermined threshold, the evaluation unit 115 determines that the training is insufficient.

[0054] The determination result of the evaluation unit 115 can be displayed, for example, on the display unit 104.

[0055] As described above, the training state of a learning model using the attention mechanism can be easily acquired according to the first embodiment. SECOND EMBODIMENT

[0056] Fig. 4 is a block diagram schematically illustrating a configuration of an information processing device 200 according to the second embodiment.

[0057] The information processing device 200 includes a storage unit 101, a communication unit 102, an input unit 103, a display unit 104, and a control unit 210.

[0058] The storage unit 101, the communication unit 102, the input unit 103, and the display unit 104 of the information processing device 200 according to the second embodiment are respectively the same as the storage unit 101, the communication unit 102, the input unit 103, and the display unit 104 of the information processing device 100 according to the first embodiment.

[0059] The control unit 210 controls the processing executed in the information processing device 200.

[0060] The control unit 210 according to the second embodiment performs the same processing as that performed by the control unit 110 according to the first embodiment and also performs the following processing.

[0061] The control unit 210 trains a learning model by using additional training data depending on the evaluation result of a training state.

[0062] The control unit 210 includes a data acquisition unit 111, a variable extraction unit 112, an attention mechanism unit 113, a decision unit 114, an evaluation unit 215, and an additional training unit 216.

[0063] The data acquisition unit 111, the variable extraction unit 112, the attention mechanism unit 113, and the decision unit 114 of the control unit 210 according to the second embodiment are respectively the same as the data acquisition unit 111, the variable extraction unit 112, the attention mechanism unit 113, and the decision unit 114 of the control unit 110 according to the first embodiment.

[0064] The evaluation unit 215 uses the result information stored in the storage unit 101 to evaluate the training state of at least the learning model used by the attention mechanism unit 113.

[0065] The evaluation unit 215 then forwards the evaluation result to the additional training unit 216. The evaluation unit 215 compares, for example, the evaluation value and a threshold value for each combination of two clusters to generate evaluation information indicating whether the training is sufficient, and forwards the evaluation information to the additional training unit 216.

[0066] The additional training unit 216 refers to the evaluation information of the evaluation unit 215 and passes additional training data to the variable extraction unit 112 to perform additional training.

[0067] If the evaluation by evaluation unit 215 is below a predetermined threshold, additional training unit 216 uses the additional training data to train at least one attention mechanism learning model. In the second embodiment, additional training unit 216 trains an extractive learning model, a decision learning model, and an attention mechanism learning model.

[0068] For example, the additional training unit 216 performs training using additional training data, i.e., training data in which decisions whose evaluations are below a predetermined threshold are determined to be correct from among multiple decisions. In other words, the additional training unit 216 can provide the variable extraction unit 112 with training data classified into the two clusters determined to be insufficiently learned as additional training data. The additional training data can be obtained, for example, from another device via the communication unit 102 or from the storage unit 101. A user can, for example, give an instruction via the input unit 103 as to where the additional training data should be obtained from.

[0069] Fig. 5 is a schematic diagram for explaining processing executed in the information processing device 200 according to the second embodiment.

[0070] The processing of steps S10 to S15 in Fig. 5 is the same as the processing of steps S10 to S15 in Fig. 3.

[0071] In the second embodiment, the evaluation unit 215 calculates, for example, the distance or similarity between clusters as an evaluation value to evaluate the training state, and thereby generates evaluation information indicating the evaluation result (step S26). The evaluation information is information indicating whether the respective combinations of two clusters have been sufficiently learned. The generated evaluation information is passed to the additional training unit 216.

[0072] The additional training unit 216 refers to the evaluation information and generates additional training data, ie, training data classified into clusters determined to be insufficiently learned (step S27), and passes this additional training data to the variable extraction unit 112 to perform additional training.

[0073] As described above, a learning model using the attention mechanism according to the second embodiment can additionally learn insufficiently learned clusters.

[0074] Here, the evaluation unit 215 can use a threshold to determine whether the training is sufficient; however, the risk of decisions can be managed, for example, by using multiple thresholds. In particular, for clusters of decisions where there is no margin for error, such as "stopping" or "accelerating" a vehicle, the distance between clusters must be large or the similarity between clusters must be low; therefore, the threshold can be adjusted to manage the risk of the decisions. THIRD EMBODIMENT

[0075] Fig. 6 is a block diagram schematically illustrating a configuration of an information processing device 300 according to the third embodiment.

[0076] The information processing device 300 includes a storage unit 101, a communication unit 102, an input unit 103, a display unit 104, and a control unit 310.

[0077] The storage unit 101, the communication unit 102, the input unit 103, and the display unit 104 of the information processing device 300 according to the third embodiment are respectively the same as the storage unit 101, the communication unit 102, the input unit 103, and the display unit 104 of the information processing device 100 according to the first embodiment.

[0078] The control unit 310 controls the processing performed in the information processing device 300.

[0079] The control unit 310 according to the third embodiment performs the same processing as that performed by the control unit 110 according to the first embodiment and also performs the following processing.

[0080] The control unit 310 selects training data in accordance with an evaluation result of a training state and uses the selected training data to train a learning model.

[0081] The control unit 310 includes a data acquisition unit 111, a variable extraction unit 112, an attention mechanism unit 113, a decision unit 114, an evaluation unit 315, a training data selection unit 317, and a training unit 318.

[0082] The data acquisition unit 111, the variable extraction unit 112, the attention mechanism unit 113, and the decision unit 114 of the control unit 310 according to the third embodiment are respectively the same as the data acquisition unit 111, the variable extraction unit 112, the attention mechanism unit 113, and the decision unit 114 of the control unit 110 according to the first embodiment.

[0083] As in the first embodiment, the evaluation unit 315 uses the result information stored in the storage unit 101 to evaluate the training state of at least the learning model used by the attention mechanism unit 113.

[0084] In the third embodiment, the evaluation unit 315 provides the training data selection unit 317 with evaluation value information indicating the evaluation value for each combination of two clusters.

[0085] The training data selection unit 317 refers to the evaluation value information of the evaluation unit 315 and selects the training data for training at least one attention mechanism learning model.

[0086] The training data selection unit 317 selects the training data such that the lower the evaluation corresponding to a decision, the larger the number of training data items for which the decision is correct. In other words, the lower the evaluation by an evaluation value specified in the evaluation value information, that is, the smaller the distance or the higher the similarity between clusters, the larger the number of training data items classified into these clusters. The training data may be stored in the storage unit 101 or in another device. If the training data is stored in another device, the training data selection unit 317 may access the device via the communication unit 102 and select the training data.

[0087] The training unit 318 uses the training data selected by the training data selection unit 317 to train at least the attention mechanism learning model.

[0088] For example, the training unit 318 performs training by passing the training data selected by the training data selection unit 317 to the variable extraction unit 112.

[0089] Fig. 7 is a schematic diagram for explaining processing executed in the information processing device 300 according to the third embodiment.

[0090] Fig. 7 shows processing executed for performing training using training data in the information processing device 300.

[0091] As a prerequisite, the training data selection unit 317 provides initial training data, that is, training data selected without reference to evaluation value information, to the training unit 318. The training unit 318 provides the initial training data to the variable extraction unit 112 to perform initial training, and then training data is selected according to the evaluation result of the initial training.

[0092] The processing of steps S11 to S15 in Fig. 7 is the same as the processing of steps S11 to S15 in Fig. 3.

[0093] In the third embodiment, the evaluation unit 315 calculates, for example, the distance or similarity between clusters as evaluation values to evaluate the training state, and thereby generates evaluation value information indicating the evaluation results for the respective combinations of two clusters (step S36). The generated evaluation value information is passed to the training data selection unit 317.

[0094] The training data selection unit 317 refers to the evaluation value information and selects the training data such that the lower the evaluation by an evaluation value indicated by the evaluation value information, the larger the number of training data items classified into the clusters (step S37). The training data selection unit 317 then forwards the selected training data to the training unit 318.

[0095] The training unit 318 performs training by passing the training data selected by the training data selection unit 317 to the variable extraction unit 112 (step S38).

[0096] As described above, according to the third embodiment, training of a learning model using an attention mechanism can be efficiently performed by selecting the training data to be intensively learned.

[0097] Although the training data selection unit 317 selects the training data such that the lower the evaluation by an evaluation value indicated by the evaluation value information, the larger the number of training data items classified into the clusters, the third embodiment is not limited to such an example. For example, in the training data selection unit 317, clusters of decisions that have no margin for error, such as "stopping" and "accelerating" of a vehicle, may be set in advance as clusters that should be intensively learned, so that the training data selection unit 317 can select training data containing many such clusters.In particular, the training data selection unit 317 may increase the number of training data items to be selected by adding or multiplying a weight value that results in a low evaluation value for clusters to be intensively learned. Such a setting may be made, for example, by a user via the input unit 103. FOURTH EMBODIMENT

[0098] Fig. 8 is a block diagram schematically illustrating a configuration of an information processing device 400 according to the fourth embodiment.

[0099] The information processing device 400 includes a storage unit 101, a communication unit 102, an input unit 103, a display unit 104, and a control unit 410.

[0100] The storage unit 101, the communication unit 102, the input unit 103, and the display unit 104 of the information processing device 400 according to the fourth embodiment are respectively the same as the storage unit 101, the communication unit 102, the input unit 103, and the display unit 104 of the information processing device 100 according to the first embodiment.

[0101] The control unit 410 controls the processing executed in the information processing device 400.

[0102] The control unit 410 according to the fourth embodiment performs the same processing as that performed by the control unit 110 according to the first embodiment and also performs the following processing.

[0103] The control unit 410 decides whether to continue the training depending on the evaluation result of the training state; if it is decided to continue the training, the training is continued; if it is decided not to continue the training, the training is terminated.

[0104] The control unit 410 includes a data acquisition unit 111, a variable extraction unit 112, an attention mechanism unit 113, a decision unit 114, an evaluation unit 115, a training unit 418, and a training continuation decision unit 419.

[0105] The data acquisition unit 111, the variable extraction unit 112, the attention mechanism unit 113, and the decision unit 114 of the control unit 410 according to the fourth embodiment are respectively the same as the data acquisition unit 111, the variable extraction unit 112, the attention mechanism unit 113, and the decision unit 114 of the control unit 110 according to the first embodiment.

[0106] The evaluation unit 215 according to the fourth embodiment is the same as the evaluation unit 215 according to the second embodiment. However, in the fourth embodiment, the evaluation unit 215 passes evaluation information to the training continuation decision unit 419.

[0107] The training continuation decision unit 419 refers to the evaluation information of the evaluation unit 215 and decides whether to continue the training of at least one attention mechanism learning model.

[0108] For example, when all or some evaluations by the evaluation values specified in the evaluation information are below a predetermined threshold, that is, when the distance is shorter than the predetermined threshold or when the similarity is higher than a predetermined threshold, the training continuation decision unit 419 decides that the training should be continued.

[0109] The term "some evaluations" may refer to a predetermined number of evaluations or evaluations of predetermined clusters. For example, if all evaluations of important clusters that have no room for error are above a threshold, the training continuation decision unit 419 may decide not to continue training.

[0110] If the training continuation decision unit 419 decides that training should continue, the training unit 418 performs training by passing the training data to the variable extraction unit 112. Conversely, if the training continuation decision unit 419 decides that training should not continue, the training unit 418 terminates training without passing the training data to the variable extraction unit 112.

[0111] The training data may be stored in storage unit 101 or in another device. If the training data is stored in another device, training unit 418 may access the device via communication unit 102 and obtain the training data.

[0112] Fig. 9 is a schematic diagram for explaining processing executed in the information processing device 400 according to the fourth embodiment.

[0113] Fig. 9 shows processing executed for performing training using training data in the information processing device 400.

[0114] As a prerequisite, the training unit 418 provides initial training data to the variable extraction unit 112 to perform initial training, and then decides whether to continue training depending on the evaluation result of the initial training.

[0115] The processing of steps S11 to S15 in Fig. 9 is the same as the processing of steps S11 to S15 in Fig. 3.

[0116] In the fourth embodiment, the evaluation unit 215 calculates, for example, the distance or similarity between clusters as evaluation values to evaluate the training state, and thereby generates evaluation information indicating the evaluation result (step S46). The evaluation information is information indicating whether the respective combinations of two clusters have been sufficiently learned. The generated evaluation information is passed to the training continuation decision unit 419.

[0117] The training continuation decision unit 419 refers to the evaluation information of the evaluation unit 215 and decides whether to continue the training (step S47).

[0118] When the training continuation decision unit 419 decides that the training should be continued, the training unit 418 performs the training by passing the training data to the variable extraction unit 112 (step S48).

[0119] As described above, according to the fourth embodiment, training can be terminated when a learning model using an attention mechanism has been sufficiently trained. Thus, training can be performed efficiently.

[0120] As in the second embodiment, the evaluation unit 215 may use a threshold to determine whether the training is sufficient; however, the risk of decisions can be managed, for example, by using multiple thresholds. In particular, for clusters of decisions where there is no margin for error, such as "stopping" or "accelerating" a vehicle, the distance between clusters must be large or the similarity between clusters must be low; therefore, the threshold can be adjusted to manage the risk of decisions. LIST OF REFERENCE SYMBOLS

[0121] 100, 200, 300, 400 information processing device; 101 storage unit; 102 communication unit; 103 input unit; 104 display unit; 110, 210, 310, 410 control unit; 111 data acquisition unit; 112 variable extraction unit; 113 attention mechanism unit; 114 decision unit; 115, 215, 315 evaluation unit; 216 additional training unit; 317 training data selection unit; 318, 418 training unit; 419 training continuation decision unit. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature

[0000] Minh-Thang Luong et al., „Effective Approaches to Attention-based Neural Machine Translation“, arXiv preprint arXiv: 1508. 04025, 18 Aug. 2015

[0003]

Claims

[1] Information processing device, comprising: an attention mechanism unit configured to calculate a context variable by weighting and summing a plurality of time series variables using an attention mechanism learning model, wherein the attention mechanism learning model is a learning model of an attention mechanism; a decision unit configured to determine a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a last variable included in the plurality of variables; a storage unit configured to store result information correlating the context variable and a decision; and an evaluation unit configured to evaluate a training state of at least the attention mechanism learning model based on the result information. [2] Information processing device according to claim 1, wherein, the decision unit is configured to determine a decision using a decision learning model, wherein the decision learning model is a learning model that determines the one decision from the context variable, and the evaluation unit is set up to evaluate the decision learning model and the attention mechanism learning model. [3] The information processing device according to claim 2, further comprising a variable extraction unit configured to extract the variables from the input data. [4] Information processing device according to claim 3, wherein, the variable extraction unit is configured to extract the variables using an extractive learning model, wherein the extractive learning model is a learning model that extracts the variables from the input data, and the evaluation unit is set up to evaluate the extractive learning model, the decision learning model and the attention mechanism learning model. [5] The information processing device according to claim 1, further comprising a variable extraction unit configured to extract the variables from the input data. [6] Information processing device according to claim 5, wherein, the variable extraction unit is configured to extract the variables using an extractive learning model, wherein the extractive learning model is a learning model that extracts the variables from the input data, and the evaluation unit is set up to evaluate the extractive learning model and the attention mechanism learning model. [7] The information processing device according to any one of claims 1 to 6, wherein the evaluation unit is arranged to assign each of the decisions to a cluster to specify a plurality of clusters, and to evaluate the clusters based on distance or similarity between the clusters. [8] The information processing device according to any one of claims 1 to 7, further comprising an additional training unit configured to train at least the attention mechanism learning model using additional training data when the evaluation is lower than a predetermined threshold. [9] The information processing device according to claim 8, wherein the additional training unit is arranged to use, as the additional training data, training data in which decisions whose evaluations are below a predetermined threshold in the decisions are set as correct. [10] Information processing device according to one of claims 1 to 7, further comprising: a training data selection unit configured to select, in accordance with the evaluation, training data to be used to train at least the attention mechanism learning model; and a training unit configured to train at least the attention mechanism learning model using the selected training data. [11] The information processing device according to claim 10, wherein the training data selection unit is arranged to perform the selection such that the number of training data items for which the one decision is correct is larger, the lower the evaluation corresponding to the one decision is. [12] Information processing device according to one of claims 1 to 7, further comprising: a training continuation decision unit configured to decide, depending on the evaluation, whether training of at least the attention mechanism learning model should be continued; and a training unit configured to continue training using training data used to train at least the attention mechanism learning model if it is decided that training will continue, and to stop training if it is decided that training will not continue. [13] The information processing device according to claim 12, wherein the training continuation decision unit is arranged to decide to continue the training when the evaluation of all the decisions or some of the decisions is lower than a predetermined threshold. [14] Program that causes a computer to act as: an attention mechanism unit configured to calculate a context variable by weighting and summing a plurality of time series variables using an attention mechanism learning model, wherein the attention mechanism learning model is a learning model of an attention mechanism; a decision unit configured to determine a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a last variable included in the plurality of variables; a storage unit configured to store result information correlating the context variable and a decision; and an evaluation unit configured to evaluate a training state of at least the attention mechanism learning model based on the result information. [15] Information processing methods comprising: Calculating a context variable by weighting and summing a plurality of time series variables using an attention mechanism learning model, wherein the attention mechanism learning model is a learning model of an attention mechanism; Determining a decision included in a plurality of decisions based on confidence levels of the plurality of decisions calculated from the context variable and a final variable included in the plurality of variables; Storing outcome information that correlates the context variable and a decision; and Evaluating a training state of at least the attention mechanism learning model based on the result information.

Citation Information

Patent Citations

  • SPATIAL AND TEMPORAL ATTENTION-BASED DEEP REINFORCEMENT LEARNING OF HIERARCHICAL LANE CHANGE STRATEGIES FOR CONTROLLING AN AUTONOMOUS VEHICLE

    DE102019115707A1

  • US000010929674B2