Deep learning device, deep learning method, and program
The deep learning device and method improve inference accuracy by extracting partial vector sequences and determining learned information within deep learning models, addressing the challenge of unlearned data during inference.
Patent Information
- Application Number
- PCT/JP2023/044396
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Deep learning models face challenges in achieving high inference accuracy due to the limitations of learning data comprehensiveness, where unlearned vector data can be input during inference, leading to meaningless results.
A deep learning device and method that includes a first extraction processing unit to extract a partial vector sequence from inference data, a second extraction processing unit to convert this partial sequence into partial sequence information using a recurrent neural network, and a determination unit to perform an approximate search and determine if the partial sequence information is learned, allowing for accurate inference processing.
The proposed solution enhances the inference accuracy of deep learning by effectively handling unlearned data through partial sequence extraction and learning determination, ensuring reliable inference results even with incomplete learning data.
Smart Images

Figure JP2023044396_19062025_PF_FP_ABST
Abstract
Description
Deep learning device, deep learning method, and program
[0001] The present disclosure relates to a deep learning device, a deep learning method, and a program.
[0002] In deep learning using neural networks and other technologies, various types of data in the real world are converted into vector format that can be handled by computer systems.
[0003] In order to improve the accuracy of inference in deep learning, it is necessary to train the system in advance with a large number of patterns of combinations of input vectors and correct answer data, and it is desirable to obtain a large amount of training vector data.
[0004] Non-Patent Document 1 discloses data augmentation to increase the amount of training data, and Non-Patent Document 2 discloses a representative data augmentation method applied to natural language processing data.
[0005] S. Yang, et al., "Image Data Augmentation for Deep Learning: A Survey, "arXiv:2204.08610v1 [cs.CV] 19 Apr 2022.B. Li, et al., "Data Augmentation Approaches in NaturalLanguage Processing: A Survey, "arXiv:2110.01852v3 [cs.CL] 27 Jun 2022.
[0006] However, even if a method of expanding and increasing the amount of learning data is adopted, as in the techniques disclosed in Non-Patent Documents 1 and 2, it is not necessarily possible to completely cover all patterns of input vectors.
[0007] When using training data that does not fully cover the vector appearance patterns, it is possible that untrained vector data will be input during inference. Generally, in deep learning, it is not assumed that the vectors of the inference data to be inferred will exactly match the trained vectors, and generalization is ensured to enable correct inference even for inference data vectors that are slightly different from the trained vectors. However, if a vector that is significantly different from all the vector data input to the neural network during training is provided as inference data, the inference results derived based on the output values of the neural network may be meaningless.
[0008] In other words, even if a large amount of training data is prepared or generalization is made possible, there is a possibility that untrained data that is significantly different from any of the trained vector data may be input during inference, which poses a problem in that it is not possible to improve the inference accuracy of deep learning.
[0009] The present disclosure has been made in consideration of the above circumstances, and its purpose is to provide a deep learning device, a deep learning method, and a program that can improve the inference accuracy of deep learning.
[0010] A deep learning device according to one aspect of the present disclosure includes a first extraction processing unit that receives inference data having a vector sequence and extracts a partial vector sequence that includes at least one vector from each of the vectors included in the vector sequence; a second extraction processing unit that receives the partial vector sequence and circulates a recurrent neural network to extract partial sequence information that reflects overall information about the partial vector sequence; a recording unit that records learned vector sequences; a determination unit that performs a vector approximation search process on the partial sequence information by referencing the learned vector sequence and determines whether the partial sequence information has been learned; and an inference processing unit that infers tag values for each vector sequence included in the inference data using the partial sequence information determined by the determination unit to have been learned.
[0011] A deep learning method according to one aspect of the present disclosure acquires inference data having a vector sequence, extracts a partial vector sequence consisting of at least one vector from each of the vectors included in the vector sequence, inputs the partial vector sequence into a recurrent neural network and circulates it to extract partial sequence information that reflects the overall information of the partial vector sequence, performs an approximate vector search process based on the partial sequence information by referencing a trained vector sequence recorded in a recording unit, determines whether the partial sequence information has been trained, and uses the partial sequence information determined to have been trained to infer tag values for each vector sequence included in the inference data.
[0012] One aspect of the present disclosure is a program for causing a computer to function as the above-described deep learning device.
[0013] According to the present disclosure, it is possible to improve the inference accuracy of deep learning.
[0014] FIG. 1 is a block diagram showing the configuration of a deep learning device according to the first embodiment. FIG. 2 is a block diagram showing the detailed configuration of a learning unit shown in FIG. 1. FIG. 3 is a block diagram showing the detailed configuration of an extraction unit shown in FIG. 1. FIG. 4 is an explanatory diagram showing the flow of data in the extraction unit shown in FIG. 1. FIG. 5A is a block diagram showing the actual configuration of an RNN. FIG. 5B is a block diagram showing the configuration of an RNN shown in the embodiment. FIG. 6 is a block diagram showing a specific configuration of an inference processing unit according to the first embodiment. FIG. 7 is an explanatory diagram showing a vector sequence of individual learning data according to the first embodiment. FIG. 8 is a flowchart showing the processing procedure of the learning unit according to the first embodiment. FIG. 9 is an explanatory diagram showing processing in the learning unit according to the first embodiment. FIG. 10 is an explanatory diagram showing the flow of data in the RNN unit of the learning unit according to the first embodiment. FIG. 11 is an explanatory diagram showing a vector sequence of individual inference data according to the first embodiment. FIG. 12 is an explanatory diagram showing a vector sequence included in individual inference data according to the first embodiment. FIG. 13A is a first subdivision of an explanatory diagram showing the flow of data in processing by the extraction unit according to the first embodiment. FIG. 13B is a second subdivision of an explanatory diagram showing the flow of data in processing by the extraction unit according to the first embodiment. FIG. 14 is a flowchart showing the processing procedure of the inference unit according to the first embodiment. FIG. 15 is a flowchart showing the specific processing procedure of step S33 shown in FIG. 14. FIG. 16 is an explanatory diagram showing the flow of data in the inference processing unit according to the first embodiment. FIG. 17 is a block diagram showing the configuration of a deep learning device according to the second embodiment. FIG. 18 is a block diagram showing the configuration of an inference processing unit according to the second embodiment. FIG. 19 is an explanatory diagram showing a vector sequence of individual training data according to the second embodiment. FIG. 20 is an explanatory diagram showing a vector sequence of individual inference data according to the second embodiment. FIG. 21 is an explanatory diagram showing the processing of recording training data of a first vector sequence and a second vector sequence included in individual training data according to the second embodiment. FIG. 22 is an explanatory diagram showing the processing of extracting learned sequences from a first vector sequence and a second vector sequence included in individual inference data according to the second embodiment.FIG. 23 is an explanatory diagram showing a process for generating learned data for all sequences from a first vector sequence and a second vector sequence included in individual inference data according to the second embodiment. FIG. 24 is a flowchart showing the processing procedure of a learning unit according to the second embodiment. FIG. 25 is an explanatory diagram showing processing in a learning unit according to the second embodiment. FIG. 26 is an explanatory diagram showing the flow of data in an RNN unit when generating learned data for all sequences in a learning unit according to the second embodiment. FIG. 27 is an explanatory diagram showing the flow of data in an RNN unit when generating first and second sequence learned data in a learning unit according to the second embodiment. FIG. 28A is a first sub-diagram of a flowchart showing the processing procedure of an inference unit according to the second embodiment. FIG. 28B is a second sub-diagram of a flowchart showing the processing procedure of an inference unit according to the second embodiment. FIG. 29 is an explanatory diagram showing the flow of data in an inference unit according to the second embodiment. FIG. 30 is a block diagram showing the hardware configuration of this embodiment.
[0015] [Description of First Embodiment] Hereinafter, an embodiment will be described with reference to the drawings. Fig. 1 is a block diagram showing the configuration of a deep learning device 100 according to the first embodiment. As shown in Fig. 1, the deep learning device 100 includes a learning unit 1, a recording unit 2, and an inference unit 3.
[0016] Learning data is input to the learning unit 1. The learning data is a collection of multiple individual learning data. The individual learning data is a combination of a vector sequence and tag values assigned to this vector sequence. Each vector sequence is assigned some classification attribute according to the task to be processed by the deep learning device 100, and the tag value is a quantified value of this attribute value. The learning unit 1 learns combinations of vector sequences and tag values based on the input individual learning data. The learning unit 1 outputs the learned combinations of vector sequences and tag values to the recording unit 2 as learned vector sequences.
[0017] 2 is a block diagram showing a specific configuration of the learning unit 1. As shown in FIG. 2, the learning unit 1 includes an RNN unit 11, a neural network 12, and a loss calculation unit 13.
[0018] The RNN unit 11 executes recurrent deep learning (RNN; Recurrent Neural Network) to learn the properties of sequential data consisting of multiple sequences. Multiple LSTMs 14 (Long Short-Term Memories) can be used as RNN cells, which are elements of the RNN. For each vector in a vector sequence included in the individual training data, the LSTM 14 executes a loop process in which it inputs the hidden state at the time of inputting the previous vector and the current vector, and outputs a new hidden state. The loop length of each LSTM 14 constituting the RNN unit 11 is variable. The RNN unit 11 outputs the output data of the hidden state of the final stage to the neural network 12 (see FIG. 9, described later).
[0019] The neural network 12 acquires the hidden state vector of the LSTM 14 at the final stage of the RNN unit 11, performs internal calculations based on this vector, and outputs the results.
[0020] The loss calculation unit 13 receives the output data of the neural network 12 and the tag values included in the individual training data as inputs and calculates the error between the two values. The loss calculation unit 13 uses a method such as backpropagation to optimize the internal parameters of the neural network 12 and the LSTMs 14 of the RNN unit 11 so as to minimize the error, thereby learning the correspondence between the data series of the individual training data and the tag values.
[0021] As shown in FIG. 1, the learning unit 1 outputs the optimized parameters as learned parameters to an inference processing unit 32, which will be described later.
[0022] The recording unit 2 records the trained vector sequence output from the learning unit 1. Hereinafter, the trained vector data recorded in the recording unit 2 will be referred to as "trained data."
[0023] The inference unit 3 includes an extraction unit 31 and an inference processing unit 32. Fig. 3 is a block diagram showing the configuration of the extraction unit 31. As shown in Figs. 1 and 3, the extraction unit 31 includes a first extraction processing unit 311, a second extraction processing unit 312, and a determination unit 313.
[0024] Inference data is input to the first extraction processing unit 311. The inference data is a collection of multiple individual inference data. The individual inference data has a vector sequence containing multiple vectors. The first extraction processing unit 311 extracts a partial vector sequence consisting of at least one vector from each vector included in the vector sequence of the individual inference data. In other words, the first extraction processing unit 311 generates a set of partial vector sequences for each individual inference data.
[0025] Specifically, the first extraction processing unit 311 receives a vector sequence of individual inference data and extracts a partial vector sequence that exists as a sequence of vectors in the vector sequence. Figure 4 is an explanatory diagram showing the flow of processing when individual inference data consisting of vectors d1 to d12 is input to the extraction unit 31.
[0026] For example, when the input vector sequence consists of a sequence of n vectors, the first extraction processing unit 311 extracts, from this sequence of vectors, vector sequences whose length is greater than or equal to 1 and less than or equal to n as partial vector sequences. FIG. 4 shows an example in which partial vector sequences d1 to d4 and d5 to d7 are extracted from 12 vectors d1 to d12. There are as many partial vector sequences that satisfy the above conditions as the total number of permutations in which 1 to n elements are extracted and arranged from n elements. In other words, theoretically, only X1, as shown in equation (1) below, exists. However, it is not necessary to extract all of these partial vector sequences, and restrictions may be imposed as necessary.
[0027]
[0028] In this embodiment, partial vector sequences are extracted under the constraints that not all partial vector sequences (X1) shown in the above formula (1) are extracted, identical vectors are not duplicated, and partial vector sequences containing reversed sequences are not handled. For example, vector d1 is extracted into only one partial vector sequence. Also, partial vector sequences containing reversed sequences such as d3, d2, and d1 are not extracted. The first extraction processing unit 311 outputs the set of partial vector sequences extracted by the above processing to the second extraction processing unit 312 shown in FIGS. 1 and 3.
[0029] The second extraction processing unit 312 acquires the set of partial vector sequences output from the first extraction processing unit 311. The second extraction processing unit 312 circulates the recurrent neural network to convert each partial vector sequence into a single vector (hereinafter referred to as "partial sequence information") that reflects the overall information of each partial vector sequence.
[0030] That is, the second extraction processing unit 312 receives a partial vector sequence and circulates it through a recurrent neural network to extract partial sequence information that reflects the overall information of the partial vector sequence. The second extraction processing unit 312 outputs the extracted partial sequence information to the determination unit 313. The second extraction processing unit 312 also outputs the partial vector sequence corresponding to this partial sequence information to the determination unit 313. This is so that when processing by the determination unit 313 and the inference processing unit 32, described below, is completed, a set of partial vector sequences and their corresponding tag values can be clearly displayed to the user. In this embodiment, an example will be described in which a recurrent neural network (RNN) is used as means for extracting partial sequence information.
[0031] FIG. 5A is a block diagram showing a specific configuration of an RNN. As shown in FIG. 5A, the RNN is composed of a single RNN cell (e.g., an LSTM) with a loop, receives a vector dt as input, and outputs a hidden state ht. In this embodiment, for ease of understanding, the RNN shown in FIG. 5A is shown with all input vectors as shown in FIG. 5B. That is, vectors d1 to dm are input to multiple RNN cells, and the hidden state ht generated by an RNN cell is input to an adjacent RNN cell. The hidden state hm generated by the final RNN cell becomes the output signal of the RNN. That is, although multiple RNN cells are shown in FIG. 5B, in reality, a recurrent neural network is composed of a single RNN cell as shown in FIG. 5A. Therefore, in FIG. 4, the RNN is represented as multiple RNN cells.
[0032] The values of the internal parameters of the RNN can be obtained by the same processing as that of the learning unit 1. Therefore, the RNN entity of the second extraction processing unit 312 may be the same as the RNN unit 11 of the learning unit 1.
[0033] 4 , the second extraction processing unit 312 converts the vector sequences "d1, d2, d3, d4" and "d5, d6, d7" of the partial vector sequence output from the first extraction processing unit 311 into hidden state sequences "h1, h2, h3, h4" and "h5, h6, h7," thereby obtaining the final-stage hidden states "h4" and "h7." The second extraction processing unit 312 outputs a single vector consisting of the final-stage hidden states "h4" and "h7" to the determination unit 313 as partial sequence information. The second extraction processing unit 312 also outputs the partial vector sequence corresponding to this partial sequence information to the determination unit 313.
[0034] That is, the second extraction processing unit 312 inputs the partial vector sequence "d1, d2, ..., dm" shown in FIG. 5B into the recurrent neural network to obtain hidden states h1 to hm. The set of final-stage hidden states hm obtained from the multiple partial vector sequences is output to the determination unit 313 as partial sequence information. Furthermore, the partial vector sequence corresponding to this partial sequence information is also output to the determination unit 313.
[0035] That is, the second extraction processor 312 includes multiple LSTMs (RNNcells) and constructs a recurrent neural network using each LSTM. The second extraction processor 312 also converts each vector included in the partial vector sequence into a hidden state vector using each LSTM, extracts the hidden state in the final-stage LSTM as partial sequence information, and also outputs the partial vector sequence corresponding to this partial sequence information.
[0036] The determination unit 313 acquires a set of pairs of partial sequence information and corresponding partial vector sequences for each individual inference data from the second extraction processing unit 312. The determination unit 313 references the learned data recorded in the recording unit 2 and performs a vector approximation search process to determine, for each partial vector sequence of the individual inference data, whether its neighboring vector is included in the learned data. If a neighboring vector is included, the determination unit 313 determines that the partial vector sequence is learned partial sequence information. Then, it outputs a set of learned partial sequence information that is the result of performing the above determination for all partial sequence information. Furthermore, it also outputs the partial vector sequence corresponding to the learned partial sequence information.
[0037] That is, the determination unit 313 performs an approximate search process for vectors based on the partial sequence information to determine whether the partial sequence information has been learned. The determination unit 313 executes the approximate search process based on the learned sequence vectors recorded in the recording unit 2. The determination unit 313 extracts learned partial sequence information for each individual inference data and outputs it to the inference processing unit 32 shown in Figure 1. The determination unit 313 also outputs the partial vector sequence corresponding to the learned partial sequence information to the inference processing unit 32.
[0038] The vector approximate search process executed by the determination unit 313 will be described below. When the determination unit 313 receives subsequence information (vector) from the second extraction processing unit 312, it calculates a value (inference conversion value) converted using a conversion function. The determination unit 313 calculates a value (learning conversion value) converted using a conversion function from learned data recorded in the recording unit 2. The determination unit 313 determines whether a member with the same value as the inference conversion value exists in the set of learning conversion values, and if not, determines that the subsequence information of this inference conversion value has not been learned. Details are described in Reference 1 below. (Reference 1) Hisashi Koga, "Similarity Search Technology Using Hashing and Its Applications," IEICE Fundamentals Review Vol. 7 No. 3, January 2014.
[0039] 1 , inference processing unit 32 performs inference processing on the learned subsequence information output from determination unit 313 and infers tag values corresponding to the vector sequences of each learned subsequence information. That is, inference processing unit 32 infers tag values of each vector sequence included in the inference data using the subsequence information determined by determination unit 313 to have been learned.
[0040] 6 is a block diagram showing a specific configuration of the inference processing unit 32. As shown in FIG.
[0041] The neural network 44 acquires the learned subsequence information output from the determination unit 313, performs internal calculations based on this data, and outputs the results.
[0042] The values of the internal parameters of the neural network 44 can be obtained by the same process as that of the learning unit 1. Therefore, the entity of the neural network 44 of the inference processing unit 32 may be the same as the neural network 12 of the learning unit 1.
[0043] In addition, the inference processing unit 32 acquires the partial vector sequence corresponding to the learned partial sequence information output from the judgment unit 313, and outputs and presents this partial vector sequence to the user together with the output result of the neural network 44.
[0044] That is, the inference processing unit 32 includes a neural network and infers the tag value of each partial vector sequence through its internal calculations.
[0045] The inference unit 3, which includes a first extraction processing unit 311, a second extraction processing unit 312, and a determination unit 313, receives inference data, which is a series of vectors, and extracts a set of learned vector series included in the vector series of the inference data by referring to the learned data recorded in the recording unit 2. The inference unit 3 infers a tag value corresponding to the extracted learned vector series.
[0046] Next, the operation of the deep learning device 100 according to this embodiment configured as described above will be described. First, the processing by the learning unit 1 shown in FIG.
[0047] 7 is an explanatory diagram showing the learning data to be processed by the learning unit 1 shown in FIG. 1. The learning data includes a plurality of (three in the figure) individual learning data D1 to D3. Each of the individual learning data D1 to D3 is data consisting of a vector sequence made up of a sequence of a plurality of vectors "l" and a tag value assigned to this vector sequence. The learning unit 1 learns the correspondence between the vector sequence and the tag value based on each of the individual learning data D1 to D3.
[0048] In FIG. 7 , the j-th vector in the vector sequence of the ith individual training data is indicated by a superscript suffix "i" of "l" and a subscript suffix "j". Hereinafter, this is represented as "l_i,j". For example, the second vector in the vector sequence of the first individual training data D1 is represented as "l_1,2". Furthermore, the length of the vector sequence of the ith individual training data is represented as "ni". For example, the length of the vector sequence of the first individual training data D1 is "n1". The lengths n1 to n3 of the vector sequences of the individual training data D1 to D3 do not have to be the same.
[0049] 8 is a flowchart showing the processing procedure performed by the learning unit 1. First, in step S11 of FIG. 8, the learning unit 1 determines whether or not unprocessed individual learning data is included in the given learning data. If unprocessed individual learning data is included (S11; YES), the process proceeds to step S12; if not (S11; NO), the process ends.
[0050] In step S12, the learning unit 1 extracts one piece of individual learning data from the unprocessed individual learning data.
[0051] In step S13, the learning unit 1 acquires the sequence length "n*" of the individual learning data extracted in the process of step S12. The serial number of the individual learning data is indicated by "*". In the example shown in FIG. 7, "n*" is "n1 to n3".
[0052] In step S14, the learning unit 1 executes a learning process. Specifically, the learning unit 1 updates the parameters of each LSTM 14 of the RNN unit 11 shown in Fig. 2 and the internal parameters of the neural network 12 by the backpropagation method. Fig. 9 is an explanatory diagram showing a specific flow of the process of step S14, and the signal flow in the backpropagation method is depicted in the block diagram shown in Fig. 2.
[0053] As shown in Figure 9, when a vector sequence (l_*,1 to l_*,n*) of individual training data is input to each LSTM 14 (14-1 to 14-n*) of the RNN unit 11, the final-stage hidden state "h_*,n*" (see Figure 10 described later) is output to the neural network 12. The "*" in Figure 9 indicates the serial number of the individual training data. For example, the vector "l_*,1" is input to LSTM 14-1, and the vector "l_*,2" is input to LSTM 14-2.
[0054] The LSTM 14-1 outputs the hidden state “h_*,1” to the LSTM 14-2. The hidden state “h_*,n*” of the final stage LSTM 14-n* is output to the neural network 12.
[0055] The neural network 12 performs learning based on the hidden state vector output from the RNN unit 11 and calculates the tag value of the individual learning data.
[0056] The loss calculation unit 13 shown in Figure 9 calculates the error between the tag value calculated by the neural network 12 and the correct tag value. The learning unit 1 repeats the operation of propagating and updating the error of the parameters of each unit in the backward direction a specified number of times, starting from the error information calculated by the loss calculation unit 13. That is, the learning unit 1 uses the error backpropagation method to optimize the internal parameters of the neural network 12 and each LSTM 14 of the RNN unit 11 so as to minimize the above-mentioned error, thereby learning the correspondence between the data sequence of the individual training data and the tag value. Thereafter, the process proceeds to step S15 shown in Figure 8.
[0057] In step S15, the learning unit 1 adds the sequence data of the individual learning data learned in this process to the learned data recorded in the recording unit 2 as a learned vector sequence. Figure 10 is an explanatory diagram showing the operation of step S15. The vector sequence of the individual learning data is input to the LSTM 14 whose internal parameter values have been determined, and the hidden state "h_*,n*" of the final stage is added to the learned data in the recording unit 2 and recorded as a learned vector. Note that the neural network 12 and loss calculation unit 13 are not used in this step S15.
[0058] As described above, the learned vector sequence learned by the learning unit 1 is additionally recorded in the recording unit 2.
[0059] Next, the processing of the inference unit 3 shown in Fig. 1 will be described. Fig. 11 is an explanatory diagram showing the inference data to be processed by the inference unit 3. The inference data includes a plurality of (three in the figure) individual inference data D11 to D13. Each of the individual inference data D11 to D13 is data consisting of a vector sequence consisting of a sequence of a plurality of vectors "d".
[0060] In Figure 11, the jth vector in the vector sequence of the i-th individual inference data is indicated by a superscript "i" and a subscript "j" of "d". Hereinafter, this will be expressed as "d_i,j". For example, the second vector in the vector sequence of the first individual inference data D11 will be expressed as "d_1,2". Furthermore, the length of the vector sequence of the i-th individual inference data will be expressed as "mi". For example, the length of the vector sequence of the first individual inference data D11 is "m1". The lengths m1 to m3 of the vector sequences of each individual inference data D11 to D13 do not have to be the same.
[0061] For example, the individual inference data D11 has a vector sequence "d_1,1", "d_1,2", ..., "d_1,m1". The inference unit 3 uses the learned parameters output from the learning unit 1 to infer tag values corresponding to the vector sequences of each of the individual inference data D11 to D13.
[0062] FIG. 12 is an explanatory diagram showing a vector sequence included in individual inference data, and the vectors included in the vector sequence are indicated by "d." As shown in FIG. 12, vector "d" is a mixture of a learned sequence and noise. "Noise" refers to a vector not included in the learned sequence.
[0063] 1 and 3 extracts a desired partial vector sequence according to the application from all partial vector sequences (the number of partial sequences shown in the above-mentioned formula (1)) included in the vector sequence of the input inference data. In this embodiment, partial sequences consisting only of consecutive vectors from the vector sequence of the original inference data are extracted, and the vectors in these partial sequences are not rearranged.
[0064] 13A and 13B are explanatory diagrams showing the flow of processing by the extraction unit 31 shown in Fig. 1. The first RNN unit 21-1 receives the vector series d1 to dm contained in the individual inference data as input to each LSTM, the second RNN unit 21-2 receives the vector series d2 to dm as input to each LSTM, and the mth RNN unit 21-m receives the vector series dm as input to its LSTM.
[0065] That is, m RNN units 21 (21-1 to 21-m) are generated for a vector sequence consisting of vectors d1 to dm contained in the individual inference data, and the LSTM of the i-th RNN unit 21-i processes the input of vectors d1 to dm. Note that the values of the internal parameters of this LSTM are those obtained as a result of the learning process performed by the learning unit 1 in step S14 of Figure 8.
[0066] The second extraction processing unit 312 outputs the hidden states hi_j (i≦j<m) of all loops to the determination unit 313. The determination unit 313 determines whether or not each vector sequence has been learned, based on the data of the hidden states.
[0067] The hidden state hi_j is a vector containing sequence information for a subsequence consisting of the i-th to j-th vectors in the vector sequence of the original inference data. Therefore, by using the m RNN units 21-1 to 21-m shown in Figures 13A and 13B to determine the learning history of the hidden states of all of the LSTM loops and extracting only the learned ones, it is possible to extract all learned subsequences subject to the above-mentioned constraints from the vector sequence of the original inference data. In other words, it is possible to extract learned subsequences using the noise-removed vector "d" shown in Figure 12. The m RNN units 21-1 to 21-m shown in Figures 13A and 13B have the functions of the first extraction processing unit 311 and the second extraction processing unit 312 shown in Figure 1.
[0068] 6, the inference processing unit 32 includes a neural network 44. The inference processing unit 32 performs inference processing based on the set of learned subsequences for each individual inference data output from the determination unit 313, and executes processing to infer a tag value for each individual inference data.
[0069] Fig. 14 is a flowchart showing the processing procedure by the inference unit 3. Fig. 15 is a flowchart showing the detailed processing procedure of step S33 shown in Fig. 14. The processing of steps S31 to S33 shown in Fig. 14 is processing for extracting learned subsequence information, and is executed by the extraction unit 31. The processing of steps S34 to S36 is processing for inferring tag values, and is executed by the inference processing unit 32.
[0070] In step S31, the second extraction processing unit 312 determines whether or not unprocessed individual inference data exists in the given inference data. If unprocessed individual inference data exists (S31; YES), the process proceeds to step S32; if not (S31; NO), the process ends.
[0071] In step S32, the second extraction processing unit 312 extracts one of the unprocessed individual inference data.
[0072] In step S33, the second extraction processing unit 312 executes a process of extracting a learned sequence from the individual inference data sequence. The detailed process procedure of step S33 will be described below with reference to the flowchart shown in FIG.
[0073] In step S51 of FIG. 15, the second extraction processing unit 312 acquires the sequence length m* of the individual inference data (the serial number is assumed to be "*") extracted in the processing of step S32 of FIG.
[0074] In step S52, the second extraction processing unit 312 sets the value of a variable p, which is used in the following processing, to 1.
[0075] In step S53, the second extraction processing unit 312 inputs the vectors "d_*,p" to "d_*,m*" contained in this individual inference data into each LSTM loop (see Figures 13A and 13B) contained in the pth RNN part, and calculates the hidden states of all LSTMs.
[0076] In step S54, the second extraction processing unit 312 sets the value of a variable q to 1, which is used in the following processing.
[0077] In step S55 , the determination unit 313 receives the hidden state “h_*, p_q” output from the second extraction processing unit 312 .
[0078] In step S56, the determination unit 313 determines whether the hidden state “h_*,p_q” has been learned. If it has been learned (S56; YES), the process proceeds to step S57. If not (S56; NO), the process proceeds to step S58.
[0079] In step S57, the determination unit 313 registers the hidden state “h_*,p_q” that has been determined to have been learned, i.e., the set of learned partial sequence information and the corresponding partial vector sequence “d_*,p” to “d_*,q”.
[0080] In step S58, the second extraction processing unit 312 determines whether q = m*, and if q = m* (S58; YES), proceeds to step S60; if not (S58; NO), proceeds to step S59.
[0081] In step S59, the second extraction processing unit 312 increments the variable q and returns the process to step S55.
[0082] In step S60, the second extraction processing unit 312 determines whether p = m*, and if p = m* (S60; YES), proceeds to step S34 in Figure 14, and if not (S60; NO), proceeds to step S61.
[0083] In step S61, the second extraction processing unit 312 increments the variable p and returns the process to step S53.
[0084] As described above, by performing the series of processes in steps S55 to S57 while incrementing the value of variable q by 1 as long as variable q does not exceed m*, it is possible to extract learned subsequences from a vector sequence consisting of vectors "d_*,p" to "d_*,m*." Furthermore, by performing the series of processes in steps S53 to S59 while incrementing the value of variable p by 1 as long as the value of variable p does not exceed m*, it is possible to generate all learned subsequences included in the vector sequence of the inference data with serial number "*," i.e., a set of learned subsequence information.
[0085] Therefore, by the process shown in FIG. 15, i.e., the process of step S33 in FIG. 14, all learned subsequence information included in the individual inference data extracted in step S32 in FIG. 14 can be extracted.
[0086] 14, it is determined whether or not there is a "pair of learned partial sequence information and partial vector sequence" extracted in step S33. That is, if there is a "pair of learned partial sequence information and partial vector sequence" (S34; YES), the process proceeds to step S35; if not (S34; NO), the process returns to step S31.
[0087] In step S35, the inference processing unit 32 extracts one "pair of learned partial sequence information and partial vector sequence" from the set of "pairs of learned partial sequence information and partial vector sequence" extracted in the process of step S33.
[0088] In step S36, inference processing unit 32 performs inference processing using the learned sequence vector extracted in step S35. Then, processing returns to step S34. That is, as long as an unprocessed learned sequence exists in the set of learned subsequence information, inference processing unit 32 extracts the unprocessed learned sequence from the set of learned subsequence information and infers the corresponding tag value.
[0089] Fig. 16 is an explanatory diagram showing the processing by the inference processing unit 32. As shown in Fig. 16, the learned subsequence information "h_*, p_q" output from the determination unit 313 is output to the neural network 42. "*" shown in Fig. 16 indicates the serial number of the individual inference data.
[0090] The neural network 42 calculates a tag value based on this learned subsequence information. The inference processing unit 32 also outputs the partial vector sequence output from the determination unit 313, and clearly indicates to the user the vector sequence corresponding to the tag value. In this way, inference processing using the learned subsequence information is executed.
[0091] As described above, the deep learning device 100 according to this embodiment includes a first extraction processing unit 311 that receives inference data having a vector sequence and extracts a partial vector sequence that includes at least one vector from each of the vectors included in the vector sequence; a second extraction processing unit 312 that receives the partial vector sequence and circulates a recurrent neural network to extract partial sequence information that reflects the overall information of the partial vector sequence; a recording unit 2 that records the learned vector sequence; a determination unit 313 that performs a vector approximation search process on the partial sequence information by referencing the learned vector sequence and determines whether the partial sequence information has been learned; and an inference processing unit 32 that infers tag values of each vector sequence included in the inference data using the partial sequence information determined by the determination unit 313 to have been learned.
[0092] In this embodiment, even if the vector series of the inference data contains a mixture of learned vectors and noise vectors, only the learned vectors are extracted and the inference process is performed, thereby making it possible to improve the inference accuracy of the tag values of each individual inference data.
[0093] For example, this is useful when dealing with conversational data written in natural language. Conversational data often contains not only important utterances with noteworthy features, but also casual conversations and redundant utterances. In such cases, this method enables inference processing that extracts only utterances with desired attributes.
[0094] In this embodiment, the second extraction processing unit 312 includes a plurality of LSTMs, and the hidden states of the vector sequence are recursively extracted using the LSTMs, making it possible to extract subsequence information with high accuracy.
[0095] In addition, the second extraction processing unit 312 converts each vector included in the partial vector sequence into a hidden state vector using each LSTM, and extracts the hidden state in the final-stage LSTM as partial sequence information, making it possible to extract partial sequence information with high accuracy.
[0096] In this embodiment, the inference processing unit 32 circulates each vector in the vector sequence in the inference data to each LSTM to infer the tag value of each vector sequence, thereby enabling highly accurate inference processing.
[0097] In this embodiment, the learned sequence vectors are recorded in the recording unit 2, and the determination unit 313 performs approximate search processing based on the learned sequence vectors recorded in the recording unit 2, making it possible to set learned subsequence information with high accuracy.
[0098] [Description of the Second Embodiment] Next, a second embodiment will be described. Fig. 17 is a block diagram showing the configuration of a deep learning device 101 according to the second embodiment. The deep learning device 101 according to the second embodiment differs from the first embodiment shown in Fig. 1 in that the inference processing unit 32A includes an RNN unit 41 (see Fig. 18) and a learning history determination unit 321, and in that the recording unit 2A stores all-sequence learned data, first-sequence learned data, and second-sequence learned data. Other components are the same as those in Fig. 1 described above, so the same reference numerals are used and a description of the components will be omitted.
[0099] 18 is a block diagram showing the configuration of the inference processing unit 32 A. As shown in FIG. 18, the inference processing unit 32 A includes an RNN unit 41 including a plurality of LSTMs 43, a learning history determination unit 321, and a neural network 42.
[0100] The learning history determination unit 321 determines whether or not all combinations of learned sequences of a first vector sequence and a second vector sequence, which will be described later, have been learned, and outputs the determination result to the neural network 42. Details will be described later.
[0101] In a deep learning device 101 according to the second embodiment, individual learning data input to a learning unit 1 includes a first vector sequence and a second vector sequence. Also, inference data input to an inference unit 3 includes a first vector sequence and a second vector sequence.
[0102] 19 is an explanatory diagram showing the individual learning data input to the learning unit 1. As shown in FIG. 19, each of the individual learning data D31 to D33 includes a first vector sequence and a second vector sequence. Furthermore, each of the individual learning data D31 to D33 includes a tag value indicating the relationship between the first vector sequence and the second vector sequence. Furthermore, the first vector sequence and the second vector sequence of each individual learning data include noise vectors other than the learned vector sequences.
[0103] The learning unit 1 of the deep learning device 101 according to the second embodiment learns the relationship between a first vector sequence and a second vector sequence included in the same individual learning data as a given tag value.
[0104] 20 is an explanatory diagram showing the individual inference data input to the first extraction processing unit 311. As shown in FIG. 20, each of the individual inference data D41 to D43 has a first vector sequence and a second vector sequence. The first vector sequence and the second vector sequence of each of the individual inference data D41 to D43 contain a mixture of vectors that become noise other than the learned vector sequence.
[0105] The inference unit 3 of the deep learning device 101 according to the second embodiment infers the relationship between the vector sequences of each individual inference data based on the learning results of the learning unit 1, and outputs a tag value that is the inference result.
[0106] 21 is an explanatory diagram showing an overview of the process of extracting learned subsequences by the learning unit 1. In FIG. 21, the suffix "i-k" in the upper right corner of the vector "l" indicates the k-th vector sequence of the i-th individual learning data. That is, the vector "l_i-k,j" indicates the j-th vector of the k-th vector sequence of the i-th individual learning data.
[0107] 21 , the learning unit 1 executes a process of inputting a first vector sequence included in the individual training data D31 into the RNN described above to generate a hidden state, thereby generating first-sequence trained data. The learning unit 1 records the generated first-sequence trained data in the recording unit 2A.
[0108] The learning unit 1 executes a process of inputting a second vector sequence included in the individual training data D31 into an RNN to generate a hidden state, thereby generating second-sequence trained data. The learning unit 1 records the generated second-sequence trained data in the recording unit 2A.
[0109] The learning unit 1 inputs a vector (see FIG. 21 ) concatenated with the first and second vector sequences of the individual learning data that is the subject of the learning process into the RNN, and executes a process of generating a hidden state of the final stage. The learning unit 1 performs a learning process using the hidden state, and generates all-series learned data. The learning unit 1 records the generated all-series learned data in the recording unit 2A. The same process is performed for individual learning data D32 and onwards.
[0110] 22 is an explanatory diagram showing the first vector sequence and the second vector sequence of the individual inference data. In FIG. 22, each vector included in the first vector sequence and the second vector sequence is indicated by "d." As shown in FIG. 22, the first vector sequence includes the learned sequences (1-1), (1-2), and noise. The second vector sequence includes the learned sequences (2-1), (2-2), and noise.
[0111] As shown in Fig. 22, the inference data includes a first vector sequence and a second vector sequence. The first extraction processing unit 311 shown in Fig. 17 extracts partial vector sequences from each of the first and second sequence vectors. The second extraction processing unit 312 extracts partial sequence information from the partial vector sequences extracted from each vector sequence. The second extraction processing unit 312 then outputs a set of the partial vector sequence and the partial sequence information to the determination unit 313.
[0112] 17 , the determination unit 313 determines whether partial sequence information corresponding to the partial vector sequence of the first vector sequence of the inference data acquired from the second extraction processing unit 312 and partial sequence information corresponding to the partial vector sequence of the second vector sequence of the inference data have been learned. If learned, the determination unit 313 outputs the corresponding partial vector sequence to the inference processing unit 32A.
[0113] The inference processing unit 32A of the inference unit 3 generates a hidden state of the final stage when the learned partial vector series of the first vector series and the second vector series of the individual inference data shown in Fig. 20, i.e., the vector obtained by concatenating the learned partial vector series of the first vector series and the learned partial vector series of the second vector series obtained from the determination unit 313, is input to the RNN. Then, by referring to the all-series learned data in the recording unit 2A shown in Fig. 17, it is determined whether the hidden state has been learned, and if it has been learned, a tag value is inferred.
[0114] That is, the first vector sequence and the second vector sequence of each individual inference data contain noise vectors other than the learned vector sequence. Therefore, the inference unit 3 first removes the noise vectors from each vector sequence. Specifically, the inference unit 3 uses the first sequence of learned data recorded in the recording unit 2A to extract a learned sequence from the first vector sequence of the individual inference data. Furthermore, the inference unit 3 uses the second sequence of learned data to extract a learned sequence from the second vector sequence of the individual inference data.
[0115] The learned sequences extracted individually from the first vector sequence and the second vector sequence are trained as the respective sequence vectors of any of the individual training data. However, not all combinations of the learned sequences of the first vector sequence and the second vector sequence are necessarily trained as a whole.
[0116] Therefore, the learning history determination unit 321 shown in Fig. 17 determines whether or not all combinations of learned sequences of the extracted first vector sequence and second vector sequence have been learned, as shown in Fig. 23, and performs tag value inference only for combinations that have been learned. This determination process is performed while referring to all sequence learned data recorded in the recording unit 2A.
[0117] In FIG. 23, it is determined whether or not each combination of learned sequences (1-1), (1-2), (2-1), and (2-2) has been learned.
[0118] That is, the inference processing unit 32A determines whether or not all vector sequences obtained by combining and concatenating the learned sequences of the first sequence vector and the learned sequences of the second sequence vector have been learned, and uses the vector sequences determined to have been learned to infer the tag values of each vector sequence included in the inference data.
[0119] Next, the operation of the learning unit 1 and the inference unit 3 in the deep learning device 101 according to the second embodiment will be described. Fig. 24 is a flowchart showing the learning processing procedure by the learning unit 1. First, in step S71 in Fig. 24, the learning unit 1 determines whether or not unprocessed individual learning data exists among each individual learning data included in the input learning data.
[0120] If unprocessed individual learning data exists (S71; YES), the process proceeds to step S72; otherwise, the process ends.
[0121] In step S72, the learning unit 1 extracts one piece of unprocessed individual learning data.
[0122] In step S73, the learning unit 1 obtains the length "n*-1" of the first vector sequence and the length "n*-2" of the second vector sequence contained in the extracted individual learning data.
[0123] In step S74, the learning unit 1 executes a learning process. The learning process is similar to the process shown in step S14 of Fig. 8, and as shown in Fig. 25, the internal parameters of the LSTM of the RNN unit and the neural network are updated by the error backpropagation method. The first and second vector sequences of the individual learning data are input consecutively to the LSTM loop of the RNN unit.
[0124] In step S75, the learning unit 1 successively inputs the first and second vector sequences of the individual learning data into the loop of the LSTM whose parameters have been updated in the processing of step S74, and records the final hidden state in the all-series learned data of the recording unit 2A, as shown in FIG. 26 .
[0125] In step S76, the learning unit 1 inputs the first vector sequence of the individual training data into the LSTM loop whose parameters have been updated in the processing of step S74, as shown in FIG. 27(a), and records the hidden state of the final output as the first sequence of trained data in the recording unit 2A.
[0126] 27(b), the learning unit 1 inputs the second vector sequence of the individual learning data to the loop of the LSTM whose parameters have been updated in the processing of step S74, and records the hidden state of the final output as second-series learned data in the recording unit 2 A. In this way, the learned data of the first vector sequence, the second vector sequence, and all vector sequences learned by the learning unit 1 can be recorded in the recording unit 2 A.
[0127] Next, a description will be given of the operation of the inference unit 3. Figures 28A and 28B are flowcharts showing the processing steps performed by the inference unit 3.
[0128] 28A, the inference unit 3 determines whether or not there is unprocessed individual inference data. If there is unprocessed individual inference data (S81; YES), the process proceeds to step S82. If there is not (S81; NO), the process ends.
[0129] In step S82, the inference unit 3 extracts one unprocessed individual inference data.
[0130] In step S83, the first extraction processing unit 311 of the inference unit 3 extracts a group of partial vector sequences from the first vector sequence included in the extracted individual inference data.
[0131] In step S84, the first extraction processing unit 311 extracts a group of partial vector sequences from the second vector sequence included in the extracted individual inference data.
[0132] In step S85, the second extraction processing unit 312 inputs the partial vector sequence group of the first vector sequence included in the individual inference data extracted by the first extraction processing unit 311 in the processing of step S83 into the LSTM and extracts the final-stage hidden state group as a partial vector sequence information group.Then, the determination unit 313 references the first sequence learned data in the recording unit 2A to determine whether each partial vector sequence information has been learned, and if it has been learned, extracts the corresponding partial vector sequence.Specifically, as shown in FIG. 22, a vector sequence that is a learned sequence is extracted for each vector "d" included in the first vector sequence.
[0133] In step S86, the second extraction processing unit 312 inputs the partial vector sequence group of the second vector sequence included in the individual inference data extracted by the first extraction processing unit 311 in the processing of step S84 into the LSTM, and extracts the final-stage hidden state group as a partial vector sequence information group.Then, the determination unit 313 references the second sequence learned data in the recording unit 2A to determine whether each partial vector sequence information has been learned, and if it has been learned, extracts the corresponding partial vector sequence.Specifically, as shown in FIG. 22, a vector sequence that is a learned sequence is extracted for each vector "d" included in the second vector sequence.
[0134] In step S87, the inference processing unit 32A generates a list of sequence pairs, which are combinations of the first and second vector sequences extracted in the processes of steps S85 and S86, as shown in FIG.
[0135] 28B, the inference processing unit 32A determines whether or not there are any unprocessed pairs in the sequence pair group list. If there are any unprocessed pairs (S88; YES), the process proceeds to step S89. If not (S88; NO), the process returns to step S81.
[0136] In step S89, the inference processing unit 32A extracts one series pair, for example, the pair (1-1) and (2-1) shown in FIG.
[0137] In step S90, the inference processing unit 32A obtains the length "n*" of the extracted sequence pair.
[0138] In step S91, the inference processing unit 32A inputs the above sequence pairs to the loop of the LSTM 43 of the RNN unit 41 shown in Fig. 18. That is, as shown in Fig. 29, the sequence pairs are input to each LSTM 43 to generate the hidden state "hn*" of the final stage.
[0139] In step S92, the learning history determination unit 321 of the inference processing unit 32A determines whether or not the learning result based on the hidden state “hn*” exists in the all-series learned data of the recording unit 2A. If it exists (S92; YES), the process proceeds to step S93; if not (S92; NO), the process returns to step S88.
[0140] In step S93, the inference processing unit 32A performs inference processing on the hidden state "hn*". That is, it outputs the tag value derived from "hn*". At the same time, it also outputs the corresponding sequence pair and presents it to the user. Thereafter, the process returns to step S88. In this way, it is possible to obtain an inference result for the pair of the first vector sequence and the second vector sequence.
[0141] In this way, in the deep learning device 101 according to the second embodiment, when the inference data includes a first vector sequence and a second vector sequence, learned subsequences of each of the first vector sequence and the second vector sequence are acquired, and tag value inference processing is performed on learned pairs of the learned subsequence of the first vector sequence and the learned sequence of the second vector sequence. Therefore, even when the inference data includes a first vector sequence and a second vector sequence, highly accurate tag value inference processing is possible.
[0142] The deep learning devices 100 and 101 of the present embodiment described above can be, for example, a general-purpose computer system including a CPU (Central Processing Unit, processor) 901, a memory 902, a storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), a communication device 904, an input device 905, and an output device 906, as shown in FIG. 30 . The memory 902 and the storage 903 are storage devices. In this computer system, the CPU 901 executes a predetermined program loaded on the memory 902, thereby realizing each function of the deep learning devices 100 and 101.
[0143] The deep learning devices 100 and 101 may be implemented on a single computer or multiple computers. Furthermore, the deep learning devices 100 and 101 may be virtual machines implemented on a computer.
[0144] The programs for the deep learning devices 100 and 101 can be stored in a computer-readable recording medium such as a HDD, SSD, USB (Universal Serial Bus) memory, CD (Compact Disc), or DVD (Digital Versatile Disc), or can be distributed via a network. The computer-readable recording medium is, for example, a non-transitory recording medium.
[0145] The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the present disclosure.
[0146] REFERENCE SIGNS LIST 1 Learning unit 2, 2A Recording unit 3 Inference unit 11, 21, 41 RNN unit 12, 42, 44 Neural network 13 Loss calculation unit 31 Extraction unit 32, 32A Inference processing unit 100, 101 Deep learning device 311 First extraction processing unit 312 Second extraction processing unit 313 Determination unit 321 Learning history determination unit
Claims
1. A deep learning device comprising: a first extraction processing unit that inputs inference data having a vector series and extracts a partial vector series including at least one vector among each vector included in the vector series; a second extraction processing unit that inputs the partial vector series, circulates a recurrent neural network, and extracts partial series information reflecting the overall information of the partial vector series; a recording unit that records a learned vector series; a determination unit that performs an approximate search process for vectors with reference to the learned vector series for the partial series information and determines whether the partial series information is learned; and an inference processing unit that infers tag values of each vector series included in the inference data using the partial series information determined to be learned by the determination unit.
2. The deep learning device according to claim 1, wherein the second extraction processing unit includes a plurality of LSTMs (Long Short-Term Memory), and each LSTM constructs a recurrent neural network.
3. The deep learning device according to claim 1 or 2, wherein the inference processing unit includes a plurality of LSTMs and a recurrent neural network, and each LSTM circulates the recurrent neural network to infer tag values of each vector series.
4. The deep learning device according to claim 1, further comprising a learning unit that inputs learning data composed of a combination of a vector series and a tag value assigned to the vector series, learns the combination of the vector series and the tag value, and records the learned vector series in the recording unit.
5. The inference data includes a first series vector and a second series vector. The first extraction processing unit extracts partial vectors for the first series vector and the second series vector, respectively. The second extraction processing unit extracts partial series information for the partial vectors extracted from each series vector. The inference processing unit determines whether each of the partial series information is learned for the combination of the first series vector and the second series vector, and infers tag values of each vector series included in the inference data using the partial series information determined to be learned. The deep learning device according to claim 1.
6. A deep learning method for obtaining inference data having a vector series, extracting a partial vector series composed of at least one vector among each vector included in the vector series, inputting and circulating the partial vector series into a recurrent neural network to extract partial series information reflecting the overall information of the partial vector series, performing an approximate search process for vectors with reference to the learned vector series recorded in a recording unit based on the partial series information, determining whether the partial series information is learned, and inferring tag values of each vector series included in the inference data using the partial series information determined to be learned.
7. A program for causing a computer to function as the deep learning device according to claim 1 or 2.
Citation Information
Patent Citations
Information processing program, information processing method, and information processing device
JP2021107967A
Data generation program and method
WO2021111630A1