Neural Networks and Object Recognition Methods
By introducing the evaluation mechanism of the first and second networks into the neural network, the high storage space requirements and long-term processing problems of deep neural networks and recurrent neural networks when processing sequence data are solved, and an efficient and low-storage neural network operation method is realized.
Patent Information
- Application Number
- CN201910222514.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-28
- Filing Date
- 2019-03-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-03-22
AI Technical Summary
When processing sequence data, deep neural networks (DNNs) and recurrent neural networks (RNNs) are prone to high storage space requirements and long-term processing, making it difficult to apply to general feature vectors.
An operation method of a neural network is proposed, including using the first network to generate state information, and the second network evaluates whether the state information meets preset conditions, and decides whether to iterate the application state information on the first network based on the evaluation results.
Through this method, high conversion capability can be maintained without increasing storage space, simplifying the processing flow of the neural network, and improving the processing efficiency of sequence data.
Smart Images

Figure CN110969239B_ABST
Abstract
Description
[0001] This application claims the benefit of Korean Patent Application No. 10-2018-0115882, filed on September 28, 2018, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein in its entirety for all purposes by reference. Technical Field
[0002] The following description relates to neural networks and methods of operating and training neural networks. Background Art
[0003] A neural network is an architecture or structure with a certain number of layers or operations, wherein a certain number of layers or operations are provided for many different machine learning algorithms to work together, process complex data inputs and recognize patterns. Typically, neural networks are in the form of deep neural networks (DNNs) to ensure high translation capabilities or high performance for feature vectors. However, DNNs include multiple layers with various weights and use a large storage space for storing multiple layers. In addition, recurrent networks (e.g., recurrent neural networks (RNNs) configured to process sequence data) perform operations for the desired number of iterations (e.g., the length of the sequence data). Therefore, in addition to sequence data, it is not easy to apply to general feature vectors. In addition, if the length of the sequence data is too long, the processing time also increases. Summary of the invention
[0004] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0005] In one general aspect, a method for operating a neural network including a first network and a second network includes: using the first network to obtain state information output based on input information; using the second network to determine whether the state information satisfies a preset condition; in response to the state information being determined to not satisfy the condition, iteratively applying the state information to the first network; in response to the state information being determined to satisfy the condition, outputting the state information.
[0006] The determining step may include comparing a preset threshold with an evaluation result output from the second network corresponding to the state information.
[0007] The determining step may further include comparing a preset number of iterations with a number of times the state information is iteratively applied to the first network.
[0008] The first network may be configured to iteratively process input information to provide a preset application service, and the second network may be configured to evaluate state information corresponding to an iterative processing result of the first network.
[0009] The operating method of the neural network may further include: using a third network to decode the state information to provide a preset application service.
[0010] The method for operating a neural network may further include: encoding input information into a dimension of state information; and applying the encoded input information to the first network.
[0011] The iteratively applying step may include: encoding the state information into a dimension of the input information; and applying the encoded state information to the first network.
[0012] The input information may include at least one of single data and sequence data.
[0013] The operating method of the neural network may also include: in response to the input information being sequence data, encoding the sequence data into an embedding vector having an input dimension of the first network; and applying the embedding vector to the first network.
[0014] The outputting may include, in response to the input information being sequence data, decoding the state information into the sequence data, and outputting the decoded state information.
[0015] The first network may include a neural network for speech recognition or a neural network for image recognition.
[0016] The first network may include at least one of a fully connected layer, a simple recurrent neural network, a long short-term memory (LSTM) network, and a gated recurrent unit (GRU).
[0017] In another general aspect, a training method for a neural network including a first network and a third network includes: generating state information for each iteration by iteratively applying input information corresponding to training data to the first network based on a preset number of iterations; predicting a result corresponding to the state information for each iteration using the third network; and training the first network based on a first loss between the result predicted for each iteration and a true value corresponding to the input information.
[0018] The neural network training method may further include: training a second network configured to evaluate state information based on an evaluation score of a result predicted at each iteration.
[0019] The step of training the second network may include determining an evaluation score by evaluating a result predicted at each iteration based on the result predicted at each iteration and a true value.
[0020] The step of training the second network may include applying noise to at least a portion of the state information at each iteration.
[0021] The neural network training method may further include: training a third network based on the first loss.
[0022] The training method of the neural network may also include: encoding input information into a dimension of state information; and applying the encoded input information to the first network.
[0023] The generating step may include: encoding the state information into a dimension of the input information; and applying the encoded state information to the first network.
[0024] In another general aspect, a neural network includes: a first network configured to generate state information based on input information; a second network configured to determine whether the state information satisfies a preset condition; and a processor configured to iteratively apply the state information to the first network in response to the state information being determined to not satisfy the condition, and output the state information in response to the state information being determined to satisfy the condition.
[0025] The second network may be further configured to compare a preset threshold with an evaluation result corresponding to the state information output from the second network.
[0026] The second network may be further configured to compare a preset number of iterations with a number of times the state information is iteratively applied to the first network.
[0027] The first network may be configured to iteratively process input information to provide a preset application service, and the second network may be configured to evaluate state information corresponding to an iterative processing result of the first network.
[0028] The neural network may further include: a third network configured to: decode the state information to provide a preset application service.
[0029] The processor may be further configured to: encode the input information into a dimension of the state information, and apply the encoded input information to the first network.
[0030] The processor may be further configured to: encode the state information into a dimension of the input information, and apply the encoded state information to the first network.
[0031] The input information may include at least one of single data and sequence data.
[0032] The neural network may further include an encoder configured to: in response to the input information being sequence data, encode the sequence data into an embedding vector having an input dimension of the first network. The processor may also be configured to apply the embedding vector to the first network.
[0033] The neural network may further include: a decoder configured to: in response to the input information being sequence data, decode the state information into sequence data. The processor may further be configured to output the decoded state information.
[0034] Other features and aspects will be apparent from the following detailed description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flow chart illustrating an example of a method of operation of a neural network.
[0036] Figure 2 and Figure 3 is a flow chart illustrating an example of a method of operation of a neural network.
[0037] Figure 4A and Figure 4B An example of a neural network is shown.
[0038] Figure 5 An example of a configuration of a neural network is shown.
[0039] Figure 6 is a flow chart illustrating an example of a training method of a neural network.
[0040] Figures 7 to 9 An example of a training method for a neural network is shown.
[0041] Fig.10 An example of a configuration of a neural network is shown.
[0042] Throughout the drawings and detailed description, unless otherwise described or provided, the same figure reference numerals will be understood to refer to the same elements, features, and structures. The drawings may not be to scale, and the relative sizes, proportions, and depictions of the elements in the drawings may be exaggerated for clarity, illustration, and convenience. DETAILED DESCRIPTION
[0043] The features described herein can be implemented in different forms and should not be interpreted as being limited to the examples described herein. On the contrary, the examples described herein are provided only to illustrate some of the many possible ways to implement the methods, devices and / or systems described herein that will be clear after understanding the disclosure of the present application.
[0044] Terms such as first, second, A, B, (a), and (b) may be used herein to describe components. However, these terms are not used to define the nature, order, or sequence of the corresponding components, but are only used to distinguish the corresponding components from other components. For example, a component referred to as a first component may be alternatively referred to as a second component, and another component referred to as a second component may be alternatively referred to as a first component.
[0045] If the specification states that a first component is “connected,” “coupled,” or “coupled” to a second component, the first component may be directly “connected,” “coupled,” or “coupled” to the second component, or a third component may be “connected,” “coupled,” or “coupled” between the first and second components. However, if the specification states that a first component is “directly connected” or “directly coupled” to a second component, the third component may not be “connected” or “coupled” between the first and second components. Similar expressions (e.g., “between” and “directly between” and “adjacent to” and “immediately adjacent to”) should also be interpreted in this manner.
[0046] The terms used herein are for the purpose of describing specific examples only and are not intended to limit the present disclosure or claims. Unless the context clearly indicates otherwise, the singular also includes the plural. The terms "include" and "comprising" indicate the presence of the stated features, quantities, operations, elements, components, or combinations thereof, but do not exclude the presence or addition of one or more other features, quantities, operations, elements, components, or combinations thereof.
[0047] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure belongs based on an understanding of the present disclosure. Unless explicitly defined as such herein, terms (such as those defined in general dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be idealized or overly formally interpreted.
[0048] Hereinafter, example embodiments are described with reference to the accompanying drawings. The same reference numerals used herein may represent the same elements throughout.
[0049] Figure 1 is a flow chart illustrating an example of a method of operation of a neural network.
[0050] Reference Figure 1 In operation 110, the neural network uses the first network to obtain state information output based on the input information. In one example, the state information may be a multi-dimensional intermediate calculation vector or an output vector according to a task for providing an application service. In other examples, the state information may be various types of information that may be output from the neural network in response to the input information. In one example, in addition to the convolutional layer, the RNN may also include a subsampling layer, a pooling layer, a fully connected layer, and the like.
[0051] The neural network may be implemented as an architecture having multiple layers including an input image, a feature map, and an output. In the neural network, a convolution operation between the input image and a filter called a kernel is performed, and a feature map is output as a result of the convolution operation. Here, the output feature map is the input feature map, and the convolution operation between the output feature map and the kernel is performed again, and as a result, a new feature map is output. Based on the convolution operation repeatedly performed in this way, the result of recognizing the characteristics of the input image via the neural network may be output.
[0052] In another example, the neural network may include an input source sentence (e.g., a voice input) instead of an input image. In such an example, a convolution operation is performed on the input source sentence using a kernel, and as a result, a feature map is output. The convolution operation is performed again on the feature map output as the input feature map using the kernel, and a new feature map is output. When the convolution operation is repeatedly performed in this way, a recognition result for the feature of the input source sentence can be finally output by the neural network.
[0053] The first network iteratively processes input information to provide a preset application service. For example, the first network may include a neural network for speech recognition or a neural network for image recognition. The first network may be configured as a single network or as a recurrent network. For example, the first network includes at least one of a fully connected layer, a simple recurrent neural network (RNN), a long short-term memory (LSTM) network, and a gated recurrent unit (GRU). The input information includes at least one of single data and sequence data. For example, the input information may be multimedia information including an object to be recognized, such as an image or voice including an object to be recognized. The input information may be information having the same dimension as the dimension of the state information, or may be information having a dimension different from the dimension of the state information. The state information corresponds to the result of processing the input information by the first network and / or the iterative processing result of the first network. For example, depending on the task for providing the application service, the state information may be a multidimensional intermediate calculation vector or an output vector. This is provided only as an example. The state information may be various types of information that may be output from the neural network in response to the input information.
[0054] In operation 120, the neural network determines whether the state information satisfies a preset condition using the second network. The second network evaluates the state information output from the first network. Any type of single network may be applied to the second network. For example, the second network may be included in an evaluation logic configured to evaluate the state information. The condition may be used to determine whether the state information is saturated to a level sufficient to perform a task for providing an application service. For example, the condition may include a condition that an evaluation result corresponding to the state information is greater than a preset threshold and / or the number of times the state information is iteratively applied to the first network corresponds to a preset number of iterations. In operation 120, the neural network evaluates the state information by comparing a threshold with an evaluation result corresponding to the state information, or by comparing the number of iterations with the number of times the state information is iteratively applied to the first network.
[0055] In operation 130, when it is determined in operation 120 that the state information does not satisfy the condition, the neural network iteratively applies the state information to the first network. In one example, by iteratively applying the state information output from the first network to the first network, high conversion capability can be ensured without using a large storage space for storing a DNN including a plurality of layers.
[0056] In operation 140, when it is determined that the state information satisfies the condition, the neural network outputs the state information. For example, the state information output in operation 140 may be output as a final result through a softmax layer.
[0057] Figure 2 1 is a flow chart showing an example of an operation method of a neural network. The neural network includes a first network denoted by f and a second network denoted by d. The first network represents a network configured to process an input vector x corresponding to input information and generate an output vector o corresponding to state information. The second network represents an evaluation network configured to evaluate the output vector o and determine whether to iteratively apply the output vector o to the first network.
[0058] Reference Figure 2 Once input information is input to the neural network (Input x) in operation 210 , the neural network applies the input information to the first network (f(x)) in operation 220 and outputs state information (o) in operation 230 .
[0059] In operation 240, the neural network evaluates the state information using the second network. For example, in operation 240, the neural network determines whether the evaluation result (d(o)) satisfies a preset threshold. Here, the neural network determines whether the evaluation result (d(o)) satisfies a threshold set as a hyperparameter. In one example, a hyperparameter is a value that can be set at a predetermined value. Figure 2The parameters are set before the method of operating the neural network starts. Optionally, when the evaluation result (d(o)) does not meet the threshold, the neural network determines whether the number of times the state information is iteratively applied to the first network reaches a preset number of iterations (e.g., a maximum number of iterations).
[0060] When it is determined in operation 240 that the evaluation result (d(o)) satisfies the threshold or the number of times the state information is iteratively applied to the first network reaches the maximum number of iterations, the neural network outputs the state information (Output o) in operation 260. For example, if the threshold = 0.7 and the evaluation result (d(o)) = 0.8, the neural network may pause the loop of the first network (f) and output the state information (o).
[0061] On the contrary, when it is determined in operation 240 that the evaluation result (d(o)) does not satisfy the threshold and the number of times the state information is iteratively applied to the first network does not reach the maximum number of iterations, the neural network encodes the state information into the dimension of the input information to iteratively process the state information (x←o) in operation 250. In operation 220, the neural network applies the encoded state information to the first network.
[0062] In one example, since it is difficult to obtain a ground truth evaluation result from the second network before training, it is necessary to train the second network using a method different from the usual method. Figure 8 and Fig. 9 To describe the training method of the second network.
[0063] In one example, the neural network may use the second network to evaluate how helpful the state information output from the first network is for the task of providing the application service. That is, the neural network may evaluate (i.e., determine) whether the state information is sufficient to perform the task. When the evaluation result of the second network is not satisfactory, the neural network may further perform iterative processing using the first network. When the evaluation result of the second network is satisfactory, the neural network may output the state information.
[0064] In one example, the neural network is also referred to as a self-determining recurrent neural network (RNN), where the neural network itself determines whether to perform an iterative process.
[0065] Figure 3 is a flowchart showing another example of a method of operating a neural network. Figure 3Once input information is input to the neural network (Input x) in operation 310, the neural network encodes the input information into a dimension of state information (x→o) in operation 320 before applying the input information to the first network. The neural network applies the encoded input information to the first network (f(o)) in operation 330, and outputs the state information (o) from the first network in operation 340.
[0066] In operation 350, the neural network evaluates the state information using the second network. For example, in operation 350, the neural network determines whether the evaluation result (d(o)) satisfies a preset threshold or whether the number of times the state information is iteratively applied to the first network reaches a preset number of iterations (e.g., a maximum number of iterations).
[0067] When it is determined in operation 350 that the evaluation result (d(o)) satisfies the threshold or the number of times the state information is iteratively applied to the first network reaches a maximum number of iterations, in operation 360, the neural network outputs the state information (Output o).
[0068] In contrast, when it is determined in operation 350 that the evaluation result (d(o)) does not satisfy the threshold and the number of times the state information is iteratively applied to the first network does not reach the maximum number of times, in operation 330, the neural network iteratively applies the state information to the first network (f(o)).
[0069] Figure 4A and Figure 4B An example of a neural network is shown. Figure 4A A perceptron system consisting of three layers is shown. Figure 4A , the input vector to the perceptron system passes through the first layer (Layer 1), the second layer (Layer 2), and the third layer (Layer 3). The probability distribution of each label is generated through the softmax layer. Here, the three layers (Layer 1, Layer 2, and Layer 3) operate independently and separate weights are stored for each layer.
[0070] Figure 4B The neural network 400 is shown in which the perceptron system is configured by the self-determining RNN 430. When the perceptron system is configured as the self-determining RNN, the maximum number of iterations preset for the self-determining RNN may be 3 or more.
[0071] Here, in Figure 4A The three layers (Layer 1, Layer 2 and Layer 3) and Figure 4BThe conversion capabilities for the input vector or input information 410 may be equal between the self-determining RNN 430 of the layer 1, the layer 2 and the layer 3. However, compared with the three layers (Layer 1, Layer 2 and Layer 3), the self-determining RNN 430 uses a relatively small memory. The self-determining RNN 430 is relatively useful for user devices with limited memory capacity. In addition, the self-determining RNN 430 can provide a faster response rate by deriving a result based on the input information 410 through less than three iterations. Optionally, the self-determining RNN 430 can provide a more accurate result by deriving a result based on the input information 410 through four or more iterations.
[0072] Reference Figure 4B In response to receiving the input information 410, the self-determining RNN 430 generates state information o through the first network 431 i . Status information i is sent to a determiner 435 including a second network 433. The determiner 435 uses the second network 433 to evaluate the state information. i For example, the determiner 435 determines the state information o i Whether the preset conditions are met. When the status information is determined i If the condition is not met, the determiner 435 will i Iteratively applied to the first network 431. When determining the state information o i When the condition is met, the determiner 435 outputs the status information o i The state information output from the self-determining RNN 430 i It can be decoded by the third network 450 and output as the final prediction result 470. Here, the third network 450 can be a decoder (eg, a softmax layer) as a network connected to the back end of the first network 431 and the second network 433 in the entire system of the application system.
[0073] Figure 5 An example of a configuration of a neural network is shown. In the following, reference is made to Figure 5 5 is a diagram to describe the structure of a neural network 500 configured to process sequence data using a self-determining RNN 530.
[0074] For example, when the input information 510 is sequence data (such as a user's speech voice, a text sentence, and a moving image), the neural network 500 encodes the sequence data into an embedding vector having the input dimension of the first network using the encoder 520. The input information 510 can be embedded by the encoder 520 and represented as a single embedding vector. The embedding vector can be iteratively applied (e.g., translated) until a satisfactory result is obtained by the self-determining RNN 530.
[0075] The state information output from the self-determining RNN 530 is decoded into sequence data through the decoder 540 and output as a final prediction result 550 .
[0076] In one example, by selectively applying the encoder 520 and the decoder 540, the self-determining RNN 530 can be applied to various scenarios, for example, {non-sequential data input, non-sequential data output}, {non-sequential data input, sequential data output}, {sequential data input, sequential data output}, and {sequential data input, non-sequential data output}.
[0077] Figure 6 is a flowchart showing an example of a method for training a neural network. Figure 6 , in operation 610, a device for training a neural network (hereinafter, a training device) generates state information for each iteration by iteratively applying input information corresponding to the training data to the first network based on a preset number of iterations. Here, the number of iterations is set to be different for each application service. For example, when an application service requires a relatively high level of recognition results (such as biometric recognition, voice recognition, and user authentication for financial transactions), the number of iterations is set to a relatively high value (e.g., 15 iterations, 20 iterations, 50 iterations, etc.). In addition, when an application requires a relatively low level of recognition results (such as simple unlocking), the number of iterations is set to a relatively low value (e.g., 2 iterations, 3 iterations, 5 iterations, etc.). For example, the input information corresponding to the training data may be an input vector. For example, the state information for each iteration generated in the first network may be Figure 7 o 1 , o 2 , o 3 , o 4 , o 5 .
[0078] In operation 610, the training device encodes the input information into the dimension of the state information and inputs the encoded input information into the first network. In this case, the training device encodes the input information into the dimension of the state information and applies the encoded input information to the first network. Optionally, the training device converts the iteratively applied state information into the dimension of the input information and applies the converted state information to the first network. In this case, the training device encodes the state information into the dimension of the input information and applies the encoded state information to the first network.
[0079] In operation 620, the training device predicts a result corresponding to the state information of each iteration using the third network. For example, the third network may include a prediction layer corresponding to a plurality of softmax layers. Here, the third network may be trained based on the first loss. For example, the result corresponding to the state information for each iteration predicted in the third network may be Figure 7 p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5 ).
[0080] In operation 630, the training device trains the first network based on a first loss between the prediction result and the true value (GT) corresponding to the input information. Figure 7 A method of training the first network and the third network by a training device is described.
[0081] Figure 7 An example of a training method for a neural network is shown. Figure 7 A method of training the first network (f) 710 and the third network (p) 730 is described.
[0082] For example, an unroll training method may be used to train the first network (f) 710. The unroll training method means that the result p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5 ) to learn a loss (e.g., a first loss 750) and perform a back-propagation method, wherein the state information is a result obtained by iteratively applying to the first network (f) 710 for a preset number of iterations (e.g., a maximum number of iterations).
[0083] For example, when the maximum number of iterations is 5, the first network (f) 710 is executed for a total of five iterations (such as the (l-1)th network (f(x)), the (l-2)th network (f(o 1 ))、the (l-3) network (f(o 2 ))、the (l-4) network (f(o 3 )) and the (l-5) network (f(o 4 )), where l represents the total number of iterations), and in response to input information (x) 701 corresponding to the training data being input to the first network (f) 710, state information (o) is generated and output at each iteration. 1 ,o 2 ,o 3 ,o4 ,o 5 ).
[0084] When the status information 1 , o 2 , o 3 , o 4 and 5 When the state information is input to the third network (p) 730, the third network (p) 730 outputs the state information in response to the state information. 1 , o 2 , o 3 , o 4 and 5 The predicted result p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5 Here, similar to the first network (f) 710, the third network (p) 730 is executed as a prediction network configured to predict a result corresponding to the state information for a preset number of iterations, and outputs a result predicted in response to the state information, for example, p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5 ).
[0085] The training device is based on the predicted results p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5 ) and the loss (e.g., the first loss 750) between the true value (GT) 705 corresponding to the input information (x) 701 to train the first network (f) 710. Here, the true value (GT) 705 corresponding to the input information (x) 701 may have the same value for all the first losses 750.
[0086] The first loss 750 is back-propagated to the third network (p) 730 and the first network (f) 710 and is used to train the third network (p) 730 and the first network (f) 710.
[0087] For example, the training device may train the first network (f) 710 to make the predicted result p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5) and the true value (GT) 705 corresponding to the input information (x) 701. In addition, the training device can train the third network (p) 730 to make the predicted result p(o 1 )、p(o 2 )、p(o 3 )、p(o 4 ) and p(o 5 ) and a first loss (GT) 705 corresponding to the input information (x) 701 is minimized. In one example, the first network (f) 710 and the third network (p) 730 may be trained together.
[0088] Figure 8 An example of a training method for a neural network is shown. Figure 8 To describe the method of training the second network (d).
[0089] Reference Figure 8 , the first network (f) 810 is trained using the same unfolding training method as the method used to train the first network (f) 710, and the third network (p) 830 is also trained in the same manner as the third network (p) 730. Here, the difference between the result (p(o)) predicted for the iteration in the third network (p) 830 and the true value (GT) 805 corresponding to the input information (x) 801 corresponds to the first loss (Loss 1) 860. The training device trains the first network (f) 810 to minimize the first loss (Loss 1) 860. Similarly, the training device trains the third network (p) 830 to minimize the first loss (Loss 1) 860.
[0090] The second network (d) 850 evaluates the state information (o) corresponding to the iterative processing result of the first network (f) 810. 1 ,o 2 ,o 3 ,o 4 ,o 5 ). For example, the second network (d) 850 may be trained to predict an evaluation value or evaluation score (d(o)) for evaluating the quality of the corresponding network. Here, various schemes for measuring the evaluation value may be used to determine the predicted evaluation value as various values. For example, the evaluation value may be determined as a continuous value between 0 and 1, or may be determined as a discontinuous value between 0 and 1.
[0091] The difference between the evaluation value determined based on the final prediction result of the third network (p) 830 and the true value (GT) 805 and the evaluation value predicted in the second network (d) 850 may correspond to the second loss (Loss 2) 870. Here, the evaluation value determined based on the final prediction result of the third network (p) 830 and the true value (GT) 805 may be referred to as an evaluation score. The training device trains the second network (d) 850 to minimize the second loss (Loss 2) 870.
[0092] Similar to the first network (f) 810, the second network (d) 850 can be trained using an unfolding training method. The second network (d) 850 is trained based on the output from each iteration point (i.e., the state information (o) of each iteration) through network unfolding. 1 ,o 2 ,o 3 ,o 4 ,o 5 )) to determine the evaluation value or evaluation score. The second network (d) 850 is trained to predict the state information (o 1 ,o 2 ,o 3 ,o 4 ,o 5 ) is used to generate an evaluation score (d(o)) of the result derived. For example, when the prediction accuracy is used for the evaluation value, in response to the specific state information (o) passing through the second network (d) 850, a predicted accuracy value between 0 and 1 may be output.
[0093] In one example, the training device trains the second network (d) 850 based on an evaluation value or evaluation score of a result (p(o)) predicted for each iteration in the third network (p) 830. The training device determines the evaluation value or evaluation score by evaluating the result predicted in the third network (p) 830 based on the result (p(o)) predicted for each iteration in the third network (p) 830 and the true value (GT) 805.
[0094] The training device trains the second network (d) 850 to minimize a second loss (Loss 2) 870 between the output of the second network (d) 850 (i.e., the predicted evaluation value of the second network (d) 850) and the evaluation result or evaluation value corresponding to the result (p(o)) predicted for each iteration in the third network (p) 830. The evaluation result or evaluation value corresponding to the result (p(o)) predicted in the third network (p) 830 may be determined based on the result (p(o)) predicted for each iteration in the third network (p) 830 and the true value (GT) 805.
[0095] The training device may train the first network (f) 810 and the second network (d) 850 together. Alternatively, the training device may train the first network (f) 810, the second network (d) 850, and the third network (p) 830 together. In this case, the first network (f) 810 is trained based on the first loss (Loss 1) 860 and the second loss (Loss 2) 870.
[0096] In one example, when it is assumed that in the later training of the second network (d) 850, the second network (d) 850 is sufficiently saturated or trained, a biased evaluation score may be output. For example, a good evaluation score may always be output. If the biased result value is derived from the second network (d) 850 regardless of the state information (o) input to the second network (d) 850, the training of the second network (d) 850 will be hindered. Here, the training device may apply noise to at least a portion of the state information (o) of each iteration by applying input information corresponding to the actual training data via the first network (f) 810, and enable the second network (d) 850 to balance the ability to assign high scores and the ability to assign low scores. The same or different noise may be assigned to each state information of each iteration.
[0097] Fig. 9 An example of a training method for a neural network is shown. Fig. 9 A process of training a neural network configured to classify an object included in an input image when the neural network is configured as a self-determining RNN will be described.
[0098] For example, when it is assumed that a neural network configured to classify pedestrians included in an input image is trained, the true value in the input image may be a pedestrian. Here, the number of iterations is 2 and the threshold is 0.65. Fig. 9 The example of identifying pedestrians included in the input image (i.e., pedestrians are objects to be identified), but it should be understood that the present invention is not limited to this, the objects to be identified are not limited to pedestrians, and the present invention is not limited to only identifying objects in images, and object recognition in any multimedia information (e.g., video, audio, image, text, etc.) is also feasible.
[0099] When input information (x) corresponding to an input image is input to the first network (f), state information (o) is output from the first network (f). 1 )(f(x)), status information(o 1 ) is applied to the second network (d)(d(o 1 )) and the third network (p)(p(o 1 )). In this example, the state information (o1 )'s third network (p(o 1 )) Predictable and state information (o 1 ) corresponding to the results, for example, each category (such as, vehicles, pedestrians and lanes). 1 ) in the third network (p(o 1 )) is predicted as a predicted value (e.g., vehicle: 0.4, pedestrian: 0.3, lane: 0.3), each predicted value may represent the ability to convert state information (o 1 ) is predicted as the probability of the corresponding category or the corresponding state information (o 1 ) belongs to the corresponding category. Here, the first network (f) can be trained to minimize the difference (Loss1) between the predicted value 0.3 of the category pedestrian and the value 1 corresponding to the true value (GT) pedestrian.
[0100] In addition, since the predicted value of the vehicle is 0.4, which is the largest among the predicted values of each category, the third network (p(o 1 )) is predicted as “vehicle”. The training device is trained based on the third network (p(o 1 ))'s prediction results "vehicle" and the true value (GT) "pedestrian" are evaluated by the third network (p(o 1 )) to determine the evaluation score. Here, since the third network (p(o 1 ))’s prediction result “vehicle” and the true value (GT) “pedestrian” are different from each other, so the training device will target the third network (p(o 1 The evaluation score of the prediction result of )) is determined to be 0.
[0101] The training device is based on the third network (p(o 1 ))'s prediction result evaluation score (0), training is configured to evaluate status information (o 1 )'s second network (d(o 1 )). For example, when the second network (d(o 1 )) is 0.5, the training device trains the second network (d(o 1 )) so that for the third network (p(o 1 ))'s prediction result evaluation score 0 and the second network (d(o 1 )) minimizes the difference (Loss 2) between the predicted evaluation value 0.5. Here, since the number of iterations is set to 2, the training device can 1 ) is iteratively applied to the first network (f).
[0102] When the status information (o 1) is iteratively applied to the first network (f), the first network (f(o 1 )) Output status information (o 2 ), and status information (o 2 ) is applied to the second network (d(o 2 )) and the third network (p(o 2 )). Here, the state information (o 2 )'s third network (p(o 2 )) can be combined with the status information (o 2 ) is predicted as the corresponding value (such as vehicle: 0.3, pedestrian: 0.4, lane: 0.3). Here, the first network (f(o 1 )) can be trained to minimize the difference (Loss 1) between the predicted value 0.4 for the class pedestrian and the value 1 corresponding to the true value (GT) pedestrian.
[0103] Here, since the third network (p(o 2 )) outputs the predicted values of each category, the predicted value of pedestrians is 0.4, which is the largest. Therefore, the third network (p(o 2 )) is predicted as “pedestrian”. The training device is trained based on the third network (p(o 2 ))'s prediction results "pedestrian" and the true value (GT) "pedestrian" evaluate the third network (p(o 2 )) to determine the evaluation score. Here, since the third network (p(o 2 ))’s predicted result “pedestrian” and the true value (GT) “pedestrian” are the same, so the training device will target the third network (p(o 2 The evaluation score of the prediction result of )) is determined to be 1.
[0104] The training device is based on the third network (p(o 2 ))'s prediction result evaluation score (1), the training is configured to evaluate the state information (o 2 )'s second network (d(o 2 )). For example, when the second network (d(o 2 )) is 0.7, the training device trains the second network (d(o 2 )) so that for the third network (p(o 2 ))'s prediction result evaluation score 1 and the second network (d(o 2 )) minimizes the difference (Loss 2) between the predicted evaluation values of 0.7.
[0105] According to an example, the neural network may output a sentence in response to speech recognition. In this example, the third network may compare the entire predicted sentence with the true value sentence, and may determine the predicted value as "1" if all words included in the single sentence match, and may determine the predicted value as "0" if none of the words match. In this case, the third network may predict a result (e.g., a predicted value) corresponding to the state information in the form of a discrete value.
[0106] Alternatively, the third network may predict a result corresponding to the state information in the form of a continuous value between 0 and 1 by assigning a partial point to each word included in a single sentence. In this case, the evaluation value may have a continuous value between 0 and 1, and the second network may be trained to predict a continuous value between 0 and 1.
[0107] In one example, in addition to the above-mentioned fields, neural networks can also be applied to various fields, such as speech recognition, biometric information recognition, text recognition, image capture, sentiment analysis, analysis of stock prices, analysis of oil prices, etc.
[0108] Fig.10 An example of a configuration of a neural network is shown. Fig.10 , the neural network 1000 includes a first network 1010, a second network 1020, and a processor 1030. The neural network 1000 also includes a third network 1040, a memory 1050, an encoder 1060, a decoder 1070, and a communication interface 1080. The first network 1010, the second network 1020, the processor 1030, the third network 1040, the memory 1050, the encoder 1060, the decoder 1070, and the communication interface 1080 communicate with each other via a communication bus 1005.
[0109] The first network 1010 generates and outputs state information based on the input information. The first network 1010 iteratively processes the input information to provide a preset application service. For example, the input information may include at least one of single data and sequence data. For example, the input information may be an image or a sound.
[0110] The second network 1020 determines whether the state information satisfies a preset condition. The second network 1020 evaluates the state information corresponding to the iterative processing result of the first network 1010. For example, the second network 1020 may compare a preset threshold with the evaluation result corresponding to the state information output from the second network 1020. In addition, the second network 1020 may compare a preset number of iterations with the number of times the state information is iteratively applied to the first network 1010.
[0111] In response to the state information being determined not to satisfy the condition in the second network 1020, the processor 1030 iteratively applies the state information to the first network 1010. In response to the state information being determined to satisfy the condition in the second network 1020, the processor 1030 outputs the state information.
[0112] The processor 1030 encodes the input information into the dimension of the state information and applies the encoded input information to the first network 1010. The processor 1030 encodes the state information into the dimension of the input information and applies the encoded state information to the first network 1010.
[0113] In addition, the processor 1030 executes the above reference Figures 1 to 9 At least one method described, or an algorithm corresponding thereto.
[0114] Processor 1030 represents a data processing device configured as hardware, wherein the hardware has circuits in a physical structure for performing a desired operation. For example, the desired operation may include a code or instruction contained in a program. For example, a data processing device configured as hardware may include a microprocessor, a central processing unit (CPU), a processor core, a multi-core processor, a multiprocessor, an application specific integrated circuit (ASIC), and a field programmable gate array (FPGA). Processor 1030 executes the program and controls neural network 1000. The program code executed by processor 1030 is stored in memory 1050.
[0115] The third network 1040 decodes the state information to provide a preset application service.
[0116] The memory 1050 stores input information and state information corresponding to the iterative processing result of the first network 1010. The memory 1050 stores the result of evaluating the state information by the second network 1020. The memory 1050 stores the embedding vector encoded by the encoder 1060 and / or the sequence data obtained by decoding the state information by the decoder 1070. The memory 1050 stores various information generated during the processing of the processor 1030. In addition, various data and programs may be stored in the memory 1050. For example, the memory 1050 may include a volatile memory or a non-volatile memory. The memory 1050 may include a large-capacity storage medium (such as a hard disk) for storing various data.
[0117] For example, when the input information is sequence data, the encoder 1060 encodes the sequence data into an embedding vector having an input dimension of the first network 1010. Here, the processor 1030 applies the embedding vector to the first network 1010.
[0118] For example, when the input information is sequence data, the decoder 1070 decodes the state information into the sequence data. Here, the processor 1030 outputs the decoded sequence data.
[0119] The communication interface 1080 receives input information from the outside of the neural network 1000. In addition, the communication interface 1080 transmits the output of the neural network 1000 to the outside of the neural network 1000.
[0120] The equipment, unit, module, device and other components described herein are implemented by hardware components. Examples of hardware components that can be used to perform the operations described in this application include: controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer may be implemented by one or more processing elements (such as logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other devices or combinations of devices configured to respond and execute instructions in a limited manner to obtain desired results). In one example, a processor or computer includes or is connected to one or more memories storing instructions or software executed by a processor or computer. The hardware components implemented by a processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) to perform the operations described in this application. Hardware components can also access, manipulate, process, create and store data in response to the execution of instructions or software. For simplicity, the singular term "processor" or "computer" can be used in the description of the examples described in this application, but in other examples, multiple processors or computers can be used, or a processor or computer can include multiple processing elements or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller can implement a single hardware component, or two or more hardware components. Hardware components can have any one or more of different processing configurations, wherein examples of different processing configurations include: a single processor, an independent processor, a parallel processor, a single instruction single data (SISD) multiprocessing, a single instruction multiple data (SIMD) multiprocessing, multiple instructions single data (MISD) multiprocessing, and multiple instructions multiple data (MIMD) multiprocessing.
[0121] The method for performing the operations described in the present application is performed by computing hardware (e.g., by one or more processors or computers), wherein the computing hardware is implemented as execution instructions or software as described above to perform the operations described in the present application performed by the method. For example, a single operation or two or more operations may be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations may be performed by one or more processors, or a processor and a controller, and one or more other operations may be performed by one or more other processors, or another processor and another controller. One or more processors, or a processor and a controller may perform a single operation, or two or more operations.
[0122] The instructions or software for controlling the processor or computer to implement the hardware components and perform the method as described above are written as computer programs, code segments, instructions or any combination thereof, to indicate or configure the processor or computer to operate as a machine or special-purpose computer individually or collectively to perform the operations performed by the hardware components and the method as described above. In one example, the instructions or software include machine code (such as, machine code generated by a compiler) directly executed by the processor or computer. In other examples, the instructions or software include high-level code executed by the processor or computer using an interpreter. Ordinary programmers in the art can easily write instructions or software based on the block diagrams and flow charts shown in the drawings and the corresponding descriptions in the specification, wherein the block diagrams and flow charts shown in the drawings and the corresponding descriptions in the specification disclose algorithms for performing the operations performed by the hardware components and the method as described above.
[0123] Instructions or software for controlling a processor or computer-implemented hardware component and performing the method described above, and any associated data, data files, and data structures are recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card-type memory (such as, multimedia card or micro card (for example, Secure Digital (SD) or Extreme Digital (XD))), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, wherein any other device is configured to store instructions or software and any associated data, data files and data structures in a non-temporary manner, and provide instructions or software and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the instructions.
[0124] Although the present disclosure includes specific examples, it will be apparent after understanding the disclosure of the present application that various changes in form and detail can be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein will be understood descriptively, rather than for the purpose of limitation. The description of the features or aspects in each example will be considered to apply to similar features or aspects in other examples. If the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in different ways, and / or replaced or supplemented by other components or their equivalents, suitable results can be achieved. Therefore, the scope of the present disclosure is not limited by specific embodiments, but by the claims and their equivalents, and all changes within the scope of the claims and their equivalents will be interpreted as included in the present disclosure.
Claims
1. A method for object recognition, the method include: obtaining state information output from the first network based on input information including an object to be identified; Determine whether the status information satisfies a preset condition using the second network; In response to the state information being determined to not satisfy the condition, iteratively applying the state information to the first network; In response to the state information being determined to satisfy the condition, outputting the state information associated with the object, The first network is trained by the following steps: generating training state information of each iteration by iteratively applying training input information corresponding to the training data to the first network based on a preset number of iterations; predicting a training result corresponding to the training state information of each iteration using a third network; training the first network based on a first loss between the training result predicted in each iteration and a true value corresponding to the training input information, The first network includes a neural network for speech recognition or a neural network for image recognition.
2. The method according to claim 1, in, The determining step includes comparing a preset threshold with an evaluation result output from the second network corresponding to the state information.
3. The method according to claim 2, in, The step of determining further includes comparing a preset number of iterations with a number of times the state information is iteratively applied to the first network.
4. The method according to claim 1, in, The first network is configured to: iteratively process input information to provide a preset application service; The second network is configured to evaluate state information corresponding to an iterative processing result of the first network.
5. The method according to claim 1, further comprising: include: The state information is decoded using a third network to provide a preset application service.
6. The method according to claim 1, further comprising: include: Encode input information into the dimension of state information; The encoded input information is applied to the first network.
7. The method according to claim 1, in, The steps applied iteratively include: Encode state information into the dimension of input information; The encoded state information is applied to the first network.
8. The method according to claim 1, in, The input information includes at least one of single data and sequence data.
9. The method of claim 1, further comprising: include: In response to the input information being sequence data, encoding the sequence data into an embedding vector having an input dimension of the first network; Apply the embedding vector to the first network.
10. The method according to claim 9, in, The step of outputting includes: in response to the input information being sequence data, decoding the state information into the sequence data, and outputting the decoded state information.
11. The method according to claim 1, in, The first network includes at least one of a fully connected layer, a simple recurrent neural network, a long short-term memory network, and a gated recurrent unit.
12. A method for object recognition, the method include: Generate state information for each iteration by iteratively applying input information corresponding to the training data to the first network based on a preset number of iterations; Using a third network to predict a result corresponding to the state information of each iteration; Training the first network based on a first loss between a result predicted at each iteration and a true value corresponding to the input information; receiving input information including an object to be identified; performing object recognition based on the received input information using the trained first network, The first network includes a neural network for speech recognition or a neural network for image recognition.
13. The method of claim 12, further comprising: include: Based on the evaluation scores of the results predicted at each iteration, a second network configured to evaluate the state information is trained.
14. The method according to claim 13, in, The step of training the second network includes determining an evaluation score by evaluating a result predicted at each iteration based on the result predicted at each iteration and a true value.
15. The method of claim 13, in, The step of training the second network includes applying noise to at least a portion of the state information at each iteration.
16. The method of claim 13, further comprising: include: The third network is trained based on the first loss.
17. The method of claim 13, further comprising: include: Encode input information into the dimension of state information; The encoded input information is applied to the first network.
18. The method of claim 13, in, The generation steps include: Encode state information into the dimension of input information; The encoded state information is applied to the first network.
19. A non-transitory computer-readable recording medium storing instructions, which, when executed by a processor, causes the processor to implement the method of claim 1 or 12.
20. A neural network device for object recognition, include: A first network configured to: generate state information based on input information including an object to be identified; The second network is configured to: determine whether the state information satisfies a preset condition; a processor configured to: in response to the state information being determined not to satisfy the condition, iteratively apply the state information to the first network; In response to the state information being determined to satisfy the condition, outputting the state information associated with the object, The processor is further configured to: generate training state information for each iteration by iteratively applying training input information corresponding to the training data to the first network based on a preset number of iterations; predict a training result corresponding to the training state information for each iteration using a third network included in the neural network device; The first network is trained based on a first loss between the training result predicted at each iteration and the true value corresponding to the training input information, The first network includes a neural network for speech recognition or a neural network for image recognition.
21. The neural network device of claim 20, in, The second network is further configured to compare a preset threshold with an evaluation result corresponding to the state information output from the second network.
22. The neural network device of claim 21, in, The second network is further configured to compare a preset number of iterations with a number of times the state information is iteratively applied to the first network.
23. The neural network device of claim 20, in, The first network is further configured to: iteratively process input information to provide a preset application service; The second network is further configured to evaluate state information corresponding to a result of the iterative processing of the first network.
24. The neural network device of claim 20, in, The third network is configured to decode the state information to provide a preset application service.
25. The neural network device of claim 20, in, The processor is further configured to encode the input information into a dimension of the state information and apply the encoded input information to the first network.
26. The neural network device of claim 20, in, The processor is further configured to encode the state information into a dimension of the input information and apply the encoded state information to the first network.
27. The neural network device of claim 20, in, The input information includes at least one of single data and sequence data.
28. The neural network device of claim 20, further comprising: include: An encoder configured to: in response to the input information being sequence data, encode the sequence data into an embedding vector having an input dimension of the first network, The processor is further configured to apply the embedding vector to the first network.
29. The neural network device of claim 27, further comprising: include: A decoder configured to: in response to the input information being sequence data, decode the state information into sequence data, The processor is further configured to output decoded status information.
Citation Information
Patent Citations
Ankle-foot orthosis to prevent foot drop and bedsore
KR1020180115882A
Deep-reinforcement-learning-based scene text detection method and system
CN108090443A