Method and apparatus for identifying an image

By extracting the similarity between the target layer metadata and the reference layer in the neural network, selecting the corresponding layer and generating a parallelization strategy, the problem of slow training and inference speed of neural network model is solved, and faster image recognition processing is achieved.

CN113408693BActive Publication Date: 2025-07-18SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011450974.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-16
Filing Date
2020-12-09
Publication Date
2025-07-18
Estimated Expiration
2040-12-09

AI Technical Summary

Technical Problem

In the prior art, in the training and inference process of neural network models, it is difficult to quickly converge to the results, and there are problems of inefficiency in the generation and application of parallel processing strategies.

Method used

By extracting the target layer metadata of the neural network, comparing it with the reference metadata of the reference layer, selecting the corresponding layer, and generating a parallelization strategy based on the similarity, the parallel processing of the target layer is achieved.

Benefits of technology

It improves the training and inference speed of neural network models, enhances the efficiency of image recognition processing, and adapts to large-scale and diverse application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113408693B_ABST
    Figure CN113408693B_ABST
Patent Text Reader

Abstract

A method and apparatus for recognizing an image are provided. The method includes: obtaining image data to be recognized as input data of a neural network; performing operations related to each layer of the neural network based on the image data to be recognized to obtain an image recognition result; and outputting the image recognition result, wherein for each of the target layers of the neural network: extracting metadata of the target layer; measuring the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; selecting a corresponding layer among the reference layers based on the similarity; and generating a parallelization strategy for the target layer based on a reference parallelization strategy matching the corresponding layer, and performing parallel processing of operations related to the target layer based on the parallelization strategy using the input data of the target layer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application claims the benefit of Korean Patent Application No. 10-2020-0032233, filed on Mar. 16, 2020, in the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field

[0002] The following description relates to a method and apparatus for identifying an image, and more particularly, to a method and apparatus for implementing image recognition through parallel processing of a neural network model. Background Art

[0003] Techniques for automating image recognition processing have used neural network models implemented by a processor, for example, as a dedicated computing structure. The neural network model can provide a computationally intuitive mapping between an input pattern and an output pattern after a large amount of training. The ability to be trained to generate such a mapping may be referred to as the "training ability of the neural network." In addition, due to specialized training, such a dedicated and trained neural network may have a generalization ability to generate a relatively accurate output for an input pattern for which it has not been trained. To process operations related to training and inference of a neural network model, methods for converging to a result more quickly may include, for example, model parallelization and / or data parallelization. Summary of the Invention

[0004] This Summary of the Invention is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.

[0005] In one general aspect, a method of identifying an image includes: obtaining image data to be identified as input data of a neural network; performing operations related to each layer of the neural network based on the image data to be identified to obtain a result of image recognition; and outputting the result of image recognition, wherein for each of target layers in the neural network: extracting metadata of the target layer; measuring a similarity between the target layer and each reference layer by comparing the metadata of the target layer with reference metadata of each reference layer; selecting a corresponding layer among the reference layers based on the similarity; and generating a parallelization strategy for the target layer based on a reference parallelization strategy matching the corresponding layer, and performing parallel processing of operations related to the target layer based on the parallelization strategy using input data of the target layer.

[0006] The target layer is one or more layers selected from a plurality of layers of the neural network.

[0007] The input data of the first layer of the neural network is the image data to be recognized, and the input data of the layers other than the first layer of the neural network is the output data of the previous layer.

[0008] In one general aspect, an apparatus for recognizing an image, the apparatus comprising: a processor; and a memory including instructions executable by the processor, wherein, in response to the instructions being executed by the processor, the processor is configured to: obtain image data to be recognized as input data of a neural network; perform operations related to each layer of the neural network based on the image data to be recognized to obtain an image recognition result; and output the image recognition result, wherein, for each of the target layers in the neural network: extract metadata of the target layer; measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; select a corresponding layer among the reference layers based on the similarity; and generate a parallelization strategy for the target layer based on the reference parallelization strategy matching the corresponding layer, and perform parallel processing of operations related to the target layer using the input data of the target layer based on the parallelization strategy.

[0009] In one general aspect, a method for training a neural network for recognizing an image, the method comprising: acquiring training image data; training the neural network based on the training image data, wherein, during the training, for each of the target layers in the neural network: extract metadata of the target layer; measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; select a corresponding layer among the reference layers based on the similarity; and generate a parallelization strategy for the target layer based on the reference parallelization strategy matching the corresponding layer, and perform parallel processing of operations related to the target layer using the input data of the target layer based on the parallelization strategy.

[0010] The target layer is one or more layers selected from multiple layers of the neural network.

[0011] The input data of the first layer of the neural network is the training image data, and the input data of the layers other than the first layer of the neural network is the output data of the previous layer.

[0012] In one general aspect, an apparatus for training a neural network for image recognition, the apparatus comprising: a processor; and a memory including instructions executable by the processor, wherein, in response to the instructions being executed by the processor, the processor is configured to: obtain training image data; train the neural network based on the training image data, wherein, in the training, for each of the target layers in the neural network: extract metadata of the target layer; measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; select a corresponding layer among the reference layers based on the similarity; and generate a parallelization strategy for the target layer based on a reference parallelization strategy matching the corresponding layer, and perform parallel processing of operations related to the target layer based on the parallelization strategy using the input data of the target layer.

[0013] In one general aspect, a parallel processing method for a target model based on a neural network includes: extracting metadata of a target layer included in the target model; measuring the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; selecting a corresponding layer among the reference layers based on the similarity; and generating a parallelization strategy for the target layer based on a reference parallelization strategy matching the corresponding layer.

[0014] The step of selecting the corresponding layer may include: in response to the presence of a first reference layer having the same reference metadata as the metadata of the target layer among the reference layers, selecting the first reference layer as the corresponding layer; and in response to the absence of the first reference layer among the reference layers, selecting a second reference layer having the reference metadata most similar to the metadata of the target layer among the reference layers as the corresponding layer.

[0015] In response to the second reference layer being selected as the corresponding layer, reference layer information corresponding to the metadata of the target layer may be added to a reference database (DB) in which the reference metadata of each reference layer is stored. The reference layer information may include link information. In response to the second reference layer being selected as the corresponding layer, the identification information of the second reference layer may be recorded as the link information in the reference layer information.

[0016] The parallel processing method may further include: in response to the reference layer information corresponding to the metadata of the target layer being added to the reference DB, generating a new parallelization strategy corresponding to the metadata of the target layer. Generation of the new parallelization strategy may be performed independently of the execution of the parallelization strategy of the target layer.

[0017] The generation of a new parallelization strategy can be performed in response to the amount of new reference layer information exceeding a threshold. The new reference layer information includes reference layer information corresponding to the metadata of the target layer and is added to the reference DB. A neural network-based similarity measurement model can be used to measure similarity. In response to the generation of the new parallelization strategy, the similarity measurement model can be retrained based on the new parallelization strategy.

[0018] In response to the generation of the new parallelization strategy, the link information of the reference layer information can be changed to an empty state. The parallel processing method may further include: performing parallel processing of operations related to the target layer based on the parallelization strategy.

[0019] In another general aspect, a parallel processing device for a neural network-based target model includes: a processor; and a memory including instructions executable by the processor, wherein, in response to the instructions being executed by the processor, the processor is configured to: extract the metadata of the target layer included in the target model; measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; select a corresponding layer among the reference layers based on the similarity; and generate a parallelization strategy for the target layer based on the reference parallelization strategy matching the corresponding layer.

[0020] In another general aspect, an electronic device includes: a processor; and a memory including instructions executable by the processor, wherein, in response to the instructions being executed by the processor, the processor is configured to: extract the metadata of the target layer included in the neural network-based target model; measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; select a corresponding layer among the reference layers based on the similarity; and generate a parallelization strategy for the target layer based on the reference parallelization strategy matching the corresponding layer.

[0021] In another general aspect, a processor-implemented method includes: extracting the metadata of the target layer included in the neural network-based target model; comparing the metadata of the target layer with the metadata of each of a plurality of reference layers; in the case where the metadata of the target layer matches the metadata of a specific reference layer among the plurality of reference layers, executing the target layer based on a first parallelization strategy corresponding to the specific reference layer; and in the case where the metadata of the target layer does not match the metadata of any of the plurality of reference layers, executing the target layer based on a second parallelization strategy corresponding to the reference layer having the closest similarity to the target layer among the plurality of reference layers, generating a new parallelization strategy for the target layer, and subsequently executing the target layer based on the new parallelization strategy.

[0022] Executing the target layer based on the second parallelization strategy can be performed independently of the generation of the new parallelization strategy for the target layer.

[0023] The steps of generating a new parallelization strategy for a target layer may include: determining whether an update condition has been met; and, after the update condition has been met, generating a new parallelization strategy for the target layer based on new reference layer information that has been added to the plurality of reference layers before the update condition was met.

[0024] The update condition may correspond to one or both of an amount of time that has passed and an amount of new reference layer information that has been added.

[0025] Other features and aspects will be apparent from the following detailed description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematically shows an example of a parallel processing process for a neural network model.

[0027] Figure 2 and Figure 3 Shows an example of parallel processing operations for each component of a parallel processing device.

[0028] Figure 4 Shows an example of adopting a parallelization strategy based on whether there is a reference layer having the same metadata as the metadata of the target layer.

[0029] Figure 5 Shows an example of operations related to the generation of a new parallelization strategy.

[0030] Figure 6 Shows examples of a reference database (DB) before and after an update.

[0031] Figure 7 Shows an example of an overall parallel processing process.

[0032] Figure 8 Shows an example of a parallel processing device.

[0033] Figure 9 Shows an example of an electronic device.

[0034] Throughout the drawings and the detailed description, unless otherwise described or provided, the same reference numerals will be understood to refer to the same elements, features, and structures. The drawings may not be to scale, and for clarity, illustration, and convenience, the relative sizes, proportions, and depictions of elements in the drawings may be exaggerated. DETAILED DESCRIPTION

[0035] The following specific embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, after understanding the disclosure of the present application, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0036] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein have been provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein that will be apparent after understanding the disclosure of the present application.

[0037] The following specific structural or functional descriptions are exemplary and are provided only to describe examples, and the scope of the examples is not limited to the descriptions provided in this specification. Those of ordinary skill in the art may make various changes and modifications thereto.

[0038] Although the terms "first" or "second" are used to explain various components, the components are not limited by the terms. These terms should only be used to distinguish one component from another. For example, within the scope of the rights according to the concept of the present disclosure, a "first" component may be referred to as a "second" component, or similarly, a "second" component may be referred to as a "first" component.

[0039] As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. It should also be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, components, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0040] Unless otherwise defined herein, all terms (including technical or scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. Unless otherwise defined herein, terms that are commonly defined in a dictionary should be construed as having a meaning that matches the context in the relevant art and will not be construed as having an idealized or overly formal meaning.

[0041] Hereinafter, examples will be described in detail with reference to the accompanying drawings, and the same reference numerals in the drawings always denote the same elements.

[0042] Figure 1An example of a parallel processing process for a neural network model is schematically shown. Referring to Figure 1 , the parallel processing device 100 can generate a parallelization strategy for the target model 110 based on the target model information. The parallel processing device 100 can perform parallel processing (e.g., training and / or inference) for the target model 110 based on the generated parallelization strategy. For example, the target model 110 can be based on an artificial neural network (hereinafter referred to as "neural network"), and operations for training and / or inference of the target model 110 can be performed in parallel based on the parallelization strategy.

[0043] By mapping input data and output data in a non-linear relationship through deep learning, the neural network can be trained to perform operations (e.g., object recognition operations or user authentication operations) according to the purpose of training. For example, training image data can be acquired, and the neural network can be trained based on the training image data. For example, a device for recognizing an image can obtain image data to be recognized as input data of the neural network; perform operations related to each layer of the neural network based on the image data to be recognized to obtain an image recognition result and output the image recognition result. The input data of the first layer of the neural network can be the image data to be recognized, and the input data of the layers other than the first layer of the neural network can be the output data of the previous layer. Deep learning can be a machine learning scheme for solving problems (such as image recognition or speech recognition) based on a large dataset. Deep learning can be understood as a process of solving an optimization problem to find a point of minimum energy while training the neural network based on the prepared training data.

[0044] Through supervised learning or unsupervised learning of deep learning, the structure of the neural network or the weights corresponding to the model can be obtained, and the input data and output data can be mapped to each other through the weights. For example, when the width and depth of the neural network are large enough, the neural network can have a large enough capacity to implement any function. When the neural network is trained on a sufficiently large amount of training data through an appropriate training process, optimal performance can be achieved.

[0045] In the following description, the neural network or network parameters (e.g., weights) can be expressed as "pre-trained", where "pre" can indicate the state before the neural network is "activated". An "activated" neural network can indicate that the neural network is ready for inference. For example, the "activation" of the neural network can include loading the neural network into the memory, or inputting input data for inference into the neural network after loading the neural network into the memory.

[0046] A neural network may include multiple layers. In this example, the neural network may be referred to as a deep neural network (DNN). The multiple layers may include an input layer, at least one hidden layer, and an output layer. The target model 110 may include multiple elements (e.g., element 111). For example, the multiple elements of the target model 110 may correspond to components of the neural network (e.g., multiple layers or multiple nodes). In the following description, each element of the target model 110 corresponds to a layer of the neural network. However, the example is not limited thereto. For example, the following description may apply to an example where each element of the target model 110 corresponds to another component of the neural network rather than a layer.

[0047] A neural network may include various types of networks, e.g., a fully connected network, a convolutional neural network (CNN), or a recurrent neural network (RNN). For example, at least part of the multiple layers in the neural network may correspond to a CNN, and another part may correspond to a fully connected network. In addition, the neural network may include multiple layers based on each type. For example, the multiple layers of the neural network may include at least one of various types of layers (e.g., a fully connected layer, a convolutional layer, or a recurrent layer).

[0048] Strategies such as model parallelization and / or data parallelization can be used as a scheme to converge to results more quickly to handle operations related to the training and / or inference of a neural network model. Parallel processing and parallelization strategies can be understood to include the concepts of distributed processing and distribution strategies. Model parallelization is a scheme of dividing a neural network model and performing calculations in different accelerators, and data parallelization is a scheme of dividing the data given as input and performing calculations in different accelerators. Model parallelization can be broadly divided into a scheme using pipelining and inter-layer parallelism and an intra-layer parallelism scheme.

[0049] Inter-layer parallelism is a scheme of assigning layers in a neural network model to different accelerators and performing calculations, and intra-layer parallelism is a scheme of assigning internal values of a layer of a neural network model to different accelerators and performing calculations. The internal values of a layer can be kernel weights and input feature maps in the case of a CNN, or can be weights of units forming a layer of an RNN or other information for calculation in the case of an RNN.

[0050] There is also a parallelization scheme between nodes. Nodes can represent servers or endpoint devices connected through a network. Distributed processing between nodes can be performed to distribute the weights of a neural network model and perform training and inference. For this purpose, a parameter sharing scheme or a message passing interface (MPI) can be mainly used. In addition, a hybrid scheme using the parameter sharing scheme and MPI can be used.

[0051] Strategies for the above parallelization schemes may include asynchronous and synchronous schemes. Synchronous training is a scheme in which the neural network model is updated when the work of the workers to learn the weights is completed and all data is collected. Asynchronous training is a scheme in which, when the computation terminates, each worker immediately updates the neural network model based on the values obtained through the computation of each worker, regardless of the actions of other workers.

[0052] In addition, the strategies may be applied differently according to whether the characteristics of the weights are dense or sparse. Dense weights indicate that few zeros exist in the weight matrix, and sparse weights indicate that multiple zeros or consecutive zeros exist in the weight matrix.

[0053] Such various parallelization schemes and strategies can be used as elements of the parallelization strategy. For example, the parallelization strategy elements may include: parallelization schemes (such as intra-layer parallel or inter-layer parallel); partition dimensions indicating the direction of partitioning the model, layer, or data (e.g., width direction, height direction, channel direction, or batch direction); partition numbers indicating the number of models to be partitioned, the number of layers to be partitioned, or the number of data to be partitioned; information on processing devices (such as processors for performing parallel processing (e.g., neural network processors (NPUs), graphics processors (GPUs), NPU#1, or NPU#2) or cores (e.g., NPU core#1 or NPU core#5)); and information on parallelization algorithms (such as asynchronous algorithms, synchronous algorithms, or all-reduce algorithms). In addition to the above elements, the parallelization strategy elements may include various methods and strategies related to parallelization known through various papers and academic research.

[0054] The parallelization strategy of the target model 110 may include various strategy elements to perform parallel processing of various operations related to the training and / or inference of the target model 110. The parallel processing device 100 may generate the above parallelization strategy for each element of the target model 110 and may execute each element of the target model 110 based on the parallelization strategy for the training and / or inference of the target model 110.

[0055] For example, the parallel processing device 100 may extract the metadata 112 of the element 111, compare the metadata 112 with each reference information in the reference database (DB) 120, select the reference information 121 corresponding to the element 111 among the multiple reference information in the reference DB 120, and generate a parallelization strategy for the element 111 based on the reference information 121. In one example, the element 111 may also be referred to as a target element or a target layer. In other words, the target layer may be one or more layers selected from multiple layers of a neural network. The reference DB 120 may include multiple reference information. Each reference information may correspond to a unit corresponding to the element 111. For example, when each element of the target model 110 corresponds to a layer of a neural network, each reference information in the reference DB 120 may also correspond to a layer of the neural network. In this example, each reference information may be referred to as "reference layer information".

[0056] Each reference information in the reference DB 120 may include reference metadata or a reference parallelization strategy. The parallel processing device 100 may compare the metadata 112 with the reference metadata of each reference information and select the reference information 121. Among all the reference information in the reference DB 120, the reference information 121 may have the reference metadata most similar to the metadata 112. For example, the reference metadata of the reference information 121 may be the same as the metadata 112, or although it is not the same as the metadata 112, the reference information 121 may be the most similar to the metadata 112 compared with the reference metadata of other reference information.

[0057] The parallel processing device 100 may generate a parallelization strategy for the element 111 based on the reference parallelization strategy of the reference information 121. The reference parallelization strategy may include application information related to various parallelization strategy elements. For example, the parallel processing device 100 may select the reference parallelization strategy of the reference information 121 as the parallelization strategy for the element 111. In this example, the parallel processing device 100 may perform parallel processing (e.g., training and / or inference) for the element 111 based on the generated parallelization strategy. The parallel processing device 100 may repeat the above process for each element of the target model 110 to generate a parallelization strategy for each element and perform parallel processing for each element, so that the neural network can converge to the result more quickly, thereby improving the processing speed of the neural network for recognizing images or the training speed of the neural network.

[0058] To optimize the parallel processing of existing neural network models, it is necessary to use a simulator and model profiling offline (e.g., when the neural network model is not being executed). This is because, although applications for training or inferring neural network models need to guarantee high training rates and fast response times (latency) to users, the training rate and response time may decrease when simulation and profiling are performed online (e.g., when the neural network model is being executed).

[0059] According to an example, the parallel processing device 100 can use the reference DB 120 to minimize the training rate and response time, so that the parallelization strategy can be generated at runtime. For example, the parallel processing device 100 can be applied to systems (e.g., autonomous vehicles configured to learn or infer data in real time, and cloud computing and data centers that require parallel processing of a large amount of data). Recently, with the development of artificial intelligence technology, neural network model applications are becoming large and diverse, and it takes a lot of time and effort to establish a parallelization strategy for neural network model applications. According to an example, the parallel processing device 100 can generate a new parallelization strategy through continuous updates, so as to cope with the above large and diverse applications.

[0060] Figure 2 and Figure 3 Shows an example of the parallel processing operations of each component of the parallel processing device. Figure 2 Shows the operations of the runtime engine 220, the reference DB 230, the policy manager 240, and the similarity measurer 250. The runtime engine 220, the reference DB 230, the policy manager 240, and the similarity measurer 250 can be implemented as at least one hardware module, at least one software module, and / or a combination thereof. The operations related to parallel processing will be described below from the perspectives of the runtime engine 220, the reference DB 230, the policy manager 240, and the similarity measurer 250, but the operations related to parallel processing do not necessarily need to be executed by separate components (such as the runtime engine 220, the reference DB 230, the policy manager 240, and the similarity measurer 250). For example, the operations described as being executed by one component can be executed by another component, or the above operations can be executed by a single integrated component (e.g., the parallel processing device).

[0061] The runtime engine 220 may execute the target model 210 given as input. For example, the runtime engine 220 may execute the target model 210 to train and / or infer the target model 210. The runtime engine 220 may execute the target model 210 based on the parallelization strategy of each target layer of the target model 210. For example, the first reference parallelization strategy of the first reference layer information 231 may be adopted as the parallelization strategy of the first target layer 211, and the third reference parallelization strategy of the third reference layer information 233 may be adopted as the parallelization strategy of the second target layer 212. In this example, the runtime engine 220 may execute the first target layer 211 based on the first reference parallelization strategy, and may execute the second target layer 212 based on the third reference parallelization strategy. In addition, the runtime engine 220 may output the execution time of the target model 210 or each target layer. The execution time may be used to evaluate the performance of each parallelization strategy.

[0062] The reference DB 230 may include reference layer information related to each of the various reference layers. The reference layer information of each reference layer may include reference metadata and a reference parallelization strategy related to each reference layer. For example, the first reference layer information 231 may include first reference metadata corresponding to the first reference layer and a first reference parallelization strategy. The reference metadata may include the metadata of the reference layer, and the reference parallelization strategy may include the parallelization strategy applied to the reference layer and the performance of the parallelization strategy (e.g., execution time).

[0063] In one example, the parallelization strategy of the layers of a predetermined neural network may have been generated in the past. In this example, extensive simulation or profiling of the layers may be performed offline, and the optimal parallelization strategy may have been established. The layer may be defined as the first reference layer, the metadata of the layer may be defined as the first reference metadata, the parallelization strategy of the layer may be defined as the first parallelization strategy, and the first reference layer, the first reference metadata, and the first parallelization strategy may be stored in the reference DB 230.

[0064] As described above, the reference DB 230 may store data related to the optimal parallelization strategy executed offline as the initial data of the reference layer information. The initial data may be used to establish a new parallelization strategy in the process of updating the parallelization strategy. For example, when there is a reference layer among the reference layers that has the same reference metadata as the metadata of the first target layer 211, the reference layer that has the same reference metadata as the metadata of the first target layer 211 may be selected as the layer corresponding to the first target layer 211. For example, when there is no reference layer among the reference layers that has the same reference metadata as the metadata of the first target layer 211, the reference layer that has the reference metadata most similar to the metadata of the first target layer 211 may be selected as the layer corresponding to the first target layer 211.

[0065] Since the metadata of the selected reference layer is different from the metadata of the first target layer 211, the parallelization strategy of the selected reference layer may not be considered optimized for the first target layer 211. Therefore, even if the parallelization strategy of the selected reference layer is applied to the first target layer 211, the policy manager 240 may generate a new parallelization strategy for the first target layer 211. For example, the policy manager 240 may generate the optimal parallelization strategy for the first target layer 211 by performing simulations or profiling related to the first target layer 211.

[0066] The similarity measurer 250 may compare the metadata of each target layer of the target model 210 with the reference metadata of each reference layer of the reference DB 230, and may measure the similarity between each target layer and each reference layer. For example, the similarity measurer 250 may extract the metadata of the first target layer 211, may compare the metadata of the first target layer 211 with each of the first reference metadata of the first reference layer information 231, the second reference metadata of the second reference layer information 232, and the third reference metadata of the third reference layer information 233, and may measure the similarity between the first target layer 211 and each of the first reference layer, the second reference layer, and the third reference layer.

[0067] For example, the metadata may include input data, output data, characteristics of weights (e.g., size or sparsity), and the type of layer (e.g., fully connected layer, convolutional layer, or recurrent layer). In a CNN, the metadata of a layer may include kernel size, padding (pad), and stride. In an RNN, the metadata of a layer may include cell information, gate information, and input embedding. The reference metadata may also include information related to the items described above. The similarity measurer 250 may measure the similarity between a target layer and a reference layer by comparing corresponding items between the metadata of the target layer and the reference metadata of the reference layer.

[0068] The policy manager 240 may select, among the reference layers, the layer corresponding to the target layer (hereinafter referred to as the "corresponding layer") based on the similarity measured by the similarity measurer 250, and may generate a parallelization strategy for the target layer based on the reference parallelization strategy matching the corresponding layer. The policy manager 240 may select the reference layer with the highest similarity as the corresponding layer. For example, when the first reference layer among the reference layers has the highest similarity to the first target layer 211, the policy manager 240 may select the first reference layer as the corresponding layer of the first target layer 211, and may generate a parallelization strategy for the first target layer 211 based on the first reference parallelization strategy.

[0069] Figure 3Shows the parallel processing operations of the reference DB 320, the similarity measurer 330, and the policy manager 340 related to the target model 310. The reference DB 320, the similarity measurer 330, and the policy manager 340 can be implemented as at least one hardware module, at least one software module, and / or a combination thereof. Additionally, as referred to above Figure 2 As described, the operations to be described below are not necessarily performed separately by the reference DB 320, the similarity measurer 330, and the policy manager 340. For example, an operation described as being performed by one component can be performed by another component, or the above operations can be performed by a single integrated component (e.g., a parallel processing device).

[0070] The similarity measurer 330 can extract the metadata 312 of the target layer 311 included in the target model 310. The metadata 312 can include the type of the layer (e.g., a fully connected layer, a convolutional layer, or a recurrent layer) or the characteristics of the layer (e.g., the characteristics of the input data, output data, and weights). The similarity measurer 330 can compare the metadata 312 with the reference metadata of each reference layer and can measure the similarity between the target layer 311 and each reference layer. The measurement result 332 can include information about the similarity between the target layer 311 and each reference layer.

[0071] The similarity measurer 330 can include a neural network-based similarity measurement model 331. The similarity measurement model 331 can be pre-trained based on the reference DB 320. For example, the similarity measurement model 331 can be pre-trained to compare the metadata 312 with the reference metadata of each reference layer in the reference DB 320 in response to the input of the metadata 312 and output the similarity between the target layer 311 and each reference layer.

[0072] For example, when the metadata 312 is very similar to the first reference metadata of the first reference layer and is the same as the second reference metadata of the second reference layer, in the measurement result 332, the similarity between the target layer 311 and the first reference layer and the similarity between the target layer 311 and the second reference layer can be "0.95" and "1.0" respectively. As described above, in response to generating a new parallelization policy, new reference layer information can be added to the reference DB 320. In this example, the similarity measurement model 331 can be updated based on a predetermined condition (e.g., the amount of the added new reference layer information exceeds a threshold).

[0073] The policy manager 340 can generate a parallelization policy for the target layer 311 based on any one or any combination of the metadata 312, the reference layer information 321, and the measurement result 332. For example, when the first reference layer among the reference layers has the highest similarity to the target layer 311, the policy manager 340 can select the first reference layer as the corresponding layer of the target layer 311, and can generate a parallelization policy for the target layer 311 based on the first reference parallelization policy.

[0074] Figure 4 An example of adopting a parallelization policy based on whether there is a reference layer having the same metadata as the metadata of the target layer is shown. Refer to Figure 4 In operation 410, the parallel processing device obtains a similarity measurement result, reference layer information, and metadata. For example, the parallel processing device can obtain the similarity measurement result from a similarity measurer, can obtain the reference layer information from a reference DB, and can obtain the metadata from the target layer.

[0075] In operation 420, the parallel processing device determines whether there is a reference layer having the same metadata as the metadata of the target layer (hereinafter, simply referred to as "same reference layer"). For example, when determining the parallelization policy of the target layer, different processing procedures can be performed based on whether the same reference layer exists in the reference DB. When there is a same reference layer, in operation 430, the reference parallelization policy of the same reference layer can be adopted, and the parallelization policy of the target layer can be determined based on the reference parallelization policy of the same reference layer. When there is no same reference layer, in operation 440, the reference layer most similar to the target layer (hereinafter, simply referred to as "most similar reference layer") can be adopted instead of the reference parallelization policy of the same reference layer, and the parallelization policy of the target layer can be determined based on the reference parallelization policy of the most similar reference layer.

[0076] The absence of a same reference layer may indicate that the selected reference parallelization policy may not be optimal for the target layer. This is because, since even though the best parallelization policy for each reference layer exists in the reference DB, there is no reference layer that matches the target layer, so even with the most similar reference layer, it cannot be guaranteed that the reference parallelization policy of the most similar reference layer is optimal for the target layer. Therefore, the task of finding the optimal parallelization policy for the target layer can be additionally performed. However, since the task of finding a new parallelization policy requires a relatively long period of time, the parallelization policy of the most similar reference layer can be applied to the target layer, and the task of finding a new parallelization policy can be performed in the future.

[0077] When there is no identical reference layer, the parallel processing device can add reference layer information corresponding to the metadata of the target layer to the reference DB. This is because the target layer corresponds to a layer with a new type of metadata, so the new policy of the target layer needs to be included in the reference DB. The reference parallelization policy of the most similar reference layer can be stored as the reference parallelization policy of the new reference layer information. The reference parallelization policy of the new reference layer information can be updated later through an optimization process.

[0078] In one example, each reference layer information in the reference DB can include link information. When the reference parallelization policy of the reference layer information is optimized for the reference layer with the reference layer information, the link information of the reference layer information can be represented as empty. For example, the link information of the reference layer information corresponding to the initial data can be represented as empty. For example, when the reference parallelization policy of the reference layer information is not optimized for the reference layer with the reference layer information, the link information of the reference layer information can be shown as the identification information of the most similar reference layer.

[0079] For example, when the target layer does not have an identical reference layer in the reference DB as described above, the new reference layer information of the target layer can be added to the reference DB. In this example, when adding the new reference layer information to the reference DB, the identification information of the most similar reference layer to the target layer can be shown as the link information of the new reference layer information. When generating the new optimal parallelization policy of the target layer, the reference parallelization policy of the new reference layer information can be updated to the new parallelization policy, and the link information of the new reference layer information can be changed to an empty state.

[0080] Figure 5 An example of an operation related to the generation of a new parallelization policy is shown. Refer to Figure 5 , in operation 510, the parallel processing device determines whether the update condition is satisfied. For example, the update condition can be set based on the amount of new reference layer information added to the reference DB (e.g., ten items) or the time period elapsed after the previous update (e.g., one day or one week). In this example, a numerical value (e.g., ten items, one day or one week) can be set as the threshold. For example, the update condition can include adding ten items of new reference layer information to the reference DB. In this example, when ten items of new reference layer information are added, operations 520, 530, and 540 can be executed.

[0081] In operation 520, the parallel processing device generates a new parallelization strategy for each new reference layer. For example, the parallel processing device can generate an optimal parallelization strategy for the new reference layer by performing profiling or simulation related to the new reference layer. In this example, compared with existing offline profiling and simulation processing, the unit of batch processing can be reduced. Since the profiling of batch processing of small units instead of the offline profiling and simulation of the entire model, relatively less time may be required compared with existing offline profiling and simulation.

[0082] When generating a new parallelization strategy, the parallel processing device can update the reference DB in operation 530 and can retrain the similarity measurer (e.g., similarity measurement module) in operation 540. For example, the parallel processing device can record the new parallelization strategy in the new reference layer information of the reference DB instead of the parallelization strategy of the most similar reference layer recorded as the reference parallelization strategy. In addition, the parallel processing device can retrain the similarity measurer based on the updated reference DB. As described above, the similarity measurement module can be trained based on the reference DB. Since the reference DB is updated in response to the generation of the new parallelization strategy, the similarity measurement module can be retrained based on the updated reference DB. When operation 540 is completed, operation 510 can be re-executed.

[0083] Regarding the generation of the new parallelization strategy Figure 5 Operations 510 to 540 can be performed independently of Figure 4 Operations 410 to 440 related to the parallel processing of the target layer. For example, when the parallel processing of the target layer is being performed, the operations related to the generation of the new parallelization strategy can be performed in the background. In addition, before the update of the reference DB or the retraining of the similarity measurer is completed, a parallelization strategy can be selected from the existing reference DB. Through the independence of the above series of operations, the speed of performing parallel processing at runtime can be ensured. In addition, the existing data can be used until the update is completely terminated, so the stability of parallel processing can be maintained.

[0084] Figure 6 An example showing the reference DB before and after the update is shown. Referring to Figure 6 , the reference DB 610 stores the first reference layer information 611 to the sixth reference layer information 616 before the update, and the reference DB 630 stores the first reference layer information 631 to the sixth reference layer information 636 after the update.

[0085] The first reference layer information 611 to the sixth reference layer information 616 and the first reference layer information 631 to the sixth reference layer information 636 may include the identification information ID#1 to ID#6 of the reference layers, the reference metadata RMD#1 to RMD#6, the reference parallelization strategies RDP#1 to RDP#6, and the link information "empty", ID#2, and ID#4. For example, the first reference layer information 611 and 631, the second reference layer information 612 and 632, and the third reference layer information 613 and 633 may correspond to the initial data. As described above, the initial data may include a parallelization strategy established offline by processing (e.g., simulation or profiling). Therefore, the link information of each of the first reference layer information 611 and 631, the second reference layer information 612 and 632, and the third reference layer information 613 and 633 may be represented as "empty".

[0086] In addition, since there is no identical reference layer, the fourth reference layer information 614 and the sixth reference layer information 616 may be newly added to the reference DB 610. In one example, the second reference layer may be the layer most similar to the fourth reference layer. Therefore, the identification information ID#2 of the second reference layer may be indicated as the link information of the fourth reference layer information 614, and the reference parallelization strategy RDP#2 of the second reference layer may be indicated as the reference parallelization strategy of the fourth reference layer information 614. Similarly, the identification information ID#4 of the fourth reference layer may be indicated as the link information of the sixth reference layer information 616, and the reference parallelization strategy RDP#2 of the second reference layer may be indicated as the reference parallelization strategy of the sixth reference layer information 616. In other words, it can be confirmed that the sixth reference layer information 616 is added to the reference DB 610 after the fourth reference layer information 614 is added to the reference DB 610 and before the reference DB 610 is updated.

[0087] As described above, the parallel processing device may update the reference DB 610 in response to the update condition being satisfied. For example, the simulator 621 of the policy manager 620 may generate a new parallelization strategy RDP#4 for the fourth reference layer and a new parallelization strategy RDP#6 for the sixth reference layer, and the parallel processing device may update the fourth reference layer information 614 and the sixth reference layer information 616 based on the new parallelization strategies RDP#4 and RDP#6. In addition, the parallel processing device may change the link information of the fourth reference layer information 614 and the sixth reference layer information 616 to an empty state. The fourth reference layer information 634 and the sixth reference layer information 636 may respectively indicate the states in which the fourth reference layer information 614 and the sixth reference layer information 616 are finally updated. The reference DB 610 also stores the fifth reference layer information 615 before the update, and the reference DB 630 stores the fifth reference layer information 635 after the update.

[0088] Figure 7 An example of the overall parallel processing process is shown. Refer toFigure 7 , in operation 710, the parallel processing device extracts metadata of a target layer included in a target model, in operation 720 measures the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer, in operation 730 selects a corresponding layer among the reference layers based on the similarity, and in operation 740 generates a parallelization strategy for the target layer based on a reference parallelization strategy that matches the corresponding layer. The above Figures 1 to 6 description also applies to the parallel processing process.

[0089] Figure 8 shows an example of a parallel processing device 800. Refer to Figure 8 , the parallel processing device 800 includes a processor 810 and a memory 820. The memory 820 can be connected to the processor 810 and can store instructions executable by the processor 810, data to be computed by the processor 810, or data processed by the processor 810. The memory 820 can include, for example, a non-transitory computer-readable storage medium (e.g., high-speed random access memory (RAM)) and / or a non-volatile computer-readable storage medium (e.g., at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state memory device).

[0090] The processor 810 can execute instructions to perform at least one of the operations described above with reference to Figures 1 to 7 . For example, the processor 810 can extract metadata of a target layer included in a target model, can compare the metadata of the target layer with the reference metadata in each reference layer, can measure the similarity between the target layer and each reference layer, can select a corresponding layer among the reference layers based on the similarity, and can generate a parallelization strategy for the target layer based on a reference parallelization strategy that matches the corresponding layer.

[0091] Figure 9 shows an example of an electronic device 900. The electronic device 900 can structurally and / or functionally include Figure 1 the parallel processing device 100 and / or Figure 8 the parallel processing device 800.

[0092] Refer to Figure 9, the electronic device 900 includes a processor 910, a memory 920, a camera 930, a storage device 940, an input device 950, an output device 960, and a network interface 970. The processor 910, the memory 920, the camera 930, the storage device 940, the input device 950, the output device 960, and the network interface 970 can communicate with each other via a communication bus 980. For example, the electronic device 900 can be implemented as at least a part of, for example, a mobile device (such as a mobile phone, a smartphone, a personal digital assistant (PDA), a netbook, a tablet computer, or a laptop computer), a wearable device (such as a smartwatch, a smart bracelet, or smart glasses), a computing device (such as a desktop or a server), a household appliance (such as a television (TV), a smart TV, or a refrigerator), a security device (such as a door lock), or a vehicle (such as a smart vehicle).

[0093] The processor 910 can execute instructions and functions in the electronic device 900. For example, the processor 910 can process instructions stored in the memory 920 or the storage device 940. The processor 910 can execute at least one of the operations described above with reference to Figures 1 to 8 one of the operations.

[0094] The memory 920 can include a non-transitory computer-readable storage medium or a non-transitory computer-readable storage device. The memory 920 can store instructions to be executed by the processor 910 and also store information related to software and / or applications when the software and / or applications are being executed by the electronic device 900.

[0095] The camera 930 can capture photos and / or videos. For example, the camera 930 can capture a face image including a user's face, an eye image including a user's eyes, or an iris image including a user's iris. In one example, the camera 930 can provide a three-dimensional (3D) image including depth information related to an object.

[0096] The storage device 940 can include a non-transitory computer-readable storage medium or a non-transitory computer-readable storage device. The storage device 940 can store various data and modules used in parallel processing (such as a runtime engine, a reference DB, a policy manager, and a similarity measurer). In one example, the storage device 940 can store a larger amount of information than the amount of information in the memory 920 over a relatively long period of time. For example, the storage device 940 can include a magnetic hard disk, an optical disk, a flash memory, a floppy disk, or other forms of non-volatile memory known in the art.

[0097] The input device 950 can receive input from a user through a conventional input scheme using a keyboard and a mouse and through new input schemes such as touch input, voice input, and image input. For example, the input device 950 can detect input from a keyboard, a mouse, a touch screen, a microphone, or the user, and can include any other device configured to transmit the detected input to the electronic device 900.

[0098] The output device 960 can provide an output of the electronic device 900 to the user through a visual channel, an auditory channel, or a tactile channel. The output device 960 can include, for example, a display, a touch screen, a speaker, a vibration generator, or any other device configured to provide an output to the user. The network interface 970 can communicate with an external device via a wired network or a wireless network.

[0099] The devices, units, modules, apparatuses, and other components described herein are implemented by hardware components. Examples of hardware components that can be used to perform the operations described in this application include, where appropriate: controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more of the hardware components that perform the operations described in this application are implemented by computing hardware (e.g., by one or more processors or computers). A processor or computer can be implemented by one or more processing elements (such as, logic gate arrays, controllers, and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond and execute instructions in a defined manner to achieve a desired result). In one example, a processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. The hardware components implemented by the processor or computer can execute instructions or software (such as an operating system (OS) and one or more software applications running on the OS) to perform the operations described in this application. The hardware components can also access, manipulate, process, create, and store data in response to the execution of the instructions or software. For simplicity, the singular terms "processor" or "computer" may be used in the description of the examples described in this application, but in other examples, multiple processors or computers may be used, or a processor or computer may include multiple processing elements or multiple types of processing elements or both. For example, a single hardware component or two or more hardware components can be implemented by a single processor, or two or more processors, or a processor and a controller. One or more hardware components can be implemented by one or more processors, or a processor and a controller, and one or more other hardware components can be implemented by one or more other processors, or an additional processor and an additional controller. One or more processors, or a processor and a controller, can implement a single hardware component or two or more hardware components. The hardware components can have any one or more of different processing configurations, examples of different processing configurations include: single processor, independent processor, parallel processor, single instruction single data (SISD) multiprocessing, single instruction multiple data (SIMD) multiprocessing, multiple instruction single data (MISD) multiprocessing, and multiple instruction multiple data (MIMD) multiprocessing.

[0100] The method of performing the operations described in this application is performed by computing hardware (e.g., by one or more processors or computers), which is implemented as described above to execute instructions or software to perform the operations performed by the method described in this application. For example, a single operation or two or more operations may be performed by a single processor, or two or more processors, or a processor and a controller. One or more operations may be performed by one or more processors, or a processor and a controller, and one or more other operations may be performed by one or more other processors, or additional processors and additional controllers. One or more processors, or a processor and a controller, may perform a single operation or two or more operations.

[0101] Instructions or software for controlling a processor or computer to implement the hardware components and perform the method described above are written as a computer program, code segment, instruction, or any combination thereof to individually or jointly direct or configure the processor or computer to operate as a machine or a special-purpose computer to perform the operations performed by the hardware components and method described above. In one example, the instructions or software include machine code (such as machine code generated by a compiler) that is directly executed by the processor or computer. In another example, the instructions or software include high-level code that is executed by the processor or computer using an interpreter. A programmer of ordinary skill in the art can easily write the instructions or software based on the block diagrams and flowcharts shown in the drawings and the corresponding descriptions in the specification, which disclose algorithms for performing the operations performed by the hardware components and method described above.

[0102] Instructions or software for controlling a processor or computer to implement the hardware components and execute the methods as described above, as well as any associated data, data files, and data structures, are recorded, stored, or fixed in one or more non-transitory computer-readable storage media or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), card memory (such as, multimedia card or micro card (e.g., secure digital (SD) or extreme digital (XD))), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device, any other device being configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner and provide the instructions or software and any associated data, data files, and data structures to the processor or computer such that the processor and computer can execute the instructions.

[0103] Although the present disclosure includes specific examples, it will be apparent to those of ordinary skill in the art that various changes in form and detail may be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein should be considered only as descriptive and not for purposes of limitation. The description of a feature or aspect in each example should be considered applicable to similar features or aspects in other examples. Appropriate results may be achieved if the described techniques are performed in a different order, and / or if the components in the described systems, architectures, devices, or circuits are combined in a different manner, and / or replaced or supplemented by other components or their equivalents. Accordingly, the scope of the disclosure is defined not by the specific embodiments but by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be construed as being included in the disclosure.

Claims

1. A method for identifying an image, the method comprising: Obtaining image data to be identified as input data for a neural network; Performing operations related to each layer of the neural network based on the image data to be identified to obtain an image recognition result; And Outputting the image recognition result, Wherein, for each of the target layers in the neural network: Extracting metadata of the target layer; Measuring the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; Selecting a corresponding layer among the reference layers based on the similarity; and Generating a parallelization strategy for the target layer based on a reference parallelization strategy matching the corresponding layer, and performing parallel processing of operations related to the target layer based on the parallelization strategy using the input data of the target layer, Wherein the step of selecting the corresponding layer includes: in response to the non-existence of a first reference layer having the same reference metadata as the metadata of the target layer among the reference layers, selecting a second reference layer having the reference metadata most similar to the metadata of the target layer among the reference layers as the corresponding layer, Wherein, in response to the second reference layer being selected as the corresponding layer, adding reference layer information corresponding to the metadata of the target layer to a reference database, in which the reference metadata of each reference layer is stored.

2. The method according to claim 1, wherein, The step of selecting the corresponding layer further includes: In response to the existence of the first reference layer among the reference layers, selecting the first reference layer as the corresponding layer.

3. The method according to claim 1, wherein, The reference layer information includes link information, and In response to the second reference layer being selected as the corresponding layer, recording the identification information of the second reference layer as the link information in the reference layer information.

4. The method according to claim 1, further comprising: Generating a new parallelization strategy corresponding to the metadata of the target layer in response to the reference layer information corresponding to the metadata of the target layer being added to the reference database.

5. The method according to claim 4, wherein Performing the generation of the new parallelization strategy independently of the execution of the parallelization strategy of the target layer.

6. The method according to claim 4, wherein, In response to the amount of the new reference layer information exceeding a threshold, performing the generation of the new parallelization strategy, the new reference layer information including the reference layer information corresponding to the metadata of the target layer and being added to the reference database.

7. The method according to claim 4, wherein, Using a similarity measurement model based on a neural network to measure the similarity, and In response to the new parallelization strategy being generated, retraining the similarity measurement model based on the new parallelization strategy.

8. The method according to claim 4, wherein In response to the new parallelization strategy being generated, changing the link information of the reference layer information corresponding to the metadata of the target layer to an empty state.

9. The method according to claim 1, wherein The target layer is one or more layers selected from multiple layers of the neural network.

10. The method according to claim 1, wherein, The input data of the first layer of the neural network is the image data to be identified, and the input data of the layers other than the first layer of the neural network is the output data of the previous layer.

11. An apparatus for identifying an image, the apparatus comprising: A processor; And A memory including instructions executable by the processor, Wherein, in response to the instructions being executed by the processor, the processor is configured to: Obtain image data to be identified as input data for a neural network; Perform operations related to each layer of the neural network based on the image data to be recognized to obtain the result of image recognition; and Output the result of image recognition, wherein, for each in the target layer of the neural network: Extract the metadata of the target layer; Measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; Select the corresponding layer among the reference layers based on the similarity; and Generate the parallelization strategy of the target layer based on the reference parallelization strategy matching the corresponding layer, and perform parallel processing of the operations related to the target layer using the input data of the target layer based on the parallelization strategy, wherein the processor is configured to: In response to the non-existence of a first reference layer having the same reference metadata as the metadata of the target layer among the reference layers, select a second reference layer having the reference metadata most similar to the metadata of the target layer as the corresponding layer; and In response to the second reference layer being selected as the corresponding layer, add the reference layer information corresponding to the metadata of the target layer to the reference database, and the reference database stores the reference metadata of each reference layer.

12. The apparatus according to claim 11, wherein The processor is configured to: In response to the existence of the first reference layer among the reference layers, select the first reference layer as the corresponding layer.

13. The apparatus according to claim 11, wherein The reference layer information includes link information, and The processor is configured to: in response to the second reference layer being selected as the corresponding layer, record the identification information of the second reference layer as the link information in the reference layer information.

14. The device according to claim 11, wherein, The processor is configured to: in response to the reference layer information corresponding to the metadata of the target layer being added to the reference database, generate a new parallelization strategy corresponding to the metadata of the target layer.

15. The device according to claim 14, wherein, The processor is configured to: Use a similarity measurement model based on the neural network to measure the similarity, and In response to the generation of the new parallelization strategy, retrain the similarity measurement model based on the new parallelization strategy.

16. The device according to claim 14, wherein, In response to the generation of the new parallelization strategy, change the link information of the reference layer information corresponding to the metadata of the target layer to an empty state.

17. An apparatus for training a neural network for image recognition, comprising: A processor; And A memory including instructions executable by the processor, wherein, in response to the instructions being executed by the processor, the processor is configured to: Obtain training image data; Train the neural network based on the training image data, wherein, during the training, for each in the target layer of the neural network: Extract the metadata of the target layer; Measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; Select the corresponding layer among the reference layers based on the similarity; and Generate the parallelization strategy of the target layer based on the reference parallelization strategy matching the corresponding layer, and perform parallel processing of the operations related to the target layer using the input data of the target layer based on the parallelization strategy, wherein the processor is configured to: In response to the non-existence of a first reference layer having the same reference metadata as the metadata of the target layer among the reference layers, select a second reference layer having the reference metadata most similar to the metadata of the target layer as the corresponding layer; and In response to the second reference layer being selected as the corresponding layer, add reference layer information corresponding to the metadata of the target layer to a reference database that stores reference metadata for each reference layer.

18. The device according to claim 17, wherein, The processor is configured to: In response to the first reference layer existing among the reference layers, select the first reference layer as the corresponding layer.

19. A method for training a neural network for image recognition, comprising: Obtain training image data; Train the neural network based on the training image data, wherein, during training, for each in the target layer of the neural network: Extract the metadata of the target layer; Measure the similarity between the target layer and each reference layer by comparing the metadata of the target layer with the reference metadata of each reference layer; Select a corresponding layer among the reference layers based on the similarity; and Generate a parallelization strategy for the target layer based on a reference parallelization strategy matching the corresponding layer, and perform parallel processing of operations related to the target layer using the input data of the target layer based on the parallelization strategy, wherein the step of selecting the corresponding layer includes: in response to the first reference layer not existing among the reference layers with reference metadata identical to the metadata of the target layer, select a second reference layer having reference metadata most similar to the metadata of the target layer as the corresponding layer, wherein, in response to the second reference layer being selected as the corresponding layer, add reference layer information corresponding to the metadata of the target layer to a reference database that stores reference metadata for each reference layer.

20. The method according to claim 19, wherein, The target layer is one or more layers selected from multiple layers of the neural network, wherein the step of selecting the corresponding layer further includes: in response to the first reference layer existing among the reference layers, select the first reference layer as the corresponding layer.

21. The method according to claim 19, wherein, The input data of the first layer of the neural network is the training image data, and the input data of the layers other than the first layer of the neural network is the output data of the previous layer.

22. A computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 10 and claims 19 to 21.

Citation Information

Patent Citations

  • vehicle systems

    KR1020200032233A

  • Information processing method, information processing apparatus, and computer readable storage medium

    US20190279087A1

  • Parallel processing method and apparatus for neutral network model

    US20210287085A1