METHOD AND SYSTEM FOR OPTIMIZING LEARNING MODEL FOR A TARGET DEVICE - Patent application
The method and system optimize learning models for embedded devices by encoding them for multiple environments and adapting to target device conditions, addressing inefficiencies in existing tools and enabling effective deployment across varied operating conditions.
Patent Information
- Application Number
- JP2024538768
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-31
- Filing Date
- 2022-12-29
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing AI development tools are inefficient for embedded devices, particularly IoT devices, and lack the ability to optimize deep learning models for various operating environments, hindering their effective incorporation into computing devices used in fields like security, transportation, manufacturing, and smart homes.
A method and system that generates a learning model using a predetermined framework, encodes it for multiple operating environments, and adapts it to target devices through inference engines suitable for their specific environments, allowing for retraining based on field data.
The system optimizes learning models to operate effectively across diverse environments, ensuring efficient deployment and adaptation to changing conditions.
Smart Images

Figure 0007790774000001 
Figure 0007790774000002 
Figure 0007790774000003
Abstract
Description
[Technical Field]
[0001] The technical idea of this disclosure relates to a method and system for optimizing a learning model for a target device. [Background technology]
[0002] Recently, there has been a demand for adding artificial intelligence (AI)-based functions to computing devices, particularly embedded devices, which are used in fields such as security, transportation, manufacturing, medicine, autonomous driving, and smart homes. However, there has not yet been sufficient effort to develop development tools such as application programs that can efficiently incorporate deep learning models created using AI into such embedded devices.
[0003] The efficiency of developing AI-based programs for embedded devices remains low, and is considered to be the biggest bottleneck hindering the development of related fields. In particular, existing AI development tools are primarily designed for the development of servers or cloud platforms, and there are currently insufficient AI development tools suited to embedded devices such as Internet of Things (IoT) devices.
[0004] In addition, to effectively incorporate deep learning models into embedded devices with various operating environments, various optimization processes are required, but conventional technologies alone have limitations in that they cannot guarantee optimal operation for deep learning models. Summary of the Invention [Problem to be solved by the invention]
[0005] The technical problem that the method and system for optimizing a learning model for a target device based on the technical idea of this disclosure aims to solve is to provide a method and system that can optimize and install a learning model to suit various operating environments of the target device.
[0006] The technical problems that the method and system according to the technical idea of this disclosure attempt to solve are not limited to the above-mentioned technical problems, and other technical problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]
[0007] According to one aspect of the technical idea of this disclosure, a method for optimizing a learning model for a target device may include the steps of: a learning model generation device training a network function based on a predetermined framework using training data to generate a learning model; a step of the learning model generation device encoding the learning model via a general-purpose encoding module so that it can be applied to multiple inference engines corresponding to multiple different operating environments or frameworks; a step of the learning model generation device transferring the encoded learning model to the target device; a step of the target device acquiring field data related to the applied process; and a step of the target device using the field data to perform inference based on the learning model via an inference engine among the inference engines that is suitable for the operating environment of the target device.
[0008] According to an exemplary embodiment, the operating environments may be differentiated based on the operating system and / or hardware of the target device.
[0009] According to an exemplary embodiment, the operating environment distinguished by the hardware may include at least one of a general-purpose central processing unit (CPU) environment, a general-purpose graphics processing unit (GPU) environment, and an embedded environment.
[0010] According to this exemplary embodiment, the field data input to the learning model and the output data of the learning model may be changed into a predetermined material structure via at least one buffer.
[0011] According to an exemplary embodiment, the method for optimizing a learning model for the target device may further include a step in which the learning model generation device retrains the learning model based on the on-site data and the output data.
[0012] According to one aspect of the technical idea of this disclosure, a learning model optimization system for a target device may include a learning model generation device that generates a learning model by training a network function based on a predetermined framework using training data, the learning model generation device encoding the learning model via a general-purpose encoding module so that the learning model can be applied to a plurality of inference engines corresponding to a plurality of different operating environments or frameworks, and the learning model generation device passing the encoded learning model to the target device, and the target device that acquires field data related to a process and uses the field data to perform inference based on the learning model via an inference engine among the inference engines that is suitable for the operating environment. [Effects of the Invention]
[0013] According to an embodiment based on the technical idea of this disclosure, a learning model generated in a learning model generation device (server) can be converted and operated to suit various operating environments of a target device.
[0014] The effects obtained by the method and system according to the technical idea of this disclosure are not limited to the effects described above, and other effects not mentioned will be clearly understood by a person having ordinary skill in the technical field to which this disclosure pertains from the following description.
[0015] To more fully understand the drawings referred to in this disclosure, a brief description of each drawing is provided. [Brief explanation of the drawings]
[0016] [Figure 1] FIG. 1 is a simplified block diagram illustrating a system for optimizing a learning model for a target device according to an embodiment of the present disclosure. [Figure 2] 1 is a flowchart illustrating a method for optimizing a learning model for a target device according to an embodiment of the present disclosure. [Figure 3] A block diagram conceptually showing the configuration of a learning model generation module of a learning model generation device and an inference module of a target device according to an embodiment of the present disclosure. [Figure 4] 1 is a flowchart illustrating a method for optimizing a learning model for a target device according to an embodiment of the present disclosure. [Figure 5] 1 is a block diagram showing a simplified configuration of a learning model generation device according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a block diagram illustrating a simplified configuration of a target device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0017] Since the technical idea of this disclosure can be variously modified and can have various embodiments, specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit the technical idea of this disclosure to the specific embodiments, and includes all modifications, equivalents, and alternatives included within the scope of the technical idea of this disclosure.
[0018] In explaining the technical ideas of this disclosure, if it is recognized that a specific description of known technologies related to the present invention may obscure the gist of this disclosure, the detailed description will be omitted. Note that numbers used in the explanation of this disclosure (e.g., "first," "second," etc.) are merely identification codes for distinguishing certain components from other components.
[0019] Furthermore, in this disclosure, when a component is referred to as being "coupled" or "connected" to another component, the component may be directly coupled or connected to the other component, but unless otherwise specified in this specification or clearly contradicted by the context, there may also be other components between them and the components may be coupled or connected via the other components.
[0020] Furthermore, the terms "unit," "device," "subsystem," "module," and the like used in this disclosure refer to a unit that processes at least one function or operation, which may be realized by hardware, software, or a combination of hardware and software, such as a processor, microprocessor, microcontroller, central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA), etc.
[0021] It should be made clear that the distinctions between components in this disclosure are merely made according to the main function that each component is responsible for. That is, two or more components described below may be combined into one component, or one component may be further divided into two or more components for each of its functions. It goes without saying that each component described below may perform some or all of the functions that are performed by other components in addition to its own main function, or that some of the main functions that each component is responsible for may be exclusively performed by other components.
[0022] The method according to the embodiment of the present disclosure may be performed in a computing device such as a personal computer, workstation, or server with computing power, or may be performed in a separate device for this purpose.
[0023] The method may also be performed in one or more computing devices. For example, at least one step of the method according to the embodiments of the present disclosure may be performed in a client device, and other steps may be performed in a server device. In such a case, the client device and the server device may be connected via a network to transmit and receive the computation results. Alternatively, the method may be performed using a distributed computing technique.
[0024] Furthermore, throughout this specification, network function may be used synonymously with neural network and / or neural network. Here, a neural network may generally be composed of a collection of interconnected computational units that may be referred to as nodes, and such nodes may also be referred to as neurons. A neural network generally comprises a plurality of nodes. The nodes that make up a neural network may be connected to each other by one or more links.
[0025] Some of the nodes that make up a neural network may form a layer based on their distance from the first input node. For example, a set of nodes that are n distances from the first input node may form n layers.
[0026] The neural networks described herein may comprise deep neural networks (DNNs) that include multiple hidden layers in addition to input and output layers.
[0027] The embodiments of this disclosure will be described in detail below.
[0028] FIG. 1 is a simplified block diagram of a system for optimizing a learning model for a target device according to an embodiment of the present disclosure.
[0029] The system may include a learning model generation device 110 and a target device 120 .
[0030] In an embodiment, the learning model generation device 110 may be a cloud server.
[0031] The learning model generation device 110 may generate a learning model and transfer it to the target device 120. To this end, the learning model generation device 110 may include a learning model generation module. The learning model generation module may be a software module.
[0032] Specifically, the learning model generation module may generate a network function based on a predetermined framework (e.g., TensorFlow), and if training data is input from outside, may perform training on the network function based on the input data and generate a learning model as a result of the training. In addition, the learning model generation module may encode the generated learning model via a general-purpose encoding module into a format that can be run in multiple operating environments and / or frameworks. The learning model may be passed to the target device 120 in the form of an encoded file.
[0033] In addition, in an embodiment, the learning model generation device 110 may retrain the learning model using the field data passed from the target device 120 and the output data output by the learning model based on the field data.
[0034] The target device 120 refers to a target device on which a learning model is to be implemented, and encompasses all types of computing devices, such as a server, a personal computer (PC), a smartphone, a smart pad, a tablet PC, etc. In an embodiment, the target device 120 may further include an IoT terminal, an edge device, an embedded board, etc. that operates in an embedded environment.
[0035] The target device 120 may receive at least one executable code, program, library, etc. for generating an encoded learning model and an inference module from the learning model generation device 110, and drive the learning model through the generated inference module. In this case, the inference module may be a software module.
[0036] The inference module may include multiple inference engines and at least one buffer (input / output buffer) corresponding to different operating environments. The inference module may run a learning model through an inference engine suitable for the operating environment, such as the operating system and / or hardware of the target device 120, and perform inference using field data acquired by the target device 120.
[0037] At this time, the field data input as the learning model and the output data of the learning model may be converted into a predetermined data structure via a buffer.
[0038] FIG. 2 is a flowchart illustrating a method for optimizing a learning model for a target device according to an embodiment of the present disclosure, and FIG. 3 is a block diagram conceptually showing the configuration of a learning model generation module of a learning model generation device according to an embodiment of the present disclosure and an inference module of a target device.
[0039] In step S210, the learning model generation device 110 may generate the learning model 312 by inputting learning data to the network function 311 of the learning model generation module 310 and performing learning.
[0040] In this case, the network function 311 may be generated based on a predetermined framework. The framework refers to a type of package configured to allow developers to easily use a plurality of verified libraries and modules, pre-trained algorithms, etc., and may include, but is not limited to, Tensorflow, Keras, Pytorch, Cafe, MXnet, etc., and a wide variety of frameworks can be used to generate the network function 311.
[0041] In this embodiment, the learning model generation module 310 may generate the learning model 312 using a network function based on tensorflow.
[0042] In an embodiment, the training data may be a plurality of images. For example, the training data may be process images acquired (or photographed) during any process such as product production, manufacturing, or processing, and may include information about a target (label) to be detected or read.
[0043] In step S220, the learning model generation device 110 may encode the learning model 312 via the general-purpose encoding module 313 of the learning model generation module 310. Through step S220, the learning model 312 can be converted into a format that can be run in multiple different operating environments and / or frameworks.
[0044] In an embodiment, the encoding module 313 may convert the training model 312 into a neural network data format such as, but not limited to, Neural Network Exchange Format (NNEF) or Open Neural Network Exchange (ONNX).
[0045] In step S230, the learning model generation device 110 may pass the encoded learning model 312 to the target device 120.
[0046] Here, the target device 120 is provided to control the process or measure or acquire specific data within the process environment and perform diagnosis or prediction related to the process based on the data, and may be provided with an inference module 320 for driving a learning model to perform inference.
[0047] In step S240, the target device 120 may acquire on-site data related to the process. For example, the target device 120 may be provided as a single component, or may acquire the on-site data from a sensing device (at least one sensor, camera, etc.) connected to the target device 120 via wired or wireless communication.
[0048] In an embodiment, the in-situ data may be at least one image (eg, a process image).
[0049] In step S250, the target device 120 may perform inference using the in-situ data via the inference module 320.
[0050] In the S250 embodiment, the inference module 320 may include multiple inference engines 321, each suitable for a different operating environment, and may perform inference based on field data by driving the learning model 312 through one of the multiple inference engines 321 that is suitable for the operating environment of the target device 120.
[0051] Here, the operating environment may be distinguished based on at least one of the operation system and the hardware (or hardware architecture) of the target device 120 .
[0052] In an embodiment, the operating environment distinguished by the operating system may include at least one of a Linux environment and a window environment, however, this is merely an example and different operating environments may be defined for a wide variety of operating systems such as iOS, Android, and Mac OS.
[0053] In addition, in an embodiment, the operating environment distinguished by hardware may include at least one of a general-purpose central processing unit (CPU) environment, a general-purpose graphics processing unit (GPU) environment, and an embedded environment. However, this is merely an example, and different operating environments may be defined according to a wide variety of hardware architectures, such as Nvidia, Jetson, Raspberry, ARM Cortex-A Family, Qcam Family, Hisilicon Family, X86-CPU, and X86-GPU.
[0054] In an embodiment, the inference engine 321 may include a first inference engine suitable for a general-purpose GPU environment and / or embedded environment and a second inference engine suitable for a general-purpose CPU environment.
[0055] In an embodiment, input data and output data of the learning model 312 driven via the inference engine 321 may be converted into a predetermined data structure via at least one buffer 322 .
[0056] For example, the buffer 322 may include an input buffer and an output buffer, and the target device 120 may allocate the on-site data to the input buffer and convert it into a data structure suitable for driving the learning model 312, and may allocate the output data generated based on the inference result of the learning model 312 to the output buffer and convert it into a specified data structure. The data structure may be a structure corresponding to a programming language that the target device 120 can parse.
[0057] FIG. 4 is a flowchart illustrating a method for optimizing a learning model for a target device according to an embodiment of the present disclosure.
[0058] Steps S410 to S450 in the method 400 are similar to steps S210 to S250 in the method 200 described above with reference to FIG. 2, and therefore a duplicated description will be omitted.
[0059] The method 400 may further include retraining the training model 313 .
[0060] Specifically, in step S460, the target device 120 may input field data and field data and pass on the output data obtained via the inference engine 321 and the learning model 313 to the learning model generation device 110, and the learning model generation device 110 may update the learning model 313 by re-training the learning model 313 based on the received field data and output data.
[0061] The updated learning model 313 may be encoded and transmitted to the target device 120 via a general-purpose encoding module 313 .
[0062] In an embodiment, step S460 may be performed periodically at regular time intervals.
[0063] FIG. 5 is a block diagram showing a simplified configuration of a learning model generation device according to an embodiment of the present disclosure.
[0064] The communication unit 510 may transmit and receive data to and from an external device. The communication unit 510 may include a wired / wireless communication unit. If the communication unit 510 includes a wired communication unit, the communication unit 510 may include one or more components that perform communication via a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, or a combination thereof. If the communication unit 510 includes a wireless communication unit, the communication unit 510 may transmit and receive data or signals wirelessly using cellular communication, wireless LAN (e.g., Wi-Fi), or the like. In an embodiment, the communication unit may transmit and receive data or signals to and from an external device or an external server under the control of the processor 540.
[0065] The input unit 520 may receive various user commands through external operations. To this end, the input unit 520 may include or be connected to one or more input devices. For example, the input unit 520 may be connected to various input interfaces, such as a keypad or a mouse, to receive user commands. To this end, the input unit 520 may include an interface such as a USB port or Thunderbolt. Furthermore, the input unit 520 may include various input devices, such as a touch screen or a button, or may be coupled to these to receive external user commands.
[0066] The memory 530 may store programs and / or program commands for the operation of the processor 540, and may temporarily or permanently store input / output data. The memory 530 may include at least one type of storage medium selected from the group consisting of a flash memory type, a hard disk type, a multimedia card micro type, a card-type memory (e.g., SD or XD memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.
[0067] Memory 530 may also store various network functions, algorithms, libraries, etc., and may store a wide variety of data, programs (one or more instructions), applications, software, instructions, code, etc. for driving and controlling device 500.
[0068] The processor 540 may execute at least one program and / or instruction to control the overall operation of the device 500. The processor 540 may execute one or more programs stored in the memory 530. The processor 540 may refer to a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor on which the methods according to the concepts of the present disclosure are performed.
[0069] In an embodiment, the processor 540 may generate the learning model generation module 310 based on at least one program and / or instruction. The learning model generation module 310 may be a software module.
[0070] In an embodiment, the processor 540 may generate a learning model by training a network function based on a predetermined framework using training data, and the learning model generation device may encode the learning model via a general-purpose encoding module so that it can be applied to multiple inference engines corresponding to multiple different operating environments or frameworks, and pass the encoded learning model to the target device.
[0071] In an embodiment, the processor 540 may retrain the learning model based on the in-situ data received from the target device and the output data of the inference engine and the learning model.
[0072] FIG. 6 is a block diagram showing a simplified configuration of a target device according to an embodiment of the present disclosure.
[0073] The communication unit 610 may receive data from an external device. The communication unit 610 may include a wired / wireless communication unit. If the communication unit 610 includes a wired communication unit, the communication unit 610 may include one or more components that perform communication via a local area network (LAN), a wide area network (WAN), a value-added network (VAN), a mobile radio communication network, a satellite communication network, or a combination thereof. If the communication unit 610 includes a wireless communication unit, the communication unit 610 may transmit and receive data or signals wirelessly using cellular communication, wireless LAN (e.g., Wi-Fi), etc. In an embodiment, the communication unit may transmit and receive data or signals to and from an external device or an external server under the control of the processor 640.
[0074] The input unit 620 may receive various user commands through external operations. To this end, the input unit 620 may include or be connected to one or more input devices. For example, the input unit 620 may be connected to various input interfaces, such as a keypad or a mouse, to receive user commands. To this end, the input unit 620 may include an interface such as a USB port or Thunderbolt. Furthermore, the input unit 620 may include various input devices, such as a touch screen or a button, or may be coupled to these to receive external user commands.
[0075] The memory 630 may store programs and / or program commands for the operation of the processor 640, and may temporarily or permanently store input / output data. The memory 630 may include at least one type of storage medium selected from the group consisting of a flash memory type, a hard disk type, a multimedia card micro type, a card-type memory (e.g., SD or XD memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk.
[0076] Memory 630 may also store various network functions, algorithms, libraries, etc., and may store a wide variety of data, programs (one or more instructions), applications, software, instructions, code, etc. for driving and controlling device 600.
[0077] The processor 640 may execute at least one program and / or instruction to control the overall operation of the device 600. The processor 640 may execute one or more programs stored in the memory 630. The processor 640 may refer to a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processor on which the methods according to the concepts of the present disclosure are performed.
[0078] In this embodiment, the processor 640 may generate the inference module 320 based on the learning model, at least one program, and / or instructions received from the learning model generation device. The inference module 320 may be a software module.
[0079] In an embodiment, the processor 540 may acquire field data associated with the process and use the field data to perform inference based on the learning model via an inference engine that is suitable for the operating environment of the device 600 among the inference engines.
[0080] In an embodiment, the processor 540 may retrain the learning model based on the in-situ data received from the target device and the output data of the inference engine and the learning model.
[0081] On the other hand, although not shown, the device 600 may further include a sensing unit and / or an imaging unit for acquiring on-site data. In this case, the sensing unit may be configured with at least one sensing module, and the imaging unit may be configured with a camera module.
[0082] Methods according to embodiments of the present disclosure may be embodied in computer-readable media and recorded as program commands that can be executed by a variety of computer means. The computer-readable media may include, alone or in combination, program commands, data files, data structures, and the like. The program commands recorded on the media may be those specially designed and constructed for this disclosure, or they may be those well known and available to those skilled in the art of computer software. Examples of computer-readable media include magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as compact discs (CDs) read-only memories (CD-ROMs) and digital versatile discs (DVDs); magneto-optical media such as floptical disks; and hardware devices specially configured to store and execute program commands, such as read-only memory (ROM), random access memory (RAM), and flash memory. Examples of program commands include not only machine language code, such as produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
[0083] Additionally, methods according to the disclosed embodiments may be provided in a computer program product, which may be traded as a commodity between sellers and buyers.
[0084] The computer program product may include a software program and a computer-readable storage medium on which the software program is stored. For example, the computer program product may include a product in the form of a software program (e.g., a downloadable application) that is electronically distributed by an electronic device manufacturer or via an online marketplace (e.g., Google Play Store, application store). For electronic distribution, at least a portion of the software program may be stored on a storage medium or may be temporarily generated. In this case, the storage medium may be a storage medium of a manufacturer's server, an online marketplace server, or a relay server that temporarily stores the software program.
[0085] In a system including a server and a client device, the computer program product may comprise a storage medium of the server or a storage medium of the client device. Alternatively, if a third device (e.g., a smartphone) is present and connected to the server or the client device via communication, the computer program product may comprise a storage medium of the third device. Alternatively, the computer program product may include the S / W program itself, which is transmitted from the server to the client device or the third device, or transmitted from the third device to the client device.
[0086] In this case, one of the server, the client device, and the third device may run the computer program product to perform the method according to the disclosed embodiments, or two or more of the server, the client device, and the third device may run the computer program product to perform the method according to the disclosed embodiments in a distributed manner.
[0087] For example, a server (e.g., a cloud server or an artificial intelligence server) may launch a computer program product stored on the server to control client devices communicatively connected to the server to perform methods according to the disclosed embodiments.
[0088] Although the embodiments have been described in detail above, the scope of the present disclosure is not limited thereto in any way, and various modifications and improvements made by those skilled in the art using the basic concepts of the present disclosure as defined in the appended claims also fall within the scope of the present disclosure.
Claims
1. 1. A method for optimizing a learning model for a target device, comprising: a step in which a learning model generation device generates a learning model by training a network function based on a predetermined framework using training data; a step in which the learning model generation device encodes the learning model via a general-purpose encoding module so that the learning model can be commonly applied to a plurality of inference engines corresponding to different operating environments; the learning model generation device transferring the encoded learning model to the target device; the target device acquiring in situ data associated with the applied process; and performing, by the target device, inference based on the learning model using the on-site data via an inference engine suitable for the operating environment of the target device, among a plurality of inference engines provided in the target device and corresponding to different operating environments distinguished based on at least one of an operation system and hardware. A method characterized by:
2. The operating environment distinguished by the hardware includes at least one of a general-purpose central processing unit (CPU) environment, a general-purpose graphics processing unit (GPU) environment, and an embedded environment. The method of claim 1.
3. The field data input to the learning model and the output data of the learning model are converted into a predetermined data structure through at least one buffer. The method of claim 1.
4. The learning model generation device further includes a step of re-learning the learning model based on the field data and the output data. The method of claim 3.
5. 1. A system for optimizing a learning model for a target device, comprising: a learning model generation device that generates a learning model by training a network function based on a predetermined framework using training data, encodes the learning model via a general-purpose encoding module so that the learning model can be applied to multiple inference engines corresponding to multiple different operating environments or frameworks, and transfers the encoded learning model to the target device; the target device acquires on-site data related to a process, and performs inference based on the learning model using the on-site data via an inference engine that is suitable for a driving environment among the inference engines; The target device performs inference based on the learning model using the on-site data via an inference engine suitable for the operating environment of the target device, out of a plurality of inference engines provided in the target device and corresponding to different operating environments distinguished based on at least one of an operating system and hardware of the target device. A system characterized by:
Citation Information
Patent Citations
Express sheet recognition model transplanting method, device and apparatus and storage medium
CN111814906A
Systems and methods for processing medical data
WO2021252384A1