Electronic device for processing image acquired using superlens and operating method thereof

By using ultra-lens encoded images in household appliances or IoT devices and combining AI models for identification, the risk of personal information leakage during monitoring is solved, and higher privacy protection and security are achieved.

CN120036006APending Publication Date: 2025-05-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202380064525.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-07-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When monitoring using artificial intelligence (AI) technology, there is a risk of personal information leakage due to hijacking of image data or hacking attacks, especially in home appliances or IoT devices.

Method used

The encoded image is obtained by using a hyperlens and input it into an AI model for identification and classification to avoid the leakage of personal information. The superlens have surfaces that form a specific pattern, encode the image by modulating the phase of light so that the original image cannot be directly recognized.

Benefits of technology

It realizes the protection of personal privacy, enhances security, and improves the recognition rate during the monitoring process, avoiding the leakage of personal information caused by image data being hijacked or hacked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036006A_ABST
    Figure CN120036006A_ABST
Patent Text Reader

Abstract

An electronic device for obtaining a result of recognizing an object from an image obtained using a superlens, and an operation method of the electronic device are provided. An electronic device according to an embodiment of the present disclosure may include: a superlens having a pattern formed on a surface thereof and composed of a plurality of pillars or pins having different shapes, heights, and areas from each other, and having an optical property of modulating a phase of light reflected from an object by the pattern on the surface; an image sensor configured to obtain an encoded image by receiving light reflected from an object and phase-modulated by penetrating the superlens and converting the received light into an electrical signal; and at least one processor configured to input the encoded image to an artificial intelligence (AI) model and to obtain a tag indicating a result of identifying the object by inferring using the AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an electronic device for processing an image obtained using a super lens and an operating method thereof. Specifically, the present disclosure provides an electronic device for performing image processing for detecting, classifying or segmenting an object from a coded image by using an artificial intelligence (AI) model. Background Art

[0002] Recently, cameras have been installed in home appliances or Internet of Things (IoT) devices such as TVs, refrigerators, or robot vacuum cleaners, and images captured by cameras are increasingly used to monitor objects. With the advancement of artificial intelligence (AI) technology based on visual recognition, when cameras are used for monitoring in places where personal privacy should not be exposed, such as at home or in indoor environments, there is a risk of personal information being leaked due to hijacking or hacker attacks on image data.

[0003] Conventional methods for preventing leakage of personal information may include methods using event cameras, which do not obtain red, green, and blue (RGB) images by receiving light reflected from an object, but obtain images based on the degree of change in the brightness (intensity) of light. Event cameras are advantageous in terms of privacy protection and prevention of leakage of personal information because they do not obtain RGB images that are visually recognizable by humans. However, the method using an event camera has a limitation in that the event camera can only obtain an image when there is a change in the intensity of light. Summary of the invention

[0004] Solution to the problem

[0005] According to one aspect of the present disclosure, an electronic device for obtaining a result of identifying an object from an image obtained by a super lens. According to an embodiment of the present disclosure, an electronic device may include a super lens having a pattern formed on its surface and composed of a plurality of columns or pins having shapes, heights and areas different from each other, and having an optical property of modulating the phase of light reflected from an object through a pattern on the surface. According to an embodiment of the present disclosure, the electronic device may include an image sensor configured to obtain a coded image by receiving light reflected from an object and phase modulated by penetrating the super lens and converting the received light into an electrical signal. According to an embodiment of the present disclosure, the electronic device may include at least one processor, at least one processor configured to input the coded image into an artificial intelligence (AI) model, and obtain a label corresponding to the result of identifying the object by performing inference using the AI ​​model. In an embodiment of the present disclosure, the AI ​​model may be a neural network model trained to obtain a simulated image by inputting a red, green and blue (RGB) image into a model reflecting the optical properties of the super lens, and outputting a label indicating the ground truth of the input RGB image as the result of identifying the obtained simulated image. In an embodiment of the present disclosure, the AI ​​model may be trained to minimize information indicating the similarity between the RGB image and the simulated image.

[0006] According to another aspect of the present disclosure, a method for identifying an object from an image obtained through a super lens performed by an electronic device is provided. The method may include obtaining a coded image by receiving light reflected from an object and phase-modulated by penetrating the super lens and converting the received light into an electrical signal. The method may include inputting the coded image to an AI model, and obtaining a label indicating a result of identifying the object by performing inference using the AI ​​model.

[0007] According to another aspect of the present disclosure, a computer program product including a computer-readable storage medium is provided. The storage medium may include instructions that can be read by an electronic device to obtain a coded image by receiving light reflected from an object and phase-modulated by penetrating a super lens and converting the received light into an electrical signal, and inputting the coded image to an AI model and performing inference using the AI ​​model to obtain a label indicating a result of recognizing the object. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present disclosure will be readily understood from the following description taken in conjunction with the accompanying drawings, in which reference numerals represent structural elements.

[0009] Figure 1 is a conceptual diagram illustrating an operation in which an electronic device obtains a coded image by using a super lens and outputs a visual recognition result by using an artificial intelligence (AI) model according to an embodiment of the present disclosure.

[0010] Figure 2 is a diagram illustrating a method of training an AI model according to an embodiment of the present disclosure.

[0011] Figure 3 is a block diagram illustrating components of an electronic device according to an embodiment of the present disclosure.

[0012] Figure 4 is a block diagram illustrating components of an electronic device and a server according to an embodiment of the present disclosure.

[0013] Figure 5 is a flowchart of a method for obtaining a visual recognition result from an image obtained using a super lens, performed by an electronic device according to an embodiment of the present disclosure.

[0014] Figure 6 is a perspective view showing the shape of a surface pattern of a superlens according to an embodiment of the present disclosure.

[0015] Figure 7 is a diagram illustrating a method of training an AI model according to an embodiment of the present disclosure.

[0016] Figure 8 is a diagram illustrating a method of training an AI model according to an embodiment of the present disclosure.

[0017] Fig. 9 is a flowchart of a method performed by an electronic device to obtain a visual recognition result by using an AI model that changes according to a purpose or use of a visual task according to an embodiment of the present disclosure.

[0018] Fig.10 is a diagram illustrating an operation in which an electronic device changes the configuration of an AI model according to a purpose or use of a visual task according to an embodiment of the present disclosure.

[0019] Fig.11 is a diagram illustrating an operation in which an electronic device changes the configuration of an AI model according to a purpose or use of a visual task according to an embodiment of the present disclosure.

[0020] Fig.12 is a flowchart of a method performed by an electronic device to change the configuration of an AI model in response to a replacement or change to a metalens being identified according to an embodiment of the present disclosure.

[0021] Fig.13 is a diagram illustrating an operation in which an electronic device changes the configuration of an AI model in response to replacement or change of a super lens being recognized according to an embodiment of the present disclosure.

[0022] Fig.14is a diagram illustrating an operation in which an electronic device obtains a visual recognition result by using a depth sensor and an audio sensor and a camera system including a super lens according to an embodiment of the present disclosure.

[0023] Fig.15 is a diagram illustrating an operation in which an electronic device obtains a visual recognition result by using a depth sensor and an audio sensor and a camera system including a super lens according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] As the terms used in the embodiments of this specification, currently widely used general terms are selected by considering the functions in the present disclosure, but these terms may change according to the intention of ordinary technicians in the field, precedents, the emergence of new technologies, etc. In addition, specific terms can be arbitrarily selected by the applicant, and in this case, the meaning of the selected terms will be described in detail in the detailed description of the corresponding embodiment. Therefore, the terms used herein should not be defined by their simple titles, but are defined based on the meaning of the terms and the overall description of the present disclosure.

[0025] Unless the context clearly indicates otherwise, singular expressions used herein are also intended to include plural expressions.All terms (including technical or scientific terms) used herein may have the same meanings as those generally understood by those of ordinary skill in the art described in this specification.

[0026] Throughout this disclosure, when a component "includes" or "comprises" an element, unless there is a specific description to the contrary, it should be understood that the component may further include other elements without excluding other elements. In addition, terms such as "part", "module", etc. used herein indicate a unit for processing at least one function or operation, and may be implemented as hardware or software or a combination of hardware and software.

[0027] Depending on the context, the expression "configured to (or set to)" used in this document may be used interchangeably with, for example, the expressions "suitable for", "capable of...", "designed to", "suitable for", "manufactured to", or "capable of". The term "configured to (or set to)" may not necessarily mean "specially designed to" in terms of hardware only. On the contrary, the expression "a system configured to..." may in some contexts mean that the system together with other devices or components "can...". For example, the expression "a processor configured to (or set to) perform A, B, and C" may mean a dedicated processor (e.g., an embedded processor) for performing the corresponding operations, or a general-purpose processor (e.g., a central processing unit (CPU) or an application processor (AP)) that can perform the corresponding operations by executing one or more software programs stored in a memory.

[0028] Furthermore, in the present disclosure, when a component is referred to as being “connected” or “coupled” to another component, it should be understood that the component may be directly connected or coupled to another component, but may also be connected or coupled to another component via another intermediate component therebetween, unless there is a specific description to the contrary.

[0029] As used herein, a "metalens" is a lens that includes a metasurface composed of a pattern of nanometer-sized posts or pins and has an optical property of modulating the phase of light reflected from an object. In an embodiment of the present disclosure, the metasurface may be formed by a plurality of posts or pins having different shapes, heights, and areas from one another.

[0030] In the present disclosure, a "coded image" is an image obtained using light whose phase is modulated by passing through a super lens. In an embodiment of the present disclosure, an image sensor may receive light whose phase is modulated by passing through a super lens, and obtain a coded image by converting the received light into an electrical signal.

[0031] In the present disclosure, functions related to artificial intelligence (AI) are performed via a processor and a memory. The processor may be configured as one or more processors. In this case, the one or more processors may be general-purpose processors such as a CPU, an AP, a digital signal processor (DSP), etc., dedicated graphics processors such as a graphics processing unit (GPU), a visual processing unit (VPU), etc., or dedicated AI processors such as a neural processing unit (NPU). One or more processors control the input data to be processed according to predefined operating rules or AI models stored in the memory. Alternatively, when one or more processors are dedicated AI processors, the dedicated AI processor may be designed with a hardware structure dedicated to processing a specific AI model.

[0032] The predefined operating rules or AI models are generated via a training process. In this case, the generation via a training process means creating a predefined operating rule or AI model that is set to perform the desired characteristics (or purpose) by training the base AI model based on a large amount of training data via a learning algorithm. The training process can be performed by the device itself on which the AI ​​according to the present disclosure is executed, or via a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.

[0033] In the present disclosure, an "AI model" may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weights, and performs a neural network calculation via calculations between the calculation results in the previous layer and the plurality of weights. The plurality of weights assigned to each of the plurality of neural network layers may be optimized by the result of training the AI ​​model. For example, the plurality of weights may be updated to reduce or minimize the loss or cost value obtained in the AI ​​model during the training process. The artificial neural network model may include a deep neural network (DNN), such as a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recursive DNN (BRDNN), or a deep Q network (DQN), but is not limited thereto.

[0034] As used in this disclosure, “visual recognition” refers to image signal processing, which involves inputting a red, green, blue (RGB) image or an encoded image into an AI model and detecting objects in the input image, classifying objects into specific categories, or segmenting objects via inference using an AI model.

[0035] The embodiments of the present disclosure will be described more fully below with reference to the accompanying drawings so that those skilled in the art can easily implement the embodiments. However, the present disclosure can be implemented in different forms and should not be construed as limited to the embodiments set forth herein.

[0036] Hereinafter, embodiments of the present disclosure are described in detail with reference to the accompanying drawings.

[0037] Figure 1 is a conceptual diagram illustrating an operation in which an electronic device obtains a coded image 12 by using a super lens 110 and outputs a visual recognition result by using an AI model 200 according to an embodiment of the present disclosure.

[0038] refer to Figure 1 , the electronic device may include a super lens 110, an image sensor 120, and an AI model 200.

[0039] The superlens 110 is a lens including a supersurface composed of nano-sized pillars or pins. The supersurface is a three-dimensional (3D) surface having a pattern composed of a plurality of pillars or a plurality of pins having different heights and areas from each other, and the distances between the plurality of pillars or the plurality of pins may be different. The superlens 110 may have an optical property of encoding light by changing the amount of light reflected from the object 10 and penetrating the supersurface. In an embodiment of the present disclosure, the superlens 110 may induce a phase delay of the light and modulate the phase of the light by changing the refractive index of the light according to a pattern composed of a plurality of pillars or pins included in the supersurface. In an embodiment of the present disclosure, the heights and areas of the plurality of pillars or pins included in the supersurface and the distances between them may be determined by parameter values ​​of a mathematical modeling model such as a point spread function (PSF) according to the result of training the AI ​​model 200. Reference Figure 6 Let’s describe the metasurface in detail.

[0040] The image sensor 120 is configured to receive light whose phase is modulated by passing through the metalens 110, and obtain the encoded image 12 by converting the received light into an electrical signal. In the embodiment of the present disclosure, the image sensor 120 may be composed of a complementary metal oxide semiconductor (CMOS), but is not limited thereto. The light reflected by the object 10 is phase-modulated by the metalens 110 due to the change in the refractive index, and the phase-modulated light is received by a specific pixel of the image sensor 120. The image sensor 120 may obtain the encoded image 12 by converting the received light into an electrical signal.

[0041] The coded image 12 is an image obtained by phase-modulating light reflected from the object 10 by the metalens 110 and receiving the modulated light by the image sensor 120. Unlike the RGB image, the coded image 12 may be a defocused or distorted or modulated image so that the shape of the object 10 cannot be recognized by the human eye.

[0042] The electronic device may input the encoded image 12 into the AI ​​model 200, and obtain a label 14 indicating a recognition result of the encoded image 12 by performing inference using the AI ​​model 200. The AI ​​model 200 may be a neural network model trained via supervised learning to obtain a simulated image by inputting an RGB image into a model reflecting the optical properties of the super lens 110, and output a label indicating the true value of the input RGB image as a recognition result of the obtained simulated image. In an embodiment of the present disclosure, the AI ​​model 200 may be implemented as a CNN model, but is not limited thereto. The AI ​​model 200 may be implemented as, for example, an RNN, an RBM, a DBN, a BRDNN, a DQN, etc. Reference Figure 2 Specific embodiments related to training AI model 200 are described in detail.

[0043] exist Figure 1In the illustrated embodiment, the electronic device can obtain a label of “bird” by inferring using the trained AI model 200 , where “bird” is the visual recognition result of the encoded image 12 input to the AI ​​model 200 .

[0044] In recent years, home appliances such as TVs, refrigerators or robot vacuum cleaners or Internet of Things (IoT) devices have been equipped with cameras, and images captured by cameras are increasingly used for target monitoring. With the advancement of AI technology based on visual recognition, when cameras are used for monitoring in places where personal privacy should not be exposed, such as at home or in indoor environments, there is a risk of personal information leakage due to hijacking or hacker attacks on image data. Event cameras used as a conventional method for preventing personal information leakage obtain images based on the degree of change in the intensity of light, so they may not be able to obtain RGB images that are visually recognizable to humans (human-readable). However, event cameras have the limitation that they can only obtain images when there is a change in the intensity of light. In addition, recent advances in AI technology make it possible to reconstruct images that are visually recognizable to humans from images obtained by event cameras, and therefore, a solution for protecting personal privacy and enhancing security is needed.

[0045] The present disclosure aims to provide an electronic device and an operating method thereof, which are used to obtain a coded image 12 that is visually unrecognizable to humans by using a super lens 110, and obtain a visual recognition result from the coded image 12, so as to protect personal privacy and enhance security when the camera is used for monitoring, etc.

[0046] according to Figure 1 The electronic device of the embodiment shown in the figure obtains a coded image 12 that cannot be visually recognized and verified by humans by modulating light reflected from an object 10 via a super lens 110, thereby providing a technical effect of preventing the risk of personal information leakage in advance and enhancing security even when used indoors such as at home. In addition, according to an embodiment of the present disclosure, the electronic device obtains a visual recognition result from the coded image 12 by inferring using a trained AI model 200, thereby improving the recognition rate compared to existing methods (e.g., event cameras, etc.). In addition, because the super lens 110 does not require multiple lenses to receive light, unlike a conventional camera lens, the electronic device according to an embodiment of the present disclosure provides the effect of reducing the lens thickness and achieving a small form factor by using the super lens 110.

[0047] Figure 2 is a diagram illustrating a method of training an AI model 200 according to an embodiment of the present disclosure.

[0048] In the embodiments of the present disclosure, Figure 2The AI ​​model 200 shown in FIG. 1 may be included on-device within an electronic device. However, the present disclosure is not limited thereto, and in another embodiment of the present disclosure, the AI ​​model 200 may be stored on a server or an external device.

[0049] The AI ​​model 200 may be a model trained using supervised learning to output a label 22 corresponding to a true value paired with the RGB image 20 when the RGB image 20 is input. In an embodiment of the present disclosure, the AI ​​model 200 may be an end-to-end neural network model trained to minimize a loss, which is the difference between a label predicted from the RGB image 20 and a label corresponding to the true value.

[0050] refer to Figure 2 , the AI ​​model 200 may include a first AI model 210 and a second AI model 220 .

[0051] The first AI model 210 may be trained to reflect the super lens ( Figure 1 110) to output a simulated image from an input RGB image 20. The first AI model 210 may include a differentiable mathematical model that mathematically models the optical properties of the metalens 110. The first AI model 210 may output a simulated image by inputting the RGB image 20 into the differentiable mathematical model. In an embodiment of the present disclosure, the first AI model 210 may model the RGB image data via a mathematical equation by reflecting the phase modulation performed by the metasurface of the metalens 110, and obtain a simulated image by performing a convolution of the modeled RGB image 20 with a PSF that mathematically models the focusing distortion or modulation of each pixel of the RGB image 20. A “simulated image” is an image generated by a differentiable mathematical model that is trained to reflect the optical properties of the metalens 110, and may be an image generated via simulation, similar to an image obtained by an image sensor ( Figure 1 120) is an image obtained by using light that has penetrated the super lens 110.

[0052] In an embodiment of the present disclosure, the first AI model 210 may obtain a simulated image by adding simulated sensor noise in the image sensor 120 to the convolution result between the RGB image 20 and the PSF. The first AI model 210 may output the obtained simulated image to the second AI model 220.

[0053] The second AI model 220 may be a neural network model trained to output a predicted label based on a visual recognition result from an input simulated image. In an embodiment of the present disclosure, the second AI model 220 may be implemented as a CNN model. However, the present disclosure is not limited thereto, and in another embodiment of the present disclosure, the second AI model 220 may be implemented as, for example, an RNN, an RBM, a DBN, a BRDNN, a DQN, etc.

[0054] The AI ​​model 200 may be a neural network model that is trained to update and optimize a plurality of weights of a neural network layer included in the first AI model 210 and the second AI model 220 via back propagation, which is performed by applying a loss 26 to the first AI model 210 and the second AI model 220, the loss 26 being the difference between a label 22 of a true value of the input RGB image 20 and a label 24 predicted from the RGB image 20. In an embodiment of the present disclosure, the AI ​​model 200 may be an end-to-end neural network model that is trained to output a label 22 that is a true value from the input RGB image 20. In an embodiment of the present disclosure, during the training process of the AI ​​model 200, a gradient obtained using a partial derivative of an error function representing the loss 26 is applied and back propagated from the second AI model 220 to the first AI model 210, and through this, a plurality of weights of a plurality of neural network layers included in the first AI model 210 and the second AI model 220 may be updated or optimized.

[0055] During the training process of the AI ​​model 200, the first AI model 210 may be trained to minimize information indicating the similarity between the input RGB image 20 and the output simulated image. In an embodiment of the present disclosure, the first AI model 210 may be trained using a loss (mutual information loss 28), which is a mathematical quantification of the correlation between the RGB image 20 and the simulated image, so that the mutual information representing the correlation between them is minimized. For example, during the training process of the first AI model 210, multiple weights in the first AI model 210 may be updated via applying the mutual information loss 28 to the back propagation of the first AI model 210. The greater the correlation (e.g., cross entropy (or CE)) between the RGB image 20 and the simulated image, the more likely it is that a human-readable image can be reconstructed from the simulated image output by the first AI model 210 when the simulated image leaks, and therefore, the first AI model 210 may be trained via back propagation using the value of the mutual information loss 28.

[0056] In an embodiment of the present disclosure, a method of minimizing the Kullback-Leibler divergence (KL-divergence) can be used to minimize the correlation between the RGB image 20 and the simulated image during the training process of the first AI model 210. The Kullback-Leibler divergence mathematically represents the dissimilarity in the correlation between the distribution of the input RGB image 20 in the vector space and the distribution of the simulated image in the vector space. Figure 7 Detailed description of a specific embodiment in which the first AI model 210 is trained by minimizing the KL-divergence.

[0057] In an embodiment of the present disclosure, the AI ​​model 200 may further include a third AI model for calculating the value of the mutual information loss 28 between the RGB image 20 and the simulated image. The third AI model may include a reconstruction model trained to generate a fake image imitating the RGB image 20 based on the simulated image, and a discriminator model that determines whether the input image is the RGB image 20 or the fake image generated by the reconstruction model. The third AI model may calculate the value of the mutual information loss 28 based on the output values ​​of the reconstruction model and the discriminator model. The first AI model 210 may be trained to minimize the value of the mutual information loss 28 via back propagation of the mutual information loss 28. Reference Figure 8 Detailed description will now be given of a specific embodiment in which the first AI model 210 is trained by applying the value of the mutual information loss 28 calculated by the third AI model.

[0058] Figure 3 is a block diagram illustrating components of the electronic device 100 according to an embodiment of the present disclosure.

[0059] The electronic device 100 may be implemented as a smart phone including a camera system, a tablet personal computer (PC), a laptop computer, a digital camera, an e-book terminal, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, an MP3 player, etc. In an embodiment of the present disclosure, the electronic device 100 may be a home appliance including a camera, such as a smart TV, an air conditioner, a robot vacuum cleaner, or a clothes manager. However, the present disclosure is not limited thereto, and in another embodiment of the present disclosure, the electronic device 100 may be implemented as a wearable device, such as a smart watch, a glasses-shaped augmented reality (AR) device (e.g., AR glasses), or a head mounted device (HMD).

[0060] refer to Figure 3 , the electronic device 100 may include a metalens 110, an image sensor 120, a processor 130, and a memory 140. The image sensor 120, the processor 130, and the memory 140 may be electrically and / or physically connected to each other. Figure 3Only necessary components for describing the operation of the electronic device 100 are shown, and the components included in the electronic device 100 are not limited to Figure 3 In an embodiment of the present disclosure, the electronic device 100 may further include a depth sensor or an audio sensor. When the electronic device 100 is implemented as a mobile device, the electronic device 100 may further include a battery that supplies power to the image sensor 120 and the processor 130.

[0061] In an embodiment of the present disclosure, the electronic device 100 may further include a communication interface ( Figure 4 150). Figure 4 An embodiment in which the electronic device 100 includes the communication interface 150 is described in more detail.

[0062] The superlens 110 is a lens having an optical property of changing the amount of light reflected from an object and penetrating therein and modulating or distorting the phase of the light by refracting or diffracting the light. The superlens 110 may include a supersurface composed of a plurality of posts or pins having a nanometer size. A supersurface is a 3D surface having a pattern composed of a plurality of posts or pins having heights and areas different from each other, and the distances between the plurality of posts or pins may be different. The superlens 110 may have an optical property of encoding light reflected from the object 10 through the supersurface. In an embodiment of the present disclosure, the superlens 110 may change the refractive index of the light according to a pattern composed of a plurality of posts or pins included in the supersurface, thereby causing a phase delay of the light and modulating the phase of the light.

[0063] The image sensor 120 is configured to receive light whose phase is modulated by penetrating the metalens 110, and obtain a coded image by converting the received light into an electrical signal. In an embodiment of the present disclosure, the image sensor 120 may be composed of a CMOS, but is not limited thereto. The light reflected by the object 10 is phase-modulated by the metalens 110 due to the change in the refractive index, and the phase-modulated light is received by a specific pixel in the image sensor 120. The image sensor 120 may obtain a coded image by converting the received light into an electrical signal. The image sensor 120 may provide image data of the obtained coded image to the processor 130.

[0064] Processor 130 may execute one or more instructions of a program stored in memory 140. Processor 130 may be composed of hardware components that perform arithmetic, logic, and input / output (I / O) operations and image processing. Figure 3140 is shown as an element, but is not limited to this. In an embodiment of the present disclosure, the processor 130 may be configured as one or more elements. The processor 130 may be a general-purpose processor such as a CPU, AP, DSP, etc.; a dedicated graphics processor such as a GPU, VPU, etc., or a dedicated AI processor such as an NPU, etc. The processor 130 may control the input data to be processed according to predefined operating rules or the AI ​​model 200 stored in the memory 140. Alternatively, when the processor 130 is a dedicated AI processor, the dedicated AI processor may be designed with a hardware structure dedicated to processing a specific AI model.

[0065] The memory 140 may include at least one type of storage medium, i.e., for example, at least one of a flash memory, a hard disk memory, a multimedia card micro memory, a card-type memory (e.g., a secure digital (SD) card or an extreme digital (XD) memory), a random access memory (RAM), a static RAM (SRAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), a PROM, or an optical disk.

[0066] The memory 140 may store instructions related to operations in which the electronic device 100 obtains a visual recognition result from an encoded image obtained through the metalens 110 and the image sensor 120. In an embodiment of the present disclosure, the memory 140 may store at least one of instructions, algorithms, data structures, program codes, and applications readable by the processor 130. The instructions, algorithms, data structures, and program codes stored in the memory 140 may be implemented in a programming or scripting language such as C, C++, Java, assembler, etc.

[0067] Hereinafter, functions or operations performed by the processor 130 by executing instructions or program codes included in the modules stored in the memory 140 are described.

[0068] The processor 130 may obtain the encoded image through the super lens 110 and the image sensor 120. The processor 130 may input the encoded image to the AI ​​model 200, and obtain a label indicating a recognition result of the encoded image by performing inference using the AI ​​model 200.

[0069] exist Figure 3 In the illustrated embodiment, the AI ​​model 200 may be stored on the electronic device 100 in an on-device manner. In an embodiment of the present disclosure, the memory 140 may store instructions, algorithms, data structures, or program codes constituting the AI ​​model 200.

[0070] The AI ​​model 200 may be a neural network model trained via supervised learning to obtain a simulated image by inputting an RGB image into a model reflecting the optical properties of the super lens 110, and outputting a label indicating the true value of the input RGB image as a recognition result of the obtained simulated image. In an embodiment of the present disclosure, the AI ​​model 200 may be implemented as a CNN model, but is not limited thereto. The AI ​​model 200 may be implemented as, for example, an RNN, an RBM, a DBN, a BRDNN, a DQN, etc. Because the specific method of training the AI ​​model 200 is similar to the reference Figure 2 The method of description is the same, so redundant description is omitted here.

[0071] In an embodiment of the present disclosure, the "label" obtained as a result of inference by the AI ​​model 200 may be classification information of an object included in the input coded image. For example, when the object whose image is captured by the super lens 110 and the image sensor 120 is a cat, the AI ​​model 200 may output a probability value indicating the possibility that the object is classified as a label indicating "cat". However, the present disclosure is not limited thereto, and as a result of visual recognition using the AI ​​model 200, the processor 130 may detect an object or segment an object from the coded image.

[0072] In an embodiment of the present disclosure, the AI ​​model 200 may include a backbone network trained to extract a feature map from an input encoded image and a head network trained to output a label indicating the result of identifying the encoded image from the feature map. The processor 130 may change at least one of the backbone network and the head network based on the purpose or use of the visual task to be identified by using the AI ​​model 200. In an embodiment of the present disclosure, the AI ​​model 200 may include multiple backbone networks and multiple head networks. The processor 130 may select a backbone network optimized according to the purpose or use of the visual task from among the multiple backbone networks, and change the backbone network of the AI ​​model 200 to the selected backbone network. In addition, the processor 130 may select a head network optimized according to the purpose or use of the visual task from among the multiple head networks, and change the head network of the AI ​​model 200 to the selected head network. In an embodiment of the present disclosure, when the head network is changed according to the purpose or use of the visual task, the same backbone network may be applied regardless of the visual task. In this case, the structure of the backbone network may also be shared by the changed head network. Reference Figures 9 to 11 Specific embodiments in which the processor 130 changes at least one of the backbone network and the head network based on the purpose or use of the visual task will be described in detail.

[0073] In an embodiment of the present disclosure, the AI ​​model 200 may further include a metalens profiler model for identifying a replacement or change of the metalens 110 by a user. The processor 130 may identify a replacement or change of the metalens 110 by a user via the metalens profiler model. When a replacement or change of the metalens 110 is identified, the processor 130 may identify the visual task corresponding to the replaced or changed metalens 110, and change the backbone network and the head network of the AI ​​model 200 to a backbone network and a head network of a model that is predetermined to be optimized for the identified visual task. Reference Fig.12 and Fig.13 Detailed description will now be given of a specific embodiment in which the processor 130 changes the backbone network and the head network of the AI ​​model 200 according to a replacement or change of the superlens 110 by a user.

[0074] In an embodiment of the present disclosure, the electronic device 100 may further include a depth sensor or an audio sensor. The processor 130 may obtain depth value information of the object from the depth sensor and obtain an audio signal emitted by the object from the audio sensor. The processor 130 may obtain a visual recognition result of the encoded image by performing inference using the AI ​​model 200. Fig.14 and Fig.15 A specific embodiment in which the processor 130 obtains a visual recognition result of an encoded image by using a depth value and an audio signal will be described in more detail.

[0075] Figure 4 is a block diagram illustrating components of the electronic device 100 and the server 300 according to an embodiment of the present disclosure.

[0076] refer to Figure 4 The electronic device 100 may further include a communication interface 150 for performing data communication with the server 300. The metalens 110, the image sensor 120, and the processor 130 included in the electronic device 100 are respectively connected to Figure 3 The superlens 110, image sensor 120, and processor 130 shown in FIG. 1 are substantially the same, and therefore, redundant descriptions are omitted.

[0077] The communication interface 150 may transmit and receive data to and from the server 300 through a wired or wireless communication network, and process the data. The communication interface 150 may perform data communication with the server 300 by using at least one of a data communication method including, for example, a wired local area network (LAN), a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), an infrared data association (IrDA), Bluetooth low energy (BLE), a near field communication (NFC), wireless broadband Internet (WiBro), world interoperability for microwave access (WiMAX), shared wireless access protocol (SWAP), wireless gigabit alliance (WiGig), and radio frequency (RF) communication. However, the present disclosure is not limited to this, and when the electronic device 100 is implemented as a mobile device, the communication interface 150 can send and receive data to and from the server 300 through a network that complies with mobile communication standards such as code division multiple access (CDMA), wideband CDMA (WCDMA), third generation (3G), fourth generation (4G), fifth generation (5G) and / or a communication method using millimeter wave (mmWave).

[0078] In an embodiment of the present disclosure, the communication interface 150 may transmit the encoded image to the server 300 according to the control performed by the processor 130, and receive a label indicating the recognition result of the encoded image performed via the AI ​​model 200 from the server 300. The communication interface 150 may provide information about the label received from the server 300 to the processor 130.

[0079] The server 300 may include a communication interface 310 for communicating with the electronic device 100 , a memory 330 for storing at least one instruction or program code, and a processor 320 configured to execute at least one instruction or program code stored in the memory 330 .

[0080] The memory 330 of the server 300 may store the trained AI model 200, which may be stored in the memory 330 of the server 300. Because the AI ​​model 200 stored in the server 300 is Figure 1 and Figure 2 The AI ​​model 200 shown and described in the figure is the same, so redundant description is omitted here. The server 300 can obtain image data of the encoded image from the electronic device 100 via the communication interface 310. The processor 320 of the server 300 can input the encoded image to the AI ​​model 200, and obtain a label indicating the recognition result of the obtained encoded image by performing inference using the AI ​​model 200. The processor 320 can control the communication interface 310 to send the data of the label to the electronic device 100.

[0081] Typically, the electronic device 100 may have limited memory ( Figure 3 140) storage capacity, processor ( Figure 3 130), the computing processing speed of the training data set, the ability to collect training data sets, etc. Therefore, after performing operations that require the storage of a large amount of data and a large amount of calculation, the server 300 can send the necessary data (e.g., image data of the encoded image) and / or data corresponding to the visual recognition result (e.g., a label) to the electronic device 100 via the communication network. Then, the electronic device 100 can receive and use the data indicating the visual recognition result from the server 300 without the need for a processor with a large capacity of memory and the ability to perform fast calculations, thereby reducing the processing time required for visual recognition of the encoded image and improving the accuracy of visual recognition.

[0082] Figure 5 is a flowchart of a method of obtaining a visual recognition result from an image obtained using a super lens, performed by the electronic device 100 according to an embodiment of the present disclosure.

[0083] In operation S510, the electronic device 100 obtains a coded image by receiving light whose phase is modulated by passing through the metalens and converting the received light into an electrical signal. Light reflected by an object to be captured may be reflected by the metalens ( Figure 1 , Figure 3 and Figure 4 The phase modulated light is transmitted to the image sensor ( Figure 1 , Figure 3 and Figure 4 120) is received by a specific pixel in the image sensor, and the electronic device 100 can obtain a coded image by converting the light received through the image sensor 120 into an electrical signal.

[0084] The “encoded image” is an image obtained by phase-modulating light reflected from an object by the metalens 110 and receiving the modulated light by the image sensor 120. In an embodiment of the present disclosure, unlike an RGB image, the encoded image may be an image that is out of focus or distorted or modulated so that the human eye cannot recognize the shape of the object.

[0085] In operation S520, the electronic device 100 inputs the encoded image to the AI ​​model, and obtains a label indicating a recognition result of the encoded image by performing inference using the AI ​​model. The "AI model" may be a neural network model trained via supervised learning to obtain a simulated image by inputting an RGB image into a model that reflects the optical properties of the superlens 110, and outputs a label indicating the true value of the input RGB image as a recognition result of the obtained simulated image. In an embodiment of the present disclosure, the AI ​​model may be implemented as a CNN model, but is not limited thereto. The AI ​​model may be implemented as, for example, RNN, RBM, DBN, BRDNN, DQN, etc. Because the method of training the AI ​​model is similar to the reference method, Figure 2 The method of description is the same, so redundant description is omitted here.

[0086] The "label" obtained as a result of the electronic device 100 performing inference using the AI ​​model may be classification information of an object included in the encoded image input to the AI ​​model. For example, when the object whose image is captured by the super lens 110 and the image sensor 120 is a cat, the AI ​​model 200 may output a probability value indicating the likelihood that the object is classified as a label indicating "cat". However, the present disclosure is not limited thereto, and therefore, the electronic device 100 may output an object detected in the encoded image, or obtain a segmentation result of the object, as a result of inference using the AI ​​model.

[0087] Figure 6 is a perspective view showing the shape of a surface pattern of a superlens 110 according to an embodiment of the present disclosure.

[0088] refer to Figure 6 , the metalens 110 may include a metasurface composed of a plurality of pillars 112 having a nanometer size. The plurality of pillars 112 protrude upward from a substrate 114 and form a 3D surface pattern. Each of the plurality of pillars 112 may be formed to have a different height and a different area. Figure 6 In the illustrated embodiment, the plurality of pillars 112 are shown as having a cylindrical shape, but the shape of the plurality of pillars 112 is not limited to a cylindrical shape. In another embodiment of the present disclosure, the shape of the plurality of pillars 112 may be a cylinder, a cone, a quadrangular pillar, or a combination thereof. The plurality of pillars 112 formed on the metasurface may also be replaced with an uneven or fin-shaped structure.

[0089] Can be based on training AI models ( Figure 1 and Figure 2The height, area, and volume of each of the plurality of pillars 112 included in the metalens 110 and the distance between the plurality of pillars 112 are determined based on the model parameter values ​​of the result of the AI ​​model 200. In an embodiment of the present disclosure, the AI ​​model 200 may include a differentiable mathematical model, and the model parameters may be updated and optimized via supervised learning using a pair of input images and true values, the differentiable mathematical model representing the optical properties of the metalens 110 and the image sensor ( Figure 1 and Figure 3 The PSF of the sensor noise of 120) is mathematically modeled. The height, area, and volume of the plurality of pillars 112 in the metasurface and the distance between the plurality of pillars 112 can be determined based on the phase function, which simulates the phase modulation caused by the optical properties of the metalens 110 among the model parameters optimized by the training of the AI ​​model 200. In an embodiment of the present disclosure, the phase function can be a formula defining the coefficients of a polynomial that mathematically represents the degree of phase modulation according to the distance from the center of the metalens 110.

[0090] Figure 7 is a diagram illustrating a method of training an AI model 200 according to an embodiment of the present disclosure.

[0091] refer to Figure 7 , the AI ​​model 200 may be a model trained via supervised learning to input a plurality of RGB images 700-1 to 700-n and output a label 740 corresponding to the true values ​​702-1 to 702-n paired with the plurality of RGB images 700-1 to 700-n. In an embodiment of the present disclosure, the AI ​​model 200 may be an end-to-end neural network model trained to minimize a loss 750, which is the difference between the label 740 predicted from the plurality of RGB images 700-1 to 700-n and the label corresponding to the true value 702-1 to 702-n.

[0092] The AI ​​model 200 may include a first AI model 210 and a second AI model 220. Figure 2 The method of description is the same, so redundant description is omitted here.

[0093] The first AI model 210 may be trained to minimize information indicating similarity between the input multiple RGB images 700-1 to 700-n and the output multiple simulated images 720-1 to 720-n. In an embodiment of the present disclosure, the first AI model 210 may be trained by calculating a mutual information loss 760 and performing back propagation using the value of the mutual information loss 760, where the mutual information loss 760 is a mathematical quantification of the correlation between the multiple RGB images 700-1 to 700-n and the multiple simulated images 720-1 to 720-n, so that the mutual information representing the correlation between them is minimized. For example, during the training process of the first AI model 210, the multiple weights included in the first AI model 210 may be updated by applying the value of the mutual information loss 760 to the back propagation of the first AI model 210.

[0094] In the present disclosure, "correlation" indicates the similarity of intensity values ​​constituting images between the plurality of RGB images 700-1 to 700-n and the plurality of simulated images 720-1 to 720-n, and may include information on whether a human-readable image similar to the plurality of RGB images 700-1 to 700-n can be reconstructed from the plurality of simulated images 720-1 to 720-n via a reconstruction model composed of a CNN model or the like. In an embodiment of the present disclosure, the electronic device 100 may obtain a value of the mutual information loss 760 by calculating a KL-divergence representing the correlation between the plurality of RGB images 700-1 to 700-n and the plurality of simulated images 720-1 to 720-n in a vector space. The KL-divergence may be calculated by using the following equation 1.

[0095] [Equation 1]

[0096]

[0097] In Equation 1, P(x) may be a probability mass function representing a distribution 710 of a plurality of RGB images 700-1 to 700-n in an n-dimensional vector space for training the first AI model 210, and Q(x) may be a probability mass function representing a distribution 730 of a plurality of simulated images 720-1 to 720-n in an n-dimensional vector space output by the first AI model 210. The cross entropy representing the correlation between the distributions 710 and 730 of the plurality of RGB images 700-1 to 700-n and the plurality of simulated images 720-1 to 720-n in an n-dimensional vector space may be calculated by using Equation 2 below.

[0098] [Equation 2]

[0099] CE(P(x), Q(x))=-ΣP(x)logQ(x)

[0100] Referring to Equations 1 and 2, when the KL-divergence is minimized, the cross entropy is maximized. An increase in cross entropy means that the dissimilarity between the predicted output data and the input data increases during the training process of the AI ​​model. The first AI model 210 can be trained to update and optimize model parameters (e.g., weights between layers) by calculating the KL-divergence, obtain the value of the mutual information loss 760 based on the calculated KL-divergence, and perform back propagation using the mutual information loss 760, thereby minimizing the KL-divergence. Because the cross entropy is maximized when the KL-divergence is minimized, the dissimilarity between the multiple RGB images 700-1 to 700-n and the multiple simulated images 720-1 to 720-n can be maximized and the correlation can be minimized during the training process of the first AI model 210.

[0101] Figure 8 is a diagram illustrating a method of training an AI model 200 according to an embodiment of the present disclosure.

[0102] refer to Figure 8 , the AI ​​model 200 may be a model trained using supervised learning to input an RGB image 800 and output a label 820 corresponding to a true value 802 paired with the input RGB image 800. In an embodiment of the present disclosure, the AI ​​model 200 may be an end-to-end neural network model trained to minimize a loss 830, which is the difference between a label 820 predicted from the input RGB image 800 and a label corresponding to the true value 802.

[0103] exist Figure 8 In the illustrated embodiment, the AI ​​model 200 may include a first AI model 210, a second AI model 220, and a third AI model 230. Because the specific method of training the first AI model 210 and the second AI model 220 is similar to the reference Figure 2 The method of description is the same, so redundant description is omitted here.

[0104] The third AI model 230 may be a neural network model trained to calculate mutual information loss 840, which is a mathematical quantification of the correlation of similarities between the input RGB image 800 and the simulated image 810 output from the first AI model 210. For example, the third AI model 230 may be implemented as a generative adversarial network (GAN), but is not limited thereto. The third AI model 230 may include a reconstruction model 232 and a discriminator model 234.

[0105] The reconstruction model 232 may be a neural network model trained to generate a fake image that imitates the RGB image 800 based on the input simulated image 810. The reconstruction model 232 may provide the generated fake image to the discriminator model 234.

[0106] The discriminator model 234 may be a neural network model trained to determine whether an input image is the RGB image 800 or a fake image generated by the reconstruction model 232 .

[0107] The third AI model 230 may calculate a value of the mutual information loss 840 based on the output values ​​of the reconstruction model 232 and the discriminator model 234. The value of the mutual information loss 840 may be calculated by using Equation 3 below.

[0108] [Equation 3]

[0109] L M.I. (X)=E[logD(X)]+E[log(1-D(R(X))]

[0110] Referring to Equation 3, L representing the value of the mutual information loss 840 can be calculated by taking the expected value (expectation (E) operation) of the value obtained by taking the logarithmic function of R(x) as the output value of the reconstruction model 232 and D(x) as the output value of the discriminator model 234. M.I ..

[0111] The first AI model 210 may be trained so that the value of the mutual information loss 840 is minimized by applying the back propagation of the mutual information loss 840 calculated by the third AI model 230. For example, during the training process of the first AI model 210, a plurality of weights included in the first AI model 210 may be updated by applying the value of the mutual information loss 840 to the back propagation of the first AI model 210.

[0112] When passing through the superlens ( Figure 1 and Figure 3 110) and image sensor ( Figure 1 and Figure 3 When the coded image obtained by the image sensor 120 is similar to the RGB image of the object captured using a general lens, the coded image may be hijacked from the image sensor 120, or hacked and leaked to the outside, so that the RGB image can be reconstructed from the coded image. In this case, there is a problem that personal information may be leaked to the outside and privacy may be violated.

[0113] exist Figure 7 and Figure 8 In the embodiment shown in FIG. 1 , because the first AI model 210 is trained to minimize the representation of the input RGB image (or multiple) ( Figure 7 700-1 to 700-n and Figure 8 800) and the simulated image (or images) obtained by simulating the optical properties of the metalens 110 ( Figure 7 720-1 to 720-n and Figure 8810), so even when the encoded image is leaked to the outside during the inference process using the AI ​​model 200, the possibility of reconstructing a human-readable image from the encoded image is low. Therefore, according to an embodiment of the present disclosure, the electronic device 100 can prevent the leakage of personal information, protect privacy, and enhance security during the inference process using the AI ​​model 200.

[0114] Fig. 9 is a flowchart of a method performed by the electronic device 100 to obtain a visual recognition result by using an AI model that changes according to a purpose or use of a visual task according to an embodiment of the present disclosure.

[0115] Fig. 9 Operations S910 and S920 are shown as Figure 5 The detailed operation of operation S520 is shown in FIG. Figure 5 After the operation S510 shown in Fig. 9 Operation S910.

[0116] Fig.10 is a diagram illustrating an operation in which the electronic device 100 changes the configuration of the second AI model 220 based on the purpose or use of a visual task according to an embodiment of the present disclosure.

[0117] In the following, refer to Fig. 9 and Fig.10 1 and 2 to describe an operation in which the electronic device 100 changes the configuration of the second AI model 220 according to the purpose or use of the visual task.

[0118] exist Fig. 9 In operation S910, the electronic device 100 changes at least one of the backbone network and the head network of the AI ​​model based on the purpose or use of the visual task. Fig. 9 For reference Fig.10 , the second AI model 220 is a visual recognition neural network model that is trained to output a visual recognition result of an encoded image when an encoded image is input. The second AI model 220 may include a backbone network 222 and a head network 224.

[0119] The backbone network 222 is a neural network model that is trained to extract feature values ​​from an input encoded image and obtain a feature map by using the extracted feature values. The backbone network 222 can be implemented as, for example, GoogleNet, ResNet, VGG, DarkNet, ImageNet, or U-Net, but is not limited thereto.

[0120] The head network 224 is a neural network model trained to output a label indicating a result of identifying an encoded image according to a feature map output from the backbone network 222. The head network 224 may be implemented as, for example, embedding & softmax, pixel-wise softmax, YOLO, or YOLO & embedding, but is not limited thereto.

[0121] The processor of the electronic device 100 ( Figure 3 130) can change at least one of the backbone network 222 and the head network 224 of the AI ​​model based on the purpose or use of the visual task. In the present disclosure, a "visual task" refers to a task that outputs a visual recognition result from an input image. For example, a visual task may include, but is not limited to, object detection, classification, segmentation, facial recognition, pose estimation, depth estimation, etc.

[0122] Multiple backbone networks 222-1 to 222-n may be models optimized and trained for different visual tasks, respectively. Here, "optimized model" refers to a neural network model that outputs prediction results with high precision and requires short processing time depending on the amount of training data and the purpose and use of the visual task. For example, the first backbone network 222-1 may be a ResNet optimized for "detection", the second backbone network 222-2 may be a VGG optimized for "classification", and the third backbone network 222-3 may be a U-Net optimized for "segmentation", but they are not limited thereto. In an embodiment of the present disclosure, the processor 130 may select a backbone network optimized for the purpose or use of a visual task from among multiple backbone networks 222-1 to 222-n. For example, when the purpose or use of a visual task is "detection" of an object, the processor 130 may select a first backbone network 222-1 optimized for "detection" from among multiple backbone networks 222-1 to 222-n. The processor 130 may change the backbone network 222 to the first backbone network 222 - 1 by replacing the existing backbone network 222 with the selected first backbone network 222 - 1 .

[0123] Multiple head networks 224-1 to 224-n may be models optimized and trained for different visual tasks, respectively. For example, the first head network 224-1 may be a pixel-by-pixel softmax optimized for "segmentation", the second head network 224-2 may be a YOLO optimized for "detection", and the third head network 224-3 may be a YOLO & embedding model optimized for "face recognition", but is not limited thereto. In an embodiment of the present disclosure, the processor 130 may select a head network optimized for the purpose or use of a visual task from among the multiple head networks 224-1 to 224-n. For example, when the purpose or use of a visual task is "detection" of an object, the processor 130 may select a second backbone network 224-2 optimized for "detection" from among the multiple head networks 224-1 to 224-n. The processor 130 may change the head network 224 to a second head network 224-2 by replacing the existing head network 224 with the selected second head network 224-2.

[0124] Return to reference Fig. 9 In operation S920, the electronic device 100 outputs a label indicating a recognition result of the encoded image by inputting the encoded image to an AI model in which at least one of the backbone network and the head network has been changed. Fig. 9 For reference Fig.10 , the second AI model 220 may include a first backbone network 222-1 and a second head network 224-2. When the encoded image is input to the second AI model 220, a feature map may be extracted from the encoded image by the first backbone network 222-1. The extracted feature map may be input to the second head network 224-2, and a label indicating a recognition result of the encoded image may be output by the second head network 224-2. Fig.10 In the illustrated embodiment, because the first backbone network 222-1 and the second head network 224-2 are both neural network models optimized for "detection", the output label may be a vector value indicating the detection result of the object.

[0125] Fig.11 is a diagram illustrating an operation in which the electronic device 100 changes the configuration of an AI model based on the purpose or use of a visual task according to an embodiment of the present disclosure.

[0126] refer to Fig.11 , the second AI model 220 is a visual recognition neural network model that is trained to output a visual recognition result of a coded image when the coded image is input. The second AI model 220 may include a first backbone network 222-1 and a head network 224. Because the first backbone network 222-1 and the head network 224 are similar to the reference Fig.10 The descriptions are the same, so redundant descriptions are omitted here.

[0127] exist Fig.11 In the illustrated embodiment, the backbone network of the second AI model 220 is determined to be the first backbone network 222-1, and may not be changed according to the purpose or use of the visual task. When the purpose or use of the visual task is changed, the head network 224 may be changed. Multiple head networks 224-1 to 224-n may be models optimized and trained for different visual tasks, respectively. For example, the first head network 224-1 may be a pixel-by-pixel softmax optimized for "segmentation", the second head network 224-2 may be a YOLO optimized for "detection", and the third head network 224-3 may be a YOLO & embedding model optimized for "face recognition". However, the present disclosure is not limited to this. In an embodiment of the present disclosure, the processor ( Figure 3 The processor 130 may select a head network optimized for the purpose or use of the visual task from among the plurality of head networks 224-1 to 224-n, and replace the existing head network 224 with the selected head network. For example, when the purpose or use of the visual task is “detection” of an object, the processor 130 may select a second head network 224-2 optimized for “detection” from among the plurality of head networks 224-1 to 224-n, and replace the existing head network 224 with the selected second head network 224-2.

[0128] When the head network 224 is changed according to the purpose or use of the visual task, the model parameters of the first backbone network 222-1 may remain unchanged regardless of the changed head network 224. For example, the weights of the first backbone network 222-1 may be shared (shared weights) for the head network 224 that is changed according to the purpose or use of the visual task.

[0129] However, the present disclosure is not limited thereto. In an embodiment of the present disclosure, when the head network 224 is changed, the model parameters of the first backbone network 222-1 may be changed according to the visual task of the changed head network 224. For example, when the purpose or use of the visual task is "face recognition", the processor 130 may change the head network 224 to the third head network 224-3 and replace the weights of the first backbone network 222-1 with weights optimized for face recognition.

[0130] exist Figures 9 to 11 In the illustrated embodiment, the electronic device 100 changes at least one of the backbone network 222 and the head network 224 according to the purpose or use of a visual task such as "object detection", "classification", "segmentation", "face recognition" or "depth estimation", thereby providing a technical effect of improving the accuracy of visual recognition results and reducing processing time.

[0131] Fig.12is a flowchart of a method of changing the configuration of an AI model in response to replacement or change of a super lens being recognized, performed by the electronic device 100 according to an embodiment of the present disclosure.

[0132] Can be Figure 5 The operation S510 shown in FIG. 1 is performed after the operation S510 is performed. Fig.12 Operations corresponding to operations S1210 to S1230 shown in FIG. Fig.12 After the operation corresponding to operation S1230 shown in FIG. 1 is performed, the Figure 5 Operation S520.

[0133] Fig.13 is a diagram illustrating an operation in which the electronic device 100 changes the configuration of the second AI model 220 in response to replacement or change of the super lens 110 being recognized according to an embodiment of the present disclosure.

[0134] In the following, refer to Fig.12 and Fig.13 hereinafter, an operation is described in which the electronic device 100 changes the configuration of the second AI model 220 in response to the replacement or change of the metalens 110 being recognized.

[0135] exist Fig.12 In operation S1210, the electronic device 100 recognizes replacement or change of the metalens.

[0136] Combination Fig.12 For reference Fig.13 , depending on the purpose or use of visual recognition, the user can replace or change the superlens 110. Depending on the purpose or use of visual recognition, the superlens may have specific optical properties. A coded image is obtained due to the optical properties of hardware such as a surface pattern of the superlens, and the accuracy of the results of visual recognition such as object detection, classification, segmentation, facial identification, and situation recognition can be determined based on the characteristics of the coded image. For example, the first superlens 110-1 may be a lens optimized for obtaining a coded image for “facial recognition”, and the second superlens 110-2 may be a lens optimized for obtaining a coded image for “situation recognition”. In Fig.13 In the embodiment shown in FIG. 1 , the user can replace the first super lens 110 - 1 with the second super lens 110 - 2 according to the change of the visual recognition purpose.

[0137] The electronic device 100 can recognize the replacement of the metalens by the user. In the embodiment of the present disclosure, the electronic device 100 may further include a metalens profiler model 160, and the processor ( Figure 3130) can identify the replacement of the metalens from the first metalens 110-1 to the second metalens 110-2 by using the metalens profiler model 160. When the image sensor 120 is receiving light reflected from the object 1300 and then phase-modulated by penetrating the metalens 110-1 and the metalens is replaced with the second metalens 110-2 by user input, the image sensor 120 can receive light reflected from the object 1320 and then phase-modulated by penetrating the second metalens 110-2. In an embodiment of the present disclosure, the metalens profiler model 160 may be a neural network model trained through supervised learning by applying an encoded image obtained by the image sensor 120 as an input and applying a vector value representing a specific metalens as a true value, the image sensor 120 receiving the light whose phase is modulated by the specific metalens. However, the present disclosure is not limited thereto, and in another embodiment of the present disclosure, the metalens may have a coding structure with a unique pattern, and the metalens profiler model 160 may be implemented as a classifier model that identifies a specific metalens by the coding structure. When a replacement for the metalens is identified, the metalens profiler model 160 may provide a signal to the processor 130 indicating that a replacement for the metalens has been identified.

[0138] Return to reference Fig.12 In operation S1220, the electronic device 100 identifies a visual task corresponding to the replaced or changed metalens. Fig.12 For reference Fig.13 , based on the identification of the replacement of the superlens, the processor 130 can identify a visual task optimized for the replaced superlens. Here, "a visual task optimized for a superlens" can refer to a visual task for which, when an encoded image obtained by receiving light whose phase is modulated by penetrating a specific superlens by the image sensor 120 is input to the second AI model 220, the output visual recognition result is highly accurate and requires a short processing time. For example, the first superlens 110-1 can be a superlens optimized for "face recognition" among visual tasks, and the second superlens 110-2 can be a superlens optimized for "situation recognition".

[0139] Return to reference Fig.12 In operation S1230, the electronic device 100 changes the backbone network and the head network of the AI ​​model to the backbone network and the head network optimized for the identified visual task. Fig.13 , the second AI model 220 may include a first backbone network 222-1 and a first head network 224-1. Because the detailed description of the backbone network and the head network is similar to that of reference Figures 9 to 11The description is the same, so redundant description is omitted here. In an embodiment of the present disclosure, the first backbone network 222-1 and the first head network 224-1 may be neural network models optimized for "face recognition" among visual tasks. For example, the first backbone network 222-1 may be GoogleNet, and the first head network 224-1 may be YOLO, but is not limited thereto. The processor 130 may identify the first backbone network 222-1 and the first head network 224-1 included in the second AI model 220. In an embodiment of the present disclosure, the processor 130 may identify which visual task the identified first backbone network 222-1 and the first head network 224-1 are optimized for. The processor 130 may replace the first backbone network 222-1 and the first head network 224-1 with a second backbone network 222-2 and a second head network 224-2 respectively optimized for the visual task "context recognition" corresponding to the second super lens 110-2 replaced by the user. For example, the second backbone network 222-2 may be ResNet, and the second head network 224-2 may be SoftMax, but they are not limited thereto.

[0140] Return to reference Fig.12 , the electronic device 100 can input the encoded image to the changed second AI model 220, and obtain a visual recognition result by performing inference using the second AI model 220. Fig.12 For reference Fig.13 , when the second AI model 220 includes the first backbone network 222-1 and the first head network 224-1, when the encoded image is input to the second AI model 220, a facial recognition result 1310 in which the face is marked with a bounding box can be obtained as an inference result. When the first super lens 110-1 is replaced by the second super lens 110-2 by the user, the second AI model 220 includes the second backbone network 222-2 and the second head network 224-2, and when the encoded image is input to the second AI model 220, a label indicating a specific situation (e.g., "fire" or "danger") can be obtained as an inference result.

[0141] exist Fig.12 and Fig.13 In the illustrated embodiment, when the metalens is changed or replaced by the user, the electronic device 100 automatically selects a backbone network and a head network optimized for the visual task corresponding to the changed or replaced metalens, and modifies the AI ​​model by replacing the existing backbone network and the head network with the selected backbone network and the head network ( Fig.13 The second AI model 220 in the image recognition system is used to provide a technical effect of improving the accuracy of visual recognition results and shortening processing time.

[0142] Fig.14is a diagram illustrating an operation in which the electronic device 100 obtains a visual recognition result by using the depth sensor 170 and the audio sensor 180 and a camera system including the super lens 110 according to an embodiment of the present disclosure.

[0143] refer to Fig.14 The electronic device 100 may further include a depth sensor 170, an audio sensor 180, and a multi-mode transformer 190, as well as a super lens 110, an image sensor 120, a processor ( Figure 3 130) and memory ( Figure 3 140). Because the metalens 110 and the image sensor 120 are aligned with the reference Figure 1 and Figure 3 Those described are the same, so redundant description is omitted herein.

[0144] The depth sensor 170 is a sensor configured to obtain depth information about the object 1400. Here, "depth information" means information about the distance from the depth sensor 170 to the specific object 1400. In an embodiment of the present disclosure, the depth sensor 170 may be configured as a time-of-flight (ToF) sensor that obtains depth value information by emitting light toward the object 1400 using a light source and based on the time taken for the emitted light to be reflected from the object 1400 and received via a light receiving sensor of the ToF sensor. However, the present disclosure is not limited thereto, and the depth sensor 170 may be configured as a sensor that obtains depth information by using at least one of a structured light method and a stereoscopic image method.

[0145] The audio sensor 180 is a sensor that receives a sound emitted by the object 1400 and converts the received sound into an audio signal, which is an electrical signal. In an embodiment of the present disclosure, the audio sensor 180 may include a microphone.

[0146] The multi-model converter 190 may obtain depth value information of the object 1400 from the depth sensor 170 and obtain an audio signal from the audio sensor 180. The multi-model converter 190 may convert the depth value information and the audio signal into an n-dimensional vector value by performing embedding on the depth value information and the audio signal, respectively. The multi-model converter 190 may input the embedding vector to the AI ​​model 200.

[0147] The AI ​​model 200 may receive the encoded image from the image sensor 120, receive the embedding vector from the multi-model transformer 190, and output a visual recognition result based on the encoded image and the embedding vector. Fig.14 In the illustrated embodiment, the AI ​​model 200 may output a label value 1410 (eg, “bird”) based on the recognition result of the object 1400 .

[0148] Fig.15 is a diagram illustrating an operation in which the electronic device 100 obtains a visual recognition result by using the depth sensor 170 and the audio sensor 180 and a camera system including the super lens 110 according to an embodiment of the present disclosure.

[0149] refer to Fig.15 The electronic device 100 may include a metalens 110, an image sensor 120, a processor ( Figure 3 130), memory ( Figure 3 140), depth sensor 170, audio sensor 180, AI model 200, depth-based visual recognition model 240, and sound-based visual recognition model 250. Because the depth sensor 170 and the audio sensor 180 are similar to the reference Fig.14 The depth sensor 170 and the audio sensor 180 described are the same, so redundant descriptions are omitted here.

[0150] The depth-based visual recognition model (or depth-based visual algorithm) 240 is a neural network model that is trained to output a label value corresponding to the recognition result of the depth value when the depth value is input. In an embodiment of the present disclosure, the depth-based visual recognition model 240 may be a neural network model trained via supervised learning by applying a vector obtained by embedding the depth value as input and applying a label corresponding to the visual recognition result as a true value. Fig.15 In the illustrated embodiment, when the depth value is input from the depth sensor 170, the depth-based visual recognition model 240 can output a label value (e.g., "fish") according to the inferred visual recognition result. The depth-based visual recognition model 240 can provide the output label value to the output aggregator model 260.

[0151] The sound-based visual recognition model (sound-based visual algorithm) 250 is a neural network model that is trained to output a label value corresponding to the recognition result of the audio signal when the audio signal is input. In an embodiment of the present disclosure, the sound-based visual recognition model 250 may be a neural network model trained via supervised learning by applying a vector obtained by embedding the audio signal as input and applying a label corresponding to the visual recognition result of the audio signal as a true value. Fig.15 In the illustrated embodiment, when an audio signal is input from the audio sensor 180, the sound-based visual recognition model 250 can output a label value (e.g., "Twitting") according to the inferred visual recognition result. The sound-based visual recognition model 250 can provide the output label value to the output aggregator model 260.

[0152] The output aggregator model 260 is a model trained to output a final recognition result about the object 1500 based on the input data from the AI ​​model 200, the depth-based visual recognition model 240, and the sound-based visual recognition model 250. In an embodiment of the present disclosure, the output aggregator model 260 may be a model trained in a rule-based manner. For example, the output aggregator model 260 may include softmax. However, the present disclosure is not limited to this, and the output aggregator model 260 may be a neural network model that is trained to convert the input data from the AI ​​model 200, the depth-based visual recognition model 240, and the sound-based visual recognition model 250 into feature vector values ​​by performing embedding on the input data, and output true values ​​from the feature vector values. Fig.15 In the illustrated embodiment, the output aggregator model 260 can receive recognition results of the object 1500 (e.g., "bird" and "yellow") from the AI ​​model 200, receive recognition results of the object 1500 (e.g., "fish") from the depth-based visual recognition model 240, receive recognition results of the object 1500 (e.g., "tweet") from the sound-based visual recognition model 250, and output a label value 1510 corresponding to the final recognition result regarding the object 1500 (e.g., "bird").

[0153] exist Fig.14 and Fig.15 In the embodiment shown in , the electronic device 100 outputs visual recognition results about objects 1400 and 1500 by using depth value information and audio signals obtained by the depth sensor 170 and the audio sensor 180, respectively, and a visual recognition result obtained by inputting an encoded image obtained via the super lens 110 and the image sensor 120 into the AI ​​model 200, thereby improving the accuracy and performance of visual recognition compared to a system including only the super lens 110, the image sensor 120 and the AI ​​model 200.

[0154] The present disclosure provides an electronic device 100 for obtaining a recognition result of an object from an image obtained through a super lens 110. According to an embodiment of the present disclosure, the electronic device 100 may include a super lens 110 having a pattern formed on its surface and composed of a plurality of columns or pins having shapes, heights, and areas different from each other, and having an optical property of modulating the phase of light reflected from an object through a pattern on the surface. According to an embodiment of the present disclosure, the electronic device 100 may include an image sensor 120 configured to obtain a coded image by receiving light reflected from an object and phase-modulated by penetrating the super lens 110 and converting the received light into an electrical signal. According to an embodiment of the present disclosure, the electronic device 100 may include at least one processor 130 configured to input the coded image to an AI model 200, and obtain a label indicating a result of recognizing an object by performing inference using the AI ​​model 200. In an embodiment of the present disclosure, the AI ​​model 200 may be a neural network model trained to obtain a simulated image by inputting a red, green, and blue (RGB) image to a model reflecting the optical properties of the super lens 110, and output a label indicating the true value of the input RGB image as a recognition result of the obtained simulated image. In an embodiment of the present disclosure, the AI ​​model 200 may be trained to minimize information indicating the similarity between the RGB image and the simulated image.

[0155] In an embodiment of the present disclosure, the shapes, heights, and areas of the plurality of posts or pins forming a pattern on the surface of the superlens 110 may be formed based on mathematical modeling parameters according to the result of training the AI ​​model 200 .

[0156] In an embodiment of the present disclosure, the AI ​​model 200 may include a first AI model 210 and a second AI model 220, the first AI model 210 being trained to output a simulated image from an RGB image by performing a convolution of the RGB image with a PSF that mathematically models the optical properties of the super lens 110, and the second AI model 220 being trained to output a label indicating a true value of the RGB image from the simulated image. At least one processor 130 may be configured to input the encoded image to the second AI model 220, and obtain a label corresponding to a recognition result of the encoded image by performing inference by the second AI model 220.

[0157] In an embodiment of the present disclosure, the first AI model 210 may be trained by updating weights via back propagation, which applies mutual information representing the similarity between the RGB image and the simulated image as a loss so that the mutual information is minimized.

[0158] In an embodiment of the present disclosure, the first AI model 210 may be trained to minimize the loss by minimizing the KL-divergence representing the correlation between the distribution of the RGB image and the simulated image in the vector space.

[0159] In an embodiment of the present disclosure, the AI ​​model 200 may further include a third AI model 230 configured to calculate a loss representing the dissimilarity between the RGB image and the simulated image. The third AI model 230 may include a reconstruction model trained to generate a fake image imitating the RGB image based on the simulated image, and a discriminator model trained to determine whether the input image is an RGB image or a fake image generated by the reconstruction model.

[0160] In an embodiment of the present disclosure, the AI ​​model 200 may include a backbone network trained to extract a feature map from an input encoded image and a head network trained to output a label indicating a result of identifying the encoded image from the feature map. At least one processor 130 may change at least one of the backbone network and the head network based on the purpose or use of the visual task to be identified by using the AI ​​model 200.

[0161] In an embodiment of the present disclosure, in the case where the head network is changed according to the purpose or use of a visual task, the structure of the applied backbone network may be the same regardless of the changed head network.

[0162] In an embodiment of the present disclosure, the AI ​​model 200 may further include a metalens profiler model configured to identify a replacement or change of the metalens 110 by a user. In the event that the replacement or change of the metalens 110 is identified via the metalens profiler model, the at least one processor 130 may be configured to identify a visual task corresponding to the replaced or changed metalens 110. The at least one processor 130 may be configured to change the backbone network and the head network of the AI ​​model 200 to a backbone network and a head network of a model predetermined to be optimized for the identified visual task.

[0163] In an embodiment of the present disclosure, the electronic device 100 may further include a depth sensor configured to measure a depth value of an object, and an audio sensor configured to obtain an audio signal from the object. The at least one processor 130 may be configured to input the depth value of the object measured by the depth sensor and the audio signal obtained by the audio sensor to the AI ​​model 200, and obtain a label corresponding to the result of recognizing the object by performing inference using the AI ​​model 200.

[0164] The present disclosure provides a method for recognizing an object from an image obtained through a super lens 110, performed by an electronic device 100. The method may include obtaining a coded image by receiving light reflected from an object and phase-modulated by penetrating the super lens 110 and converting the received light into an electrical signal (S510). The method may include inputting the coded image to an AI model 200, and obtaining a label indicating a result of recognizing the object by performing inference using the AI ​​model 200 (S520).

[0165] In an embodiment of the present disclosure, a pattern consisting of a plurality of pillars or pins having shapes, heights, and areas different from each other can be formed on the surface of the superlens 110, and the shapes, heights, and areas of the plurality of pillars or pins forming the pattern on the surface of the superlens 110 can be formed based on mathematical modeling parameters according to the result of training the AI ​​model 200.

[0166] In an embodiment of the present disclosure, the AI ​​model 200 may include a first AI model 210 and a second AI model 220, the first AI model 210 being trained to output a simulated image from an RGB image by performing a convolution of the RGB image with a PSF that mathematically models the optical properties of the super lens 110, and the second AI model 220 being trained to output a label indicating the true value of the RGB image from the simulated image. When obtaining a label (S520), the electronic device 100 may be configured to input the encoded image to the second AI model 220, and obtain a label corresponding to the recognition result of the encoded image by inference by the second AI model 220.

[0167] In an embodiment of the present disclosure, the first AI model 210 may be trained by updating weights via back propagation, which applies mutual information representing the similarity between the RGB image and the simulated image as a loss so that the mutual information is minimized.

[0168] In an embodiment of the present disclosure, the first AI model 210 may be trained to minimize the loss by minimizing the KL-divergence representing the correlation between the distribution of the RGB image and the simulated image in the vector space.

[0169] In an embodiment of the present disclosure, the AI ​​model 200 may further include a third AI model 230 configured to calculate a loss representing the dissimilarity between the RGB image and the simulated image. The third AI model 230 may include a reconstruction model trained to generate a fake image imitating the RGB image based on the simulated image, and a discriminator model trained to determine whether the input image is an RGB image or a fake image generated by the reconstruction model.

[0170] In an embodiment of the present disclosure, the AI ​​model 200 may include a backbone network trained to extract a feature map from an input encoded image and a head network trained to output a label indicating a result of identifying the encoded image from the feature map. Obtaining a label (S520) may include changing at least one of the backbone network and the head network (S910) based on the purpose or use of a visual task to be identified by using the AI ​​model 200. Obtaining a label (S520) may include outputting a label indicating a recognition result of the encoded image by inputting the encoded image into the AI ​​model 200 in which at least one of the backbone network and the head network has been changed (S920).

[0171] In an embodiment of the present disclosure, the method may further include identifying a replacement or change of the superlens 110 by a user (S1210). The method may further include identifying a visual task corresponding to the replaced or changed superlens 110 (S1220). The method may further include changing the backbone network and the head network of the AI ​​model 200 to a backbone network and a head network of a model predetermined to be optimized for the identified visual task (S1230).

[0172] In an embodiment of the present disclosure, the method may further include obtaining a depth value of the object from a depth sensor and obtaining an audio signal from the object from an audio sensor. Obtaining a label (S520) may include inputting the depth value of the object measured by the depth sensor and the audio signal obtained by the audio sensor to the AI ​​model 200, and obtaining a label corresponding to the result of identifying the object by performing inference using the AI ​​model 200.

[0173] The present disclosure provides a computer program product including a computer-readable storage medium. The storage medium may include instructions readable by the electronic device 100 to obtain a coded image by receiving light reflected from an object and phase-modulated by penetrating the metalens 110 and converting the received light into an electrical signal, and inputting the coded image to the AI ​​model 200 and performing inference using the AI ​​model 200 to obtain a label indicating a result of recognizing the object.

[0174] The program executed by the electronic device 100 described in this specification may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. The program may be executed by any system capable of executing computer-readable instructions.

[0175] The software may include a computer program, a code segment, an instruction, or a combination of one or more thereof, and configures the processing device to operate as desired or instructs the processing device, either independently or collectively.

[0176] Software may be implemented as a computer program including instructions stored in a computer-readable storage medium. Examples of computer-readable recording media include magnetic storage media (e.g., ROM, RAM, floppy disk, hard disk, etc.), optical recording media (e.g., compact disk (CD)-ROM and digital versatile disk (DVD)), etc. Computer-readable recording media may be distributed on computer systems connected via a network so that computer-readable code may be stored and executed in a distributed manner. The medium may be readable by a computer, stored in a memory, and executed by a processor.

[0177] The computer-readable storage medium may be provided in the form of a non-transitory storage medium. In this regard, the term "non-transitory" simply means that the storage medium does not include a signal and is a tangible device, and the term does not distinguish between where data is semi-permanently stored in the storage medium and where data is temporarily stored in the storage medium. For example, a "non-transitory storage medium" may include a buffer that temporarily stores data.

[0178] In addition, the program according to the embodiment disclosed in this specification may be included in a computer program product when provided. The computer program product may be traded between a seller and a buyer as a product.

[0179] The computer program product may include a software program and a computer-readable storage medium on which the software program is stored. For example, the computer program product may include a computer program that is provided by the manufacturer of the electronic device 100 or through an electronic market (e.g., Samsung Galaxy Store). TM ) A product in the form of a software program that is electronically distributed (e.g., a downloadable application). For such electronic distribution, at least a portion of the software program may be stored in a storage medium or may be temporarily generated. In this case, the storage medium may be a storage medium of a server of a manufacturer of the electronic device 100, a server of an electronic market, or a relay server for temporarily storing software programs.

[0180] In a system including the electronic device 100 and / or a server, the computer program product may include a storage medium of the server or a storage medium of the electronic device 100. Alternatively, in the presence of a third device (e.g., a wearable device) communicatively connected to the electronic device 100, the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include the software program itself that is sent from the electronic device 100 to the third device or from the third device to the electronic device.

[0181] In this case, one of the electronic device 100 and the third device may execute the computer program product to perform the method according to the disclosed embodiment. Alternatively, at least one of the electronic device 100 or the third device may execute the computer program product to perform the method according to the disclosed embodiment in a distributed manner.

[0182] For example, the electronic device 100 may execute a program stored in the memory ( Figure 2 130) to control another electronic device communicatively connected to the electronic device 100 to perform a method according to the disclosed embodiment.

[0183] In another example, the third device may execute the computer program product to control an electronic device communicatively connected to the third device to perform a method according to the disclosed embodiments.

[0184] In the case where the third device executes the computer program product, the third device may download the computer program product from the electronic device 100 and execute the downloaded computer program product. Alternatively, the third device may execute a computer program product preloaded therein to perform the method according to the disclosed embodiment.

[0185] Although the embodiments have been described above with reference to limited examples and drawings, it will be understood by those skilled in the art that various modifications and changes in form and detail may be made from the above description. For example, even when the above-described techniques are performed in a different order than described above, and / or the aforementioned components such as computer systems or modules are coupled or combined in a different form and mode than described above or replaced or supplemented by other components or their equivalents, sufficient effects may be achieved.

Claims

1. An electronic device (100), include: A metalens (110) having a pattern formed on a surface thereof and consisting of a plurality of pillars or pins having shapes, heights, and areas different from each other, and having an optical property of modulating a phase of light reflected from an object by the pattern on the surface; An image sensor (120) configured to obtain a coded image by receiving light reflected from an object and phase-modulated by passing through the metalens (110) and converting the received light into an electrical signal; as well as at least one processor (130) configured to input the encoded image to an artificial intelligence (AI) model (200) and obtain a label indicating a result of recognizing the object by performing inference using the AI ​​model (200), in, The AI ​​model (200) is a neural network model that is trained to obtain a simulated image by inputting a red, green, and blue (RGB) image to a model reflecting the optical properties of the metalens 110, and outputs a label indicating a true value of the input RGB image as a recognition result of the obtained simulated image, and The AI ​​model (200) is trained to minimize information indicating the similarity between the RGB image and the simulated image.

2. The electronic device (100) according to claim 1, in: The AI ​​model (200) includes: a first AI model (210) and a second AI model (220), the first AI model (210) being trained to output a simulated image from an RGB image by performing a convolution of the RGB image with a point spread function (PSF) that mathematically models the optical properties of the metalens (110), and the second AI model (220) being trained to output a label indicating a true value of the RGB image from the simulated image, and The at least one processor (130) is further configured to: The encoded image is input to the second AI model (220), and a label corresponding to the recognition result of the encoded image is obtained by inference by the second AI model (220).

3. The electronic device (100) according to claim 2, further comprising: include: The first AI model (210) is trained by updating weights via back-propagation, which applies mutual information representing the similarity between the RGB image and the simulated image as a loss so that the mutual information is minimized.

4. The electronic device (100) according to claim 3, in, The first AI model (210) is trained to minimize the loss by minimizing the Kullback-Leibler (KL) divergence representing the correlation between the distribution of the RGB image and the simulated image in the vector space.

5. The electronic device (100) according to claim 3, in: The AI ​​model (200) further includes a third AI model (230) configured to calculate a loss representing dissimilarity between the RGB image and the simulated image, and The third AI model (230) includes: a reconstruction model trained to generate fake images that mimic the RGB images based on the simulated images; and The discriminator model is trained to determine whether the input image is an RGB image or a fake image generated by the reconstruction model.

6. The electronic device (100) according to any one of claims 1 to 5, in, The AI ​​model (200) includes: A backbone network, trained to extract feature maps from the input encoded image; and A head network, trained to output labels indicating the result of identifying the encoded image from the feature maps, and The at least one processor (130) is further configured to: At least one of the backbone network and the head network is changed based on the purpose or use of the visual task to be identified by using the AI ​​model (200).

7. The electronic device (100) according to any one of claims 1 to 6, in, The AI ​​model (200) also includes: a metalens profiler model configured to recognize a replacement or change of the metalens (110) by a user, and The at least one processor (130) is further configured to, Where a replacement or change to the metalens (110) is identified via the metalens profiler model, identifying the visual task corresponding to the replaced or changed metalens (110), and The backbone network and head network of the AI ​​model (200) are changed to the backbone network and head network of a model predetermined to be optimized for the identified visual task.

8. A method for obtaining a result of identifying an object from an image obtained through a super lens (110), performed by an electronic device (100), the method include: Obtaining a coded image (S510) by receiving light reflected from an object and phase-modulated by passing through a super lens (110) and converting the received light into an electrical signal; as well as The encoded image is input to an artificial intelligence (AI) model (200), and a label indicating a result of recognizing an object is obtained by performing inference using the AI ​​model (200) (S520), in The AI ​​model (200) is a neural network model that is trained to obtain a simulated image by inputting a red, green, and blue (RGB) image to a model reflecting the optical properties of the metalens (110), and outputs a label indicating a true value of the input RGB image as a recognition result of the obtained simulated image, and The AI ​​model (200) is trained to minimize information indicating the similarity between the RGB image and the simulated image.

9. The method according to claim 8, in, The AI ​​model (200) includes: a first AI model (210) and a second AI model (220), the first AI model (210) being trained to output a simulated image from an RGB image by performing a convolution of the RGB image with a point spread function (PSF) that mathematically models the optical properties of the metalens (110), and the second AI model (220) being trained to output a label indicating a true value of the RGB image from the simulated image, and Obtaining a label (S520) includes: The encoded image is input to the second AI model (220), and the second AI model (220) performs inference to obtain a label corresponding to the recognition result of the encoded image.

10. The method according to claim 9, in, The first AI model (210) is trained by updating weights via back-propagation, which applies mutual information representing the similarity between the RGB image and the simulated image as a loss so that the mutual information is minimized.

11. The method according to claim 10, in, The first AI model (210) is trained to minimize the loss by minimizing the Kullback-Leibler (KL) divergence representing the correlation between the distribution of the RGB image and the simulated image in the vector space.

12. The method according to claim 10, in, The AI ​​model (200) further includes a third AI model (230) configured to calculate a loss representing dissimilarity between the RGB image and the simulated image, and The third AI model (230) includes: A reconstruction model, trained to generate fake images that mimic RGB images based on the simulated images, and The discriminator model is trained to determine whether the input image is an RGB image or a fake image generated by the reconstruction model.

13. The method according to any one of claims 8 to 12, in, The AI ​​model (200) includes: A backbone network, trained to extract feature maps from the input encoded image; and A head network, trained to output labels indicating the result of identifying the encoded image from the feature maps, and Obtaining a label (S520) includes: Changing at least one of the backbone network and the head network based on the purpose or use of the visual task to be identified by using the AI ​​model (200) (S910); and A label indicating a recognition result of the encoded image is output by inputting the encoded image to the AI ​​model (200) in which at least one of the trunk network and the head network has been changed (S920).

14. The method according to any one of claims 8 to 13, further comprising: include: Recognizing replacement or change of the metalens (110) by a user (S1210); identifying a visual task corresponding to the replaced or altered metalens (110) (S1220); and The backbone network and the head network of the AI ​​model (200) are changed to the backbone network and the head network of the model predetermined to be optimized for the identified visual task (S1230).

15. A computer program product comprising a computer-readable storage medium, in, The storage medium includes instructions related to a method executed by an electronic device (100) for obtaining a result of identifying an object from an image obtained through a super lens (110), the method comprising: Obtaining a coded image by receiving light reflected from an object and phase-modulated by passing through a super lens (110) and converting the received light into an electrical signal; and The encoded image is input to an artificial intelligence (AI) model (200), and a label indicating a result of recognizing an object is obtained by performing inference using the AI ​​model (200).

Citation Information

Cited By

  • Multi-channel privacy imaging device and design method and device thereof

    CN121357294A

  • Multichannel privacy imaging device and method and apparatus for designing the same

    CN121357294B

  • De-recognition camera device and monitoring method using same

    CN121967636A