Electronic device for acquiring image and operation method therefor
The electronic device enhances images by using a neural network model with variable frame processing, addressing inefficiencies in existing technologies by dynamically adjusting frame usage and improving image quality under varying conditions.
Patent Information
- Application Number
- PCT/KR2024/020655
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-17
AI Technical Summary
Existing image enhancement technologies require large amounts of training data and fixed numbers of image frames, leading to inefficient use of storage space and inflexibility in processing variable image conditions.
An electronic device and method that utilizes a neural network model to enhance images by variably applying the number of input frames, employing a first sub-neural network model to extract feature information from each frame and a second sub-neural network model to reconstruct an enhanced image, reducing storage needs and adapting to varying image conditions.
The solution allows for improved image enhancement by dynamically adjusting the number of frames based on conditions, reducing storage requirements and enhancing image quality under varying environmental conditions.
Smart Images

Figure KR2024020655_17072025_PF_FP_ABST
Abstract
Description
Electronic device for acquiring images and method of operating the same
[0001] The present disclosure relates to an electronic device for acquiring an image and a method of operating the same.
[0002] Deep learning can be used to train neural networks and enhance images using the resulting neural network models. Super-resolution imaging technology can generate high-resolution images corresponding to input low-resolution images. AI neural networks can learn from training data using deep learning algorithms. Training data can consist of pairs of data and labels (correct answers to the data). Training an AI neural network can require a large amount of training data. For example, an AI neural network can be configured to predict enhanced images from degraded images by learning training data consisting of pairs of ground truth images (labels) and degraded images (data) that are degraded compared to the ground truth images.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art in connection with the present disclosure.
[0004] In one embodiment, an electronic device may include at least one processor and a memory storing one or more instructions. The memory may further store a first sub-neural network model and a second sub-neural network model. The one or more instructions, when executed by the at least one processor, may cause the electronic device to acquire N image frames. The one or more instructions, when executed by the at least one processor, may cause the electronic device to input the image frames one by one to the first sub-neural network model to acquire feature information for each of the image frames. The one or more instructions, when executed by the at least one processor, may cause the electronic device to input integrated feature information, which is obtained by merging feature information for each of the image frames, to the second sub-neural network model to acquire an enhanced image. The first sub-neural network model may be configured to receive one image frame and output feature information. The second sub-neural network model may be configured to output an enhanced image based on the input feature information.
[0005] An operating method of an electronic device according to one embodiment may include an operation of acquiring N image frames. The operating method of the electronic device may include an operation of inputting the image frames one by one into the first sub-neural network model to acquire feature information for each of the image frames. The operating method of the electronic device may include an operation of acquiring integrated feature information that merges feature information for each of the image frames. The operating method of the electronic device may include an operation of inputting the integrated feature information into a second sub-neural network model to acquire an enhanced image. The first sub-neural network model may be configured to input one image frame and output feature information. The second sub-neural network model may be configured to output an enhanced image based on the input feature information.
[0006] In one embodiment, a computer-readable, non-transitory recording medium may have recorded thereon a computer program. The computer program may cause an electronic device to perform a method of operating the electronic device when executed. The method may include an operation of acquiring N image frames. The method may include an operation of inputting the image frames one by one into the first sub-neural network model to acquire feature information for each of the image frames. The method may include an operation of acquiring integrated feature information that merges the feature information for each of the image frames. The method may include an operation of inputting the integrated feature information into a second sub-neural network model to acquire an enhanced image. The first sub-neural network model may be configured to input one image frame and output feature information. The second sub-neural network model may be configured to output an enhanced image based on the input feature information.
[0007] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.
[0008] FIG. 2 is a block diagram illustrating a camera module according to one embodiment.
[0009] FIG. 3 illustrates learning data for training a neural network model to obtain an enhanced image according to one embodiment.
[0010] Figure 4 illustrates a process for obtaining an enhanced image based on a neural network model.
[0011] Figure 5 illustrates a configuration of an electronic device according to one embodiment.
[0012] FIG. 6 is a flowchart illustrating a process for configuring a neural network model mounted on an electronic device according to one embodiment.
[0013] FIG. 7 illustrates a process for an electronic device to acquire an enhanced image according to one embodiment.
[0014] FIG. 8 illustrates a process by which an electronic device according to one embodiment determines the number of image frames to be processed through neural network operations.
[0015] FIG. 9 illustrates a process by which an electronic device acquires an enhanced image based on a reference image frame and neighboring image frames, in one embodiment.
[0016] Hereinafter, embodiments are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the disclosed embodiments may be implemented in various different forms and are not limited to the embodiments described herein.
[0017] When training a neural network model to obtain an enhanced image from multiple image frames, the neural network model may be configured to process neural network operations on a fixed number of image frames. For example, a neural network model trained on training data consisting of pairs of five image frames and an enhanced image may require input of five image frames. In one embodiment, an electronic device and an operating method thereof may be provided for obtaining an enhanced image through a neural network model by varying the number of input image frames.
[0018] To process neural network operations on variable image frames, a separate neural network model trained based on training data containing a different number of image frames may be required. In one embodiment, an electronic device and method of operating the same can be provided that can reduce the storage space required to store a neural network model for processing variable image frames.
[0019] The technical problem to be achieved in this document is not limited to the technical problem mentioned above, and other technical problems that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs. FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) through a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0020] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0021] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0022] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0023] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0024] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0025] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0026] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0027] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0028] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0029] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0030] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0031] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0032] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0033] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0034] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0035] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0036] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0037] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0038] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0039] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0040] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0041] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0042] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0043] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0044] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0045] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0046] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0047] FIG. 2 is a block diagram (200) illustrating a camera module (180) according to various embodiments. Referring to FIG. 2, the camera module (180) may include a lens assembly (210), a flash (220), an image sensor (230), an image stabilizer (240), a memory (250) (e.g., a buffer memory), or an image signal processor (260). The lens assembly (210) may collect light emitted from a subject that is a target of image capturing. The lens assembly (210) may include one or more lenses. According to one embodiment, the camera module (180) may include a plurality of lens assemblies (210). In this case, the camera module (180) may form, for example, a dual camera, a 360-degree camera, or a spherical camera. Some of the plurality of lens assemblies (210) may have the same lens properties (e.g., angle of view, focal length, autofocus, f-number, or optical zoom), or at least one lens assembly may have one or more lens properties that are different from the lens properties of the other lens assemblies. A lens assembly (210) may include, for example, a wide-angle lens or a telephoto lens.
[0048] The flash (220) can emit light used to enhance light emitted or reflected from a subject. According to one embodiment, the flash (220) can include one or more light-emitting diodes (e.g., red-green-blue (RGB) LED, white LED, infrared LED, or ultraviolet LED), or a xenon lamp. The image sensor (230) can acquire an image corresponding to the subject by converting light emitted or reflected from the subject and transmitted through the lens assembly (210) into an electrical signal. According to one embodiment, the image sensor (230) can include one image sensor selected from among image sensors having different properties, such as an RGB sensor, a black and white (BW) sensor, an IR sensor, or a UV sensor, a plurality of image sensors having the same property, or a plurality of image sensors having different properties. Each image sensor included in the image sensor (230) can be implemented using, for example, a CCD (charged coupled device) sensor or a CMOS (complementary metal oxide semiconductor) sensor.
[0049] The image stabilizer (240) can move at least one lens or image sensor (230) included in the lens assembly (210) in a specific direction or control the operating characteristics of the image sensor (230) (e.g., adjusting the read-out timing, etc.) in response to the movement of the camera module (180) or the electronic device (101) including the same. This allows compensating for at least some of the negative effects of the movement on the captured image. In one embodiment, the image stabilizer (240) can detect such movement of the camera module (180) or the electronic device (101) using a gyro sensor (not shown) or an acceleration sensor (not shown) disposed inside or outside the camera module (180). In one embodiment, the image stabilizer (240) can be implemented as, for example, an optical image stabilizer. The memory (250) can temporarily store at least a portion of the image acquired through the image sensor (230) for the next image processing task. For example, when image acquisition is delayed due to the shutter, or when multiple images are acquired at high speed, the acquired original image (e.g., a Bayer-patterned image or a high-resolution image) is stored in the memory (250), and a corresponding copy image (e.g., a low-resolution image) can be previewed through the display module (160). Thereafter, when a specified condition is satisfied (e.g., a user input or a system command), at least a portion of the original image stored in the memory (250) can be acquired and processed, for example, by the image signal processor (260). According to one embodiment, the memory (250) can be configured as at least a portion of the memory (130) or as a separate memory that operates independently therefrom.
[0050] The image signal processor (260) can perform one or more image processing operations on an image acquired through an image sensor (230) or an image stored in a memory (250). The one or more image processing operations may include, for example, depth map generation, 3D modeling, panorama generation, feature extraction, image synthesis, or image compensation (e.g., noise reduction, resolution adjustment, brightness adjustment, blurring, sharpening, or softening). Additionally or alternatively, the image signal processor (260) may perform control (e.g., exposure time control, read-out timing control, etc.) for at least one of the components included in the camera module (180) (e.g., image sensor (230)). An image processed by the image signal processor (260) may be stored back in the memory (250) for further processing or provided to an external component of the camera module (180) (e.g., memory (130), display module (160), electronic device (102), electronic device (104), or server (108)). According to one embodiment, the image signal processor (260) may include at least one of the processors (120). It may be configured as a separate processor that is configured as a part of the processor (120) or operates independently of the processor (120). If the image signal processor (260) is configured as a separate processor from the processor (120), at least one image processed by the image signal processor (260) may be displayed through the display module (160) as is or after undergoing additional image processing by the processor (120).
[0051] According to one embodiment, the electronic device (101) may include a plurality of camera modules (180), each having different properties or functions. In this case, for example, at least one of the plurality of camera modules (180) may be a wide-angle camera, and at least another may be a telephoto camera. Similarly, at least one of the plurality of camera modules (180) may be a front camera, and at least another may be a rear camera.
[0052] FIG. 3 illustrates learning data (300) for training a neural network model to obtain an enhanced image according to one embodiment.
[0053] In one embodiment, a neural network model may be trained based on learning data (300) through machine learning, so that the neural network model can perform inference based on the learning data (300). Referring to FIG. 3, a neural network model according to one embodiment may be configured based on learning data (300) including a ground truth image (320) and a pair of M degraded images (310) that are degraded compared to the ground truth image (320). The ground truth image (320) may be included in the learning data (300) as a label for the degraded images (310). In one embodiment, the process of training the neural network model may be performed by an electronic device (e.g., the electronic device of FIG. 1). A neural network model trained by an external device of the electronic device may also be stored in the electronic device. In the present disclosure, for the convenience of explanation, the process of training the neural network model is described as being performed by the electronic device, but the entity performing the process of training the neural network model is not limited thereto.
[0054] In one embodiment, the electronic device may configure a feature extraction unit to extract feature information from M image frames by training a neural network model based on learning data (300). The electronic device may determine a reference image frame from among the M image frames. The feature extraction unit may extract feature points from the reference image frame. For example, the feature extraction unit may extract feature points from the reference image frame through a scale-invariant feature transform (SIFT) algorithm. However, the method of extracting feature points is not limited thereto. The feature extraction unit may be configured to output feature information based on similar and dissimilar feature points between the reference image frame and other image frames. For example, a neural network model may be configured through machine learning based on learning data (300) including at least one of global motion information representing the overall motion of an image with respect to the reference image frame or local motion information representing the motion of a portion of the image. However, the present invention is not limited thereto. For example, the electronic device may extract feature points individually from each image frame, without distinction with respect to the reference image frame.
[0055] Figure 4 illustrates a process of obtaining an improved image based on a neural network model (400).
[0056] In one embodiment, the neural network model (400) may be a base neural network model for configuring a sub-neural network model mounted on an electronic device (e.g., the electronic device (101) of FIG. 1). The neural network model (400) may be generated by learning learning data (e.g., the learning data (300) of FIG. 3). The neural network model (400) may be configured to receive M image frames (410-1, 420-2, ..., 420-M) according to the configuration of the learning data (e.g., the learning data (300) of FIG. 3) and output an enhanced image (470). The image frames (410-1, 410-2, ..., 410-M) may include images acquired through multiple shooting operations performed through at least one camera (e.g., the camera module (180) of FIGS. 1 and 2). For example, at least some of the image frames (410-1, 420-2, ..., 420-M) may include images captured based on different exposure values than other image frames. At least one sub-neural network model configured based on at least a portion of the neural network model (400) may be mounted on the electronic device.
[0057] The neural network model (400) may include a plurality of layers for performing neural network operations. The plurality of layers may include at least one of an operation layer (e.g., a convolution layer, a deconvolution layer) that performs an operation operation or an activation layer that performs an activation operation according to an activation function. The operation layer may refer to a layer that performs a convolution operation or a deconvolution operation. The activation layer may refer to a layer that performs an activation operation and determines whether to activate based on input data. However, the operation layer or the activation layer illustrated in FIG. 4 are merely conceptually illustrating the structure of the neural network model (400) including a plurality of layers for the purpose of explaining the invention, and the structure of the neural network model (400) of the present disclosure is not limited to the structure illustrated in FIG. 4. The neural network model (400) may include a feature extraction portion, which is composed of first layers (420-1, 420-2, ..., 420-M) for extracting feature information (430-1, 430-2, ..., 430-M) from a fixed number (M) of image frames (410-1, 420-2, ..., 420-M). The first layers (420-1, 420-2, ..., 420-M) may include at least one 1-1 layer (420-1) configured to extract first feature information (430-1) from the first image frame (410-1). The first layers (420-1, 420-2, ..., 420-M) may include at least one first-second layer (420-2) configured to extract second feature information (430-2) from a second image frame (410-2). The first layers (420-1, 420-2, ..., 420-M) may include at least one 1-M layer (420-M) configured to extract M feature information (430-M) from an M image frame (410-M).
[0058] The neural network model (400) may be configured to perform a merge operation (440) to obtain integrated feature information (450) by merging feature information (430-1, 430-2, ..., 430-M) obtained from a plurality of image frames (410-1, 410-2, ..., 410-M). The neural network model (400) may include at least one second layer (460) configured to perform a neural network operation on the integrated feature information (450). The neural network model (400) may be configured to perform an operation on the integrated feature information (450) according to at least one second layer (460). The neural network model (400) may be configured to reconstruct an enhanced image (470) from a result of performing a neural network operation on the integrated feature information (450) through at least one second layer (460). For example, whenever an operation based on a layer included in at least one second layer (460) is performed, feature information for an image with stepwise noise reduction or artifact removal can be acquired from the integrated feature information (450). In FIG. 4, the output of at least one second layer (460) is illustrated as an enhanced image (470), but is not limited thereto. For example, the output of at least one second layer (460) may include feature information corresponding to an enhanced image (470), and the neural network model (400) may be configured to output an image reconstructed from the feature information output from at least one second layer (460) as the enhanced image (470). For example, the neural network model (400) may output feature information output from the second layer (460), and a module for reconstructing the image may be configured separately from the neural network model (400). For example, the second layer (460) may be configured to perform an operation on feature information based on the input integrated feature information (450) and output an enhanced image (470) corresponding to the result of the operation. The enhanced image (470) may include a plurality of image frames (410-1, 410-2, ..., 410-M) may include images with reduced noise or artifacts removed.
[0059] In one example, among the plurality of layers included in the neural network model (400), the first layers (420-1, 420-2, ..., 420-M) may include a layer configured to extract feature information (first feature information (430-1, 430-2, ..., 430-M)) from each of the M image frames (410-1, 410-2, ..., 410-M) that are input. In one example, at least one second layer (460) among the plurality of layers included in the neural network model (400) may include a layer configured to predict feature information for an image from which noise is reduced or artifacts are removed based on feature information for the image frame.
[0060] FIG. 5 illustrates a configuration of an electronic device (101) (e.g., the electronic device (101) of FIG. 1) according to one embodiment.
[0061] In one embodiment, the electronic device (101) may include one or more processors (520) (e.g., the processor of FIG. 1) and a memory (530) (e.g., the memory 130 of FIG. 1, the memory 250 of FIG. 2). The electronic device (101) may perform operations of the electronic device (101) by having the one or more processors (520) execute one or more instructions stored in the memory (530). The electronic device (101) may further include at least one camera (580) (e.g., the camera module (180) of FIG. 1, the camera module (180) of FIG. 2). The one or more processors (520) may include, but are not limited to, a neural processing unit (NPU) configured to perform neural network operations. For example, one or more processors (520) may include a general-purpose processor such as an application processor (AP), a graphic processing unit (GPU), a digital signal processor (DSP), or an image signal processor (ISP).
[0062] In one embodiment, a neural network device may be configured based on at least one of one or more processors (520) or memory (530). The neural network device may analyze input data based on the neural network to extract valid information, and make a judgment on the input data or control a component of the electronic device (101) based on the extracted information. The neural network device may be configured to store a neural network model generated as a result of performing machine learning on another device in the memory (530), and have one or more processors (520) execute the stored neural network model. The operation of analyzing input data to extract valid information and making a judgment on the input data based on the extracted information may be referred to as “predicting” or “inferring” a result for the input data. The process of generating a neural network model as a result of performing machine learning may be referred to as “training” the neural network model.
[0063] In one embodiment, the neural network device may be configured to perform not only a prediction operation but also a training operation using a neural network model trained in another device. For example, the neural network device may generate a neural network, train (or learn) a neural network, perform a neural network-based operation based on input data, generate an information signal based on the operation result, or retrain a neural network. At least a portion of the neural network device may be implemented by an external device (e.g., the electronic device (102), the electronic device (104), or the server (108) of FIG. 1 ). For example, the electronic device (101) may receive a neural network generated or trained by the external device from the external device and store it in the memory (530). For example, the electronic device (101) may transmit an image frame or information extracted from the image frame to the external device, and may receive the result of performing a neural network operation on the image frame from the external device. For example, the electronic device (101) may receive a neural network retrained by the external device from the external device and update the neural network model stored in the memory (530). One or more processors (520) may include a hardware accelerator for executing the neural network model. For example, the hardware accelerator may include, but is not limited to, a neural processing unit (NPU), a tensor processing unit (TPU), or an artificial intelligence engine.
[0064] In one embodiment, one or more processors (520) may execute a neural network model. The neural network model may include a deep learning model that performs a specific purpose operation based on the result of learning training data. For example, the neural network model may include at least one of various types of neural network models, such as a convolution neural network (CNN), a region with convolution neural network (R-CNN), a region proposal network (RPN), a recurrent neural network (RNN), a stacking-based deep neural network (S-DNN), a state-space dynamic neural network (S-SDNN), a deconvolution network, a deep belief network (DBN), a restricted boltzmann machine (RBM), a fully convolutional network, a long short-term memory (LSTM) network, or a classification network.
[0065] In one embodiment, the memory (530) may include a first sub-neural network model (541) and a second sub-neural network model (542). The first sub-neural network model (541) may be configured to receive one image frame as input and output feature information. The second sub-neural network model (542) may be configured to output an enhanced image based on the feature information input to the second sub-neural network model (542).
[0066] In one embodiment, the first sub-neural network model (541) may include at least one layer (420) configured based on at least one first layer (e.g., the first layers (420-1, 420-2, ..., 420-M) of FIG. 4) configured to extract feature information from an image frame from among a plurality of layers included in the base neural network model (e.g., the plurality of layers including the first layers (420-1, 420-2, ..., 420-M) of FIG. 4 and at least one second layer (460). For example, the at least one layer (420) may be selected from among the 1-1 layer (420-1), the 1-2 layer (420-2), ..., or the 1-M layer (420-M) of FIG. 4 included in the base neural network model. For example, at least one layer (420) may be configured by merging the 1-1 layer (420-1), the 1-2 layer (420-2), ..., or the 1-M layer (420-M) of FIG. 4 included in the base neural network model. At least one layer (420) may be configured to perform an operation on an input image frame. For example, at least one layer (420) may include at least one of an operation layer (e.g., a convolution layer, a deconvolution layer) that performs an operation operation or an activation function.
[0067] In one embodiment, the second sub-neural network model (542) may include at least one second layer (460) configured to receive feature information (e.g., integrated feature information) from among a plurality of layers included in the base neural network model (e.g., a plurality of layers including the first layers (420-1, 420-2, ..., 420-M) and at least one second layer (460) of FIG. 4) and determine feature information corresponding to the enhanced image or the enhanced image. The electronic device (101) may reconstruct the enhanced image based on feature information output as a result of an operation performed based on at least one second layer (460) on feature information acquired through the first sub-neural network model (541).
[0068] In one embodiment, the electronic device (101) may execute the first sub-neural network model (541) a number of times corresponding to the number of image frames used to obtain an enhanced image. For example, when an image is obtained based on N image frames, the electronic device (101) may input the image frames one by one into the first sub-neural network model (541) and execute the first sub-neural network model N times. The electronic device (101) may operate so that N has the same value as the number M of image frames that the base neural network model receives as input, but may also operate so that it has a different value. For example, when the first sub-neural network model (541) and the second sub-neural network model (542) are configured based on a base neural network model configured to output an enhanced image from three image frames, and the enhanced image is configured based on image frames with a lot of noise, the electronic device (101) may perform an operation of executing the first sub-neural network model (541) five times to obtain feature information from the five image frames.
[0069] In one embodiment, the electronic device (101) can acquire integrated feature information by merging feature information acquired from multiple image frames. The electronic device (101) can acquire an enhanced image by inputting the integrated feature information into a second sub-neural network model (542).
[0070] In one embodiment, the electronic device (101) may acquire image frames to be input to the first sub-neural network model (541) by performing a photographing operation through at least one camera (580). However, the present invention is not limited thereto. For example, the image frames to be input to the first sub-neural network model (541) may be stored in the memory (530) or received from an external device. The electronic device (101) may determine the number (N) of image frames used to acquire an enhanced image based on a specified condition. For example, if it is determined that the level of noise occurring in an image captured through at least one camera (580) is high, the electronic device (101) may increase the number of times images are captured through at least one camera (580). The electronic device (101) may execute the first sub-neural network model (541) as many times as the number (N) of image frames acquired through at least one camera (580).
[0071] In one embodiment, the first sub-neural network model (541) and the second sub-neural network model (542) may be executed by one or more processors (520). The electronic device (101) may execute at least a part of the first sub-neural network model (541) or the second sub-neural network model (542) by different processors according to predefined conditions. For example, the electronic device (101) may execute at least a part of the first sub-neural network model (541) or the second sub-neural network model (542) by a CPU, a GPU, or a DSP when the NPU is occupied by another process.
[0072] FIG. 6 is a flowchart (600) illustrating a process for configuring a neural network model mounted on an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) according to one embodiment.
[0073] In the present disclosure, the operation of the electronic device may be understood as being performed by one or more processors (e.g., processor (120) of FIG. 1, processor (520) of FIG. 5) executing one or more instructions stored in a memory (e.g., memory (130) of FIG. 1, memory (530) of FIG. 5). In the present disclosure, at least some of the operations of the electronic device may also be performed by an external electronic device (e.g., electronic device (102), electronic device (104), server (108) of FIG. 1).
[0074] In operation 610, an electronic device or an external electronic device (e.g., the electronic device (102), the electronic device (104), the server (108) of FIG. 1) according to one embodiment may train a neural network model (e.g., the neural network model (400) of FIG. 4) to be embedded in the electronic device based on learning data (e.g., the learning data (300) of FIG. 3). The learning data may include pairs of a plurality of degraded images (e.g., the degraded images (310) of FIG. 3) and a true image (320). The neural network model trained based on the learning data may be configured to receive image frames corresponding to the number of degraded images constituting a pair of learning data.
[0075] In one embodiment, the electronic device or the external electronic device can train a neural network model by further including motion information between images. The electronic device or the external electronic device can configure the feature extraction part of the neural network model to extract feature points from a reference image frame among a plurality of images, and output feature information based on feature points that are similar to and dissimilar to the feature points of the reference image frame among feature points of other image frames. For example, the electronic device or the external electronic device can configure the neural network model through machine learning based on at least one of global motion information representing the overall motion of another image based on the reference image frame, or local motion information representing the motion of a portion of the image.
[0076] In one embodiment, the electronic device or the external electronic device may train a neural network model to perform preprocessing to determine whether to synthesize information from another image with information from a reference image even when the multiple image frames have different characteristics. For example, the electronic device or the external electronic device may train the neural network model based on multiple image frames that have different brightness values, were captured based on different exposure values, or have noise with different characteristics. In operation 610, the electronic device or the external electronic device may train the neural network model by inputting multiple image frames in parallel.
[0077] In operation 620, the electronic device or the external electronic device according to an embodiment may obtain a first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) and a second neural network model (e.g., the second sub-neural network model (542) of FIG. 5) from the neural network model generated in operation 610. For example, the electronic device or the external electronic device may generate the first sub-neural network model based on at least one first layer (e.g., at least one of the first layers (420-1, 420-2, ..., 420-M) of FIG. 4) configured to perform an operation of inferring feature information from an image frame among a plurality of layers included in the neural network model. The electronic device or the external electronic device may generate the second sub-neural network model based on at least one second layer (e.g., at least one second layer (460) of FIG. 4) configured to perform an operation of inferring feature information corresponding to an enhanced image based on integrated feature information among a plurality of layers included in the neural network model. The first sub-neural network model and the second sub-neural network model can be stored in the memory of the electronic device (e.g., the memory (120) of FIG. 1, the memory (530) of FIG. 5).
[0078] FIG. 7 is a flowchart (700) illustrating a process by which an electronic device (e.g., the electronic device (101) of FIG. 1 or FIG. 5) acquires an enhanced image according to one embodiment.
[0079] In operation 710, the electronic device according to one embodiment may acquire N image frames. For example, the electronic device may perform an operation of capturing an image through at least one camera (e.g., the camera module (180) of FIGS. 1 and 2, at least one camera (580) of FIG. 5) N times. At least some of the N image frames may be captured based on different settings from other image frames. For example, the electronic device may capture some of the image frames based on an appropriate exposure value, some of the image frames based on an overexposure value, and some of the remaining image frames based on an underexposure value.
[0080] In operation 720, an electronic device according to an embodiment may input N image frames one by one into a first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) to obtain feature information about the image frames. By extracting feature information for each image frame using the first sub-neural network model, the electronic device may obtain feature information for a variable number of image frames using the first sub-neural network model.
[0081] In operation 730, an electronic device according to an embodiment may acquire integrated feature information by merging feature information for image frames. For example, the electronic device may determine whether to merge feature information extracted from another image frame with reference feature information extracted from a reference image frame among the image frames. However, the present invention is not limited thereto. For example, the electronic device may merge feature information extracted from image frames without distinction of the reference image frame (e.g., merge information by interpolating each feature point). The electronic device may perform a merging operation to merge feature information determined to be merged with the reference feature information. The merging operation may be composed of an operation without weights. For example, the merging operation may include at least one of a softmax function, a sigmoid operation, an average operation, or a sum operation. For example, the merging operation may be performed based on the softmax operation of Mathematical Expression 1 below.
[0082]
[0083] In mathematical expression 1, n may correspond to the number N of image frames. In one embodiment, the softmax function may include an activation function that passes the N outputs of the first sub-neural network model as inputs of an exponential function so that the sum of the values becomes 1. By passing the feature information through the softmax function and merging it, the learning and inference process of the neural network model can be performed in a state in which the characteristic that the output value increases exponentially as the input value increases is reflected. By applying the softmax function, the electronic device can cause the neural network model to extract feature information from an image frame suitable for obtaining an enhanced image from among a plurality of images.
[0084] In one embodiment, the operation for merging feature information may be learned without parameters. The merging operation performed in operation 730 may be the same as the merging operation configured during the training process of the neural network model (e.g., the merging operation (440) of FIG. 4), but another merging operation may also be performed.
[0085] In operation 740, an electronic device according to one embodiment may input integrated feature information into a second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5 ). The electronic device may construct an enhanced image (e.g., the enhanced image (470) of FIG. 5 ) based on the feature information output from the second sub-neural network model.
[0086] FIG. 8 is a flowchart (800) illustrating a process in which an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) according to one embodiment determines the number of image frames to be processed through neural network operations.
[0087] In one embodiment, the electronic device may determine the number N of image frames to be processed through neural network operations based on specified conditions. For example, if a large number of image frames are required to obtain an image of adequate quality, the electronic device may increase the value of N. For example, if an image of adequate quality can be obtained with only a small number of image frames, the electronic device may decrease the value of N.
[0088] In operation 810, an electronic device according to an embodiment may determine a noise level for at least some of the image frames. For example, the electronic device may determine the illuminance value through a sensor that detects an illuminance value (e.g., an illuminance sensor of the sensor module (176) of FIG. 1, an image sensor (230) of FIG. 2). If the illuminance value is below a threshold, the electronic device may determine that the noise level is high. For example, the electronic device may determine information that an image frame contains noise from an image stream output from an image sensor (e.g., an image sensor (230) of FIG. 2).
[0089] In operation 820, the electronic device according to one embodiment may determine the number N of image frames corresponding to the determined noise level. For example, if the determined noise level is higher than a first threshold, the electronic device may determine the value of N to be greater than a specified value (e.g., M in FIGS. 3 and 4 ). For example, if the determined noise level is lower than a second threshold, the electronic device may determine the value of N to be less than a specified value (e.g., M in FIGS. 3 and 4 ). The first threshold may be a different value from the second threshold, but the first threshold and the second threshold may be the same value.
[0090] In operation 830, an electronic device according to one embodiment may acquire image frames based on the determined N value. For example, if the determined value is 5, the electronic device may acquire five image frames captured through at least one camera.
[0091] FIG. 9 illustrates a process in which an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) acquires an enhanced image based on a reference image frame and a neighboring image frame, in one embodiment.
[0092] In one embodiment, the electronic device can obtain an enhanced image from a plurality of image frames. For example, referring to FIG. 9, the electronic device can obtain an enhanced image (950) based on a first image frame (901), a second image frame (902), and a third image frame (903). The electronic device can obtain preprocessing results for the first image frame (901), the second image frame (902), and the third image frame (903) through a header (910) trained to perform preprocessing. The preprocessing performed based on the header (910) may refer to image processing for allowing a neural network model to perform operations on the image frames. For example, the preprocessing may include an image processing operation for resizing an image. For example, the electronic device can determine a reference image frame among the first image frame (901), the second image frame (902), and the third image frame (903), and obtain a preprocessed first image frame (921), a preprocessed second image frame (922), and a preprocessed third image frame (923). The number of the plurality of image frames can be variable.
[0093] In one embodiment, the electronic device may determine a reference image frame from among a plurality of image frames. The reference image frame may refer to an image frame that serves as a reference when synthesizing the plurality of image frames. For example, referring to FIG. 9, the electronic device may determine a first image frame (901) as a reference image frame. The electronic device may obtain feature information including information about extracted feature points by processing a preprocessed first image frame (921) corresponding to the reference image frame through a neural network (930). For example, the electronic device may extract feature points from the reference image frame by executing a scale-invariant feature transform (SIFT) algorithm on the reference image frame. However, the method of extracting feature points is not limited thereto. The neural network (930) may be a first sub-neural network model (541) (e.g., the first sub-neural network model (541) of FIG. 5), but may also be a separately configured neural network model.
[0094] In one embodiment, the electronic device may acquire feature information by processing each of a plurality of image frames based on a first sub-neural network model (541). Referring to FIG. 9, the electronic device may execute the first sub-neural network model (541) a total of three times by executing the first sub-neural network model (541) for each of a preprocessed first image frame (921), a preprocessed second image frame (922), and a preprocessed third image frame (923). The first sub-neural network model (541) may be executed multiple times depending on the number of the plurality of image frames.
[0095] In one embodiment, the electronic device may perform a product operation (940) on feature information for a reference image frame (e.g., the first image frame (901)) and feature information for each of the image frames (e.g., the first image frame (901), the second image frame (902), and the third image frame (903)). During the product operation (940), the electronic device may determine, among the feature information for each of the image frames, which information will be merged with the feature information of the reference image frame and which information will be ignored. For example, since each of the image frames is captured at a different time, when synthesizing image frames in which a moving subject is captured, the subject may appear blurry, like a ghost effect. Accordingly, among the feature information of other image frames (e.g., the second image frame (902), the third image frame (903)), feature information for a subject whose movement exceeds a certain level may be excluded from being merged with the feature information for the reference image frame (e.g., the first image frame (901)).
[0096] In one embodiment, the electronic device may obtain integrated feature information by performing a merge operation (440) that merges feature information for image frames (e.g., a first image frame (901), a second image frame (902), and a third image frame (903)). The merge operation (440) of FIG. 9 may be the same as the merge operation (e.g., the merge operation (440) of FIG. 4) of the underlying neural network model (e.g., the neural network model (400) of FIG. 4), but may also be configured through separate learning. The electronic device may select an operation for performing the merge operation (440) from among a plurality of operations according to conditions related to the image frames.
[0097] In one embodiment, the electronic device can obtain an enhanced image (950) that integrates a plurality of preprocessed image frames (e.g., the preprocessed first, second, and third image frames (921, 922, 923)) by executing a second sub-neural network model (542) based on the integrated feature information. The electronic device can construct the enhanced image (950) based on the feature information obtained through the neural network operation of the second sub-neural network model (542).
[0098] In one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (101) of FIG. 5) may include at least one processor (e.g., processor (120) of FIG. 1, processor (520) of FIG. 3) and a memory (e.g., memory (130) of FIG. 1, memory (530) of FIG. 5) that stores one or more instructions. The memory (e.g., memory (130) of FIG. 1, memory (530) of FIG. 5) may further store a first sub-neural network model (e.g., first sub-neural network model (541) of FIG. 5) and a second sub-neural network model (e.g., second sub-neural network model (542) of FIG. 5). The one or more instructions may be executed by the at least one processor (e.g., the processor (120) of FIG. 1, the processor (520) of FIG. 3) to cause the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to obtain N image frames. The one or more instructions may be executed by the at least one processor (e.g., the processor (120) of FIG. 1, the processor (520) of FIG. 3) to cause the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to input the image frames one by one into the first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) to obtain feature information for each of the image frames. The one or more instructions may be executed by the at least one processor (e.g., the processor (120) of FIG. 1, the processor (520) of FIG. 3) to cause the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to input integrated feature information that merges feature information for each of the image frames into the second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) to obtain an enhanced image.The first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) may be configured to input one image frame and output feature information. The second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) may be configured to output an enhanced image based on the input feature information.
[0099] In one embodiment, the first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) may be configured based on at least one first layer configured to extract feature information from an image frame among a plurality of layers included in the base neural network model. The second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) may include at least one second layer configured to determine feature information corresponding to the enhanced image from input feature information among the plurality of layers included in the base neural network model. The base neural network model may be generated through a machine learning process that inputs M image frames and outputs an inference result for the enhanced image.
[0100] In one embodiment, the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) may further include at least one camera. The one or more instructions may be executed by the at least one processor (e.g., the processor (120) of FIG. 1, the processor (520) of FIG. 3) to cause the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to capture the N image frames through the at least one camera.
[0101] In one embodiment, the one or more instructions may be executed by the at least one processor (e.g., the processor 120 of FIG. 1, the processor 520 of FIG. 3) to cause the electronic device (e.g., the electronic device 101 of FIG. 1, the electronic device 101 of FIG. 5) to determine the value of N based on shooting condition information related to an operation of capturing an image by the at least one camera. The one or more instructions may be executed by the at least one processor (e.g., the processor 120 of FIG. 1, the processor 520 of FIG. 3) to cause the electronic device (e.g., the electronic device 101 of FIG. 1, the electronic device 101 of FIG. 5) to capture N images through the at least one camera based on the determined value of N.
[0102] In one embodiment, the one or more instructions may be executed by the at least one processor (e.g., the processor (120) of FIG. 1, the processor (520) of FIG. 3) to cause the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to determine the value of N as a first value when it is determined that the noise in the image captured through the at least one camera is low, and to determine the value of N as a second value greater than the first value when it is determined that the noise in the image is high.
[0103] In one embodiment, the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) may further include a light sensor that detects light information. The one or more instructions may be executed by the at least one processor (e.g., the processor (120) of FIG. 1, the processor (520) of FIG. 3) to cause the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to determine that the image has a lot of noise when the light information detected through the light sensor is below a threshold.
[0104] In one embodiment, N and M may have different values.
[0105] In one embodiment, the one or more instructions may be executed by the at least one processor (e.g., processor (120) of FIG. 1, processor (520) of FIG. 3) to cause the electronic device (e.g., electronic device (101) of FIG. 1, electronic device (101) of FIG. 5) to obtain the integrated feature information based on at least one of a softmax function, a sigmoid operation, an average operation, or a sum operation for feature information for each of the image frames.
[0106] According to one embodiment, a method of operating an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) may include an operation of acquiring N image frames. The method of operating an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) may include an operation of inputting the image frames one by one into the first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) to acquire feature information for each of the image frames. The method of operating an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) may include an operation of acquiring integrated feature information that merges feature information for each of the image frames. A method of operating an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) may include an operation of inputting the integrated feature information into a second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) to obtain an enhanced image. The first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) may be configured to input one image frame and output feature information. The second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) may be configured to output an enhanced image based on the input feature information.
[0107] In one embodiment, the first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) may be configured based on at least one first layer among a plurality of layers included in a base neural network model. The second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) may include at least one second layer among the plurality of layers included in the base neural network model, configured to determine feature information corresponding to the enhanced image from input feature information. The base neural network model may be generated through a machine learning process that inputs M image frames and outputs an inference result for the enhanced image.
[0108] In one embodiment, the act of obtaining the image frames may include an act of capturing the image frames through at least one camera of the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5).
[0109] In one embodiment, the operation of capturing the image frames may include an operation of determining the value of N based on shooting condition information related to the operation of capturing an image by the at least one camera. The operation of capturing the image frames may include an operation of capturing N images through the at least one camera based on the determined value of N.
[0110] In one embodiment, the operation of capturing the image frames may include an operation of determining the value of N as a first value when it is determined that the image captured through the at least one camera has little noise, and an operation of determining the value of N as a second value greater than the first value when it is determined that the image has a lot of noise.
[0111] In one embodiment, the operation of capturing the image frames may include an operation of determining that the image has a lot of noise when the illumination information detected through the illumination sensor of the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) is below a threshold.
[0112] In one embodiment, N and M may have different values.
[0113] In one embodiment, the operation of obtaining the integrated feature information may include an operation of obtaining the integrated feature information based on at least one of a softmax operation, a sigmoid operation, an average operation, or a sum operation on feature information for each of the image frames.
[0114] In one embodiment, a computer-readable, non-transitory recording medium may have recorded thereon a computer program. The computer program may cause an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) to perform an operating method of the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5) when executed. The method may include an operation of acquiring N image frames. The method may include an operation of inputting the image frames one by one into the first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) to acquire feature information for each of the image frames. The method may include an operation of acquiring integrated feature information that merges the feature information for each of the image frames.
[0115] The method may include an operation of obtaining an enhanced image by inputting the integrated feature information into a second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5). The first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) may be configured to input one image frame and output feature information. The second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) may be configured to output an enhanced image based on the input feature information.
[0116] In one embodiment, the first sub-neural network model (e.g., the first sub-neural network model (541) of FIG. 5) may be configured based on at least one first layer among a plurality of layers included in a base neural network model. The second sub-neural network model (e.g., the second sub-neural network model (542) of FIG. 5) may include at least one second layer among the plurality of layers included in the base neural network model, configured to determine feature information corresponding to the enhanced image from input feature information. The base neural network model may be generated through a machine learning process that inputs M image frames and outputs an inference result for the enhanced image.
[0117] In one embodiment, the act of obtaining the image frames may include an act of capturing the image frames through at least one camera of the electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (101) of FIG. 5).
[0118] In one embodiment, the operation of capturing the image frames may include an operation of determining the value of N based on shooting condition information related to the operation of capturing an image by the at least one camera. The operation of capturing the image frames may include an operation of capturing N images through the at least one camera based on the determined value of N.
[0119] An electronic device and an operating method thereof according to one embodiment can obtain an improved image through a neural network model by variably applying a quantity of input image frames.
[0120] An electronic device and a method of operating the same according to one embodiment can reduce storage space used to store a neural network model for processing a variable number of image frames.
[0121] An electronic device and an operating method thereof according to one embodiment can acquire an improved image even when conditions deteriorate by adjusting the number of image frames according to conditions for acquiring an image.
[0122] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description of the present disclosure.
[0123] The methods according to the embodiments described in the claims or specification of the present disclosure may be implemented in the form of hardware, software, or a combination of hardware and software.
[0124] When implemented in software, a computer-readable storage medium storing one or more programs (software modules) may be provided. The one or more programs stored in the computer-readable storage medium are configured for execution by one or more processors within an electronic device. The one or more programs include instructions that cause the electronic device to execute methods according to embodiments described in the claims or specification of the present disclosure.
[0125] In the present disclosure, a function or operation performed by an electronic device may be performed by one or more processors executing one or more instructions stored in a memory. The function or operation of the electronic device mentioned in the present disclosure may be performed by one processor executing one or more instructions, or may be performed by a combination of multiple processors executing one or more instructions. The processor mentioned in the present disclosure may be understood to include a circuit for performing an operation or controlling other components of the electronic device. For example, the one or more processors may include a central processing unit (CPU), a microprocessor unit (MPU), an application processor (AP), a communication processor (CP), a neural processing unit (NPU), a system on chip (SoC), an integrated circuit (IC), or an application-specific integrated circuit (ASIC) configured to execute one or more instructions. The one or more processors may be configured to perform the operations of the electronic device described above.
[0126] In the present disclosure, a program (software module, software) may be stored in a non-volatile memory including a random access memory (RAM), a flash memory, a read only memory (ROM), an electrically erasable programmable read only memory (EEPROM), a magnetic disc storage device, a compact disc ROM (CD-ROM), digital versatile discs (DVDs) or other forms of optical storage devices, a magnetic cassette. Or, it may be stored in a memory formed by a combination of some or all of these. The memory may be formed by a single storage medium, or may be formed by a combination of a plurality of storage media. The one or more commands may be stored in a single storage medium, or may be distributed and stored in a plurality of storage media.
[0127] Additionally, the program may be stored on an attachable storage device that is accessible via a communication network such as the Internet, an intranet, a local area network (LAN), a wide LAN (WLAN), or a storage area network (SAN), or a combination thereof. Such a storage device may be connected to a device performing an embodiment of the present disclosure via an external port. Additionally, a separate storage device on the communication network may be connected to a device performing an embodiment of the present disclosure.
[0128] In the specific embodiments of the present disclosure described above, components included in the disclosure are expressed in the singular or plural form, depending on the specific embodiment presented. However, the singular or plural expressions are selected to suit the presented situation for convenience of explanation, and the present disclosure is not limited to singular or plural components. Components expressed in the plural form may be composed of singular elements, or components expressed in the singular form may be composed of plural elements.
[0129] Additionally, in the present disclosure, terms such as “part”, “module”, etc. may refer to a hardware component such as a processor or circuit, and / or a software component executed by a hardware component such as a processor.
[0130] A "component" or "module" may be implemented by a program stored in an addressable storage medium and executed by a processor. For example, a "component" or "module" may be implemented by components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.
[0131] The specific implementations described in this disclosure are merely exemplary and do not limit the scope of the present disclosure in any way. For the sake of brevity, descriptions of conventional electronic components, control systems, software, and other functional aspects of the systems may be omitted.
[0132] Additionally, in the present disclosure, “comprising at least one of a, b, or c” may mean “comprising only a, including only b, including only c, or including a combination of two or more (including a and b, including b and c, including a and c, or including all of a, b, and c).
[0133] While the detailed description of this disclosure has described specific embodiments, it should be understood that various modifications are possible without departing from the scope of this disclosure. Therefore, the scope of this disclosure should not be limited to the described embodiments, but should be defined not only by the scope of the claims described below, but also by equivalents thereof.
[0134] As used herein, the term "if" will be understood to mean "when, upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, "if it is determined to," or "if [the stated condition or event] is detected," will optionally be understood to mean "upon determining," or "in response to determining," "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]."
[0135] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0136] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0137] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0138] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0139] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In electronic devices, at least one processor; and A memory storing a first sub-neural network model, a second sub-neural network model, and one or more instructions, The one or more instructions are executed by the at least one processor, such that the electronic device: Obtain N image frames, By inputting the image frames one by one into the first sub-neural network model, feature information for each of the image frames is obtained, The integrated feature information, which is the feature information for each of the image frames, is input into the second sub-neural network model to obtain an enhanced image. The above first sub-neural network model is configured to input one image frame and output feature information, An electronic device, wherein the second sub-neural network model is configured to output an enhanced image based on input feature information.
2. In claim 1, The above first sub-neural network model is configured based on at least one first layer configured to extract feature information from an image frame among a plurality of layers included in the base neural network model, The second sub-neural network model includes at least one second layer configured to determine feature information corresponding to the enhanced image from the input feature information among the plurality of layers included in the base neural network model, An electronic device wherein the above-mentioned underlying neural network model is generated through a machine learning process that inputs M image frames and outputs inference results for an enhanced image.
3. In claim 2, An electronic device wherein the above N and the above M have different values.
4. In claim 1, The electronic device further comprises at least one camera, An electronic device, wherein said one or more instructions are executed by said at least one processor to cause said electronic device to capture said N image frames through said at least one camera.
5. In claim 4, The one or more instructions are executed by the at least one processor so that the electronic device: The value of N is determined based on shooting condition information related to the operation of capturing an image by at least one camera, An electronic device configured to capture N images through at least one camera based on the determined value of N.
6. In claim 4, An electronic device wherein the one or more instructions are executed by the at least one processor to cause the electronic device to determine the value of N as a first value when the image captured through the at least one camera is determined to have little noise, and to determine the value of N as a second value greater than the first value when the image is determined to have a lot of noise.
7. In claim 6, Further comprising a light sensor for detecting light information, An electronic device wherein the one or more instructions are executed by the at least one processor to cause the electronic device to determine that the image has a lot of noise when the illuminance information detected through the illuminance sensor is below a threshold.
8. In claim 1, An electronic device, wherein the one or more instructions are executed by the at least one processor to cause the electronic device to obtain the integrated feature information based on at least one of a softmax function, a sigmoid operation, an average operation, or a sum operation on feature information for each of the image frames.
9. In the method of operating an electronic device, The action of obtaining N image frames; An operation of inputting the image frames one by one into the first sub-neural network model to obtain feature information for each of the image frames; An operation of obtaining integrated feature information by merging the feature information for each of the image frames; and Including an operation of obtaining an enhanced image by inputting the above integrated feature information into a second sub-neural network model, The above first sub-neural network model is configured to input one image frame and output feature information. A method wherein the second sub-neural network model is configured to output an enhanced image based on input feature information.
10. In claim 9, The above first sub-neural network model is configured based on at least one first layer among multiple layers included in the base neural network model, The second sub-neural network model includes at least one second layer configured to determine feature information corresponding to the enhanced image from the input feature information among the plurality of layers included in the base neural network model, A method wherein the above-mentioned underlying neural network model is generated through a machine learning process that inputs M image frames and outputs inference results for an enhanced image.
11. In claim 10, A method wherein the above N and the above M have different values.
12. In claim 9, A method wherein the act of obtaining the image frames comprises an act of photographing the image frames through at least one camera of the electronic device.
13. In claim 12, The action of capturing the above image frames is: An operation of determining the value of N based on shooting condition information related to an operation of capturing an image by at least one camera, and A method comprising an operation of capturing N images through at least one camera based on the determined value of N.
14. In claim 12, The action of capturing the above image frames is: A method comprising: determining the value of N as a first value when it is determined that the noise in the image captured through the at least one camera is low, and determining the value of N as a second value greater than the first value when it is determined that the noise in the image is high.
15. In a non-transitory computer-readable recording medium, the recording medium is capable of executing, when an electronic device is running, The action of obtaining N image frames; An operation of inputting the image frames one by one into the first sub-neural network model to obtain feature information for each of the image frames; An operation of obtaining integrated feature information by merging feature information for each of the above image frames; and A computer program is recorded that performs a method including an operation of obtaining an enhanced image by inputting the above integrated feature information into a second sub-neural network model, The above first sub-neural network model is configured to input one image frame and output feature information. A recording medium, wherein the second sub-neural network model is configured to output an enhanced image based on input feature information.
Citation Information
Patent Citations
Face Recognition System For Real Image Judgment Using Face Recognition Model Based on Deep Learning
KR102294574B1
Image display method and device
KR102574141B1
Method, and device for training an image processing model
KR102582738B1
KR20200055760A
KR20200134813A