Electronic device for identifying region of interest in image and method for controlling the same
Through the neural network model, the region of interest and background areas in the image are identified and processed, the problem of increasing power consumption of large screen electronic devices is solved, and the power consumption optimization and image quality balance is achieved.
Patent Information
- Application Number
- CN202380073980.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-07-28
- Publication Date
- 2025-06-03
AI Technical Summary
With the development of large screens, the power consumption of electronic devices increases, resulting in carbon emission problems. At the same time, the prior art is difficult to effectively identify and separate areas of interest and background areas, resulting in poor power consumption optimization results.
Using neural network model, the areas of interest in the image are identified through convolutional networks and long short-term memory networks (LSTMs) training, and the brightness of the background area is reduced through image processing technology, thereby reducing power consumption.
Accurate identification and processing of areas of interest and background areas is achieved, power consumption of electronic devices is reduced, carbon emissions are reduced, and image quality is maintained.
Smart Images

Figure CN120092260A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an electronic device and a control method of the electronic device, and more particularly, to an electronic device for identifying a region of interest in an image and a control method of the region of interest. Background Art
[0002] With the development of electronic technology, electronic devices providing various functions are being developed. Recently, in the display field, screens are actively becoming larger. The demand for large screens is increasing not only in the home TV market but also in the outdoor industrial / advertising display (large format display (LFD) and LED sign) market.
[0003] Since power consumption increases as the screen size increases, problems such as carbon emissions may occur. Recently, leading countries have provided carbon emission regulations, requiring companies to conduct environmental, social, and corporate governance (ESG) management. In this case, it is necessary for display devices to also consume power effectively.
[0004] As a method for improving power consumption efficiency, for example, a display device may reduce power consumption by reducing the brightness of a background area, which is a remaining area outside a region of interest that is excluded content, while minimizing the perceived degradation of image quality. That is, it is necessary to sufficiently distinguish a region of interest and a background area from content. Summary of the Invention
[0005] Provided herein is an electronic device including: a memory configured to store a neural network model including a first network and a second network, where the neural network model includes weights; and at least one processor connected to the memory and configured to control the electronic device, where the at least one processor is configured to: obtain first description information corresponding to a first image by inputting the first image into the first network using the weights, obtain a second image based on the first description information, and obtain a third image representing a first region of interest of the first image by inputting the first image and the second image into the second network using the weights, where the weights of the neural network model are trained based on: i) a plurality of sample images, ii) a plurality of sample description information corresponding to the plurality of sample images, and iii) a sample region of interest of each sample image among the plurality of sample images.
[0006] In some embodiments, the at least one processor is further configured to: obtain a third image by inputting the first image into an input layer of the second network using the weights, and input the second image into an intermediate layer of the second network using the weights.
[0007] In some embodiments, the first description information includes at least one word, and at least one processor is further configured to obtain a second image by converting each of the at least one word into a corresponding color.
[0008] In some embodiments, at least one processor is further configured to: obtain a first image by reducing the original image to a preset resolution or by reducing the original image to a preset scaling ratio.
[0009] In some embodiments, at least one processor is further configured to enlarge a third image to correspond to the resolution of the original image.
[0010] In some embodiments, the third image depicts a first region of interest of the first image in a first color, and the third image depicts a background region in a second color, where the background region is the remaining region excluding the first region of interest of the first image, and the resolution of the third image is the same as the resolution of the first image.
[0011] In some embodiments, a first network is configured to train a first relationship of a plurality of sample description information of a plurality of sample images through an artificial intelligence algorithm, and a second network is configured to train a second relationship between each sample image of the plurality of sample images and a sample region of interest through an artificial intelligence algorithm, where each sample image corresponds to the sample description information among the plurality of sample description information.
[0012] In some embodiments, the first network and the second network are trained simultaneously.
[0013] In some embodiments, the first network includes a convolutional network and a plurality of long short-term memory networks (LSTMs), and the plurality of LSTMs are configured to output the first description information.
[0014] In some embodiments, at least one processor is further configured to: identify the remaining region excluding the first region of interest from the first image as the background region, and perform different image processing on the first region of interest and the background region.
[0015] This document also provides a control method for an electronic device, the control method including: obtaining first description information corresponding to a first image by inputting the first image into a first network included in a neural network model; obtaining a second image based on the first description information; and obtaining a third image showing a region of interest of the first image by inputting the first image and the second image into a second network included in the neural network model; and wherein, the neural network model is a model trained based on a plurality of sample images, a plurality of sample description information corresponding to the plurality of sample images, and a sample region of interest of each sample image among the plurality of sample images.
[0016] In some embodiments, obtaining the third image includes obtaining the third image by inputting the first image into the input layer of the second network and inputting the second image into the intermediate layer of the second network.
[0017] In some embodiments, the first description information includes at least one word, and obtaining the second image includes obtaining the second image by converting each word in the at least one word into a corresponding color.
[0018] In some embodiments, the control method includes obtaining the first image by scaling down the original image to a preset resolution or scaling down the original image to a preset scaling ratio.
[0019] In some embodiments, the control method includes enlarging the third image to correspond to the resolution of the original image.
[0020] In some embodiments, the preset resolution is 320 by 240.
[0021] In some embodiments, scaling down is configured to reduce power consumption by using the first image, where the first image has a low resolution.
[0022] In some embodiments, scaling down maintains the horizontal width of the original image and the vertical height of the original image.
[0023] In some embodiments, the resolution of the first image after scaling down is a first resolution of 320 by 240, and the third image has a full high definition (FHD) resolution after enlargement, where the original image has an FHD resolution.
[0024] In some embodiments, the at least one processor is further configured to: read the weights of the neural network model from the memory; and implement the first network and the second network in the at least one processor using the weights. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figures 1a to 1f is a diagram showing identifying a region of interest to assist in understanding the method of the present disclosure;
[0026] Figure 2 is a block diagram showing the configuration of an electronic device according to one or more embodiments of the present disclosure;
[0027] Figure 3 is a block diagram showing the detailed configuration of an electronic device according to one or more embodiments of the present disclosure;
[0028] Figure 4 is a diagram showing the operations and effects according to one or more embodiments of the present disclosure;
[0029] Figure 5is a diagram showing a detailed method for identifying a region of interest according to one or more embodiments of the present disclosure;
[0030] Figure 6 is a diagram showing a learning method of a neural network model for identifying a region of interest according to one or more embodiments of the present disclosure;
[0031] Figure 7 is a flowchart showing a method for identifying a region of interest according to one or more embodiments of the present disclosure;
[0032] Figures 8 to 10 is a diagram showing an effect according to one or more embodiments of the present disclosure; and
[0033] Figure 11 is a flowchart showing a control method of an electronic device according to one or more embodiments of the present disclosure. Detailed Description
[0034] Various modifications can be made to the exemplary embodiments of the present disclosure. Therefore, specific exemplary embodiments are shown in the drawings and described in detail in the detailed description. However, it should be understood that the present disclosure is not limited to specific exemplary embodiments, but includes all modifications, equivalents, and alternatives that do not depart from the scope and spirit of the present disclosure. In addition, well-known functions or structures are not described in detail because they would obscure the present disclosure with unnecessary details.
[0035] An object of the present disclosure is to provide an electronic device and a control method thereof that can more effectively identify a region of interest from an image.
[0036] The present disclosure will be described in detail below with reference to the drawings.
[0037] The terms used to describe one or more embodiments of the present disclosure are selected general terms that are currently widely used in consideration of their functions herein. However, the terms can change according to the intentions of those skilled in the relevant art, legal or technical interpretations, the emergence of new technologies, etc. In addition, in some cases, there may be arbitrarily selected terms, and in such cases, the meanings of the terms will be disclosed in more detail in the corresponding descriptions. Therefore, the terms used herein should not be simply understood as their names, but based on the meanings of the terms and the overall context of the present disclosure.
[0038] In the present disclosure, expressions such as "having", "may have", "including", "may include", etc. are used to specify the existence of corresponding characteristics (e.g., elements such as numerical values, functions, operations, or components), and do not exclude the existence or possibility of additional characteristics.
[0039] The expression "at least one of A and / or B" should be understood to indicate any one of "A" or "B" or "A and B".
[0040] Expressions such as "first", "second", "1st", "2nd", etc. used in this document can be used to refer to various elements, regardless of order and / or importance. In addition, it should be noted that these expressions are only used to distinguish one element from another element, rather than limiting the relevant elements.
[0041] Unless otherwise specified, singular expressions include plural expressions. It should be understood that terms such as "formed" or "including" are used in this document to specify the existence of characteristics, quantities, steps, operations, elements, components, or combinations thereof, and do not exclude the existence or possibility of adding one or more of other characteristics, quantities, steps, operations, elements, components, or combinations thereof.
[0042] In this disclosure, the term "user" may refer to a person who uses an electronic device or a device that uses an electronic device (e.g., an artificial intelligence electronic device).
[0043] Various embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0044] Figures 1a to 1f is a diagram showing a method of identifying a region of interest to help understand the present disclosure.
[0045] Recently, technologies for various methods have been developed to achieve a resolution that can improve immersion. For example, as Figure 1a shown, if a user views content, when a region of interest (salience), which is the most focused and viewed region, is detected, a more immersive and stereoscopic image quality experience can be provided to the user, and the image quality of the region of interest can be enhanced.
[0046] In addition, recently, as part of ESG management, the demand for improving the power consumption of display devices has been increasing. For this purpose, as Figure 1b shown, a method of reducing power consumption by reducing the brightness of a non-region of interest (background region) rather than the region of interest without a perceived deterioration in image quality is being used.
[0047] Considering the above, detecting the region of interest is important for achieving an immersive image quality and improving the effective power consumption, and it is necessary to improve the detection accuracy.
[0048] One method among the methods for detecting the region of interest is a method of finding a region having characteristics that a user can focus on by analyzing the feature information of an image within a single image. For example, the region of interest can be detected by image processing (e.g., but not limited to histogram, frequency domain analysis, etc.).
[0049] Alternatively, as Figure 1cAs shown, an artificial intelligence (AI)-based method using deep learning can be used to detect a region of interest.
[0050] Alternatively, as Figure 1d shown, to enhance detection performance, a multi-modal method that collects and uses feature information of a single image, flow (motion) information obtained from multiple images, speech information, etc. can be used to detect a region of interest. In Figure 1d , the two-dimensional information marked as spatial saliency is an example of the output of the second image and the text-to-image encoder 540-1 as described with respect to Figure 5 . The corresponding text description is not shown in Figure 1d .
[0051] When only using the feature information of a single image, there is a problem that the detection accuracy is relatively low. For example, if the speaker is the person on the right among the three people shown at the upper end of Figure 1e , if only using the feature information of a single image, the region of interest can be recognized as the lower left end of Figure 1e , but if speech information is further used, the region of interest can be recognized as the lower right end of Figure 1e . That is, the performance may be relatively low when only using the feature information of a single image compared to when speech information is additionally used.
[0052] Alternatively, for example, as shown at the upper end of Figure 1f , if the person on the right among two people is playing golf, if only using the feature information of a single image, the region of interest can be recognized as the lower left end of Figure 1f , but if motion information is additionally used, the region of interest can be recognized as the lower right end of Figure 1f . That is, the performance may be relatively low when only using the feature information of a single image compared to when motion information is additionally used.
[0053] As described above, the multi-modal method can improve the detection accuracy, but additional devices for information extraction are required. For example, to use speech information, a device for separating speech signals from an image is required, and to use flow information, a device for comparing memories storing multiple images and extracting information such as motion is required. In addition, a comprehensive device for comparing / analyzing information with each other and finally outputting it in one form is required. Ultimately, due to the increase in the cost and computational amount of the additional devices, the detection time may increase compared to when only using a single image.
[0054] Figure 2 is a block diagram showing the configuration of an electronic device 100 according to one or more embodiments of the present disclosure.
[0055] The electronic device 100 may identify a region of interest from an image. For example, the electronic device 100 may be a device for identifying a region of interest from an image, such as but not limited to the body of a computer, a set-top box (STB), a server, an AI speaker, etc. Specifically, the electronic device 100 may include a display, such as but not limited to a television (TV), a desktop personal computer (PC), a notebook, a smartphone, a tablet PC, a pair of smart glasses, a smart watch, etc., and may be a device for identifying a region of interest from a displayed image.
[0056] Reference Figure 2 , the electronic device 100 may include a memory 110 and a processor 120.
[0057] The memory 110 may refer to hardware that stores information such as data in an electrical or magnetic form for access by the processor 120, etc. To this end, the memory 110 may be implemented as at least one of non-volatile memory, volatile memory, flash memory, a hard disk drive (HDD), or a solid state drive (SSD), random access memory (RAM), read-only memory (ROM), etc.
[0058] In the memory 110, at least one instruction required for the operation of the electronic device 100 or the processor 120 may be stored. Here, the instruction may be a code unit indicating the operation of the electronic device 100 or the processor 120, and may be prepared in machine language, which is a language that a computer can understand. Alternatively, the memory 110 may store a plurality of instructions for performing a specific task of the electronic device 100 or the processor 120 as an instruction set.
[0059] The memory 110 may store data, which is information in units of bits or bytes that can represent characters, numbers, images, etc. For example, the memory 110 may store a neural network model, etc. Here, the neural network model may include a first network and a second network, and may be a model trained based on a plurality of sample images, a plurality of sample description information corresponding to the plurality of sample images, and sample regions of interest of the plurality of sample images.
[0060] The memory 110 may be accessed by the processor 120, and reading, writing, modifying, deleting, updating, etc. of the instruction, instruction set, or data may be performed by the processor 120.
[0061] The processor 120 may control the overall operation of the electronic device 100. Specifically, the processor 120 may control the overall operation of the electronic device 100 by connecting to each configuration of the electronic device 100. For example, the processor 120 may be connected to configurations such as the memory 110, a display (not shown), a communication interface (not shown), etc., and control the operation of the electronic device 100.
[0062] At least one processor 120 may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an integrated many-core (MIC), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. At least one processor 120 may control one or a random combination of other elements of the electronic device 100 and perform operations associated with communication or data processing. At least one processor 120 may execute at least one program or instruction stored in the memory. For example, at least one processor 120 may execute a method according to one or more embodiments of the present disclosure by executing at least one instruction stored in the memory.
[0063] If a method according to one or more embodiments of the present disclosure includes a plurality of operations, the plurality of operations may be executed by one processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are executed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be executed by a first processor, or the first operation and the second operation may be executed by a first processor (e.g., a general-purpose processor), and the third operation may be executed by a second processor (e.g., an artificial intelligence dedicated processor).
[0064] At least one processor 120 may be implemented as a single-core processor including one core or as at least one multi-core processor including a plurality of cores (e.g., homogeneous multi-core or heterogeneous multi-core). If at least one processor 120 is implemented as a multi-core processor, each of the plurality of cores included in the multi-core processor may include memories inside the processor, such as cache memory and on-chip memory, and a common cache shared by the plurality of cores may be included in the multi-core processor. In addition, each of the plurality of cores (or a part of the plurality of cores) included in the multi-core processor may independently read and execute program commands for implementing a method according to one or more embodiments, or read and execute program commands for implementing a method according to one or more embodiments of the present disclosure due to the interconnection of the whole (or a part) of the plurality of cores.
[0065] When a method according to one or more embodiments of the present disclosure includes a plurality of operations, the plurality of operations may be executed by one of the plurality of cores or by the plurality of cores included in the multi-core processor. For example, when a first operation, a second operation, and a third operation are executed by a method according to one or more embodiments, the first operation, the second operation, and the third operation may all be executed by a first core included in the multi-core processor, or the first operation and the second operation may be executed by a first core included in the multi-core processor, and the third operation may be executed by a second core included in the multi-core processor.
[0066] According to one or more embodiments, at least one processor 120 may refer to a system-on-chip (SoC) in which at least one processor and other electronic components are integrated, a single-core processor or a multi-core processor, or a core included in a single-core processor or a multi-core processor, and the core herein may be implemented as a CPU, GPU, APU, MIC, NPU, hardware accelerator, machine learning accelerator, etc., but is not limited to one or more embodiments of the present disclosure. However, for ease of description, the expression "processor 120" will be used hereinafter to describe the operation of the electronic device 100.
[0067] The processor 120 may obtain description information corresponding to the first image by inputting the first image into a first network included in the neural network model. Here, the description information may include at least one word. For example, the processor 120 may input the first image into the first network included in the neural network model and obtain description information corresponding to the first image, such as "The first person on the left is speaking among two people." Here, the first image may be an image directly displayed by the electronic device 100 or an image corresponding to the screen data provided by the electronic device 100 to the display device.
[0068] However, the above is not limited thereto, and the description information may include various languages.
[0069] The processor 120 may obtain a second image based on the description information. For example, if the description information includes at least one word, the processor 120 may obtain a second image by converting each of the at least one word into a corresponding color. In the example, the processor 120 may obtain a second image including color information of eleven colors by converting each word of "The first person on the left is speaking among two people" into a corresponding color.
[0070] The processor 120 may obtain a third image representing the region of interest of the first image by inputting the first image and the second image into a second network included in the neural network model. Here, the third image may represent the region of interest of the first image in a first color and represent the background region, which is the remaining region excluding the region of interest of the first image, and its resolution may be the same as that of the first image. In the example, the region of interest in the third image may be shown in white, and the background region, which is the remaining region excluding the region of interest, may be shown in black.
[0071] The neural network model can be a model trained based on multiple sample images, multiple sample description information corresponding to the multiple sample images, and sample regions of interest of the multiple sample images. For example, the first network can be configured such that the relationships of the multiple sample description information of the multiple sample images are learned through an artificial intelligence algorithm, and the second network can be configured such that the relationships between the sample regions of interest of the multiple sample images and the multiple sample description information for the multiple sample images can be trained through an artificial intelligence algorithm. In addition, the first network and the second network can be trained simultaneously.
[0072] The processor 120 can input the first image into the input layer of the second network and obtain a third image by inputting the second image into the intermediate layer of the second network. However, the above is not limited thereto, and the neural network model can be trained to obtain a third image by inputting the first image and the second image into the input layer of the second network.
[0073] The processor 120 can obtain the first image by shrinking the original image to a preset resolution or by shrinking the original image to a preset scaling ratio. For example, the processor 120 can obtain the first image by shrinking the original image to a resolution of 320×240. Since the operation of identifying the region of interest can use any image with a low resolution, the computational amount can be reduced through operations such as the above to reduce power consumption.
[0074] However, the above is not limited thereto, and the processor 120 can shrink the original image while maintaining the horizontal and vertical widths of the original image.
[0075] The processor 120 can enlarge the third image to correspond to the resolution of the original image. For example, if the original image has an FHD resolution and the first image obtained from the original image is shrunk to a resolution of 320×240, the third image can also be 320×240 resolution. The FHD (Full High Definition) resolution is 1920×1080 pixels. In this case, the processor 120 can enlarge the third image with a resolution of 320×240 to the FHD resolution.
[0076] The first network can include a convolutional network and multiple long short-term memory (LSTM), and the multiple LSTM can output description information. For example, each of the multiple LSTM can output a word. In this case, the processor 120 can obtain a sentence based on the multiple words output from the multiple LSTM. Alternatively, each of the multiple LSTM can output words, but the neural network model can be trained to sequentially output words to form a sentence.
[0077] The processor 120 may identify the remaining area outside the region of interest excluded from the first image as the background area, and perform different image processing on the region of interest and the background area. For example, the processor 120 may maintain the brightness of the region of interest and reduce the brightness of the background area, thereby reducing power consumption while minimizing the perceived degradation of image quality.
[0078] Functions associated with artificial intelligence according to the present disclosure may be operated by the processor 120 and the memory 110.
[0079] The processor 130 may be configured by one or more processors. The one or more processors may be a general-purpose processor such as a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP), a graphics dedicated processor such as a graphics processing unit (GPU) or a vision processing unit (VPU), or an artificial intelligence dedicated processor such as a neural processing unit (NPU).
[0080] The one or more processors may control the input data to be processed according to predefined operation rules or artificial intelligence models stored in the memory 110. Alternatively, if the one or more processors are artificial intelligence dedicated processors, the artificial intelligence dedicated processors may be designed as hardware structures dedicated to the processing of specific artificial intelligence models. The predefined operation rules or artificial intelligence models are characterized by structures created through learning.
[0081] The structure created through learning mentioned herein refers to a set of predefined operation rules or artificial intelligence models to perform desired features (or purposes), and is created as a basic artificial intelligence model trained by a learning algorithm using multiple learning data. The learning may be performed in the machine itself where artificial intelligence according to the present disclosure is executed, or through a separate server and / or system. Examples of learning algorithms may include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the above examples.
[0082] The artificial intelligence model may be formed by multiple neural network layers. Each of the multiple neural network layers may include multiple weight values (also referred to as weights), and perform neural network calculations through the calculation between the calculation results of the previous layer and the multiple weight values. The multiple weight values included in the multiple neural network layers may be optimized through the learning results of the artificial intelligence model. For example, the multiple weight values may be updated to reduce or minimize the loss value or cost value obtained by the artificial intelligence model during the learning process. In some embodiments, the processor 120 is configured to: obtain a third image by inputting the first image into the input layer of the second network using weights, and input the second image into the intermediate layer of the second network using weights.
[0083] An artificial neural network may include a deep neural network (DNN), and examples thereof may include a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, etc., but are not limited thereto.
[0084] Figure 3 is a block diagram showing a detailed configuration of an electronic device 100 according to one or more embodiments of the present disclosure.
[0085] The electronic device 100 may include a memory 110 and a processor 120. Additionally, referring to Figure 3 , the electronic device 100 may further include a display 130, a communication interface 140, a user interface 150, a microphone 160, a speaker 170, and a camera 180. A detailed description of parts that are repeated with the parts of the elements shown in Figure 3 will be omitted. Figure 2 the elements shown in
[0086] The display 130 may be a configuration for displaying an image and may be implemented as various types of displays, such as a liquid crystal display (LCD), an organic light emitting diode (OLED) display, and a plasma display panel (PDP). In the display 130, a driving circuit, a backlight unit, etc., which may be implemented in the form of an a-si TFT, a low temperature polycrystalline silicon (LTPS) TFT, an organic TFT (OTFT), etc., may be included. The display 130 may be implemented as a touch screen coupled with a touch sensor, a flexible display, a three-dimensional display (3D display), etc.
[0087] The communication interface 140 may be a configuration for performing communication with various types of external devices according to various types of communication methods. For example, the electronic device 100 may perform communication with a content server or a user terminal device through the communication interface 140.
[0088] The communication interface 140 may include a Wi-Fi module, a Bluetooth module, an infrared communication module, a wireless communication module, etc. Here, each communication module may be implemented in the form of at least one hardware chip.
[0089] The Wi-Fi module and the Bluetooth module may perform communication by the Wi-Fi method and the Bluetooth method, respectively. When using the Wi-Fi module or the Bluetooth module, various joining information such as a service set identifier (SSID) and a session key may be first sent and received, and after communicatively joining using it, various information may be sent and received. The infrared communication module may perform communication according to infrared communication (Infrared Data Association (IrDA)) technology, which wirelessly transmits data over a short distance by using infrared rays existing between visible light and millimeter waves.
[0090] In addition to the above communication methods, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards, such as but not limited to ZigBee, third generation (3G), 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), LTE-Advanced (LTE-A), fourth generation (4G), fifth generation (5G), etc.
[0091] Alternatively, the communication interface 140 may include a wired communication interface, such as but not limited to HDMI, DP, Thunderbolt interface, USB, RGB, D-SUB, DVI, etc.
[0092] In addition to this, the communication interface 140 may include at least one of a wired communication module that performs communication using a local area network (LAN) module, an Ethernet module, or a paired cable, a coaxial cable, an optical fiber cable, etc.
[0093] The user interface 150 may be implemented with buttons, a touchpad, a mouse, and a keyboard, or may also be implemented as a touch screen capable of performing a display function and an operation input function together therewith. Here, the buttons may be various types of buttons, such as mechanical buttons, touchpads, or wheels, which are formed at random areas on the front surface portion, side surface portion, rear surface portion, etc. of the exterior of the main body of the electronic device 100.
[0094] The microphone 160 may be configured to receive sound and convert it into an audio signal. The microphone 160 may be electrically connected to the processor 120 and may receive sound under the control of the processor 120.
[0095] For example, the microphone 160 may be formed as an integrated type integrated into the upper side, front surface direction, side surface direction, etc. of the electronic device 100. Alternatively, the microphone 160 may be provided in a remote controller or the like separated from the electronic device 100. In this case, the remote controller may receive sound through the microphone 160 and provide the received sound to the electronic device 100.
[0096] The microphone 160 may include various configurations, such as a microphone that collects sound in an analog form, an amplifier circuit that amplifies the collected sound, an A / D converter circuit that samples the amplified sound and converts it into a digital signal, a filter circuit that removes noise components from the converted digital signal, etc.
[0097] The microphone 160 may be implemented in the form of a sound sensor and may be any method as long as it is a configuration capable of collecting sound.
[0098] The speaker 170 may be an element that outputs not only various audio data processed in the processor 120 but also various notification sounds, voice messages, and the like.
[0099] The camera 180 may be a configuration for capturing a still image or a moving image. The camera 180 may capture a still image at a specific point in time, but may also capture still images continuously.
[0100] The camera 180 may capture the actual environment in the front direction of the electronic device 100 by capturing the front direction of the electronic device 100. The processor 120 may recognize an area of interest from an image captured by the camera 180.
[0101] The camera 180 may include a lens, a shutter, an aperture, a solid-state imaging device, an analog front end (AFE), and a timing generator (TG). The shutter may be configured to adjust the time when light reflected from a target enters the camera 180, and the aperture may be configured to adjust the amount of light incident on the lens by mechanically increasing or decreasing the size of the opening portion through which the light enters. The solid-state imaging device may be configured to output an image as an electrical signal based on the accumulation of light reflected from the target as photocharges through the photocharges. The TG may be configured to output a timing signal for reading out pixel data of the solid-state imaging device, and the AFE may be configured to digitize the electrical signal output from the solid-state imaging device by sampling.
[0102] The electronic device 100 as described above can reduce manufacturing costs and power consumption by identifying the region of interest from one image and comparing with the multi-mode method. In addition, the electronic device 100 can further identify the region of interest by using the description information corresponding to one image to improve the recognition function of the region of interest.
[0103] The following will be Figures 4 to 10 The operation of the electronic device is described in more detail. Figures 4 to 10 In the present invention, for the convenience of description, a single embodiment will be described. However, Figures 4 to 10 The individual embodiments of the present invention may be implemented in any combination.
[0104] Figure 4 are diagrams illustrating operations and effects according to one or more embodiments of the present disclosure.
[0105] Figure 4 The upper end shows a multimodal method, and the multimodal method can extract a single image from an input image, extract voice information, and extract a stream image.
[0106] When the region of interest is detected from a single image, the region of interest is detected from voice information, and motion information of a stream is detected, the integrated output module may output the region of interest from the detected information.
[0107] In this case, compared with the method of extracting a region of interest from a single image, extracting speech information and storing a stream image may further require memory capacity. That is, the manufacturing cost may increase. In addition, as the calculations for detecting the region of interest from the speech information and the calculations for detecting motion information from the stream image are added, the processing delay and power consumption can increase according to the increase in the amount of calculation.
[0108] Figure 4 The lower end shows the method of the present disclosure, and according to the present disclosure, a single image can be extracted from the input image and reduced. Therefore, when compared with the multi-mode method, since speech information is not extracted or since a stream image does not need to be extracted, the manufacturing cost on the hardware side can be reduced.
[0109] The processor 120 can generate description information (image description information) from a single image. The above operation can be a software process. Alternatively, separate hardware for neural network calculations, such as a neural processing unit (NPU), can be provided, but the hardware can be mass-produced without significantly increasing the manufacturing cost, and recently, for more devices that basically include an NPU, no further manufacturing cost may be incurred.
[0110] The processor 120 can extract image feature information from a single image. The processor 120 can output a region of interest from the image feature information and the image description information through an integration output module.
[0111] The operations of obtaining the image feature information and the image description information can be performed through neural network calculations, and a fast operation rate can be ensured when using a dedicated processor such as an NPU. In addition, since the calculations are performed for a single image, the power consumption can be reduced compared with the multi-mode method that performs calculations for multiple images.
[0112] However, by comparing with the operation of simply detecting a region of interest from a single image, the performance of generating image description information from a single image and identifying the region of interest can be further improved because the above operations are used in the operation of identifying the region of interest.
[0113] Figure 5 is a diagram showing a detailed method of identifying a region of interest according to one or more embodiments of the present disclosure.
[0114] The processor 120 can extract a single image from the input image and reduce the extracted image (510). For example, the processor 120 can capture a frame of the input image and reduce the frame with a resolution of 1920×1080 to a resolution of 320×240.
[0115] The processor 120 may generate description (image description) information (520) from an image. For example, the processor 120 may obtain the description information by inputting the image into a first network of a neural network model. Here, the first network may include a convolutional network and multiple LSTMs. The convolutional network may extract features of the image, and each of the multiple LSTMs may output a word from the features of the image as a text generator.
[0116] The processor 120 may extract image feature information (530) from the image. The processor 120 may generate an image (540-1) from the description information by executing an integrated output module (540), and obtain a black-and-white image (540-2) showing the region of interest from the image feature information and the image description information. For example, Figure 5 the operations 530 and 540-2 may be implemented using the first network of the neural network model, where the operation 530 may be a saliency encoder, and the operation 540-2 may be a saliency decoder. The saliency encoder may extract features in the image, and the saliency decoder may extract a black-and-white image showing the region of interest corresponding to the image.
[0117] That is, the operations after the reduction of the image may be operations of the neural network model. In some embodiments, the neural network model includes weights. The weights are part of the information for implementing the first network and the second network. As Figure 5 shown, in some embodiments, the first network is a CNN, followed by several LSTMs. The number of LSTMs may be adjusted to correspond to the typical number of words required to represent the scene in the image. In some embodiments, the second network is a CNN implementing the saliency encoder 530, which has an intermediate input into the saliency decoder 540-2 (implemented as a CNN).
[0118] In some embodiments, the CNN 530 is implemented with N1 layers, and each layer has N2 neurons. In some embodiments, the CNN 540-2 is implemented with N3 layers, where each layer has N4 neurons. In some embodiments, the text-to-image encoder 540-1 is implemented by a CNN having N5 layers, and each layer has N6 neurons. The CNN part of 520 is implemented with N7 layers, and each layer has N8 neurons. The example LSTM part of 520 is a recurrent neural network implemented with N9 nodes and N10 memory blocks in a chain structure.
[0119] Figure 6 is a diagram showing a learning method of a neural network model for identifying a region of interest according to one or more embodiments of the present disclosure.
[0120] The neural network model may include a first network and a second network. The neural network model may be a model trained based on multiple sample images, multiple sample description information corresponding to the multiple sample images, and sample regions of interest of the multiple sample images.
[0121] The first network may be trained by an artificial intelligence algorithm for the relationships of the multiple sample description information of the multiple sample images, and the second network may be trained by an artificial intelligence algorithm for the relationships between the multiple sample images and the sample regions of interest of the multiple sample images corresponding to the multiple sample description information.
[0122] In addition, the first network and the second network may be trained simultaneously.
[0123] For example, as Figure 6 shown, an image ① of two people including the person on the left being the speaker may be a sample image having ① "the first person on the left among the two people is speaking" as the sample description information corresponding to image ①, and a black-and-white image ① in which only the form of the person on the left is white and the remaining area is black is the sample region of interest corresponding to image ①. The first network may learn the relationship between image ① and ① "the first person on the left among the two people is speaking", and the second network may learn the relationship between image ① and black-and-white image ① for the information of imaging ① "the first person on the left among the two people is speaking". The difference between the correct answer of the predicted description information output by the first network and each predicted black-and-white image output by the second network may be defined as a loss, and learning may be performed in the direction of minimizing the two losses.
[0124] The above operations may be repeated with various sample learning data to perform the learning by the first network and the second network.
[0125] Figure 7 is a flowchart showing a method for identifying a region of interest according to one or more embodiments of the present disclosure.
[0126] The processor 120 may receive an input of a video image (S710). However, the above is not limited thereto, and the video image may be pre-stored in the memory 110 or captured in real time by the camera 180.
[0127] The processor 120 may capture a frame of the video image and downsize the captured image (S720). However, the above is not limited thereto, and the processor 120 may obtain screen data of one frame among the multiple frames included in the video image from the data corresponding to the video image.
[0128] The processor 120 may extract features from the downsized image (S730-1) and obtain description information such as a description statement or words from the downsized image (S730-2).
[0129] The processor 120 may integrate the image feature information and the description information through the integrated output module, and output an image showing the region of interest from it.
[0130] Figures 8 to 10 is a diagram showing the effects according to one or more embodiments of the present disclosure.
[0131] As Figure 8 At the upper end of, when the person on the right among two people faces the front direction and the person on the left faces the back direction, if only the feature information of a single image is used, the region of interest can be identified as Figure 8 at the lower left end of, but by additionally using the description information that only the person on the left among the two people faces the front direction, the region of interest can be identified as Figure 8 at the lower right end of.
[0132] As Figure 9 At the upper end of, when the region occupied by about two people among multiple people is large, when only the feature information of a single image is used, the region of interest can be identified as Figure 9 at the lower left end of, but the region of interest can be identified as Figure 9 at the lower right end of by additionally using the description information that the person on the left among the two people is the speaker.
[0133] If as Figure 10 At the upper end of, if only the person in the middle among multiple people is the speaker, then if only the feature information of a single image is used, the region of interest including multiple people can be identified as Figure 10 at the lower left end of, but by additionally using the description information that only the person in the middle is the speaker, the region of interest can be identified as Figure 10 at the lower right end of.
[0134] Figure 11 is a flowchart showing a control method of an electronic device according to one or more embodiments of the present disclosure.
[0135] First, description information corresponding to the first image may be obtained by inputting the first image into a first network included in the neural network model (S1110). Then, a second image may be obtained based on the description information (S1120). Then, a third image representing the region of interest of the first image may be obtained by inputting the first image and the second image into a second network included in the neural network model (S1130). Here, the neural network model may be a model trained based on a plurality of sample images, a plurality of sample description information corresponding to the plurality of sample images, and sample regions of interest of the plurality of sample images.
[0136] Here, obtaining the third image (S1130) may include obtaining the third image by inputting the first image into the input layer of the second network and inputting the second image into the intermediate layer of the second network.
[0137] The description information may include at least one word, and obtaining the second image (S1120) may include obtaining the second image by converting each of the at least one word into a corresponding color. For example, in some embodiments, the processor 120 is configured to obtain the second image by converting each of the at least one word into a corresponding color.
[0138] In addition, the method may further include obtaining the first image by reducing the original image to a preset resolution or reducing the original image proportionally to a preset scaling ratio.
[0139] Here, it may further include a step of enlarging the third image to correspond to the resolution of the original image.
[0140] The third image may represent the region of interest of the first image in a first color and represent the background region, which is the remaining region excluding the region of interest of the first image, in a second color, and its resolution may be the same as that of the first image. For example, the third image depicts the first region of interest of the first image in a first color, and the third image depicts the background region, which is the remaining region excluding the first region of interest of the first image, in a second color, and the resolution of the third image is the same as the resolution of the first image.
[0141] In addition, the first network may be trained by an artificial intelligence algorithm for the relationship of multiple sample description information of multiple sample images, and then the second network may be trained by an artificial intelligence algorithm for the relationship of multiple sample images of multiple sample description information and the sample regions of interest of multiple sample images.
[0142] Here, the first network and the second network may be trained simultaneously.
[0143] The first network may include a convolutional network and multiple LSTMs, and the multiple LSTMs may output description information. LSTM may also be referred to as a long short-term memory network.
[0144] In addition, the method may further include: identifying the remaining region excluding the region of interest from the first image as a background image, and performing different image processing on the region of interest and the background region.
[0145] For example, the control method may include: obtaining first description information corresponding to a first image by inputting the first image into a first network included in a neural network model (S1110); obtaining a second image based on the first description information (S1120); and obtaining a third image showing a region of interest of the first image by inputting the first image and the second image into a second network included in the neural network model (S1130). In some embodiments, the processor 120 executes the Figure 11 logical flow by reading the weights of the neural network model from the memory 110; and uses the weights of the neural network model to implement the first network 520 and the second networks (530 and 540-2) and the text-to-image encoder 540-1.
[0146] According to various embodiments of the present disclosure as described above, the electronic device may reduce manufacturing costs and power consumption by identifying a region of interest from one image and comparing it with a multi-mode method.
[0147] In addition, the electronic device may identify the region of interest by further using the description information corresponding to one image, and may improve the recognition performance of the region of interest.
[0148] According to one or more embodiments of the present disclosure, the above various embodiments may be implemented by software including instructions stored in a machine-readable storage medium (e.g., a computer). The machine may call the instructions stored in the storage medium, and as a device operable according to the called instructions, may include an electronic device according to the above embodiments (e.g., electronic device (A)). Based on the instructions executed by the processor, the processor may directly or using other elements execute functions corresponding to the instructions under the control of the processor. The instructions may include code generated by a compiler or executed by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Herein, "non-transitory" only means that the storage medium is tangible and does not include signals, and the term does not distinguish between data stored semi-permanently or temporarily stored in the storage medium.
[0149] According to one or more embodiments, a method according to the above various embodiments may be provided included in a computer program product. The computer program product may be exchanged as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or distributed online through an application store (e.g., PLAYSTORE TM ). In the case of online distribution, at least a part of the computer program product (e.g., a downloadable app) may be stored at least temporarily in a storage medium such as a memory of a manufacturer's server, an application store's server, or a relay server, or generated temporarily.
[0150] According to one or more embodiments of the present disclosure, the various embodiments described above can be implemented in a recordable medium readable by a computer or a computer-like device using software, hardware, or a combination of software and hardware. In some cases, the embodiments described herein can be implemented by the processor itself. According to a software implementation, embodiments such as the processes and functions described herein can be implemented with separate software. The corresponding software can perform one or more of the functions and operations described herein.
[0151] Computer instructions for performing processing operations in a device according to the various embodiments described above can be stored in a non-transitory computer-readable medium. When the computer instructions stored in such a non-transitory computer-readable medium are executed by a processor of a specific device, they can cause the specific device to perform processing operations from a device according to the various embodiments described above. A non-transitory computer-readable medium can refer to a medium that stores data semi-permanently rather than for a very short period of time, such as registers, caches, memories, etc., and can be read by a device. Specific examples of non-transitory computer-readable media can include, for example but not limited to, compact discs (CDs), digital versatile discs (DVDs), hard disks, Blu-ray discs, USBs, memory cards, ROMs, etc.
[0152] In addition, each of the elements (e.g., modules or programs) according to the various embodiments described above can be formed by a single entity or multiple entities, and some of the sub-elements described above can be omitted, or other sub-elements can be further included in the various embodiments. Alternatively or additionally, some elements (e.g., modules or programs) can be integrated into one entity to perform the same or similar functions as performed by the corresponding elements before integration. According to the various embodiments, the operations performed by modules, programs, or other elements can be executed sequentially, in parallel, repeatedly, or heuristically, or at least some of the operations can be executed in a different order, omitted, or different operations can be added.
[0153] Although the present disclosure has been shown and described with reference to various example embodiments of the present disclosure, it should be understood that the various example embodiments are intended to be illustrative and not restrictive. Those skilled in the art will understand that various changes in form and detail can be made therein without departing from the true spirit and full scope of the present disclosure (including the appended claims and their equivalents).
Claims
1. An electronic device, comprising: a memory configured to store a neural network model including a first network and a second network, wherein the neural network model includes weights; and at least one processor connected to the memory and configured to control the electronic device, wherein the at least one processor is configured to: obtain first description information corresponding to a first image by inputting the first image into the first network using the weights, obtain a second image based on the first description information; and obtain a third image representing a first region of interest of the first image by inputting the first image and the second image into the second network using the weights, wherein the weights of the neural network model are trained based on: i) a plurality of sample images, ii) a plurality of sample description information corresponding to the plurality of sample images, and iii) a sample region of interest of each sample image among the plurality of sample images.
2. The electronic device according to claim 1, wherein the at least one processor is further configured to: obtain a third image by inputting the first image into an input layer of the second network using the weights, and input the second image into an intermediate layer of the second network using the weights.
3. The electronic device according to claim 1, wherein the first description information includes at least one word, and the at least one processor is further configured to obtain the second image by converting each word of the at least one word into a corresponding color.
4. The electronic device according to claim 1, wherein the at least one processor is further configured to: obtain the first image by reducing the original image to a preset resolution, or reduce the original image to a preset scaling ratio.
5. The electronic device according to claim 4, wherein the at least one processor is further configured to enlarge the third image to correspond to the resolution of the original image.
6. The electronic device according to claim 1, wherein the third image depicts the first region of interest of the first image in a first color, and the third image depicts the background region in a second color, wherein the background region is the remaining region excluding the first region of interest of the first image, and the resolution of the third image is the same as the resolution of the first image.
7. The electronic device according to claim 1, wherein the first network is configured to train a first relationship of a plurality of sample description information of a plurality of sample images through an artificial intelligence algorithm, the second network is configured to train a second relationship of a plurality of sample images and a sample region of interest of each sample image among the plurality of sample images through an artificial intelligence algorithm, wherein each sample image corresponds to the sample description information among the plurality of sample description information.
8. The electronic device according to claim 7, wherein the first network and the second network are trained simultaneously.
9. The electronic device according to claim 1, wherein the first network includes a convolutional network and a plurality of long short-term memory networks LSTM, and the plurality of LSTM are configured to output the first description information.
10. The electronic device according to claim 1, wherein the at least one processor is further configured to: Identifying a remaining area outside a first region of interest in a first image as a background area, and Performing different image processing on the first region of interest and the background area.
11. A control method for an electronic device, the control method comprising: Obtaining first description information corresponding to a first image by inputting the first image into a first network included in a neural network model; Obtaining a second image based on the first description information; and Obtaining a third image showing a region of interest of the first image by inputting the first image and the second image into a second network included in the neural network model; and wherein the neural network model is a model trained based on a plurality of sample images, a plurality of sample description information corresponding to the plurality of sample images, and sample regions of interest of each of the plurality of sample images.
12. The control method according to claim 11, wherein obtaining the third image comprises: Obtaining the third image by inputting the first image into an input layer of the second network, and Inputting the second image into an intermediate layer of the second network.
13. The control method according to claim 11, wherein the first description information includes at least one word, and obtaining the second image comprises: obtaining the second image by converting each of the at least one word into a corresponding color.
14. The control method according to claim 11, further comprising: Obtaining the first image by reducing the original image to a preset resolution, or Reducing the original image to a preset scaling ratio.
15. The control method according to claim 14, further comprising enlarging the third image to correspond to the resolution of the original image.