Electronic device and method for displaying image search results
The electronic device enhances image search by allowing users to enlarge and display object areas within thumbnail images based on user input, addressing the challenge of confirming object details in constrained spaces, thereby simplifying object identification in multiple-image displays.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-06-04
AI Technical Summary
Existing image search systems struggle to effectively display and enlarge thumbnail images of objects corresponding to keywords, especially when multiple images are displayed simultaneously, making it difficult for users to confirm the appearance of the objects due to limited space and size constraints.
An electronic device and method that allows users to enlarge and display an area containing an object corresponding to a keyword in a thumbnail image based on user input, maintaining the layout of multiple thumbnail images on the screen, using techniques such as cropping and outpainting to adjust the image size and aspect ratio.
Enables users to easily and intuitively verify the detailed features of objects within thumbnail images by enlarging the relevant areas, while preserving the number of visible images, thus simplifying the identification of objects in a cluttered display.
Smart Images

Figure KR2025019834_04062026_PF_FP_ABST
Abstract
Description
Electronic device and method for displaying image search results
[0001] An electronic device and method for displaying image search results are disclosed. Specifically, an electronic device and method for providing a thumbnail image of an image containing an object corresponding to a keyword as an image search result are disclosed.
[0002] Image search using keywords can be utilized for the efficient search and management of multiple images stored in an electronic device. In this case, the electronic device may provide a thumbnail image of an image containing an object corresponding to the keyword as the result of the image search.
[0003] According to one aspect of the present disclosure, a method for displaying image search results is disclosed. In one embodiment, the method may include the step of obtaining a keyword for image search. In one embodiment, the method may include the step of identifying an image containing at least one object corresponding to the keyword among a plurality of previously stored images. In one embodiment, the method may include the step of displaying a thumbnail image of the identified image. In one embodiment, the method may include the step of obtaining user input for enlarging the thumbnail image. In one embodiment, the method may include the step of enlarging and displaying an area containing at least one object corresponding to the keyword of the thumbnail image based on the user input.
[0004] According to one aspect of the present disclosure, an electronic device for displaying image search results is disclosed. In one embodiment, the electronic device may include a memory for storing one or more instructions, a display, and at least one processor for executing one or more instructions stored in the memory. In one embodiment, by executing one or more instructions by at least one processor, the electronic device may obtain keywords for image search. In one embodiment, by executing one or more instructions by at least one processor, the electronic device may display a thumbnail image of an identified image through the display. In one embodiment, by executing one or more instructions by at least one processor, the electronic device may obtain user input for enlarging the thumbnail image. In one embodiment, by executing one or more instructions by at least one processor, the electronic device may enlarge and display an area containing at least one object corresponding to the keyword of the thumbnail image through the display based on user input.
[0005] According to one aspect of the present disclosure, a computer-readable recording medium may be provided having a program recorded thereon for executing any one of the methods described above and below for performing the operation of an electronic device.
[0006] FIG. 1 is a diagram for schematically explaining the operation of electronic devices according to one embodiment of the present disclosure.
[0007] FIG. 2 is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure displaying image search results.
[0008] FIG. 3 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to obtain an object detection result from an image.
[0009] FIG. 4 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure identifying the region of at least one object corresponding to a keyword.
[0010] FIG. 5 is a flowchart illustrating the operation of selecting a method for acquiring a thumbnail image by an electronic device according to one embodiment of the present disclosure.
[0011] FIG. 6 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to acquire a thumbnail image.
[0012] FIG. 7 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure cropping an area of an image.
[0013] FIGS. 8a, FIGS. 8b, and FIGS. 8c are drawings for explaining the operation of an electronic device identifying a crop area according to one embodiment of the present disclosure.
[0014] FIGS. 9a and 9b are drawings for illustrating crop areas for a plurality of objects according to one embodiment of the present disclosure.
[0015] FIG. 10 is a drawing for explaining the operation of an electronic device outpainting an image according to one embodiment of the present disclosure.
[0016] FIG. 11 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to select an outpainting method.
[0017] FIG. 12 is a drawing for explaining an extended area according to one embodiment of the present disclosure.
[0018] FIGS. 13a and FIGS. 13b are drawings for illustrating a thumbnail GUI according to one embodiment of the present disclosure.
[0019] FIG. 14 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.
[0020] FIG. 15 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.
[0021] The terms used in the embodiments of this specification have been selected to be as widely used as possible, taking into account the functions of the present disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description section of the relevant embodiments. Therefore, terms used in this specification should be defined not merely by their names, but based on their meanings and the overall content of the present disclosure.
[0022] Unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" may be understood to include plural objects. Thus, for example, the description "constituent surface" may include cases where it refers to one or more of such surfaces.
[0023] Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by a person skilled in the art as described in this specification.
[0024] In this disclosure, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "...part," "...module," etc., as used in this specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.
[0025] As used herein, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean “specifically designed to” in hardware. Instead, in some situations, the expression “system configured to” may mean that the system is “capable of” in conjunction with other devices or components. For example, the phrase “processor configured to perform A, B, and C” may mean a dedicated processor for performing the said operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or an application processor) capable of performing said operations by executing one or more software programs stored in memory.
[0026] It should be understood that in this disclosure, the blocks in each flowchart and combinations of flowcharts may be executed by one or more computer programs comprising computer-executable instructions. One or more computer programs may be stored all in a single memory or may be divided and stored in a plurality of different memories.
[0027] All functions or operations described in this disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors is a circuitry that performs processing and may include circuitry such as an AP (Application Processor), CP (Communication Processor), GPU (Graphical Processing Unit), NPU (Neural Processing Unit), MPU (Microprocessor Unit), SoC (System on Chip), IC (Integrated Chip), etc.
[0028] In the present disclosure, "model" or "Altificial Intelligence (AI) model" may refer to a set of functions or algorithms configured to perform a desired characteristic (or objective) by being learned using multiple learning data by a learning algorithm. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not necessarily limited thereto. In one embodiment, the AI model may be stored in the memory of an electronic device. However, it is not limited thereto, and the AI model may be stored on an external server, and the electronic device may transmit data input to the AI model to the server and receive data output from the AI model from the server.
[0029] In the present disclosure, a 'model' or 'artificial intelligence model' may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and can perform neural network operations through operations between the results of operations of a previous layer and the multiple weights. The multiple weights possessed by the multiple neural network layers may be optimized by the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Examples of models including multiple neural network layers include, but are not limited to, Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Restricted Boltzmann Machines (RBM), Deep Belief Networks (DBN), Bidirectional Recurrent Deep Neural Networks (BRDNN), and Deep Q-Networks.
[0030] In the present disclosure, "thumbnail image" or "thumbnail" may refer to an image that briefly displays the context of an original image, such as an image that has been reduced in size from the original image or has an aspect ratio different from that of the original image. In one embodiment, the thumbnail image may be used to represent or preview the original image in web pages, galleries, search results, etc. In one embodiment, the thumbnail image may include a continuous image composed of a plurality of images. For example, the thumbnail image may include an animated or GIF format thumbnail containing two or more images, but is not necessarily limited to the examples described above. Meanwhile, in the present disclosure, the thumbnail image may be referred to by various expressions representing the same or similar concept. The thumbnail image may be replaced by expressions such as "preview image," "representative image," "search result image," "tile image," etc., and is not limited to the examples described above.
[0031] In the present disclosure, 'crop' or 'cropping' may refer to an image processing method in which a specific area within an original image is cut out to leave only that specific area and remove the remaining part. In one embodiment, the portion remaining through image cropping may result in an image having an aspect ratio different from that of the original image. In one embodiment, the image cropping may be performed together with a resizing process in which only the size is changed while maintaining the aspect ratio of the portion remaining through cropping.
[0032] In the present disclosure, 'outpainting' may refer to image processing that generates new content by extending the boundaries or outer areas of an original image, or generates a new image containing new content. In one embodiment, through outpainting, new visual elements may be predicted and added to a new area extending beyond the boundaries of the original image. In one embodiment, the visual elements added through outpainting may form a natural extended image while maintaining continuity with the elements of the original image.
[0033] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement the invention. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.
[0034] The present disclosure will be described below with reference to the attached drawings.
[0035] FIG. 1 is a diagram for schematically explaining the operation of electronic devices according to one embodiment of the present disclosure.
[0036] In one embodiment, the electronic device (2000) may display a GUI (Graphic User Interface) (30) that displays the results of an image search. In one embodiment, the GUI (30) may be displayed on a screen (20) of an application (e.g., a gallery application) for searching or managing a plurality of images stored in advance on the electronic device (2000). In one embodiment, the electronic device (2000) may display a thumbnail image containing at least one object on the GUI (30). Here, the at least one object may refer to an object that is the target of the image search. For example, if the object that the user wants to search for is a dog, the electronic device (2000) may display a thumbnail image containing a dog as the result of the image search.
[0037] In one embodiment, the electronic device (2000) can obtain a keyword for image search. For example, if a user wants to obtain an image containing a puppy among a plurality of images pre-stored in the electronic device (2000) as an image search result, the user can obtain a text input containing the keyword "puppy" through a search input window (120) included in the application screen (20). The electronic device (2000) can obtain "puppy" as a keyword based on the obtained text input.
[0038] In one embodiment, the electronic device (2000) can identify an image containing at least one object corresponding to a keyword among a plurality of previously stored images. For example, an image DB (10) containing a plurality of images may be stored in the memory of the electronic device (2000). The image DB (10) may include object detection results corresponding to each of the plurality of images. The object detection results may include information about the objects included in each of the plurality of images. For example, the image DB (10) may store an object detection result, which includes information about a first object (Ob1) corresponding to a puppy included in the first image (101) and a second object (Ob2) corresponding to a ball, together with the first image (101) as metadata of the first image (101) in the image DB (10). The electronic device (2000) can identify a first image (101) containing a first object (Ob1) corresponding to the keyword "dog" in the metadata as an object detection result among a plurality of images stored in the image DB (10), as an image containing an object corresponding to the keyword. In the same way, other images other than the first image (101) containing an object corresponding to the keyword containing a dog among a plurality of images stored in the image DB (10) can be further identified.
[0039] In one embodiment, the electronic device (2000) may display a thumbnail image of an identified image. Here, the thumbnail image may be an image that includes an area of at least one object corresponding to a keyword in the identified image. For example, the electronic device (2000) may obtain a thumbnail image (111) (hereinafter referred to as the first thumbnail image) of the first image by cropping an area (CR1) in the first image (101) in which the area of the first object (Ob1) occupies more than a preset ratio (e.g., 0.3). In the same way, the electronic device (2000) may obtain thumbnail images (112, 113, 114) (hereinafter referred to as the second to fourth thumbnail images) of other identified images by cropping an area in which the area of the puppy occupies more than a preset ratio from other identified images other than the first image (101). The electronic device (2000) can display the acquired first to fourth thumbnail images (111, 112, 113, 114) on the GUI (30). Accordingly, the user can verify the appearance of the dog, which is the target they wish to search for among the multiple images stored in the image DB (10), through the first to fourth thumbnail images (111, 112, 113, 114) displayed on the GUI (30).
[0040] Meanwhile, if the size of the object corresponding to the keyword in the thumbnail image displayed through the electronic device (2000) is small, the user may have difficulty confirming the appearance of the object corresponding to the keyword. In particular, when multiple thumbnail images are displayed together on the GUI (30), the size of each thumbnail image is limited due to space constraints on the application screen (20), so the user may find it more difficult to confirm the appearance of the object corresponding to the keyword in the thumbnail images.
[0041] In one embodiment, the electronic device (2000) may obtain user input for enlarging a thumbnail image. Here, the enlarging of the thumbnail image may correspond to an increase in the proportion of the area occupied by the object corresponding to the keyword in the thumbnail image. For example, the electronic device (2000) may obtain input for enlarging the thumbnail image in the first to fourth thumbnail images (111, 112, 113, 114) through a user's gesture input, such as pinch-out, on the application screen (20).
[0042] In one embodiment, the electronic device (2000) can enlarge and display an area containing at least one object corresponding to a keyword in the thumbnail image based on user input for enlarging the thumbnail image. In other words, the electronic device (2000) can identify an area in the identified image where the area containing at least one object corresponding to the keyword occupies a proportion corresponding to the user input, and can display an enlarged thumbnail image obtained based on cropping the identified area. Here, the enlarged thumbnail image may be an image in which the proportion of the area of at least one object corresponding to the keyword occupied in the entire image area is increased compared to the thumbnail image. For example, the electronic device (2000) may identify a second area (CR2) of the first image (101) based on user input to enlarge the first thumbnail image (111), and display an enlarged thumbnail image (121) obtained based on cropping the second area (CR2) on the GUI (30) where the first thumbnail image (111) was displayed. In this case, the enlarged first thumbnail image (121) may be an image in which the proportion of the area occupied by the first object (Ob1) is increased compared to the first thumbnail image (111). In the same way, the electronic device (2000) may display enlarged second to fourth thumbnail images (122, 123, 124) on the GUI (30) where the second to fourth thumbnail images (112, 113, 114) were displayed. That is, the electronic device (2000) can display enlarged first to fourth thumbnail images (121, 122, 123, 124) in the areas where the first to fourth thumbnail images (111, 112, 113, 114) on the GUI (30) are displayed, while maintaining the layout of the areas where the first to fourth thumbnail images (111, 112, 113, 114) on the GUI (30) are displayed.
[0043] In this way, an electronic device (2000) according to one embodiment of the present disclosure can enlarge and display an object corresponding to a keyword in a thumbnail image based on user input. Accordingly, the user can more easily and simply check detailed features of an object corresponding to a keyword that are difficult to verify through the thumbnail image by looking at the enlarged thumbnail image.
[0044] In addition, when an electronic device (2000) according to one embodiment of the present disclosure displays a plurality of thumbnail images together, it can display enlarged thumbnail images while maintaining the layout of each area where the plurality of thumbnail images are displayed. Accordingly, the number of multiple thumbnail images that a user can view at once through the screen of the electronic device (2000) does not change, and the user can easily identify the appearance of objects corresponding to keywords in the plurality of thumbnail images.
[0045] FIG. 2 is a flowchart illustrating the operation of an electronic device according to one embodiment of the present disclosure displaying image search results.
[0046] With reference to FIG. 2, the operations of an electronic device (2000) displaying image search results are schematically described, and a detailed description of each of the operations will be described with reference to the drawings that follow.
[0047] In step S210, the electronic device (2000) can obtain keywords for image search.
[0048] In one embodiment, keywords for image search may include words, phrases, or clauses for referring to or identifying specific objects. In one embodiment, keywords for image search may be obtained based on user input containing at least one keyword. For example, keywords for image search may be obtained by extracting from text input or voice input containing one or more words and / or one or more sentences. As another example, keywords for image search may be obtained based on user input selecting a specific object included in an image. However, the foregoing examples are not limited to the foregoing examples, and keywords for image search may include various forms of data for identifying or determining specific objects.
[0049] In one embodiment, the electronic device (2000) can identify at least one object corresponding to a keyword for image search. In one embodiment, a plurality of object candidates composed of a plurality of objects included in the object detection results of a plurality of images stored in advance may be pre-set. In one embodiment, the electronic device (2000) can identify at least one object corresponding to a keyword by matching a keyword for image search to one of the plurality of object candidates. For example, the electronic device (2000) can match a keyword word for image search to one of the plurality of object candidates through a synonym dictionary or similar word mapping. As another example, the electronic device (2000) can identify the object with the highest similarity among the plurality of object candidates as at least one object by calculating a similarity based on word embedding for a keyword for image search. However, it is not necessarily limited to the examples described above, and the electronic device (2000) can identify at least one object corresponding to a keyword for image search by utilizing semantic analysis or a natural language processing model.
[0050] In one embodiment, at least one object corresponding to a keyword for image search may include a first object and a second object. In one embodiment, the keyword for image search may include a keyword corresponding to the first object and a keyword corresponding to the second object. In one embodiment, the electronic device (2000) may identify the first object and the second object as at least one object corresponding to the keyword based on the keyword corresponding to the first object and the keyword corresponding to the second object. In other words, the electronic device (2000) may acquire a plurality of keywords for image search based on user input including a plurality of keywords, and may identify an object corresponding to each of the acquired plurality of keywords. For example, if the electronic device (2000) acquires a text input of "dog and ball," it may acquire "dog" and "ball" as keywords for image search based on the acquired text input, and identify the objects corresponding to the keywords as "dog" and "ball."
[0051] In step S220, the electronic device (2000) can identify an image containing at least one object corresponding to a keyword among a plurality of previously stored images.
[0052] In one embodiment, the electronic device (2000) may store a plurality of images captured through a camera of the electronic device (2000) or a plurality of images received from an external electronic device in the memory of the electronic device (2000). In one embodiment, an object detection result obtained from a plurality of images of the electronic device (2000) may be mapped to or associated with each of the plurality of images and stored in the memory of the electronic device (2000). For example, the object detection result of each of the plurality of images may be stored as metadata of the plurality of images.
[0053] In one embodiment, the electronic device (2000) can identify an image containing at least one object corresponding to a keyword among a plurality of previously stored images based on the object detection result of each of the plurality of previously stored images. In one embodiment, the electronic device (2000) can identify a first image and a second image containing at least one object corresponding to a keyword among a plurality of previously stored images. In other words, among a plurality of previously stored images, there may be multiple images containing at least one object corresponding to a keyword as an object detection result. In this case, the electronic device (2000) can identify a plurality of images containing at least one object corresponding to a keyword. Hereinafter, the operation and function of the electronic device (2000) for the identified image described in this disclosure may be understood as an operation and function for each of the identified plurality of images.
[0054] In one embodiment, the electronic device (2000) can obtain object detection results for objects included in a plurality of pre-stored images from a plurality of pre-stored images. In one embodiment, the object detection result may include identification information of an object included in an image and location information of an object. For example, the object detection result may include identification information of an object, such as a label indicating the type or name of an object included in an image. Additionally, the object detection result may include location information of an object, such as the coordinates of the vertices of a bounding box surrounding an object or the coordinates of pixels constituting an object. However, it is not necessarily limited to the examples described above, and the object detection result may include data of various formats for identifying objects and locations detected from an image.
[0055] In one embodiment, the electronic device (2000) can obtain object detection results for objects included in a plurality of previously stored images based on user input specifying tags for objects included in an image. In one embodiment, the electronic device (1000) can obtain user input specifying tags for objects included in an image. In one embodiment, the user input specifying tags for objects may include user input specifying identification information of objects included in an image (e.g., name, type, etc. of the object) and location information of objects (e.g., coordinates of the object within the image, coordinates of a bounding box surrounding the object). For example, the electronic device (2000) can obtain user input specifying location information of objects, such as touching a specific object included in an image, dragging along the outline of a specific object, or creating a bounding box containing a specific object. Additionally, the electronic device (2000) can obtain user input specifying identification information of objects, such as selecting an object, designating the label of the selected object as 'puppy', or entering the text "puppy". The electronic device (2000) can obtain an object detection result based on the location information of the object and the identification information of the object obtained through user input that tags the object.
[0056] Specific details regarding how the electronic device (2000) obtains an object detection result from an image will be explained again below through FIG. 3.
[0057] In step S230, the electronic device (2000) can display a thumbnail image of the identified image.
[0058] In one embodiment, the thumbnail image may be an image containing an area of an object corresponding to a keyword. In one embodiment, the electronic device (2000) may identify an area of at least one object corresponding to a keyword in the identified image based on the identified image and the object detection result of the identified image. In one embodiment, the area of at least one object corresponding to a keyword may refer to an area of a bounding box surrounding at least one object corresponding to a keyword or a plurality of pixels constituting at least one object corresponding to a keyword. In one embodiment, if at least one object corresponding to a keyword includes a first object and a second object, the electronic device (2000) may identify an area containing both the first object and the second object in the identified image.
[0059] Specific details regarding the operation of the electronic device (2000) identifying the region of at least one object corresponding to a keyword in an identified image will be explained again below through FIG. 4.
[0060] In one embodiment, the electronic device (2000) can identify an area in an identified image in which the area of at least one object corresponding to a keyword occupies a preset ratio. In one embodiment, the electronic device (2000) can display a thumbnail image obtained based on cropping the identified area. Here, the statement that the area of at least one object corresponding to a keyword occupies more than a preset ratio in a specific area may mean that the area of at least one object corresponding to a keyword is entirely included in the specific area, while at the same time the ratio occupied by the area of at least one object corresponding to a keyword in the specific area is greater than or equal to a preset ratio.
[0061] In one embodiment, the electronic device (2000) can crop an identified area from an identified image. In one embodiment, the electronic device (2000) can obtain a thumbnail image based on adjusting the size of the cropped area. In one embodiment, the identified area may have a preset aspect ratio. Here, the preset aspect ratio may correspond to the aspect ratio of the area where the thumbnail image is displayed in the electronic device (2000). Hereinafter, for convenience of describing the invention, the preset aspect ratio may be referred to as the thumbnail aspect ratio. The thumbnail aspect ratio may correspond to a value obtained by dividing the width of the area where the thumbnail image is displayed by its height. In other words, when the aspect ratio of the identified area corresponds to the thumbnail aspect ratio, the electronic device (2000) can crop the identified area and adjust the size of the cropped area to obtain a thumbnail image.
[0062] In one embodiment, adjusting the size of the cropped area may mean an action or function of enlarging or shrinking the cropped area to match the size of the area where the thumbnail image was previously displayed. In other words, the electronic device (2000) can adjust the size of the cropped area to correspond to the size of the area where the thumbnail image is displayed.
[0063] Specific details regarding the method of obtaining a thumbnail image by cropping an identified image using an electronic device (2000) will be explained again below through FIGS. 7 and FIGS. 8a to 8c.
[0064] In one embodiment, the electronic device (2000) can crop an identified area from an identified image. In one embodiment, the electronic device (2000) can obtain a thumbnail image based on performing outpainting on the cropped area. In one embodiment, the electronic device (2000) can outpaint the remaining area excluding the cropped area in the thumbnail image so that the center of the area of at least one object corresponding to the keyword is located at the center of the thumbnail image. Here, the aspect ratio of the identified area may not correspond to the thumbnail aspect ratio. In other words, if the aspect ratio of the identified area does not correspond to the thumbnail aspect ratio, the electronic device (2000) can obtain a thumbnail image by cropping the identified area and performing outpainting on the cropped area.
[0065] Specific details regarding the method of obtaining a thumbnail image by performing outpainting on a cropped area of an electronic device (2000) will be explained again below through FIGS. 10 to 12.
[0066] In one embodiment, the electronic device (2000) may display a acquired thumbnail image. In one embodiment, the electronic device (2000) may display the acquired thumbnail image in an area on a GUI (hereinafter referred to as a thumbnail GUI) that displays image search results. In one embodiment, the thumbnail GUI may be displayed on the screen of an application (hereinafter referred to as an application) for searching or managing a plurality of pre-stored images. In one embodiment, the area on the thumbnail GUI where the thumbnail image is displayed may be an area having a preset aspect ratio and a preset size. For example, the area where the thumbnail image is displayed may be a square area of 50 x 50 pixels in size, but is not necessarily limited to the example described above.
[0067] In one embodiment, the electronic device (2000) may display a visual indicator on the thumbnail image that distinguishes the area of at least one object corresponding to a keyword included in the thumbnail image from other areas other than the area of at least one object. For example, the visual indicator may include an overlay element displayed on the area of the object corresponding to the keyword included in the thumbnail image, visual markup indicating the location of the object (highlight box, text label, icon, etc.), and is not necessarily limited to the examples described above. In one embodiment, the visual indicator may be displayed based on the ratio occupied by the area of at least one object corresponding to the keyword in the thumbnail image. For example, the visual indicator may be displayed when the ratio occupied by the area of at least one object corresponding to the keyword in the thumbnail image is less than a preset value, and may not be displayed when it is greater than or equal to the preset value.
[0068] In one embodiment, the electronic device (2000) may display a first thumbnail image of a first image and a second thumbnail image of a second image in a first area and a second area on a GUI representing image search results. In one embodiment, the thumbnail GUI may include a plurality of areas where a plurality of thumbnail images are displayed. For example, the thumbnail GUI may include a grid view layout composed of a plurality of areas where a plurality of thumbnail images are displayed in a grid form. However, it is not necessarily limited to the examples described above, and the thumbnail GUI may include a plurality of areas of various forms where a plurality of thumbnail images are displayed.
[0069] In step S240, the electronic device (2000) can obtain user input to enlarge a thumbnail image.
[0070] In one embodiment, the electronic device (2000) may acquire a user's touch input for a thumbnail GUI on which a thumbnail image is displayed as user input for enlarging the thumbnail image. For example, the electronic device (2000) may acquire a gesture input, such as a double tap or pinch-out, for a thumbnail GUI on which a thumbnail image is displayed as user input for enlarging the thumbnail image. However, it is not necessarily limited to the examples described above, and the electronic device (2000) may acquire various forms of user input input through an input interface as user input for enlarging the thumbnail image.
[0071] In one embodiment, the electronic device (2000) may obtain user input for reducing a thumbnail image. Here, the reduction of the thumbnail image may correspond to a decrease in the proportion of the area occupied by an object corresponding to a keyword in the thumbnail image. In other words, the electronic device (2000) may obtain user input for adjusting the proportion of the area occupied by at least one object corresponding to a keyword in a thumbnail image, including one of an input for enlarging the thumbnail image and an input for reducing the thumbnail image.
[0072] In one embodiment, user input for enlarging a thumbnail image (or user input for reducing a thumbnail image) may include information regarding a scale factor for determining a ratio corresponding to the user input to a keyword in the thumbnail image. For example, when the electronic device (2000) obtains a user's pinch-out (or pinch-in) gesture input for an application screen where a thumbnail GUI is displayed, it may determine a scale factor based on the distance of the pinch-out gesture input. The electronic device (2000) may determine a ratio corresponding to the user input based on the scale factor. For example, the electronic device (2000) may identify a value obtained by multiplying or adding the scale factor to the ratio occupied by the area of at least one object in the displayed thumbnail image as a ratio corresponding to the user input. However, it is not limited to the examples described above, and the electronic device (2000) may identify the scale factor itself obtained based on the user input as a ratio corresponding to the user input. In one embodiment, the ratio corresponding to the user input may correspond to the ratio occupied by the area of at least one object in the enlarged thumbnail image (or reduced thumbnail image) described later.
[0073] In step S250, the electronic device (2000) can enlarge and display an area containing at least one object corresponding to a keyword based on user input. In one embodiment, enlarge and displaying an area containing at least one object corresponding to a keyword may mean displaying an enlarged thumbnail image in which the proportion of the area occupied by at least one object corresponding to the keyword is increased.
[0074] In one embodiment, the electronic device (2000) can identify an area in an identified image where the area of at least one object corresponding to a keyword occupies a proportion corresponding to user input. In one embodiment, the electronic device (2000) can display an enlarged thumbnail image obtained based on cropping the identified area. Here, the fact that the area containing at least one object corresponding to a keyword occupies a proportion corresponding to user input may mean that the area of at least one object corresponding to a keyword is entirely contained within a specific area, while at the same time, the proportion of the area of at least one object corresponding to a keyword occupies within the specific area is greater than or equal to the proportion corresponding to user input.
[0075] In one embodiment, the electronic device (2000) can identify a region centered on a point closest to the center of the region of at least one object corresponding to the keyword. In other words, if the electronic device (2000) can identify multiple regions in which the region of at least one object corresponding to the keyword occupies a proportion corresponding to the user input in the identified image, the electronic device (2000) can identify the region among the multiple regions whose center is closest to the center of the region of at least one object.
[0076] In one embodiment, the electronic device (2000) can crop an identified area from an identified image. In one embodiment, the electronic device (2000) can obtain an enlarged thumbnail image based on adjusting the size of the cropped area. In one embodiment, the identified area may have a preset aspect ratio. Here, the preset aspect ratio may correspond to the aspect ratio of the area where the thumbnail image is displayed in the electronic device (2000). Hereinafter, for convenience of describing the invention, the preset aspect ratio may be referred to as the thumbnail aspect ratio. The thumbnail aspect ratio may correspond to a value obtained by dividing the width of the area where the thumbnail image is displayed by the height. In other words, when the aspect ratio of the identified area corresponds to the thumbnail aspect ratio, the electronic device (2000) can crop the identified area and adjust the size of the cropped area to obtain an enlarged thumbnail image. In one embodiment, the electronic device (2000) can adjust the size of the cropped area to correspond to the size of the display area of the thumbnail image. Accordingly, when the enlarged thumbnail image is displayed in the area where the previous thumbnail image is displayed, the user can perceive that the area containing at least one object corresponding to the keyword is enlarged through the area where the thumbnail image is displayed.
[0077] In one embodiment, the electronic device (2000) can crop an identified area from an identified image. In one embodiment, the electronic device (2000) can obtain an enlarged thumbnail image based on performing outpainting on the cropped area. In one embodiment, the electronic device (2000) can outpaint the remaining area excluding the cropped area in the enlarged thumbnail image so that the center of the area of at least one object corresponding to the keyword is located at the center of the enlarged thumbnail image. Here, the aspect ratio of the identified area may not correspond to the thumbnail aspect ratio. In other words, if the aspect ratio of the identified area does not correspond to the thumbnail aspect ratio, the electronic device (2000) can obtain an enlarged thumbnail image by cropping the identified area and performing outpainting on the cropped area.
[0078] In one embodiment, the enlarged thumbnail image may be displayed in the area where the thumbnail image is displayed. Accordingly, the user can observe the enlarged view of the object corresponding to the keyword through the enlarged thumbnail image.
[0079] In one embodiment, when the electronic device (2000) receives user input to reduce the thumbnail image, it may display a reduced thumbnail image instead of an enlarged thumbnail image according to the method described above. Here, the reduced thumbnail image may be an image in which the proportion of the area occupied by at least one object corresponding to the keyword is reduced compared to the thumbnail image. Accordingly, the user can more easily grasp the context of the image formed by elements not observed in the previously displayed thumbnail image through the reduced thumbnail image.
[0080] The operation of the electronic device (2000) cropping a portion of an image to obtain an enlarged thumbnail image (or a reduced thumbnail image) may correspond to the operation of the electronic device (2000) cropping a portion of an image to obtain a thumbnail image in step S220. Accordingly, specific details regarding this will be explained again below through FIGS. 7, 8a to 8c.
[0081] In one embodiment, the electronic device (2000) may display enlarged areas containing at least one object corresponding to a keyword of a first thumbnail image and a second thumbnail image in the first area and the second area, while maintaining the layout on the first area and the second area on the GUI representing image search results. Here, the first thumbnail image and the second thumbnail image may be thumbnail images of a first image and a second image identified as containing at least one object corresponding to a keyword among a plurality of images stored in advance. Additionally, the first area and the second area may correspond to areas on the thumbnail GUI where the first thumbnail image and the second thumbnail image are displayed. Here, maintaining the layout may mean that the size, shape, and distance from other areas of the plurality of areas where the thumbnail images are displayed do not change. In other words, the electronic device (2000) may display a plurality of enlarged thumbnail images as they are in the plurality of areas on the thumbnail GUI where the thumbnail images are displayed. Accordingly, users can more easily identify the appearance of the object being searched within the thumbnail images, without changing the number of thumbnail images that can be viewed at once through the application screen.
[0082] Specific details regarding the method of displaying multiple thumbnail images in multiple areas on the thumbnail GUI by the electronic device (2000) will be explained again below through FIG. 13a and FIG. 13b.
[0083] In one embodiment, the electronic device (2000) may display a thumbnail image containing both a first object and a second object and obtain user input for enlarging the thumbnail image. In one embodiment, the electronic device (2000) may identify a first area containing the first object and a second area containing the second object in the thumbnail image. In one embodiment, the electronic device (2000) may display an animation including a first frame corresponding to the first area and a second frame corresponding to the second area. Here, the first object and the second object may be included in a plurality of objects identified as corresponding to a keyword.
[0084] In one embodiment, the electronic device (2000) may identify, based on user input, an area in a thumbnail image where the area of a first object occupies a proportion corresponding to the user input as a first area, and an area where the area of a second object occupies a proportion corresponding to the user input as a second area. In one embodiment, the electronic device (2000) may crop the first area and the second area identified in the thumbnail image. In one embodiment, the electronic device (2000) may obtain a first frame and a second frame by adjusting the size of the cropped first area and the second area to correspond to the size of the area where the thumbnail image is displayed.
[0085] In one embodiment, the animation may further include a third frame corresponding to a third area of the thumbnail image. In one embodiment, the electronic device (2000) may identify a line connecting the center of the area of the first object and the center of the area of the second object in the thumbnail image. In one embodiment, the electronic device (2000) may identify the area where the center is located on the line identified in the thumbnail image as the third area. In one embodiment, the electronic device (2000) may crop the third area identified in the thumbnail image. In one embodiment, the electronic device (2000) may obtain a third frame by adjusting the size of the cropped third area to correspond to the size of the area where the thumbnail image is displayed.
[0086] In one embodiment, the animation may further include a fourth frame and a fifth frame corresponding to a fourth area and a fifth area of the thumbnail image. In one embodiment, the electronic device (2000) may identify a fourth area between a first area and a third area and a fifth area between a second area and a third area, with the center located on an identified line. In one embodiment, the electronic device (2000) may crop the fourth area and the fifth area identified in the thumbnail image. In one embodiment, the electronic device (2000) may obtain the fourth frame and the fifth frame by adjusting the size of the cropped fourth area and the fifth area to correspond to the size of the area where the thumbnail image is displayed. However, it is not necessarily limited to the examples described above, and the electronic device (2000) may obtain a plurality of other frames with the center located on an identified line in addition to the first to fifth frames in the same manner, and may display an animation that further includes the obtained plurality of frames. As a result, users can experience their gaze moving more naturally from the first object to the second object through animations containing more frames.
[0087] Specific details regarding how the electronic device (2000) displays animation will be explained again below through FIGS. 9a and 9b.
[0088] FIG. 3 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure to obtain an object detection result from an image.
[0089] In one embodiment, the electronic device (2000) can obtain an object detection result (320) of an image (310) from an image (310). In one embodiment, the electronic device (2000) can store the object detection result (320) of the image (310) together with the image (310) in an image DB (10) included in the memory of the electronic device (2000).
[0090] For example, the image (310) may be an image containing a first object (Ob1) corresponding to a 'puppy' and a second object (Ob2) corresponding to a 'ball'. The electronic device (2000) may obtain an object detection result (320-1) for the first object (Ob1) and an object detection result (320-2) for the second object (Ob2) based on the image (310). The object detection result (310-1) for the first object (Ob1) may include object identification information indicating that the first object (Ob1) corresponds to a 'puppy' and object location information indicating the location of a bounding box surrounding the first object (Ob1). Additionally, the object detection result (320-2) for the second object (Ob2) may include object identification information indicating that the second object (Ob2) corresponds to a 'ball' and object location information indicating the location of the bounding box surrounding the second object (Ob2).
[0091] In one embodiment, the location information of an object may include the coordinates of the top-left and bottom-right vertices of a bounding box. Here, the coordinates within the image may have the origin (0, 0) located at the top-left of the image, the x-axis value increasing as it moves to the right, and the y-axis value increasing as it moves downward. For example, the object detection result (320-1) for the first object (Ob1) may include the coordinates of the top-left vertex (350, 250) and bottom-right vertex (700, 400) of the bounding box surrounding the first object (Ob1). Additionally, the object detection result (320-2) for the second object (Ob2) may include the coordinates of the top-left vertex (300, 400) and bottom-right vertex (350, 450) of the bounding box surrounding the second object (Ob2). However, the object location information is not necessarily limited to the examples described above, and may include the center coordinates and width and height values of a bounding box surrounding the object, or the coordinates of multiple pixels constituting the object. In one embodiment, the object detection result (320) of the image (310) may be mapped to or associated with the image (310) and stored together in the image DB (10). For example, an electronic device (2000) may acquire the object detection result (320) of the image (310), store the object detection result (320) in the form of metadata of the image (310), and store it together with the image (310) in the image DB. However, it is not necessarily limited to the examples described above, and the object detection result (320) of the image (310) may be stored in the image DB (10) together with the image (310) as a separate external file (e.g., a JSON file, etc.) containing unique identification information of the image (310), so that the image (310) and the object detection result (320) may be mapped or associated with each other.
[0092] In one embodiment, an electronic device (2000) can obtain object detection results through an artificial intelligence model trained to detect objects included in an image using an image as input. For example, the electronic device (2000) can obtain object detection results as output by inputting a plurality of previously stored images into a deep learning-based object detection model. In one embodiment, the deep learning-based object detection model may include a Convolutional Neural Network (CNN)-based model and a Transformer-based model, but is not necessarily limited to the examples described above. In one embodiment, the deep learning-based object detection model may be trained based on a training dataset that includes a plurality of training images and ground-truth values of the training images. Here, the ground-truth values of the training images may include the actual labels of objects included in the training images and actual location information of the object's region (e.g., the location of a bounding box or the location of pixels). In one embodiment, a deep learning-based object detection model can be trained by predicting object detection results from each of the training images and using a loss function calculated based on the difference between the predicted object detection result and the correct value.
[0093] In one embodiment, an object detection result for an image in an electronic device (2000) may be obtained by an external electronic device that communicates with the electronic device (2000). In other words, according to the method described above, the external electronic device may obtain an object detection result (320) from an image (310) and transmit the image (310) and the object detection result (320) to the electronic device (2000). The electronic device (2000) may map or associate the image (300) and the object detection result (320) received from the external electronic device and store them in an image DB (10).
[0094] FIG. 4 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure identifying the region of at least one object corresponding to a keyword.
[0095] In one embodiment, the electronic device (2000) can identify an area (430) of at least one object corresponding to a keyword in the image (410) based on the image (410) and the object detection result (420) of the image. Here, the image (410) and the object detection result (420) of the image may correspond to the image (300) and the object detection result (320) of FIG. 3.
[0096] In one embodiment, the electronic device (2000) can identify an image containing at least one object corresponding to a keyword among a plurality of previously stored images. For example, if the object corresponding to the keyword is identified as containing at least one of 'dog' and 'ball', the electronic device (2000) can identify an image (410) containing 'dog' and 'ball' as object detection results (420) among a plurality of previously stored images.
[0097] In one embodiment, the electronic device (2000) can identify the region of at least one object corresponding to a keyword in the identified image based on the identified image and the object detection result of the identified image. Here, the region of at least one object corresponding to the keyword may mean a bounding box surrounding the object corresponding to the keyword or a plurality of pixels constituting the object corresponding to the keyword.
[0098] For example, if the object corresponding to the keyword is 'dog', the electronic device (2000) obtains bounding box location information (420-1) surrounding the first object (Ob1) corresponding to 'dog' from the object detection result (420), and can identify the bounding box (Tb1) surrounding the first object (Ob1) as the area of the object corresponding to the keyword based on the obtained bounding box location information (420-1). As another example, if the object corresponding to the keyword is 'ball', the electronic device (2000) obtains bounding box location information (420-2) surrounding the second object (Ob2) corresponding to 'ball' from the object detection result (420), and can identify the bounding box (Tb2) surrounding the second object (Ob2) as the area of the object corresponding to the keyword based on the obtained bounding box location information (420-2).
[0099] In one embodiment, when there are multiple objects corresponding to a keyword, the electronic device (2000) can identify an area that includes all of the areas of each of the multiple objects. For example, when the objects corresponding to the keyword are 'dog' and 'ball', the electronic device (2000) obtains bounding box location information (420-1) surrounding a first object (Ob1) corresponding to 'dog' and bounding box location information (420-2) surrounding a second object (Ob2) corresponding to 'ball' from the object detection result (420), and based thereon, can identify a bounding box (UTb) that includes both the first object (Ob1) and the second object (Ob2) as the area of the object corresponding to the keyword. Here, the bounding box (UTb) that includes both the first object (Ob1) and the second object (Ob) may be referred to as the area of multiple objects corresponding to the keyword. The area of multiple objects corresponding to the keyword can be determined as the smallest bounding box that includes both the bounding box (Tb1) surrounding the first object (Ob1) and the bounding box (Tb2) surrounding the second object (Ob2). Additionally, the electronic device (2000) can identify each of the bounding box (Tb1) surrounding the first object (Ob1) and the bounding box (Tb2) surrounding the second object (Ob2) as the area of the object corresponding to the keyword. Here, each of the bounding box (Tb1) surrounding the first object (Ob1) and the bounding box (Tb2) surrounding the second object (Ob2) can be referred to as the area of each individual object corresponding to the keyword. In other words, if there are multiple objects corresponding to the keyword, the electronic device (2000) can identify the area of the multiple objects corresponding to the keyword and the areas of the individual objects corresponding to the keyword.
[0100] Meanwhile, in the aforementioned examples, the region of the object corresponding to the keyword was described as being identified as the bounding box surrounding the object corresponding to the keyword within the image, but it is not necessarily limited thereto. For example, if the object detection result includes the coordinates of the pixels constituting the object corresponding to the keyword, the region of the object corresponding to the keyword may be identified as the pixels constituting the object corresponding to the keyword. Furthermore, if there are multiple objects corresponding to the keyword, the region of the multiple objects corresponding to the keyword may be determined as all the pixels constituting the multiple objects or as the smallest bounding box surrounding all the pixels constituting the multiple objects.
[0101] In one embodiment, information regarding the area (430) of at least one object corresponding to a keyword may be stored in association with or mapped to an image (410). For example, information regarding the area (430) of at least one object corresponding to an identified keyword may include information for identifying the location of the area (430) of at least one object corresponding to the keyword, such as the coordinates of the vertices of a bounding box corresponding to the area (430) of at least one object corresponding to the keyword, the coordinates of the center, the width, and the height, and may be stored as metadata of the image (410) and used to perform the operation and function of the electronic device (2000) of the present disclosure.
[0102] FIG. 5 is a flowchart illustrating the operation of selecting a method for acquiring a thumbnail image by an electronic device according to one embodiment of the present disclosure. Since steps S220 and S230 of FIG. 5 have been described above in FIG. 2, a redundant description will be omitted.
[0103] In step S510, the electronic device (2000) can identify whether the base thumbnail image of the identified image contains an area of at least one object corresponding to a keyword.
[0104] In one embodiment, the electronic device (2000) may obtain a basic thumbnail image of a plurality of images that are stored in advance. In one embodiment, the basic thumbnail image may be an image obtained from a specific image and stored in the electronic device (2000) when a specific image is stored in the electronic device (2000). In one embodiment, the basic thumbnail image may be an image obtained from a previously stored image and stored in the electronic device (2000) when an application that displays a screen including a thumbnail GUI is executed on the electronic device (2000). In one embodiment, the electronic device (2000) may display the basic thumbnail image on the thumbnail GUI before a keyword for image search is obtained. For example, when an application that displays a screen including a thumbnail GUI is executed, the basic thumbnail images of a plurality of images that are stored in advance in the electronic device (2000) may be displayed in a plurality of areas on the thumbnail GUI.
[0105] In one embodiment, the basic thumbnail image may be an image obtained by cropping a preset area based on the center of the original image. For example, the basic thumbnail image may be an image obtained by cropping a square area with a size of 50 x 50 pixels around the center of the original image.
[0106] In one embodiment, the basic thumbnail image may be an image obtained by cropping a square area having a size of 50 x 50 pixels around the center of the area of a specific object included in the original image. For example, the basic thumbnail image may be an image obtained by cropping a pre-set area centered on a representative person, who is one of a plurality of people included in the original image, or, if there is no person in the original image, an image obtained by cropping a pre-set area centered on an object. In one embodiment, the electronic device (2000) may identify the location of a person or object based on the object detection result included in the metadata of the original image, and identify a pre-set area centered on the identified person or object.
[0107] However, it is not necessarily limited to the examples described above, and the basic thumbnail image is an image obtained or displayed regardless of whether a keyword for image search is obtained, and may refer to an image representing a portion of an original image obtained in various ways. In one embodiment, information regarding the area cropped from the original image to the basic thumbnail image may be stored together with the original image and the basic thumbnail image. For example, location information such as the center, width, height, and coordinates of the vertices of the bounding box of the area cropped from the original image to the basic thumbnail image may be stored in the metadata of the original image and the basic thumbnail image.
[0108] In one embodiment, the electronic device (2000) can identify the ratio in which the area of at least one object corresponding to the keyword is included in the area cropped into the base thumbnail image from the identified image. Here, information regarding the area cropped into the base thumbnail image and information regarding the area of at least one object corresponding to the keyword can be obtained based on information stored in the metadata of the identified image. In one embodiment, the electronic device (2000) can compare the ratio in which the area of at least one object corresponding to the keyword is included in the area cropped into the base thumbnail image with a threshold value, and identify whether the area of at least one object corresponding to the keyword is included in the base thumbnail image based on the comparison result. For example, if the threshold value is 0.8 and the area of at least one object corresponding to the keyword is included in the area cropped into the base thumbnail image by a ratio of 0.6, the electronic device (2000) can identify that the area of at least one object corresponding to the keyword is not included in the base thumbnail image (S510-No). Conversely, if the area of at least one object corresponding to the keyword is included in the area cropped into the basic thumbnail image by a ratio of 0.9, the electronic device (2000) can identify that the area of at least one object corresponding to the keyword is included in the basic image (S510-Yes).
[0109] In step S520, the electronic device (2000) can acquire a basic thumbnail image as a thumbnail image that includes an area of at least one object corresponding to a keyword. In other words, if the basic thumbnail image contains an area of at least one object corresponding to a keyword above a threshold value, the user can easily identify the appearance of at least one object corresponding to the keyword using only the basic thumbnail image. Accordingly, the electronic device (2000) can use the basic thumbnail image without generating a new thumbnail image.
[0110] In step S530, the electronic device (2000) can obtain a thumbnail image containing at least one object corresponding to a keyword based on the identified image. The method by which the electronic device (2000) obtains the thumbnail image based on the identified image may correspond to the operation of the electronic device (2000) described in step S230 of FIG. 2.
[0111] In this way, an electronic device (2000) according to one embodiment of the present disclosure can utilize a basic thumbnail image of an image containing at least one object corresponding to a keyword. Accordingly, the electronic device (2000) can minimize the waste of hardware resources by avoiding unnecessary data processing and efficiently utilizing existing data.
[0112] FIG. 6 is a flowchart illustrating a method for an electronic device according to an embodiment of the present disclosure to acquire a thumbnail image. Step S610 of FIG. 6 may be an operation performed after Step S510 of FIG. 5 (S510-No). Additionally, the thumbnail image of FIG. 6 may be understood as a concept that includes both an enlarged thumbnail image and a reduced thumbnail image. That is, steps S610 to S640 of FIG. 6 may be performed in at least one of steps S230 and S250 of FIG. 2.
[0113] In step S610, the electronic device (2000) can identify a crop area in the identified image where the area of at least one object corresponding to the keyword occupies an area equal to the crop ratio. Here, the identified image may correspond to an image identified among a plurality of previously stored images as containing at least one object corresponding to the keyword. Additionally, the crop ratio may be a preset ratio for obtaining a thumbnail image or a ratio corresponding to user input for enlarging (or reducing) the thumbnail image. The specific operation of the electronic device (2000) identifying the crop area in the identified image will be explained again below in FIG. 7.
[0114] In step S620, the electronic device (2000) can identify whether the aspect ratio of the crop area corresponds to a preset aspect ratio. Here, the preset aspect ratio may correspond to a thumbnail aspect ratio. For example, the electronic device (2000) can identify that the aspect ratio of the crop area of the identified image corresponds to the preset aspect ratio if the value obtained by dividing the width of the crop area by the height matches the preset aspect ratio (S620-Yes). As another example, the electronic device (2000) can identify that the aspect ratio of the crop area does not correspond to the preset aspect ratio if the value obtained by dividing the width of the crop area by the height does not match the preset aspect ratio (S620-No).
[0115] In step S630, the electronic device (2000) can obtain a thumbnail image based on adjusting the size of the crop area if the aspect ratio of the crop area corresponds to a preset aspect ratio (S620-Yes). That is, the fact that the aspect ratio of the crop area corresponds to a preset aspect ratio means that a thumbnail image can be obtained in which the entire appearance of at least one object corresponding to the keyword can be verified by adjusting the size while maintaining the aspect ratio of the crop area. Accordingly, the electronic device (2000) can obtain a thumbnail image having a preset aspect ratio by adjusting the size of the crop area to correspond to the size of the thumbnail image, so that the area of at least one object corresponding to the keyword occupies an area equal to the crop ratio.
[0116] In step S640, if the aspect ratio of the crop area does not correspond to a preset aspect ratio (S620-No), the electronic device (2000) can obtain a thumbnail image by performing out-painting on the crop area. That is, the fact that the aspect ratio of the crop area does not correspond to a preset aspect ratio may mean that it is not possible to obtain a thumbnail image that can verify the full appearance of at least one object corresponding to the keyword by adjusting the size while maintaining the aspect ratio of the crop area. Accordingly, the electronic device (2000) can obtain a thumbnail image having a preset aspect ratio while including the area of at least one object corresponding to the keyword by performing out-painting on the crop area.
[0117] FIG. 7 is a diagram illustrating the operation of an electronic device according to one embodiment of the present disclosure cropping an area of an image.
[0118] In one embodiment, the electronic device (2000) can crop a crop area (720) from an image (710) to obtain an image (730) corresponding to the crop area (720) (S700). Here, the image (710) may correspond to an image identified as containing at least one object corresponding to a keyword. Additionally, the image (730) corresponding to the crop area (720) may correspond to a thumbnail image. Here, the thumbnail image may be understood as a concept that includes both an enlarged thumbnail image and a reduced thumbnail image.
[0119] In one embodiment, the electronic device (2000) can identify a crop area (720) from an image (710). In one embodiment, the crop area (720) can be identified as a region within the image (710). In one embodiment, the crop area (720) can be identified as a region where the region (711) of at least one object corresponding to the keyword occupies more than the crop ratio (713). In one embodiment, the crop area (720) can be identified such that the center of the crop area (720) is located closest to the center of the region (711) of at least one object corresponding to the keyword. Here, the crop ratio (713) may be the minimum ratio among the ratios preset to obtain a thumbnail image, or a ratio corresponding to user input for enlarging (or reducing) the thumbnail image. In one embodiment, the electronic device (2000) can identify the crop area (720) and obtain location information of the crop area (720). For example, the electronic device (2000) can obtain the center coordinates, width and height of the crop area (720), the coordinates of the vertices of the bounding box, etc.
[0120] In one embodiment, the electronic device (2000) can identify the area (711) of at least one object corresponding to a keyword based on information (712) regarding the area of at least one object corresponding to a keyword included in the image (710). In one embodiment, the electronic device (2000) can determine the maximum area of the crop area (720) by dividing the area (711) of at least one object corresponding to a keyword by a crop ratio (713). In one embodiment, the electronic device (2000) can determine the maximum width of the crop area (720) by the square root of the value obtained by multiplying the maximum area of the crop area (720) by a preset aspect ratio, and determine the maximum height of the crop area (720) by the value obtained by dividing the maximum width of the crop area (720) by a preset aspect ratio.
[0121] In one embodiment, the electronic device (2000) may determine the width of the crop area (720) to be the smaller value between the width of the image (710) and the maximum width of the crop area (720). In one embodiment, the electronic device (2000) may determine the height of the crop area (720) to be the smaller value between the height of the image (710) and the maximum height of the crop area (720). Accordingly, the crop area (720) may be identified within the internal area of the image (710).
[0122] In one embodiment, the electronic device (2000) can determine the location of the crop area (720) based on the width and height of the determined crop area (720). In one embodiment, the electronic device (2000) can position the center of the crop area (720) having the determined width and height at the center of the area (711) of at least one object corresponding to the keyword. In one embodiment, the electronic device (2000) can identify whether there is a vertex among the vertices of the bounding box of the crop area (720) that extends beyond the image (710). In one embodiment, if there is no vertex among the vertices of the bounding box of the crop area (720) that extends beyond the image (710), the electronic device (2000) can determine the location of the crop area (720) as is. In one embodiment, if there is a vertex among the bounding box vertices of the crop area (720) that extends beyond the image (710), the electronic device (2000) can move the position of the crop area (720) so that the vertex of the crop area (720) that extends beyond the image (710) is located within the image (710). Accordingly, the crop area (720) can be identified as an area centered on the point closest to the center of at least one object's area (711).
[0123] In one embodiment, the electronic device (2000) can identify the crop area (720) such that the center of the crop area (720) is closer to the center of the image (710). In one embodiment, after positioning the crop area (720) according to the method described above, the electronic device (2000) can move the crop area (720) within a range that includes at least one object area (711) inside the crop area (720). In this case, the electronic device (2000) can adjust the distance between the center of the crop area (720) and the center of the image (710) by moving the direction in which the crop area (720) is moved so that the center of the crop area (720) is closer to the center of the image (710).
[0124] In one embodiment, the aspect ratio of the identified crop area (720) may correspond to a preset aspect ratio. In one embodiment, the electronic device (2000) may crop the crop area (720) having a preset aspect ratio from the image (710). In one embodiment, the electronic device (2000) may adjust the size of the crop area (720) to correspond to the size of the area where the thumbnail image is displayed, thereby obtaining an image (730) corresponding to the crop area. In one embodiment, the electronic device (2000) may display the image (730) corresponding to the crop area as a thumbnail image or as an enlarged thumbnail image (or a reduced thumbnail image).
[0125] FIGS. 8a, FIGS. 8b, and FIGS. 8c are drawings for explaining the operation of an electronic device identifying a crop area according to one embodiment of the present disclosure.
[0126] Referring to FIG. 8a, the electronic device (2000) can identify a first crop area (810a) or a second crop area (810b) from an image (710). Here, the first crop area (810a) and the second crop area (810b) may be areas with a crop ratio of 0.3. In one embodiment, the electronic device (2000) can crop the first crop area (810a) or the second crop area (810b) from the image (710) to obtain an image (820a) corresponding to the first crop area or an image (820b) corresponding to the second crop area. In one embodiment, the electronic device (2000) may display an image (820a) corresponding to a first crop area or an image (820b) corresponding to a second crop area as a thumbnail image, or display it as an enlarged thumbnail image (or a reduced thumbnail image).
[0127] In one embodiment, the electronic device (2000) can identify a first crop area (810a) centered on the point closest to the center of the image (710) from the image (710). In other words, the electronic device (2000) can identify a first crop area (810a) centered on the point closest to the center of the image (710) where the ratio of the area (711) of at least one object corresponding to the keyword in the image (710) is 0.3. In one embodiment, elements highly relevant to the context of the image (710) are likely to be located close to the center of the image. In this case, the user can easily grasp the context of the image (710) while observing the appearance of at least one object corresponding to the keyword through the image (820a) corresponding to the first crop area.
[0128] In one embodiment, the electronic device (2000) can identify a second crop area (810b) centered on the point closest to the center of the area of at least one object from the image (710). In other words, the electronic device (2000) can identify a second crop area (810b) centered on the point closest to the center of the area (711) of at least one object corresponding to the keyword in the image (710), where the ratio of the area (711) of at least one object corresponding to the keyword is 0.3. In one embodiment, the object located in the center of the image is more likely to be easily observed by the user. In this case, the user can observe the appearance of at least one object corresponding to the keyword more closely through the image (820b) corresponding to the second crop area (810b).
[0129] In one embodiment, the electronic device (2000) may determine one of a first crop area (810a) or a second crop area (810b) as a crop area based on user input. For example, the electronic device (2000) may determine the first crop area (810a) as a crop area based on user input in which the center of the crop area is positioned at the center of the image (710), or determine the second crop area (810b) as a crop area based on user input in which the center of the crop area is positioned at the center of the area (711) of at least one object corresponding to the keyword.
[0130] In one embodiment, the electronic device (2000) may determine either a first crop area (810a) or a second crop area (810b) as a crop area based on a crop ratio. For example, the electronic device (2000) may determine the first crop area (810a) as a crop area if the crop ratio is less than a preset value, and determine the second crop area (810b) as a crop area if the crop ratio is greater than or equal to a preset value.
[0131] Referring to FIG. 8b, the electronic device (2000) can identify a third crop area (810c) or a fourth crop area (810d) from an image (710). Here, the third crop area (810c) may be an area with a crop ratio of 0.3, and the fourth crop area (810d) may be an area with a crop ratio of 0.5. In one embodiment, the electronic device (2000) can crop the third crop area (810c) or the fourth crop area (810d) from the image (710) to obtain an image (820c) corresponding to the third crop area or an image (820d) corresponding to the fourth crop area. In one embodiment, the electronic device (2000) may display an image (820c) corresponding to a third crop area or an image (820d) corresponding to a fourth crop area as a thumbnail image, or display an enlarged thumbnail image (or a reduced thumbnail image).
[0132] In one embodiment, the electronic device (2000) can identify a crop area centered on a point closer to the center of the area of at least one object corresponding to the keyword in the image (710) as the crop ratio increases. In other words, as the crop ratio increases, the distance between the center of the crop area in the image (710) and the center of the area of at least one object corresponding to the keyword (711) can become closer. For example, a third crop area (810c) identified when the crop ratio is 0.3 may be an area centered on a point that is a first distance away from the center of the area of at least one object corresponding to the keyword (711), and a fourth crop area (810d) identified when the crop ratio is 0.5 may be an area centered on a point that is a second distance shorter than the first distance away from the center of the area of at least one object corresponding to the keyword (711).
[0133] In this way, an electronic device (2000) according to one embodiment of the present disclosure can display the position of the center of at least one object in a thumbnail image as changing according to the crop ratio. Accordingly, a user who wants to enlarge the thumbnail image can focus more intently on the appearance of the object corresponding to the keyword, and a user who wants to reduce the thumbnail image can focus more intently on other elements excluding the object corresponding to the keyword through the thumbnail image.
[0134] Referring to FIG. 8c, the electronic device (2000) can identify a fifth crop area (810e) or a sixth crop area (810f) from an image (710). Here, the fifth crop area (810e) may be an area with a crop ratio of 0.7, and the sixth crop area (810f) may be an area with a crop ratio of 1.2. In one embodiment, the electronic device (2000) can crop the fifth crop area (810e) or the sixth crop area (810f) from the image (710) to obtain an image (820e) corresponding to the fifth crop area or an image (820f) corresponding to the sixth crop area. In one embodiment, the electronic device (2000) may display the image (820e) corresponding to the fifth crop area or the image (820f) corresponding to the sixth crop area as a thumbnail image, or display it as an enlarged thumbnail image (or a reduced thumbnail image).
[0135] In one embodiment, the electronic device (2000) can identify an area inside the region (711) of at least one object corresponding to a keyword in the image (710) as a crop area. In one embodiment, the crop ratio may have a value greater than 1. In one embodiment, if the crop ratio exceeds 1, the electronic device (2000) can identify an area inside the region (711) of at least one object corresponding to a keyword in the image (710) as a crop area. In this case, the value obtained by dividing the area of at least one object region (711) by the area of the identified crop area may correspond to the crop ratio. For example, the value obtained by dividing the area of at least one object region (711) corresponding to a keyword by the area of the fifth crop area (810e) identified when the crop ratio is 0.7 may be 0.7. As another example, the value obtained by dividing the area of at least one object area (711) corresponding to the keyword by the area of the sixth crop area (810f) identified when the crop ratio is 1.2 may be 1.2.
[0136] In this way, an electronic device (2000) according to one embodiment of the present disclosure can display a thumbnail image by expanding it into the area of at least one object corresponding to a keyword. Accordingly, the user can check even more detailed features of the object corresponding to the keyword through the thumbnail image.
[0137] FIGS. 9a and 9b are drawings for illustrating crop areas for a plurality of objects according to one embodiment of the present disclosure.
[0138] Referring to FIG. 9a, at least one object corresponding to a keyword may include a first object (Ob1) and a second object (Ob2). In this case, the electronic device (2000) may display a thumbnail image that includes both the first object (Ob1) and the second object (Ob2).
[0139] In one embodiment, the electronic device (2000) acquires a plurality of keywords for image search and can identify a plurality of objects corresponding to the keywords based on the acquired plurality of keywords. For example, if the electronic device (2000) acquires "dog" and "ball" as keywords for image search, it can identify at least one object corresponding to the keywords as "dog" and "ball". Here, the object corresponding to "dog" in the image (710) may be the first object (Ob1), and the object corresponding to "ball" may be the second object (Ob2).
[0140] In one embodiment, the electronic device (2000) can identify the area of a plurality of objects, which is a bounding box (UTb) that includes both the area of a first object (Ob1) and the area of a second object (Ob2), based on information regarding the area of an object corresponding to a keyword in the image (710). In one embodiment, the electronic device (2000) can identify the area of a plurality of objects corresponding to a keyword in the image (710) as a crop area, which is an area of the image (710) corresponding to a crop ratio.
[0141] For example, if the crop ratio is 0.5, the electronic device (2000) can identify a first crop area (910a) in the image (710) in which the ratio of the area of multiple objects corresponding to the keyword is 0.5. The electronic device (2000) can crop the first crop area (910a) from the image (710) to obtain an image (920a) corresponding to the first crop area. The electronic device (2000) can display the image (920a) corresponding to the first crop area as a thumbnail image or as an enlarged thumbnail image (or a reduced thumbnail image).
[0142] As another example, if the crop ratio is 0.8, the electronic device (2000) can identify a second crop area (910b) in the image (710) in which the ratio of the area of multiple objects corresponding to the keyword is 0.8. The electronic device (2000) can crop the second crop area (910b) from the image (710) to obtain an image (920b) corresponding to the second crop area. The electronic device (2000) can display the image (920b) corresponding to the second crop area as a thumbnail image or as an enlarged thumbnail image (or a reduced thumbnail image).
[0143] In this way, an electronic device (2000) according to one embodiment of the present disclosure can display a thumbnail image that includes all the regions of the plurality of objects when there are multiple objects corresponding to a keyword. Accordingly, even if the user enlarges or reduces the thumbnail image, the user can view all the multiple objects corresponding to the keyword together through the thumbnail image.
[0144] Referring to FIG. 9b, at least one object corresponding to a keyword may include a first object (Ob1) and a second object (Ob2). In this case, the electronic device (2000) may acquire a first frame (940a) containing the first object (Ob1) of the image (710) and a second frame (940b) containing the second object (Ob2), and may display an animation (940) containing the first frame (940a) and the second frame (940b). In one embodiment, the animation (940) may be displayed in an area where a thumbnail image is displayed.
[0145] In one embodiment, the animation (940) may include an image in which a plurality of frames arranged in chronological order are played continuously at preset time intervals. In one embodiment, the animation (940) may include a GIF image that returns to the first frame after the last frame ends and plays repeatedly. In one embodiment, the animation (940) may include an image in which a first frame (940a) and a second frame (940b) are displayed alternately. However, it is not necessarily limited to the examples described above, and the animation (940) may refer to an image or graphic representation in which a plurality of frames including a first frame (940a) and a second frame (940b) are displayed continuously in various ways.
[0146] In one embodiment, the electronic device (2000) may display an animation (940) including a first frame (940a) and a second frame (940b) after displaying a thumbnail image containing both the first object (Ob1) and the second object (Ob2) of FIG. 9a, when user input for enlarging the thumbnail image containing both the first object (Ob1) and the second object (Ob2) is obtained. In one embodiment, the electronic device (2000) may display a thumbnail image containing both the first object (Ob1) and the second object (Ob2) as in FIG. 9a, if the ratio corresponding to the user input is less than the maximum ratio (e.g., 0.8) that the area of the plurality of objects corresponding to the keyword in the thumbnail image can occupy. In one embodiment, the electronic device (2000) can display an animation (940) including a first frame (940a) and a second frame (940b) if the ratio corresponding to the user input is greater than or equal to the maximum ratio (e.g., 0.8) that the area of a plurality of objects corresponding to the keyword in the thumbnail image can occupy. In this case, the crop ratio for identifying the first frame (940a) and the second frame (940b) described later may be determined to have a value smaller than the maximum ratio that the area of a plurality of objects corresponding to the keyword in the thumbnail image can occupy.
[0147] In one embodiment, the electronic device (2000) may identify a bounding box (Tb1) surrounding a first object (Ob1) as the area of the first object (Ob1) and a bounding box (Tb2) surrounding a second object (Ob2) as the area of the second object (Ob2) based on information regarding the area of an object corresponding to a keyword of the image (710). In one embodiment, the electronic device (2000) may identify a first area (930a) and a second area (930b) from the image (710) in which the area of the first object (Ob1) and the area of the second object (Ob2) occupy a crop ratio. In one embodiment, the electronic device (2000) may crop the first area (930a) and the second area (930b) from the image (710). In one embodiment, the electronic device (2000) can obtain a first frame (940a) corresponding to the first area (930a) and a second frame (940b) corresponding to the second area (930b) by adjusting the sizes of the cropped first area (930a) and the second area (930b) to correspond to the size of the area where the thumbnail image is displayed. In one embodiment, the electronic device (2000) can generate an animation (940) including the first frame (940a) and the second frame (940b) and display the generated animation (940).
[0148] In one embodiment, the animation (940) may further include a third frame (940c) corresponding to a third region (930c) of the image (710). In one embodiment, the electronic device (2000) may identify a line (950) connecting the center of the region of the first object (Ob1) and the center of the region of the second object (Ob2) in the image (710). In one embodiment, the line (950) may be the shortest line connecting the center of the region of the first object (Ob1) and the center of the region of the second object (Ob2). In one embodiment, the electronic device (2000) may identify a third region (930c) whose center is located on the identified line (950). For example, the line (950) connecting the center of the area of the first object (Ob1) and the center of the area of the second object (Ob2) may be a line connecting P1, which is the center of the area of the first object (Ob1), and P2, which is the center of the area of the second object (Ob2) in the image (710). In one embodiment, the electronic device (2000) may adjust the size of the cropped third area (930c) to correspond to the size of the area where the thumbnail image is displayed, thereby obtaining a third frame (940c) corresponding to the third area (930c).
[0149] In one embodiment, the third region (930c) may be identified to have a size between that of the first region (930a) and the second region (930b). In one embodiment, the size of the third region (930c) may be determined such that as the position of P3 approaches the position of P1, it approaches the size of the first region (930a), and as the position of P3 approaches the position of P2, it approaches the size of the second region (930b). In the same way, the electronic device (2000) may identify regions centered between P1 and P3 and regions centered between P3 and P2 on the line (950), and obtain frames corresponding to the identified regions. Accordingly, the animation (940) can naturally express the process of a user moving their gaze or interest from the first object (Ob1) to the second object (Ob2) on the image (710).
[0150] In one embodiment, the animation (940) may include a plurality of frames obtained by cropping an area of the same size from the image (710). In other words, the electronic device (2000) may obtain a plurality of frames of the animation (940) by cropping an area of the same size from the image (710). In one embodiment, the size of the area cropped from the image (710) may be determined as the largest area among the areas of a plurality of objects corresponding to the keyword.
[0151] For example, the electronic device (2000) can compare the size of the area of the first object (Ob1) and the size of the area of the second object (Ob2), and identify a first area (930a) in which the area of the first object (Ob1), which is larger in size, occupies an area equal to the crop ratio. Then, the electronic device (2000) can identify a fourth area (not shown) that includes the area of the second object (Ob2) while having the same size as the first area (930a). Additionally, the electronic device (2000) can identify a fifth area (not shown) that has the same size as the first area (930a) and has its center located on the line (950).
[0152] In one embodiment, the electronic device (2000) can crop a first area (930a), a fourth area (not shown), and a fifth area (not shown) to obtain a first frame (940a) corresponding to the first area (930a), a fourth frame (not shown) corresponding to the fourth area (not shown), and a fifth frame corresponding to the fifth area (not shown). In one embodiment, the electronic device (2000) can display an animation (940) including the first frame (940a), the fourth frame (not shown), and the fifth frame (not shown). In other words, the electronic device (2000) can display an animation (940) including the fourth frame (not shown) instead of the second frame (940b) and the fifth frame (not shown) instead of the third frame (940c). In this case, the animation (940), including the first frame (940a), the fourth frame (not shown), and the fifth frame (not shown), includes multiple frames obtained by cropping an area of the same size from the image (710), thereby preventing visual fatigue that may occur from the user's perspective due to changes in the size of the cropped area.
[0153] In one embodiment, the electronic device (2000) may determine the order in which multiple frames are arranged in an animation (940) based on the order in which multiple keywords are included in a user input containing multiple keywords. Here, the order in which multiple keywords are included in the user input may be referred to by various expressions representing the same or similar concepts, such as the order in which multiple keywords appear, the order in which multiple keywords emerge, or the order in which multiple keywords appear. In one embodiment, the electronic device (2000) may arrange multiple frames containing each of the objects corresponding to the multiple keywords in the animation (940) so as to correspond to the order in which multiple keywords are included in the user input. For example, the electronic device (2000) may acquire a text input "puppy and ball" containing multiple keywords and, based on the acquired text input, identify that the multiple objects corresponding to the keywords are 'puppy' and 'ball'. In this case, the order in which multiple keywords are included in the user input may be such that 'ball' is positioned after 'puppy'. The electronic device (2000) can determine the order in which a plurality of frames are arranged in the animation (940) such that a second frame (940b) containing a second object (Ob2) corresponding to a 'ball' is positioned after a first frame (940a) containing a first object (Ob1) corresponding to a 'puppy'. Accordingly, the electronic device (2000) can display the animation (940) in the order of the first frame (940a), the third frame (940C), and the second frame (940b).
[0154] Accordingly, an electronic device (2000) according to one embodiment of the present disclosure can display an animation composed of frames including each of the plurality of objects when there are multiple objects corresponding to a keyword. Accordingly, the user can view the detailed appearance of each of the plurality of objects in more detail through the animation.
[0155] FIG. 10 is a drawing for explaining the operation of an electronic device outpainting an image according to one embodiment of the present disclosure.
[0156] In one embodiment, the electronic device (2000) can obtain an outpainted image (1020) based on performing outpainting on a crop area (1010) (S1000). In other words, the electronic device (2000) can obtain an outpainted image (1020) by extending the boundary of the crop area (1010). Here, the crop area (1010) may correspond to the crop area (720) cropped from the image (710) in FIG. 7. Additionally, the outpainted image (1020) corresponds to a thumbnail image, and the thumbnail image may be understood as a concept that includes both an enlarged thumbnail image and a reduced thumbnail image.
[0157] In one embodiment, the electronic device (2000) can outpaint the remaining area excluding the crop area (1010) such that the center of the area (1011) of at least one object corresponding to the keyword is located at the center of the outpainted image (1020).
[0158] In one embodiment, the electronic device (2000) can identify whether the aspect ratio of the crop area (1010) corresponds to a preset aspect ratio (1012). Here, the preset aspect ratio (1012) may correspond to a thumbnail aspect ratio. In one embodiment, if the crop area (1010) does not correspond to a preset aspect ratio (1012), the electronic device (2000) can perform outpainting on the crop area (1010).
[0159] In one embodiment, the electronic device (2000) can adjust the size of the crop area (1010) to correspond to the size of the area where the thumbnail image is displayed. For example, the electronic device (2000) can calculate a resizing scale which is the value obtained by dividing the larger value between the width and height of the thumbnail image by the larger value between the width and height of the crop area (1010), and adjust the size of the crop area (1010) by the resizing scale.
[0160] In one embodiment, the electronic device (2000) can align the position of the resized crop area (1010) such that the center of the area (1011) of at least one object corresponding to a keyword in the resized crop area (1010) is located at the center of the outpainted image (1020).
[0161] In one embodiment, the electronic device (2000) can identify an extension area (1021) such that the aspect ratio of the outpainted image (1020) corresponds to a preset aspect ratio (1012) based on a cropped area (1010) that is aligned in position. Here, the extension area (1021) may correspond to the remaining area of the outpainted image (1020) excluding the cropped area (1010). In one embodiment, the electronic device (2000) can obtain an outpainted image (1020) by performing image outpainting on the extension area (1021).
[0162] In one embodiment, the electronic device (2000) can perform outpainting on a crop area (1010) through image matching based on a wide-angle shot of an image identified as including an area (1011) of at least one object corresponding to a keyword. In one embodiment, the electronic device (2000) can store a wide-angle shot of each of a plurality of previously stored images together with the plurality of images. In one embodiment, the wide-angle shot may be acquired at the time when a specific image is captured. For example, if the camera of the electronic device (2000) includes a multi-lens camera, the electronic device (2000) can acquire a basic image through the basic lens of the multi-lens camera and acquire a wide-angle image of the basic image through the wide-angle lens. In one embodiment, the wide-angle shot may be an image captured with a wider field of view than the basic image. Accordingly, the wide-angle shot may include additional information about an area not included in the basic image. In one embodiment, the electronic device (2000) can map or associate a wide-angle image with a acquired base image and store the base image and the wide-angle image together in an image DB.
[0163] In one embodiment, the electronic device (2000) can acquire a wide-angle image of an image identified as including an area (1011) of at least one object corresponding to a keyword. In one embodiment, the electronic device (2000) can identify an area corresponding to an extended area (1021) in the acquired wide-angle image. For example, the electronic device (2000) can perform image alignment to set the crop area (1010) and the wide-angle image to the same coordinate system. In one embodiment, the electronic device (2000) can identify the coordinates of an area corresponding to the extended area (1021) of the wide-angle image based on the coordinates of the extended area (1021) for the crop area (1010) where image alignment was performed. In one embodiment, the electronic device (2000) can perform outpainting on the crop area (1010) by determining the pixel value of the expanded area (1021) in the area corresponding to the expanded area (1021) in the acquired wide-angle captured image.
[0164] In one embodiment, an electronic device (2000) can obtain an outpainted image (1020) through an artificial intelligence model trained to perform outpainting on a cropped area (1010) using a cropped area (1010) and an extended area (1021) as inputs. For example, the electronic device (2000) can obtain an outpainting result as an output by inputting a plurality of previously stored images into a deep learning-based outpainting model. In one embodiment, the deep learning-based outpainting model may include a Generative Adversarial Network (GAN)-based model, a Diffusion-based model, and a Transformer-based model, but is not necessarily limited to the examples described above. In one embodiment, the deep learning-based outpainting model may be trained based on a training dataset containing a plurality of training images and ground-truth values of the training images. Here, the ground-truth values of the training images may include actual pixel values of missing areas included in the training images and actual location information of the corresponding areas. In one embodiment, a deep learning-based outpainting model can be trained by predicting outpainting results from each of the training images and through a loss function calculated based on the difference between the predicted outpainting result and the correct value.
[0165] FIG. 11 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to select an outpainting method.
[0166] In step S1110, the electronic device (2000) can identify an extended area. In one embodiment, the electronic device (2000) can identify a cropped area from an image and identify whether the aspect ratio of the identified cropped area corresponds to a preset aspect ratio (thumbnail aspect ratio). In one embodiment, if the aspect ratio of the cropped area does not correspond to a preset aspect ratio, the electronic device (2000) can identify an extended area that causes the aspect ratio of the outpainted image to correspond to a preset aspect ratio.
[0167] In step S1120, the electronic device (2000) can identify whether a wide-angle shot of the identified image exists. In one embodiment, the wide-angle shot of the identified image can be captured together with the identified image and stored by mapping to or associating with the identified image.
[0168] In step S1130, if the electronic device (2000) has a wide-angle image of the identified image (S1120-Yes), it can perform outpainting through image matching based on the wide-angle image. In one embodiment, the electronic device (2000) can identify an area of the wide-angle image corresponding to the expanded area and determine the pixel value of the expanded area based on the pixel value of the identified area of the wide-angle image. In one embodiment, step S1140 may not be included. In this case, the electronic device (2000) can perform step S1160 after step S1130.
[0169] In step S1140, the electronic device (2000) can identify whether there is a remaining area in the extended region where outpainting has not been performed. In one embodiment, some of the extended region may represent an area that is not included in the wide-angle image. Accordingly, some of the extended region may not be outpainted through image matching based on the wide-angle image.
[0170] In step S1150, if the electronic device (2000) does not have a wide-angle shot image of the identified image (S1120-No) or if there is a remaining area where outpainting has not been performed (S1140-Yes), it can outpaint the extended area or the remaining area where outpainting has not been performed in the extended area through an image generation model.
[0171] In step S1160, if there is no remaining area in the expanded region where outpainting has not been performed (S1140-No), the electronic device (2000) can acquire the outpainted image as a thumbnail image. Here, the thumbnail image can be understood as a concept that includes both the enlarged thumbnail image and the reduced thumbnail image.
[0172] FIG. 12 is a drawing for explaining an extended area according to one embodiment of the present disclosure.
[0173] Referring to FIG. 12, the electronic device (2000) can identify an extension area that corresponds to a preset aspect ratio of the aspect ratio of the outpainted image based on the crop area (1010).
[0174] In one embodiment, the electronic device (2000) can identify a first extended area (1210a) having a minimum width, wherein the aspect ratio of the first outpainted image (1220a) corresponds to a preset aspect ratio. For example, an outpainted image (1220a) having a preset aspect ratio can be obtained by extending the boundary only in the vertical direction without extending the boundary in the horizontal direction relative to the crop area (1010). Accordingly, the electronic device (2000) can identify a first extended area (1210a) having a minimum width, wherein the boundary is extended only in the vertical direction from the crop area (1010). In one embodiment, when outpainting is performed on the first extended area (1210a) having a minimum width, an image with a shorter processing speed and higher consistency with the relationship with the crop area (1010) can be obtained, given that the area where outpainting is performed is small.
[0175] In one embodiment, the electronic device (2000) can identify a second extended area (1010b) such that the aspect ratio of the second outpainted image (1220b) corresponds to a preset aspect ratio, and the area (1011) of at least one object corresponding to the keyword in the outpainted image (1220b) occupies a ratio equal to the crop ratio. For example, if the crop ratio is 0.5, the second outpainted image (1220b) obtained by performing outpainting on the second extended area (1010b) may be an image in which the area (1011) of at least one object corresponding to the keyword occupies 0.5. Accordingly, the electronic device (2000) can cause the area (1011) of at least one object corresponding to the keyword to display a preset ratio or a ratio corresponding to user input even on the thumbnail image obtained through outpainting.
[0176] FIGS. 13a and FIGS. 13b are drawings for illustrating a thumbnail GUI according to one embodiment of the present disclosure.
[0177] Referring to FIG. 13a, the electronic device (22000) can display an application screen. In one embodiment, the application screen may include an execution screen of an application for searching or managing a plurality of images previously stored in the electronic device (2000).
[0178] In one embodiment, the electronic device (2000) may display an input field (1310) on an application screen. In one embodiment, the input field (1310) may represent a UI for receiving text containing keywords from a user. In one embodiment, the user may input text containing keywords through the input field (1310). For example, the user may input the text "puppy" into the input field (1310), and the electronic device (2000) may identify the object corresponding to the keyword as 'puppy' based on the user's text input.
[0179] In one embodiment, the electronic device (2000) can acquire a voice signal of a user uttering a keyword through an input interface (e.g., a microphone). In one embodiment, the electronic device can convert the user's voice signal into text using Automatic Speech Recognition. In one embodiment, the electronic device (2000) can extract a keyword from the user's voice signal converted into text and identify an object corresponding to the keyword. For example, the user may utter a voice containing the keyword 'dog', such as "search target dog," and the electronic device (2000) can acquire the user's voice signal containing the keyword 'dog' through an input interface. In one embodiment, the electronic device can extract the word "dog" corresponding to the keyword from the user's voice signal converted into text, and identify the object corresponding to the keyword as 'dog' based on the extracted word "dog". However, it is not necessarily limited to the examples described above, and the electronic device (2000) can extract keywords from a user's voice signal through various deep learning-based models and identify objects corresponding to the extracted keywords.
[0180] In one embodiment, the electronic device (2000) may display a magnification ratio interface (1320) on an application screen. In one embodiment, the magnification ratio interface (1320) may mean a UI that determines or indicates the ratio (or crop ratio) occupied by the area of at least one object corresponding to a keyword in a thumbnail image displayed on an application screen. For example, the magnification ratio interface (1320) may include a slide UI, and the ratio (or crop ratio) occupied by the area of at least one object corresponding to a keyword in the thumbnail image may be determined or expressed based on the position of a button on the track of the slide UI.
[0181] In one embodiment, the electronic device (2000) may display a thumbnail GUI (1330) on an application screen. In one embodiment, the thumbnail GUI (1330) may include a plurality of areas where a plurality of thumbnail images are displayed. In one embodiment, the thumbnail GUI (1330) may include a grid view layout composed of a plurality of areas where a plurality of thumbnail images are displayed in a grid form. In one embodiment, the electronic device (2000) may identify a plurality of images containing at least one object corresponding to a keyword among a plurality of images stored in advance, and display thumbnail images of the identified plurality of images on the thumbnail GUI (1330). For example, the electronic device (2000) may identify a plurality of images containing a 'dog' based on the object detection results of a plurality of images stored in advance, and display a plurality of thumbnail images obtained by cropping the area where the 'dog' area occupies a preset ratio in the identified images on the thumbnail GUI (1330).
[0182] In one embodiment, the electronic device (2000) may obtain a first user input (1350a) for enlarging a thumbnail image. Here, the first user input (1350a) for enlarging a thumbnail image may refer to a user input for enlarging a plurality of thumbnail images displayed on a thumbnail GUI (1330). For example, the electronic device (2000) may obtain a pinch-out gesture input for the thumbnail GUI (1330) as the first user input (1350a). As another example, the electronic device (2000) may obtain a touch input for adjusting the position of a button on the magnification ratio interface (1320) as the first user input (1350a). In one embodiment, the electronic device (2000) may identify a ratio corresponding to the first user input (1350a) based on the obtained first user input (1350a).
[0183] In one embodiment, the electronic device (2000) can enlarge and display an area containing at least one object corresponding to a keyword of a plurality of thumbnail images based on a first user input (1350a). For example, the electronic device (2000) can identify an area in which the region of 'dog' in a plurality of images identified as containing 'dog' occupies a ratio corresponding to the first user input (1350a), and can display enlarged multiple thumbnail images obtained based on cropping the identified area on the thumbnail GUI (1330). In this case, the enlarged multiple thumbnail images can be displayed in the area where the previous multiple thumbnail images are displayed. In other words, the electronic device (2000) can enlarge and display areas containing at least one object corresponding to a keyword while maintaining the layout of the areas where thumbnail images are displayed on the thumbnail GUI (1330).
[0184] Referring to FIG. 13b, the electronic device (2000) can obtain a second user input (1350b) for enlarging a thumbnail image. Here, the second user input (1350b) for enlarging a thumbnail image may refer to a user input for enlarging a selected thumbnail image among a plurality of thumbnail images displayed on the thumbnail GUI (1330). For example, the electronic device (2000) can obtain a pinch-out gesture input for a surrounding area (1331) of a first thumbnail image (1341a) on the thumbnail GUI (1330) as the second user input (1350b). As another example, the electronic device (2000) may obtain a pinch-out gesture input for the thumbnail GUI (1330) as a second user input (1350b) while the first thumbnail image (1341a) is selected based on user input (e.g., touch input for the first thumbnail image (1341a)). In this case, the first thumbnail image (1341a) selected based on user input may be displayed with various visual elements indicating that a specific image has been selected, such as having its border highlighted or icons such as check boxes displayed around it.
[0185] In one embodiment, the electronic device (2000) may enlarge and display an area containing at least one object corresponding to a keyword of the first thumbnail image (1341a) based on the second user input (1350b). For example, the electronic device (2000) may identify an area in the original image of the first thumbnail image (1341a) (a pre-stored image identified as containing 'dog') in which the area of 'dog' occupies a ratio corresponding to the second user input (1350b), and display the enlarged first thumbnail image (1341b) obtained based on cropping the identified area on the thumbnail GUI (1330). In this case, among the plurality of thumbnail images displayed on the thumbnail GUI (1330), the unselected thumbnail images may not have an enlarged area containing at least one object corresponding to the keyword.
[0186] FIG. 14 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.
[0187] Referring to FIG. 14, the electronic device (2000) may include a memory (2100), a display (2200), a communication interface (2300), an input interface (2400), an output interface (2500), a camera (2600), and a processor (2700). The memory (2100), the display (2200), the communication interface (2300), the input interface (2400), the output interface (2500), the camera (2600), and the processor (2700) may each be electrically and / or physically connected to each other.
[0188] The components illustrated in FIG. 14 are merely according to one embodiment of the present disclosure, and the components included in the electronic device (2000) are not limited to those illustrated in FIG. 14. The electronic device (2000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 14 and may further include components not illustrated in FIG. 14.
[0189] The memory (2100) may store instructions or program code for performing functions or operations of the electronic device (2000). In one embodiment, one or more instructions, algorithms, data structures, program codes, and application programs stored in the memory (2100) may be implemented in a programming or scripting language such as, for example, C, C++, Java, assembler, etc.
[0190] In one embodiment, the memory (2100) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), Mask ROM, Flash ROM, etc.), a hard disk drive (HDD), or a solid-state drive (SSD). The memory (2100) may not exist separately but may be configured to be included in the processor (2700). The memory (2100) may be composed of volatile memory, non-volatile memory, or a combination of volatile memory and non-volatile memory. The memory (2100) may store a program or at least one instruction for performing operations according to the embodiments described below. The memory (2100) may also provide stored data to the processor (2700) upon the request of the processor (2700).
[0191] In one embodiment, the memory (2100) may include at least one of a cropping module (2110) and an outpainting module (2120). In one embodiment, the modules stored in the memory (2100) represent a unit that processes a function or operation of an electronic device (2000), which may be implemented as hardware included in the electronic device (2000) or software stored in the electronic device (2000), or as a combination of hardware and software. That is, the operation and function of the modules stored in the memory (2100) can be understood as the operation and function of the electronic device (2000).
[0192] In one embodiment, the cropping module (2110) can identify a crop area from an image and crop the identified crop area. In one embodiment, the cropping module (2110) can obtain a thumbnail image based on adjusting the size of the cropped area. Since the operation and function of the cropping module (2110) correspond to the operation and function of the electronic device (2000) described in FIG. 7, a redundant description will be omitted.
[0193] In one embodiment, the outpainting module (2120) can perform outpainting on an image (or a region of the image). In one embodiment, the outpainting module (2120) can identify an extended region of the image based on the image and a preset aspect ratio such that the aspect ratio of the outpainted image corresponds to the preset aspect ratio. In one embodiment, the outpainting module (2120) can obtain an outpainted image by outpainting the extended region of the image. Since the operation and function of the outpainting module (2120) correspond to the operation and function of the electronic device (2000) described in FIG. 10, a redundant description will be omitted. In one embodiment, the memory (2100) may include a training data set for training the outpainting module (2120).
[0194] In one embodiment, the memory (2100) may include at least one of a plurality of images, wide-angle images of the plurality of images, metadata of the plurality of images, and basic thumbnail images of the plurality of images. Here, the metadata may include an object detection result for at least one object included in the plurality of images and / or information about the area of the object corresponding to a keyword. However, it is not necessarily limited to the examples described above, and the memory (2100) may further include various data necessary to perform the operation and function of the electronic device (2000) disclosed in this specification.
[0195] In one embodiment, the display (2200) is a component for displaying images and / or videos. In one embodiment, the display (2200) may be composed of a physical device comprising at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display. In one embodiment, the display (2200) may display an application screen related to image search, a thumbnail image, and a thumbnail GUI in which the thumbnail image is displayed, based on a signal received from the processor (2700). However, it is not necessarily limited to the examples described above, and the display (2200) may display various graphic elements necessary for performing the operation and function of the electronic device (2000) disclosed in this specification based on a signal received from the processor (2700).
[0196] A communication interface (2300) is a component for an electronic device (2000) to perform communication with an external electronic device. In one embodiment, the communication interface (2300) can perform data communication between the electronic device (2000) and an external electronic device using at least one of a data communication method including wired LAN, wireless LAN, Wi-Fi, Bluetooth, Zigbee, WFD (Wi-Fi Direct), infrared communication (IrDA, infrared Data Association), BLE (Bluetooth Low Energy), NFC (Near Field Communication), Wibro (Wireless Broadband Internet), WiMAX (World Interoperability for Microwave Access), SWAP (Shared Wireless Access Protocol), WiGig (Wireless Gigabit Alliance), and RF communication.
[0197] In one embodiment, the communication interface (2300) may receive at least one of a plurality of images, a wide-angle shot of the plurality of images, and metadata of the plurality of images from an external electronic device, or transmit it to an external electronic device. In one embodiment, the communication interface may receive data input to or output to at least one of the cropping module (2110) and the outpainting module (2120) from an external electronic device or transmit it to an external electronic device. However, it is not necessarily limited to the examples described above, and the communication interface (2300) may receive various data necessary to perform the operation and function of the electronic device (2000) disclosed in this specification from an external electronic device or transmit it to an external electronic device.
[0198] The input interface (2400) is a component for receiving various user inputs. In one embodiment, the input interface (2400) may include a touch panel, a physical button, a microphone, etc. In one embodiment, information input through the input interface (2400) may be provided to the processor (2700). In one embodiment, the input interface (2400) may receive various user inputs including keywords. For example, the input interface (2400) may receive text input or voice input including keywords, or receive user input selecting one of a plurality of displayed keywords. In one embodiment, the input interface (2400) may receive user input for enlarging a thumbnail image or user input for reducing a thumbnail image. Here, user input for enlarging a thumbnail image or user input for reducing a thumbnail image may be understood as user input for adjusting the crop ratio in an identified image. However, it is not necessarily limited to the examples described above, and the input interface (2400) can acquire various data necessary to perform the operation and function of the electronic device (2000) disclosed in this specification.
[0199] The output interface (2500) is a component for the electronic device (2000) to provide various information to the user. In one embodiment, the electronic device (2000) may include a speaker, which is a component for outputting sound. In one embodiment, the output interface (2500) may output a voice representing a ratio corresponding to user input based on a signal received from the processor (2700), or output a voice representing a keyword identified based on user input. However, it is not necessarily limited to the examples described above, and the output interface (2500) may output various voices or information for performing the operation and function of the electronic device (2000) disclosed in this specification.
[0200] A camera (2600) is configured to acquire information about a real environment by having an electronic device (2000) capture the real environment. In one embodiment, the camera (2600) may include a lens module, an image sensor, and an image processing module. The camera (2600) may acquire a still image or video obtained by an image sensor (e.g., CMOS or CCD). In one embodiment, the image processing module may process the still image or video obtained through the image sensor to extract necessary information. In one embodiment, the information obtained through the camera (2600) may be provided to a processor (2700). In one embodiment, a plurality of images including at least one object corresponding to a keyword may be captured through the camera (2600) and stored in a memory (2100). In one embodiment, the camera (2600) may include a multi-lens camera. In one embodiment, the multi-lens camera may include a basic lens for capturing a basic image and a wide-angle lens for capturing a wide-angle image. In one embodiment, the camera (2600) may store in memory (2100) a basic image obtained by shooting through a basic lens and a wide-angle image obtained by shooting through a wide-angle lens by mapping or associating them with each other. However, it is not necessarily limited to the example described above, and the electronic device (2000) may obtain various images and videos to perform the operation and function of the electronic device (2000) disclosed in this specification through the camera (2600).
[0201] The processor (2700) can control the overall operations of the electronic device (2000). In one embodiment, the processor (2700) may include a plurality of processors. In one embodiment, at least one processor (2700) can perform the operation and function of the electronic device (2000) disclosed herein by executing one or more instructions of a program stored in memory (2100).
[0202] The processor (2700) may be composed of at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), an Application Processor, a Neural Processing Unit, or an AI-dedicated processor designed with a hardware structure specialized for processing AI models, but is not limited thereto.
[0203] When a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to one embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor and the third operation may be performed by a second processor. However, the embodiments of the present disclosure are not limited thereto.
[0204] One or more processors according to the present disclosure may be implemented as a single-core processor or as a multi-core processor. In the case where a method according to one embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single core or by a plurality of cores included in one or more processors.
[0205] In one embodiment, at least one processor (2700) can obtain a keyword for image search by executing one or more instructions. In one embodiment, at least one processor (2700) can identify an image containing at least one object corresponding to the keyword among a plurality of previously stored images by executing one or more instructions. In one embodiment, at least one processor (2700) can display a thumbnail image of the identified image through a display (2200) by executing one or more instructions. In one embodiment, at least one processor (2700) can obtain user input for enlarging the thumbnail image by executing one or more instructions. In one embodiment, at least one processor (2700) can enlarge and display an area containing at least one object corresponding to the keyword of the thumbnail image through a display (2200) based on user input by executing one or more instructions.
[0206] In one embodiment, at least one processor (2700) can identify a first image and a second image containing at least one object corresponding to a keyword among a plurality of images by executing one or more instructions. In one embodiment, at least one processor (2700) can display a first thumbnail image of the first image and a second thumbnail image of the second image in a first area and a second area on a GUI (Graphic User Interface) that displays the results of an image search through a display (2200) by executing one or more instructions. In one embodiment, at least one processor (2700) can display a first thumbnail image of the first image and a second thumbnail image of the second image in a first area and a second area on a GUI (Graphic User Interface) that displays the results of an image search through a display (2200) by executing one or more instructions.
[0207] In one embodiment, at least one object corresponding to the keyword may include a first object and a second object. In one embodiment, at least one processor (2700) may display a thumbnail image including both the first object and the second object through a display (2200) by executing one or more instructions.
[0208] In one embodiment, at least one processor (2700) can identify a first region containing a first object and a second region containing a second object in an identified image by executing one or more instructions. In one embodiment, at least one processor (2700) can display an animation including a first frame corresponding to the first region and a second frame corresponding to the second region through a display (2200) by executing one or more instructions.
[0209] In one embodiment, the animation may further include a third frame corresponding to a third area of the identified image. In one embodiment, at least one processor (2700) can identify a line connecting the center of the area of the first object and the center of the area of the second object in the thumbnail image by executing one or more instructions. In one embodiment, at least one processor (2700) can identify a third area with a center located on the line identified in the thumbnail image by executing one or more instructions.
[0210] In one embodiment, at least one processor (2700) can identify an area in an identified image where the area of at least one object corresponding to a keyword occupies a proportion corresponding to the user input by executing one or more instructions. In one embodiment, at least one processor (2700) can display an enlarged thumbnail image obtained by cropping the identified area through a display (2200) by executing one or more instructions.
[0211] In one embodiment, at least one processor (2700) can identify a region centered on the point closest to the center of the region of at least one object in an identified image by executing one or more instructions.
[0212] In one embodiment, at least one processor (2700) can crop an identified area in an identified image by executing one or more instructions. In one embodiment, at least one processor (2700) can obtain an enlarged thumbnail based on adjusting the size of the cropped area by executing one or more instructions.
[0213] In one embodiment, at least one processor (2700) can crop an identified area in an identified image by executing one or more instructions. In one embodiment, at least one processor (2700) can obtain an enlarged thumbnail image based on performing outpainting on the cropped area by executing one or more instructions.
[0214] In one embodiment, at least one processor (2700) can outpaint the remaining area excluding the cropped area by executing one or more instructions so that the center of the area of at least one object corresponding to the keyword is positioned at the center of the enlarged thumbnail image.
[0215] However, it is not necessarily limited to the examples described above, and at least one processor (2700) can perform various operations and functions of the electronic device (2000) disclosed in this specification by executing one or more instructions.
[0216] FIG. 15 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.
[0217] The server (3000) may include memory (3100), a communication interface (3200), and a processor (3300). Each may be electrically and / or physically connected to the others.
[0218] The components illustrated in FIG. 15 are merely according to one embodiment of the present disclosure, and the components included in the server (3000) are not limited to those illustrated in FIG. 15. The server (3000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 15 and may further include components not illustrated in FIG. 15.
[0219] Instructions or program code for performing functions or operations of the server (3000) may be stored in the memory (3100). In one embodiment, one or more instructions, algorithms, data structures, program codes, and application programs stored in the memory (2100) may be implemented in a programming or scripting language such as, for example, C, C++, Java, assembler, etc.
[0220] In one embodiment, the memory (3100) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), Mask ROM, Flash ROM, etc.), a hard disk drive (HDD), or a solid-state drive (SSD).
[0221] In one embodiment, the memory (3100) may include at least one of a cropping module (3110) and an outpainting module (3120). In one embodiment, since the data stored in the memory (3100) may correspond to the information stored in the memory (2100) of the electronic device (2000) of FIG. 14, a redundant description will be omitted.
[0222] A communication interface (3200) is a component for a server (3000) to perform communication with an external electronic device. In one embodiment, the communication interface (3200) can perform data communication between the server (3000) and an external electronic device using at least one of a data communication method including wired LAN, wireless LAN, Wi-Fi, Bluetooth, Zigbee, WFD (Wi-Fi Direct), infrared communication (IrDA, infrared Data Association), BLE (Bluetooth Low Energy), NFC (Near Field Communication), Wibro (Wireless Broadband Internet), WiMAX (World Interoperability for Microwave Access), SWAP (Shared Wireless Access Protocol), WiGig (Wireless Gigabit Alliance), and RF communication.
[0223] In one embodiment, since the operation and function of the communication interface (3200) can correspond to the operation and function of the communication interface (2300) of the electronic device (2000) of FIG. 14, a redundant description will be omitted.
[0224] The processor (3300) can control the overall operations of the server (3000). In one embodiment, the processor (3300) may include a plurality of processors. In one embodiment, at least one processor (3300) can perform the operation and function of the electronic device (2000) disclosed herein by executing one or more instructions of a program stored in memory (3100).
[0225] In one embodiment, since the operation and function performed by at least one processor (3300) by executing one or more instructions stored in memory (3100) can correspond to the operation and function of the processor (2700) of the electronic device (2000) of FIG. 14, redundant descriptions will be omitted.
[0226] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules executed by a computer. A computer-readable medium may be any available medium accessible by a computer and includes both volatile and non-volatile media, and both removable and non-removable media. Additionally, a computer-readable medium may include computer storage media and communication media. Computer storage media include both volatile and non-volatile, removable and non-removable media implemented by any method or technique for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include other data of modulated data signals, such as computer-readable instructions, data structures, or program modules.
[0227] Additionally, computer-readable storage media may be provided in the form of non-transitory storage media. Here, 'non-transitory storage media' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, 'non-transitory storage media' may include a buffer in which data is stored temporarily.
[0228] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0229] The foregoing description of the present disclosure is for illustrative purposes only, and those skilled in the art will understand that other specific forms can be easily modified without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single unit may be implemented in a distributed manner, and components described as distributed may likewise be implemented in a combined form.
[0230] The scope of the present disclosure is defined by the claims set forth below rather than by the detailed description above, and all modifications or variations derived from the meaning and scope of the claims and equivalent concepts thereof should be interpreted as being included within the scope of the present disclosure.
Claims
1. In a method for displaying image search results, Step of obtaining keywords for image search (S210); A step (S220) of identifying an image containing at least one object corresponding to the keyword among a plurality of previously stored images; A step of displaying a thumbnail image of the identified image (S230); A step of obtaining user input to enlarge the above thumbnail image (S240); A method comprising the step (S250) of expanding and displaying an area containing at least one object corresponding to the keyword of the thumbnail image based on the user input.
2. In Paragraph 1, The step of identifying an image containing at least one object corresponding to the above keyword is: The method includes the step of identifying a first image and a second image comprising at least one object corresponding to the keyword among the plurality of images; and The step of displaying the thumbnail image above is, The method includes the step of displaying a first thumbnail image of the first image and a second thumbnail image of the second image in a first area and a second area on a GUI (Graphic User Interface) representing the results of the image search; The step of expanding and displaying an area containing at least one object corresponding to the above keyword is: A method comprising the step of expanding and displaying areas including at least one object corresponding to the keyword of the first thumbnail image and the second thumbnail image in the first area and the second area, while maintaining the layout of the first area and the second area on the GUI.
3. In any one of paragraphs 1 to 2, At least one object corresponding to the above keyword is, Includes a first object and a second object, The step of displaying the thumbnail image above is, A method comprising the step of displaying the thumbnail image including both the first object and the second object.
4. In Paragraph 3, The step of expanding and displaying an area containing at least one object corresponding to the above keyword is: The method includes the step of identifying a first region containing the first object and a second region containing the second object in the identified image; A method comprising the step of displaying an animation including a first frame corresponding to the first area and a second frame corresponding to the second area.
5. In Paragraph 4, The above animation is, It further includes a third frame corresponding to a third region of the identified image, and The step of expanding and displaying an area containing at least one object corresponding to the above keyword is: A step of identifying a line connecting the center of the area of the first object and the center of the area of the second object in the thumbnail image above; A method further comprising the step of identifying the third region whose center is located on the identified line in the thumbnail image.
6. In any one of paragraphs 1 through 5, The step of expanding and displaying an area containing at least one object corresponding to the above keyword is: A step of identifying an area in the identified image in which the area of at least one object corresponding to the keyword occupies a proportion corresponding to the user input; and A method comprising the step of displaying an enlarged thumbnail image obtained based on cropping the identified area.
7. In Paragraph 6, The step of identifying the above region is, A method for identifying a region centered on the point closest to the center of the region of at least one object in the identified image.
8. In any one of paragraphs 6 through 7, The step of displaying the enlarged thumbnail image above is, A step of cropping the identified area in the identified image; and A method comprising the step of obtaining the enlarged thumbnail based on adjusting the size of the cropped area.
9. In any one of paragraphs 6 through 7, The step of displaying the enlarged thumbnail image above is, A step of cropping the identified area in the identified image; and A method comprising the step of obtaining the enlarged thumbnail image based on performing outpainting on the cropped area.
10. In Paragraph 9, The step of acquiring the enlarged thumbnail image above is, A method comprising the step of outpainting the remaining area excluding the cropped area so that the center of the area of at least one object corresponding to the above keyword is positioned at the center of the enlarged thumbnail image.
11. In an electronic device (2000) that displays image search results, Memory (2100) for storing one or more instructions;, Display (2200); and It includes at least one processor (2700) that executes one or more instructions stored in the memory; and By the above at least one processor executing the above one or more instructions, the electronic device (2000) is, Obtain keywords for image search, Identify an image containing at least one object corresponding to the keyword among a plurality of previously stored images, and A thumbnail image of the identified image is displayed through the display (2200), and Obtaining user input to enlarge the above thumbnail image, and An electronic device that enlarges and displays an area containing at least one object corresponding to the keyword of the thumbnail image based on user input through the display (2200).
12. In Paragraph 11, By the above at least one processor executing the above one or more instructions, the electronic device, Identifying a first image and a second image containing at least one object corresponding to the keyword among the plurality of images above, and A first thumbnail image of the first image and a second thumbnail image of the second image are displayed in a first area and a second area on a GUI (Graphic User Interface) that displays the results of the image search through the display, and An electronic device that maintains the layout of the first area and the second area on the GUI through the display, and expands and displays areas including at least one object corresponding to the keyword of the first thumbnail image and the second thumbnail image in the first area and the second area.
13. In any one of paragraphs 11 to 12, At least one object corresponding to the above keyword is, Includes a first object and a second object, By the above at least one processor executing the above one or more instructions, the electronic device, An electronic device that displays the thumbnail image including both the first object and the second object through the display.
14. In Paragraph 13, By the above at least one processor executing the above one or more instructions, the electronic device, Identifying a first region containing the first object and a second region containing the second object in the identified image above, and An electronic device that displays an animation including a first frame corresponding to the first area and a second frame corresponding to the second area through the display.
15. A computer-readable recording medium having a program recorded thereon for performing the method of any one of paragraphs 1 through 10 on a computer.