Three-dimensional space object search method, and server for performing same

The method employs a server-based system that uses feature extraction and matching to efficiently search for objects within a 3D space, addressing the challenges of indoor object localization and content retrieval.

WO2025127332A1PCT designated stage expired Publication Date: 2025-06-193I INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/013879
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-09-12
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently searching for objects within a 3D space, particularly in indoor environments, due to the complexity of spatial representation and the need for precise location information.

Method used

A method utilizing a server to receive a query image, generate global and local features, and match these features against a database of 360-degree images associated with a 3D space, allowing for the identification of candidate images and the determination of object locations within the space.

Benefits of technology

This approach enables accurate and efficient object search within a 3D space by leveraging feature matching and similarity calculations, providing users with the ability to locate objects and access associated content in an indoor environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024013879_19062025_PF_FP_ABST
    Figure KR2024013879_19062025_PF_FP_ABST
Patent Text Reader

Abstract

This object search method comprises: receiving a query image including a target object from a user terminal connected to a server; generating a target global feature for the query image; using the target global feature to determine one or more candidate 360-degree images from among a plurality of 360-degree images associated with a target space; converting the one or more candidate 360-degree images into one or more planar image sets; generating a target local feature for the query image; and using the target local feature to determine a target planar image that includes at least some of reference object images corresponding to target object images from among the one or more planar image sets.
Need to check novelty before this filing date? Find Prior Art

Description

Method for searching objects in 3D space and server for performing the same

[0001] The present disclosure relates to an image-based object retrieval method, and more particularly, to a technique for retrieving objects within a query image based on images of an indoor space. Furthermore, the present disclosure relates to a technique for providing stored content by associating the retrieved objects with the images.

[0002] In order to record a three-dimensional space, multiple 360-degree images can be created by taking pictures in a 360-degree direction from each of multiple points in the space, and a virtual three-dimensional space can be created for the three-dimensional space by associating the multiple points with each of the multiple 360-degree images.

[0003] A virtual three-dimensional space may include location information for multiple points where multiple 360-degree images were captured, and the location information for the multiple points may be mapped onto a plan view of the three-dimensional space and provided to the user. The virtual three-dimensional space may be provided to the user as virtual reality (VR) or augmented reality (AR) content.

[0004] The background technology described above is something that the inventor possessed or acquired in the process of deriving the disclosure of the present application, and cannot necessarily be said to be a publicly known technology disclosed to the general public prior to the present application.

[0005] One embodiment may provide a method for searching for objects performed by a server providing a virtual three-dimensional space.

[0006] One embodiment may provide a content providing method performed by a server providing a virtual three-dimensional space.

[0007] According to one embodiment, a method for searching for an object performed by a server includes the steps of: receiving a query image from a user terminal connected to the server, wherein the query image includes a target object image; generating a target global feature for the query image; determining, using the target global feature, one or more candidate 360-degree images from among a plurality of 360-degree images associated with a target space; converting the one or more candidate 360-degree images into one or more planar image sets; generating a target local feature for the query image; and determining, using the target local feature, a target planar image from among the one or more planar image sets, the target planar image including at least a portion of a reference object image corresponding to the target object image.

[0008] The step of determining one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space using the target global feature may include the step of calculating a first similarity between the target global feature for the query image and a global feature for each of the plurality of 360-degree images, and the step of determining the one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity.

[0009] The step of determining one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity may include the step of determining a 360-degree image among the plurality of 360-degree images whose first similarity is greater than or equal to a first threshold as a candidate 360-degree image.

[0010] The step of converting the one or more candidate 360-degree images into one or more sets of planar images may include the steps of dividing a first candidate 360-degree image among the one or more candidate 360-degree images into a plurality of regions, and converting images corresponding to the plurality of regions into a first set of planar images for the first candidate 360-degree image.

[0011] The plurality of planar images within the first planar image sets may include overlapping areas with each other.

[0012] The step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets using the local feature may include the step of calculating a second similarity between the target local feature for the query image and a local feature for each of the planar images of the one or more planar image sets, and the step of determining the target planar image including at least a portion of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity.

[0013] The step of determining the target planar image including at least a portion of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity may include the step of determining a planar image having the second similarity greater than or equal to a second threshold among the planar images of the one or more planar image sets as the target planar image.

[0014] The step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets using the target local feature may include the step of determining the target planar image using the target local feature and posture information of the user terminal from which the query image was captured.

[0015] The above object search method may further include a step of providing information on an indoor location corresponding to the target plane image to the user terminal in association with the target space.

[0016] According to one embodiment, a computer-readable recording medium can store a program for performing the object search method.

[0017] According to one embodiment, a server includes a communication unit, a processor, and a memory, and the memory may store instructions that, when executed by the processor, cause the server to receive a query image from a user terminal connected to the server through the communication unit, the query image including a target object image, generate a target global feature for the query image, determine one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space using the target global feature, transform the one or more candidate 360-degree images into one or more planar image sets, generate a target local feature for the query image, and determine, using the target local feature, a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets.

[0018] According to one embodiment, a content providing method performed by a server includes the steps of: receiving a query image from a user terminal connected to the server, wherein the query image includes a target object image; generating a target global feature for the query image; determining, using the target global feature, one or more candidate 360-degree images from among a plurality of 360-degree images associated with a target space; converting the one or more candidate 360-degree images into one or more planar image sets; generating a target local feature for the query image; determining, using the target local feature, a target planar image including at least a portion of a reference object image corresponding to the target object image from among the one or more planar image sets; determining a target area of ​​the target planar image within a target 360-degree image corresponding to the target planar image from among the plurality of 360-degree images; determining whether target content exists within the target area; and, if the target content exists within the target area, providing the target content to the user terminal.

[0019] The above target content may be overlapped with the query image and output to the user terminal.

[0020] The above content providing method may further include a step of receiving information about a target location of the target 360-degree image to be associated with the target content from an electronic device connected to the server, a step of receiving the target content from the electronic device, and a step of storing the target content in association with the target location.

[0021] The target content may include at least one of a photo, video, audio, text, or a uniform resource locator (URL).

[0022] The step of determining one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space using the target global feature may include the step of calculating a first similarity between the target global feature for the query image and a global feature for each of the plurality of 360-degree images, and the step of determining the one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity.

[0023] The step of determining one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity may include the step of determining a 360-degree image among the plurality of 360-degree images whose first similarity is greater than or equal to a first threshold as a candidate 360-degree image.

[0024] The step of converting the one or more candidate 360-degree images into one or more sets of planar images may include the steps of dividing a first candidate 360-degree image among the one or more candidate 360-degree images into a plurality of regions, and converting images corresponding to the plurality of regions into a first set of planar images for the first candidate 360-degree image.

[0025] The plurality of planar images within the first planar image sets may include overlapping areas with each other.

[0026] The step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets using the local feature may include the step of calculating a second similarity between the target local feature for the query image and a local feature for each of the planar images of the one or more planar image sets, and the step of determining the target planar image including at least a portion of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity.

[0027] The step of determining the target planar image including at least a portion of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity may include the step of determining the planar image having the highest second similarity among the planar images of the one or more planar image sets as the target planar image.

[0028] The step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets using the target local feature may include the step of determining the target planar image using the target local feature and posture information of the user terminal from which the query image was captured.

[0029] According to one embodiment, a computer-readable recording medium can store a program for performing the content providing method.

[0030] According to one embodiment, a server includes a communication unit, a processor, and a memory, wherein the memory, when executed by the processor, causes the server to receive a query image from a user terminal connected to the server through the communication unit, the query image including a target object image, generate a target global feature for the query image, determine one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space using the target global feature, transform the one or more candidate 360-degree images into one or more planar image sets, generate a target local feature for the query image, determine a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets using the target local feature, determine a target area of ​​the target planar image within a target 360-degree image corresponding to the target planar image among the plurality of 360-degree images, determine whether target content exists within the target area, and, if the target content exists within the target area, transmit the target content to the user terminal. You can save the instructions that are set to be provided.

[0031] FIG. 1 is a schematic diagram of a virtual three-dimensional space providing system according to one embodiment.

[0032] Figure 2 is a flowchart illustrating an object search method according to one embodiment.

[0033] FIG. 3 is a flowchart illustrating a method for determining one or more candidate 360-degree images using global features according to one embodiment.

[0034] Figure 4 is a flowchart illustrating a method for obtaining a set of planar images according to one embodiment.

[0035] FIG. 5 is a flowchart illustrating an object search method using local features according to one embodiment.

[0036] FIG. 6 is a drawing for explaining a target plane image according to one embodiment.

[0037] FIG. 7 is a block diagram of a server performing object search according to one embodiment.

[0038] Figure 8 is a flowchart illustrating a content provision method according to one embodiment.

[0039] FIG. 9 is a flowchart illustrating a method for determining one or more candidate 360-degree images using global features according to one embodiment.

[0040] FIG. 10 is a flowchart illustrating a method for obtaining a set of planar images according to one embodiment.

[0041] Fig. 11 is a flowchart illustrating an object search method using local features according to one embodiment.

[0042] Figure 12 is a flowchart illustrating a method of storing content according to one embodiment.

[0043] Figure 13 is a drawing for explaining a target area according to one embodiment.

[0044] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Therefore, the actual implementation is not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or alternatives within the technical concepts described in the embodiments.

[0045] Although terms such as "first" or "second" may be used to describe various components, these terms should be interpreted solely to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.

[0046] When it is said that a component is "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.

[0047] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprises" or "has" should be understood to indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but not to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0048] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0049] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.

[0050] FIG. 1 is a schematic diagram of a virtual three-dimensional space providing system according to one embodiment.

[0051] According to one embodiment, a virtual three-dimensional (dimension: D) space provision system (1) (hereinafter, system (1)) may include a camera (10), a server (20), and a user terminal (30).

[0052] According to one embodiment, the camera (10) may be a 360-degree camera. The camera (10) may capture all surrounding directions, including 360 degrees up and down and left and right, using one or more wide-angle lenses. A user may use the camera (10) to create a 360-degree image (e.g., a panoramic image) from a specific point in 3D space. The user may create a plurality of 360-degree images through the camera (10) while moving to a plurality of points in 3D space. The created plurality of 360-degree images may be transmitted and stored in the server (20). For example, a first 360-degree image may be associated with information about a first point where the first 360-degree image was captured. The location information of the plurality of points may include information about the first point.

[0053] According to one embodiment, the camera (10) can transmit and receive data to and from the server (20) via a user terminal (30) or any electronic device (not shown). Alternatively, the camera (10) can transmit and receive data directly to and from the server (20) via a network.

[0054] According to one embodiment, the server (20) may generate a 3D tour for a virtual 3D space corresponding to the 3D space based on a plurality of 360-degree images captured at a plurality of points in the 3D space and positional information of the plurality of points. Hereinafter, the virtual 3D space may be referred to as a target space. The server (20) may provide the target space to the user through a 3D tour in the form of virtual reality or augmented reality. For example, the server (20) may map positional information of each point in the target space onto a floor plan and provide it to the user terminal (30). Movement information of the camera (10) that appears during the process of capturing a plurality of 360-degree images may be used to determine the user's movement path for the 3D tour. For example, the movement information may be determined by the user terminal connected to the camera (10). For example, the user terminal may determine the movement information using simultaneous localization and map-building (SLAM). For example, the user's movement path may be displayed on the floor plan and provided to the user terminal (30).

[0055] According to one embodiment, when a user of a user terminal (30) selects a first point among a plurality of points through a 3D tour, the server (20) that receives the first point from the user terminal (30) may provide a first 360-degree image corresponding to the first point of the target space to the user terminal (30). The server (20) may provide the first 360-degree image to the user terminal (30). When the server (300) receives a target direction of the first point from the user terminal (30), the server (300) may determine a target area corresponding to the target direction among the first 360-degree images, generate a target plane image for the determined target area, and provide the generated target plane image to the user terminal (30).

[0056] According to one embodiment, the server (20) can provide an augmented reality experience using content associated with a target space. The content can include at least one of a photo, video, audio, text, or a uniform resource locator (URL). The content can be stored for a specific location in the target space (e.g., around an object, on an indoor wall, etc.). Information related to the content can be input to the server (20) through an electronic device of a 3D tour manager or other authorized user (hereinafter, “user”). When a user selects a location (or point) in a 360-degree image to store content from among a plurality of 360-degree images through the electronic device, the server (20) can receive information about the location to be associated with the content from the electronic device. The server (20) can receive the content from the electronic device and store it in association with the location.

[0057] According to one embodiment, the server (20) may receive a query image from a user terminal (30). For example, the server (20) may receive a query image such as a captured photo or a video shot in real time from the user terminal (30).

[0058] According to one embodiment, the server (20) can search for an object corresponding to a target object image included in a query image in a target space. The target object image can represent an image of a specific object included in the query image.

[0059] According to one embodiment, the server (20) may search for an object identical to or similar to a target object image and provide an image of the searched object to the user terminal (30). The server (20) may provide a portion of a 360-degree image including the image of the searched object to the user terminal (30) in the form of a 2D image (or, flat image).

[0060] According to one embodiment, the server (20) can search for an object identical to or similar to a target object image and determine location information of the searched object.

[0061] According to one embodiment, the server (20) can search for an object and determine location information of the object by comparing features (e.g., global features or local features) of a plurality of 360-degree images associated with a target space for an indoor space with features (e.g., global features or local features) of a query image. That is, the server (20) can determine the indoor location of the object or the user terminal (200) by using location information of a 360-degree image including the searched object, without using location information that can be received via GPS. The server (20) can provide location information of the searched object to the user terminal (30).

[0062] In one embodiment, retrieving an object in a query image may mean determining the (indoor) location of the object or user terminal (30), and vice versa.

[0063] According to one embodiment, the server (20) may provide content associated with a location determined using a query image to the user terminal (30). The server (20) may determine whether target content stored in association with the determined location exists among content pre-stored in association with any location in the target space. If target content associated with the determined location exists, the server (20) may provide the target content to the user terminal (30), and an augmented reality screen displaying the target content on the query image, such as a captured photo or a video shot in real time, may be output to the user terminal (30).

[0064] According to one example, a method for presetting content describing how to use a manufacturing machine located within a target space within the target space may be described. When a user selects a target location of a target 360-degree image in which a manufacturing machine is captured from among multiple 360-degree images of a 3D tour through an electronic device, a server (20) may receive information about the target location from the electronic device. For example, the information about the target location may be coordinates of the target location within the target 360-degree image. The server (20) may receive target content, such as a video or photo, regarding how to use the manufacturing machine from the electronic device, and store the received target content in association with the target location.

[0065] According to an example, a method of providing a user of a user terminal (30) with content describing a method of using a manufacturing machine located within a target space may be described. A server (20) may receive a query image including at least a portion of an image of a manufacturing machine from the user terminal (30), and the server (20) may search for an object image corresponding to the image of the manufacturing machine included in the query image (or any target object image included in the query image) within the target space. The server (20) may determine whether there is target content stored in association with the location of the searched object image. For example, if there is target content associated with the location of the searched object image (i.e., the manufacturing machine), the server (20) may provide the target content to the user terminal (30). That is, in response to a user inputting a query image including a manufacturing machine into the server (20) through the user terminal (30), the user terminal (30) may output target content describing a method of using the manufacturing machine by overlapping the query image.

[0066] Figure 2 is a flowchart illustrating an object search method according to one embodiment.

[0067] According to one embodiment, the operations 210 to 260 below may be performed by a server (e.g., server (20) of FIG. 1).

[0068] In operation 210, the server may receive a query image from a user terminal connected to the server (e.g., the user terminal (30) of FIG. 1). The query image may include an image captured in real time using a camera of the user terminal or an image stored in the user terminal. The query image may include a target object image. The target object image may be an image including at least a portion of an object to be searched in the target space.

[0069] At operation 220, the server can generate a target global feature for the query image.

[0070] According to one embodiment, the server can generate (or extract) a global feature in vector format for a query image using a pre-trained deep learning model. The deep learning model includes a neural network model and can be trained to generate the same global feature for images taken of the same location, for example. The deep learning model can output one global feature for one image. The method for generating the global feature can include a handcrafted feature extraction method or any of the conventional feature extraction methods previously disclosed.

[0071] In operation 230, the server may use the target global feature to determine one or more candidate 360-degree images among a plurality of 360-degree images associated with the target space.

[0072] According to one embodiment, the server may calculate a first similarity between a target global feature and a global feature for each of a plurality of 360-degree images, and determine one or more candidate 360-degree images based on the first similarity. A method for determining candidate 360-degree images is described in detail with reference to FIG. 3.

[0073] At operation 240, the server can convert one or more candidate 360-degree images into one or more sets of planar images.

[0074] According to one embodiment, the server may divide a first candidate 360-degree image among one or more candidate 360-degree images into a plurality of regions, and convert images corresponding to the plurality of regions into a first planar image set for the first candidate 360-degree image. A method for obtaining the planar image set is described in detail with reference to FIG. 4.

[0075] At operation 250, the server can generate target local features for the query image.

[0076] According to one embodiment, a server may generate (or extract) local features in vector format for an image using a pre-trained deep learning model. The local features may include vector values ​​and location information (e.g., coordinates and scale) for patches derived around each key point extracted from the image. The deep learning model may be trained to generate the same local features for patches of the same object. The deep learning model may output multiple local features for a single image. The method for generating the local features may include a hand-crafted feature extraction method or any of the conventional feature extraction methods previously disclosed.

[0077] According to one embodiment, the server can generate target local features in vector format for a query image using a pre-trained deep learning model.

[0078] In operation 260, the server may use the target local feature to determine a target planar image including at least a portion of a reference object image corresponding to the target object image among one or more sets of planar images. The reference object image may represent an image of any corresponding object that can be classified into the same class as the object corresponding to the target object image.

[0079] According to one embodiment, the server may calculate a second similarity between a target local feature and each of the local features of planar images of one or more planar image sets, and determine a target planar image among the planar images of one or more planar image sets based on the second similarity. The server may determine a target planar image in which an object corresponding to a target object image included in a query image is captured. A method for searching for an object using a local feature is described in detail with reference to FIG. 5.

[0080] According to one embodiment, the server may determine a target planar image using target local features and attitude information of the user terminal from which the query image was captured. The attitude information of the user terminal may include information related to an angle or acceleration measured by a sensor (e.g., a gyro sensor) included in the user terminal. The server may refer to attitude information of the user terminal from which the query image was captured to determine a target planar image using the target local features. For example, if the query image was captured while the user terminal was tilted toward the ground (or in the direction of gravitational acceleration), the server may assign a greater weight to local features of planar images closer to the ground among the first planar image set for the first candidate 360-degree image when calculating the second similarity using the target local features.

[0081] In one embodiment, the server may provide a target plane image to the user terminal. For example, if there are multiple target plane images, the server may provide the user terminal with a split screen composed of each target plane image.

[0082] According to one embodiment, the number of objects corresponding to a target object image within a target space may be counted based on a target plane image. For example, the server may count the number of objects corresponding to a target object image within a target space based on a target plane image. For example, a user may count the number of objects corresponding to a target object image within a target space based on a target plane image provided to a user terminal.

[0083] According to one embodiment, the server may provide information about an indoor location corresponding to a target plane image to the user terminal in association with the target space. For example, the server may display the location of a 360-degree image corresponding to the target plane image among a plurality of 360-degree images on a floor plan of the target space and provide the display to the user terminal. For example, the server may display the direction corresponding to the target plane image from the location of the 360-degree image of the target space on the floor plan of the target space and provide the display to the user terminal based on the relative location of the target plane image within the 360-degree images.

[0084] FIG. 3 is a flowchart illustrating a method for determining one or more candidate 360-degree images using global features according to one embodiment.

[0085] According to one embodiment, the operations 310 and 320 below may be performed by a server (e.g., server (20) of FIG. 1).

[0086] According to one embodiment, operation 230 of FIG. 2 may include operations 310 and 320.

[0087] At operation 310, the server may calculate a first similarity between a target global feature for the query image and a global feature for each of the plurality of 360-degree images. As described with reference to operation 220 of FIG. 2 , the server may generate a target global feature for the query image and a global feature for each of the plurality of 360-degree images associated with the target space. The server may generate one global feature for each of the plurality of 360-degree images.

[0088] According to one embodiment, the server may calculate a first similarity by performing an operation (e.g., cosine distance) on a target global feature in vector format and a global feature for each of a plurality of 360-degree images.

[0089] In operation 320, the server can determine one or more candidate 360-degree images from among the plurality of 360-degree images based on the first similarity.

[0090] In one embodiment, the server may determine a 360-degree image among multiple 360-degree images with a first similarity greater than or equal to a first threshold as a candidate 360-degree image. In other words, the server may determine one or more candidate 360-degree images that are similar to the query image and thus highly relevant. The first threshold may be adjusted based on the accuracy of object search.

[0091] Figure 4 is a flowchart illustrating a method for obtaining a set of planar images according to one embodiment.

[0092] According to one embodiment, the operations 410 and 420 below may be performed by a server (e.g., server (20) of FIG. 1).

[0093] According to one embodiment, operation 240 of FIG. 2 may include operations 410 and 420.

[0094] In operation 410, the server can divide a first candidate 360-degree image from among one or more candidate 360-degree images into a plurality of regions.

[0095] In one embodiment, the server may project one or more candidate 360-degree images onto a cube in an orthogonal coordinate system. For example, the server may divide the first candidate 360-degree image into six zones by projecting it onto the cube.

[0096] In one embodiment, the server may project one or more candidate 360-degree images into spherical coordinates. For example, the server may project a first candidate 360-degree image into spherical coordinates and divide it into 18 regions.

[0097] In operation 420, the server may convert images corresponding to a plurality of regions into a first planar image set for a first candidate 360-degree image. The plurality of planar images within the first planar image sets may include overlapping areas. For example, the server may convert images corresponding to six regions of the first candidate 360-degree image projected onto a cube into a first planar image set including six 2D images.

[0098] FIG. 5 is a flowchart illustrating an object search method using local features according to one embodiment.

[0099] According to one embodiment, the operations 510 and 520 below may be performed by a server (e.g., server (20) of FIG. 1).

[0100] According to one embodiment, operation 260 of FIG. 2 may include operations 510 and 520.

[0101] In operation 510, the server may determine a second similarity between a target local feature for the query image and a local feature for each of the planar images of one or more planar image sets. As described with reference to operation 250 of FIG. 2, the server may generate a target local feature for the query image and a local feature for each of the planar images of the one or more planar image sets. The server may generate multiple local features for each of the planar images.

[0102] According to one embodiment, the server can compute a second similarity by performing an operation (e.g., cosine distance) on a target local feature in vector format and a local feature for each of the planar images of one or more planar image sets.

[0103] In operation 520, the server may determine a target planar image that includes at least a portion of a reference object image corresponding to the target object image among planar images of one or more planar image sets based on the second similarity. For example, if the target object image is an image of a ladder, the reference object image may represent an image of any ladder that can be classified into the same class as the target object image. The server may determine a planar image among the planar images of the one or more planar image sets as the target planar image if the planar image includes a partially cropped portion or all of the reference object image (i.e., the image of the ladder).

[0104] According to one embodiment, the server may determine a planar image among one or more planar image sets having a second similarity greater than or equal to a second threshold as a target planar image. The second threshold may be adjusted based on the accuracy of object search.

[0105] In one embodiment, any planar image that includes at least a portion of a reference object image may be determined as a target planar image. Since multiple 360-degree images associated with the target space may be generated by photographing overlapping points, the same object may be included in one or more 360-degree images. That is, when multiple planar images are determined as target planar images, the target planar image may include some target planar images that photograph the same object from different positions and angles.

[0106] In one embodiment, a single target plane image may include one or more reference object images that can be classified into the same class. For example, if the target object image is an image of a ladder, a single target plane image may include at least some of images of different ladders.

[0107] FIG. 6 is a drawing for explaining a target plane image according to one embodiment.

[0108] As described above, a server (e.g., server (20) of FIG. 1) may receive a query image from a user terminal (e.g., user terminal (30) of FIG. 1) connected to the server, generate a target global feature for the query image, and use the target global feature to determine one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space.

[0109] According to one embodiment, the server may divide a first candidate 360-degree image (61) from among one or more candidate 360-degree images into a plurality of zones, and convert images corresponding to the plurality of zones into a first planar image set (63) for the first candidate 360-degree image (61). For example, the server may divide the first 360-degree image (61) into six zones by projecting it onto a cube in an orthogonal coordinate system, and convert images corresponding to the plurality of zones into the first planar image set (63).

[0110] As described above, the server can generate a target local feature for the query image and a local feature for each of the planar images of the first planar image set (63), and calculate a similarity (e.g., a second similarity) between the target local feature and the local feature for each of the planar images of the first planar image set (63). The server can determine a target planar image (65) that includes at least a portion of a reference object image (610) corresponding to a target object image included in the query image among the planar images of the first planar image set (63). Information about an indoor location of a target space corresponding to the target planar image (65) can be determined based on a relative position of the target planar image (65) within the first candidate 360-degree image (61).

[0111] FIG. 7 is a block diagram of a server performing object search according to one embodiment.

[0112] A server (20) according to one embodiment may represent the server (20) of FIG. 1. The server (20) may include a communication unit (710), a processor (720), and a memory (730).

[0113] The communication unit (710) is connected to the processor (720), the memory (730), and the data collection module (240) to transmit and receive data. The communication unit (710) may be connected to another external device (e.g., the camera (10) or the user terminal (30) of FIG. 1) to transmit and receive data. Hereinafter, the expression "transmitting and receiving "A" may refer to transmitting and receiving "information or data representing A."

[0114] The communication unit (710) may be implemented as a circuitry within the server (20). For example, the communication unit (710) may include an internal bus and an external bus. As another example, the communication unit (710) may be an element that connects the server (20) to an external device. The communication unit (710) may be an interface. The communication unit (710) may receive data from an external device and transmit the data to the processor (720) and the memory (730).

[0115] The processor (720) processes data received by the communication unit (710) and data stored in the memory (730). A “processor” may be a hardware-implemented data processing device having a circuit with a physical structure for executing desired operations. For example, the desired operations may include code or instructions included in a program. For example, a hardware-implemented data processing device may include a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array).

[0116] The processor (720) executes computer-readable code (e.g., software) stored in memory (e.g., memory (730)) and instructions generated by the processor (720).

[0117] The memory (730) stores data received by the communication unit (710) and data processed by the processor (720). For example, the memory (730) may store a program (or application, software). The stored program may be a set of syntaxes that are coded based on collected data and can be executed by the processor (720).

[0118] According to one aspect, the memory (730) may include one or more volatile memory, non-volatile memory, and random access memory (RAM), flash memory, a hard disk drive, and an optical disk drive.

[0119] The memory (730) stores a set of instructions (e.g., software) that operate the server (20). The set of instructions that operate the server (20) is executed by the processor (720).

[0120] According to one embodiment, the memory (730) may store a plurality of 360-degree images of a 3D space and information about points at which each of the plurality of 360-degree images was captured. The memory (730) may store a 3D tour of a virtual 3D space corresponding to the 3D space based on the plurality of 360-degree images captured at a plurality of points in the 3D space and positional information of the plurality of points. For example, the memory (730) may store a 3D tour of a target space corresponding to the 3D space based on the plurality of 360-degree images captured at a plurality of points in the 3D space and positional information of the plurality of points. The memory (730) may store target content associated with the target space in association with the target position of the target 360-degree image.

[0121] Figure 8 is a flowchart illustrating a content provision method according to one embodiment.

[0122] According to one embodiment, the operations 810 to 890 below may be performed by a server (e.g., server (20) of FIG. 1).

[0123] In operation 810, the server may receive a query image from a user terminal connected to the server (e.g., the user terminal (30) of FIG. 1). The query image may include an image captured in real time using a camera of the user terminal or an image stored in the user terminal. The query image may include a target object image. The target object image may be an image including at least a portion of an object to be searched in the target space.

[0124] At operation 820, the server can generate a target global feature for the query image.

[0125] According to one embodiment, the server can generate (or extract) a global feature in vector format for a query image using a pre-trained deep learning model. The deep learning model includes a neural network model and can be trained to generate the same global feature for images taken of the same location, for example. The deep learning model can output one global feature for one image. The method for generating the global feature can include a handcrafted feature extraction method or any of the conventional feature extraction methods previously disclosed.

[0126] At operation 830, the server may use the target global feature to determine one or more candidate 360-degree images from among a plurality of 360-degree images associated with the target space.

[0127] According to one embodiment, the server may calculate a first similarity between a target global feature and a global feature for each of a plurality of 360-degree images, and determine one or more candidate 360-degree images based on the first similarity. A method for determining candidate 360-degree images is described in detail with reference to FIG. 9.

[0128] At operation 840, the server can convert one or more candidate 360-degree images into one or more sets of planar images.

[0129] According to one embodiment, the server may divide a first candidate 360-degree image among one or more candidate 360-degree images into a plurality of regions, and convert images corresponding to the plurality of regions into a first planar image set for the first candidate 360-degree image. A method for obtaining the planar image set is described in detail with reference to FIG. 10.

[0130] At operation 850, the server can generate target local features for the query image.

[0131] According to one embodiment, a server may generate (or extract) local features in vector format for an image using a pre-trained deep learning model. The local features may include vector values ​​and location-related information (e.g., coordinates and scale) for patches derived around each keypoint extracted from the image. The deep learning model may be trained to generate the same local features for patches of the same object. The deep learning model may output multiple local features for a single image. The method for generating the local features may include a hand-crafted feature extraction method or any of the conventional feature extraction methods previously disclosed.

[0132] According to one embodiment, the server can generate target local features in vector format for a query image using a pre-trained deep learning model.

[0133] In operation 860, the server may use the target local feature to determine a target planar image including at least a portion of a reference object image corresponding to the target object image among one or more sets of planar images. The reference object image may represent an image of any corresponding object that can be classified into the same class as the object corresponding to the target object image. The server may determine a single planar image with the highest similarity to the searched object as the target planar image.

[0134] According to one embodiment, the server may calculate a second similarity between a target local feature and each of the local features of planar images of one or more planar image sets, and determine a target planar image among the planar images of one or more planar image sets based on the second similarity. The server may determine a target planar image in which an object corresponding to a target object image included in a query image is captured. A method for searching for an object using a local feature is described in detail with reference to FIG. 11.

[0135] According to one embodiment, the server may provide information about an indoor location corresponding to a target plane image to the user terminal in association with the target space. For example, the server may display the location of a 360-degree image corresponding to the target plane image among a plurality of 360-degree images on a floor plan of the target space and provide the display to the user terminal. For example, based on the relative location of the target plane image within the 360-degree images, the server may display the direction corresponding to the target plane image from the location of the 360-degree image of the target space on the floor plan of the target space and provide the display to the user terminal.

[0136] In operation 870, the server can determine a target area of ​​a target plane image within a target 360-degree image corresponding to the target plane image among a plurality of 360-degree images.

[0137] In one embodiment, as described with reference to FIG. 1, when a user selects a target direction of a target 360-degree image through a 3D tour, the server can determine a target area corresponding to the target direction within the target 360-degree image and generate a target planar image for the determined target area. Conversely, the server can determine a target area within the target 360-degree image corresponding to the target planar image determined through object search.

[0138] At operation 880, the server can determine whether target content exists within the target area.

[0139] According to one embodiment, the server may store content associated with the target space prior to operation 810. A method for storing content is described in detail with reference to FIG. 12. In summary, the server may receive information regarding the location of a 360-degree image to be associated with the content among a plurality of 360-degree images from the user's electronic device, as well as the content, and store the content in association with the location. The location information may indicate an area corresponding to an arbitrary direction of the 360-degree image or an arbitrary location (e.g., coordinates) within a planar image corresponding to the area.

[0140] In one embodiment, the server can determine whether target content stored in association with a target area exists among content pre-stored in association with any location within the target space. In other words, the server can determine whether target content stored in association with any location within the target area (or a target plane image corresponding to the target area) exists.

[0141] In operation 890, the server can provide target content to the user terminal if target content exists within the target area.

[0142] In one embodiment, target content may be output to a user terminal by overlapping the query image. The target content may be provided to the user terminal in response to a user inputting a query image (e.g., a captured photo) through the user terminal, or while the query image (e.g., a video captured in real time by the user terminal) is output to the user terminal by overlapping the query image. For example, the server may provide the user terminal with a composite image generated by overlapping the target content with the query image. For example, the user terminal may receive information related to the target content from the server, and output (or display) target content generated based on the information or target content received directly from the server by overlapping the query image.

[0143] According to one embodiment, target content may be overlapped with a query image based on a target location stored in association with the target content. The target location may represent any location (or point) within the target plane image or a target area corresponding to the target plane image. For example, the target content may be overlapped with a specific location within the query image corresponding to the relative location of the target location within the target plane image or the target area and output to the user terminal. For example, if content describing a method of using a manufacturing machine located in a target space is stored near the top of the manufacturing machine, the content describing the method of using the manufacturing machine may be overlapped with a location near the top of the manufacturing machine included in the query image and output to the user terminal.

[0144] According to one embodiment, the target content may include at least one of a photo, video, audio, text, or URL. For example, visual target content such as a photo, video, or text may be displayed on the user terminal by overlapping the query image. For example, auditory target content such as audio may be played while the query image is displayed on the user terminal. For example, when target content such as a URL (or an icon representing a URL) is displayed on the user terminal by overlapping the query image, other linked content may be provided to the user terminal based on a user input for the URL.

[0145] In one embodiment, the server can simultaneously provide information about target content and an indoor location corresponding to the target planar image to the user terminal. For example, a first screen in which the target content overlaps the query image and a second screen in which the location of a 360-degree image corresponding to the target planar image among multiple 360-degree images is displayed on a plan view of the target space can be simultaneously output to the user terminal.

[0146] FIG. 9 is a flowchart illustrating a method for determining one or more candidate 360-degree images using global features according to one embodiment.

[0147] According to one embodiment, the operations 910 and 920 below may be performed by a server (e.g., server (20) of FIG. 1).

[0148] According to one embodiment, operation 830 of FIG. 8 may include operations 910 and 920.

[0149] At operation 910, the server may calculate a first similarity between a target global feature for the query image and a global feature for each of the plurality of 360-degree images. As described with reference to operation 820 of FIG. 8, the server may generate a target global feature for the query image and a global feature for each of the plurality of 360-degree images associated with the target space. The server may generate one global feature for each of the plurality of 360-degree images.

[0150] According to one embodiment, the server may calculate a first similarity by performing an operation (e.g., cosine distance) on a target global feature in vector format and a global feature for each of a plurality of 360-degree images.

[0151] In operation 920, the server can determine one or more candidate 360-degree images from among the plurality of 360-degree images based on the first similarity.

[0152] In one embodiment, the server may determine a 360-degree image among multiple 360-degree images with a first similarity greater than or equal to a first threshold as a candidate 360-degree image. In other words, the server may determine one or more candidate 360-degree images that are similar to the query image and thus highly relevant. The first threshold may be adjusted based on the accuracy of object search.

[0153] FIG. 10 is a flowchart illustrating a method for obtaining a set of planar images according to one embodiment.

[0154] According to one embodiment, the operations 1010 and 1020 below may be performed by a server (e.g., server (20) of FIG. 1).

[0155] According to one embodiment, operation 840 of FIG. 8 may include operations 1010 and 1020.

[0156] In operation 1010, the server can divide a first candidate 360-degree image from among one or more candidate 360-degree images into a plurality of regions.

[0157] In one embodiment, the server may project one or more candidate 360-degree images onto a cube in an orthogonal coordinate system. For example, the server may divide the first candidate 360-degree image into six zones by projecting it onto the cube.

[0158] In one embodiment, the server may project one or more candidate 360-degree images into spherical coordinates. For example, the server may project a first candidate 360-degree image into spherical coordinates and divide it into 18 regions.

[0159] In operation 1020, the server may convert images corresponding to a plurality of regions into a first planar image set for a first candidate 360-degree image. The plurality of planar images within the first planar image sets may include overlapping areas. For example, the server may convert images corresponding to six regions of the first candidate 360-degree image projected onto a cube into a first planar image set including six 2D images.

[0160] Fig. 11 is a flowchart illustrating an object search method using local features according to one embodiment.

[0161] According to one embodiment, the operations 1110 and 1120 below may be performed by a server (e.g., server (20) of FIG. 1).

[0162] According to one embodiment, operation 860 of FIG. 8 may include operations 1110 and 1120.

[0163] In operation 1110, the server may determine a second similarity between a target local feature for the query image and a local feature for each of the planar images of one or more planar image sets. As described with reference to operation 850 of FIG. 8, the server may generate a target local feature for the query image and a local feature for each of the planar images of the one or more planar image sets. The server may generate multiple local features for each of the planar images.

[0164] According to one embodiment, the server can compute a second similarity by performing an operation (e.g., cosine distance) on a target local feature in vector format and a local feature for each of the planar images of one or more planar image sets.

[0165] In operation 1120, the server may determine a target planar image that includes at least a portion of a reference object image corresponding to the target object image among planar images of one or more planar image sets based on the second similarity. For example, if the target object image is an image of a ladder, the reference object image may represent an image of any ladder that can be classified into the same class as the target object image. The server may determine a planar image among the planar images of the one or more planar image sets as the target planar image if the planar image includes a partially cropped portion or all of the reference object image (i.e., the image of the ladder).

[0166] According to one embodiment, the server may determine a planar image having the second highest similarity among planar images of one or more sets of planar images as a target planar image.

[0167] In one embodiment, the server may determine a planar image having the highest second similarity among the planar images of one or more sets of planar images and exceeding a second threshold as the target planar image. The second threshold may be adjusted based on the accuracy of object search.

[0168] According to one embodiment, the server may determine a target planar image using a target local feature and attitude information of a user terminal from which the query image was captured. The attitude information of the user terminal may include information related to an angle or acceleration measured by a sensor (e.g., a gyro sensor) included in the user terminal. The server may refer to attitude information of the user terminal from which the query image was captured to determine a target planar image using the target local feature. For example, if the query image was captured while the user terminal was tilted toward the ground (or in the direction of gravitational acceleration), the server may assign a greater weight to local features of planar images closer to the ground when calculating a second similarity between the target local feature and local features of each of the planar images of one or more planar image sets.

[0169] Figure 12 is a flowchart illustrating a method of storing content according to one embodiment.

[0170] According to one embodiment, operations 1210 to 1230 below may be performed by a server (e.g., server (20) of FIG. 1). Operations 1210 to 1230 may be performed before operation 810 of FIG. 8.

[0171] In operation 1210, the server may receive information about a target location of a target 360-degree image to be associated with target content from an electronic device connected to the server.

[0172] According to one embodiment, when a user selects a target 360-degree image (or a point in the target space corresponding to the target 360-degree image) from among a plurality of 360-degree images associated with the target space through an electronic device, the server may provide the target 360-degree image to the electronic device. When the user selects a target direction of the target 360-degree image through the electronic device, the server may determine a target area corresponding to the target direction among the target 360-degree images. The server may generate a target plane image for the determined target area and provide the image to the electronic device. That is, the server may provide the target 360-degree image to the electronic device in the form of a target plane image for the target area corresponding to the target direction. When the user selects a target location of the target plane image through the electronic device, the server may receive information about the target location. The target location may represent an arbitrary location (or point) within the target plane image or the target area corresponding to the target plane image. For example, the information about the target location may include coordinate information of the target 360-degree image projected onto a 3D coordinate system or the target location for the target area. For example, information about a target location may include coordinate information of the target location with respect to a target plane image projected onto a 2D coordinate system.

[0173] At operation 1220, the server may receive target content from the electronic device.

[0174] According to one embodiment, the server may receive a multimedia file or connection link of target content, such as a photo, video, audio, text, table, or graph, from an electronic device.

[0175] In one embodiment, the server may receive information, such as the type, shape, size, characters, or URL of the target content from the electronic device. For example, the server may provide the electronic device with a basic template for storing the target content and receive input information based on that template.

[0176] At operation 1230, the server may store target content in association with a target location.

[0177] According to one embodiment, the server can store and manage multiple contents associated with multiple locations in the target space.

[0178] Figure 13 is a drawing for explaining a target area according to one embodiment.

[0179] As described above, a server (e.g., server (20) of FIG. 1) may receive a query image including a target object image from a user terminal (e.g., user terminal (30) of FIG. 1) connected to the server, and may convert one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space into one or more planar image sets. For example, the server may divide a target 360-degree image (75) among one or more candidate 360-degree images into a plurality of zones, and may convert images corresponding to the plurality of zones into a target planar image set (71).

[0180] According to one embodiment, the server may determine a target planar image (73) including at least a portion of a reference object image corresponding to a target object image among one or more sets of planar images. As illustrated in FIG. 13, the target planar image (73) may be included in a target planar image set (71) among one or more sets of planar images.

[0181] According to one embodiment, the server may determine a target area (77) of a target plane image (73) within a target 360-degree image (75) corresponding to the target plane image (73) among a plurality of 360-degree images. The server may determine whether target content exists within the target area (77), and if target content exists within the target area (77), provide the target content to a user terminal (e.g., user terminal (30) of FIG. 1).

[0182] In one embodiment, target content may be output to a user terminal based on a target location stored in association with the target content. For example, if the target content is stored in association with a target location (79) within a target area (77), the target content may be output to the user terminal by overlapping a specific location within the query image corresponding to the relative location of the target location (79) within the target area (77).

[0183] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0184] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.

[0185] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may store program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0186] The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.

[0187] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0188] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.

[0189]

[0190] - Assignment information

[0191] Assignment Number: 20023322

[0192] Ministry Name: Ministry of Trade, Industry and Energy

[0193] Research Management Specialist Organization: Korea Institute of Industrial Technology Evaluation and Planning

[0194] Research Project Name: Excellent Corporate Research Institute Promotion Project (ATC+)

[0195] Research Project Title: Advancing an Intelligent Digital Twin Service Platform Based on a Convergent Visual Localization Algorithm that Supports Various Types of Collected Data

[0196] Host organization: 3RII Co., Ltd.

[0197] Research period: January 1, 2024 - December 31, 2024

Claims

1. In a method of searching for an object performed by a server, A step of receiving a query image from a user terminal connected to the above server, wherein the query image includes a target object image; A step of generating a target global feature for the above query image; A step of determining one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space using the above target global features; A step of converting said one or more candidate 360 ​​degree images into one or more sets of planar images; A step of generating a target local feature for the above query image; and A step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets by using the target local feature. Including, How to search for objects.

2. In paragraph 1, The step of determining one or more candidate 360-degree images among a plurality of 360-degree images associated with the target space using the above target global feature is as follows. A step of calculating a first similarity between the target global feature for the query image and the global feature for each of the plurality of 360-degree images; and A step of determining one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity. Including, How to search for objects.

3. In paragraph 2, The step of determining one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity is: A step of determining a 360-degree image among the above multiple 360-degree images having a first similarity greater than or equal to a first threshold as a candidate 360-degree image. Including, How to search for objects.

4. In paragraph 1, The step of converting one or more of the above candidate 360-degree images into one or more sets of planar images comprises: A step of dividing a first candidate 360-degree image among the one or more candidate 360-degree images into a plurality of regions; and A step of converting images corresponding to the above plurality of regions into a first planar image set for the first candidate 360-degree image. Including, How to search for objects.

5. In paragraph 4, A plurality of planar images within the first set of planar images include overlapping areas with each other, How to search for objects.

6. In paragraph 1, The step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more sets of planar images using the target local feature is as follows: a step of calculating a second similarity between the target local feature for the query image and a local feature for each of the planar images of the one or more planar image sets; and A step of determining the target planar image including at least a part of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity. Including, How to search for objects.

7. In paragraph 6, The step of determining the target planar image including at least a part of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity is: A step of determining a planar image among the one or more planar image sets having a second similarity greater than or equal to a second threshold as the target planar image. Including, How to search for objects.

8. In paragraph 1, A step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more sets of planar images using the target local feature is as follows: A step of determining the target plane image by using the target local feature and the detailed information of the user terminal where the query image was captured. Including, How to search for objects.

9. In paragraph 1, A step of providing information on an indoor location corresponding to the target plane image to the user terminal in association with the target space. Including more, How to search for objects.

10. Computer readable recording medium storing a program for performing the method of paragraph 1 11. On the server, Department of Communications; processor; and Memory Including, The above memory, when executed by the processor, causes the server to: Receive a query image from a user terminal connected to the server through the above communication unit, wherein the query image includes a target object image; Generate target global features for the above query image, Using the above target global feature, one or more candidate 360-degree images are determined among a plurality of 360-degree images associated with the target space, Converting one or more of the above candidate 360-degree images into one or more sets of planar images, Generate target local features for the above query image, Using the above target local features, a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets is determined. Stores instructions that are set to be performed, Server.

12. In a content provision method performed by a server, A step of receiving a query image from a user terminal connected to the above server, wherein the query image includes a target object image; A step of generating a target global feature for the above query image; A step of determining one or more candidate 360-degree images among a plurality of 360-degree images associated with a target space using the above target global features; A step of converting said one or more candidate 360 ​​degree images into one or more sets of planar images; A step of generating target local features for the above query image; A step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more sets of planar images using the target local features; A step of determining a target area of ​​the target plane image within the target 360-degree image corresponding to the target plane image among the plurality of 360-degree images; A step of determining whether target content exists within the target area; and If the target content exists within the target area, a step of providing the target content to the user terminal is included. How to provide content.

13. In paragraph 12, The above target content overlaps the above query image and is output to the user terminal. How to provide content.

14. In paragraph 12, A step of receiving information about a target location of the target 360-degree image to be associated with the target content from an electronic device connected to the server; A step of receiving the target content from the electronic device; and A step of storing the above target content in association with the above target location. Including more, How to provide content.

15. In paragraph 14, The target content includes at least one of a photo, video, audio, text, or a uniform resource locator (URL). How to provide content.

16. In paragraph 12, The step of determining one or more candidate 360-degree images among a plurality of 360-degree images associated with the target space using the above target global feature is as follows. A step of calculating a first similarity between the target global feature for the query image and the global feature for each of the plurality of 360-degree images; and A step of determining one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity. Including, How to provide content.

17. In paragraph 16, The step of determining one or more candidate 360-degree images among the plurality of 360-degree images based on the first similarity is: A step of determining a 360-degree image among the above multiple 360-degree images having a first similarity greater than or equal to a first threshold as a candidate 360-degree image. Including, How to provide content.

18. In paragraph 12, The step of converting one or more of the above candidate 360-degree images into one or more sets of planar images comprises: A step of dividing a first candidate 360-degree image among the one or more candidate 360-degree images into a plurality of regions; and A step of converting images corresponding to the above plurality of regions into a first planar image set for the first candidate 360-degree image. Including, How to provide content.

19. In Article 18, A plurality of planar images within the first set of planar images include overlapping areas with each other, How to provide content.

20. In paragraph 12, Using the local features, the step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more sets of planar images is: a step of calculating a second similarity between the target local feature for the query image and a local feature for each of the planar images of the one or more planar image sets; and A step of determining the target planar image including at least a part of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity. Including, How to provide content.

21. In paragraph 20, The step of determining the target planar image including at least a part of a reference object image corresponding to the target object image among the planar images of the one or more planar image sets based on the second similarity is: A step of determining a planar image having the second highest similarity among the planar images of the one or more planar image sets as the target planar image. Including, How to provide content.

22. In paragraph 12, A step of determining a target planar image including at least a portion of a reference object image corresponding to the target object image among the one or more sets of planar images using the target local feature is as follows: A step of determining the target plane image by using the target local feature and the detailed information of the user terminal where the query image was captured. Including, How to provide content.

23. Computer readable recording medium storing a program for performing the method of Article 12 24. On the server, Department of Communications; processor; and Memory Including, The above memory, when executed by the processor, causes the server to: Receiving a query image from a user terminal connected to the server through the above communication unit - the query image includes a target object image -, Generate target global features for the above query image, Using the above target global feature, one or more candidate 360-degree images are determined among a plurality of 360-degree images associated with the target space, Converting one or more of the above candidate 360-degree images into one or more sets of planar images, Generate target local features for the above query image, Using the above target local features, a target planar image is determined that includes at least a portion of a reference object image corresponding to the target object image among the one or more planar image sets, Determine a target area of ​​the target plane image within the target 360-degree image corresponding to the target plane image among the plurality of 360-degree images, Determine whether target content exists within the target area, If the target content exists within the target area, the target content is provided to the user terminal. Stores instructions that are set to be performed, Server.

Citation Information

Patent Citations

  • Image processing system, program, and image processing method

    JP6686547B2

  • Image tracking system using object recognition information based on Virtual Reality, and image tracking method thereof

    KR101634966B1

  • Pharmaceutical composition for preventing or treating heart failure comprising anchovy processed product and / or soybean processed product as an active ingredient

    KR1020220052561A

  • Method for indoor localization using deep learning

    KR102449031B1

  • Object tracking method using multi-camera based drone and object tracking system using the same

    KR102610397B1