Search and guidance system
The system uses a three-dimensional model and server-based object extraction to guide users to desired objects efficiently, overcoming inefficiencies in conventional methods by superimposing route and object marks on captured images without additional markers.
Patent Information
- Application Number
- JP2023100547
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-06-20
AI Technical Summary
Conventional systems require users to manually search for and locate objects, such as books or products, which is inefficient, and existing image-based systems may fail to accurately display the desired object's location if the camera's position and direction are not fixed, necessitating the use of additional markers.
A system equipped with a camera and display that uses a three-dimensional model to identify the user's location and generate route instructions, superimposing them on the captured image, and employs a server to extract and mark the desired object's location without requiring separate markers.
Provides efficient and accurate route guidance to the object's location, allowing users to easily find desired items by superimposing route instructions and marks on the captured image, even with changing camera positions.
Smart Images

Figure 0007775257000001 
Figure 0007775257000002 
Figure 0007775257000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a system for searching for an object such as a book, guiding to the location of the object such as a bookshelf, and extracting and displaying the object from among a plurality of objects. [Background technology]
[0002] 2. Description of the Related Art Libraries and bookstores are currently using systems that display the location of books.
[0003] For example, Patent Document 1 discloses a terminal device that, when a user searches for a book, outputs the location of the shelf where the book is stored, allowing the user to efficiently find the book they are looking for. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2003-72919 [Patent Document 2] Patent 7023338 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the conventional technology such as that disclosed in Patent Document 1, the user has to search for the shelf where the book is stored, which is inefficient.
[0006] Furthermore, even if the shelf on which the desired book is stored is identified, the user must find the location of the book on the shelf on their own. In this regard, Patent Document 2 discloses a system for inventory management that, when an image of a group of books lined up on a shelf is taken, frames are displayed around the book to be checked for inventory. When this system is applied to a user's search for a book, if the image capturing position and direction of the camera capturing the group of books are not fixed, the frame of the desired book may not be displayed correctly. This problem can be solved by capturing an image of a marker at the same time as capturing the image and correcting the position, but there is a problem in that the marker must be prepared.
[0007] The above-mentioned problems occur not only when the object is a book, but also when the object is a product, a part, or the like.
[0008] SUMMARY OF THE INVENTION An object of the present invention is to solve any of the above problems and to provide an efficient search guidance system and object extraction system. [Means for solving the problem]
[0009] Several independently applicable features of the present invention are listed below.
[0010] (1)(2) The search guidance device of the present invention is equipped with a camera and a display, and is configured to be able to access a three-dimensional model for identifying a target area and a model for a route. Based on an image of the target area captured by the camera, the device identifies an imaging position in the target area by referring to the three-dimensional model for identification, generates route instructions from the identified imaging position to a location where a desired object is placed based on the model for the route, and displays guidance to the location where the desired object is placed on the display based on the generated route instructions.
[0011] Therefore, the imaging position can be identified based on the captured image, and the route to the position where the desired object is placed can be shown, so that effective route guidance can be provided with a simple configuration.
[0012] (3) The search and guidance device of the present invention is characterized in that the three-dimensional identification model is three-dimensional feature point data generated based on an image of the target area, and the route model is data that can distinguish at least between passageways and non-passageways.
[0013] Therefore, it is easy to generate a three-dimensional identification model and a route model.
[0014] (4) The search and guidance device of the present invention is characterized in that the guidance acquires the imaging position and imaging direction in the identified target area in real time by referring to the identification three-dimensional model based on an image of the target area captured by a camera, determines how the route instructions generated based on the route model will be displayed in two dimensions in the imaging direction at the imaging position, and displays the two-dimensional display of the route instructions superimposed on the image captured by the camera.
[0015] Therefore, the route instructions are displayed on an image that represents the surroundings that the user is currently viewing, so the user can easily know which direction to go.
[0016] (5) When the search and guidance device of the present invention determines that it has reached the location where the desired object is placed, it transmits a multiple item display image of multiple items captured by a camera to an object extraction server device, and displays an area mark on the display superimposed on the multiple item display image of multiple items captured by the camera based on area mark information of the object display of the target object, which is the object, received from the object extraction server device.The object extraction server device extracts an area of the object display of the target object from the multiple item display captured image received from the search and guidance terminal device, and transmits area mark information indicating the area of the object display of the extracted target object to the object extraction terminal device.
[0017] Therefore, it is possible to identify the specific location of the target item.
[0018] (6) The search guidance device of the present invention is characterized in that it uses a multiple item display image as a marker and controls the area mark so that it is displayed superimposed on the item display of the target item even if the current multiple item display image captured by the camera changes.
[0019] Therefore, even without preparing a separate marker, the area mark can be displayed superimposed on the target object in accordance with the change in the image capturing position by the user.
[0020] (7)-(10) A search guidance system according to the present invention is a search guidance system comprising a search guidance terminal device and a search guidance server device that is capable of communicating with the search guidance terminal device, the search guidance terminal device is equipped with a camera and a display, transmits an image of the target area captured by the camera to the search guidance server device, and displays on the display guidance to a location where a desired target object is placed based on a route instruction received from the search guidance server device; The search guidance server device is configured to be able to access a three-dimensional identification model and a route model for the target area, and is characterized in that it refers to the three-dimensional identification model based on the captured image received from the search guidance terminal device to identify the imaging position and imaging direction in the target area, and generates route instructions from the identified imaging position to the location where the desired object is placed based on the route model and transmits them to the search guidance terminal device.
[0021] Therefore, the imaging position can be identified based on the captured image, and the route to the position where the desired object is placed can be shown, so that effective route guidance can be provided with a simple configuration.
[0022] (11) The search and guidance system of the present invention is characterized in that the three-dimensional identification model is three-dimensional feature point data generated based on an image of the target area, and the route model is data that can distinguish at least between passageways and non-passageways.
[0023] Therefore, it is easy to generate a three-dimensional identification model and a route model.
[0024] (12) The search and guidance system of the present invention is characterized in that the guidance in the search and guidance terminal device acquires the imaging position and imaging direction in the identified target area in real time by referring to the identification three-dimensional model based on an image of the target area captured by a camera, determines how the route instructions generated based on the route model will be displayed in two dimensions in the imaging direction at the imaging position, and displays the two-dimensional display of the route instructions superimposed on the image captured by the camera.
[0025] Therefore, the route instructions are displayed on an image that represents the surroundings that the user is currently viewing, so the user can easily know which direction to go.
[0026] (13) The search guidance system of the present invention is characterized in that the search guidance terminal device acquires identification information of the target item, transmits a multiple item display image of multiple items captured by a camera to the object extraction server device, identifies the target item based on the target item identification information received from the object extraction server device, and displays an area mark for the target item on the display, superimposing it on the multiple item display image of multiple items captured by the camera, and the object extraction server device generates target item identification information for identifying the target item based on the multiple item display image received from the object extraction terminal device and transmits it to the object extraction terminal device.
[0027] Therefore, it is possible to identify the specific location of the target item.
[0028] (14) The search guidance system according to the present invention is characterized in that the search guidance terminal device uses a multiple item display image as a marker and controls the area mark to be displayed superimposed on the item display of the target item even if the current multiple item display image captured by the camera changes.
[0029] Therefore, even if a special marker is not separately prepared and even if the imaging position is changed by the user, the area mark can be displayed superimposed on the target object.
[0030] (15) In the search guidance system according to the present invention, the target area is a bookshelf on which books are placed.
[0031] Therefore, it is possible to guide the user along the route to the desired bookshelf.
[0032] (16)-(19) An object extraction system according to the present invention is an object extraction system including an object extraction terminal device and an object extraction server device that is capable of communicating with the object extraction terminal device, the object extraction terminal device includes a camera and a display, acquires identification information of the object, transmits a multiple object display image of the multiple objects captured by the camera to the object extraction server device, identifies the object based on the object identification information received from the object extraction server device, and displays an area mark for the object on the display, superimposed on the multiple object display image of the multiple objects captured by the camera; The object extraction server device is characterized in that it generates object identification information for identifying the object based on the multiple object display captured image received from the object extraction terminal device, and transmits the information to the object extraction terminal device.
[0033] The object extraction server device is characterized by extracting an object display area of the target object from a captured image displaying multiple objects received from the object extraction terminal device, and transmitting area mark information indicating the object display area of the extracted target object to the object extraction terminal device.
[0034] Therefore, it is possible to find a desired item from among a large number of target items.
[0035] (20) The object extraction system of the present invention is characterized in that the object extraction server device generates correspondence information including image position information and identification information of each object included in the multiple-object display captured image as information for identifying the object, and transmits the correspondence information to the object extraction terminal device, and the object extraction terminal device identifies image position information that matches the identification information of the object based on the correspondence information received from the object extraction server device as image position information of the object, and displays the area mark based on the image position information.
[0036] Therefore, it is possible to find a desired item from among a large number of target items.
[0037] (21) The object extraction system according to the present invention is characterized in that the object extraction server device extracts an object display image of each object based on the multiple object display images, estimates image position information of each object, and estimates identification information of the corresponding object based on each extracted object display image.
[0038] Therefore, the identification information can be estimated more accurately.
[0039] (22) The object extraction system according to the present invention is characterized in that an object extraction terminal device transmits identification information of the target object to an object extraction server device, the object extraction server device acquires a target object display image based on the received identification information of the target object, identifies the target object display image from the multiple object display images and estimates its image position information, and transmits the image position information to the object extraction terminal device as information for identifying the target object, and the object extraction terminal device displays the area mark based on the received image position information.
[0040] Therefore, it is possible to find a desired item from among a large number of target items.
[0041] (23) The object extraction system according to the present invention is characterized in that the object extraction terminal device uses the multiple item display image as a marker and controls the area mark to be displayed superimposed on the item display of the target item even if the current multiple item display image captured by the camera changes.
[0042] Therefore, even if a special marker is not separately prepared and even if the imaging position is changed by the user, the area mark can be displayed superimposed on the target object.
[0043] (24) In the object extraction system according to the present invention, the articles are books, and the multiple article display image is a multiple spine image obtained by capturing images of spines of a plurality of the books arranged side by side.
[0044] Therefore, the user can easily find the desired book.
[0045] The concept of "device" includes not only what is constituted by one computer, but also what is constituted by multiple computers connected via a network, etc. Therefore, when the means of the present invention (or even a part of the means) is distributed among multiple computers, these multiple computers correspond to the device.
[0046] The term "program" is a concept that includes not only programs that can be executed directly by a CPU, but also programs in source format, compressed programs, encrypted programs, and programs that work in conjunction with an operating system to perform their functions. [Brief explanation of the drawings]
[0047] [Figure 1] 1 shows a functional configuration of a search and guidance device according to an embodiment of the present invention. [Figure 2] 1 shows the system configuration of a search and guidance system using a search and guidance device. [Figure 3] 1 shows the hardware configuration of a search and guidance device. [Figure 4] 1 shows the hardware configuration of a server device. [Figure 5] 10 is a flowchart of a search and guidance process. [Figure 6] FIG. 6A is an example of three-dimensional feature point data, FIG. 6B is an example of a three-dimensional model, and FIG. 6C is an example of a route model. [Figure 7] FIG. 2 is a diagram illustrating a captured image and feature points. [Figure 8] FIG. 10 is a diagram illustrating a route setting using a route model. [Figure 9] FIG. 10 is a diagram showing route instructions displayed on a smartphone SP. [Figure 10] 10 is a functional configuration of a search and guidance system according to another embodiment. [Figure 11] 10 is a flowchart of a search and guidance system according to another embodiment. [Figure 12] 10 is a functional configuration of an object extraction system according to a second embodiment. [Figure 13] 10 is a flowchart of an object extraction process. [Figure 14] 10 is an example of a multiple spine image (reference image). [Figure 15] 10 is an example of an extracted spine image and image position information. [Figure 16] 10A and 10B are diagrams illustrating coordinates of image position information and a method for specifying the coordinates. [Figure 17] FIG. 17A is a diagram showing a plurality of spine images (reference images) as learning data, and FIG. 17B is a diagram showing a schematic diagram of extraction of spine images of each book. [Figure 18] FIG. 10 is a diagram showing a spine image and an estimated ISBN code. [Figure 19] FIG. 10 is a diagram showing spine images and ISBN codes used as learning data. [Figure 20] 10 is an example of correspondence data. [Figure 21a] FIG. 10 is a diagram showing a captured image with an area mark M added thereto. [Figure 21b] FIG. 10 is a diagram showing a captured image with an area mark M added thereto. [Figure 22]10 is a functional configuration of an object extraction system according to another example. [Figure 23] 10 is a flowchart of another example of an object extraction process. [Figure 24] FIG. 10 is a diagram showing a correspondence table between spine images and ISBN codes. [Figure 25] FIG. [Figure 26] 26A is a diagram showing a plurality of spine images (FIG. 26A), a spine image (FIG. 26B), and image position information (FIG. 26C) as learning data. DETAILED DESCRIPTION OF THE INVENTION
[0048] 1. First embodiment 1.1 Functional configuration 1 shows the functional configuration of a search guidance device according to one embodiment of the present invention. A user captures an image of a target area using a camera 2 of a search guidance terminal device T. In a captured image acquisition process 4, the search guidance terminal device T acquires the captured image captured by the camera 2.
[0049] The search and guidance device stores a three-dimensional identification model 14, which is a three-dimensional model of the target area. The three-dimensional identification model 14 is a three-dimensional model that includes at least the feature points of the target area.
[0050] In the position / direction specification process 12, the specification three-dimensional model 14 is referenced to specify from which location in the target area and in which direction the acquired captured image was captured.
[0051] The search and guidance device also stores a route model 18 used to generate route instructions. The route model 18 is a model configured to be able to distinguish at least between passages and non-passages.
[0052] In route instruction creation processing 16, the search and guidance device generates a route from the location identified in the position / direction identification processing to the location where the desired object is placed, by referring to a three-dimensional route model 18. In route guidance display processing 8, the search and guidance device superimposes the route on the image currently being captured by camera 2 and displays it on display 6.
[0053] As described above, the system uses a three-dimensional identification model to identify the current location based on the captured image, generates a route to the target location, and displays it superimposed on the captured image, thereby providing accurate and easy-to-understand route guidance to the target.
[0054] 1.2 System configuration and hardware configuration Figure 2 shows the system configuration of a search and guidance system using a search and guidance device. In this example, a smartphone SP is used as the search and guidance device. The smartphone SP is configured to be able to communicate with a server device S via the Internet.
[0055] The hardware configuration of the smartphone SP is shown in Figure 3. Connected to a CPU 20 are a memory 22, a touch display 24, a non-volatile memory 26, a camera 28, and a communication circuit 30. Note that the call circuitry and the like are omitted.
[0056] The communication circuit 30 is a circuit for connecting to the Internet. The non-volatile memory 26 stores an operating system 34 and a search and guidance program 36. The search and guidance program 36 works in cooperation with the operating system 34 to perform its functions.
[0057] 4 shows the hardware configuration of the server device S. A CPU 40 is connected to a memory 42, an SSD 44, and a communication circuit 46.
[0058] The communication circuit 46 is a circuit for connecting to the Internet. The SSD 44 stores an operating system 48 and a server program 50. The server program 50 cooperates with the operating system 48 to perform its functions.
[0059] 1.3 Search Guidance Processing The following describes a case where a user with a smartphone SP searches for a shelf in a library where a desired book is located.
[0060] A flowchart of the search guidance process is shown in Fig. 5. The CPU 40 of the server device S (hereinafter sometimes abbreviated as server device S) transmits three-dimensional feature point data and a route model of the library to the smartphone SP (step S21).
[0061] The three-dimensional feature point data is data that constructs feature points (such as vertices) in three dimensions based on a large number of captured images of the inside of a library taken from different locations (taken with some images overlapping and shifted). Therefore, the three-dimensional feature point data is data that shows the feature points inside the library in three dimensions, as shown in Figure 6A. The mapping app immersal (trademark) by Immersal Inc. can be used to generate the three-dimensional feature point data.
[0062] As shown in Figure 6C, the route model is a floor plan of the building that distinguishes between aisles where people can move and non-aisles where people cannot. In this embodiment, the route model is created based on a three-dimensional model such as that shown in Figure 6B. The three-dimensional feature point data and the three-dimensional model are superimposed based on their shapes, and the floor of the three-dimensional model is extracted to generate the route model, so that the positions of the three-dimensional feature point data and the route model on the coordinate system match. Note that the route model records the identification codes of shelves in association with the positions where the shelves are located.
[0063] The CPU 20 of the smartphone SP (sometimes abbreviated as smartphone SP) receives the three-dimensional feature point data and the route model, and records them in the nonvolatile memory 26 (step S1).
[0064] Next, the smartphone SP acquires the identification code of the shelf where the desired book is located (step S2). For example, the user can search for a book using a book search system installed in the library and input the shelf identification code obtained by the user into the smartphone SP. Alternatively, the smartphone SP can be used to read a QR code (trademark) containing the shelf identification code displayed on the display of the book search system.
[0065] Next, the user operates the camera 28 of the smartphone SP to capture an image of the surroundings. Note that the image capture referred to here refers to the act of launching the camera application of the smartphone SP to capture an image, and there is no need to press the shutter (the shutter may be pressed). Since the camera application is automatically launched when the search guidance program is launched, the camera 28 continues to capture images, and the captured image changes from moment to moment depending on the orientation of the smartphone SP by the user, etc. The smartphone SP acquires this captured image (step S3). Next, as shown in FIG. 7, the smartphone SP calculates feature points FT (points that indicate the characteristics of the object, such as the vertices of the object) in this two-dimensional captured image (step S4). It is preferable to calculate these feature points using the same calculation method as used when generating the three-dimensional feature point data described above. Note that since the surroundings are captured, two-dimensional captured images and feature points in different directions can be obtained.
[0066] Next, the smartphone SP refers to the three-dimensional feature point data based on the feature points of the surrounding two-dimensional captured image, and identifies the user's location (captured position) and direction (captured direction) (step S5). If a certain position in the three-dimensional feature point data matches the user's position (captured position), the feature point data when viewing the surroundings from that position in the three-dimensional feature point data matches the feature point data generated in step S4. Therefore, the smartphone SP can identify the user's location. Such location identification can be achieved using the aforementioned immersal (trademark).
[0067] Next, the smartphone SP refers to the route model based on the identification code of the shelf and identifies the position of the shelf (step S6). As described above, the route model indicates the position of each shelf by its identification code, so the position of the shelf can be easily identified.
[0068] The smartphone SP refers to the route model, calculates a route from the current position identified in step S5 to the desired shelf position (the shortest route along the aisle), and generates route instructions along that route (step S7). For example, as shown in Fig. 8, a route from the current position CP to the shelf position TP is calculated. This process can be realized by using NavMesh from Unity Technologies.
[0069] The smartphone SP acquires an image captured by the camera 28 (step S8) and displays route instructions superimposed on the captured image. At this time, the current position and direction are identified in the same manner as in steps S4 and S5, and the route instructions are superimposed (step S9). As a result, the route instructions GL are displayed on the touch display 24 of the smartphone SP superimposed on the captured image, as shown in FIG.
[0070] The smartphone SP determines whether the user's current location has reached the location of the desired shelf (step S10). If not, steps S4 to S9 are repeated to update the route display in real time. If the user has reached the desired shelf, the guidance process ends.
[0071] In this way, the user can be provided with guidance to the shelf where the desired book is located.
[0072] If the route is in a direction different from the direction captured by the user (for example, the opposite direction), the route cannot be displayed on the captured image. In this case, instructions such as "Please turn around" or "Please look in the opposite direction" can be displayed on the smartphone SP using text or arrows.
[0073] 1.4 Variations (Other) (1) In the above embodiment, the target objects are books stored on shelves. However, the target objects may be things other than books, such as merchandise. That is, although libraries and bookstores are used as the target areas, factories, warehouses, shops, etc. may also be used as the target areas. Furthermore, in addition to shelves, objects on which multiple targets can be placed, such as stands, may also be used.
[0074] (2) In the above embodiment, the three-dimensional feature point data is generated based on the captured image of the target area. However, it is also possible to obtain three-dimensional point cloud data using Lidar or the like and generate the three-dimensional feature point data based on this data.
[0075] (3) In the above embodiment, the shelf identifiers are recorded in the route model, and the shelf location is identified by acquiring the identifier of the target shelf. However, the shelf location may also be identified by acquiring the coordinates.
[0076] (4) In the above embodiment, the identification three-dimensional model and the route model are downloaded from the server device S to the smartphone SP. However, instead of downloading, the server device S may be accessed and referenced each time.
[0077] (5) In the above embodiment, route guidance is provided by superimposing route instructions in real time on an image captured by the smartphone SP. However, a map or the like showing the route to the target shelf may also be displayed on the smartphone SP.
[0078] (6) In the above embodiment, a book is searched for using a search terminal device (such as an OPAC terminal) installed in the library, and the identification code of the shelf on which the book is stored is read into the smartphone SP using a QR code (trademark) or the like (or a coded code such as a two-dimensional barcode or barcode).
[0079] In this case, if the coded code contains the location of the search terminal device, the smartphone SP can easily and quickly identify the location and direction in step S5.
[0080] (7) In the above embodiment, the current location is determined based on three-dimensional feature point data. However, by referring to current information obtained by GPS, Wi-Fi, etc., it becomes easier to determine the location using three-dimensional feature point data.
[0081] Alternatively, a QR code (trademark) containing a code for identifying a location may be attached to a bookshelf, and the code may be captured by a smartphone to identify the current location. In this case, the current location can be identified without using three-dimensional feature point data.
[0082] (8) In the above embodiment, a book is searched for using a search terminal device (such as an OPAC terminal) installed in the library, and the identification code of the shelf on which the book is stored is read into the smartphone SP using a QR code (trademark) or the like.
[0083] However, it is also possible to connect to the search server device from the smartphone SP, search for books, and obtain the identification code of the storage shelf, which eliminates the need for the user to input the identification code or read the QR code (trademark).
[0084] (9) In the above embodiment, the processes of steps S1 to S10 in Fig. 5 are performed by the smartphone SP. However, some or all of the processes may be performed by a server device, and the smartphone SP may receive and display the processing results.
[0085] FIG. 10 shows an example of the functional configuration of a system in which some of the processing is executed by the server device S in this way.
[0086] The user captures an image of the target area using the camera 2 of a search guidance terminal device T such as a smartphone SP. In a captured image transmission process 4, the search guidance terminal device T transmits the captured image captured by the camera 2 to the search guidance server device S.
[0087] The captured image is received by search guidance server device S. A three-dimensional identification model 14, which is a three-dimensional model of the target area, is recorded in search guidance server device S. The three-dimensional identification model 14 is a three-dimensional model that includes at least the feature points of the target area.
[0088] In the position / direction identification process 12, the search guide server device S refers to the identification three-dimensional model 14 and identifies from which location in the target area and in which direction the received captured image was captured.
[0089] A three-dimensional route model 18 used to generate route instructions is recorded in the search guidance server device S. The three-dimensional route model 18 is a three-dimensional model configured to be able to distinguish at least between passages and non-passages.
[0090] In the route instruction creation process 16, the search guide server device S generates a route from the location identified by the position / direction identification process to the location where the desired object is placed, by referring to the three-dimensional route model 18. The search guide server device S transmits the generated route to the terminal device T.
[0091] This route is received by the search and guidance terminal device T. In the route guidance display process 8, the search and guidance terminal device T superimposes the route on the image currently being captured by the camera 2 and displays it on the display 6.
[0092] A processing flowchart of the system shown in Fig. 10 is shown in Fig. 11. In steps S31 and S8, a captured image is sent from the smartphone SP, and the server device S identifies the position and direction, generates route instructions, and sends them back to the smartphone SP.
[0093] The captured image transmitted from the smartphone SP may be the captured image itself, or may be a captured image from which only feature points have been extracted.
[0094] The process allocation between the smartphone SP and the server device S shown in FIG. 11 is an example, and any process allocation may be adopted.
[0095] (10) In the above embodiment, a smartphone SP is used. However, instead of this, a tablet computer, a portable PC such as a PDA, or a wearable device such as smart glasses may be used.
[0096] (11) The above-described embodiments and their modifications may be implemented in combination with each other, or may be implemented in combination with other embodiments or modifications.
[0097] 2. Second embodiment 2.1 Functional configuration 12 shows the functional configuration of the object extraction system according to the second embodiment. A user uses the camera 102 of the terminal device T to capture an image of a plurality of items arranged in a row (for example, the spines of a plurality of books arranged on a shelf) as a multiple-item display image. In a captured image transmission process 104, the terminal device T transmits the multiple-item display image captured by the camera 102 to the server device S.
[0098] The server device S receives the multiple item display image, and in a correspondence information generation process 112, extracts each item display image included in the multiple item display image using an estimation model 114, and estimates identification information of the item shown by each item display image. Furthermore, the server device S generates position information (image position information) on the image of each item display image in the multiple item display image, and transmits correspondence information in which the identification information of each item is associated with this information to the terminal device T as information for identifying the target item. In this embodiment, the correspondence information generation process 112, the correspondence information transmission process 116, and the estimation model 114 configure a target item identification information generation and transmission process 111.
[0099] Here, the estimation model 114 can be a trained model that has been machine-learned to extract individual item display images based on multiple item display images, and a trained model that has been machine-learned to estimate item identification information based on item display images.
[0100] Terminal device T receives correspondence information associating image position information and identification information of each item, and in area mark superimposition display processing 108, if the correspondence information includes the identification information of the target item acquired in identification information acquisition processing 109, identifies the corresponding image position information. Based on the image position information of the target item, an area mark is superimposed on the target item in the multiple item display image. This superimposed image is displayed on display 106.
[0101] As described above, the area mark of the target item is displayed superimposed on the captured multiple item display image, so that the user can easily find the target item.
[0102] 2.2 System and hardware configuration The system configuration and hardware configuration are the same as those in Figures 2 to 4 in the first embodiment. However, an object extraction terminal program is recorded in the nonvolatile memory 26 of the smartphone SP instead of or in addition to the search guide program 36. Also, an object extraction server program is recorded in the SSD 44 of the server device S instead of or in addition to the server program 50.
[0103] 2.3 Object extraction processing A flowchart of the object extraction process is shown in Fig. 13. In this embodiment, a case where books stored in a bookshelf with multiple shelves is to be found will be described.
[0104] The user inputs the ISBN code of the book to be found into the smartphone SP as identification information of the book. The ISBN code of the target book is obtained, for example, by the user searching for the book using a book search system installed in a library. Alternatively, this can be achieved by reading a QR code (trademark) containing the ISBN code of the book displayed on the display of the book search system. In this way, the smartphone SP obtains the ISBN code of the book (step S101).
[0105] The user goes to the shelf where the desired book is placed. The route to the shelf may be guided by the guidance according to the first embodiment (guidance by another method), or the user may find the shelf by themselves.
[0106] A user who comes in front of a shelf captures an image of one shelf level with the camera 28 of the smartphone SP. Note that the image capture referred to here is captured by activating the camera application of the smartphone SP, and there is no need to press the shutter (the shutter may be pressed). Since the camera application is automatically activated when the object extraction program is activated, the camera 28 continues to capture images, and the captured image changes from moment to moment depending on the orientation of the smartphone SP by the user, etc. This captured image shows the spines of multiple books lined up on the shelf (multiple spine display image). The captured multiple spine image (multiple item display image) is sent to the server device S (step S102). The smartphone SP also records the multiple spine image at this time as a reference image (step S103).
[0107] The server device S receives the spine images and records them as reference images (step S123). Fig. 16 shows an example of the spine images (reference images).
[0108] The server device S extracts the spine image of each book and its image position (position coordinates on the image) based on the spine images acquired in step S122 (step S122).
[0109] The extraction results are shown in Figure 15. The spine images of each book extracted from the multiple spine images and their image positions are recorded. As shown in Figure 16, the image positions are expressed as coordinates of points A, B, C, and D at the four corners of the spine image, with the origin O at the bottom left of the multiple spine images, the horizontal axis as X, and the vertical axis as Y. Note that different coordinate methods may be used, or the image may be specified by two points on the spine image (for example, A and C).
[0110] In order to obtain the extraction results shown in FIG. 15, in this embodiment, a trained estimation model is used.
[0111] This estimation model was trained as follows. An example of the training data used for training is shown in FIG. 17. FIG. 17A is the input data, and FIG. 17B is the correct data (data for identifying the spine (for example, data indicated by the coordinates of the four corners)). Note that while FIG. 17 shows the same multiple background images as FIG. 14, multiple background images of different books are actually used. A large amount of such training data is prepared for various books, and the estimation model is trained. If multiple spine images are input to the model trained in this way, the spine of each book can be extracted. Note that it is preferable to use a CNN or a Vision Transformer as the estimation model.
[0112] The server device S inputs the reference image of FIG. 14 into this trained estimation model, and obtains spine images of individual books as shown in FIG. 15 (step S122).
[0113] Next, the server device S estimates the ISBN code of each book based on the spine image of the book (step S123).
[0114] The estimation result is shown in Figure 18. The ISBN code of each book is shown in association with the spine image of each book in Figure 15. As shown in Figure 18, in this embodiment, a trained estimation model (a trained estimation model different from the trained estimation model used in step S122) is used to estimate the ISBN code in association with each spine image.
[0115] This estimation model was trained as follows. An example of the training data used for training is shown in FIG. 19. The spine images of each book in FIG. 19 are the input data, and the ISBN code is the correct answer data. A large number of such training data are prepared for various books, and the estimation model is trained using them. By inputting the spine image into the model trained in this way, the ISBN code of the book can be estimated. It is preferable to use CNN as the estimation model.
[0116] The server device S inputs the spine image of each book in FIG. 15 into this trained estimation model, and obtains the ISBN code of each book as shown in FIG. 18 (step S123).
[0117] Next, based on the results of Figures 15 and 18, the server device S generates correspondence information that associates the image position of each book included in the reference image (multiple spine images) with the ISBN code (step S124), as shown in Figure 20. The server device S transmits the generated correspondence information to the smartphone SP as target item identification information (step S125).
[0118] The smartphone SP receives the correspondence information (step S104). Then, it determines whether the ISBN code of the target book is included in this correspondence information (step S105). If not, it means that the target book is not included in the group of books imaged in step S102. In this case, the process returns to step S102, and multiple spine images of the next shelf are imaged and sent to the server device S. Thereafter, the same process as above is performed.
[0119] Furthermore, if the transmitted correspondence information includes the ISBN code of the target book, the smartphone SP displays the target book with an area mark M superimposed on it, as shown in FIG. 21a. This process is performed as follows: First, the image position corresponding to the ISBN code of the target book in the correspondence data is acquired. This image position indicates the position of the spine of the target book in the multiple spine image. Therefore, the smartphone SP displays the area mark M superimposed on this image position.
[0120] However, although the user points the smartphone SP at the shelf, its angle and position are not completely fixed, so the captured image changes from moment to moment. Therefore, if the area mark M is displayed at the above image position, the area mark M will be out of alignment with the spine image of the target book. Therefore, in this embodiment, the reference image is used as a marker (using feature points in the reference image) to estimate changes in the position and angle of the multiple spine images currently captured by the smartphone SP (by comparing the feature points of the reference image with the feature points of the current multiple spine images), and the area mark M is controlled to be displayed correctly on the spine image of the target book accordingly (steps S106 and S107).
[0121] For example, even if the orientation of the smartphone SP changes and the captured image becomes as shown in FIG. 21b, the area mark will still be displayed on the spine of the target book.
[0122] Therefore, even if a special marker is not provided on the bookshelf, the area mark M can be displayed correctly in response to changes in the orientation of the smartphone SP.
[0123] In this way, the location of the target book on the bookshelf can be communicated to the user in an easy-to-understand manner.
[0124] It should be noted that which of the terminal device T and the server device S performs each process shown in FIG. 12 is merely an example, and the division of the processes is arbitrary.
[0125] 2.4 Variations (Other) (1) In the above embodiment, as shown in Fig. 23, processing is shared between the smartphone SP and the server device S. However, some or all of the processing performed by the smartphone SP may be performed by the server device S. Also, some or all of the processing performed by the server device S may be performed by the smartphone SP.
[0126] In the above embodiment, the server device S generates correspondence information that associates the image position of each book included in multiple spine images (reference images) with the ISBN code, and transmits this to the smartphone SP. Based on this correspondence information, the smartphone SP obtains the image position corresponding to the ISBN code of the target book.
[0127] However, the server device S may also be configured to identify the image position corresponding to the ISBN code of the target book. For example, this can be achieved as follows. When the smartphone SP transmits the reference image to the server device S (step S102 in FIG. 13), it also transmits the identification information of the target book. After generating the correspondence information, the server device S identifies the image position of the target book based on the identification information of the target book. The server device S transmits the image position of the target book to the smartphone SP as information for identifying the target item. In this embodiment, the image position of the target book becomes the information for identifying the target item.
[0128] (2) In the above embodiment, the smartphone SP identifies the image position of the target book based on the correspondence information, which is information for identifying the target item, and generates the area mark M. However, the server device S may identify the image position of the target book by the following process and transmit this to the smartphone SP as information for identifying the target item.
[0129] The functional configuration of the system in this case is shown in Fig. 22. The user uses the camera 102 of the terminal device T to capture an image of multiple items arranged in a row (for example, the spines of multiple books lined up on a shelf) as a multiple-item display image. In a captured image transmission process 104, the terminal device T transmits the multiple-item display image captured by the camera 102 to the server device S.
[0130] The terminal device T acquires the identification information of the target book in the identification information acquisition process 109 and transmits it to the server device S.
[0131] The server device S receives the multiple item display image, and in a target item region extraction process 117, extracts a target item from the multiple item display image using an estimation model 119. Here, as the estimation model 119, a trained model that has been machine-learned to estimate and identify a target item from an image of multiple items can be used.
[0132] When the target item is identified, the server device S transmits image position information of the item in the multiple item display image to the terminal device T in an image position information transmission process 118. In this embodiment, the target item identification information generation and transmission process 111 is configured by the corresponding item area extraction process 117, the image position information transmission process 118, and the estimation model 119. Upon receiving the image position information, the terminal device T superimposes an area mark on the target item in the multiple item display image captured by the camera based on the image position information in area mark superimposition display processing 108. This superimposed image is displayed on the display 106.
[0133] As described above, the area mark of the target item is displayed superimposed on the captured multiple item display image, so that the user can easily find the target item.
[0134] A flowchart of the above process is shown in Figure 23. The user inputs the ISBN code of the book to be found into the smartphone SP. The smartphone SP transmits the acquired ISBN code of the book to the server device S (step S101).
[0135] The server device S receives the ISBN code of the book (step S131). A table as shown in FIG. 24, which associates the ISBN code of the book with the image of the spine, is recorded in advance in the server device S. The server device S obtains the image of the spine corresponding to the received ISBN code from the table (step S132). Here, it is assumed that the image of the spine of the target book as shown in FIG. 25 has been obtained.
[0136] A user who comes to a shelf takes an image of one shelf with the camera 28 of the smartphone SP. The captured image (image showing multiple items) is sent to the server device S (step S102). The smartphone SP also records the captured image as a reference image (step S103).
[0137] The server device S receives the captured image and records it as a reference image (see FIG. 14) (step S133).
[0138] If the spine image acquired in step S132 is included in the reference image, the server device S extracts it (step S134). In this embodiment, this process is performed using a trained estimation model.
[0139] This estimation model was trained as follows. An example of the training data used for training is shown in FIG. 26. FIGS. 26A and 26B are input data, and FIG. 26C is correct answer data (coordinate data for identifying the frame (see FIG. 16)). The correct answer data in FIG. 26C indicates the position of the input data in FIG. 26B in the input data in FIG. 26A. Note that although the training data in FIG. 26 is the same as that in FIG. 14, images of different books are actually used.
[0140] A large amount of such training data is prepared for various books, and an estimation model is trained. By inputting a spine image and a captured image (an image of the spines of multiple books lined up in a row) into the trained model, it is possible to extract the spine image contained in the captured image. Note that YOLO, for example, can be used as the estimation model.
[0141] The server device S inputs the reference image of FIG. 14 and the spine image of FIG. 25 into this trained estimation model (step S134). If the spine image is included in the reference image, the server device S transmits the coordinates (image position) of the spine in the reference image to the smartphone SP as information for identifying the target item (steps S136 and S138). If the spine image is not included in the reference image, the server device S transmits the result information indicating that identification was unsuccessful (step S137) to the smartphone SP (step S138).
[0142] The smartphone SP receives result information from the server device S (step S104). If the result information indicates that the identification was successful and includes the image position, an area mark is generated based on the coordinates, superimposed on the captured image, and displayed on the touch display 24. As a result, an image in which the spine of the target book is surrounded by the area mark M can be displayed on the display 24, as shown in FIG. 21. Therefore, the user can easily find the location of the target book by looking at this image of the bookshelf in front of them. In this case, the area mark M is made to follow the reference image as a marker, as described above.
[0143] On the other hand, if the result information received in step S104 indicates that identification was unsuccessful, the smartphone SP knows that the target book is not on the shelf that was imaged. In this case, the smartphone SP displays an instruction on the touch display 24 to instruct the user to image another shelf. In response to this, when the user images the next shelf on the bookshelf, the smartphone SP transmits this to the server device S (step S102).
[0144] Thereafter, this captured image is used as a reference image and step S102 and subsequent steps are repeatedly executed.
[0145] In this way, the location of the target book on the bookshelf can be communicated to the user in an easy-to-understand manner.
[0146] It should be noted that which of the terminal device T and the server device S performs each process shown in FIG. 22 is merely an example, and the allocation of the processes is arbitrary.
[0147] (3) In the above embodiment, the user is required to consider the route to reach the desired bookshelf.
[0148] However, the route to the bookshelf may be displayed on the smartphone SP using the search guidance process of the first embodiment. In this case, when the server device S determines that the user has reached the target bookshelf, it may display a message prompting the user to capture an image of the bookshelf (for example, "Turn your smartphone sideways and capture an image of one shelf") on the touch display 24.
[0149] (4) In the above embodiment, the case of locating a book has been described, but the present invention can also be applied to locating general objects such as merchandise. In this case, the target merchandise can be located based on an image of the back of a merchandise storage box, for example.
[0150] In the above embodiment, the ISBN code is used as the identification information for the book, but a combination of the book title, author name, publisher name, etc. may also be used as the identification information. For products, the serial number, etc. may be used as the identification number.
[0151] (5) In the above embodiment, the target book is found based on the spine of the book. However, in cases where the book is laid out flat, the target book may be found based on the front or back cover of the book.
[0152] (6) In the above embodiment, it is not known which shelf of a bookshelf the target book is located on, so images of each shelf are taken in order. However, if it is known which shelf the target book is located on, it is sufficient to take an image of only that shelf.
[0153] (7) In the above embodiment, images are taken for each shelf of the bookshelf. However, images of the first several shelves (for example, all shelves) may be taken, and the target book may be found from among them.
[0154] (8) In the above embodiment, the spine of the target book is identified using a trained estimation model based on the image of the spine. However, the spine of the target book may be identified by converting the characters on the spine into text and determining whether the converted text matches the characters in the captured image.
[0155] (9) In the above embodiment, the area mark M is displayed as a frame that surrounds the entire spine. However, the area mark M may be displayed in any form that allows the user to identify the area of the spine of the desired book. For example, an arrow may be displayed as the area mark at the top or bottom of the spine.
[0156] (10) In the above embodiment, the area mark M is displayed on the spine of the found book. In addition, clicking a details display button (for example, the area mark M can be used) on the smartphone SP may display a website page showing details of the found book (publisher, author, summary, etc.).
[0157] The server device S may record the URL of the page showing the details of each book in association with the ISBN code of each book, and may transmit this URL when transmitting the area mark M to the smartphone SP. The smartphone SP may then use the URL to create a link on the details display button.
[0158] (11) In the above embodiment, the ISBN is estimated from the spine image using a machine learning model in step S123. However, a table that associates spine images of many books with ISBNs may be prepared in advance, and the spine image of the target book may be matched with the spine image in the table through image processing to obtain the corresponding ISBN code.
[0159] (12) The above-described embodiments and their modifications may be implemented in combination with each other, or may be implemented in combination with other embodiments or modifications.
Claims
1. A search and guidance system comprising a search and guidance terminal device and an object extraction server device that is capable of communicating with the search and guidance terminal device, The search and guidance terminal device A camera and a display are provided, configured to have access to a three-dimensional model for identifying an area of interest and a model for a route; identifying an imaging position in the target area based on an image of the target area captured by a camera, with reference to the identifying three-dimensional model; generating a route instruction from the identified imaging position to a location where the desired object is placed based on the route model; Based on the generated route instructions, a guide to the location where the desired object is placed is displayed on a display; When it is determined that the robot has reached the location where the desired objects are placed, it transmits a multiple object display image of the multiple objects captured by the camera to the object extraction server device; displaying an area mark on a display superimposed on a multiple item display image of multiple items captured by a camera based on area mark information of the item display of the target item, which is the target item, received from the object extraction server device; The object extraction server device extracting an area of the object display of the target object from the captured image of the multiple object display received from the search guidance terminal device; In a search and guidance system, area mark information indicating an area of the item display of the extracted target item is transmitted to a search and guidance terminal device, The search and guidance system is characterized in that the object display area of the target object is determined based on image features of each of the object display captured images after separating the multiple object display captured image into individual object display captured images.
2. A search guidance terminal device that configures a search guidance system together with an object extraction server device. A camera and a display are provided, configured to have access to a three-dimensional model for identifying an area of interest and a model for a route; identifying an imaging position in the target area based on an image of the target area captured by a camera, with reference to the identifying three-dimensional model; generating a route instruction from the identified imaging position to a location where the desired object is placed based on the route model; Based on the generated route instructions, a guide to the location where the desired object is placed is displayed on a display; When it is determined that the robot has reached the location where the desired objects are placed, it transmits a multiple object display image of the multiple objects captured by the camera to the object extraction server device; A search guidance terminal device that displays area marks on a display by superimposing them on a multiple item display image of multiple items captured by a camera based on area mark information of an item display of a target item that is an object received from an object extraction server device, The search and guidance terminal device is characterized in that the object display area of the target object is determined based on image features of each object display captured image after separating the multiple object display captured image into each object display captured image.
3. A search guidance program for implementing a search guidance terminal device that configures a search guidance system together with an object extraction server device by a computer, the search guidance program comprising: Identifying an imaging position in the target area by referring to the three-dimensional identification model based on an image of the target area captured by the camera; generating a route instruction from the identified imaging position to a location where the desired object is placed based on the route model; Based on the generated route instructions, a guide to the location where the desired object is placed is displayed on a display; When it is determined that the robot has reached the location where the desired objects are placed, it transmits a multiple object display image of the multiple objects captured by the camera to the object extraction server device; A search and guidance program including instructions for displaying area marks on a display by superimposing the area marks on a multiple item display image of multiple items captured by a camera based on area mark information of the item display of the target item, which is the target, received from an object extraction server device, The search and guidance terminal device is characterized in that the object display area of the target object is determined based on image features of each object display captured image after separating the multiple object display captured image into each object display captured image.
4. In the system, device, or program according to any one of claims 1 to 3, the three-dimensional identification model is three-dimensional feature point data generated based on an image of a target area, The route model is data that can distinguish at least between passages and non-passages.
5. In the system, device, or program according to any one of claims 1 to 3, The guidance is as follows: acquiring an imaging position and an imaging direction in the identified target area in real time by referring to the three-dimensional identification model based on an image of the target area captured by a camera; determining how the route instructions generated based on the route model are displayed in two dimensions in the imaging direction at the imaging position; A device or program characterized in that the two-dimensional display of the route instructions is superimposed on an image captured by the camera.
6. In the system, device, or program according to any one of claims 1 to 3, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; A system, device, or program characterized in that if the desired object is not included in the multiple-item display captured image, a message is displayed indicating that the desired object is not included in the multiple-item display captured image, prompting the user to capture multiple objects in another group.
7. In the system, device, or program according to any one of claims 1 to 3, the target area is a bookshelf containing books in a library or bookstore; the object is a book, The search and guidance terminal device A system, device or program characterized in that a starting position of the route instruction is determined based on the installation position of a book search terminal device provided in the library or bookstore.
8. In the system, device, or program according to any one of claims 1 to 3, A system, device, or program characterized by using the multiple item display image as a marker and controlling the area mark so that it is displayed superimposed on the item display of the target item even if the current multiple item display image captured by the camera changes.
9. A search and guidance system comprising a search and guidance terminal device and a search and guidance server device that is capable of communicating with the search and guidance terminal device, The search and guidance terminal device A camera and a display are provided, transmitting an image of the target area captured by the camera to the search guidance server device; displaying on a display a guide to the location where the desired object is placed based on the route instruction received from the search guide server device; Obtain the identification information of the target item, transmitting a plurality of item display images of the plurality of items captured by the camera to an object extraction server device; Identifying a target item based on the target item identification information received from the search guidance server device, and displaying an area mark for the target item on the display by superimposing it on a multiple item display image of the multiple items captured by the camera; The search guide server device configured to have access to a three-dimensional model for identifying an area of interest and a model for a route; identifying an imaging position and an imaging direction in a target area by referring to the three-dimensional identification model based on the captured image received from the search guidance terminal device; generating a route instruction from the identified imaging position to a location where the desired object is placed based on the route model and transmitting the generated route instruction to the search guidance terminal device; A search guidance system characterized in that target item identification information for identifying a target item is generated based on a multiple item display captured image received from a search guidance terminal device, and transmitted to the search guidance terminal device, In a search and guidance system, the article display area of the target article is determined based on image features of each article display captured image by separating the multiple article display captured image into individual article display captured images, The search and guidance system is characterized in that the object display area of the target object is determined based on image features of each of the object display captured images after separating the multiple object display captured image into individual object display captured images.
10. A search guidance terminal device that is provided so as to be able to communicate with the search guidance server device, A camera and a display are provided, transmitting an image of the target area captured by the camera to the search guidance server device; based on the captured image received from the search guidance terminal device, referring to a three-dimensional identification model, identifying an imaging position and imaging direction in a target area, receiving route instructions from the search guidance server device generated based on the route model from the identified imaging position to a location where a desired object is placed, and displaying guidance to the location where the desired object is placed on a display based on the route instructions; Obtain the identification information of the target item, transmitting a plurality of item display images of the plurality of items captured by the camera to an object extraction server device; A search guidance terminal device that identifies a target item based on target item identification information received from a search guidance server device, and displays an area mark for the target item on a display by superimposing it on a multiple item display image of multiple items captured by a camera, In a search and guidance terminal device, the object display area of the target object is determined based on image features of each object display image by separating the multiple object display captured image into each object display captured image, The search and guidance terminal device is characterized in that the object display area of the target object is determined based on image features of each object display captured image after separating the multiple object display captured image into each object display captured image.
11. A search guide terminal program for realizing a search guide terminal device by a computer, the program comprising: transmitting an image of the target area captured by the camera to the search guidance server device; based on the captured image received from the search guidance terminal device, referring to a three-dimensional identification model, identifying an imaging position and imaging direction in a target area, receiving route instructions from the identified imaging position to a location where a desired object is placed based on a route model from a search guidance server device that has generated the route, and displaying guidance to the location where the desired object is placed on a display based on the route; Obtain the identification information of the target item, transmitting a plurality of item display images of the plurality of items captured by the camera to an object extraction server device; 1. A search guidance terminal program comprising: a command to identify a target object based on target object identification information received from a target object extraction server device; and to display an area mark for the target object on a display by superimposing the area mark on a multiple object display image of multiple objects captured by a camera, a search guide terminal program for determining an area of the object display of the target object by separating the multiple object display captured image into individual object display captured images and determining the area of the object display of the target object based on image features of each object display captured image; The search guide terminal program is characterized in that the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into each object display captured image.
12. A search guidance server device that is provided so as to be able to communicate with the search guidance terminal device, configured to provide access to a three-dimensional model of the target area; identifying an imaging position and an imaging direction in the target area by referring to a three-dimensional identification model based on the captured image of the target area received from the search guidance terminal device; generating, based on a route model, route instructions from the identified imaging position to the location where the desired object is placed, and transmitting the generated route instructions to the search guidance terminal device, so that the search guidance terminal device can display guidance to the location where the desired object is placed on a display; In the search guidance terminal device, in order to be able to display an area mark for the target item on the display by superimposing it on a multiple item display image of multiple items captured by a camera, a search guidance server device generates target item identification information for identifying the target item based on the multiple item display captured image received from the search guidance terminal device and transmits the information to the search guidance terminal device, In the search guide server device, the article display area of the target article is determined by separating the multiple article display captured image into individual article display captured images and based on image features of each article display captured image, The search guide server device is characterized in that the object display area of the target object is determined based on image features of each object display captured image after separating the multiple object display captured image into each object display captured image.
13. In the system, device, or program according to any one of claims 9 to 12, the three-dimensional identification model is three-dimensional feature point data generated based on an image of a target area, The route model is data that can distinguish at least between passages and non-passages.
14. In the system, device, or program according to any one of claims 9 to 12, The guidance in the search guidance terminal device is acquiring an imaging position and an imaging direction in the identified target area in real time by referring to the three-dimensional identification model based on an image of the target area captured by a camera; determining how the route instructions generated based on the route model are displayed in two dimensions in the imaging direction at the imaging position; A system, device, or program characterized in that the two-dimensional display of the route instructions is superimposed on an image captured by the camera.
15. In the system, device, or program according to any one of claims 9 to 12, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; A system, device, or program characterized in that if the desired object is not included in the multiple-item display captured image, a message is displayed indicating that the desired object is not included in the multiple-item display captured image, prompting the user to capture multiple objects in another group.
16. In the system, device, or program according to any one of claims 9 to 12, the target area is a bookshelf containing books in a library, the object is a book, The search guide server device A system, device or program characterized in that a starting position of the route instruction is determined based on the installation position of a book search terminal device provided in the library obtained from the search guidance terminal device.
17. In the system, device, or program according to any one of claims 9 to 12, The search and guidance terminal device A system, device, or program characterized by using the multiple item display image as a marker and controlling the area mark so that it is displayed superimposed on the item display of the target item even if the current multiple item display image captured by the camera changes.
18. In the system, device, or program according to any one of claims 1 to 3 and 9 to 12, A system, device, or program, characterized in that the target area is a bookshelf on which books are placed.
19. An object extraction system including an object extraction terminal device and an object extraction server device that is communicable with the object extraction terminal device, The object extraction terminal device A camera and a display are provided, Obtain the identification information of the target item, transmitting a plurality of item display images of the plurality of items captured by the camera to an object extraction server device; Identifying a target object based on the target object identification information received from the target object extraction server device, and displaying an area mark for the target object on the display by superimposing it on a multiple object display image of the multiple objects captured by the camera; The object extraction server device 1. An object extraction system, comprising: generating object identification information for identifying an object based on a multiple object display image received from an object extraction terminal device; and transmitting the information to the object extraction terminal device; In the object extraction system, the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into individual object display captured images, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, a message is displayed indicating that the desired object is not included in the multiple-item display captured image, prompting the user to capture multiple objects in another group.
20. An object extraction terminal device that is provided so as to be able to communicate with the object extraction server device, A camera and a display are provided, Obtain the identification information of the target item, transmitting a plurality of item display images of the plurality of items captured by the camera to an object extraction server device; an object extraction terminal device that identifies a target object based on target object identification information received from an object extraction server device that generates target object identification information for identifying a target object based on the multiple object display image, and displays an area mark for the target object on a display by superimposing it on the multiple object display image of the multiple objects captured by a camera; In an object extraction terminal device, the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into each of the object display captured images, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, the object extraction terminal device displays a message indicating that the desired object is not included in the multiple-item display captured image, prompting the user to capture multiple objects in another group.
21. An object extraction terminal program for realizing an object extraction terminal device by a computer, the program comprising: Obtain the identification information of the target item, transmitting a plurality of item display images of the plurality of items captured by the camera to an object extraction server device; an object extraction terminal program that identifies a target object based on target object identification information received from an object extraction server device that generates target object identification information for identifying a target object based on the multiple object display image, and displays an area mark for the target object on a display by superimposing it on the multiple object display image of the multiple objects captured by a camera; an object extraction terminal program for extracting an object, the object display area of the object being extracted being determined based on image features of each of the object display captured images by separating the multiple object display captured image into each of the object display captured images; At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, the object extraction terminal program displays a message indicating that the desired object is not included in the multiple-item display captured image, prompting the user to capture multiple objects in another group.
22. An object extraction server device that is provided so as to be able to communicate with an object extraction terminal device, receiving a captured image showing multiple items from the object terminal device; an object extraction server device that generates object identification information for identifying the object based on the multiple object display image received from the object terminal device and transmits the information to the object extraction terminal device so that the object extraction terminal device can identify the object and display an area mark for the object on a display by superimposing the area mark on a multiple object display image of the multiple objects captured by a camera; In the object extraction server device, the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into each of the object display captured images, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, the object extraction server device displays a message indicating that the desired object is not included in the multiple-item display captured image, prompting the user to capture multiple objects in another group.
23. In the system, device, or program according to any one of claims 19 to 22, the object extraction server device generates correspondence information including image position information and identification information of each object included in the multiple object display image as the object item identification information, and transmits the correspondence information to the object extraction terminal device; the object extraction terminal device specifies, based on the correspondence information received from the object extraction server device, image location information that matches the identification information of the target object as image location information of the target object; A system, device, or program that displays the area mark based on the image position information.
24. 24. The system, device or program of claim 23, The object extraction server device extracts an object display image of each object based on the multiple object display images, estimates image position information of each object, and estimates identification information of the corresponding object based on each extracted object display image.
25. In the system, device, or program according to any one of claims 19 to 22, the object extraction terminal device transmits identification information of the object to the object extraction server device; the object extraction server device acquires a target object display image based on the received identification information of the target object, identifies the target object display image from the plurality of object display images, estimates image position information thereof, and transmits the image position information to the object extraction terminal device as the target object identification information; The object extraction terminal device displays the area mark based on the received image position information.
26. In the system, device, or program according to any one of claims 19 to 22, The object extraction terminal device A system, device, or program characterized by using the multiple item display image as a marker and controlling the area mark so that it is displayed superimposed on the item display of the target item even if the current multiple item display image captured by the camera changes.
27. In the system, device, or program according to any one of claims 19 to 22, the article is a book, The system, device, or program is characterized in that the multiple item display image is a multiple spine image obtained by capturing images of spines of multiple books lined up.
28. An object extraction system including an object extraction terminal device and an object extraction server device that is communicable with the object extraction terminal device, The object extraction terminal device Equipped with a display, Obtain the identification information of the target item, Transmitting a multiple item display image of the multiple items to an object extraction server device; displaying a multiple object display image with an area mark for the target object on a display based on information including at least an area mark for the target object received from the object extraction server device; The object extraction server device 1. An object extraction system, comprising: transmitting, to an object extraction terminal device, information including at least a multiple object display image received from the object extraction terminal device and an area mark for the object generated based on identification information of the object; In the object extraction system, the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into individual object display captured images, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, a message is displayed indicating that the desired object is not included in the multiple-item display captured image, prompting the user to use an image of multiple objects from another group.
29. An object extraction terminal device that is provided so as to be able to communicate with the object extraction server device, Equipped with a display, Obtain the identification information of the target item, Transmitting a multiple item display image of the multiple items to an object extraction server device; an object extraction terminal device that displays a multiple object display image with area marks for the target objects on a display based on information including at least area marks for the target objects received from an object extraction server device, In an object extraction terminal device, the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into each of the object display captured images, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, the object extraction terminal device displays a message indicating that the desired object is not included in the multiple-item display captured image, prompting the user to use an image of multiple objects from another group.
30. An object extraction terminal program for realizing an object extraction terminal device by a computer, the program comprising: Obtain the identification information of the target item, Transmitting a multiple item display image of the multiple items to an object extraction server device; an object extraction terminal program for displaying a multiple object display image with area marks for the target objects on a display based on information including at least area marks for the target objects received from an object extraction server device, an object extraction terminal program for extracting an object, the object display area of the object being extracted being determined based on image features of each of the object display captured images by separating the multiple object display captured image into each of the object display captured images; At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; An object extraction terminal program characterized in that, if the desired object is not included in the multiple item display captured image, a message is displayed indicating that the desired object is not included in the multiple item display captured image, prompting the user to use an image of multiple objects from another group.
31. An object extraction server device that is provided so as to be able to communicate with an object extraction terminal device, receiving a multiple item display image and identification information of the target items from the target terminal device; an object extraction server device that transmits information including at least an area mark for the target object generated based on the multiple object display image and the identification information of the target object to an object extraction terminal device; In the object extraction server device, the object display area of the target object is determined based on image features of each of the object display captured images by separating the multiple object display captured image into each of the object display captured images, At the location where the desired object is placed, a plurality of groups are placed as groups each including a plurality of objects, and the desired object is included in at least one of the groups; If the desired object is not included in the multiple-item display captured image, the object extraction server device displays a message indicating that the desired object is not included in the multiple-item display captured image, prompting the user to use an image of multiple objects from another group.
32. In the system, device, or program of claims 28 to 31, The object extraction server device transmits an image in which an area mark for the object is added to an image showing multiple objects, as information including the area mark.
33. In the system, device, or program of claims 28 to 31, The object extraction server device transmits coordinates of the area mark as information including the area mark.
34. In the system, device, or program according to any one of claims 28 to 31, the article is a book, The system, device, or program is characterized in that the multiple item display image is a multiple spine image obtained by capturing images of spines of multiple books lined up.
Citation Information
Patent Citations
Article management system, non-contact distinguishing method, antenna unit and article management shelf
JP2003072919A
Picking assisting device and program
JP2015160696A
Picking system
JP2016188123A
Route guide device, route guide system, and program
JP2022055218A
Information processing apparatus, method for supporting book arrangement, and program for supporting book arrangement
JP2023072545A