3D model generating device and 3D model generation method

WO2026203677A1PCT designated stage Publication Date: 2026-10-01PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/000662
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-01-13
Publication Date
2026-10-01

Smart Images

  • Figure JP2026000662_01102026_PF_FP_ABST
    Figure JP2026000662_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This 3D model generating device executes shelf skeleton model generation processing (P2) for generating a 3D model of installed objects with objects present in the surrounding area excluded, storage information acquisition processing (P3) for acquiring arrangement information indicating the arrangement state of accommodated objects in an area of interest, and shelf attribute information acquisition processing (P5) for acquiring attribute information of the installed objects, wherein the arrangement information of the accommodated objects and the attribute information of the installed objects are stored in a memory unit in association with each other together with the 3D model of the installed objects. Furthermore, in accordance with an operation by a user, the 3D model generating device switches between a first mode for executing the shelf skeleton model generation processing, storage information acquisition processing and shelf attribute information acquisition processing, and a second mode for executing area model generation processing (P6) for generating a 3D model of the entire area of interest.
Need to check novelty before this filing date? Find Prior Art

Description

3D Model Generation Apparatus and 3D Model Generation Method

[0001] The present disclosure relates to a 3D model generation apparatus and a 3D model generation method for generating a 3D model that reproduces the state of installations in a target area.

[0002] By reviewing the layout in a warehouse, it is possible to improve the efficiency of operations in the warehouse. In this case, using a 3D model that reproduces the arrangement state of objects such as shelves arranged in the warehouse enables efficient simulation work for examining various modified layouts of the warehouse.

[0003] As a technique related to such 3D models targeting warehouses, conventionally, 3D information (point cloud data) of a target area is acquired by performing 3D sensing on the target area (such as a warehouse), objects arranged in the target area are detected from the 3D information of the target area, and a 3D model of the detected object is arranged in the 3D space of the target area so as to correspond to the actual arrangement state of the objects in the target area. Such a technique is known (see Patent Document 1).

[0004] Japanese Patent No. 7226534

[0005] According to the conventional technique, a 3D model of an object (such as a shelf) that is expected to be arranged in a target area (such as a warehouse) is created in advance and registered in a database, and the 3D model of the object detected from the point cloud data of the target area is acquired from the database and arranged, thereby generating a 3D model of the target area with high reproducibility.

[0006] On the other hand, when a large number of shelves are installed in a warehouse, or when a large number of stored items (such as product boxes) are placed on one shelf, there are cases where the state of shelves and stored items in the warehouse cannot be appropriately checked just by visually observing the 3D model of the warehouse. Therefore, there is a demand for a technique that allows a user to appropriately check the state of installations and stored items in a target area such as a warehouse.

[0007] Furthermore, 3D models (e.g., detailed mesh models) generated from point cloud data within a warehouse are modeled with the contents in their original positions, resulting in relatively large data sizes. This creates a problem where the 3D models of warehouse shelves cannot be smoothly manipulated during simulation. Therefore, there is a need for technology that generates lightweight 3D models to enable users to perform simulations smoothly.

[0008] Therefore, the main purpose of this disclosure is to provide a 3D model generation apparatus and a 3D model generation method that can generate lightweight 3D models that allow users to appropriately check the status of installed objects and contents in the target area, and that enable users to perform simulation work smoothly.

[0009] Furthermore, the primary objective of this disclosure is to provide a 3D model generation apparatus and a 3D model generation method that can generate lightweight 3D models that have sufficient reproducibility to enable users to perform simulation work appropriately, while also enabling users to perform simulation work smoothly.

[0010] The 3D model generation apparatus of this disclosure is a 3D model generation apparatus that uses a processor to perform a process of generating a 3D model that reproduces the state of an installation in a target area, wherein the processor performs a first process of generating a 3D model of the installation in a state in which surrounding objects have been removed, a second process of acquiring arrangement information representing the arrangement of contents in the target area, and a third process of acquiring attribute information of the installation, and stores the 3D model of the installation, the arrangement information of the contents, and the attribute information of the installation together in a storage unit in relation to each other.

[0011] Furthermore, the 3D model generation method of this disclosure is a method for causing a processor to execute a process to generate a 3D model that reproduces the state of an installation in a target area, and is configured to execute a first process to generate a 3D model of the installation in a state in which surrounding objects have been removed, a second process to acquire arrangement information representing the arrangement of contents in the target area, and a third process to acquire attribute information of the installation, and to store the 3D model of the installation, the arrangement information of the contents, and the attribute information of the installation together in a storage unit in relation to each other.

[0012] Furthermore, the 3D model generation apparatus of this disclosure is a 3D model generation apparatus that uses a processor to perform a process of generating a 3D model that reproduces the state of an installed object in a target area, wherein the processor uses an image recognition engine that performs image segmentation to detect the installed object from a captured image of the target area, extracts depth data of the installed object from depth data within the target area acquired in each capture based on the region information of the installed object on the captured image, generates point cloud data of the installed object based on the depth data of the installed object, and generates a 3D model of the installed object in a state in which objects present around the installed object have been excluded based on the point cloud data of the installed object.

[0013] Furthermore, the 3D model generation method of this disclosure is a method for causing a processor to execute a process to generate a 3D model that reproduces the state of an installed object in a target area, and is configured to detect an installed object from a captured image of the target area using an image recognition engine that performs image segmentation, extract depth data of the installed object from depth data within the target area acquired in each capture based on the region information of the installed object on the captured image, generate point cloud data of the installed object based on the depth data of the installed object, and generate a 3D model of the installed object in a state in which objects present around the installed object have been excluded based on the point cloud data of the installed object.

[0014] According to this disclosure, a 3D model of the installation is generated in a state where objects surrounding the installation have been removed. In addition, the user can check the placement information of the contents and the attribute information of the installation along with the 3D model of the installation. This allows the user to properly check the status of the installation and contents in the target area, and generates a lightweight 3D model that enables the user to perform simulation work smoothly.

[0015] Furthermore, according to this disclosure, an installation object is detected from a captured image (e.g., a color image), and based on the detection result, depth data of the installation object is extracted to obtain point cloud data of the installation object and generate a 3D model of the installation object. As a result, a 3D model can be generated in which objects surrounding the installation object are excluded. This makes it possible to generate a 3D model that has sufficient reproducibility for the user to perform simulation work appropriately, while also being lightweight to allow the user to perform simulation work smoothly.

[0016]

[0017] The first invention made to solve the aforementioned problems is a 3D model generation device that uses a processor to generate a 3D model that reproduces the state of an installation in a target area, wherein the processor performs a first process of generating a 3D model of the installation in a state where surrounding objects have been removed, a second process of acquiring arrangement information representing the arrangement of contents in the target area, and a third process of acquiring attribute information of the installation, and stores the 3D model of the installation, the arrangement information of the contents, and the attribute information of the installation together in a storage unit in relation to each other.

[0018] This generates a 3D model of the installation object with all surrounding objects removed. Furthermore, the user can view the placement information of the contents and the attribute information of the installation object along with the 3D model. This allows the user to properly check the status of the installation object and its contents in the target area, and generates a lightweight 3D model that facilitates the user's simulation work.

[0019] Furthermore, the second invention is configured such that the processor switches between a first mode, which executes the first, second, and third processes, and a second mode, which executes a fourth process, which generates a 3D model of the entire target area, in response to user operation.

[0020] According to this, users can specify a mode as needed to view the 3D model of the installation, the placement information of the contents, and the attribute information of the installation generated in the first mode, or view the 3D model of the entire target area generated in the second mode.

[0021] Furthermore, the third invention is configured such that, in the first processing, the processor uses an image recognition engine that performs image segmentation to detect the installed object from the captured image of the target area, extracts depth data of the installed object from the depth data of the target area acquired in each capture based on the region information of the installed object on the captured image, generates point cloud data of the installed object based on the depth data of the installed object, and generates a 3D model of the installed object in a state in which objects present around the installed object have been excluded based on the point cloud data of the installed object.

[0022] According to this method, the installed object is detected from a captured image (e.g., a color image), and based on the detection result, depth data of the installed object is extracted to obtain point cloud data of the installed object and generate a 3D model of the installed object. As a result, a 3D model can be generated in a state where objects surrounding the installed object have been excluded. This makes it possible to generate a 3D model that has sufficient reproducibility for the user to perform simulation work appropriately, and is lightweight so that the user can perform simulation work smoothly. The depth data of the target area may be generated by capturing (3D sensing) the target area, or it may be generated from existing point cloud data or a 3D model (mesh model) of the target area.

[0023] Furthermore, the fourth invention is configured such that the processor generates a CAD model of the installation object to which feature information relating to the appearance of the installation object extracted from the captured image is applied, generates a plurality of training images with different 2D imaging processing conditions from the CAD model of the installation object, and constructs the image recognition engine by machine learning using the training images.

[0024] According to this, machine learning (e.g., deep learning) is performed using training images generated from a CAD model of the installed object to which the object's characteristic information (e.g., texture) has been applied, making it easy to build a highly accurate image recognition engine.

[0025] Furthermore, the fifth invention is configured such that the processor is connected to a language modeling device that analyzes instruction sentences written in natural language by the user and obtains processing conditions for 2D imaging, and the processor generates the training images based on the processing conditions for 2D imaging obtained from the language modeling device.

[0026] According to this, the processing conditions for generating training images can be easily changed by the user modifying the instruction text.

[0027] Furthermore, the sixth invention is configured such that, in the second processing, the processor acquires, as arrangement information, identification information and location information of the contained object, and identification information and location information of the installation in which the contained object is contained.

[0028] According to this, the user can be presented with information regarding the placement of the contents, including identification and location information of the contents, and identification and location information of the installation in which the contents are housed. This allows the user to easily confirm the placement status of the contents within the target area.

[0029] Furthermore, the seventh invention is configured such that the processor detects labels that identify installed objects and contained objects from captured images of the target area, obtains identification information of the installed objects and contained objects included in the labels, extracts label depth data from depth data of the target area obtained in each capture based on the region information of the labels on the captured images, generates point cloud data of the labels based on the label depth data, and obtains position information of the labels in 3D space based on the label point cloud data.

[0030] According to this, identification information and location information of the contents, as well as identification information and location information of the installed objects, can be appropriately acquired.

[0031] Furthermore, the eighth invention is configured such that, in the third process, the processor acquires information regarding the configuration of the installed object as attribute information of the installed object.

[0032] According to this, users can easily check information about the configuration (specifications) of the installed items, such as the number of shelves and their placement height.

[0033] Furthermore, the ninth invention is a 3D model generation method that causes a processor to execute a process to generate a 3D model that reproduces the state of an installation in a target area, comprising: a first process to generate a 3D model of the installation in a state where surrounding objects have been removed; a second process to acquire arrangement information representing the arrangement of contents in the target area; and a third process to acquire attribute information of the installation, and the 3D model of the installation, along with the arrangement information of the contents and the attribute information of the installation, are stored in a storage unit in association with each other.

[0034] According to this, similar to the first invention, it is possible to generate a lightweight 3D model that allows the user to appropriately check the status of installed objects and contents in the target area, and to enable the user to perform simulation work smoothly.

[0035] The tenth invention, made to solve the aforementioned problems, is a 3D model generation device that uses a processor to perform a process of generating a 3D model that reproduces the state of an installed object in a target area, wherein the processor uses an image recognition engine that performs image segmentation to detect the installed object from a captured image of the target area, extracts depth data of the installed object from depth data within the target area acquired in each capture based on the region information of the installed object on the captured image, generates point cloud data of the installed object based on the depth data of the installed object, and generates a 3D model of the installed object in a state in which objects present around the installed object have been excluded based on the point cloud data of the installed object.

[0036] According to this method, the installed object is detected from a captured image (e.g., a color image), and based on the detection result, depth data of the installed object is extracted to obtain point cloud data of the installed object and generate a 3D model of the installed object. As a result, a 3D model can be generated in a state where objects surrounding the installed object have been excluded. This makes it possible to generate a 3D model that has sufficient reproducibility for the user to perform simulation work appropriately, and is lightweight so that the user can perform simulation work smoothly. The depth data of the target area may be generated by capturing (3D sensing) the target area, or it may be generated from existing point cloud data or a 3D model (mesh model) of the target area.

[0037] Furthermore, the eleventh invention is configured such that the processor generates a CAD model of the installation object to which feature information relating to the appearance of the installation object extracted from the captured image is applied, generates a plurality of training images with different 2D imaging processing conditions from the CAD model of the installation object, and constructs the image recognition engine by machine learning using the training images.

[0038] According to this approach, machine learning (e.g., deep learning) is performed using training images generated from a CAD model of the installed object to which feature information (e.g., texture information) of the installed object has been applied, making it easy to build a highly accurate image recognition engine.

[0039] Furthermore, the twelfth invention is configured to be connected to a language modeling device that analyzes instruction sentences written in natural language by the user and obtains processing conditions for 2D imaging, and the processor generates the training images based on the processing conditions for 2D imaging obtained from the language modeling device.

[0040] According to this, the processing conditions for generating training images can be easily changed by the user modifying the instruction text.

[0041] Furthermore, in a thirteenth invention, the processor detects a label for identifying an installation and a stored item from the captured image, acquires identification information of the installation and the stored item included in the label, extracts depth data of the label from depth data of the target area acquired in each imaging operation based on area information of the label on the captured image, generates point cloud data of the label based on the depth data of the label, and acquires position information of the label in a 3D space based on the point cloud data of the label.

[0042] According to this, as arrangement information, identification information and position information of a stored item, and identification information and position information of an installation accommodating the stored item can be presented to a user. This allows the user to easily check the arrangement status of the stored item in the target area.

[0043] Furthermore, in a fourteenth invention, the processor is configured to superimpose and arrange an image representing the identification information of the installation and the stored item at a corresponding position on the 3D model of the installation based on the position information of the label.

[0044] According to this, the user can easily check the arrangement status of stored items on the 3D model of the installation.

[0045] Furthermore, a fifteenth invention is a 3D model generation method that causes a processor to execute processing for generating a 3D model reproducing the status of an installation in a target area, wherein an image recognition engine that performs image segmentation is used to detect the installation from a captured image of the target area, depth data of the installation is extracted from depth data in the target area acquired in each imaging operation based on area information of the installation on the captured image, point cloud data of the installation is generated based on the depth data of the installation, and the 3D model of the installation in a state where objects existing around the installation are excluded is generated based on the point cloud data of the installation.

[0046] According to this, similar to the tenth invention, it is possible to generate a lightweight 3D model that has sufficient reproducibility for a user to appropriately perform simulation work and is lightweight for the user to smoothly perform simulation work.

[0047] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0048] FIG. 1 is an overall configuration diagram of a 3D model generation system according to the present embodiment.

[0049] The present system generates a 3D model that reproduces the state of installations in a target area. In the present embodiment, for example, a 3D model that reproduces the state of shelves (installations) in a warehouse (target area) is generated. In the present embodiment, an example where the target area is a warehouse will be mainly described, but the target area is not limited to a warehouse, and may be other facilities such as a factory. Further, in the present embodiment, an example where products are stored on shelves in a warehouse of a retail store will be described, but items stored in the warehouse are not limited to products.

[0050] The present system includes an imaging terminal 1 (imaging device), a model generation server 2 (3D model generation device), and a user terminal 3 (user device). The model generation server 2 is connected to the imaging terminal 1 and the user terminal 3 via a network.

[0051] The imaging terminal 1 includes a visible camera 11 and a depth sensor 12 as sensors for three-dimensional measurement of a target area. The visible camera 11 (color camera) outputs color captured images. The depth sensor 12, for example, measures the depth (distance) to a subject by imaging with left and right infrared cameras (not shown) and outputs depth data. A user holds the imaging terminal 1, moves within the target area, and performs an imaging operation that causes the imaging terminal 1 to perform imaging (3D sensing) of the target area. In the imaging terminal 1, processing based on the SLAM (Simultaneous Localization and Mapping) method is performed. The imaging terminal 1 may be provided with an IMU (Inertial Measurement Unit) to improve the accuracy of position and orientation estimation information. The imaging terminal 1 may be a terminal having a LiDAR sensor mounted with a visible camera. Further, the imaging terminal 1 may be a device in which a sensor is mounted on a remotely or automatically operated cart, drone, or the like.

[0052] The shooting terminal 1 transmits (uploads) the shooting data (SLAM results) to the model generation server 2. The shooting data includes color images, depth data, and position and orientation estimation information acquired during each shooting session within the warehouse (target area).

[0053] The model generation server 2 generates a shelf skeleton model, an area model, storage information, and shelf attribute information based on the image data acquired from the image capture terminal 1. The shelf skeleton model is a mesh model of the shelf's skeleton state. The area model is a mesh model of the entire captured area. The storage information (arrangement information) represents the storage status (arrangement status) of goods (contained items) in the warehouse (target area). Specifically, the storage information includes identification information and location information of goods stored in the warehouse, and identification information and location information of the shelves in which the goods are stored. The shelf attribute information is information about the shelf configuration (number of shelves, shelf height, etc.). The model generation server 2 may be configured as a cloud computer (virtual machine).

[0054] The model generation server 2 sends (downloads) the shelf skeleton model, area model, storage information, and shelf attribute information to the user terminal 3 as processing results.

[0055] User terminal 3 is used by the user to perform operations that instruct the model generation server 2 to carry out processing, and also to view information transmitted from the model generation server 2. Specifically, the user performs operations that instruct the model generation server 2 to generate shelf skeleton models, area models, storage information, and shelf attribute information, and as a result of the processing, the shelf skeleton models, area models, storage information, and shelf attribute information are displayed on the screen, and the user can view this information.

[0056] Next, we will explain the skeletalization process performed by the model generation server 2. Figures 2A and 2B are explanatory diagrams showing the state of the 3D model after skeletalization.

[0057] In this embodiment, the shelf (installed object), which is the main object to be modeled, is subjected to skeletonization (abstraction) by removing unnecessary objects that exist around it.

[0058] Here, the example shown in Figure 2A is the case where no skeletonization is performed. In a warehouse in operation, goods remain on the shelves, so in a 3D model (mesh model) generated by taking images (3D sensing) of a warehouse in operation, the goods on the shelves and the floor on which the shelves are installed will be visible.

[0059] On the other hand, the example shown in Figure 2B is the case when the model is skeletonized. In this case, objects surrounding the shelves are excluded from the 3D model. Specifically, the products placed on the shelves and the floor on which the shelves are installed are excluded from the 3D model. In addition, in the skeletonized 3D model, parts that are hidden by products are not reproduced, and only the parts that are not hidden by products, namely the support columns, side panels, and the front edges of the shelves, are reproduced. In this way, the skeletonized 3D model does not perfectly reproduce the layout of the shelves in the warehouse, but in simulation work it is lightweight, easy to operate, and reproduces with sufficient accuracy to check the configuration and installation of the shelves.

[0060] Next, we will explain the process of generating training images for building the image recognition engine, which is performed on the model generation server 2. Figure 3 is an explanatory diagram showing an overview of the process performed on the model generation server 2.

[0061] In the model generation server 2, an image recognition engine is constructed that performs image segmentation to detect the area of ​​the shelf (object) from a color captured image. At this time, texture information of the shelf (feature information regarding the appearance of the shelf) is extracted from the color captured image shown in Figure 3(B). Next, the extracted texture information is applied to the material model shown in Figure 3(A) to generate the training model shown in Figure 3(C). Next, the training model is subjected to a 2D image conversion process to generate the training image shown in Figure 3(D). Finally, an image recognition engine is constructed using machine learning (deep learning) with the training image.

[0062] The material model (CAD model used as material) shown in Figure 3(A) is a model created in advance using CAD. The training model (CAD model for training) shown in Figure 3(C) is created by applying the shelf texture (e.g., color) extracted from a color photograph to the material model shown in Figure 3(A).

[0063] Image segmentation using an image recognition engine primarily recognizes shelves (objects) based on texture (e.g., color), so the training model only needs to have the same texture as the shelves being recognized. The training model and its source material model do not need to have the same shape as the shelves actually installed in the target warehouse.

[0064] Furthermore, if the filming is conducted on a warehouse that is currently in operation, the goods (storage items) will be placed on the shelves inside the warehouse. Therefore, it would be beneficial to create the training model with the goods placed on the shelves in that state.

[0065] The training images shown in Figure 3(D) are generated by performing a 2D image conversion process on the training model. By changing the 2D image conversion processing conditions in various ways, a large number of diverse training images (for example, 90 images) can be generated from a single training model.

[0066] The training images are created in a format suitable for building an image recognition engine that performs image segmentation, particularly semantic segmentation (region classification). Specifically, training images are created by labeling pixels on a pixel-by-pixel basis in 2D images generated by creating 2D images of the training model. In the example shown in Figure 3(D), pixels within the region of the shelf (installed object) are labeled as shelves. Similarly, pixels within the region of the product (contained item) are labeled as products (contained item).

[0067] In this embodiment, the model generation server 2 performs the process of generating training images (images of the training shelf) from the training model (CAD model of the training shelf). However, the user may manually create the training images one by one by operating an application installed on the user terminal 3. In this case, the user can generate a variety of training images by changing the conditions for generating the training images.

[0068] Furthermore, in this embodiment, the model generation server 2 performs the process of applying textures to a material model (a pre-existing CAD model of a shelf) to generate a training model. However, the training model may also be created by the user operating an application installed on the user terminal 3.

[0069] Next, we will describe the general configuration of the model generation server 2. Figure 4 is a block diagram showing the general configuration of the model generation server 2.

[0070] The model generation server 2 comprises a communication unit 21, a storage unit 22, and a processor 23.

[0071] The communication unit 21 communicates with the shooting terminal 1 and the user terminal 3 via the network.

[0072] The memory unit 22 stores programs and other data executed by the processor 23. The memory unit 22 also stores captured data uploaded from the capture terminal 1.

[0073] The processor 23 performs various processes by executing programs stored in the memory unit 22. In this embodiment, the processor 23 performs image recognition engine construction process P1, shelf skeleton model generation process P2, storage information acquisition process P3, shelf plan view creation process P4, shelf attribute information acquisition process P5, area model generation process P6, and display process P7, etc.

[0074] Next, we will explain the processing performed by the processor 23 of the model generation server 2. Figure 5 is a block diagram illustrating the overview of the processing performed by the processor 23 of the model generation server 2.

[0075] In the image recognition engine construction process P1, the processor 23 constructs an image recognition engine for detecting the area of ​​a shelf (object) from the color captured image contained in the captured data acquired from the shooting terminal 1.

[0076] In the shelf skeleton model generation process P2 (first process), the processor 23 generates a shelf skeleton model (an abstract model of the shelf), that is, a mesh model in a skeleton state from which unnecessary objects surrounding the shelf have been removed, based on the captured data (color captured image, depth data, and position and orientation estimation information) acquired from the shooting terminal 1. In the shelf skeleton model generation process P2, the image recognition engine built in the image recognition engine construction process P1 is used to detect the shelf region from the color captured image.

[0077] In the storage information acquisition process P3 (second process), the processor 23 generates storage information based on the image data (color image, depth data, and position / orientation estimation information) acquired from the shooting terminal 1. The storage information (arrangement information) includes identification information and location information of the goods (containers) stored in the warehouse, and identification information and location information of the shelves (installations) in which the goods are stored.

[0078] In the shelf plan drawing creation process P4, the processor 23 detects the positions of the vertical members of the shelf (posts and side panels) and the overall position of the shelf based on the point cloud data of the shelf acquired in the shelf skeleton model generation process P2. Based on these detection results, it creates a plan drawing (2D drawing) of the shelf.

[0079] In the shelf attribute information acquisition process P5 (third process), the processor 23 acquires attribute information of the shelf (installed object). In this embodiment, the shelf attribute information acquired includes information about the shelf's configuration (specifications), such as the number of shelves and the height at which each shelf is positioned.

[0080] In area model generation process P6 (fourth process), the processor 23 generates an area model, that is, a mesh model of the entire area, based on the captured data (color captured image, depth data, and position and orientation estimation information) acquired from the shooting terminal 1. The area model includes not only shelves (installed objects) but also structures such as the warehouse floor, walls, and columns.

[0081] In display processing P7, the processor 23 generates a predetermined screen and displays it on the user terminal 3. Specifically, a screen including the shelf skeleton model generated in shelf skeleton model generation processing P2 is displayed. In addition, a screen is displayed in which the storage information of the products acquired in storage information acquisition processing P3 is superimposed on the shelf skeleton model. Furthermore, a screen is displayed in which the shelf attribute information acquired in shelf attribute information acquisition processing P5 is superimposed on the shelf skeleton model. By viewing these screens, the user can confirm that the shelf skeleton model, product storage information, and shelf attribute information have been generated appropriately.

[0082] In this embodiment, the user can switch between shelf skeleton mode (first mode) and normal mode (second mode). In shelf skeleton mode, shelf skeleton model generation process P2 (first process), storage information acquisition process P3 (second process), and shelf attribute information acquisition process P5 (third process) are executed to generate a shelf skeleton model, storage information, and shelf attribute information. On the other hand, in normal mode, area model generation process P6 (fourth process) is executed to generate an area model.

[0083] Next, we will explain the image recognition engine construction process P1 performed by the processor 23 of the model generation server 2. Figure 6 is a block diagram showing the contents of the image recognition engine construction process P1.

[0084] In the image recognition engine construction process P1, the processor 23 constructs an image recognition engine for detecting the area of ​​a shelf (object) from a color captured image using machine learning such as deep learning. The image recognition engine construction process P1 includes a texture extraction process P11, a training model generation process P12, a training image generation process P13, and a training process P14.

[0085] In the texture extraction process P11, the processor 23 extracts texture information (characteristic information regarding the appearance of the shelf) of the main components of the shelf (e.g., support columns, side panels, etc.) from the color captured image. The user can specify the range from which to extract the texture on the color captured image. In this embodiment, the range from which to extract the texture is set manually, but the range from which to extract the texture may be set automatically.

[0086] In the training model generation process P12, the processor 23 generates a training model (a CAD model for training) by applying the texture information extracted in the texture extraction process P11 to the material model (the CAD model that will serve as the material).

[0087] In the training image generation process P13, the processor 23 generates training images by performing a 2D image conversion process on the training model. At this time, by changing various conditions for the 2D image conversion process, a large number of diverse training images can be generated from a single training model. The conditions for the 2D image conversion process include, for example, the position of the viewpoint in 3D space, the angle of the line of sight, and the brightness. The training images are created in a format suitable for building an image recognition engine that performs image segmentation. Specifically, labels representing shelves (objects) are assigned to the training images on a pixel-by-pixel basis.

[0088] In the learning process P14, the processor 23 constructs an image recognition engine using machine learning (such as deep learning) with the training images generated in the training image generation process P13.

[0089] Next, we will explain the shelf skeleton model generation process P2 performed by the processor 23 of the model generation server 2. Figure 7 is a block diagram showing the contents of the shelf skeleton model generation process P2.

[0090] In the shelf skeleton model generation process P2, the processor 23 generates a shelf skeleton model (an abstraction model of the shelf), that is, a mesh model in a skeleton state from which unnecessary objects present around the shelf have been removed. The shelf skeleton model generation process P2 includes image recognition processing P21, shelf depth data extraction processing P22, shelf point cloud generation processing P23, and meshing processing P24.

[0091] In the image recognition process P21, the processor 23 uses the image recognition engine constructed in the image recognition engine construction process P1 to detect shelves (installed objects) from each color image captured in the captured data and obtain area information of the shelves on the color image. At this time, image segmentation is performed on the color image, and it is determined whether each pixel of the color image is a shelf or not.

[0092] In the shelf depth data extraction process P22, the processor 23 extracts shelf depth data from the depth data of each capture included in the captured data, based on the shelf area information acquired in the image recognition process P21. At this time, depth data corresponding to the shelf area on the color captured image is extracted from the depth data detected simultaneously with the capture of that color captured image.

[0093] In each shooting session within the warehouse, the visible light camera 11 photographs the target object (shelf in the warehouse) from various viewpoints, and simultaneously, the depth sensor 12 detects the target object from the same direction as the visible light camera 11's shooting. Therefore, the area of ​​the shelf (installed object) in each color image captured by the visible light camera 11 corresponds to the area of ​​the shelf in the depth data acquired by the depth sensor 12.

[0094] In the shelf point cloud generation process P23, the processor 23 generates shelf point cloud data based on the shelf depth data extracted in the shelf depth data extraction process P22 and the position and orientation estimation information. At this time, the shelf point cloud data is generated by projecting the shelf depth data for each step into 3D space based on the position and orientation estimation information for each step.

[0095] In the meshing process P24, the processor 23 meshes the point cloud data of the shelves to generate a shelf skeleton model (a mesh model of the shelves in a skeleton state).

[0096] In this embodiment, the mesh model (shelf skeleton model) of the shelf is generated using the depth data included in the captured data. However, if a mesh model of the entire area has been previously generated, the mesh model of the shelf may be generated using that mesh model of the entire area. In this case, the depth data of the entire area can be generated from the mesh model of the entire area. Alternatively, if point cloud data of the entire area has been previously generated, the mesh model of the shelf may be generated using that point cloud data of the entire area. In this case, the point cloud data of the shelf can be extracted from the point cloud data of the entire area.

[0097] Next, we will explain the storage information acquisition process P3 performed by the processor 23 of the model generation server 2. Figure 8 is a block diagram showing the contents of the storage information acquisition process P3.

[0098] In the storage information acquisition process P3, the processor 23 generates storage information based on the captured data (color captured image, depth data, and position and orientation estimation information) acquired from the shooting terminal 1. This information includes identification information and location information for products stored in the warehouse, and identification information and location information for the shelves where the products are stored. The storage information acquisition process P3 includes label detection processing P31, character recognition processing P32, label depth data extraction processing P33, label point cloud generation processing P34, and label position identification processing P35.

[0099] In the label detection process P31, the processor 23 detects labels (product labels and shelf labels) that identify products and shelves from the color captured image and obtains region information of the labels on the color captured image. In this embodiment, product labels placed near products placed on the shelf boards of the shelf are detected. In addition, shelf labels placed on the side panels of the shelf are detected.

[0100] In character recognition processing P32, the processor 23 extracts a label image from the color image based on the label area information acquired in label detection processing P31, and recognizes the characters written on the label (product label and shelf label) based on that label image, and acquires the character information of the label (product number and shelf number) as identification information for the shelf (installed item) and product (contained item).

[0101] In the label depth data extraction process P33, the processor 23 extracts the label depth data from the depth data of each capture included in the image data, based on the label region information obtained in the label detection process P31.

[0102] In the label point cloud generation process P34, the processor 23 generates label point cloud data based on the label depth data extracted in the label depth data extraction process P33.

[0103] In the label position identification process P35, the processor 23 identifies the 3D position of the label based on the label point cloud data generated in the label point cloud generation process P34, and obtains the 3D position information of the label.

[0104] Alternatively, depth data may be calculated using the principle of triangulation, based on correspondence information obtained by associating character features between color images captured in each session, and position and orientation estimation information.

[0105] Next, we will explain the shelf plan creation process P4 performed by the processor 23 of the model generation server 2. Figure 9 is a block diagram showing the contents of the shelf plan creation process P4. Figures 10A, 10B, and 10C are explanatory diagrams showing an overview of the shelf plan creation process P4.

[0106] In the shelf plan creation process P4, the processor 23 detects the positions of the vertical members (posts and side panels) of the shelf and the overall position of the shelf, based on the point cloud data of the shelf generated in the shelf point cloud generation process P23 of the shelf skeleton model generation process P2. Based on these detection results, it creates a plan view (2D drawing) of the shelf. The shelf plan creation process P4 includes point cloud projection processing P41, vertical member position acquisition processing P42, overall shelf position acquisition processing P43, and drawing processing P44.

[0107] In the point cloud projection process P41, the processor 23 projects the point cloud data of the shelf onto a horizontal plane (XY plane). At this time, the Z coordinate value of each point included in the point cloud data is replaced with 0, so that each point is projected in the vertical direction (Z axis direction). The horizontal plane is divided into multiple cells, and the points included in each cell are counted to determine the point cloud density (number of points in each cell). The cells are divided into a grid of identical rectangles (see Figures 10A, 10B, and 10C).

[0108] In the vertical member position acquisition process P42, the processor 23 acquires the planar position of the vertical members (columns and side panels) of the shelf based on the point cloud density (number of points per cell) acquired in the point cloud projection process P41. The vertical member position acquisition process P42 includes a first point cloud density determination process P45 and a vertical member region setting process P46. When the point cloud of the shelf is projected perpendicularly to the horizontal plane, the number of projected points increases in regions where vertical members extending vertically, such as the columns and side panels of the shelf, exist, making it possible to detect the position of the vertical members (columns and side panels).

[0109] Furthermore, statistical information such as the distribution of points in each cell in the vertical direction may be used to detect the position of vertical members (support columns and side plates).

[0110] In the first point cloud density determination process P45, the processor 23 determines cells containing vertical members (support columns and side plates) by comparing the point cloud density (number of points per cell) with a predetermined first threshold (see Figure 10A). Specifically, cells with a point cloud count equal to or greater than the first threshold are determined to contain vertical members.

[0111] In the vertical member area setting process P46, the processor 23 sets a rectangular frame (bounding box) surrounding the cell where the vertical members (posts and side panels) of the shelf are located as the area of ​​the vertical members (posts and side panels) (see Figure 10A).

[0112] In the shelf-wide position acquisition process P43, the processor 23 acquires the planar position of the entire shelf based on the point cloud density (number of points per cell) acquired in the point cloud projection process P41. The shelf-wide position acquisition process P43 includes a second point cloud density determination process P47 and a shelf-wide area setting process P48. When the point cloud of the shelf is projected perpendicular to the horizontal plane, the area where the entire shelf exists will have more projected points than its surroundings, thus allowing the position of the entire shelf to be detected.

[0113] In the second point cloud density determination process P47, the processor 23 determines which cells contain an entire shelf by comparing the point cloud density (number of points per cell) with a predetermined second threshold (see Figure 10B). Specifically, cells where the number of points is equal to or greater than the second threshold are determined to contain an entire shelf.

[0114] In the shelf-wide area setting process P48, the processor 23 sets a rectangular bounding box surrounding the cell containing the entire shelf as the area of ​​the entire shelf (see Figure 10B).

[0115] In the drawing process P44, the processor 23 draws a plan view of the shelf based on the regions of the vertical members (columns and side panels) acquired in the vertical member position acquisition process P42 and the region of the entire shelf acquired in the entire shelf position acquisition process P43 (see Figure 10C). Furthermore, in the drawing process P44, the processor 23 projects the 3D mesh model of the shelf (shelf skeleton model) onto the horizontal plane (XY plane) to generate a 2D mesh model of the shelf, and superimposes rectangular frames representing the vertical members (columns and side panels) of the shelf and rectangular frames representing the entire shelf onto the 2D mesh model of the shelf.

[0116] Furthermore, character images representing the shelf and product identification information (product number and shelf number) obtained in the storage information acquisition process P3 may be superimposed and drawn on the shelf plan view.

[0117] Next, we will explain the shelf attribute information acquisition process P5 performed by the processor 23 of the model generation server 2. Figure 11 is a block diagram showing the contents of the shelf attribute information acquisition process P5. Figures 12A and 12B are explanatory diagrams showing an overview of the shelf attribute information acquisition process P5.

[0118] In the shelf attribute information acquisition process P5, the processor 23 acquires information about the shelf's configuration (specifications) as attribute information of the shelf (installed object), specifically the number of shelves and the placement height of each shelf. The shelf attribute information acquisition process P5 includes the support column determination process P51 and the shelf detection process P52.

[0119] In the support column determination process P51, the processor 23 determines whether a column is a support column based on whether the area of ​​the vertical member (support column and side plate) acquired in the vertical member position acquisition process P42 (see Figure 9) of the shelf plan drawing creation process P4, that is, the rectangular frame (bounding box) surrounding the cell where the vertical member of the shelf exists, is close to a square. If the area of ​​the vertical member is close to a square, that area is determined to be a support column, and the position information of the support column is acquired based on the position of that area. If the area of ​​the vertical member is not close to a square, that area is determined to be a side plate.

[0120] In the shelf detection process P52, the processor 23 detects shelves based on the distribution of the point cloud data in the height direction and obtains the number of shelves and the placement height of each shelf as shelf attribute information. In this embodiment, shelves on both the left and right sides of the support column are detected, and shelf attribute information (number of shelves and placement height) is obtained relative to the support column.

[0121] Specifically, as shown in Figure 12A, first, a cylindrical measurement target area of ​​a predetermined diameter is set around the support column. Next, point cloud data within the measurement target area is extracted from the shelf's point cloud data. Then, for each point included in the point cloud within the measurement target area, the distance from the central axis (Z-axis) of the support column is calculated.

[0122] Next, the height direction (Z-axis direction) is divided into multiple sections, and the maximum point cloud distance is calculated for each section. The maximum point cloud distance is the maximum distance from the central axis (Z-axis) of the support column for each point in the point cloud included in each section. Figure 12B shows a histogram with the maximum point cloud distance as the frequency. Here, in the point cloud data for the shelf, points are distributed even in areas far from the support column at the height position (Z-coordinate value) where the shelf board exists. Therefore, at the height position where the shelf board exists, the maximum point cloud distance becomes large, and a peak where the maximum point cloud distance is larger than the surrounding area appears at the position of the shelf board.

[0123] Therefore, the maximum point cloud distance at the peak position is compared with a predetermined threshold, and the peak position where the maximum point cloud distance is greater than or equal to the threshold is determined to be the shelf position. The number of peak positions where the maximum point cloud distance is greater than or equal to the threshold is then set as the number of shelves. In addition, the Z coordinate values ​​of the peak positions where the maximum point cloud distance is greater than or equal to the threshold are set as the shelf height.

[0124] In the example shown in Figure 12B, four peaks of maximum distance are detected in the point cloud on the left side of the support column, and if no products are stored on the top shelf (top shelf), the number of shelves will be set to three. On the other hand, three peaks of maximum distance are detected in the point cloud on the right side of the support column, so the number of shelves will be set to two.

[0125] Furthermore, the shelf and product identification information (product number and shelf number) acquired in the storage information acquisition process P3 may be associated with the shelf attribute information (number of shelves and height) based on the support column near the label (shelf label and product label) from which it originates, and stored in the storage unit 22.

[0126] Furthermore, statistical information such as variance may be used instead of the maximum point cloud distance to detect the number of shelves.

[0127] Next, we will explain the area model generation process P6 performed by the processor 23 of the model generation server 2. Figure 13 is a block diagram showing the contents of the area model generation process P6.

[0128] In the area model generation process P6, the processor 23 generates an area model, that is, a mesh model of the entire area. The area model generation process P6 includes the area point cloud generation process P61 and the meshing process P62.

[0129] In the area point cloud generation process P61, the processor 23 generates point cloud data for the entire area based on the color images, depth data, and position / orientation estimation information included in the captured data. At this time, the point cloud data for the entire area is generated by projecting the depth data for the entire area into 3D space based on the position / orientation estimation information.

[0130] In the meshing process P62, the processor 23 meshes the point cloud data of the entire area to generate a mesh model (area model) of the entire area.

[0131] In this embodiment, a mesh model (area model) of the entire area is generated by meshing the point cloud data of the entire area. However, a mesh model of the entire area excluding shelves may also be generated by excluding the point cloud data of shelves from the point cloud data of the entire area and then meshing it.

[0132] Next, we will explain the storage information and shelf attribute information generated by the model generation server 2. Figures 14A and 14B are explanatory diagrams showing the contents of the product location information and shelf attribute information.

[0133] In the storage information acquisition process P3 (see Figure 8), storage information (location information) of the goods (contained items) is acquired. The storage information of the goods includes the identification information and location information of the goods stored in the warehouse, and the identification information and location information of the shelves in which the goods are stored. The storage information is stored in a tabular file.

[0134] Specifically, as shown in Figure 14A, the product storage information includes product name, product label location, shelf number, and shelf label location information. The product name is the product number or identification name assigned to the product. The product label location represents the 3D location of the product label corresponding to the product. The shelf number is the shelf number or identification name assigned to the shelf on which the product is placed. The shelf label location represents the 3D location of the shelf label on the shelf on which the product is placed. The product label location and shelf label location are represented by XYZ coordinate values ​​in a 3D Cartesian coordinate system set in the space of the target area.

[0135] In the shelf attribute information acquisition process P5 (see Figure 11), attribute information of the shelf (installed object) is acquired. Shelf attribute information is information about the shelf's configuration. Shelf attribute information is stored in a tabular file. Shelf attribute information is saved in association with the shelf skeleton model.

[0136] Specifically, as shown in Figure 14B, the shelf attribute information includes the following information: column position, number of shelves, shelf height, nearby product names, and shelf number. The column position represents the position of the reference column on the horizontal plane (floor). The column position is expressed as XY coordinate values ​​in the 2D Cartesian coordinate system of the target area. The number of shelves represents the number of shelves on both the left and right sides of the column. The shelf height represents the height at which each shelf is positioned on both the left and right sides of the column. The shelf height is expressed as the height from the floor (Z coordinate value). Nearby product names are product numbers or identification names assigned to products in the vicinity of the reference column. The shelf number is the shelf number or identification name assigned to the shelf.

[0137] Next, we will describe the shooting data input screen 101 and the mode selection screen 111 displayed on the user terminal 3. Figures 15A and 15B are explanatory diagrams showing the shooting data input screen 101 and the mode selection screen 111.

[0138] As shown in Figure 15A, the shooting data input screen 101 is provided with a file specification unit 102. The file specification unit 102 allows the user to specify the shooting data file. Specifically, for example, when the user operates the file specification unit 102, a file selection screen (not shown) is displayed, and on that file selection screen, the user can select the folder where the shooting data file is stored and then select the shooting data file. The user can also select previously selected folders or files by operating the pull-down menu.

[0139] The shooting data input screen 101 is also provided with a "Load" button 103 and a "Cancel" button 104. When the user operates the "Load" button 103, the shooting data file stored in the memory unit 22 is read, and the system transitions to the mode selection screen 111 (see Figure 15B). When the user operates the "Cancel" button 104, the system returns to a menu screen (not shown).

[0140] As shown in Figure 15B, the mode selection screen 111 is provided with a captured image display unit 112. The captured image display unit 112 displays the color captured images contained in the loaded capture data file in a row. The user can view all the color captured images on the captured image display unit 112 by operating the scroll bar.

[0141] Furthermore, the mode selection screen 111 is provided with a "shelf skeleton mode" button 113 and a "normal mode" button 114. When the user operates the "shelf skeleton mode" button 113, the screen transitions to the shelf skeleton mode start screen 201 (see Figure 16). On the other hand, when the user operates the "normal mode" button 114, the screen transitions to the normal mode start screen 301 (see Figure 22A). In shelf skeleton mode, a shelf skeleton model is generated, that is, a mesh model in a skeleton state with unnecessary objects around the shelf removed. In normal mode, an area model is generated, that is, a mesh model of the entire area.

[0142] Next, we will explain the shelf skeleton mode start screen 201 displayed on the user terminal 3. Figure 16 is an explanatory diagram showing the shelf skeleton mode start screen 201.

[0143] The shelf skeleton mode start screen 201 is equipped with a "Start" button 202 and a "Cancel" button 203. When the user operates the "Start" button 202, the process proceeds to the shelf skeleton mode and transitions to the texture selection screen 211 (see Figures 17A and 17B). When the user operates the "Cancel" button 203, the user returns to the mode selection screen 111 (see Figure 15B).

[0144] Next, we will explain the texture selection screen 211 displayed on the user terminal 3. Figures 17A and 17B are explanatory diagrams showing the texture selection screen 211.

[0145] The texture selection screen 211 is provided with a status display unit 212. The status display unit 212 includes a "shelf texture" box 213, a "storage information" box 214, and a "shelf attribute information" box 215. The display form (color, density, pattern, etc.) of boxes 213, 214, and 215 changes according to the processing stage.

[0146] Furthermore, the texture selection screen 211 is provided with a captured image display unit 216. The captured image display unit 216 displays each of the color captured images included in the loaded capture data.

[0147] The captured image display unit 216 allows the user to specify the area from which to extract texture (mainly color). Specifically, the user can specify the area from which to extract texture using a polygon (e.g., a rectangle) by performing a predetermined operation (e.g., dragging with the mouse) on the color captured image displayed on the captured image display unit 216.

[0148] Here, the texture selection screen 211 shown in Figure 17A allows the user to specify the texture of the shelf's side panels, and a message guiding the user to do so, specifically, the words "Please select the texture of the shelf's side panels," is displayed. The user then selects the texture of the shelf's side panels according to the instructions. The texture selection screen 211 shown in Figure 17B allows the user to specify the texture of the shelf's support columns, and a message guiding the user to do so, specifically, the words "Please select the texture of the shelf's support columns," is displayed. The user then selects the texture of the shelf's support columns according to the instructions.

[0149] The texture selection screen 211 also has a "Next" button 217, an "OK" button 218, and a "Cancel" button 219. When the user operates the "Next" button 217 on the first texture selection screen 211 (see Figure 17A), the screen transitions to the next texture selection screen 211 (see Figure 17B). When the user operates the "Cancel" button 219, the screen returns to a predetermined screen (for example, the shelf skeleton mode start screen 201 (see Figure 16)). When the user operates the "OK" button 218, the texture extraction process P11 and the training model generation process P12 (see Figure 6) are started. Once the texture extraction process P11 and the training model generation process P12 are completed, the screen transitions to the training model confirmation screen 221 (see Figure 18).

[0150] Next, we will explain the training model confirmation screen 221 displayed on the user terminal 3. Figure 18 is an explanatory diagram showing the training model confirmation screen 221.

[0151] The learning model confirmation screen 221 is provided with a learning model display unit 222. The learning model display unit 222 displays the learning model (CAD model of the learning shelf) generated by the learning model generation process P12. This allows the user to confirm whether the learning model that will serve as the basis for the learning images has been generated appropriately.

[0152] The learning model confirmation screen 221 is also provided with an "OK" button 223 and a "Cancel" button 224. The learning model confirmation screen 221 also displays a message informing the user to start the process, specifically, "Pressing OK will start the process of modeling the shelves, acquiring product location information, and acquiring shelf attribute information." When the user operates the "OK" button 223, the learning image generation process P13 and the learning process P14 (see Figure 6) are started, and the screen transitions to the processing screen 231 (see Figure 19A). After the learning image generation process P13 and the learning process P14 are completed, the shelf skeleton model generation process P2, the storage information acquisition process P3, and the shelf attribute information acquisition process P5 (see Figure 5) are executed.

[0153] Next, the processing screen 231 and the shelf skeleton model confirmation screen 241 displayed on the user terminal 3 will be described. Figures 19A and 19B are explanatory diagrams showing the processing screen 231 and the shelf skeleton model confirmation screen 241.

[0154] As shown in Figure 19A, the processing screen 231 displays a progress bar 232 (progress display unit) showing the progress of the shelf skeleton model generation process P2, a progress bar 233 (progress display unit) showing the progress of the storage information acquisition process P3, and a progress bar 234 (progress display unit) showing the progress of the shelf attribute information acquisition process P5. Once the shelf skeleton model generation process P2, the storage information acquisition process P3, and the shelf attribute information acquisition process P5 are completed, the screen transitions to the shelf skeleton model confirmation screen 241 (see Figure 19B).

[0155] As shown in Figure 19B, the shelf skeleton model confirmation screen 241 is provided with a shelf skeleton model display unit 242. The shelf skeleton model display unit 242 displays the shelf skeleton model, that is, a mesh model in a skeleton state from which unnecessary objects surrounding the shelf have been removed. This allows the user to confirm whether the shelf skeleton model has been generated correctly.

[0156] The shelf skeleton model confirmation screen 241 is equipped with an "OK" button 243 and a "Cancel" button 244. The shelf skeleton model confirmation screen 241 also displays a message instructing the user to save the shelf skeleton model, specifically, for example, "Saving the shelf skeleton model to Shelf.ply." If the user operates the "OK" button 243, the file containing the shelf skeleton model is saved to the storage unit 22 with a predetermined file name. On the other hand, if the user operates the "Cancel" button 244, the user transitions to a predetermined screen (for example, a screen for changing processing conditions) and can redo appropriate processing such as the shelf skeleton model generation process P2.

[0157] Furthermore, the status display unit 212 displays a message regarding the presentation of storage information, specifically, the message, "Click to display the product's storage information." Also, a message regarding the presentation of shelf attribute information, specifically, the message, "Click to display shelf attribute information," is displayed. When the user operates the "Storage Information" box 214, the screen transitions to the storage information presentation screen 251 (see Figures 20A and 20B), and the storage information is presented to the user. When the user operates the "Shelf Attribute Information" box 215, the screen transitions to the shelf attribute information presentation screen 261 (see Figure 21), and the shelf attribute information is presented to the user.

[0158] Next, we will explain the storage information display screen 251 that is displayed on the user terminal 3. Figures 20A and 20B are explanatory diagrams showing the storage information display screen 251.

[0159] The storage information display screen 251 is provided with a shelf skeleton model display unit 242, similar to the shelf skeleton model confirmation screen 241 (see Figure 19B). The shelf skeleton model display unit 242 displays the shelf skeleton model as a 2D image based on the specified viewpoint.

[0160] Furthermore, the shelf skeleton model display unit 242 overlays character images representing product and shelf identification information (product number, shelf number) as storage information onto the shelf skeleton model. These character images representing product and shelf identification information are displayed on the shelf skeleton model at positions corresponding to the actual placement of labels (product labels, shelf labels) based on the product and shelf location information.

[0161] In the example shown in Figure 20A, the shelf skeleton model display unit 242 displays an enlarged view of the support columns and shelves in the shelf skeleton model, and on top of that, a character image representing the product number (product identification information) written on the product label is displayed as storage information.

[0162] In the example shown in Figure 20B, the shelf skeleton model display unit 242 displays an enlarged view of the side panel portion of the shelf skeleton model, and on top of that, the shelf number (shelf identification information) written on the shelf label placed on the side panel is displayed as storage information.

[0163] Furthermore, the storage information display screen 251 is equipped with an "OK" button 243 and a "Cancel" button 244. The storage information display screen 251 also displays a message instructing the user to save the storage information, specifically, the text, "The product location information linked to the shelf number will be saved to productposinfo.csv." When the user operates the "OK" button 243, a file containing the storage information is given a predetermined file name and saved to the storage unit 22. At this time, the storage information is saved in association with the shelf skeleton model. On the other hand, when the user operates the "Cancel" button 244, the user is redirected to a predetermined screen (for example, a screen for changing processing conditions), and can redo appropriate processing such as the storage information acquisition process P3.

[0164] Next, we will explain the shelf attribute information display screen 261 that is displayed on the user terminal 3. Figure 21 is an explanatory diagram showing the shelf attribute information display screen 261.

[0165] The shelf attribute information display screen 261 is provided with a shelf skeleton model display unit 242, similar to the shelf skeleton model confirmation screen 241 (see Figure 19B). The shelf skeleton model display unit 242 displays the shelf skeleton model as an image based on the specified viewpoint.

[0166] Furthermore, the shelf skeleton model display unit 242 overlays character images representing shelf attribute information onto the shelf skeleton model. Specifically, the character representing the number of shelves and their height is displayed as shelf attribute information at positions on the shelf skeleton model corresponding to the actual shelf positions.

[0167] Furthermore, the shelf attribute information display screen 261 is equipped with an "OK" button 243 and a "Cancel" button 244. The shelf attribute information display screen 261 also displays a message instructing the user to save the shelf attribute information, specifically, the text, "Shelf attribute information will be saved to the ShelfInfo.csv file." When the user operates the "OK" button 243, a file containing the shelf attribute information is saved to the storage unit 22 with a predetermined file name. At this time, the shelf attribute information is saved in association with the shelf skeleton model. On the other hand, when the user operates the "Cancel" button 244, the user is redirected to a predetermined screen (for example, a screen for changing processing conditions), and can redo appropriate processing such as the shelf attribute information acquisition process P5.

[0168] In this embodiment, the shelf attribute information (number of shelves and height) is superimposed on an image of the shelf skeleton model viewed from a specified viewpoint. However, the shelf attribute information (number of shelves and height) may also be superimposed on a plan view (2D drawing) of the shelf.

[0169] Next, the normal mode start screen 301 and the area model confirmation screen 311 displayed on the user terminal 3 will be described. Figures 22A and 22B are explanatory diagrams showing the normal mode start screen 301 and the area model confirmation screen 311.

[0170] As shown in Figure 22A, the normal mode start screen 301 is provided with a voxel size setting unit 302. In the voxel size setting unit 302, the user can specify the voxel size by operating a pull-down menu. In the area model generation process P6, prior to meshing the point cloud data, voxelization is performed by downsampling the point cloud using a voxel grid filter, and the size of the voxel grid filter (voxel size) at this time is specified by the user.

[0171] Furthermore, the normal mode start screen 301 is equipped with a "Start" button 303 and a "Cancel" button 304. When the user operates the "Start" button 303, the normal mode processing, specifically the area model generation process P6, is started. Once the area model generation process P6 is completed, the screen transitions to the area model confirmation screen 311 (see Figure 22B). When the user operates the "Cancel" button 304, the screen returns to the mode selection screen 111 (see Figure 15B).

[0172] As shown in Figure 22B, the area model confirmation screen 311 is provided with an area model display unit 312. The area model display unit 312 displays the area model generated in the area model generation process P6, that is, the mesh model of the entire target area, as an image based on the specified viewpoint.

[0173] The area model screen also has an "OK" button 313 and a "Cancel" button 314. The area model screen also displays a message instructing the user to save the area model, specifically, "Saving mesh model general_01.ply." When the user operates the "OK" button 313, the file containing the area model is saved to the storage unit 22 with a predetermined file name. On the other hand, when the user operates the "Cancel" button 314, the user is redirected to a predetermined screen (for example, a screen for changing processing conditions) and can restart the area model generation process P6.

[0174] In this embodiment, in response to user operation, the image recognition engine construction process P1 is performed first, followed by the shelf skeleton model generation process P2. However, the image recognition engine construction process P1 and the shelf skeleton model generation process P2 may be performed separately. For example, an image recognition engine may already be constructed, in which case only the shelf skeleton model generation process P2 may be performed using the constructed image recognition engine.

[0175] (First Modification) Next, the first modification will be described. Points that are not specifically mentioned here are the same as in the above embodiment. Figure 23 is an overall configuration diagram of the 3D model generation system according to the first modification.

[0176] In the above embodiment, the user performs an operation on the user terminal 3 to instruct the model generation server 2 to perform processing, and the model generation server 2 executes the processing according to the pre-set conditions. On the other hand, in this modified example, Large Language Models (LLMs) are used to instruct the processing performed on the model generation server 2. That is, an LLM server 4 intervenes between the model generation server 2 and the user terminal 3, interprets prompts (instructions) from the user terminal 3, and instructs the model generation server 2 to perform the necessary processing.

[0177] Specifically, LLM is used in the process of generating training images from the training model (training image generation process P13). Furthermore, multimodal LLM may be used to perform necessary image processing. When using LLM to instruct the model generation server 2 to perform the training image generation process P13, the prompt sent from the user terminal 3 to the LLM server 4 may be, for example, "Use this texture to generate a large amount of training data for semantic segmentation under various lighting conditions."

[0178] (Second Modification) Next, a second modification will be described. Points that are not specifically mentioned here are the same as in the above embodiment. Figure 24 is a block diagram showing the contents of the integrated model generation process P8 performed by the processor 23 of the model generation server 2 according to the second modification.

[0179] In this modified version, only the shelves (installed objects) within the target area are reproduced as a skeleton mesh model (shelf skeleton model), while objects other than the shelves within the target area are reproduced as a normal mesh model. The processor 23 of the model generation server 2 performs the integrated model generation process P8. The integrated model generation process P8 includes the point cloud generation process for objects other than the shelves P81, the meshing process P82, and the model integration process P83.

[0180] In the non-shelf point cloud generation process P81, the processor 23 generates non-shelf point cloud data by excluding the shelf point cloud data generated in the shelf point cloud generation process P23 of the shelf skeleton model generation process P2 from the point cloud data of the entire area generated in the area point cloud generation process P61 of the area model generation process P6. The non-shelf point cloud data includes structures such as the warehouse floor, walls, and columns.

[0181] In the meshing process P82, the processor 23 meshes the point cloud data other than shelves generated in the non-shelf point cloud generation process P81 to generate a mesh model of non-shelf data (non-shelf model).

[0182] In the model integration process P83, the processor 23 integrates the mesh model of non-shelf elements generated in the mesh generation process P82 with the shelf skeleton model (mesh model of the shelf skeleton) generated in the shelf skeleton model generation process P2 to generate a mesh model of the entire area (integrated model). At this time, the shelf skeleton model is placed at the location of the shelf on the non-shelf model.

[0183] In the resulting mesh model (integrated model) of the entire area, it becomes possible to simulate changing the layout of the warehouse by moving the shelf skeleton model relative to the mesh model of structural elements such as the warehouse floor, walls, and columns.

[0184] As described above, embodiments have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these embodiments and can be applied to embodiments that have been modified, replaced, added, or omitted. Furthermore, it is possible to create new embodiments by combining the components described in the above embodiments.

[0185] The 3D model generation apparatus and 3D model generation method relating to this disclosure have the effect of generating lightweight 3D models that allow users to appropriately confirm the status of installed objects and contents in a target area, and that enable users to perform simulation work smoothly. They are useful as a 3D model generation apparatus and 3D model generation method for generating 3D models that reproduce the status of installed objects in a target area.

[0186] 1: Shooting terminal 2: Model generation server (3D model generation device) 3: User terminal (user device) 4: LLM server (language model device) 23: Processor P1: Image recognition engine construction process P2: Shelf skeleton model generation process (first process) P3: Storage information acquisition process (second process) P4: Shelf plan view creation process P5: Shelf attribute information acquisition process (third process) P6: Area model generation process (fourth process) P7: Display process P8: Integrated model generation process P11: Texture extraction process P12: Training model generation process P13: Training image generation process P14: Training process P21: Image recognition process P22: Shelf depth data extraction process P23: Shelf point cloud generation process P24, P62, P82: Meshing process 101: Shooting data input screen 201: Shelf skeleton mode start screen 211: Texture specification screen 221: Learning model confirmation screen 231: Processing screen 241: Shelf skeleton model confirmation screen 251: Storage information display screen 261: Shelf attribute information display screen 301: Normal mode start screen 311: Area model confirmation screen

Claims

1. A 3D model generation device that uses a processor to generate a 3D model that reproduces the state of an installation in a target area, wherein the processor performs a first process of generating a 3D model of the installation in a state where surrounding objects have been removed, a second process of acquiring arrangement information representing the arrangement of contents in the target area, and a third process of acquiring attribute information of the installation, and stores the 3D model of the installation, the arrangement information of the contents, and the attribute information of the installation together in a storage unit in relation to each other.

2. The 3D model generation apparatus according to claim 1, characterized in that the processor switches between a first mode for executing the first process, the second process, and the third process, and a second mode for executing a fourth process for generating a 3D model of the entire target area, in accordance with user operation.

3. The 3D model generation apparatus according to claim 1, characterized in that the processor, in the first processing, detects an installed object from a captured image of the target area using an image recognition engine that performs image segmentation, extracts depth data of the installed object from the depth data of the target area acquired in each capture based on the region information of the installed object on the captured image, generates point cloud data of the installed object based on the depth data of the installed object, and generates a 3D model of the installed object in a state in which objects present around the installed object have been excluded based on the point cloud data of the installed object.

4. The 3D model generation apparatus according to claim 3, characterized in that the processor generates a CAD model of the installation object to which feature information relating to the appearance of the installation object extracted from the captured image is applied, generates a plurality of training images with different 2D imaging processing conditions from the CAD model of the installation object, and constructs the image recognition engine by machine learning using the training images.

5. The 3D model generation device according to claim 4, wherein the device is connected to a language modeling device that analyzes a natural language instruction written by a user to obtain the processing conditions for 2D imaging, and the processor generates the training image based on the processing conditions for 2D imaging obtained from the language modeling device.

6. The 3D model generation apparatus according to claim 1, characterized in that the processor, in the second processing, acquires, as arrangement information, identification information and location information of the contained object and identification information and location information of the installation in which the contained object is contained.

7. The 3D model generation apparatus according to claim 6, characterized in that the processor detects labels that identify installed objects and contained objects from captured images of the target area, obtains identification information of the installed objects and contained objects included in the labels, extracts label depth data from depth data of the target area obtained in each capture based on the region information of the labels on the captured images, generates point cloud data of the labels based on the label depth data, and obtains position information of the labels in 3D space based on the label point cloud data.

8. The 3D model generation apparatus according to claim 1, characterized in that the processor acquires information regarding the configuration of the installed object as attribute information of the installed object in the third processing.

9. A 3D model generation method for causing a processor to execute a process to generate a 3D model that reproduces the state of an installation in a target area, characterized in that the method performs a first process to generate a 3D model of the installation in a state where surrounding objects have been removed, a second process to acquire arrangement information representing the arrangement of contents in the target area, and a third process to acquire attribute information of the installation, and stores the 3D model of the installation, the arrangement information of the contents, and the attribute information of the installation together in a memory unit in relation to each other.

10. A 3D model generation device that uses a processor to perform a process of generating a 3D model that reproduces the state of an installed object in a target area, wherein the processor uses an image recognition engine that performs image segmentation to detect the installed object from a captured image of the target area, extracts depth data of the installed object from depth data within the target area acquired in each capture based on the region information of the installed object on the captured image, generates point cloud data of the installed object based on the depth data of the installed object, and generates a 3D model of the installed object in a state in which objects present around the installed object have been removed based on the point cloud data of the installed object.

11. The 3D model generation apparatus according to 10, characterized in that the processor generates a CAD model of the installation object to which feature information relating to the appearance of the installation object extracted from the captured image is applied, generates a plurality of training images with different 2D imaging processing conditions from the CAD model of the installation object, and constructs the image recognition engine by machine learning using the training images.

12. The 3D model generation device according to claim 11, wherein the device is connected to a language modeling device that analyzes a natural language instruction written by a user to obtain processing conditions for 2D imaging, and the processor generates the training image based on the processing conditions for 2D imaging obtained from the language modeling device.

13. The 3D model generation apparatus according to claim 10, characterized in that the processor detects labels that identify the installation and contents from the captured image, obtains identification information of the installation and contents included in the label, extracts depth data of the label from the depth data of the target area obtained in each capture based on the region information of the label on the captured image, generates point cloud data of the label based on the depth data of the label, and obtains position information of the label in 3D space based on the point cloud data of the label.

14. The 3D model generation apparatus according to 13, characterized in that the processor superimposes images representing identification information of the installation and contents onto corresponding positions on the 3D model of the installation, based on the positional information of the labels.

15. A 3D model generation method that causes a processor to execute a process to generate a 3D model that reproduces the state of an installed object in a target area, characterized in that: an image recognition engine that performs image segmentation is used to detect the installed object from a captured image of the target area; depth data of the installed object is extracted from depth data within the target area acquired in each capture based on the region information of the installed object on the captured image; point cloud data of the installed object is generated based on the depth data of the installed object; and a 3D model of the installed object is generated based on the point cloud data of the installed object, with objects present around the installed object excluded.