Information processing device, information processing method, and program

By incorporating multimodal information to update labels in 3D spatial recognition maps, the technology addresses the challenge of accurately categorizing difficult furniture shapes, enabling correct AR character behavior.

WO2025173529A1PCT designated stage Publication Date: 2025-08-21SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/002524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-01-28
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing 3D spatial recognition maps struggle to accurately categorize furniture with simple or unique shapes, leading to incorrect AR character behavior due to unrecognized furniture categories.

Method used

Utilize multimodal information, such as human actions and positional relationships, to update labels in 3D spatial recognition maps, improving label accuracy for furniture by complementing 2D image-based recognition.

Benefits of technology

Enhances the accuracy of furniture categorization in 3D spatial recognition maps, ensuring AR characters perform intended actions by correctly identifying challenging furniture types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025002524_21082025_PF_FP_ABST
    Figure JP2025002524_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing device, an information processing method, and a program that make it possible to acquire a 3D spatial recognition map including a label indicating a correct category of an object. An information processing device according to one aspect of the present technology: subjects a first 2D image to segmentation; generates a 3D spatial recognition map in which a label indicating a category of each of a plurality of objects in real space is set for each region; determines whether the plurality of objects include an unknown object for which the first label was not recognized in the segmentation; determines a category of the objects corresponding to the movement of a person appearing in a second 2D image and a positional relationship between the person and the unknown object; and, if the plurality of objects include an unknown object, updates the label of the unknown object on the basis of the determination result of the category of the object corresponding to the movement of the person and the determination result of the positional relationship. The present technology can be applied to information processing devices that control the display of AR characters.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present technology relates to an information processing device, an information processing method, and a program, and in particular to an information processing device, an information processing method, and a program that enable acquisition of a 3D spatial recognition map including a label indicating the correct category of an object.

[0002] There is a technology that uses a camera or a ToF sensor to scan a real space and generate a 3D spatial recognition map that describes the meaning of the real space in a format such as a Scene Graph or a Panoptic Map. 3D spatial recognition maps are used in applications that make an AR (Augmented Reality) character take actions that are in line with the meaning of the real space (Non-Patent Document 1). By using a 3D spatial recognition map, it is possible to make an AR character take actions such as "sitting in a chair facing the TV and watching a TV program" in various spaces.

[0003] Japanese Patent Application Laid-Open No. 2019-192145

[0004] Tomu Tahara, et al., “Retargetable AR: Context-aware Augmented Reality in Indoor Scenes based on 3D Scene Graph”, 2020 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct), 2020

[0005] The categories of furniture in the real world are recognized when creating a 3D spatial recognition map, but there is furniture in the real environment that is difficult to recognize as a category. Examples of furniture that is difficult to correctly recognize as a category include: - Furniture with a simple shape that is difficult to distinguish from other objects - Furniture that emphasizes design - Furniture that the user has made themselves

[0006] If such furniture exists in the real space where the AR character is to be operated, the furniture whose category cannot be recognized cannot be correctly reflected in the 3D spatial recognition map, and in this case, the AR character will not be able to operate as intended.

[0007] This technology was developed in light of these circumstances, and makes it possible to obtain a 3D spatial recognition map that includes labels indicating the correct category of objects.

[0008] An information processing device according to one aspect of the present technology includes: a 3D space recognition map generation unit that performs segmentation on a first 2D image and generates a 3D space recognition map in which a first label indicating a category of each of a plurality of objects in real space is set in each region; a first determination unit that determines whether an unknown object whose first label was not recognized in the segmentation is included in the plurality of objects; a second determination unit that determines a category of the object according to the movement of a person appearing in the second 2D image and a positional relationship between the person and the unknown object; and a label update unit that updates the first label of the unknown object based on a determination result of the category of the object according to the movement of the person and a determination result of the positional relationship when the unknown object is included in the plurality of objects.

[0009] In one aspect of the present technology, a first 2D image is segmented, a 3D space recognition map is generated in which a first label indicating a category of each of a plurality of objects in real space is set for each region, and it is determined whether an unknown object whose first label was not recognized in the segmentation is included in the plurality of objects. In addition, a category of the object according to a movement of a person appearing in the second 2D image and a positional relationship between the person and the unknown object are determined, and if the unknown object is included in the plurality of objects, the first label of the unknown object is updated based on the determination result of the category of the object according to the movement of the person and the determination result of the positional relationship.

[0010] 1 is a diagram illustrating an example of a user experience realized by an AR application. FIG. 2 is a diagram illustrating an example of panoptic segmentation. FIG. 3 is a diagram illustrating an example of furniture that may not be recognized correctly. A block diagram illustrating an example of the configuration of an information processing unit. FIG. 4 is a diagram illustrating an example of a 3D Scene Graph. FIG. 5 is a block diagram illustrating an example of the configuration of a label accuracy improvement processing unit. FIG. 6 is a flowchart illustrating low-reliability labeled area detection processing. FIG. 7 is a flowchart illustrating label update processing. FIG. 8 is a diagram illustrating a flow of label completion using human action recognition. A block diagram illustrating an example of the configuration of an information processing unit. A block diagram illustrating an example of the configuration of a label accuracy improvement processing unit. FIG. 9 is a flowchart illustrating label correspondence relationship list generation processing. A flowchart illustrating label update processing. A block diagram illustrating an example of the hardware configuration of an AR display device. A diagram illustrating an example of an information processing system including an external device. A block diagram illustrating an example of the configuration of a computer.

[0011] Hereinafter, embodiments of the present technology will be described. The description will be made in the following order: 1. Overview of the present technology 2. Label completion using multimodal information 3. Label completion using human action recognition 4. Configuration example 5. Modified example

[0012] <<Outline of the Present Technology>> <Example of Application> FIG. 1 is a diagram showing an example of a user experience realized by an AR application to which the present technology is applied.

[0013] An AR application is an application in which a real person wearing an AR display device becomes a user and communicates with a virtual character. In the example of Fig. 1, an AR character, which is a humanoid virtual character, is presented to a user wearing an AR display device 1.

[0014] The AR display device 1 is, for example, an optical see-through head-mounted display (HMD). When an image of an AR character is displayed on the display unit of the AR display device 1, the user sees the AR character superimposed on the scenery in front of them. When using an AR application indoors as shown in Figure 1, the user can communicate with the AR character, which acts in the same space as the user. The display of the AR character is controlled so that it behaves in accordance with the meaning of the space, such as sitting in a chair in the space or reading a book on a table.

[0015] The image used to display the AR character is generated by the AR display device 1 or by an external device connected to the AR display device 1. The following mainly describes the case where the image used to display the AR character is generated by the AR display device 1. The AR display device 1 is an information processing device that performs various processes such as recognizing the spatial situation and planning the AR character's behavior. In addition to a display unit used to display the image, the AR display device 1 is provided with various sensor devices such as a camera that captures the state of the space and a ToF sensor that measures the distance to objects such as furniture in the space. The AR character's behavior plan uses a 3D spatial recognition map that describes the meaning of the real space.

[0016] <Generation of 3D spatial recognition map> A 3D spatial recognition map is generated by performing panoptic segmentation on a 2D image such as an RGB image or a depth image, and projecting the labels for each pixel, which are the recognition results of panoptic segmentation, onto a 3D spatial map based on the position and orientation of the camera. The 3D spatial map is composed of a 3D mesh that indicates the shape of each object in the space. A 3D spatial map with a label set in each region becomes the 3D spatial recognition map.

[0017] Fig. 2 shows an example of panoptic segmentation of a 2D image. Fig. 2A shows the space where the user is located, and Fig. 2B shows an example of the results of panoptic segmentation on a 2D image obtained by capturing the space where the user is located.

[0018] In panoptic segmentation, semantic segmentation and instance segmentation are performed on 2D images to identify labels for each pixel. The label includes the ID of the object region that each pixel constitutes (object ID and ID indicating the object category (type)). A label including the object ID of that object and an ID indicating the object category is set for the region of a certain object.

[0019] Label recognition in Panoptic Segmentation is performed by classifying objects using a DNN (Deep Neural Network). The accuracy of the labels in the 3D spatial recognition map depends on the accuracy of the DNN's class classification based on 2D images. Furniture with shapes that are difficult to include in the furniture shown in the 2D images of the training data, or furniture with shapes that make class classification ambiguous, may not be labeled correctly.

[0020] Figure 3 shows examples of furniture whose categories may not be correctly recognized. The cubic object shown in Figure 3A and the spherical object shown in Figure 3B are both chairs. Furniture with a simple shape that is difficult to distinguish from other objects, such as the chair (stool) shown in Figure 3A, may not be correctly recognized. In addition, furniture with a unique shape that emphasizes design, such as the chair shown in Figure 3B, may not be correctly recognized.

[0021] If the label is not correctly recognized even though a chair like the one shown in Figure 3A is in front of the TV, the AR application will determine that there is no chair facing the TV, even if the situation is intended for the AR character to "sit in a chair facing the TV and watch a TV program." As a result, the AR character will move to another location instead of sitting in the chair that actually exists, which is different from the behavior expected by the user.

[0022] In this technology, for objects whose shapes are difficult to label based on 2D images, label recognition is performed using multimodal information, which is a type of information different from 2D images, and the label recognition results are reflected in a 3D spatial recognition map. For example, the following processing is performed: 1. For objects whose labels are difficult to recognize based on 2D images, i.e., objects with low-confidence labels, label recognition is performed using multimodal information. 2. Information indicating the correspondence between the label recognition results using multimodal information and low-confidence label areas in the 3D spatial recognition map is generated, and the low-confidence labels are updated using the label recognition results using multimodal information.

[0023] Specifically, if the reliability of the label for area A in the 3D spatial recognition map is determined to be low, label recognition for area A is performed by using multimodal information acquired in synchronization with the image capture by the camera. If label B is acquired as the label recognition result using multimodal information, information indicating the correspondence between area A and label B is generated, and the label for area A in the 3D spatial recognition map is updated using label B to complement the label. This makes it possible to complement the label of an object that is difficult to correctly recognize using image data alone with the label recognition result using multimodal information.

[0024] <<Label Completion Using Multimodal Information>> <Configuration of Information Processing Unit> Fig. 4 is a block diagram showing an example configuration of the information processing unit 11. The information processing unit 11 having the configuration shown in Fig. 4 is provided in the AR display device 1, for example.

[0025] As shown in FIG. 4, the information processing unit 11 is composed of an image acquisition unit 21, a self-position estimation unit 22, a 3D spatial recognition map generation unit 23, a multimodal information acquisition unit 24, a label accuracy improvement processing unit 25, a 3D Scene Graph generation unit 26, and a behavior planning unit 27.

[0026] The image acquisition unit 21 controls a sensor device including an RGB camera to acquire 2D images such as RGB images and depth images. The 2D image data acquired by the image acquisition unit 21 is supplied to the self-position estimation unit 22, the 3D space recognition map generation unit 23, and the other modal information acquisition unit 24.

[0027] The sensor devices used to acquire the RGB image and the depth image may be an RGB camera and a ToF sensor, respectively, or a device such as an RGB-D camera that combines an RGB camera and a ToF sensor. Also, a device that acquires a depth image using a monocular depth estimation technique based on an RGB image captured by an RGB camera may be used.

[0028] The self-position estimation unit 22 acquires measurement results from the acceleration sensor, gyro sensor, and positioning sensor, and estimates the position and orientation of the AR display device 1 in the space where the user is located. The 2D image supplied from the image acquisition unit 21 is used as appropriate to estimate the position and orientation of the AR display device 1. The information on the position and orientation of the AR display device 1 estimated by the self-position estimation unit 22 is output to the 3D space recognition map generation unit 23 and the label accuracy improvement processing unit 25 as camera position information.

[0029] The 3D space recognition map generation unit 23 performs panoptic segmentation based on the 2D images supplied from the image acquisition unit 21 and the camera position information supplied from the self-position estimation unit 22, and recognizes the labels of each object in the space. Specifically, the 3D space recognition map generation unit 23 performs semantic segmentation and instance segmentation on the 2D images to recognize the labels for each pixel. The 3D space recognition map generation unit 23 also integrates the labels into a 3D space map generated by SLAM using depth images to generate a 3D space recognition map. PanopticFusion, one method of panoptic segmentation, is described in, for example, Literature 1. Literature 1: Gaku Narita, et al., “PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things”, Proc. IROS, 2019

[0030] In this way, the 3D space recognition map generation unit 23 has a function of segmenting the 2D image supplied from the image acquisition unit 21 and generating a 3D space recognition map. A label (first label) indicating each category of multiple objects, such as furniture, in the real space is set in each area of ​​the 3D space recognition map generated by the 3D space recognition map generation unit 23. During execution of the AR application, the camera system operates, and the 3D space recognition map is generated using each frame of the 2D image. Furthermore, estimation of the self-position in the 3D space recognition map is repeatedly performed.

[0031] The multimodal information acquisition unit 24 acquires multimodal information used to complement the labels. The multimodal information is generated and acquired by the multimodal information acquisition unit 24 based on, for example, 2D images supplied from the image acquisition unit 21. If the 2D image used to generate the 3D spatial recognition map is the first 2D image, the 2D image used to acquire the multimodal information is the second 2D image. 2D images captured at the same time may be used as the first 2D image and the second 2D image, or a 2D image captured after the 3D spatial recognition map is generated using the first 2D image may be used as the second 2D image.

[0032] The multi-modal information may be generated using information other than 2D images, such as audio detected by a microphone. The multi-modal information acquired by the multi-modal information acquisition unit 24 is output to the label accuracy improvement processing unit 25.

[0033] The label accuracy improvement processing unit 25 updates low-reliability labels in the 3D space recognition map based on the camera position information supplied from the self-position estimation unit 22, the 3D space recognition map supplied from the 3D space recognition map generation unit 23, and the other-modal information supplied from the other-modal information acquisition unit 24. Information on the 3D space recognition map whose labels have been updated by the label accuracy improvement processing unit 25 is output to the 3D Scene Graph generation unit 26. Details of the configuration and operation of the label accuracy improvement processing unit 25 will be described later.

[0034] The 3D Scene Graph generation unit 26 generates a 3D Scene Graph based on the 3D spatial recognition map supplied from the label accuracy improvement processing unit 25. Various types of sensor data, such as RGB images captured by an RGB camera and depth images detected by a depth sensor, are also input to the 3D Scene Graph generation unit 26 and are used as appropriate to generate the 3D Scene Graph.

[0035] FIG. 5 is a diagram showing an example of a 3D Scene Graph.

[0036] If a user's environment contains a sofa, a table, a television, chair A, and chair B, the 3D Scene Graph representing the user's environment will contain five nodes representing these objects, as shown in Figure 5.

[0037] In the example of FIG. 5, the sofa node and the TV node share an edge E 1 Edge E 1 The label "on_right" indicates that the sofa is in front of the TV. The sofa node and the table node share an edge E with the label "on_right". 2 Edge E 2 The label indicates that the table is to the right of the sofa.

[0038] The TV node and the table node have an edge E with the label “on_left”. 3 Edge E 3 The label indicates that the table is on the left side of the TV. The edge E between the table node and the chair A node, and the edge E between the table node and the chair B node are also labeled with a label indicating their respective positional relationships. 4 , E 5 The labels assigned to the edges are labels that represent spatial positional relationships (front / behind / left / right / on / above / under / near, etc.).

[0039] In this way, a 3D Scene Graph is information with a graph structure in which real objects, such as furniture, that exist in the environment are represented as nodes, and the relationship between two objects, such as their positional relationship, is represented as edges. The generation of a 3D Scene Graph is described, for example, in Literature 2 (the same as Non-Patent Document 1 above). Literature 2: Tomu Tahara, et al., "Retargetable AR: Context-aware Augmented Reality in Indoor Scenes based on 3D Scene Graph," 2020 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct), 2020

[0040] The information of the 3D Scene Graph generated by the 3D Scene Graph generating unit 26 is supplied to the action planning unit 27 shown in FIG.

[0041] The behavior planning unit 27 plans the behavior of the AR character, which is an autonomous moving body, based on the 3D Scene Graph supplied from the 3D Scene Graph generation unit 26. Behaviors such as sitting in a chair that is actually in the space where the user is located or reading a book on a table are planned. An image of the AR character performing the behavior planned by the behavior planning unit 27 is generated in a subsequent processing unit and presented to the user.

[0042] <Configuration and Operation of Label Accuracy Improvement Processing Unit> FIG. 6 is a block diagram showing an example of the configuration of the label accuracy improvement processing unit 25 in FIG.

[0043] 6 , the label accuracy improvement processing unit 25 is composed of a multimodal label recognition unit 51, a low-reliability label area detection unit 52, a label correspondence acquisition unit 53, and a label update unit 54. The multimodal information output from the multimodal information acquisition unit 24 is input to the multimodal label recognition unit 51, and the camera position information output from the self-position estimation unit 22 is input to the label correspondence acquisition unit 53. The 3D space recognition map generated by the 3D space recognition map generation unit 23 based on the 2D image is input to the low-reliability label area detection unit 52, the label correspondence acquisition unit 53, and the label update unit 54. The 3D space recognition map input to the label accuracy improvement processing unit 25 is a 3D space recognition map that includes areas to which low-reliability labels are set.

[0044] Multimodal Label Recognition Unit The multimodal label recognition unit 51 performs label recognition based on the multimodal information and outputs information on the multimodal label recognition result, which is the label recognition result. For label recognition by the multimodal label recognition unit 51, a 3D spatial recognition map and camera position information are used as appropriate. The multimodal label recognition result output from the multimodal label recognition unit 51 is supplied to the label correspondence acquisition unit 53 and the label update unit 54.

[0045] Low-reliability label area detection unit The low-reliability label area detection unit 52 detects low-reliability label areas, which are areas to which a low-reliability label is set, from each segmented area in the 3D space recognition map. The detection of low-reliability label areas is performed by determining whether or not the label of each area is a low-reliability label for each area in the 3D space recognition map.

[0046] For example, the following labels are judged to be low-confidence labels: (1) Labels indicating recognition failure, such as unknown / unlabeled (2) General labels indicating a rough category, such as Other Furniture (3) Labels where the time series of labels for the same region changes significantly, making it unclear which recognition result is correct

[0047] Labels in areas where the recognition results change significantly over time, such as being recognized as "sofa" at one timing, "chair" at the next timing, and "table" at the timing after that, are determined to be low-reliability labels based on the above condition (3).

[0048] Labels whose reliability is below a threshold may simply be determined as low-reliability labels, or labels whose reliability changes significantly over time may be determined as low-reliability labels. Furthermore, a combination of multiple conditions may be used to determine whether a label is low-reliability. In this case, labels that satisfy a combination of two or more of the multiple conditions described above are determined as low-reliability labels.

[0049] In this way, the low-reliability label area detection unit 52 functions as a determination unit (first determination unit) that determines whether an object with a low-reliability label is included among multiple objects in real space. Objects with a low-reliability label include objects recognized with a label indicating a recognition failure, such as unknown / unlabeled (the label in (1) above). An object recognized with a label indicating a recognition failure is an object for which no label indicating any object was recognized in the segmentation of the 2D image, i.e., an unknown object.

[0050] In addition to objects recognized with a label indicating a recognition failure, it is possible to determine as unknown objects at least any of the following: objects recognized with a generic label indicating a broad category such as Other Furniture (the label (2) above); objects recognized with a label where the labels in the same area change significantly over time and it is unclear which recognition result is correct (the label (3) above); objects recognized with a label whose reliability is below a threshold; and objects recognized with a label whose reliability changes significantly over time.

[0051] FIG. 7 is a flowchart illustrating the low-reliability label area detection process.

[0052] In step S1, the low-reliability labeled region detection unit 52 selects any one region (region x) from the region list. The region list is a list of segmented regions in the 3D spatial recognition map. The following process is performed for each region in the 3D spatial recognition map.

[0053] In step S2, the low-reliability label area detection unit 52 determines whether the label of area x is a label indicating a recognition failure, such as unknown / unlabeled. If it is determined in step S2 that the label of area x is a label indicating a recognition failure, the process proceeds to step S3.

[0054] In step S3, the low-reliability label area detection unit 52 adds area x to the reliability label area ID list, which is a list of IDs of areas to which low-reliability labels are set.

[0055] In step S4, the low-reliability label area detection unit 52 determines whether all areas have been selected. If it is determined in step S4 that all areas have not been selected, the process returns to step S1, and the same process is repeated for the next area.

[0056] On the other hand, if it is determined in step S2 that the label of region x is not a label indicating a recognition failure, then in step S5, the low-reliability label region detection unit 52 determines whether the label of region x is a label indicating a broad category such as "Other Furniture." If it is determined in step S5 that the label of region x is a label indicating a broad category, the process proceeds to step S3. After region x is added to the low-reliability label region ID list, subsequent processes are performed.

[0057] If it is determined in step S5 that the label of area x is not a label indicating a broad category, then in step S6, the low-reliability label area detection unit 52 determines whether the reliability of the label of area x is equal to or less than a threshold. In this example, the low-reliability label determination is performed based on whether the reliability is equal to or less than a threshold, rather than on the magnitude of the time-series change of the label. If it is determined in step S6 that the reliability of the label of area x is equal to or less than the threshold, the process proceeds to step S3. After area x is added to the low-reliability label area ID list, the subsequent processes are performed.

[0058] If it is determined in step S6 that the reliability of the label of area x is not below the threshold, that is, if it is determined that the label of area x is not a label indicating a recognition failure, a label indicating a general category, or a label with a reliability below the threshold, the label is not added to the low-reliability label area ID list, and processing from step S4 onwards is carried out.

[0059] If it is determined in step S4 that all areas have been selected, the processing in Fig. 7 ends. Information on the low-reliability label area ID list generated by the above processing is supplied from the low-reliability label area detection unit 52 to the label correspondence relationship acquisition unit 53 in Fig. 6.

[0060] Label Correspondence Acquisition Unit The label correspondence acquisition unit 53 generates label correspondence information based on the multimodal label recognition results supplied from the multimodal label recognition unit 51 and the camera position at the same time as the time the multimodal information was acquired.

[0061] For example, when it is determined based on camera position information and a 3D spatial recognition map that area B, which has been determined to be a low-reliability label area, is within the camera's field of view, and label recognition result A is obtained based on other modal information acquired at the same time, label correspondence information is generated indicating that label recognition result A corresponds to area B.

[0062] Since each region in the 3D spatial recognition map is segmented, the label correspondence information is expressed in the form of [“region ID”, “other-modal label recognition result”]. The label correspondence information is information that indicates the correspondence relationship between the other-modal label recognition result and the label recognition result of which region in the 3D spatial recognition map. Information in the label correspondence list, which is a list of label correspondence information, is supplied to the label update unit 54.

[0063] The label correspondence information registered in the label correspondence list is managed by linking it to the entire 3D spatial recognition map, rather than to the frame used to determine the correspondence. This makes it possible to complete labels at any time after label recognition based on multimodal information has been achieved.

[0064] Label Update Unit The label update unit 54 updates the low-reliability labels of the 3D spatial recognition map with the corresponding multi-modal label recognition results, using the label correspondence information registered in the label correspondence list supplied from the label correspondence acquisition unit 53. The 3D spatial recognition map with updated labels is a 3D spatial recognition map with improved label accuracy.

[0065] FIG. 8 is a flowchart illustrating the label update process.

[0066] In step S11, the label update unit 54 selects one piece of label correspondence information from the label correspondence list. For example, [“area ID”, “other-modal label recognition result”] = [“area x”, “label a”] is acquired. “Area x” identified by “area ID” is an area to which a low-confidence label has been set. “Label a” is the other-modal label recognition result.

[0067] In step S12, the label update unit 54 updates the label by changing the label of "area x" in the 3D space recognition map to "label a."

[0068] In step S13, the label update unit 54 determines whether or not all of the label correspondence information has been selected. If it is determined in step S13 that all of the label correspondence information has not been selected, the process returns to step S11, and the same process is repeated for the next label correspondence information.

[0069] If it is determined in step S13 that all label correspondence information has been selected, the process in Fig. 8 ends. As a result, the labels of the low-reliability labeled areas (all meshes) identified by the respective area IDs registered in the label correspondence list are updated with the corresponding multimodal label recognition results.

[0070] The label update by the label update unit 54 does not need to be performed at the same cycle as other processes. For example, the label update by the label update unit 54 is performed when the label correspondence relationship list is updated.

[0071] <<Label Completion Using Human Action Recognition>> <Flow of Label Completion> Human pose information is used as multimodal information. In this example, the actions of a person such as a user are recognized based on the human pose information, and object label recognition is performed using the relationship between the action recognition result and the object in space. The label recognition result using the relationship between the action recognition result and the object in space corresponds to the above-mentioned multimodal label recognition result.

[0072] FIG. 9 is a diagram showing the flow of label completion using human action recognition.

[0073] The space shown in the upper left of Fig. 9 shows an example of a space in a 3D spatial recognition map. When generating the 3D spatial recognition map, the floor of the room, the rug laid on the floor, the side table placed on the floor, and the cup placed on the side table are recognized, and labels indicating "Floor," "Mat," "Table," and "Cup" are set for each object (for the object's area). As shown in speech bubble #1, the category of object O, which has a cube shape and is placed on the rug, cannot be recognized, and a label indicating "Unlabeled" is set.

[0074] The lower left of Figure 9 shows an example of a person's movement. For example, skeleton estimation is performed based on 2D images of each frame, and skeleton point information at each time is acquired as multi-modal information. Furthermore, the person's movement is recognized based on the skeleton point information acquired as multi-modal information. The 2D images used for skeleton estimation may be acquired by the AR display device 1 or an external device. The person whose skeleton is to be estimated may be a user using an AR application, or another person in the same space as the user.

[0075] 9, as shown in speech bubble #2, it is recognized that a person is sitting on object O, which is labeled "Unlabeled," and the category of object O used in the action of "sitting" is recognized as a "chair." In this way, the object label recognition result that utilizes the relationship between the action recognition result and objects in space is used to complement the label of object O, as shown by the white arrow.

[0076] The right side of Fig. 9 shows an example of a space in a 3D spatial recognition map generated by completing the label. In the example on the right side of Fig. 9, a label indicating "Chair" is set for an object O.

[0077] In this way, even for furniture with shapes that are difficult to label with high reliability from 2D images alone, it is possible to detect scenes in which people are actually using the furniture and estimate the furniture category from how it is used. When using AR applications in a private home, it is expected that the same space will be photographed repeatedly. By photographing the same space repeatedly, it is possible to obtain footage of scenes in which the user is using the furniture.

[0078] <Configuration of Information Processing Unit> Fig. 10 is a block diagram showing an example configuration of the information processing unit 11. Of the configuration shown in Fig. 10, the same components as those described with reference to Fig. 4 are denoted by the same reference numerals. Duplicate descriptions will be omitted as appropriate.

[0079] In the example of Fig. 10, a human pose acquisition unit 101 is provided as a component corresponding to the multimodal information acquisition unit 24 in Fig. 4. The human pose acquisition unit 101 performs skeleton estimation on the 2D image supplied from the image acquisition unit 21, and acquires skeleton point information, which is human pose information. The human pose information acquired by the human pose acquisition unit 101 is output to the label accuracy improvement processing unit 25. A known method used in Kinect (trademark) or the like can be used as a method for acquiring skeleton point information. Information on feature points other than skeleton points may also be acquired as human pose information.

[0080] The label accuracy improvement processing unit 25 updates the low-reliability labels in the 3D space recognition map based on the camera position information supplied from the self-position estimation unit 22, the 3D space recognition map supplied from the 3D space recognition map generation unit 23, and the human pose information supplied from the human pose acquisition unit 101. Information on the 3D space recognition map whose labels have been updated by the label accuracy improvement processing unit 25 is output to the 3D Scene Graph generation unit 26.

[0081] <Configuration and Operation of Label Accuracy Improvement Processing Unit> Fig. 11 is a block diagram showing a configuration example of the label accuracy improvement processing unit 25 in Fig. 10. Of the configuration shown in Fig. 11, the same components as those described with reference to Fig. 6 are denoted by the same reference numerals. Duplicate descriptions will be omitted as appropriate.

[0082] 11, a pose label recognition unit 111 is provided as a component corresponding to the multimodal label recognition unit 51 in Fig. 6. The human pose information output from the human pose acquisition unit 101 is input to the pose label recognition unit 111.

[0083] Pose Label Recognition Unit The pose label recognition unit 111 recognizes how a person uses an object (the person's actions) based on human pose information. Human pose information is information acquired based on 2D images. The pose label recognition unit 111 functions as a recognition unit that recognizes human actions based on the 2D images used to acquire the human pose information. Various actions such as "walking," "sitting," "standing," and "grabbing" are recognized by the pose label recognition unit 111. Known methods such as Skeleton-Based Action Recognition can be used to recognize human actions based on skeleton point information. The pose label recognition unit 111 is provided with an inference model that inputs skeleton point information and outputs information on human actions.

[0084] Furthermore, the pose label recognition unit 111 recognizes the category of an object based on how the object is used (performs object label recognition). For example, if a person is performing the action of "sitting" near a certain object, the category of the object is recognized as a "chair." The pose label recognition unit 111 outputs information on the pose label recognition result, which is the result of recognizing the object category based on the person pose information. The pose label recognition result output from the pose label recognition unit 111 is supplied to the label correspondence acquisition unit 53 and the label update unit 54.

[0085] Label Correspondence Acquisition Unit The label correspondence acquisition unit 53 generates label correspondence information based on the pose label recognition results supplied from the pose label recognition unit 111 and the camera position at the same time as the human pose information was acquired.

[0086] 12 is a flowchart illustrating the processing of the pose label recognition unit 111 and the label correspondence acquisition unit 53. The processing of steps S101 to S103 is the processing of the pose label recognition unit 111, and the processing of steps S104 to S106 is the processing of the label correspondence acquisition unit 53.

[0087] In step S101 , the pose label recognition unit 111 recognizes the movement of a person based on the person pose information supplied from the person pose acquisition unit 101 .

[0088] In step S102, the pose label recognition unit 111 determines whether the person's action is "sitting." If it is determined in step S102 that the person's action is "sitting," then in step S103, the pose label recognition unit 111 estimates that there is a chair nearby. Categories of other objects are estimated in a similar manner. For example, if the action of a person placing their hands on a flat surface is recognized, the object used in that action is recognized as a "table."

[0089] In step S104, the label correspondence acquisition unit 53 selects, from the low-reliability label regions, region A, which is a low-reliability label region that is closest to the person who performed the "sit" action.

[0090] In step S105, the label correspondence acquisition unit 53 determines whether the positional relationship between the person who performed the "sit" action and the area A satisfies a predetermined condition. For example, if a skeleton point of the person's lower body is in contact with a low-reliability labeled area, or if the distance between the low-reliability labeled area and a skeleton point of the person's lower body is within a distance dist as a threshold, it is determined that the predetermined condition is met. The coordinates of the n-th skeleton point of the lower body of the human body are defined as bone n = (x bn , y bn , z bn ), the k-th vertex in the low-confidence label region is unstable k = (x uk , y uk , z uk ), the condition of the positional relationship is expressed by the following equation (1).

[0091] If it is determined in step S105 that the positional relationship between the person who performed the "sit" action and area A satisfies a predetermined condition, in step S106, the label correspondence acquisition unit 53 generates label correspondence information ["area A", "chair"] and registers it in the label correspondence list.

[0092] On the other hand, if it is determined in step S102 that the person's action is not "sitting," or if it is determined in step S105 that the positional relationship between the person who performed the "sitting" action and area A does not satisfy the specified conditions, the process returns to step S101 and the same process is repeated.

[0093] Here, an example of heuristic determination has been described, in which if a person is performing the motion of "sitting" near a low-reliability label area, the category of the object corresponding to that area is determined to be a "chair." However, the object category may also be determined using an inference model generated by machine learning. In this case, for example, an inference model generated by machine learning based on the motion corresponding to each object and the positional relationship between the person and the object is prepared in the pose label recognition unit 111.

[0094] In order to increase the robustness of the determination, the pose label recognition result may be confirmed when the time series of the reliability of the pose label recognition result exceeds a threshold value for a predetermined period of time.

[0095] In this way, the label correspondence acquisition unit 53 determines a category according to the person's movement represented by the pose label recognition result supplied from the pose label recognition unit 111. The label correspondence acquisition unit 53 also determines the positional relationship between the person and an object to which a low-confidence label is set in accordance with the above conditions. The label correspondence acquisition unit 53 functions as a determination unit (second determination unit) that determines the category of the object according to the person's movement and the positional relationship between the person and an unknown object.

[0096] Label Update Unit The label update unit 54 updates the low-confidence labels of the 3D space recognition map with the corresponding pose label recognition results, using the label correspondence information registered in the label correspondence list supplied from the label correspondence acquisition unit 53. The 3D space recognition map with updated labels is a 3D space recognition map with improved label accuracy.

[0097] FIG. 13 is a flowchart illustrating the label update process.

[0098] In step S111, the label update unit 54 selects one piece of label correspondence information from the label correspondence list. For example, [“Area ID”, “Pose label recognition result”] = [“Area A”, “Chair”] is acquired. “Area A” identified by the “Area ID” is an area to which a low-confidence label has been set. “Chair” is the pose label recognition result.

[0099] In step S112, the label update unit 54 updates the label by changing the label of "area A" in the 3D space recognition map to "chair."

[0100] In step S113, the label update unit 54 determines whether or not all of the label correspondence information has been selected. If it is determined in step S113 that all of the label correspondence information has not been selected, the process returns to step S111, and the same process is repeated for the next label correspondence information.

[0101] If it is determined in step S113 that all label correspondence information has been selected, the process in Fig. 13 ends. As a result, the labels of the low-confidence labeled regions identified by the respective region IDs registered in the label correspondence list are updated with the corresponding pose label recognition results.

[0102] In this way, when an object with a low-confidence label is found among the objects for which label recognition has been performed by segmentation based on a 2D image, the label update unit 54 updates the low-confidence label with the pose label recognition result (second label). The pose label recognition result used to update the low-confidence label is selected based on the result of determining the object category based on the person's movement and the positional relationship between the person and the unknown object. When an object with a low-confidence label is found among the objects for which label recognition has been performed by segmentation based on a 2D image, the label update unit 54 has a function of updating the low-confidence label based on the result of determining the object category based on the person's movement and the positional relationship between the person and the unknown object.

[0103] The above process makes it possible to correctly identify the category of furniture that is difficult to recognize from the results of segmentation of 2D images, thereby obtaining a 3D spatial recognition map that includes labels indicating the correct furniture category.

[0104] Furthermore, once the labels of the 3D spatial recognition map are updated, the updated, correct label information can be used in subsequent processing. For example, if a certain piece of furniture is recognized as a chair based on human pose information, it is possible to make an AR character act using that chair even if no one uses that chair after that.

[0105] <<Configuration Example>> <Hardware Configuration of AR Display Device> FIG. 14 is a block diagram showing an example of the hardware configuration of the AR display device 1. As shown in FIG.

[0106] As shown in FIG. 14, the AR display device 1 is configured by connecting a camera 202, a sensor 203, a communication unit 204, a display unit 205, and a memory 206 to a control unit 201.

[0107] The control unit 201 is configured with a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc. The control unit 201 executes programs stored in the ROM and memory 206, and controls the overall operation of the AR display device 1. The control unit 201 executes a predetermined program to realize the information processing unit 11 having the above-described functions.

[0108] Furthermore, the control unit 201 executes a predetermined program to implement the display control unit 12. The display control unit 12 controls the display unit 205 to present an AR character to the user. For example, an image of the AR character performing planned actions based on the 3D spatial recognition map with updated labels is displayed under the control of the display control unit 12.

[0109] The camera 202 captures an image in front of the AR display device 1 and outputs an RGB image to the control unit 201 .

[0110] The sensor 203 is composed of various sensors such as a ToF sensor. The ToF sensor constituting the sensor 203 measures the distance to each position in front of the AR display device 1 and outputs a depth image to the control unit 201. The sensor 203 appropriately includes various sensors such as an acceleration sensor, a gyro sensor, a positioning sensor, a microphone, and a temperature sensor. In this case, the measurement results of the acceleration sensor, gyro sensor, and positioning sensor are output to the control unit 201. The measurement results of the acceleration sensor, gyro sensor, and positioning sensor are used to estimate the position and orientation of the AR display device 1, etc.

[0111] The communication unit 204 is configured by a communication module such as a wireless LAN. The communication unit 204 communicates with external devices via a network and transmits data supplied from the control unit 201 to the external devices. The communication unit 204 also receives data transmitted from external devices and outputs the data to the control unit 201. External devices include not only devices present in the environment of the AR display device 1 but also servers connected via the Internet.

[0112] The display unit 205 is provided, for example, in a lens portion of the AR display device 1, which is a glasses-type wearable device. The display unit 205 displays various information including the AR character according to the control of the display control unit 12, and presents it to the user.

[0113] The memory 206 is a storage medium such as a flash memory, etc. The memory 206 stores various data such as programs executed by the CPU of the control unit 201.

[0114] <Example of Device> Instead of a glasses-type wearable device, a video see-through type HMD or a mobile terminal such as a smartphone may be used as a display device for an AR character.

[0115] When a video-transparent HMD is used, the image of the AR character is displayed superimposed on a video see-through image of the scenery in front of the user, captured by a camera installed in the HMD. A display is provided in front of the eyes of the user wearing the HMD, which displays the image of the AR character superimposed on the image captured by the camera.

[0116] When a smartphone is used, the image of the AR character is superimposed on the image of the scenery in front of the smartphone, captured by a camera on the back of the smartphone. A display is provided on the front of the smartphone. Various devices, such as a tablet terminal or a television receiver, can be used as a display device for the AR character.

[0117] <Example in which display of AR character is controlled by external device> The information processing section 11 described with reference to FIGS. 4 and 10 may be realized by a device external to the AR display device 1.

[0118] Fig. 15 is a diagram showing an example of an information processing system including an external device. The information processing system shown in Fig. 15 is configured by connecting an AR display device 1 and a server 301 via a network such as the Internet. In the information processing system shown in Fig. 15, the AR display device 1 functions as an edge device and displays an AR character under the control of the server 301.

[0119] The information processing unit 11 of the server 301 receives image data transmitted from the AR display device 1 and performs the various processes described above, such as updating the labels of the 3D space recognition map. The server 301 plans the AR character's behavior based on the 3D space recognition map with updated labels, and generates an image of the AR character. The image generated by the server 301 is transmitted to the AR display device 1 and used to display the AR character. In this way, the server 301 functions as an information processing device on the cloud.

[0120] 4 and 10 are realized in the AR display device 1, and the other functional units are realized in the server 301. In place of the server 301, the display of the AR character may be controlled by an information processing device such as a personal computer installed in the user's home or the user's smartphone.

[0121] <<Modifications>>> Information other than person pose information may be used as the multimodal information. For example, information on the recognition results of a person's voice may be used as the multimodal information. In this case, a voice uttered by a person near an object to which a low-reliability label has been assigned is recognized, and the category of the object is recognized based on the recognition result. For example, an AR character near an object to which a low-reliability label has been assigned may ask a question to the user, such as "What is this?", and the user may respond with "It's a chair," thereby making it possible to recognize the category of the object to which a low-reliability label has been assigned.

[0122] Furthermore, in the above, information on the pose of a person is used to recognize the category of an object, but the movement of an animal such as a dog or cat may be recognized based on the pose of the animal, and the category of the object may be recognized based on the movement of the animal. Instead of an animal, the movement of a robot capable of autonomous action may be recognized and used to recognize the category of the object. In this way, information on the pose of various moving objects such as people, animals, and robots can be used to recognize the category of an object.

[0123] This technology can be applied to the generation of 3D spatial recognition maps for various autonomous mobile objects that behave autonomously. For example, this technology can be used in robots equipped with various sensors such as RGB cameras and ToF sensors. In this case, segmentation is performed on 2D images acquired by the RGB camera and ToF sensor mounted on the robot, and a 3D spatial recognition map is generated. The labels of objects with low confidence labels in the 3D spatial recognition map are updated with labels indicating categories corresponding to the recognition results of human actions.

[0124] Although the above description concerns the case where the object to be updated is furniture, the categories of various objects present in the space, such as electrical appliances, books, tableware, and writing implements, may also be recognized based on other modal information, and the labels may be updated accordingly.

[0125] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware, or into a general-purpose personal computer, etc.

[0126] 16 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program. For example, the control unit 201 in FIG. 14 is configured by a computer having the configuration shown in FIG.

[0127] A CPU (Central Processing Unit) 1001 , a ROM (Read Only Memory) 1002 , and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004 .

[0128] An input / output interface 1005 is also connected to the bus 1004. An input unit 1006 including a keyboard, a mouse, etc., and an output unit 1007 including a display, a speaker, etc. are connected to the input / output interface 1005. In addition, a storage unit 1008 including a hard disk, a nonvolatile memory, etc., a communication unit 1009 including a network interface, etc., and a drive 1010 that drives removable media 1011 are also connected to the input / output interface 1005.

[0129] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0130] The program executed by the CPU 1001 is installed in the storage unit 1008 by being recorded on, for example, a removable medium 1011 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.

[0131] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0132] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.

[0133] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0134] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.

[0135] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0136] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0137] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0138] <Examples of Combinations of Configurations> The present technology can also have the following configurations.

[0139] (1) An information processing device comprising: a 3D space recognition map generation unit that performs segmentation on a first 2D image and generates a 3D space recognition map in which a first label indicating a category of each of a plurality of objects in a real space is set in each region; a first determination unit that determines whether an unknown object whose first label was not recognized in the segmentation is included in the plurality of objects; a second determination unit that determines a category of the object according to a movement of a person appearing in a second 2D image and a positional relationship between the person and the unknown object; and a label update unit that updates the first label of the unknown object based on a determination result of the category of the object according to the movement of the person and a determination result of the positional relationship when the unknown object is included in the plurality of objects. (2) The information processing device described in (1), wherein the first determination unit further determines, as the unknown object, at least one of the object to which the first label indicating a general category is set, the object whose category indicated by the first label changes, and the object whose reliability of the first label is equal to or lower than a threshold. (3) The information processing device according to (1) or (2), further comprising a recognition unit that recognizes the person's movement based on the second 2D image. (4) The information processing device according to (3), wherein the second determination unit generates correspondence information indicating a correspondence between the unknown object and a category of the object corresponding to the person's movement relative to the unknown object. (5) The label update unit updates the first label of the unknown object using a second label indicating the category of the object corresponding to the unknown object in the correspondence information. (6) The information processing device according to any of (3) to (5), wherein the recognition unit recognizes the person's movement based on a result of skeletal estimation for the second 2D image and recognizes the category of the object corresponding to the recognized person's movement. (7) The information processing device according to (6), wherein the recognition unit determines the category of the object corresponding to the person's sitting movement as a chair.(8) The information processing device according to any one of (1) to (7), further comprising: a 3D Scene Graph generation unit that generates a 3D Scene Graph, which is graph-structured information in which nodes representing the objects including the unknown object are connected by edges that represent positional relationships of the objects, based on the 3D space recognition map in which the first labels have been updated; and a behavior planning unit that performs behavior planning for an autonomous moving body based on the 3D Scene Graph. (9) The information processing device according to any one of (1) to (8), wherein the second 2D image is an image captured after the first 2D image used to generate the 3D space recognition map. (10) An information processing method in which an information processing device performs segmentation on a first 2D image, generates a 3D spatial recognition map in which a label indicating the category of each of a plurality of objects in real space is set in each region, determines whether an unknown object whose label was not recognized in the segmentation is included in the plurality of objects, determines the category of the object according to the movement of a person appearing in a second 2D image and the positional relationship between the person and the unknown object, and if the unknown object is included in the plurality of objects, updates the label of the unknown object based on the determination result of the category of the object according to the movement of the person and the determination result of the positional relationship. (11) A program that causes a computer to execute the following processes: segmenting a first 2D image, generating a 3D spatial recognition map in which a label indicating the category of each of a plurality of objects in real space is set in each region; determining whether an unknown object whose first label was not recognized in the segmentation is included in the plurality of objects; determining the category of the object according to the movement of a person appearing in a second 2D image and the positional relationship between the person and the unknown object; and updating the label of the unknown object based on the determination result of the category of the object according to the movement of the person and the determination result of the positional relationship if the unknown object is included in the plurality of objects.

[0140] DESCRIPTION OF SYMBOLS 1 AR display device, 11 Information processing unit, 21 Image acquisition unit, 22 Self-position estimation unit, 23 3D space recognition map generation unit, 24 Other modal information acquisition unit, 25 Label accuracy improvement processing unit, 26 3D Scene Graph generation unit, 27 Action planning unit, 51 Other modal label recognition unit, 52 Low-reliability label area detection unit, 53 Label correspondence acquisition unit, 54 Label update unit, 101 Human pose acquisition unit, 111 Pose label recognition unit

Claims

1. An information processing device comprising: a 3D space recognition map generation unit that performs segmentation on a first 2D image and generates a 3D space recognition map in which a first label indicating the category of each of a plurality of objects in real space is set in each region; a first determination unit that determines whether an unknown object whose first label was not recognized in the segmentation is included in the plurality of objects; a second determination unit that determines the category of the object according to the movement of a person appearing in a second 2D image and the positional relationship between the person and the unknown object; and a label update unit that updates the first label of the unknown object based on the determination result of the category of the object according to the movement of the person and the determination result of the positional relationship when the unknown object is included in the plurality of objects.

2. The information processing device of claim 1, wherein the first determination unit further determines, as the unknown object, at least one of the object to which the first label indicating a general category is set, the object whose category indicated by the first label changes, and the object whose reliability of the first label is below a threshold.

3. The information processing device according to claim 1, further comprising a recognition unit that recognizes the movement of the person based on the second 2D image.

4. The information processing device according to claim 3, wherein the second determination unit generates correspondence information indicating a correspondence between the unknown object and a category of the object according to the action of the person relative to the unknown object.

5. The information processing device according to claim 4, wherein the label updating unit updates the first label of the unknown object with a second label indicating a category of the object corresponding to the unknown object in the correspondence information.

6. The information processing device according to claim 3, wherein the recognition unit recognizes the movement of the person based on the result of skeletal estimation for the second 2D image, and recognizes the category of the object according to the recognized movement of the person.

7. The information processing device according to claim 6, wherein the recognition unit determines the category of the object according to the person's sitting action as a chair.

8. The information processing device according to claim 1, further comprising: a 3D Scene Graph generation unit that generates a 3D Scene Graph, which is information with a graph structure in which nodes representing each of the objects including the unknown object are connected by edges that represent the positional relationships of the objects, based on the 3D spatial recognition map in which the first label has been updated; and a behavior planning unit that performs behavior planning for an autonomous moving body based on the 3D Scene Graph.

9. The information processing device according to claim 1, wherein the second 2D image is an image captured after the first 2D image used to generate the 3D spatial awareness map.

10. An information processing method in which an information processing device performs segmentation on a first 2D image, generates a 3D spatial recognition map in which a label indicating the category of each of a plurality of objects in real space is set in each region, determines whether an unknown object whose label was not recognized in the segmentation is included in the plurality of objects, determines the category of the object according to the movement of a person appearing in a second 2D image and the positional relationship between the person and the unknown object, and if the unknown object is included in the plurality of objects, updates the label of the unknown object based on the determination result of the object category according to the movement of the person and the determination result of the positional relationship.

11. A program that causes a computer to execute the following processes: segmenting a first 2D image, generating a 3D spatial recognition map in which a label indicating the category of each of a plurality of objects in real space is set in each region; determining whether an unknown object whose label was not recognized in the segmentation is included in the plurality of objects; determining the category of the object according to the movement of a person appearing in a second 2D image and the positional relationship between the person and the unknown object; and if the unknown object is included in the plurality of objects, updating the label of the unknown object based on the determination result of the object category according to the movement of the person and the determination result of the positional relationship.

Citation Information

Patent Citations

  • Identification system and identification method

    JP2019128804A

  • Display control device, display control method, and program

    WO2022224522A1

  • Information processing device, information processing method, and recording medium

    WO2024009748A1