Information processing device, information processing method, and program
A UI for specifying conditions in a virtual space facilitates easy content retrieval and avatar-based behavior replication, addressing the challenge of finding content and avatars in large datasets.
Patent Information
- Application Number
- PCT/JP2025/017302
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-27
- Filing Date
- 2025-05-13
- Publication Date
- 2025-12-04
AI Technical Summary
Users face difficulty in finding desired content from a large amount of content, and when actions of people are reproduced as avatars, it becomes challenging to identify the desired avatar from many similar ones.
A UI is provided for specifying conditions to narrow down content in a virtual space, and when selected, an avatar reproduces the behavior of a person associated with the content.
Enables easy retrieval of desired content and behavior replication of content creators, enhancing user interaction and efficiency in content search and retrieval.
Smart Images

Figure JP2025017302_04122025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present technology relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program that enable a user to easily search for desired content from among a large amount of content.
[0002] Conventionally, when equipment trouble occurs at a manufacturing site or the like, a supervisor in a remote location away from the manufacturing site (for example, an expert in an office or at home) may instruct workers at the manufacturing site on how to resolve the equipment trouble, etc. In such cases, the supervisor can provide instructions and support remotely while checking a 3D model of the equipment or images taken by a worker at the manufacturing site using a camera.
[0003] Patent document 1 describes a technology that synchronizes the real space (AR space) where the worker is located with the VR space where the supporter is located, and virtually places and presents photographed images, etc., within the spaces where the worker and instructor are located.
[0004] International Publication No. 2023 / 181887
[0005] When a large amount of content, such as images and memos, is placed in a space, it is difficult for users to find the desired content from among the large amount of content. Also, when the actions of people related to the content are reproduced as the actions of avatars, many similar avatars are placed in the space, making it difficult for users to find the desired avatar from among the large amount of avatars.
[0006] The present technology has been made in view of such circumstances, and enables users to easily search for desired content from among a large amount of content.
[0007] An information processing device according to one aspect of the present technology includes a UI for specifying conditions for narrowing down a plurality of pieces of content virtually arranged in a space, and an output control unit that presents to a user the pieces of content that satisfy the conditions from among the plurality of pieces of content arranged in the space, and when the content presented to the user is selected by the user, the output control unit presents to the user an avatar that reproduces the behavior of a person associated with the selected content.
[0008] An information processing method according to one aspect of the present technology includes: presenting to a user a UI for specifying conditions for narrowing down a plurality of pieces of content virtually arranged in a space; presenting to the user the pieces of content that satisfy the conditions from among the plurality of pieces of content arranged in the space; and, when the content presented to the user is selected by the user, presenting to the user an avatar that reproduces the behavior of a person associated with the selected piece of content.
[0009] A program according to one aspect of the present technology causes a computer to execute processing including: presenting to a user a UI for specifying conditions for narrowing down a plurality of pieces of content virtually arranged in a space; presenting to the user the pieces of content that satisfy the conditions from among the plurality of pieces of content arranged in the space; and, when the user selects one of the pieces of content presented to the user, presenting to the user an avatar that reproduces the behavior of a person associated with the selected piece of content.
[0010] In one aspect of the present technology, a UI for specifying conditions for narrowing down a plurality of pieces of content virtually arranged in a space and a piece of content that satisfies the conditions among the plurality of pieces of content arranged in the space are presented to a user, and when the user selects the piece of content presented to the user, an avatar that reproduces the behavior of a person related to the selected piece of content is presented to the user.
[0011] 1 is a first diagram illustrating an overview of a three-dimensional space sharing system according to the present technology. FIG. 2 is a second diagram illustrating an overview of a three-dimensional space sharing system according to the present technology. FIG. 3 is a diagram illustrating an example of presentation of a captured image. FIG. 4 is a diagram illustrating an example configuration of a three-dimensional space sharing system according to an embodiment of the present technology. FIG. 5 is a diagram illustrating an example of a screen displayed on a device. FIG. 6 is a diagram illustrating an example of an image including a detail window. FIG. 7 is a diagram illustrating an example of a screen when content presented on a device is narrowed down based on a condition related to a range in a space. FIG. 8 is a diagram illustrating a use case in which a worker creates a trouble report using a device. FIG. 9 is a first diagram illustrating a use case in which a supervisor creates a trouble report using a device. FIG. 10 is a second diagram illustrating a use case in which a supervisor creates a trouble report using device 1. FIG. 11 is a diagram illustrating a use case in which a search for a trouble report is based on a condition related to a range in a space. FIG. 12 is a diagram illustrating an example of a screen when content is presented in the form of a point object. FIG. 13 is a diagram illustrating another use case in which a search for a trouble report is based on a condition related to a range in a space. FIG. 14 is a diagram illustrating a use case in which a search for a trouble report is based on a condition related to the creation date and time of the content. FIG. 15 is a diagram illustrating a use case in which a search for a trouble report is based on a condition related to the person responsible for the trouble. FIG. 16 is a diagram illustrating a use case in which a search for a trouble report is based on a plurality of conditions. 1 is a diagram illustrating a use case in which a user learns past responses to problems. FIG. 1 is a diagram illustrating a flow of recording behavioral data. FIG. 2 is a diagram illustrating the extraction of behavioral data. FIG. 3 is a diagram illustrating the relationship between the trajectory of a worker's movement at a site and the extracted behavioral data. FIG. 4 is a block diagram illustrating an example configuration of a smartphone. FIG. 5 is a block diagram illustrating an example configuration of a PC. FIG. 6 is a block diagram illustrating an example configuration of an HMD. FIG. 7 is a flowchart illustrating the process by which the three-dimensional space sharing system of the present technology extracts behavioral data. FIG. 8 is a first flowchart illustrating the process by which the three-dimensional space sharing system of the present technology automatically selects representative content and creates a trouble report.29 is a flowchart illustrating an example of a configuration of computer hardware.
[0012] Hereinafter, embodiments of the present technology will be described in the following order: 1. Overview of the present technology 2. Use cases 3. Configuration and operation of each device
[0013] 1. Overview of the Present Technology FIGS. 1 and 2 are diagrams showing an overview of a three-dimensional space sharing system according to the present technology.
[0014] The three-dimensional space sharing system in FIG. 1 is a system that shares content placed in a VR (Virtual Reality) space and a real space (AR (Augmented Reality) space) between the two spaces.
[0015] When equipment trouble occurs at a manufacturing site or the like, an instructor in a remote location (such as an office or home) accesses a VR space that recreates the manufacturing site by wearing device 1, which may be a VR head-mounted display (HMD), as shown on the left side of Figure 1. Similarly, an operator at the manufacturing site accesses the AR space by wearing device 2, which may be an AR HMD, as shown on the right side of Figure 1.
[0016] The instructor may access the VR space by using a device 1 configured as, for example, a PC, as shown on the left side of Fig. 2. The worker may access the AR space by using a device 2 configured as, for example, a smartphone, as shown on the right side of Fig. 2.
[0017] A user (instructor) in the VR space is presented with captured images (still images or video images) taken by a worker at a manufacturing site using, for example, a camera mounted on device 2. The instructor can remotely provide instructions and support to resolve equipment problems while checking the captured images. The user (worker) in the AR space operates the equipment and takes further photos according to the instructor's instructions.
[0018] FIG. 3 is a diagram showing an example of how a captured image is presented.
[0019] As shown on the left side of Figure 3, in the VR space, image objects POV1 to POV3 corresponding to photographed images of the equipment are virtually placed around a virtual object 10V corresponding to the equipment to be operated by the worker.
[0020] The virtual object 10V is, for example, a 3D model (3D CAD (Computer Aided Design) data, scan data, etc.) that indicates the three-dimensional shape of the equipment to be operated. The image objects POV1 to POV3 are rectangular planar virtual objects onto which captured images are projected, and are arranged in the VR space at positions and orientations that correspond to the capture positions and capture orientations of the captured images.
[0021] In this way, the captured images are arranged in the form of image objects in the VR space.
[0022] 3, in the AR space, image objects POA1 to POA3 corresponding to captured images of the equipment are virtually arranged around a real object 10R, which is the equipment itself that is the target of operation by the worker. The image objects POA1 to POA3 are virtual objects with a rectangular planar shape onto which the captured images are projected, and are arranged in the AR space at positions and orientations that correspond to the shooting positions and orientations of the captured images.
[0023] In this way, the captured images are arranged in the AR space in the form of image objects.
[0024] The same photographed image is projected onto the image objects POA1 and POV1, the image objects POA2 and POV2, and the image objects POA3 and POV3. Therefore, by looking at the image objects POA1 to POA3, the instructor can confirm the state of the equipment from the same viewpoint as the worker.
[0025] The content placed in the space is not limited to photographs. Moving images such as movies, animations, and games, still images such as paintings, manga, and illustrations, text such as books, instruction manuals, and memos, and audio such as music and audiobooks may also be placed in the VR or AR space.
[0026] FIG. 4 is a diagram illustrating an example of the configuration of a three-dimensional space sharing system according to an embodiment of the present technology.
[0027] The three-dimensional space sharing system in Fig. 4 is composed of a device 1 used by a support user (instructor), a device 2 used by a site user (worker), a server 3, and a recording unit 4. In the three-dimensional space sharing system, device 1 and device 2 each communicate with server 3 via a network such as the Internet, either wired or wirelessly.
[0028] The device 1 is configured as an information processing device such as an AR or VR HMD, PC, smartphone, or tablet terminal. The device 1 acquires content data and a 3D model of the equipment to be operated from the server 3, and places the content and the 3D model of the equipment in a VR or AR space and presents it to the instructor. The device 1 estimates user position and orientation information indicating the position and orientation of the instructor's head and hands, and transmits the user position and orientation information to the server 3. The position and orientation of the instructor is expressed, for example, by three-dimensional coordinates or three-dimensional vectors with the position of the equipment to be operated as the origin.
[0029] The device 2 is configured as an information processing device such as an AR HMD, a smartphone, or a tablet terminal. The device 2 transmits a captured image of the equipment to be operated, obtained by using, for example, a camera mounted on the device 2 itself, together with camera position and orientation information indicating the capture position and capture orientation of the captured image, to the server 3. The capture position and capture orientation of the captured image are expressed, for example, by three-dimensional coordinates or a three-dimensional vector with the position of the equipment to be operated as the origin.
[0030] The device 2 acquires content data from the server 3, places the content in the AR space, and presents it to the worker. The device 2 estimates user position and orientation information indicating the positions and orientations of the worker's head and hands, and transmits the user position and orientation information to the server 3. The worker's position and orientation are expressed, for example, by three-dimensional coordinates or three-dimensional vectors with the position of the device to be operated as the origin.
[0031] The server 3 records the camera position and orientation information transmitted from the device 2 in the recording unit 4 as metadata corresponding to content such as captured images. The metadata corresponding to the content includes the creation date and time of the content (for example, the capture date and time of the captured image) in addition to the camera position and orientation information. The server 3 records the captured image captured by the device 2 in the recording unit 4 in association with the metadata. The camera position and orientation information can also be said to be content position and orientation information that indicates the position and orientation of the content in the VR space or AR space.
[0032] Furthermore, the server 3 records the history of the user position and orientation information transmitted from the device 1 or device 2 as behavior data indicating the user's behavior in the recording unit 4. The device 1 or device 2 can reproduce, as the behavior of the avatar, how the instructor or worker acted to solve the equipment trouble by moving the avatar based on the behavior data.
[0033] The recording unit 4 records the content and behavioral data supplied from the server 3 in the form of, for example, a trouble report. A trouble report is a compilation of images taken from the time a certain trouble occurs until the time the trouble is resolved and recorded behavioral data. The trouble report includes information indicating the title, a 3D model of the equipment to be operated, location information of the equipment, the content, metadata of the content, information indicating the order of the content, information indicating the person who handled the trouble, etc.
[0034] For example, a user who accessed the VR space or AR space from the time a problem occurred until the time the problem was resolved is registered in the problem report as a person who responded to the problem.
[0035] An instructor or worker can use device 1 or device 2 to give instructions or operate equipment while sharing trouble reports arranged in the space.
[0036] It should be noted that the content presented and the content layout may vary depending on the characteristics and usage of the device, rather than the exact same content being presented on device 1 and device 2. For example, device 1 presents a 3D model of the equipment to be operated, but because the equipment actually exists at the site where the worker is located, the 3D model of the equipment is not presented.
[0037] In the above-described three-dimensional space sharing system, when a large amount of content is placed in the VR space or the AR space, it is difficult for the user to find the desired content from among the large amount of content.
[0038] Therefore, the three-dimensional space sharing system of the present technology presents the user with a UI (User Interface) for specifying conditions to narrow down multiple pieces of content virtually placed in the space, and presents the user with content from the multiple pieces of content placed in the space that meets those conditions, thereby enabling the user to easily find the desired content from among a large amount of content.
[0039] FIG. 5 is a diagram showing an example of a screen displayed on the device.
[0040] The screen shown in Fig. 5 is a screen that provides a bird's-eye view of space Sp1 in which multiple pieces of content are arranged, and is a screen that is displayed on, for example, a PC serving as device 1 used by an instructor. In the example of Fig. 5, the pieces of content included in the trouble report for troubles A to C are arranged in space Sp1. In Fig. 5, the pieces of content included in the trouble report for trouble A are indicated by white rectangles, and the pieces of content included in the trouble report for trouble B are indicated by gray rectangles. The pieces of content included in the trouble report for trouble C are indicated by hatched rectangles. Note that the numbers written inside the rectangles representing each piece of content indicate the content number.
[0041] The user can narrow down the content presented on the device based on conditions related to, for example, at least one of the equipment to be operated, the location where the equipment is installed, the three-dimensional coordinates within space Sp1, the range within space Sp1, the date and time the content was created, and the person who will respond to the problem.
[0042] A pull-down menu PM1, which is an equipment designation UI for designating the equipment to be operated, is presented in the upper left portion of the screen in Fig. 5. The user can select (designate) from the pull-down menu PM1 the equipment trouble report for which the user wants to view. In the example in Fig. 5, the contents included in the trouble reports for troubles A to C that occurred in device α are arranged in space Sp1.
[0043] A time bar B1, which is a time specification UI for specifying a range of content creation dates and times, is presented at the bottom of the screen in Fig. 5. The user can specify the time at which they want to view content created by moving the black dots at the left and right ends of the time bar B1 to the left or right in the figure. In the example in Fig. 5, content created at times within a certain week is arranged in space Sp1.
[0044] In the upper right portion of the screen in FIG. 5 , icons I1 to I3 are displayed, indicating the responders for troubles A to C. Icon I1 indicates that one responder responded to the trouble using an HMD, and icon I2 indicates that another responder responded to the trouble using a smartphone. Icon I3 indicates that yet another responder created a trouble report. The user can specify which responder responded to the trouble and view the trouble report for that trouble by operating each of icons I1 to I3.
[0045] Icons I1 to I3 can be said to be a person designation UI for designating a person related to the content (trouble report).
[0046] A menu icon MI1 for setting options is presented to the right of the icon I3. For example, the user can select the menu icon MI1 and perform a predetermined operation to display a details window W1 on the screen, as shown in FIG.
[0047] The details window W1 presents the contents included in one trouble report arranged in numerical order. While the contents are arranged three-dimensionally in the center of the screen, the contents are arranged two-dimensionally in the details window W1. Lines connect the same contents between the space Sp1 and the details window W1. While checking the details of each content in the details window W1, the user can easily grasp the location in the space Sp1 where that content is arranged.
[0048] The date and time when the trouble occurred (date and time when the content was created) is displayed in the upper part of the details window W1.
[0049] FIG. 7 is a diagram showing an example of a screen when content presented on a device is narrowed down based on a condition related to the range within the space Sp1.
[0050] The screen in FIG. 7 is, for example, a screen that allows the space Sp1 to be viewed from a certain viewpoint within the space Sp1, and is a screen that is displayed on, for example, a VR HMD that serves as the device 1 used by the instructor.
[0051] As shown in FIG. 7, when a position in a space Sp1 is designated by the user, a pin Pi1 is presented at that position, and a sphere Sph1 centered at the position of the pin Pi1 is also presented.
[0052] The content (photographed image) placed inside the sphere Sph1 is presented in the form of, for example, an image object so that the user can check the content of the photographed image. In the example of Fig. 7, image objects PO11 to PO14 are placed inside the sphere Sph1.
[0053] On the other hand, the content placed outside the sphere Sph1 is presented in a simplified manner so that the user can understand that the content is placed there. In the example of Fig. 7, the content placed outside the sphere Sph1 is presented in the form of a point object.
[0054] The size of the point object varies depending on the distance from the pin Pi1. For example, the point object closer to the pin Pi1 is presented larger. A heat map showing the number of contents arranged at each position outside the sphere Sph1 by color may be displayed outside the sphere Sph1, thereby providing a simple presentation of the contents. The contents arranged inside the sphere Sph1 may be arranged and presented in a detail window.
[0055] The position of the pin Pi1 may be set by the user moving the pin Pi1, or may be set based on the user's line of sight. For example, if the user continues to direct their gaze at a certain piece of content for three seconds or more, the pin Pi1 is presented at the position of the content. The user's point of gaze may be estimated based on the user's gaze as well as the user's focus, and the pin Pi1 may be presented at the point of gaze.
[0056] The size of the sphere Sph1 is set, for example, by the user moving the spherical surface of the sphere Sph1.
[0057] In this way, the user can specify (range specification) the range in which they want to view the content placed therein by operating the pin Pi1 and the sphere Sph1. The sphere Sph1 and the pin Pi1 can be considered as a range specification UI for specifying a range within the space Sp1. Hereinafter, the range within the space specified by the range specification UI will also be referred to as a specified range.
[0058] If the desired content is hidden by the content in front, it is likely that the user will try to move the content in front, so when the user makes a gesture of waving their hand, the content located within the range of the hand movement and inside the sphere Sph1 may be simply presented.
[0059] Pins may be set at multiple positions within the space Sp1. Furthermore, a history of the pin positions and the specified range may be recorded. In this case, the method of setting the pin may also be recorded, for example, such as the pin being set because the user directed their gaze for three seconds or more. When pins are set at multiple positions, only the content within the specified range included in the user's field of view may be presented in the form of an image object, and other content may be presented simply.
[0060] The device may predict and pre-load content that the user is likely to look at based on the user's head and eye movements, and may also predict content that the user is likely to look at based on the user's hand movements and gestures.
[0061] When content is presented to the user by a PC, the user can specify a range in space using a mouse. For example, if the user moves the mouse cursor in a circle, a pin is set at the position of the content closest to the center of the circle, and the radius of the circle is set as the radius of the sphere.
[0062] As described above, a user can narrow down the content presented on the device and search for desired content by specifying conditions related to the range of the VR space or the AR space, etc. In the device 1 or the device 2 of the present technology, when the user selects content presented to the user, for example, the behavior of the content creator during the period including the time of creation of the content is reproduced as the behavior of an avatar.
[0063] 2. Use Cases Possible use cases for the three-dimensional space sharing system of the present technology include, for example, creating a trouble report, searching for past trouble reports, and learning past responses to trouble.
[0064] First, a use case in which an operator creates a trouble report using the device 2 will be described with reference to FIG.
[0065] The worker takes an image of the equipment where the problem has occurred, for example, using a camera mounted on the device 2. The captured image is placed in a space Sp1, as shown on the left side of FIG.
[0066] The worker operates the device 2 to group together, among the multiple captured images arranged in the space Sp1, those that have a high degree of similarity. For example, multiple captured images that have been continuously shot are grouped together. The worker selects representative content that represents the group from the multiple captured images included in the group.
[0067] For example, as described above, the worker can specify a range within space Sp1 to narrow down the captured images presented on device 2, group highly similar captured images into one group, or determine representative content.
[0068] After the representative content is determined, the trouble report is completed. When the content included in the trouble report is presented to the user, only the image objects PO21 to P24 corresponding to the representative content are presented in the space Sp1, as shown on the right side of FIG.
[0069] Next, a use case in which an instructor creates a trouble report using the device 1 will be described with reference to FIGS.
[0070] As shown in A of Fig. 9 , captured images obtained by an operator are arranged two-dimensionally and presented on the device 1. The instructor selects representative content from the two-dimensionally arranged captured images. In the example of Fig. 9 , a check mark is added to the captured image selected by the instructor as the representative content.
[0071] After the representative content is determined, as shown in Fig. 9B, a screen that provides a bird's-eye view of the space Sp1 in which the representative content is arranged is displayed on the device 1. In the example of Fig. 9B, image objects PO31 to PO34 corresponding to the representative content are arranged in the space Sp1.
[0072] When the user selects the menu icon MI1 and performs a predetermined operation, a details window W1 is displayed on the screen, as shown in Fig. 10C. In the details window W1, for example, photographed images, which are representative contents, are arranged in the order in which they were taken. In the example of Fig. 10C, the photographed images corresponding to the image objects PO31, PO32, PO33, and PO34 are arranged in the details window W1 in this order.
[0073] As shown by the double arrow in D of Figure 10, when the instruction person rearranges the order of the representative contents in the detailed window W1 into a desired order, the order of the representative contents in the detailed window W1 is recorded in the trouble report as the order of the contents.
[0074] In the example of Fig. 10D, the content corresponding to the image object PO32 is the first content, the content corresponding to the image object PO31 is the second content, the content corresponding to the image object PO33 is the third content, and the content corresponding to the image object PO34 is the fourth content.
[0075] After the order of the content is determined, the trouble report is completed.
[0076] When representative content is selected, the photographed images may not be arranged and presented, but a screen may be displayed that provides a bird's-eye view of the space Sp1 in which the photographed images are arranged.
[0077] In this case, as described above, the instructor specifies the range within the space Sp1 to narrow down the captured images presented on the device 1 and determine the representative content. The content pointed at by the instructor may be determined as the representative content, or the representative content may be determined based on the frequency and duration of the instructor's gaze or face directed toward the content.
[0078] Next, a use case in which an instructor searches for a trouble report using the device 1 will be described with reference to FIGS.
[0079] FIG. 11 is a diagram illustrating a use case in which a trouble report is searched for based on conditions related to the equipment to be operated.
[0080] As shown in FIG. 11, the instructor can narrow down the desired trouble report from among multiple trouble reports created in the past by selecting from a pull-down menu PM1 the trouble report for which equipment the instructor wants to view.
[0081] FIG. 12 is a diagram illustrating a use case of searching for a trouble report based on a condition related to a range in space.
[0082] As shown in the upper part of Fig. 12, if the instructor selects, for example, the second representative content for trouble A, a pin Pi11 is presented at the position of the representative content, and a sphere Sph11 centered on the pin Pi11 is presented, as shown in the lower part of Fig. 12. When the pin Pi11 is presented, the content located inside the sphere Sph11 and included in the same trouble report (trouble report for trouble A) as the representative content selected by the instructor is presented in the form of an image object, as shown by a solid white rectangle in the lower part of Fig. 12.
[0083] In the lower part of Figure 12, representative contents other than the second representative content for trouble A and representative contents for troubles B and C are illustrated as dashed rectangles, indicating that these representative contents are presented in the form of translucent image objects.
[0084] In this way, the instructor can select one representative content from among multiple representative contents placed in the space, thereby narrowing down the desired trouble report from multiple trouble reports created in the past and specifying a range within space Sp1.
[0085] The pin Pi11 may be set at the position of the representative content that is closest to the position of the instructor within the space Sp1 and that is the representative content toward which the instructor's line of sight or face is directed.
[0086] Here, the shape of the specified range (range specification UI) has been described as spherical, but the shape of the specified range is not limited to a sphere and may be cubic or conical with the position of the instructor as the apex. The shape of the specified range may be changed depending on the shape of the site or equipment. For example, if the shape of the equipment is a rectangular parallelepiped, the shape of the specified range can be rectangular, or the contact surface with the equipment can be shaped to follow the shape of the equipment. Furthermore, the shape of the specified range can be shaped to follow the path of a site, for example.
[0087] In addition, the representative contents other than the second representative content for trouble A and the representative contents for troubles B and C may be presented in the form of point objects, as shown in Figure 13, rather than being presented in the form of translucent image objects.
[0088] FIG. 14 is a diagram illustrating another use case of searching for trouble reports based on a condition related to a range in space.
[0089] As shown in the upper part of Figure 14, if the instructor specifies a position near the second representative content for, for example, Trouble A, a pin Pi12 is presented at the position specified by the instructor, and a sphere Sph12 centered on the pin Pi12 is presented, as shown in the lower part of Figure 14.
[0090] When the pin Pi12 is presented, the content arranged inside the sphere Sph12 is presented in the form of image objects, as shown by the solid-line rectangle in the lower part of Fig. 14. In the example in the lower part of Fig. 14, among the content for trouble A, the content that belongs to the same group as the second representative content, and the third representative content for trouble C are presented in the form of image objects.
[0091] In the lower part of Figure 14, representative contents other than the second representative content for Trouble A, representative contents other than the third representative content for Trouble B, and representative contents for Trouble C are illustrated as dashed rectangles, indicating that these representative contents are presented in the form of translucent image objects.
[0092] In this way, by specifying a range within the space Sp1, the instructor can narrow down the desired content from among the content contained in each of a plurality of trouble reports created in the past.
[0093] FIG. 15 is a diagram illustrating a use case in which a trouble report is searched for based on a condition related to the creation date and time of the content.
[0094] The upper screen of Fig. 15 presents representative content created over a certain week. As shown in the upper part of Fig. 15, when the instructor moves the position of the black circle at the right end of the time bar B1 to the left in the figure, the representative content presented on device 1 is narrowed down to representative content created over the three days from August 20, 2023 to August 23, 2023, as shown in the lower part of Fig. 15. In the example at the bottom of Fig. 14, representative content for troubles A and B is presented.
[0095] In this way, the instructor can operate the time bar B1 to specify conditions regarding the date and time of content creation, thereby narrowing down the desired trouble report from among multiple trouble reports created in the past.
[0096] FIG. 16 is a diagram illustrating a use case in which a trouble report is searched for based on a condition related to a trouble responder.
[0097] The upper screen of Fig. 16 presents representative contents for troubles for which at least one of the three people represented by icons I1 to I3 is registered as the responder. When the instructor selects, for example, icon I2 as shown in the upper part of Fig. 16, the representative contents presented on device 1 are narrowed down to representative contents for troubles for which the person represented by icon I2 is registered as the responder, as shown in the lower part of Fig. 16. In the example shown in the lower part of Fig. 16, representative contents for troubles A and C are presented.
[0098] In this way, the person in charge can select an icon and specify conditions regarding the person who will handle the problem, thereby narrowing down the desired trouble report from among multiple trouble reports that have been created in the past.
[0099] The color of the icon may be changed depending on the attributes of the person who responded to the problem or the type of device used by the person who responded to the problem. It may also be possible to search for trouble reports using the attributes of the person who responded to the problem or the type of device used as conditions. It may also be possible to search for trouble reports using the eye position or height of the person who responded to the problem as conditions.
[0100] FIG. 17 is a diagram illustrating a use case in which a trouble report is searched for based on a plurality of conditions.
[0101] As shown in the upper part of Fig. 17, if the instructor selects, for example, icon I2 and then moves the position of the black circle at the right end of the time bar B1 to the left in the figure, the representative content presented on device 1 will be narrowed down to representative content for troubles for which the person represented by icon I2 is registered as the responder and that was created during the three days from 20 / 08 / 2023 to 2-23 / 08 / 23, as shown in the lower part of Fig. 17. In the example at the bottom of Fig. 17, representative content for trouble A is presented.
[0102] In this way, the person in charge can further narrow down the desired trouble report from among multiple trouble reports created in the past by operating multiple UIs to specify different conditions.
[0103] Next, a use case in which a user learns past responses to a problem will be described with reference to FIG.
[0104] When the user finds a desired trouble report or content from among multiple trouble reports and then selects the play button B11 shown to the left of the time bar B1 in Fig. 18, the behavior of the content creator related to the user's desired content is reproduced as the behavior of an avatar. Hereinafter, having an avatar reproduce the behavior of a person is also referred to as "playing behavior data."
[0105] 18 shows a screen that appears when the user selects the play button B11 after the trouble reports presented on the device have been narrowed down to those about Trouble A. In the example of Fig. 18, an avatar A1 pointing at the second representative content is presented, followed by an avatar A2 making a thumbs-up gesture. Here, avatars that reproduce the actions of people associated with each representative content are presented in order of the content recorded in the trouble report.
[0106] By watching the avatar's movements, the user can learn what actions the responder took to resolve the problem. Here, the time bar B1 indicates the playback position of the behavioral data. The user can adjust the playback position of the behavioral data by operating the time bar B1. Note that a display may be provided that shows the relationship between the position on the time bar B1 and the creation time of each content.
[0107] The user can switch the presentation of the avatar on and off by operating an icon I11 presented at the top center of the screen in FIG.
[0108] FIG. 19 is a diagram illustrating the flow of recording behavioral data.
[0109] As shown in the upper part of Fig. 19, for example, assume that a worker photographs a real object 10R in the AR space and captures images corresponding to image objects PO51 to PO53. All actions of the worker (user A) and instructor (user B) while they are accessing the AR space and the VR space, respectively, are automatically recorded in the form of action data in an action DB (DataBase) 51, as indicated by the white arrow #1 in Fig. 19.
[0110] When a worker photographs the real object 10R, the photographed image is recorded in the image DB 52 as indicated by the white arrow #2 in Fig. 19. At this time, the photographing time of the photographed image (the creation time of the content) is recorded as metadata in the image DB 52. In the image DB 52, a content ID that can identify the content is assigned to each photographed image.
[0111] After the worker and the instructor exit the AR space and the VR space, respectively, in other words, after the recording of the behavioral data is completed, the behavioral data related to the captured image is extracted from the behavioral data recorded in the behavior DB 51 based on the shooting time of the captured image.
[0112] 19, the extracted behavioral data is recorded for each user in a behavioral DB 53 that is different from the behavioral DB 51. In the behavioral DB 53, each behavioral data is recorded in association with a content ID assigned to the related captured image.
[0113] The behavior DB 51, the image DB 52, and the behavior DB 53 are provided, for example, in the recording unit 4 ( FIG. 4 ). Note that the behavior data recorded in the behavior DB 51 is deleted after a predetermined period has elapsed since the worker and the instructor exited the AR space and the VR space, respectively.
[0114] FIG. 20 is a diagram for explaining the extraction of behavioral data.
[0115] The horizontal arrow at the top of FIG. 20 indicates the time from when an operator starts to deal with a problem until when the worker finishes, and the vertical line above the horizontal arrow indicates the time when the image was taken.
[0116] When a worker photographs representative content A while standing still, for example, behavioral data while the worker is standing still, behavioral data for three steps before the worker starts to stand still, and behavioral data for three steps after the worker starts to move are extracted as one piece of behavioral data related to representative content A.
[0117] When a worker photographs representative contents B and C while performing a task, for example, behavioral data while the worker is performing the task, behavioral data for three steps before the worker starts the task, and behavioral data for three steps after the worker finishes the task are extracted as one piece of behavioral data related to representative contents B and C.
[0118] When a worker photographs representative content D while moving, for example, behavioral data while the worker is moving, behavioral data for 30 seconds before the worker starts moving, and behavioral data for 30 seconds after the worker stops are extracted as one piece of behavioral data related to representative content D.
[0119] 21 is a diagram illustrating the relationship between the trajectory of a worker's movement at the site and the extracted behavioral data. In FIG. 21, thin arrows indicate the trajectory of a worker's movement at the site.
[0120] For example, if a worker takes a photo while stationary at time t1, the history of the worker's actions within a range of three steps centered on the stationary position (history of user position and posture information) is extracted as behavioral data A.
[0121] For example, if a worker takes a photo while moving at time t3, the worker's behavior history (user position and posture information history) from 30 seconds before time t2, when the worker starts moving, to 30 seconds after time t4, when the worker stops, is extracted as behavioral data B.
[0122] If the stationary period, the working period, and the moving period are longer than a predetermined threshold, the extracted behavioral data may be divided into sections corresponding to the threshold.
[0123] In this way, the behavioral history of the content creator (the worker who took the photographed image) during the period including the time of content creation (the time the photographed image was taken) is extracted as behavioral data, and the behavior of the content creator during that period is reproduced as the behavior of an avatar.
[0124] The three-dimensional space sharing system of this technology can be applied to sharing content created not only at manufacturing sites, but also at construction sites, agricultural sites (farms), medical sites, video production sites, and other locations.
[0125] 3. Configuration and Operation of Each Device FIG. 22 is a block diagram showing an example configuration of the smartphone 101. As shown in FIG.
[0126] The device 1 (FIG. 4) and the device 2 (FIG. 4) of the present technology are configured as, for example, a smartphone 101 having the configuration shown in FIG. 22 .
[0127] The smartphone 101 in FIG. 22 is configured from a sensor unit 111, a control unit 112, a storage unit 113, an image output unit 114, an audio output unit 115, and an external communication unit 116.
[0128] The sensor unit 111 is composed of a self-position estimation unit 121 , a voice acquisition unit 122 , and an image acquisition unit 123 .
[0129] The self-position estimation unit 121 acquires sensor data used to estimate the self-position of the smartphone 101 (user) using an acceleration sensor, a gyro sensor, a direction sensor, SLAM (Simultaneous Localization and Mapping), a depth sensor, etc., and supplies the sensor data to the control unit 112.
[0130] The voice acquisition unit 122 records the voice of the user using a microphone or the like and supplies the recorded data to the control unit 112. The voice recorded by the voice acquisition unit 122 is presented to, for example, another user using an external device. That is, the user can talk to another user using the smartphone 101.
[0131] The image acquisition unit 123 uses a camera or the like to capture an image of the equipment to be operated, and supplies the captured image to the control unit 112 .
[0132] The control unit 112 is configured as a processor such as a CPU, and controls each unit of the smartphone 101. The control unit 112 has a self-position estimation processing unit 131, an attitude detection processing unit 132, an equipment information processing unit 133, an application execution unit 134, a recording processing unit 135, an output control unit 136, and a communication control unit 137.
[0133] The self-position estimation processing unit 131 estimates the self-position of the smartphone 101 based on the sensor data (e.g., acceleration data, angular velocity data, SLAM data, orientation data, etc.) supplied from the self-position estimation unit 121. The self-position of the smartphone 101 estimated by the self-position estimation processing unit 131 is used, for example, to place an image object corresponding to a captured image in space.
[0134] The posture detection processing unit 132 analyzes the sensor data supplied from the self-position estimation unit 121 and the captured image supplied from the image acquisition unit 123, and detects the user's posture, the posture of the user's hands, etc. The posture of the user and the posture of the user's hands detected by the posture detection processing unit 132 are recorded as behavioral data related to the captured image.
[0135] The equipment information processing unit 133 processes information about the equipment to be operated. For example, the equipment information processing unit 133 acquires a 3D model of the equipment to be operated.
[0136] The application execution unit 134 executes a predetermined application to realize the presentation of the above-mentioned content and avatars.
[0137] The recording processing unit 135 records various data in the storage unit 113 and reads various data from the storage unit 113 .
[0138] The output control unit 136 controls the image output by the image output unit 114 and the audio output by the audio output unit 115 .
[0139] The communication control unit 137 controls wireless communication by the external communication unit 116 to transmit and receive data to and from the server 3 .
[0140] The storage unit 113 is configured as, for example, a flash memory, and stores various data required for the processing performed by the control unit 112 .
[0141] The image output unit 114 uses a display or the like to present the captured image or the like as content to the user.
[0142] The audio output unit 115 uses a speaker or the like to present audio or the like as content to the user.
[0143] The external communication unit 116 performs wireless communication with external devices using, for example, a wireless LAN (Local Area Network).
[0144] Fig. 23 is a block diagram showing an example of the configuration of the PC 151. In Fig. 23, the same components as those in Fig. 22 are denoted by the same reference numerals. Duplicate explanations will be omitted where appropriate.
[0145] The PC 151 in FIG. 23 differs from the smartphone 101 in FIG. 22 in that an operation position estimation unit 161 is provided in place of the self-position estimation unit 121 in the sensor unit 111, and an operation position estimation processing unit 171 is provided in place of the self-position estimation processing unit 131 in the control unit 112.
[0146] The device 1 of the present technology is configured as a PC 151 having the configuration shown in FIG. 23, for example.
[0147] The operation position estimation unit 161 acquires an operation signal used to estimate the user's operation position on the screen (e.g., the position of the mouse cursor) using an input device such as a mouse, and supplies the operation signal to the control unit 112.
[0148] The image acquisition unit 123 uses a camera or the like to capture an image of the user, for example, and supplies the captured image to the control unit 112 .
[0149] The operation position estimation processing unit 171 estimates the user's operation position on the screen based on the operation signal supplied from the self-position estimation unit 121 .
[0150] The posture detection processing unit 132 analyzes the captured image supplied from the image acquisition unit 123 and detects the posture of the user and the posture of the user's hands.
[0151] Fig. 24 is a block diagram showing an example of the configuration of the HMD 201. In Fig. 24, the same components as those in Fig. 22 are denoted by the same reference numerals. Duplicate explanations will be omitted where appropriate.
[0152] The HMD 201 in FIG. 24 differs from the smartphone 101 in FIG. 22 in that the image acquisition unit 123 is not provided in the sensor unit 111 and that a hand posture estimation unit 211 is provided.
[0153] The device 1 and the device 2 of the present technology are configured as, for example, an HMD 201 for VR or an HMD 201 for AR having the configuration shown in FIG. 24 .
[0154] The hand posture estimation unit 211 acquires sensor data used to estimate the posture of the user's hand using a depth sensor, an infrared camera, or the like, and supplies the sensor data to the control unit 112 .
[0155] The posture detection processing unit 132 analyzes the sensor data supplied from the self-position estimation unit 121 and the hand posture estimation unit 211, and detects the posture of the user and the posture of the user's hands.
[0156] Next, a process of extracting behavioral data by the three-dimensional space sharing system of the present technology will be described with reference to the flowchart of Fig. 25. Here, for example, an example will be described in which a worker as a user controls the extraction of behavioral data using the device 2.
[0157] In step S1 , the communication control unit 137 of the device 2 acquires from the server 3 the content arranged in the space.
[0158] In step S2, the output control unit 136 of the device 2 presents the content arranged in the space to the user.
[0159] In step S3, the application execution unit 134 of the device 2 determines whether or not the user has selected content.
[0160] If it is determined in step S3 that no content has been selected by the user, the process returns to step S1, and the presentation of the content arranged in the space continues.
[0161] If it is determined in step S3 that content has been selected by the user, the communication control unit 137 of the device 2 notifies the server 3 of the content ID of the content selected by the user. Then, in step S4, the server 3 obtains from the behavior DB 51 behavior data that records all of the user's behavior during the period when the user was responding to the problem.
[0162] In step S5, the server 3 determines whether the user was moving at the time the content selected by the user was created.
[0163] If it is determined in step S5 that the user was not moving, the server 3 determines in step S6 that the user created the content while stationary, and extracts the behavioral data. Specifically, the server 3 extracts the behavioral data that records the user's behavior within a range of three steps from the stationary position.
[0164] On the other hand, if it is determined in step S5 that the user has been moving, the server 3 determines in step S7 whether the user has been moving while remaining in the same position.
[0165] If it is determined in step S7 that the user did not stay in the same position but moved, the server 3 determines in step S8 that the user created the content while moving, and extracts the behavioral data. Specifically, the server 3 extracts the behavioral data that records the worker's behavior from 30 seconds before the worker started moving to 30 seconds after the worker stopped moving.
[0166] If it is determined in step S7 that the user moved while remaining in the same position, the server 3 determines in step S9 that the user created the content while working, and extracts the behavioral data. Specifically, the server 3 extracts the behavioral data that records the user's behavior within a range of three steps from the work position.
[0167] After the processes of steps S6, S8, and S9 are performed, in step S10, the server 3 registers the extracted behavior data in the behavior DB 53. In the behavior DB 53, a behavior data ID that can identify the behavior data is assigned to each behavior data.
[0168] In step S11, the server 3 associates the behavior data ID with the content ID.
[0169] In step S12, the server 3 determines whether or not to end the extraction of the behavioral data. For example, if the user instructs to end the extraction of the behavioral data, it is determined that the extraction of the behavioral data is to be ended.
[0170] If it is determined in step S12 that the extraction of behavioral data is not to be completed, the process returns to step S1, and the subsequent processes are carried out.
[0171] On the other hand, if it is determined in step S12 that the extraction of behavioral data has been completed, the process ends.
[0172] Next, a process in which the three-dimensional space sharing system of the present technology automatically selects representative content and creates a trouble report will be described with reference to the flowcharts of Figures 26 and 27. Here, an example will be described in which representative content is automatically selected based on the response made by a user to a trouble.
[0173] In step S31, the application execution unit 134 of the device 1 sets the number of times of pointing for each content to Count=0.
[0174] In step S32, the application executing unit 134 sets the cumulative pointing time t2 for each piece of content to 0.
[0175] In step S33, the equipment information processing unit 133 of the device 1 acquires a 3D model of the equipment to be operated.
[0176] In step S34, the communication control unit 137 of the device 1 obtains from the server 3 the content IDs of all content relating to the trouble associated with the equipment to be operated, and obtains from the server 3 all content to which the content IDs are assigned.
[0177] In step S35, the output control unit 136 of the device 1 arranges all content relating to the troubles associated with the equipment to be operated within the space and presents it to the user. Here, a 3D model of the equipment to be operated is also arranged within the space and presented to the user.
[0178] In step S36, the orientation detection processing unit 132 of the device 1 determines whether the user has started pointing at the content. For example, in order to solve a problem, the instructor issues instructions to a worker on-site while pointing at the image object POV3, as shown in FIG.
[0179] If it is determined in step S36 that pointing at the content has not started, the process returns to step S34, and presentation of the content continues.
[0180] On the other hand, if it is determined in step S36 that pointing at the content has started, the application execution unit 134 sets the pointing duration t1 to 0 in step S37.
[0181] In step S38, the application execution unit 134 acquires the content ID of the content pointed to by the user.
[0182] In step S39, the application execution unit 134 increments the number of times Count, which is the number of times the user has pointed at the content, by 1 (Count=Count+1).
[0183] In step S40, the application executing unit 134 starts measuring a pointing duration t1.
[0184] In step S41, the orientation detection processing unit 132 determines whether or not the pointing of the finger on the content has ended, and continues measuring the pointing duration t1 until the pointing of the finger on the content has ended.
[0185] If it is determined in step S41 that the pointing on the content has ended, then in step S42, the application execution unit 134 ends measurement of the pointing duration t1.
[0186] In step S43, the application execution unit 134 adds the pointing duration t1 to the pointing cumulative time t2 of the content pointed at by the user. For example, if the past pointing cumulative time t2 was 8 seconds and the pointing duration t1 was 2 seconds, the current pointing cumulative time t2 is 10 seconds.
[0187] In step S44, the application executing unit 134 determines whether or not the troubleshooting has been completed.
[0188] If it is determined in step S44 that the troubleshooting has not been completed, the process returns to step S34, and the subsequent steps are carried out.
[0189] If it is determined in step S44 that the troubleshooting has been completed, in step S45, the application execution unit 134 sorts the content IDs in descending order of the cumulative pointing time t2.
[0190] In step S46, the application execution unit 134 sorts the content IDs in descending order of the number of times the finger has been pointed at (Count).
[0191] In step S47, the application execution unit 134 assigns a representative flag to the top N content IDs, indicating that the content is representative content. Basically, the content with the largest number of pointing counts is selected as the representative content. For content with the same number of pointing counts, the content with the longer cumulative pointing time t2 is selected as the representative content.
[0192] In step S48, the output control unit 136 arranges in the space only the content (representative content) that has been assigned a representative flag from among all the content regarding the trouble related to the equipment to be operated, and presents it to the user.
[0193] In step S49, the application executing unit 134 determines whether or not to end the presentation of the representative content.
[0194] If it is determined in step S49 that the presentation of the representative content should not be ended, the process returns to step S48, and the presentation of the representative content arranged in the space continues.
[0195] On the other hand, if it is determined in step S49 that the presentation of the representative content is to be ended, the process ends.
[0196] Next, a process in which the three-dimensional space sharing system of the present technology presents a trouble report to a user will be described with reference to the flowchart of FIG. 29 .
[0197] In step S61, the equipment information processing unit 133 of the device acquires a 3D model of the equipment to be operated. The communication control unit 137 of the device acquires from the server 3 all content relating to the trouble associated with the equipment to be operated.
[0198] In step S62, the output control unit 136 of the device arranges all content related to the troubles associated with the equipment to be operated in the space and presents it to the user. Here, a 3D model of the equipment to be operated is also arranged in the space and presented to the user.
[0199] In step S63, the application execution unit 134 of the device determines whether or not the user has specified a range within the space.
[0200] If it is determined in step S63 that the range within the space has not been specified, the process returns to step S62, and the presentation of the content arranged within the space continues.
[0201] On the other hand, if it is determined in step S63 that a range within space has been specified, in step S64, the application execution unit 134 sets the center position of the specified range based on the operations input by the user, the user's line of sight, etc., and the output control unit 136 presents a pin to the user.
[0202] In step S65, the application execution unit 134 sets the radius (size of the specified range) from the center position based on operations input by the user, and the output control unit 136 presents a range specification UI (sphere) to the user.
[0203] In step S66, the application execution unit 134 acquires the content ID of the content placed within the specified range.
[0204] In step S67, the output control unit 136 presents the content placed within the specified range to the user. The content placed outside the specified range is presented in the form of a semi-transparent image object or a point object, as described above.
[0205] In step S68, the application execution unit 134 determines whether the content presented to the user has been selected by the user.
[0206] If it is determined in step S68 that no content has been selected by the user, the process returns to step S63, and the subsequent processes are carried out.
[0207] On the other hand, if it is determined in step S68 that content has been selected by the user, the control unit 112 of the device executes behavior data playback processing in step S69. Through the behavior data playback processing, the behavior of the trouble solver to solve the trouble is reproduced as the behavior of the avatar. Details of the behavior data playback processing will be described later with reference to FIG. 30.
[0208] In step S70, the application executing unit 134 determines whether or not to end the presentation of the trouble report.
[0209] If it is determined in step S63 that the presentation of the trouble report should not be terminated, the process returns to step S63, and the subsequent processes are carried out.
[0210] On the other hand, if it is determined in step S64 that the presentation of the trouble report is to be terminated, the process ends.
[0211] Next, the behavior data reproducing process performed in step S69 of FIG. 29 will be described in detail with reference to the flowchart of FIG.
[0212] In step S81, the application execution unit 134 acquires the content ID of the content selected by the user.
[0213] In step S82, the communication control unit 137 acquires from the server 3 a behavior data ID linked to the content ID of the content selected by the user, and acquires from the server 3 the behavior data to which the behavior data ID has been assigned.
[0214] In step S83, the application execution unit 134 plays the behavioral data acquired from the server 3. Specifically, the output control unit 136 places an avatar that moves based on the behavioral data in a space and presents it to the user.
[0215] In step S84, the application executing unit 134 determines whether the playback of the behavior data has ended.
[0216] If it is determined in step S84 that the playback of the behavior data has not ended, the process returns to step S83, and the playback of the behavior data continues.
[0217] On the other hand, if it is determined in step S84 that the playback of the behavior data has ended, the process returns to step S69 in FIG. 29, and the subsequent processes are carried out.
[0218] As described above, in the device 1 and the device 2 of the present technology, a UI for specifying conditions for narrowing down a plurality of pieces of content virtually arranged in a space and a content that satisfies the conditions among the plurality of pieces of content arranged in the space are presented to the user, thereby enabling the user to easily find a desired piece of content from a large amount of content.
[0219] Furthermore, in device 1 or device 2 of the present technology, when a user selects content presented to the user, an avatar that reproduces the behavior of a person related to the selected content is presented to the user. After the content is narrowed down based on conditions such as the range within the space, the behavior data linked to the narrowed down content is played back, eliminating the need for the user to search for a desired avatar from among the large number of avatars placed.
[0220] <Regarding the Computer> The above-described series of processes can be executed by hardware or software. When the series of processes are executed by software, the program constituting the software is installed from a program recording medium into a computer incorporated in dedicated hardware, or into a general-purpose personal computer, etc.
[0221] FIG. 31 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0222] A CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .
[0223] An input / output interface 505 is also connected to the bus 504. An input unit 506 including a keyboard, a mouse, etc., and an output unit 507 including a display, a speaker, etc. are connected to the input / output interface 505. Also connected to the input / output interface 505 are a storage unit 508 including a hard disk, a nonvolatile memory, etc., a communication unit 509 including a network interface, etc., and a drive 510 that drives removable media 511.
[0224] In a computer configured as described above, the CPU 501 performs the above-described series of processes by, for example, loading a program stored in the storage unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executing it.
[0225] The program executed by the CPU 501 is installed in the storage unit 508 by being recorded on, for example, a removable medium 511 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.
[0226] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0227] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0228] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0229] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0230] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0231] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0232] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0233] <Examples of Combinations of Configurations> The present technology can also have the following configurations.
[0234] (1) An information processing device comprising: a UI for specifying conditions for narrowing down a plurality of pieces of content virtually arranged in a space; and an output control unit that presents to a user the pieces of content that satisfy the conditions from among the plurality of pieces of content arranged in the space, wherein the output control unit, when the user selects the content presented to the user, presents to the user an avatar that reproduces the behavior of a person related to the selected content. (2) The information processing device described in (1), wherein the UI is a range designation UI for designating a range within the space, and the output control unit presents to the user the pieces of content arranged within the range designated by the range designation UI. (3) The information processing device described in (2), wherein the content includes an image, and the output control unit presents the image arranged within the range designated by the range designation UI in the form of an image object onto which the image is projected, by arranging it in the space. (4) The information processing device described in (3), wherein the output control unit presents the image arranged outside the range designated by the range designation UI in the form of a point object or a semi-transparent image object by arranging it in the space. (5) The information processing device according to (3) or (4), wherein the output control unit arranges and presents the plurality of images two-dimensionally, and arranges in the space an image selected by the user from the plurality of images arranged and presented two-dimensionally. (6) The information processing device according to any of (3) to (5), wherein the output control unit arranges and further presents the images arranged within a range specified by the range designation UI two-dimensionally. (7) The information processing device according to any of (3) to (6), wherein the images are photographed images taken using a camera, and the output control unit arranges and presents the image object in the space at a position and attitude corresponding to the photographing position and photographing attitude of the photographed image. (8) The information processing device according to (7), wherein the avatar reproduces the behavior of the person who took the photographed image selected by the user from the photographed images presented to the user.(9) The information processing device according to any one of (8), wherein the avatar reproduces the behavior of the person during a period including the shooting time of the captured image selected by the user from the captured images presented to the user. (10) The information processing device according to any one of (1) to (9), wherein the UI is a time designation UI for designating a range of creation times of the content, and the output control unit presents to the user the content created at a time within the range designated by the time designation UI. (11) The information processing device according to any one of (1) to (10), wherein the UI is a person designation UI for designating the person associated with the content, and the output control unit presents to the user the content associated with the person designated in the person designation UI. (12) The information processing device according to any one of (1) to (11), wherein the output control unit presents a plurality of the UIs for designating different conditions. (13) An information processing method comprising: presenting to a user a UI for specifying conditions for narrowing down multiple pieces of content virtually arranged in a space; and presenting to the user, among the multiple pieces of content arranged in the space, the pieces of content that satisfy the conditions; and, when the user selects the content presented to the user, presenting to the user an avatar that reproduces the behavior of a person related to the selected content. (14) The information processing method described in (13), wherein the UI is a range specification UI for specifying a range in the space, and the content arranged within the range specified by the range specification UI is presented to the user. (15) The information processing method described in (14), wherein the content includes an image, and the image arranged within the range specified by the range specification UI is presented in the space in the form of an image object onto which the image is projected. (16) The information processing method described in (15), further comprising presenting, in the space, the image arranged outside the range specified by the range specification UI in the form of a point object or a semi-transparent image object.(17) The information processing method according to (15) or (16), further including: arranging and presenting a plurality of the images two-dimensionally; and arranging, in the space, an image selected by the user from the plurality of images arranged two-dimensionally and presented. (18) The information processing method according to any of (15) to (17), further including arranging and presenting a plurality of the images arranged within a range specified by the range specification UI two-dimensionally. (19) The information processing method according to any of (15) to (18), wherein the images are photographed images taken using a camera, and the image objects are presented by arranging them in the space at positions and attitudes corresponding to the photographing positions and photographing attitudes of the photographed images. (20) A program for causing a computer to execute processing including: presenting to a user a UI for specifying conditions for narrowing down a plurality of contents virtually arranged in the space, and content that satisfies the conditions from among the plurality of contents arranged in the space; and, when the content presented to the user is selected by the user, presenting to the user an avatar that reproduces the behavior of a person associated with the selected content.
[0235] DESCRIPTION OF SYMBOLS 1, 2 Device, 3 Server, 4 Recording unit, 101 Smartphone, 111 Sensor unit, 112 Control unit, 113 Memory unit, 114 Image output unit, 115 Audio output unit, 116 External communication unit, 121 Self-position estimation unit, 122 Audio acquisition unit, 123 Image acquisition unit, 131 Self-position estimation processing unit, 132 Posture detection processing unit, 133 Equipment information processing unit, 134 Application execution unit, 135 Recording processing unit, 136 Output control unit, 137 Communication control unit, 151 PC, 161 Operation position estimation unit, 171 Operation position estimation processing unit, 201 HMD, 211 Hand posture estimation unit
Claims
a UI for specifying a condition for narrowing down a plurality of pieces of content virtually arranged in a space, and an output control unit for presenting to a user the pieces of content that satisfy the condition among the plurality of pieces of content arranged in the space; When the content presented to the user is selected by the user, the output control unit presents to the user an avatar that reproduces the behavior of a person related to the selected content. Information processing device. the UI is a range designation UI for designating a range within the space, The output control unit presents the content arranged within the range specified by the range specification UI to the user. The information processing device according to claim 1 . the content includes an image; The output control unit arranges and presents the image arranged within the range specified by the range designation UI in the space in the form of an image object onto which the image is projected. The information processing device according to claim 2 . The output control unit arranges and presents the image placed outside the range specified by the range specification UI in the form of a point object or a semi-transparent image object within the space. The information processing device according to claim 3 . The output control unit A plurality of the images are arranged two-dimensionally and presented; The image selected by the user from the plurality of images arranged and presented two-dimensionally is arranged in the space. The information processing device according to claim 3 . The output control unit arranges the images arranged within the range specified by the range specification UI two-dimensionally and further presents them. The information processing device according to claim 3 . The image is a photographed image taken using a camera, The output control unit arranges and presents the image object in the space at a position and orientation corresponding to the photographing position and photographing orientation of the photographed image. The information processing device according to claim 3 . The avatar reproduces the behavior of the person who took the photograph of the photographed image selected by the user from the photographed images presented to the user. The information processing device according to claim 7 . The avatar reproduces the behavior of the person during a period including the shooting time of the photographed image selected by the user from the photographed images presented to the user. The information processing device according to claim 8 . the UI is a time specification UI for specifying a range of creation times of the content, The output control unit presents the content created at a time within a range specified by the time specification UI to the user. The information processing device according to claim 1 . the UI is a person designation UI for designating the person associated with the content, The output control unit presents the content related to the person designated in the person designation UI to the user. The information processing device according to claim 1 . The output control unit presents a plurality of UIs for specifying different conditions. The information processing device according to claim 1 . presenting to a user a UI for specifying a condition for narrowing down a plurality of pieces of content virtually arranged in a space, and the pieces of content that satisfy the condition among the plurality of pieces of content arranged in the space; When the content presented to the user is selected by the user, presenting to the user an avatar that reproduces the behavior of a person related to the selected content; An information processing method including: the UI is a range designation UI for designating a range within the space, The content arranged within the range specified by the range specification UI is presented to the user. The information processing method according to claim 13. the content includes an image; The image placed within the range specified by the range specification UI is arranged and presented in the space in the form of an image object onto which the image is projected. The information processing method according to claim 14. The image placed outside the range specified by the range specification UI is displayed in the space in the form of a point object or a semi-transparent image object. The information processing method according to claim 15. presenting a plurality of said images in a two-dimensional array; arranging the image selected by the user from the plurality of images arranged and presented two-dimensionally in the space; The information processing method according to claim 15, further comprising: The method further includes presenting the images arranged within the range specified by the range specification UI in a two-dimensional arrangement. The information processing method according to claim 15. The image is a photographed image taken using a camera, The image object is arranged in the space at a position and orientation corresponding to the photographing position and photographing orientation of the photographed image and is presented. The information processing method according to claim 15. On the computer, presenting to a user a UI for specifying a condition for narrowing down a plurality of pieces of content virtually arranged in a space, and the pieces of content that satisfy the condition among the plurality of pieces of content arranged in the space; When the content presented to the user is selected by the user, presenting to the user an avatar that reproduces the behavior of a person related to the selected content; A program for executing a process including:
Citation Information
Patent Citations
Virtual field training device
JP2002366021A
Work support system, work support method, and program
JP2020149143A
Content provision system, content provision method, and content provision program
JP2022183944A
Program and information processing apparatus
JP2024047954A
Information processing device, information processing method, and program
WO2019187747A1