Virtual three-dimensional space sharing system, virtual three-dimensional space sharing method, and virtual three-dimensional space sharing server
The virtual three-dimensional space sharing system addresses the challenge of real-time sharing of on-site situations and actions, facilitating effective guidance from remote locations by using sensors and a server to map and transmit data in real-time.
Patent Information
- Application Number
- JP2022156516
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2025-11-26
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Existing systems fail to provide real-time sharing of on-site situations and actions of multiple people in remote locations, making it difficult to offer appropriate guidance from a remote location.
A virtual three-dimensional space sharing system that includes first and second sensors observing objects and users at different locations, a server mapping their data onto a virtual three-dimensional space, and display devices transmitting real-time information.
Enables real-time sharing of on-site situations and actions, allowing for effective guidance from remote locations.
Smart Images

Figure 0007776397000001 
Figure 0007776397000002 
Figure 0007776397000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a virtual three-dimensional space sharing system. [Background technology]
[0002] There are situations where multiple people in different locations need to share information. For example, if equipment at a site breaks down, an experienced maintenance technician may travel to the site to provide maintenance instructions. However, traveling to a remote site requires scheduling, which delays repair work and incurs travel costs. However, when receiving instructions from an experienced maintenance technician using a remote conferencing system, there is a problem in that it is difficult to provide accurate instructions verbally or through image sharing.
[0003] Meanwhile, the following prior art exists as a system for understanding work conditions using a virtual space. Patent Document 1 (JP 2021-47610 A) describes a situation understanding support system in which a worker wearing an MR-HMD observes a construction object within a space that is a construction site from various positions and directions, and a terminal device measures the three-dimensional shape of the construction object from images captured by the MR-HMD. The terminal device receives three-dimensional shape data representing the three-dimensional shape of the construction object, and generates an image in which an input field for inspection results related to the construction of the construction object is superimposed on the three-dimensional shape of the construction object as seen by the inspector in a virtual space with a common coordinate system, determined based on the three-dimensional shape data and the position and orientation of the VR-HMD worn by the inspector, and displays the image on the VR-HMD. The system describes a situation understanding support system in which the inspector enters the results of the inspection conducted into the input field while viewing the three-dimensional shape of the construction object displayed on the VR-HMD.
[0004] Furthermore, Patent Document 2 (JP 2006-349578 A) describes a finished form confirmation system that uses a 3D laser scanner to scan the finished form surface and synthesizes 3D point cloud data of the finished form surface into a virtual space constructed in a computer. Next, information about the center line defined in the work site is synthesized into the virtual space, and a vertical virtual plane is constructed and moved to set a virtual skeleton surface. The system then changes the display format of the finished form surface and displays it on the front or back side of the set virtual skeleton surface. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 2021-47610 [Patent Document 2] JP 2006-349578 A Summary of the Invention [Problem to be solved by the invention]
[0006] The situation understanding support system described in Patent Document 1 and the finished product confirmation system described in Patent Document 2 mentioned above do not have a mechanism for sharing the real-time situation on site and the actions of multiple people in remote locations in real time, which makes it difficult to provide appropriate guidance to the site from a remote location.
[0007] The present invention aims to share the real-time situation of an on-site location and the actions of multiple people in remote locations in real time. [Means for solving the problem]
[0008] A representative example of the invention disclosed in the present application is as follows: That is, a virtual three-dimensional space sharing system, comprising: a first display device visible to a first user at a first location; It is a dynamic object whose shape and / or position change.The system includes a first sensor that observes an object and the first user, a second sensor that observes the movement of a second user at a second location different from the first location, and a server that collects data from the first sensor and the second sensor, wherein the server maps the object and the first user observed by the first sensor and the second user observed by the second sensor onto a virtual three-dimensional space, and generates a virtual three-dimensional map of the object and the first user mapped onto the virtual three-dimensional space. For the dynamic object The second user's movement and location information In real time The display device is characterized in that the information is transmitted to the first display device. [Effects of the Invention]
[0009] According to one aspect of the present invention, it is possible to share the real-time situation of an on-site location and the actions of multiple people in remote locations in real time. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating a configuration of an information sharing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the physical configuration of a computer provided in the information sharing system of the present embodiment. [Figure 3] 1 is a logical block diagram of an information sharing system according to an embodiment of the present invention. [Figure 4] FIG. 10 is a diagram showing details of on-site sensing processing in this embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of a database configuration according to the present embodiment. [Figure 6] 10A and 10B are diagrams illustrating an example of an image displayed on the MR glasses of the present embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of an overhead image displayed on the administrator terminal of the present embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] FIG. 1 is a diagram showing the configuration of an information sharing system according to an embodiment of the present invention.
[0012] The information sharing system of this embodiment includes a plurality of three-dimensional sensors 10, an edge processing device 20 connected to the three-dimensional sensors 10, an MEC server 40 that processes the observation results from the three-dimensional sensors 10, a network 30 that connects the edge processing device 20 to the MEC server 40, MR glasses 50, VR glasses 60, a three-dimensional sensor 61 that observes a wearer of the VR glasses 60, and an edge processing device 62 connected to the three-dimensional sensor 61. The information sharing system may also include an administrator terminal 70.
[0013] The three-dimensional sensor 10 is a sensor that observes the situation of a work site to be shared in a virtual three-dimensional space (metaverse space) 100. The three-dimensional sensor 10 can acquire three-dimensional point cloud data. For example, a time-of-flight (TOF) camera that outputs a distance-attached image in which the distance D for each pixel is added to RGB data can be used. Multiple three-dimensional sensors 10 are provided to cover a wide area of the work site, including the worker's work area, and the observation ranges of each three-dimensional sensor 10 should be overlapped. The three-dimensional sensor 10 observes static objects, whose shape and position do not change, such as equipment installed on the site or room structures, as well as dynamic objects, whose shape and position change, such as vehicles, construction machinery, robots, workers, tools, and work targets. The three-dimensional sensor 10 observes the situation of the worker (e.g., the movement and position of a remote worker).
[0014] The edge processing device 20 is a computer that generates three-dimensional information including multiple pieces of three-dimensional model data and a human skeletal model from the point cloud data acquired by the three-dimensional sensor 10. By having the edge processing device 20 generate three-dimensional information from the point cloud data, the amount of communication between the edge processing device 20 and the MEC server 40 can be reduced, and congestion on the network 30 can be alleviated. Note that if there is no problem with the bandwidth of the network 30, the point cloud data may be transmitted directly to the MEC server 40 and then the three-dimensional information may be generated.
[0015] The MEC server 40 is a computer that is provided on the network 30 and realizes edge computing, and in this embodiment, generates a virtual three-dimensional space 100 from three-dimensional information collected from one or more edge processing devices 20.
[0016] The network 30 is a wireless network suitable for data communication that connects the edge processing device 20 and the MEC server 40, and may be, for example, a high-speed, low-latency 5G network. Note that if the edge processing device 20 is installed in a fixed location, a wired network may also be used.
[0017] The MR glasses 50 are display devices that can be seen by workers on-site, and are preferably worn on the head of the worker to share the virtual three-dimensional space 100. The MR glasses 50 have a processor that executes programs, a memory that stores programs and data, a network interface that communicates with the MEC server 40, and a display that displays images (described later with reference to FIG. 6) transmitted from the MEC server 40. The display is a see-through type, and the wearer can see the surroundings through the display. Video of It is preferable that the image can be viewed superimposed on the image transmitted from the MEC server 40. The MR glasses 50 may also have a camera that captures an image in front of the wearer, and transmit the image captured by the camera to the MEC server 40. The MR glasses 50 may also display an image captured by a camera that captures an image in front of the wearer, superimposed on the image transmitted from the MEC server 40. The MR glasses 50 may also have a camera that captures an image of the wearer's eyes, and detect the wearer's line of sight from the image captured by the camera. The MR glasses 50 may also have a microphone that detects sounds that the wearer is hearing.
[0018] Furthermore, on-site workers may wear wearable sensors (for example, tactile gloves). The tactile gloves detect the worker's sense of touch and transmit the information to the MEC server 40. Furthermore, the wearable sensors may detect the movements of the worker's fingers, and a skeletal model of the worker may be generated from the movements of the fingers detected by the wearable sensors to detect the worker's actions.
[0019] The VR glasses 60 are display devices that can be viewed by a person in a remote location away from the work site (hereinafter referred to as a remote person, for example, an experienced worker), and are preferably worn on the head of the worker to share the virtual three-dimensional space 100. The VR glasses 60 include a processor that executes programs, memory that stores programs and data, a network interface that communicates with the MEC server 40, and a display that displays images transmitted from the MEC server 40 (described later with reference to FIG. 6). The VR glasses 60 may also include a camera that captures an image in front of the wearer and transmits the image captured by the camera to the MEC server 40. When the VR glasses 60 are installed outside the network in which the MEC server 40 is installed, the VR glasses 60 and the MEC server 40 may be connected via a public network such as the Internet 80 or another dedicated network. The VR glasses 60 receive motion data, including the movements and positions of the on-site worker represented by a skeletal model, from the MEC server 40, and display the virtual three-dimensional space 100 including an avatar of the on-site worker. The information on the virtual three-dimensional space 100 that the VR glasses 60 receive from the MEC server 40 includes information on the object observed by the three-dimensional sensor 10 in addition to the worker's avatar.
[0020] The three-dimensional sensor 61 is a sensor that observes the situation (for example, the movement and position of the remote person) of the remote person wearing the VR glasses 60 to be shared in the virtual three-dimensional space 100. Like the three-dimensional sensor 10, the three-dimensional sensor 61 may be capable of acquiring three-dimensional point cloud data, and may use, for example, a TOF camera that outputs a distance image in which RGB data is assigned a distance D for each pixel. The remote person may wear a wearable sensor that detects the movement of the fingers of the remote person. The wearable sensor detects the movement of the fingers of the remote person and transmits it to the MEC server 40. The MEC server 40 may generate a skeletal model of the worker from the movement of the fingers detected by the wearable sensor and detect the worker's behavior.
[0021] The edge processing device 62 generates a plurality of three-dimensional model data from the point cloud data acquired by the three-dimensional sensor 61. orIt is a computer that generates three-dimensional information including a human skeletal model. The edge processing device 62 generates three-dimensional information from point cloud data, which reduces the amount of communication between the edge processing device 62 and the MEC server 40. If there is no problem with the amount of communication, the three-dimensional information may be generated after transmitting the point cloud data directly to the MEC server 40.
[0022] The administrator terminal 70 is a computer used by an on-site administrator who uses the information sharing system, and can display information about the virtual three-dimensional space 100 (for example, an overhead image).
[0023] The information sharing system of this embodiment may include a cloud 90 that forms a large-scale virtual three-dimensional space for sharing three-dimensional information collected from multiple MEC servers 40. The large-scale virtual three-dimensional space formed in the cloud 90 is an integration of the virtual three-dimensional spaces formed by the multiple MEC servers 40, and can form a large-scale virtual three-dimensional space over a wide area.
[0024] Access to the MEC server 40 from the MR glasses 50, the VR glasses 60, and the administrator terminal 70 can be authenticated using an ID and password, or the unique address (e.g., MAC address) of these devices to ensure the security of the information sharing system.
[0025] 2 is a block diagram showing the physical configuration of a computer provided in the information sharing system of this embodiment. In FIG. 2, the MEC server 40 is shown as an example of a computer, but the edge processing devices 20, 62 and the administrator terminal 70 may also have the same configuration.
[0026] The MEC server 40 of this embodiment is configured by a computer having a processor (CPU) 1, a memory 2, an auxiliary storage device 3, and a communication interface 4. The MEC server 40 may also have an input interface 5 and an output interface 8.
[0027] The processor 1 is a computing device that executes programs stored in the memory 2. The processor 1 executes various programs to realize various functional units of the MEC server 40 (for example, the metaverse analysis function 400, etc.). Note that some of the processing performed by the processor 1 by executing the programs may be executed by other computing devices (for example, hardware such as a GPU, ASIC, or FPGA).
[0028] The memory 2 includes a ROM, which is a non-volatile storage element, and a RAM, which is a volatile storage element. The ROM stores unchanging programs (e.g., BIOS), etc. The RAM is a high-speed, volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor 1 and data used when the programs are executed.
[0029] The auxiliary storage device 3 is a large-capacity, non-volatile storage device such as a magnetic storage device (HDD) or flash memory (SSD). The auxiliary storage device 3 also stores data used by the processor 1 when executing a program and the program executed by the processor 1. That is, the program is read from the auxiliary storage device 3, loaded into the memory 2, and executed by the processor 1 to realize each function of the MEC server 40.
[0030] The communication interface 4 is a network interface device that controls communication with other devices (for example, the edge processing device 20, the cloud 90) in accordance with a predetermined protocol.
[0031] The input interface 5 is an interface to which input devices such as a keyboard 6 and a mouse 7 are connected and which receives input from an operator. The output interface 8 is an interface to which output devices such as a display device 9 and a printer (not shown) are connected and which outputs the results of program execution in a format that can be viewed by the user. Note that a user terminal connected to the MEC server 40 via a network may provide the input and output devices. In this case, the MEC server 40 may have web server functionality, and the user terminal may access the MEC server 40 using a predetermined protocol (e.g., http).
[0032] The programs executed by the processor 1 are provided to the MEC server 40 via removable media (CD-ROM, flash memory, etc.) or a network, and are stored in a non-volatile auxiliary storage device 3, which is a non-transitory storage medium. For this reason, the MEC server 40 should preferably have an interface for reading data from removable media.
[0033] The MEC server 40 is a computer system configured on a single physical computer or on multiple logically or physically configured computers, and may operate on a virtual computer constructed on multiple physical computer resources. For example, each functional unit may operate on a separate physical or logical computer, or multiple functional units may be combined to operate on a single physical or logical computer.
[0034] FIG. 3 is a logical block diagram of the information sharing system of this embodiment.
[0035] The processing by the information sharing system of this embodiment is executed by an on-site sensing function 200, a remote sensing function 300, a metaverse analysis function 400, and a feedback function 500.
[0036] In the on-site sensing function 200, in on-site sensing and transmission processing 210, the three-dimensional sensor 10 observes the on-site situation and transmits the observed point cloud data to the edge processing device 20. Then, in three-dimensional information generation processing 220, the edge processing device 20 generates three-dimensional information including the point cloud data and three-dimensional model data observed by the three-dimensional sensor 10. The three-dimensional sensor 10 may capture images of a dynamic object installed at the site, and the edge processing device 20 may transmit differential data between the frame of the image of the dynamic object captured by the three-dimensional sensor 10 and a frame from a previous time to the MEC server 40.
[0037] 4, the on-site sensing function 200 is configured such that the edge processing device 20 integrates (221) point cloud data observed by the multiple 3D sensors 10 based on the relationship between the positions and observation directions of the multiple 3D sensors 10. When integrating the point cloud data, an image of the front of the wearer captured by the MR glasses 50 may also be integrated.
[0038] Then, a high-speed 3D modeling process for static objects is performed (222). For example, the outer surface of a static object can be constructed using an algorithm that generates a surface based on the positional relationship of adjacent point clouds. A high-speed 3D modeling process for dynamic objects is also performed (223). For example, a range where shape and position change is extracted from the point cloud data, a skeletal model obtained by skeletal estimation is generated, and a person is modeled. The generated skeletal model represents the position of the person (worker), and the time-series changes in the skeletal model represent the person's movement. Modeling of static objects and modeling of dynamic objects can be performed sequentially, and the order can be any.
[0039] The 3D model is then segmented according to the continuity of the constructed surfaces and the extent of dynamic objects, distinguishing between dynamic and static objects and determining the extent to which they are meaningful as objects (224).
[0040] In addition, the edge processing device 20 collects the gaze direction of the wearer and the sounds the wearer is hearing from the MR glasses 50 and transmits them to the MEC server 40. In the MEC server 40, a metaverse analysis function 400 (described later) recognizes static objects and dynamic objects, and a virtual three-dimensional space 100 is generated.
[0041] In the remote sensing function 300, in motion sensing processing 310, the three-dimensional sensor 61 observes the situation of the remote person and transmits the observed point cloud data to the edge processing device 62. The edge processing device 62 then performs high-speed three-dimensional modeling of dynamic objects on the point cloud data observed by the three-dimensional sensor 61 (310). For example, it extracts from the point cloud data a range in which shape and position change occurs, generates a skeletal model obtained by skeletal estimation, and models the person. The generated skeletal model represents the position of the person (worker), and time-series changes in the skeletal model represent the person's movement. The three-dimensional sensor 61 may capture video on the remote side, and the edge processing device 20 may transmit differential data between the frame of the video captured by the three-dimensional sensor 61 and a frame from an earlier time to the MEC server 40.
[0042] The edge processing device 62 then generates an avatar from the generated skeletal model (320). The edge processing device 62 also collects the wearer's line of sight and the sounds the wearer is listening to from the VR glasses 60 and transmits them to the MEC server 40. The generated skeletal model is transmitted to the MEC server 40 and treated as action B of the remote person. The generated avatar is also transmitted to the MEC server 40 together with sound data listened to by the wearer of the VR glasses 60, incorporated into the virtual three-dimensional space 100, and fed back to the MR glasses 50. The generated avatar may also be fed back directly to the MR glasses 50. The wearer of the MR glasses 50 can share with the remote person the virtual three-dimensional space 100 that incorporates the actions and sensations represented by the movements and positions of the remote person, allowing them to understand the movements of the remote person and even converse with the remote person.
[0043] In the metaverse analysis function 400, the MEC server 40 generates an avatar of a field worker from a skeletal model of a dynamic object recognized by the field-side sensing function 200, and generates an avatar of a remote person from a skeletal model of the remote person generated by the remote-side sensing function 300. A virtual three-dimensional space 100 is generated by mapping these generated avatars and the three-dimensional model data of static objects recognized by the field-side sensing function 200.
[0044] In the object recognition process 410, the MEC server 40 recognizes the segmented three-dimensional model and identifies the object. For example, the type of object can be estimated using a machine learning model that has learned from images of objects installed at the site, or a model that records the three-dimensional shape of the object installed at the site.
[0045] In the action recognition process 420, the MEC server 40 recognizes the worker's action A (type of action) from motion data including the movement and position of the worker on-site represented by the skeletal model. For example, the worker's action can be estimated using motion data based on past changes in the worker's skeletal model and a machine learning model learned from the worker's actions.
[0046] In the skill detection process 430, the MEC server 40 detects the skill level of a worker based on the direction of the worker's line of sight and the sounds heard by the worker. For example, the skill level of a worker can be estimated by a machine learning model trained based on the direction of the worker's line of sight and the sounds heard by the worker while working, and the skill level of the worker. The skill level of a worker may also be estimated by comparing the working time of the worker with the standard working time. For example, the working time but If the time is shorter than the standard work time, it can be determined that the worker is highly skilled.
[0047] In the action recognition process 440, the MEC server 40 recognizes the action B (type of action) of the remote person from changes in the skeletal model of the remote person. For example, the action of the remote person can be estimated using a machine learning model learned from past changes in the skeletal model of the remote person and the actions of the remote person. The action recognition process 420 and the action recognition process 440 may use the same estimation model.
[0048] In task recognition processing 450, the MEC server 40 recognizes task A of the worker from the object identified in object recognition processing 410 and the worker's action A recognized in action recognition processing 420. For example, task A of the worker can be estimated using a machine learning model trained on the object and action A, or a knowledge graph associating objects and actions. Furthermore, task A of the worker may be recognized using action B of a remote worker recognized in action recognition processing 440.
[0049] In the structuring and storage process 460, the MEC server 40 records task A recognized in the task recognition process 450 in a database 470. In the database 470, the object used to recognize task A, action A, motion data resulting from changes in the skeletal model in action A, action B, and motion data including the movement and position of the worker on-site represented by the skeletal model in action B are registered as related information. A detailed configuration example of the database 470 will be described with reference to FIG. 5.
[0050] In the feedback function 500, the MEC server 40 searches the database 470 using the recognized worker's action A as a key, and transmits the feedback information obtained from the database 470 to the MR glasses 50. The information fed back to the MR glasses 50 includes an avatar generated from motion data of the same task in the same process previously performed, video of the same task previously performed, and work instructions for the next process of the task. In particular, the avatar and video of the task may provide data of the same task performed by a remote worker. The information fed back to the MR glasses 50 may be changed according to the proficiency level and worker attributes estimated by the proficiency detection process 430. For example, detailed information may be provided to less skilled workers, and general information may be provided to more skilled workers. The feedback function 500 allows a worker wearing the MR glasses 50 to automatically obtain information related to their own action A.
[0051] The feedback function 500 not only provides feedback to the MR glasses 50 but also provides feedback to equipment (e.g., robots, construction machinery, vehicles) to Feedback can be used to give commands, which allows changes in the virtual three-dimensional space to be reflected in the real world, enabling various machines to be controlled.
[0052] Fig. 5 is a diagram showing an example of the configuration of the database 470 of this embodiment. Although Fig. 5 shows the database 470 in a table format, it may be configured in another data structure.
[0053] The database 470 includes pre-recorded work-related information 471 and work acquisition information 472 acquired in accordance with the actions of the worker.
[0054] The work-related information 471 stores a work ID, a work reference time, a work manual, work video content, and work text content in association with each other. The work ID is pre-recorded identification information for the work. The work reference time is the standard time for the work performed by the worker. The work manual is an instruction manual for the work performed by the worker, and link information for accessing the instruction manual may be recorded. The work video content is a video of the work performed by the worker, performed by an expert or previously performed by the worker, and link information for accessing the video may be recorded. The work text content is text information related to the work performed by the worker, and link information for accessing the text information may be recorded.
[0055] The task acquisition information 472 stores an action ID, actual task time, environmental objects, worker motion, worker position, worker viewpoint, worker sound field, worker tactile sense, worker vitals, worker proficiency, task ID, and task log in association with each other. The task ID is identification information for a task, which is a series of actions performed by a worker. The actual task time is the time required for the worker's task. The environmental objects are objects (e.g., a room, a floor, a device, a tool, a screw) photographed in relation to the worker's task. The worker motion is stored as a log of the feature points (fingers, arms, etc.) of the worker's skeletal model. jointThe worker position is the position of the worker's feature points (head, left and right hands, etc.) and their positional relationship (distance, direction) to environmental objects. The worker viewpoint is the worker's line of sight and the intersection of the line of sight and the surface of an object in the line of sight. The worker sound field is the sound heard by the worker, and link information for accessing the sound data may be recorded. The worker tactile sense is the worker's tactile sense obtained with tactile gloves. The worker vital signs include the worker's voice, facial expression, and pulse rate estimated from changes in blood flow, and are used to estimate the worker's emotions and attributes. The worker proficiency level is the worker's proficiency level detected by the proficiency detection process 430. The task ID is the worker's task recognized by the task recognition process 450. The task log is the results of the task, and records whether it was completed successfully, re-tasked, or abnormally.
[0056] FIG. 6 is a diagram showing an example of an image that is fed back to the field worker in the information sharing system of this embodiment and displayed on the MR glasses 50.
[0057] As shown in Fig. 6, the worker video is displayed so that the remote worker's VR glasses (head position) 601 and hands 602 are superimposed on the real landscape. In Fig. 6, the remote worker's avatar is configured with VR glasses and hands, but an avatar representing the remote worker's entire body may be generated and displayed. Furthermore, the worker video may display worker attributes 611, a work manual 612, and work instructions 613. Furthermore, the worker video may display an avatar (not shown) generated from a skeletal model of another person present at the work site.
[0058] On-site workers can see the actions of remote workers through avatars via worker video, and can receive appropriate work instructions from skilled workers in remote locations.
[0059] FIG. 7 is a diagram showing an example of an overhead image displayed on the administrator terminal 70 in the information sharing system of this embodiment.
[0060] The overhead image displayed on the administrator terminal 70 shows the VR glasses (head position) 701 and hands 702 of a remotely located expert, the avatar 711 of a field worker, and an environmental object (work object) 721 superimposed on the image of three-dimensional space.
[0061] The bird's-eye view allows on-site managers to monitor events in the virtual three-dimensional space, confirm the guidance that workers are receiving from experts, and manage the work of workers.
[0062] The MEC server 40 may adjust the range of the image so that the feedback video (FIG. 6) or the overhead image (FIG. 7) includes only the necessary image range. The overhead image shown in FIG. 7 shows an example in which only the work object, the worker's avatar, the remote worker's avatar, and the surrounding background are displayed, with the other background removed. For example, the range of the feedback image may be adjusted to exclude static and dynamic objects recognized in the segmentation process (224) that are not related to the work, or a mosaic process may be applied to the area that does not include objects that are not related to the work, thereby concealing the unrelated objects. When used in a factory or other on-site environment, this can prevent the leakage of factory space and customer information to the outside.
[0063] As described above, according to the embodiment of the present invention, the real-time situation at the site and the actions of multiple people in remote locations can be shared in real time, making it possible to provide appropriate guidance to the site from a remote location.
[0064] The present invention is not limited to the above-described embodiments, but includes various modifications and equivalent configurations within the spirit and scope of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to configurations including all of the described configurations. Furthermore, part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment may be added, deleted, or replaced with other configurations.
[0065] Furthermore, the aforementioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole in hardware, for example by designing them as integrated circuits, or may be realized in software by having a processor interpret and execute a program that realizes each function.
[0066] Information such as programs, tables, and files that realize each function can be stored in a storage device such as a memory, a hard disk, or an SSD (Solid State Drive), or in a recording medium such as an IC card, an SD card, or a DVD.
[0067] In addition, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily represent all the control lines and information lines that are necessary for implementation. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0068] 1 processor 2. Memory 3 Auxiliary storage 4. Communication Interface 5 Input Interface 6 Keyboard 7. Mouse 8 Output Interfaces 9 Display Devices 10 Three-dimensional sensor 20 Edge processing equipment 30 Network 40 MEC Server 50 MR Glass 60 VR Glasses 61 Three-dimensional sensor 62 Edge Processing Equipment 70 Administrator terminal 80 Internet 90 Cloud 100 Virtual 3D Space 200 On-site sensing function 210 Transmission Processing 220 Three-dimensional information generation processing 300 Remote sensing function 310 Motion Sensing Processing 400 Metaverse Analysis Function 410 Object Recognition Processing 420 Motion Recognition Processing 430 Skilled Sensing Processing 440 Motion Recognition Processing 450 Task Recognition Processing 460 Accumulation Processing 470 databases 471 Work-related information 472 Work Acquisition Information 500 Feedback Function
Claims
1. A virtual three-dimensional space sharing system, a first display device viewable by a first user at a first location; a first sensor for observing an object that is a dynamic object whose shape and / or position changes at the first location and the first user; a second sensor that observes the movement of a second user at a second location different from the first location; a server that collects data from the first sensor and the second sensor; The server mapping the object and the first user observed by the first sensor and the second user observed by the second sensor into a virtual three-dimensional space; A virtual three-dimensional space sharing system characterized by transmitting information on the movement and position of the second user relative to the dynamic object mapped in the virtual three-dimensional space to the first display device in real time.
2. 2. The virtual three-dimensional space sharing system according to claim 1, a second display device viewable by the second user at the second location; A virtual three-dimensional space sharing system characterized in that the server transmits information on the object mapped in the virtual three-dimensional space and the movement and position of the first user to the second display device.
3. 2. The virtual three-dimensional space sharing system according to claim 1, a third sensor configured to detect at least one of a sound perceived by the first user, a line of sight of the first user, and a tactile sense of the first user; the third sensor transmits detected information to the server; A virtual three-dimensional space sharing system, characterized in that the first sensor observes the movement of the first user.
4. 2. The virtual three-dimensional space sharing system according to claim 1, a first edge device to which the first sensor is connected; the first sensor captures an image of the object installed at the first location; A virtual three-dimensional space sharing system characterized in that the first edge device transmits differential data between a frame of the image of the object captured by the first sensor and a frame from a previous time to the server.
5. 2. The virtual three-dimensional space sharing system according to claim 1, a first edge device to which the first sensor is connected; the first sensor acquires information about the movement and position of the first user; A virtual three-dimensional space sharing system characterized in that the first edge device transmits to the server a skeletal model generated from information on the movement and position of the first user acquired by the first sensor.
6. 2. The virtual three-dimensional space sharing system according to claim 1, a fourth sensor configured to detect at least one of a sound perceived by the second user, a line of sight of the second user, and a tactile sense of the second user; The virtual three-dimensional space sharing system is characterized in that the fourth sensor transmits detected information to the server.
7. 2. The virtual three-dimensional space sharing system according to claim 1, a second edge device to which the second sensor is connected; the second sensor captures an image of the second location; A virtual three-dimensional space sharing system characterized in that the second edge device transmits differential data between a frame of video captured by the second sensor and a frame from a previous time to the server.
8. 2. The virtual three-dimensional space sharing system according to claim 1, a second edge device to which the second sensor is connected; the second sensor acquires information on the movement and position of the second user; The second edge device A virtual three-dimensional space sharing system characterized in that a skeletal model generated from information on the movement and position of the second user acquired by the second sensor is transmitted to the server.
9. 2. The virtual three-dimensional space sharing system according to claim 1, The virtual three-dimensional space sharing system is characterized in that the server records skeletal models generated from images of the first user and the second user in a database.
10. The virtual three-dimensional space sharing system according to claim 9, The server Recognizing the object from the result of observation of the object by the first sensor; Identifying an operation of the first user based on a relationship between a skeletal model generated from the video of the first user and the recognized object; A virtual three-dimensional space sharing system characterized in that the specified work is recorded in the database.
11. 2. The virtual three-dimensional space sharing system according to claim 1, a fifth sensor that detects at least one of the voice, blood flow, and facial expression of the first user; A virtual three-dimensional space sharing system characterized in that at least one of the proficiency and attributes of the first user is estimated from at least one of the voice, blood flow, and facial expression detected by the fifth sensor.
12. The virtual three-dimensional space sharing system according to claim 11, A virtual three-dimensional space sharing system, wherein the server changes information to be transmitted to the first display device in accordance with at least one of the estimated skill level and attribute.
13. The virtual three-dimensional space sharing system according to claim 10, The server retrieves information related to the identified task from the database; A virtual three-dimensional space sharing system characterized in that information acquired from the database is transmitted to the first display device.
14. 2. The virtual three-dimensional space sharing system according to claim 1, a terminal connected to the server, A virtual three-dimensional space sharing system characterized in that the server transmits data of the virtual three-dimensional space in which the object, the first user, and the second user are mapped to the terminal.
15. A computer-executed virtual three-dimensional space sharing method, comprising: The computer A computing device that executes predetermined computational processing and a storage device that can be accessed by the computing device, a first display device visible to a first user at a first location, a first sensor installed at the first location, and a second sensor installed at a second location different from the first location; The virtual three-dimensional space sharing method includes: the computing device collects data of the first user and an object that is a dynamic object whose shape and / or position change and that is observed by the first sensor at the first location, and data of the second user that is observed by the second sensor at the second location; the computing device maps the object and the first user observed by the first sensor and the second user observed by the second sensor into a virtual three-dimensional space; A virtual three-dimensional space sharing method characterized in that the computing device transmits to the first display device in real time information on the movement and position of the second user relative to the object, which is a dynamic object, mapped in the virtual three-dimensional space.
16. A virtual three-dimensional space sharing server, A computing device that executes predetermined computational processing and a storage device that can be accessed by the computing device, a first display device visible to a first user at a first location, a first sensor installed at the first location, and a second sensor installed at a second location different from the first location; collecting data of the first user and an object that is a dynamic object whose shape and / or position change and that is observed by the first sensor at the first location, and data of the second user that is observed by the second sensor at the second location; mapping the object and the first user observed by the first sensor and the second user observed by the second sensor into a virtual three-dimensional space; A virtual three-dimensional space sharing server characterized by transmitting information on the movement and position of the second user relative to the object, which is a dynamic object, mapped in the virtual three-dimensional space to the first display device in real time.
Citation Information
Patent Citations
Finished form confirmation system, method and program
JP2006349578A
Image generation device and image generation method
JP2014017776A
Method and system for emotion and behavior recognition
JP2015130151A
Operating method and system for participating in a virtual reality scene
JP2019522856A
Remote work support system
JP2021010101A