A method for building a cloud-edge-device collaborative mixed reality environment for industrial applications
By building a cloud-edge-end collaborative computing task scheduling model and a mixed reality co-location collaborative shared space, and combining PUN2, Azure Spatial Anchors, and WebRTC technologies, the problems of unclear technical details and imperfect human-computer interaction in mixed reality collaboration in industrial applications are solved, and an efficient and stable multi-terminal collaboration environment is achieved.
Patent Information
- Application Number
- CN202310205419.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-03-06
AI Technical Summary
Existing mixed reality collaboration methods have unclear technical details in industrial applications, are not closely related to reality, are not universal, have imperfect human-computer interaction models, and lack research on co-location collaboration, resulting in low collaboration efficiency in complex industrial scenarios.
Build a cloud-edge-end collaborative computing task scheduling model, establish a mixed reality co-location collaborative shared space, realize multimodal interaction, and optimize the collaboration process through a multi-terminal conflict resolution model. Combine PUN2, Azure Spatial Anchors and WebRTC technologies to ensure synchronization and coordination between terminals.
It improves the real-time performance of industrial collaboration, enhances the collaborative experience of virtual-reality integration, solves the problems of high latency and poor interactive intuitiveness in existing technologies, and achieves stability and efficiency in multi-terminal collaboration.
Smart Images

Figure CN116489149B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mixed reality industrial application technology, and specifically relates to a method for constructing a cloud-edge-end collaborative mixed reality collaborative environment for industrial applications. Background Art
[0002] Multi-person collaboration is a technology that allows multiple people to work together to complete a task, whether in the same or different spatial contexts. For complex tasks, multi-person collaboration can effectively execute tasks through communication between team members. In recent years, with the rapid development of next-generation information technology, multi-person collaboration is often achieved through computer-supported collaborative work (CSCW).
[0003] Due to the complexity of tasks and processes, industrial production activities require division of labor and collaboration in all aspects of planning, design, production, operation and maintenance, and monitoring. Therefore, industry is one of the sectors where CSCW is most widely used.
[0004] Industrial sites often feature complex environments and numerous devices. When designing and developing collaborative environments, ensuring the rapid and accurate transmission of information between team members is crucial. Traditional CSCW methods, such as text, voice, and video, have limitations in communication efficiency. Mixed reality (MR) collaboration promises to address this pain point. MR technology can seamlessly integrate interactive virtual information with the physical world in real time within the same three-dimensional space, enhancing visualization and transcending geographical and spatial limitations, making collaboration more natural and intuitive.
[0005] Publication number CN111553974A, "A data visualization remote assistance method and system based on mixed reality", includes the following steps: starting a communication connection, obtaining live video and audio information and sharing it in real time, the remote terminal user analyzing the problems of the on-site equipment and prompting feasible solutions, sending the feasible solutions to the on-site terminal, the on-site terminal user operating according to the guidance prompts of the remote terminal user, and disconnecting the communication connection. It can greatly improve the efficiency and effectiveness of remote assistance.
[0006] Publication No. CN112667179A, titled "A Remote Synchronous Collaboration System Based on Mixed Reality," includes a remote expert client and a local user client. The local user client, located in a work environment, consists of an augmented reality head-mounted display (HMD), a local tracking module, a depth camera, and a local computer. The remote expert client also consists of a virtual reality HMD, a handheld controller, a remote tracking module, and a remote computer. The remote expert can issue suggestions and instructions within the virtual reality environment. These instructions are displayed from the local user's perspective via the HMD, providing intuitive guidance for the local user's work.
[0007] The above two solutions provide mixed reality remote collaboration methods that can realize real-time synchronous collaboration between the local and remote ends, but both have the following limitations: 1) They only involve the components of the remote collaboration system and the implementation process of remote collaboration, and do not involve technical details such as how to establish communication between terminals and how remote terminal users guide on-site terminal users; 2) They only provide the theoretical framework and technical route of mixed reality remote collaboration, and do not combine mixed reality collaboration with actual industrial application scenarios.
[0008] Authorization announcement number CN106339094B, "Interactive remote expert collaborative maintenance system and method based on augmented reality technology", the on-site terminal captures video or takes images of the equipment to be inspected and repaired, and transmits them to the remote server. The remote server receives the video or image sent by the on-site terminal, receives the expert's delineation processing of the video or image, and sends the delineated target area image to the on-site terminal; uses augmented reality technology to synthesize augmented reality video, and completes fault diagnosis and maintenance of on-site equipment with the collaboration of experts.
[0009] The above solution applies augmented reality remote collaboration to equipment inspection and maintenance, freeing customers' hands and helping them find problems quickly. However, it does not address the issue of human-computer interaction in remote collaboration, does not regard people as the core elements of the industrial production system, and cannot establish a deep connection between people and the industrial environment.
[0010] Publication number CN111679740A, "Method for Remote Intelligent Diagnosis of Power Plant Equipment Using Augmented Reality (AR) Technology," achieves "human-machine" connection for remote diagnosis through gestures and postures; utilizes "end-cloud," AI data modeling, and visualization technologies to achieve data connectivity and application; and utilizes AR technology to integrate the physical and virtual environments, allowing all collaborating parties to share first-person perspective images, providing a scenario where a virtual environment and the real world are superimposed, making collaboration more accurate and efficient. This has positive significance for improving the safety, reliability, and availability of power plant equipment and reducing diagnostic and maintenance costs.
[0011] The above solution applies augmented reality remote collaboration to intelligent diagnosis of power plant equipment, but has the following limitations: 1) The human-computer interaction model needs to be improved. The "human-computer" connection for remote diagnosis is achieved through gestures and postures, but it does not take into account the situation where the technician's hands are occupied; 2) The remote diagnosis center is connected to the AR field terminal through "end-to-cloud" technology, but "end-to-cloud" collaboration will bring high latency. Applying it to industrial scenarios with strict latency requirements may affect timeliness.
[0012] The above two solutions combine mixed reality remote collaboration with specific industrial application scenarios, but the solutions provided are only aimed at collaboration in a certain task or link in a certain industrial category, and are limited in terms of versatility and scalability.
[0013] All of the above solutions only involve remote mixed reality collaboration, not co-located mixed reality collaboration. In fact, throughout the industrial production process, there are many scenarios that require face-to-face communication and collaboration, such as collaborative product design, collaborative operation plan planning, and discussion and analysis of production operations. Therefore, the establishment of a mixed reality co-located collaboration environment is also very necessary. Summary of the Invention
[0014] In order to solve the problems raised in the background technology, such as unclear technical details of existing mixed reality collaboration methods or systems, loose connection with industrial reality, weak versatility, imperfect human-computer interaction models, and lack of research on co-location collaboration, the present invention provides a method for constructing a cloud-edge-end collaborative mixed reality collaboration environment for industrial applications.
[0015] The present invention adopts the following technical solution: a method for constructing a cloud-edge-end collaborative mixed reality collaborative environment for industrial applications, including the following steps: S1: constructing a computing task scheduling model for cloud-edge-end collaboration; S2: constructing a mixed reality co-location collaborative shared space; S3: constructing a mixed reality remote collaboration environment; S4: constructing a multimodal interaction model; S5: constructing a multi-terminal conflict resolution model.
[0016] The step S1 includes the following steps:
[0017] S101: Build a mixed reality computing cloud platform; S102: Deploy edge nodes; S103: Build a mixed reality collaborative terminal system; S104: Task offloading and resource allocation.
[0018] The mixed reality computing cloud platform described in S101 is used to provide edge node management and provide core business logic processing related services for edge applications; the edge nodes described in S102 are deployed at the industrial site, including edge cloud, edge gateway and edge controller; the mixed reality collaboration terminal system described in S103 is a collection of all terminals used by users to complete mixed reality collaboration; the task offloading and resource allocation described in S104 refers to deploying tasks with high real-time requirements such as audio and video encoding and decoding, core image rendering, hand tracking and motion tracking on the edge side, while deploying non-core tasks that are not sensitive to delay such as image rendering and natural voice interaction on the cloud.
[0019] The step S2 comprises the following steps:
[0020] S201: Construction of a cross-platform multi-terminal network communication framework based on PUN2; S202: Spatial state synchronization based on Azure SpatialAnchors.
[0021] PUN2 described in S201 is a high-performance state synchronization network library for Unity3d, which can be naturally integrated into the common Unity workflow. The default protocol is UDP, with a reliability protocol on top, and it also supports TCP and Websocket. Azure Spatial Anchors described in S202 is a spatial cloud anchor that can save information such as the position and orientation of virtual content on the anchor so that the virtual images in each terminal have the same position and state in the world coordinate system, ensuring consistency of collaboration and enabling cross-platform use.
[0022] The step S201 specifically includes the following steps:
[0023] The mixed reality co-location collaborative shared space includes a cross-platform multi-terminal network communication framework and a spatial state synchronization system, wherein the cross-platform multi-terminal network communication framework is used for communication between multiple terminals, and the spatial state synchronization system is used to achieve spatial alignment between multiple terminals; the mixed reality remote collaboration environment includes a mixed reality remote audio and video communication module and a mixed reality remote holographic annotation module, the mixed reality remote audio and video communication module realizes remote audio and video communication between the communicating parties, and the mixed reality remote holographic annotation module is used for remote experts and on-site staff to jointly annotate spatial anchor points; the multimodal interaction model is used for interaction between on-site staff and remote experts; the multi-terminal conflict resolution model is used to process and analyze the operation data of multiple terminals, determine whether there is a conflict, and sort the priorities when a conflict exists.
[0024] S2011: Connect multiple terminals to the main server of the mixed reality computing cloud platform. The main server will be responsible for all terminal-to-server transmissions, and the load balancing function of the mixed reality computing cloud platform will be responsible for coordinating all available rooms; S2012: The main terminal creates a room in the main server, sets the room name and maximum connection number parameters for the room, and waits for connections from other terminals; S2013: Other terminals access and join the room by indexing the room name, so that all terminals are in the same room; S2014: Use the event system to try to synchronize content, and all necessary data can be distributed among the participants in the room; when the remaining terminals receive this information, a direct connection between the terminals is established; whenever a terminal connects to the main server, it synchronizes a delay-corrected timestamp; which can be used in the room to synchronize the time of events.
[0025] The step S202 specifically includes the following steps:
[0026] S2021: Create a spatial anchor resource in the Azure portal; S2022: Deploy a shared anchor service, that is, deploy an ASP.NET Core Web application in Azure that can be used for shared anchors; S2023: Configure and deploy the Unity project. First, import the ASA SDK and OpenXR plug-in into the project, then connect the Unity 3D scene to the Azure resource, and finally add and configure the SpatialAnchorManager interface to call the ASA service; S2024: Synchronize spatial states by creating anchors, retrieving anchors, and sharing anchors. First, one of the multiple terminals starts a session by calling the StartSessionAsync() method; then, an anchor is created through the CreateAnchor() method, and the position, rotation data, and environmental data around the anchor are collected, and the anchor is saved. Finally, the terminal shares the anchor ID to the network through the ShareAzureAnchorIdToNetwork() function, and the remaining terminals obtain the shared anchor ID from the network through the GetAzureAnchorIdFromNetwork() function, thereby achieving spatial alignment between multiple terminals.
[0027] The step S3 comprises the following steps:
[0028] S301: Implementation of remote audio and video communication function of mixed reality based on WebRTC; S302: Implementation of remote holographic annotation function of mixed reality.
[0029] The step S301 specifically includes the following steps:
[0030] S3011: Set up signaling server;
[0031] S3012: The communicating parties establish a connection with the signaling server via Websocket.
[0032] S3013: The communicating parties send signals to the signaling server through the RTCPeerConnection API and establish a peer-to-peer connection by exchanging Session Description Protocol information.
[0033] S3014: The communicating parties collect audio and video streaming media or other data through the MediaStream API and transmit them using the RTP / SRTP protocol.
[0034] The step S302 specifically includes the following steps:
[0035] S3021: On-site staff establish audio and video communication with remote experts, and on-site staff transmit real-time first-person video streams of the industrial site to remote experts;
[0036] S3022: The on-site staff sends a remote guidance request to the remote expert. After the remote expert agrees to the request, the video stream is automatically frozen and the annotation tool is called up.
[0037] S3023: The remote expert uses a mouse or touchscreen to click on a frozen still frame in a 2D device, transmits the relative click position within the video to the on-site worker, and creates a spatial anchor point at this position for storing annotation data.
[0038] S3024: Remote experts annotate frozen still frames;
[0039] S3025: The coordinates of the 2D annotation are converted to 3D space, and the converted annotation information is stored in the Azure spatial anchor. The on-site staff reads the annotation information stored in the spatial anchor, performs collision detection with the reconstructed 3D scene, and calculates the actual position of the annotation in the 3D space.
[0040] S3026: The converted annotations are positioned in the real space, forming a remote guidance that integrates virtual and real space;
[0041] S3027: Synchronize the annotated video stream from the terminal worn by the on-site staff to the remote expert. Both parties can now see the annotations in the real-time video.
[0042] The step S4 comprises the following steps:
[0043] S401: Determine the interaction mode; S402: Develop a multimodal interaction strategy.
[0044] Preferably, the interaction modalities in step S401 include gesture, voice, gaze, BCI, and keyboard and mouse; wherein, gesture, voice, gaze, and BCI are used by on-site staff and are implemented based on head-mounted augmented reality devices and brain-computer interface devices; PC is used by remote experts to interact with the mouse and keyboard to provide guidance to on-site staff;
[0045] Preferably, the multimodal interaction strategy in step S402 is: under normal conditions, on-site workers give priority to using gesture recognition to complete the required interaction; when the on-site workers' hands are occupied, a combination of gaze and voice is used as a backup interaction strategy, using gaze rays to select the operation object, and then using voice commands to confirm the operation, thereby reducing the probability of misoperation; when an emergency occurs at the industrial site and there is no time to perform operations through other interaction modes, the BCI system is used to analyze the instinctive EEG signals of the on-site workers and execute corresponding commands.
[0046] The multi-terminal conflict resolution model in step S5 is as follows: the system processes and analyzes the operation data of multiple terminals to determine whether there is a conflict. If there is no conflict between multiple operation instructions, each instruction can be executed directly; if there is a conflict between multiple operation instructions, the operation with low timeliness requirements and can be postponed is defined as a normal operation, and the operation that must be executed immediately is defined as an emergency operation. If the operations performed by each user are all normal operations, a negotiation-based conflict resolution method is adopted, and each user can fully discuss and make a new decision; if there is an emergency operation among the operations performed by each user, the number of emergency operations must be further determined; if there is only one emergency operation, this operation is executed first; if there are multiple emergency operations, the conflict resolution is performed according to the principle of priority of the high-priority subject, that is, the operation priority of each user is different, and the strategy with the high priority takes precedence.
[0047] The principle of the above priority division is: on-site staff have higher priority than remote experts, and among all on-site staff, the more experienced they are, the higher their priority, and the same applies to remote experts.
[0048] In the present invention, the cloud-edge-end collaborative computing task scheduling model constructed in step S1 is the basis for the efficient completion of subsequent steps; the mixed reality co-location collaborative shared space and the mixed reality remote collaborative shared space constructed in steps S2 and S3 form the main body of the mixed reality collaborative environment; the multimodal interaction model constructed in step S4 provides specific operation instruction issuance methods and strategies for on-site staff and remote experts in the mixed reality collaborative environment; the multi-terminal conflict resolution model constructed in step S5 can analyze the operation instructions issued by multiple users and sort the priorities when there is a conflict.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] (1) The present invention constructs a computing task scheduling model for cloud-edge-end collaboration, which overcomes the disadvantage of high latency in the end-cloud collaboration model. It not only effectively utilizes the cloud computing power, but also reduces the latency of collaborative tasks, greatly improving the real-time performance of mixed reality collaboration, making it more suitable for delay-sensitive industrial scenarios.
[0051] (2) The present invention builds a mixed reality co-location collaborative shared space based on PUN2 and Azure Spatial Anchors, enabling multiple workers to observe and discuss together in a virtual-reality integrated workspace, communicate with each other naturally, work in parallel, expand the dimension of perceived information, and improve the interaction process.
[0052] (3) The present invention realizes mixed reality audio and video communication based on WebRTC, and further realizes mixed reality remote holographic annotation. The clear and concise visual assistance can greatly save communication costs, while accurately locating problem points and avoiding misoperation, which effectively solves the problems of intuitiveness, guidance, immersion and other problems of existing remote solutions.
[0053] (4) The present invention constructs a multimodal human-computer interaction model for a mixed reality collaborative environment, fully utilizing and integrating the functions of each modality, selecting the appropriate modality according to the scenario and task, achieving coordination between humans and machines, and solving the problem that a single modality input cannot meet the needs of complex and changing real industrial scenarios.
[0054] (5) The present invention constructs a multi-terminal conflict resolution model, which effectively avoids data inconsistency caused by multi-terminal collaboration, enables the industrial system to perform the most appropriate operation at the moment during collaborative work, and improves the stability of the collaboration process.
[0055] (6) The method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications provided by this invention provides a specific and universal solution for applying mixed reality technology to industrial scenarios. It can be effective throughout the entire process of industrial system planning, design, production, operation and maintenance, and monitoring. The mixed reality collaborative environment is human-centric in terms of experience and also serves people, truly achieving the effects of improving quality, increasing efficiency, and reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a flow chart of a method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications provided by the present invention;
[0057] Figure 2 This is a schematic diagram of a computing task scheduling model for cloud-edge-end collaboration in a method for constructing a cloud-edge-end collaborative mixed reality collaborative environment for industrial applications provided by the present invention;
[0058] Figure 3 This is a flow chart of establishing a cross-platform multi-terminal network communication framework based on PUN2 in a method for constructing a cloud-edge-end collaborative mixed reality collaborative environment for industrial applications provided by the present invention;
[0059] Figure 4 This is a flowchart of spatial state synchronization based on Azure Spatial Anchors in a method for building a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications provided by the present invention;
[0060] Figure 5This is a schematic diagram of the implementation principle of the mixed reality remote audio and video communication function based on WebRTC in the cloud-edge-end collaborative mixed reality collaborative environment construction method for industrial applications provided by the present invention;
[0061] Figure 6 This is a flow chart of a mixed reality remote holographic annotation function in a method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications provided by the present invention;
[0062] Figure 7 This is a schematic diagram of a multimodal human-computer interaction model for a mixed reality collaborative environment in a method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications provided by the present invention;
[0063] Figure 8 This is a flow chart of a multi-terminal conflict resolution model in a cloud-edge-end collaborative mixed reality collaborative environment construction method for industrial applications provided by the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention, but are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without making creative work shall fall within the scope of protection of the present invention.
[0065] As attached Figure 1 As shown, a method for constructing a cloud-edge-end collaborative mixed reality collaborative environment for industrial applications includes the following steps: S1: constructing a computing task scheduling model for cloud-edge-end collaboration; S2: constructing a mixed reality co-location collaborative shared space; S3: constructing a mixed reality remote collaboration environment; S4: constructing a multimodal interaction model; and S5: constructing a multi-terminal conflict resolution model.
[0066] The computing task scheduling model of cloud-edge-end collaboration in step S1 is as shown in the attached Figure 2 As shown, the specific steps include:
[0067] S101: Build a mixed reality computing cloud platform; S102: Deploy edge nodes; S103: Build a mixed reality collaborative terminal system; S104: Task offloading and resource allocation.
[0068] The mixed reality computing cloud platform described in S101 is used to provide edge node management and provide core business logic processing related services for edge applications; the edge nodes described in S102 are deployed at the industrial site, including edge cloud, edge gateway and edge controller; the mixed reality collaboration terminal system described in S103 is a collection of all terminals used by users to complete mixed reality collaboration; the task offloading and resource allocation described in S104 refers to deploying tasks with high real-time requirements such as audio and video encoding and decoding, core image rendering, hand tracking and motion tracking on the edge side, while deploying non-core tasks that are not sensitive to delay such as image rendering and natural voice interaction on the cloud.
[0069] The step S2 comprises the following steps:
[0070] S201: Construction of a cross-platform multi-terminal network communication framework based on PUN2; S202: Spatial state synchronization based on Azure SpatialAnchors.
[0071] PUN2 described in S201 is a high-performance state synchronization network library for Unity3d, which can be naturally integrated into the common Unity workflow. The default protocol is UDP, with a reliability protocol on top, and it also supports TCP and Websocket. Azure Spatial Anchors described in S202 is a spatial cloud anchor that can save information such as the position and orientation of virtual content on the anchor so that the virtual images in each terminal have the same position and state in the world coordinate system, ensuring consistency of collaboration and enabling cross-platform use.
[0072] The process of building a cross-platform multi-terminal network communication framework based on PUN2 in step S201 is shown in the attached figure. Figure 3 As shown, the specific steps include:
[0073] The mixed reality co-location collaborative shared space includes a cross-platform multi-terminal network communication framework and a spatial state synchronization system, wherein the cross-platform multi-terminal network communication framework is used for communication between multiple terminals, and the spatial state synchronization system is used to achieve spatial alignment between multiple terminals; the mixed reality remote collaboration environment includes a mixed reality remote audio and video communication module and a mixed reality remote holographic annotation module, the mixed reality remote audio and video communication module realizes remote audio and video communication between the communicating parties, and the mixed reality remote holographic annotation module is used for remote experts and on-site staff to jointly annotate spatial anchor points; the multimodal interaction model is used for interaction between on-site staff and remote experts; the multi-terminal conflict resolution model is used to process and analyze the operation data of multiple terminals, determine whether there is a conflict, and sort the priorities when a conflict exists.
[0074] S2011: Connect multiple terminals to the main server of the mixed reality computing cloud platform. The main server will be responsible for all terminal-to-server transmissions, and the load balancing function of the mixed reality computing cloud platform will be responsible for coordinating all available rooms; S2012: The main terminal creates a room in the main server, sets the room name and maximum connection number parameters for the room, and waits for connections from other terminals; S2013: Other terminals access and join the room by indexing the room name, so that all terminals are in the same room; S2014: Use the event system to try to synchronize content, and all necessary data can be distributed among the participants in the room; when the remaining terminals receive this information, a direct connection between the terminals is established; whenever a terminal connects to the main server, it synchronizes a delay-corrected timestamp; which can be used in the room to synchronize the time of events.
[0075] The process of spatial state synchronization based on Azure Spatial Anchors in step S202 is as shown in the attached figure. Figure 4 As shown, the specific steps include:
[0076] S2021: Create a spatial anchor resource in the Azure portal; S2022: Deploy a shared anchor service, that is, deploy an ASP.NET Core Web application in Azure that can be used for shared anchors; S2023: Configure and deploy the Unity project. First, import the ASA SDK and OpenXR plug-in into the project, then connect the Unity 3D scene to the Azure resource, and finally add and configure the SpatialAnchorManager interface to call the ASA service; S2024: Synchronize spatial states by creating anchors, retrieving anchors, and sharing anchors. First, one of the multiple terminals starts a session by calling the StartSessionAsync() method; then, an anchor is created through the CreateAnchor() method, and the position, rotation data, and environmental data around the anchor are collected, and the anchor is saved. Finally, the terminal shares the anchor ID to the network through the ShareAzureAnchorIdToNetwork() function, and the remaining terminals obtain the shared anchor ID from the network through the GetAzureAnchorIdFromNetwork() function, thereby achieving spatial alignment between multiple terminals.
[0077] The step S3 comprises the following steps:
[0078] S301: Implementation of remote audio and video communication function of mixed reality based on WebRTC; S302: Implementation of remote holographic annotation function of mixed reality.
[0079] The principle of implementing the mixed reality remote audio and video communication function based on WebRTC in step S301 is as shown in the attached Figure 5 As shown, the specific steps include:
[0080] S3011: Set up signaling server;
[0081] S3012: The communicating parties establish a connection with the signaling server via Websocket.
[0082] S3013: The communicating parties send signals to the signaling server through the RTCPeerConnection API and establish a peer-to-peer connection by exchanging Session Description Protocol information.
[0083] S3014: The communicating parties collect audio and video streaming media or other data through the MediaStream API and transmit them using the RTP / SRTP protocol.
[0084] The process of the mixed reality remote holographic annotation function in step S302 is as shown in the attached figure. Figure 6 As shown, the specific steps include:
[0085] S3021: On-site staff establish audio and video communication with remote experts, and on-site staff transmit real-time first-person video streams of the industrial site to remote experts;
[0086] S3022: The on-site staff sends a remote guidance request to the remote expert. After the remote expert agrees to the request, the video stream is automatically frozen and the annotation tool is called up.
[0087] S3023: The remote expert uses a mouse or touchscreen to click on a frozen still frame in a 2D device, transmits the relative click position within the video to the on-site worker, and creates a spatial anchor point at this position for storing annotation data.
[0088] S3024: Remote experts annotate frozen still frames;
[0089] S3025: The coordinates of the 2D annotation are converted to 3D space, and the converted annotation information is stored in the Azure spatial anchor. The on-site staff reads the annotation information stored in the spatial anchor, performs collision detection with the reconstructed 3D scene, and calculates the actual position of the annotation in the 3D space.
[0090] S3026: The converted annotations are positioned in the real space, forming a remote guidance that integrates virtual and real space;
[0091] S3027: Synchronize the annotated video stream from the terminal worn by the on-site staff to the remote expert. Both parties can now see the annotations in the real-time video.
[0092] The multimodal human-computer interaction model for the mixed reality collaborative environment in step S4 is shown in FIG. 7 , and includes the following steps:
[0093] S401: Determine the interaction mode; S402: Develop a multimodal interaction strategy.
[0094] Preferably, the interaction modalities in step S401 include gesture, voice, gaze, BCI, and keyboard and mouse; wherein, gesture, voice, gaze, and BCI are used by on-site staff and are implemented based on head-mounted augmented reality devices and brain-computer interface devices; PC is used by remote experts to interact with the mouse and keyboard to provide guidance to on-site staff;
[0095] Preferably, the multimodal interaction strategy in step S402 is: under normal conditions, on-site workers give priority to using gesture recognition to complete the required interaction; when the on-site workers' hands are occupied, a combination of gaze and voice is used as a backup interaction strategy, using gaze rays to select the operation object, and then using voice commands to confirm the operation, thereby reducing the probability of misoperation; when an emergency occurs at the industrial site and there is no time to perform operations through other interaction modes, the BCI system is used to analyze the instinctive EEG signals of the on-site workers and execute corresponding commands.
[0096] The multi-terminal conflict resolution model in step S5 is shown in the attached Figure 8 As shown, the specific process is as follows: the system processes and analyzes the operation data of multiple terminals to determine whether there is a conflict. If there is no conflict between multiple operation instructions, all instructions can be executed directly; if there is a conflict between multiple operation instructions, the operation with lower timeliness requirements and can be postponed is defined as a normal operation, and the operation that must be executed immediately is defined as an emergency operation. If the operations performed by each user are all normal operations, a negotiation-based conflict resolution method is adopted, and each user can fully discuss and make a new decision; if there is an emergency operation among the operations performed by each user, the number of emergency operations must be further determined; if there is only one emergency operation, it will be executed first; if there are multiple emergency operations, the conflict resolution is carried out according to the principle of priority of the subject with higher priority, that is, the operation priority of each user is different, and the strategy with higher priority will take precedence;
[0097] The principle of the above priority division is: on-site staff have higher priority than remote experts, and among all on-site staff, the more experienced they are, the higher their priority, and the same applies to remote experts.
[0098] In the present invention, the cloud-edge-end collaborative computing task scheduling model constructed in step S1 is the basis for the efficient completion of subsequent steps; the mixed reality co-location collaborative shared space and the mixed reality remote collaborative shared space constructed in steps S2 and S3 form the main body of the mixed reality collaborative environment; the multimodal interaction model constructed in step S4 provides specific operation instruction issuance methods and strategies for on-site staff and remote experts in the mixed reality collaborative environment; the multi-terminal conflict resolution model constructed in step S5 can analyze the operation instructions issued by multiple users and sort the priorities when there is a conflict.
[0099] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications, characterized by: The following steps are included: S1: Build a cloud-edge-device collaborative computing task scheduling model; S2: Building a mixed reality co-location collaborative shared space; The steps include: S201: Building a cross-platform multi-terminal network communication framework based on PUN2; S2011: Connect multiple terminals to the main server of the mixed reality computing cloud platform. The main server will be responsible for all terminal-to-server transmissions, and the load balancing function of the mixed reality computing cloud platform will be responsible for coordinating all available rooms; S2012: The master terminal creates a room in the master server, sets the room name and maximum number of connections for the room, and waits for connections from other terminals. S2013: Other terminals access and join the room by indexing the room name, so that all terminals are in the same room; S2014: Using the event system to try to synchronize content, all necessary data can be distributed among the participants in the room; when the other terminals receive this information, a direct connection between the terminals is established; every time a terminal connects to the master server, it synchronizes a delay-corrected timestamp; this can be used in the room to synchronize the time of events; S202: Spatial state synchronization based on Azure Spatial Anchors; S2021: Create a spatial anchor resource in the Azure portal; S2022: Deploy a shared anchor service, that is, deploy an ASP.NET Core web application in Azure that can be used for shared anchors; S2023: Configure and deploy a Unity project. First, import the ASA SDK and OpenXR plug-in into the project. Then connect the Unity 3D scene to Azure resources. Finally, add and configure the SpatialAnchorManager interface to call the ASA service. S2024: Synchronize spatial states by creating anchor points, retrieving anchor points, and sharing anchor points; S3: Constructing a mixed reality remote collaboration environment. The mixed reality co-located collaboration shared space and the mixed reality remote collaboration shared space constructed in steps S2 and S3 form the main body of the mixed reality collaboration environment. Step S3 includes the following steps: S301: Implementation of remote audio and video communication functions for mixed reality based on WebRTC; S302: Implementation of mixed reality remote holographic annotation function; S4: Build a multimodal interaction model that provides specific methods and strategies for issuing operational instructions to on-site staff and remote experts in a mixed reality collaborative environment. S5: Construct a multi-terminal conflict resolution model. The multi-terminal conflict resolution model can analyze the operation instructions issued by multiple users and sort the priorities when there is a conflict.
2. The method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications according to claim 1 is characterized by: The step S1 includes the following steps: S101: Build a mixed reality computing cloud platform, which is used to provide edge node management and core business logic processing services for edge applications; S102: Deploy edge nodes. Edge nodes are deployed at industrial sites, including edge clouds, edge gateways, and edge controllers. S103: Building a mixed reality collaboration terminal system, where the mixed reality collaboration terminal system is a collection of all terminals used by users to complete mixed reality collaboration; S104: Task offloading and resource allocation. Task offloading and resource allocation refers to deploying audio and video encoding and decoding, core image rendering, hand tracking and motion tracking tasks with high real-time requirements on the edge side, while deploying non-core image rendering and natural voice interaction tasks that are not sensitive to delay on the cloud.
3. The method for constructing a cloud-edge-device collaborative mixed reality environment for industrial applications according to claim 1 is characterized by: The step S301 specifically includes the following steps: S3011: Set up signaling server; S3012: The communicating parties establish a connection with the signaling server via Websocket. S3013: The communicating parties send a signal to the signaling server through the RTCPeerConnection API and establish a peer-to-peer connection by exchanging Session Description Protocol information; S3014: The communicating parties collect audio and video streaming media or other data through the MediaStream API and transmit them using the RTP / SRTP protocol.
4. The method for constructing a cloud-edge-device collaborative mixed reality environment for industrial applications according to claim 1 is characterized by: The step S302 specifically includes the following steps: S3021: On-site staff establish audio and video communication with remote experts, and on-site staff transmit real-time first-person video streams of the industrial site to remote experts; S3022: The on-site staff sends a remote guidance request to the remote expert. After the remote expert agrees to the request, the video stream is automatically frozen and the annotation tool is called up. S3023: The remote expert uses a mouse or touchscreen to click on a frozen still frame in a 2D device, transmits the relative click position within the video to the on-site worker, and creates a spatial anchor point at this position for storing annotation data. S3024: Remote experts annotate frozen still frames; S3025: The coordinates of the 2D annotation are converted to 3D space, and the converted annotation information is stored in the Azure spatial anchor. The on-site staff reads the annotation information stored in the spatial anchor, performs collision detection with the reconstructed 3D scene, and calculates the actual position of the annotation in the 3D space. S3026: The converted annotations are positioned in the real space, forming a remote guidance that integrates virtual and real space; S3027: Synchronize the annotated video stream from the terminal worn by the on-site staff to the remote expert. Both parties can now see the annotations in the real-time video.
5. The method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications according to claim 1, characterized in that: The step S4 comprises the following steps: S401: Determine interaction modalities; the interaction modalities include gesture, voice, gaze, BCI, and keyboard and mouse. Among them, gesture, voice, gaze, and BCI interaction modalities are used by on-site staff and are implemented based on head-mounted augmented reality devices and brain-computer interface devices; PCs are used by remote experts to interact with the mouse and keyboard to provide guidance to on-site staff; S402: Develop a multimodal interaction strategy. The multimodal interaction strategy is as follows: under normal conditions, on-site workers give priority to using gesture recognition to complete the required interaction; when the on-site workers' hands are occupied, a combination of gaze and voice is used as a backup interaction strategy, using gaze rays to select the operation object, and then using voice commands to confirm the operation, thereby reducing the probability of misoperation; when an emergency occurs at the industrial site and there is no time to perform operations through other interaction modes, the BCI system is used to analyze the instinctive EEG signals of the on-site workers and execute corresponding commands.
6. The method for constructing a cloud-edge-device collaborative mixed reality collaborative environment for industrial applications according to claim 1, characterized in that: The multi-terminal conflict resolution model in step S5 is as follows: the system processes and analyzes the operation data of multiple terminals to determine whether there is a conflict. If there is no conflict between the multiple operation instructions, all instructions can be directly executed. If there is a conflict between the multiple operation instructions, the operation with lower timeliness requirements and can be postponed is defined as a normal operation, and the operation that must be executed immediately is defined as an emergency operation. If the operations performed by each user are all normal operations, a negotiation-based conflict resolution method is adopted, and each user can fully discuss and make a new decision. If there are emergency operations among the operations performed by each user, the number of emergency operations is further determined; if there is only one emergency operation, it is executed first; if there are multiple emergency operations, conflicts are resolved according to the principle of priority of the higher priority subject, that is, the operation priority of each user is different, and the strategy with the higher priority takes precedence.
Citation Information
Patent Citations
An Interactive Remote Expert Collaborative Maintenance System and Method Based on Augmented Reality Technology
CN106339094B
Data visualization remote assistance method and system based on mixed reality
CN111553974A
Method for carrying out remote intelligent diagnosis on power station equipment by utilizing augmented reality AR technology
CN111679740A
Remote synchronous cooperation system based on mixed reality
CN112667179A