A mixed reality multi-person space alignment and task collaboration method based on digital twinning
Patent Information
- Application Number
- CN202610927193.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-15
Smart Images

Figure CN122760894A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mixed reality and digital twin technology, specifically a mixed reality multi-person spatial alignment and task collaboration method based on digital twins. Background Technology
[0002] In equipment maintenance, assembly guidance, spatial inspection, and multi-person collaborative operations, mixed reality technology typically requires first aligning the terminal with the real work space, and then overlaying virtual content such as task prompts, operation annotations, and path guidance onto the vicinity of the corresponding physical objects. In existing solutions, mixed reality terminals mostly achieve positioning through environmental images, depth data, inertial measurement data, visual features, or spatial anchor points. Some solutions further introduce building information models, 3D point cloud models, or digital twin 3D models to register on-site physical objects with model objects, thereby improving the accuracy of the virtual-real overlay position.
[0003] In multi-person mixed reality collaborative scenarios, there is a continuous data relationship between spatial alignment, task status synchronization, and virtual content display. Spatial alignment determines whether virtual content can be superimposed on the correct entity position, task status determines what kind of guidance content is displayed on different terminals, and changes in the occlusion, pose, disassembly, or connection relationships of target objects on site will affect whether the original spatial registration benchmark remains reliable. Existing technologies mostly divide the above processes into two independent processes: one process focuses on terminal positioning, spatial anchor point sharing, or 3D model registration; the other process focuses on task allocation, execution status updates, and collaborative prompt synchronization. There is a lack of closed-loop processing between the two to maintain the spatial alignment benchmark based on changes in the task object's state.
[0004] In actual operation, objects such as equipment shells, cabinet doors, pipeline accessories, brackets, and temporary tooling may change their state due to maintenance, disassembly, or obstruction; personnel, tools, and materials may also enter the field of view of different terminals. If these objects continue to be used as spatial registration references, or if their state changes do not affect the credibility of shared spatial anchor points, different mixed reality terminals may continue to use invalid or low-credibility registration relationships, resulting in offset of task labels at the same target object, inconsistency between the virtual guidance content and the front and back occlusion relationship of the physical object, or inconsistency in the path guidance and operation prompts received by different personnel.
[0005] Therefore, the main shortcomings of existing digital twin-assisted mixed reality collaborative technologies are that digital twin models are often only used as spatial models, object indexes, or task data carriers, and there is a lack of calculable correlation between the state of task objects and the credibility of spatial anchor points; the spatial alignment results of multiple terminals are also difficult to dynamically correct as the occlusion state, pose state, disassembly and assembly state of the target object and the topological relationship of the object change. This problem will directly affect the consistency of virtual and real superposition in multi-person mixed reality collaborative operations, the stability of task guidance, and the verifiability of spatial registration results. Summary of the Invention
[0006] The purpose of this invention is to provide a mixed reality multi-person spatial alignment and task collaboration method based on digital twins, which is used to solve the problem of the disconnect between spatial registration benchmark and task status update in multi-person mixed reality collaborative operations. In multi-person operation scenarios, mixed reality terminals need to stably overlay task prompts, operation annotations and path guidance near physical objects. However, during the operation, the target object may be occluded, disassembled, or have its pose changed, or its connection relationship with surrounding objects may change. If spatial alignment still uses the original spatial anchor point, it is easy for the guidance content in different mixed reality terminals to have positional offset, incorrect occlusion relationship or inconsistent collaborative guidance.
[0007] To achieve the above objectives, the present invention adopts the following technical solution.
[0008] The method of this invention first performs objectification processing on the digital twin 3D model of the target work space to generate a twin object library and a twin coordinate system. The digital twin 3D model is generated from one or more of the following: building information model, 3D point cloud model, equipment ledger data, and computer-aided design space drawings. During objectification processing, the physical components, equipment, pipelines, cabinets, platform boundaries, and passage boundaries in the target work space are mapped to twin objects respectively. Each twin object is configured with a unique object identifier, spatial geometric description, object semantic category, object pose, and object topological relationship. The spatial geometric description is used to represent the 3D contour, corners, edges, planes, or feature point clouds of the twin object; the object topological relationship is used to represent the relative positional relationship, connection relationship, containment relationship, or adjacency relationship between twin objects.
[0009] After generating the twin object library, based on the pose change records of the twin objects and the movement markers, disassembly markers, and occlusion markers corresponding to the twin objects in the current task configuration data, twin reference objects are selected from the twin object library. The twin reference object is a twin object whose pose change is not greater than the reference threshold and does not have movement markers, disassembly markers, and occlusion markers. The reference threshold is determined based on the model verification data before the task starts. The model verification data includes one or more of the following: the pose difference of the same twin object in adjacent model versions, and the registration residual between the field scan point cloud and the digital twin 3D model. Through the above screening, the reference objects participating in spatial registration are twin objects that are stable in position within the current task cycle, have not been disassembled, and have not been occluded.
[0010] The twin reference feature set is extracted from the twin reference object. The twin reference feature set includes the geometric reference features, semantic category and object topological relationship of the twin reference object. The geometric reference features are extracted from the three-dimensional contour, corner points, edges, planes and feature point cloud of the twin reference object and stored in association with the object unique identifier of the corresponding twin reference object. The semantic category is used to limit the type of candidate matching object. The object topological relationship is used to verify whether the spatial relationship between the local observation object and the twin reference object is consistent in subsequent matching.
[0011] Multiple mixed reality terminals acquire environmental images, depth data, and terminal motion data within the target workspace. Each mixed reality terminal extracts the local observation object corresponding to the entity object based on the environmental image and depth data, and generates local geometric features, local semantic category, and observation weights for the local observation object. The local geometric features include edge, corner, plane, contour, or point cloud features obtained from the environmental image and depth data. The local semantic category is used to represent the entity object type corresponding to the local observation object. The observation weights are used in the subsequent calculation of matching confidence and terminal observation edge weights.
[0012] The observation weight is determined based on the displacement deviation between the spatial displacement of the local observed object in consecutive frames and the terminal pose change corresponding to the terminal motion data. Specifically, the spatial displacement of the local observed object in consecutive frames is converted to the same mixed reality terminal local coordinate system and then compared with the terminal pose change corresponding to the terminal motion data. When the displacement deviation between the two is greater than the dynamic judgment threshold, the observation weight of the local observed object is updated according to the preset attenuation coefficient. The preset attenuation coefficient is greater than 0 and less than 1. The dynamic judgment threshold is determined based on the displacement deviation statistics obtained by the mixed reality terminal from continuous sampling of the twin reference object before the task starts. This processing is used to reduce the impact of personnel, mobile tools, temporary obstructions, or objects being disassembled or assembled on spatial registration.
[0013] After obtaining the local observation objects and twin reference feature sets, the local observation objects and twin reference objects in each mixed reality terminal are matched. During the matching process, candidate twin reference objects are first determined based on the semantic categories of the local and twin reference objects. Then, a matching confidence score is generated based on the geometric deviation between the local geometric features and the geometric reference features, the topological deviation between the relative positional relationship between the local observation objects and the object topological relationship between the twin reference objects, and the observation weight. The geometric deviation is determined by one or more of the following: edge direction deviation, plane normal deviation, corner relative position deviation, and contour scale deviation. The topological deviation is determined by one or more of the following: relative orientation deviation, difference in the number of adjacent objects, difference in the identifier of connected objects, and difference in the inclusion relationship.
[0014] When the matching confidence is not less than the matching threshold, a matching relationship is generated between the local observation object and the twin reference object. The rotation and translation from the local coordinate system to the twin coordinate system of the mixed reality terminal are calculated based on the matching relationship, which serves as the initial spatial transformation relationship of the mixed reality terminal. The matching threshold is determined based on the lower limit of the matching confidence obtained by the mixed reality terminal performing initial sampling on no less than three labeled twin reference objects before the task starts.
[0015] After obtaining initial spatial transformation relationships from multiple mixed reality terminals, a shared spatial anchor point constraint graph is constructed based on the initial spatial transformation relationships, matching relationships, matching confidence, observation weights, and object topological relationships. The shared spatial anchor point constraint graph includes terminal nodes, reference object nodes, terminal observation edges, and reference object topological edges. Terminal nodes correspond to mixed reality terminals, reference object nodes correspond to twin reference objects, terminal observation edges are formed by the mixed reality terminal and its matched twin reference object, and reference object topological edges are formed by the relative positional relationships, connection relationships, inclusion relationships, or adjacency relationships between twin reference objects. Each terminal observation edge is assigned an edge weight, which is determined based on at least two of the following: matching confidence, observation weight, reprojection bias, observation distance, observation angle, and effective depth point ratio. The reprojection bias is the deviation between the twin reference object and the corresponding local observation object position after projecting it onto the mixed reality terminal's field of view according to the current spatial transformation relationship; the effective depth point ratio is the ratio between the number of effective depth points in the local observation object area and the number of sampling points in that area.
[0016] After the shared spatial anchor point constraint graph is generated, the initial spatial transformation relationship of each mixed reality terminal is corrected using this graph. During the correction, the reprojection deviation of the same twin reference object in different mixed reality terminals and the topological deviation between twin reference objects are used as the correction objects. The edge weight of the terminal observation edge is used as the constraint strength of the terminal observation edge. For terminal observation edges with edge weights lower than the retention threshold, the weights are reduced or deleted. For the retained terminal observation edges and reference object topological edges, the unified spatial transformation relationship from the corresponding mixed reality terminal to the twin coordinate system is recalculated. The retention threshold is determined based on the edge weight distribution of the terminal observation edges obtained by continuous sampling after the first consistency correction is completed.
[0017] After spatial alignment is completed, the task to be executed is divided into multiple task units, and the task units are bound with task identifiers, execution terminal identifiers, unique identifiers of target objects, target object poses, target object states, and task dependencies to generate a task object state graph. The task object state graph includes task nodes, target object nodes, and task dependency edges. Task nodes correspond to task units, target object nodes correspond to target objects in the twin object library, and task dependency edges correspond to task dependencies. The target object state includes occlusion state, pose state, disassembly / assembly state, and task execution state. The task dependency relationship includes one or more of the following: operation sequence relationship, target object connection relationship, spatial occlusion relationship, and safety distance relationship.
[0018] Based on a unified spatial transformation relationship and a task object state diagram, the mixed reality terminal overlays mixed reality guidance content into its display field of view. The mixed reality guidance content includes one or more of the following: task prompts, operation annotations, path guidance, and execution terminal location identifiers. When overlaying mixed reality guidance content, the terminal determines the front-to-back occlusion relationship between the entity object and the mixed reality guidance content in the mixed reality terminal's field of view based on the geometric information of the target object in the digital twin 3D model and the unified spatial transformation relationship. The terminal then occludes and clips the mixed reality guidance content according to the front-to-back occlusion relationship, so that the mixed reality guidance content corresponds to the entity spatial position of the target object.
[0019] When any mixed reality terminal updates the state of a target object, the updated state of the target object is written into the task object state diagram, and the affected task units are identified. When identifying the affected task units, the corresponding task units are first identified in the task object state diagram based on the unique identifier of the target object whose state has changed. Then, the target objects that are associated with the target object whose state has changed are identified based on the task dependencies, and the task units that are bound to the associated target objects are identified as the affected task units.
[0020] When the target object undergoing a state change is a twin reference object, or when the target object has an object topological relationship with at least one twin reference object, the terminal observation edges connected to the reference object nodes corresponding to the at least one twin reference object are determined as relevant terminal observation edges. The edge weights of the relevant terminal observation edges are adjusted according to the occlusion state, pose state, or topological state in the target object's state. Specifically, the occlusion state is used to reduce the edge weights of the terminal observation edges corresponding to the occluded twin reference object; the pose state is used to recalculate the reprojection bias of the terminal observation edges corresponding to the twin reference object whose pose has changed; and the topological state is used to update the reference object topological edges related to the twin reference object whose topological relationship has changed. If the adjusted edge weights are lower than the retention threshold, twin reference object replacement or local reregistration is triggered.
[0021] When replacing a twin reference object, a twin reference object with an edge weight not lower than the retention threshold and meeting the twin reference object selection conditions is selected from the current field of view to replace the twin reference object with an edge weight lower than the retention threshold. During local reregistration, the unified spatial transformation relationship of the corresponding mixed reality terminal is recalculated based on the replaced twin reference object, and the corresponding terminal observation edge in the shared spatial anchor point constraint graph is updated.
[0022] Compared with existing technologies, this invention objectifies the digital twin 3D model into a twin object library and selects twin reference objects based on pose change records, movement markers, disassembly markers, and occlusion markers. This allows the spatial registration reference to be determined from the twin objects available in the current task cycle, reducing the probability of dynamic objects, temporary occlusion objects, and disassembly objects participating in spatial registration.
[0023] This invention achieves joint matching of semantic categories, geometric features, and object topological relationships between local observed objects and twin reference objects, and constructs a shared spatial anchor point constraint graph by combining observation weights. This enables the spatial transformation relationships of multiple mixed reality terminals to be uniformly corrected based on the twin coordinate system, reducing the impact of mismatching of similar-shaped objects and single-terminal positioning drift on multi-person spatial alignment.
[0024] This invention links the task object state diagram with the shared spatial anchor point constraint diagram. When the target object is occluded, changes its pose, is disassembled or assembled, or changes its topological relationship, the edge weights of the observation edges of the relevant terminals are updated with the target object state. When the edge weights are lower than the retention threshold, twin reference object replacement or local re-registration is triggered, so that the task state update can participate in the maintenance of the spatial registration reference, thereby reducing the display offset and guidance errors caused by the disconnect between task collaboration and spatial alignment.
[0025] The beneficial effects of this invention are as follows: 1. This invention objectifies the digital twin 3D model of the target workspace into a twin object library, and assigns each twin object a unique identifier, spatial geometric description, semantic category, pose, and topological relationship. This allows the spatial alignment of the mixed reality terminal to no longer rely solely on temporarily acquired visual features or ordinary spatial anchor points, but instead uses twin objects with clear object identities and spatial relationships in the digital twin model as the registration basis. Furthermore, this invention selects twin reference objects based on pose change records, movement markers, disassembly markers, and occlusion markers, and determines a reference threshold based on model validation data. This excludes objects with unstable positions, disassembly / reassembly, or occlusion within the current task cycle from participating in registration. Through these processes, the reliability of the spatial registration reference is improved, and erroneous registration caused by dynamic objects and temporarily occluded objects is reduced.
[0026] 2. This invention extracts local observation objects from multiple mixed reality terminals and performs matching based on local semantic categories, local geometric features, twin reference feature sets, object topological relationships, and observation weights. This ensures that the correspondence between local observation objects and twin reference objects depends not only on geometric similarities such as contours, corners, and edges, but also on the joint constraints of semantic categories and object topological relationships. Furthermore, this invention constructs a shared spatial anchor point constraint graph including terminal nodes, reference object nodes, terminal observation edges, and reference object topological edges. It also uses matching confidence, observation weights, reprojection bias, observation distance, observation angle, and the proportion of effective depth points to determine the edge weights of terminal observation edges. Through the above processing, it achieves the effect of uniformly correcting spatial transformation relationships across multiple terminals, reducing mismatches of similar objects, and mitigating the impact of single-terminal positioning drift.
[0027] 3. This invention generates a task object state diagram by binding the task unit with the unique identifier of the target object, the target object pose, the target object state, and the task dependency relationship. The task object state diagram is linked with the shared space anchor point constraint diagram, so that the change of the target object state during the task execution process can have a reverse effect on the spatial registration reference. When the target object is occluded, its pose changes, its disassembly / assembly state changes, or its topological relationship changes, this invention determines the relevant terminal observation edge and adjusts the edge weight according to the target object state. When the edge weight is lower than the retention threshold, it triggers the replacement of the twin reference object or local re-registration. Through the above processing, the effect of updating the task state and maintaining spatial alignment are carried out simultaneously, reducing the display offset, occlusion error, and inconsistent collaborative guidance caused by failed anchor points or low-reliability registration relationships in multi-person mixed reality collaborative operations. Attached Figure Description
[0028] Figure 1 This is a diagram illustrating the closed-loop mechanism for multi-person mixed reality spatial alignment and task collaboration driven by digital twins, as described in this invention. Figure 2This is the overall flowchart of the mixed reality multi-person spatial alignment and task collaboration method based on digital twins of the present invention; Figure 3 This is a flowchart of the twin reference object selection process of the present invention; Figure 4 This is a flowchart of the closed-loop process for state linkage and reregistration in this invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] like Figures 1 to 4 As shown, this embodiment provides a mixed reality multi-person spatial alignment and task collaboration method based on digital twins. It is applicable to scenarios where multiple people perform equipment maintenance, production line assembly, space inspection, operation training, or emergency response in the same target work space through a mixed reality terminal. The mixed reality terminal includes an environmental image acquisition unit, a depth acquisition unit, a terminal motion data acquisition unit, a processing unit, and a display unit. The environmental image acquisition unit is used to acquire environmental images within the target work space. The depth acquisition unit is used to acquire depth data of the corresponding area of the environmental image. The terminal motion data acquisition unit is used to acquire pose change data of the mixed reality terminal at continuous time. The display unit is used to display mixed reality guidance content superimposed on the physical object.
[0031] In this embodiment, the target workspace is a space with fixed physical objects and task objects. The fixed physical objects include physical components, equipment, pipelines, cabinets, platform boundaries, and passage boundaries. The task objects are physical objects to be inspected, assembled, patrolled, or operated. The terminal motion data of the mixed reality terminal is obtained by fusing one or more of inertial measurement data, visual odometry data, and depth odometry data. The terminal motion data includes at least the displacement change and attitude change of the mixed reality terminal between adjacent acquisition times.
[0032] This embodiment first obtains a digital twin 3D model of the target work space. The digital twin 3D model is generated from one or more of the following: building information model, 3D point cloud model, equipment ledger data, and computer-aided design spatial drawings. When generating the digital twin 3D model using multi-source data, the fixed building axis, fixed spatial reference point, or on-site calibration reference within the target work space is first used as the coordinate reference to register the multi-source data, resulting in a unified twin coordinate system. The twin coordinate system serves as the unified coordinate reference for subsequent spatial transformation, task object positioning, and mixed reality guidance content overlay.
[0033] When objectifying the digital twin 3D model, the physical components, equipment, pipelines, cabinets, platform boundaries, and passage boundaries in the target work space are divided into twin objects, and a twin object library is generated. Each twin object is configured with a unique object identifier, spatial geometric description, object semantic category, object pose, and object topological relationship. The unique object identifier is used to uniquely point to the corresponding twin object during spatial registration, task binding, and state update. The spatial geometric description includes one or more of the twin object's 3D contour, corners, edges, planes, and feature point clouds. The object semantic category is used to indicate the category to which the twin object belongs. The object pose includes the position and orientation of the twin object in the twin coordinate system. The object topological relationship includes one or more of the relative positional relationship, connection relationship, containment relationship, and adjacency relationship between twin objects.
[0034] After generating the twin object library, a twin reference object is selected from the twin object library. The twin reference object is used as the reference object for spatial registration of the mixed reality terminal. When selecting a twin reference object, the pose change record of the twin object is read, and the movement marker, disassembly marker, and occlusion marker corresponding to the twin object in the current task configuration data are read. The pose change record comes from the comparison of adjacent model versions, on-site scan verification data, or terminal scan data before the task starts. The current task configuration data comes from the task plan, work ticket, equipment ledger, or on-site confirmation data, and is used to mark twin objects that need to be moved, disassembled, installed, turned on, turned off, or have occlusion risks within the current task cycle.
[0035] Specifically, twin objects whose pose change is no greater than the reference threshold and which do not have movement markers, disassembly markers, or occlusion markers are selected as twin reference objects. The reference threshold is determined based on model verification data before the task starts. The model verification data includes one or more of the following: pose difference between the same twin object in adjacent model versions, and registration residual between the field scanned point cloud and the digital twin 3D model. The registration residual is the distance deviation between corresponding points, corresponding edges, or corresponding planes of the corresponding twin objects in the field scanned point cloud and the digital twin 3D model after spatial registration. If the pose change of a twin object is greater than the reference threshold, or if it has any of the movement markers, disassembly markers, or occlusion markers, then the twin object will not be used as a twin reference object.
[0036] After selecting a twin reference object, a twin reference feature set is extracted from the twin reference object. The twin reference feature set includes the geometric reference features, semantic category, and object topological relationship of the twin reference object. The geometric reference features are extracted from the three-dimensional contour, corner points, edges, planes, and feature point cloud of the twin reference object, and are associated with and stored with the object unique identifier of the corresponding twin reference object. The semantic category is used to limit the category range of subsequent candidate matching objects. The object topological relationship is used to verify whether the spatial relationship between local observed objects is consistent with the spatial relationship between twin reference objects. For example, two cabinets with similar shapes are close in three-dimensional contours, but their relative positional relationship with pipelines, walls, or passage boundaries is different. The object topological relationship can distinguish between the two.
[0037] After multiple mixed reality terminals enter the target work space, they respectively collect environmental images, depth data, and terminal motion data within their respective fields of view. Each mixed reality terminal extracts the local observation object corresponding to the entity object based on the environmental image and depth data. The local observation object is the entity object region or entity object feature set detected by the mixed reality terminal in the current field of view and participating in spatial matching. For each local observation object, local geometric features, local semantic category, and observation weight are generated. The local geometric features include edge, corner, plane, contour, or point cloud features obtained from the environmental image and depth data. The local semantic category is obtained by the object recognition model, the geometric rule-based recognition method, or the on-site object identification recognition method. The on-site object identification includes equipment nameplate, object code identifier, or location identifier.
[0038] Observation weights are used to represent the weight of local observation objects in spatial registration calculations.
[0039] To facilitate the updating of observation weights by those skilled in the art, this embodiment can calculate the displacement deviation of the local observation object in the following manner; Let the position of the a-th local observation object observed by the i-th mixed reality terminal in the k-th frame be denoted as . The rotational change obtained from the terminal motion data between frame k and frame (k+1) is: The amount of translation change is If the local observed object is a stationary object, then the predicted observation position in the (k+1)th frame is: , The corresponding displacement deviation is: Where i represents the mixed reality terminal number, a represents the local observation object number, and k represents the frame number. This represents the three-dimensional position of the locally observed object in the local coordinate system of the i-th mixed reality terminal. This represents the amount of rotation change of the i-th mixed reality terminal between adjacent frames. This represents the amount of translation change of the i-th mixed reality terminal between adjacent frames. This represents the observation position in the (k+1)th frame predicted based on terminal motion data. This indicates the displacement deviation of the locally observed object. This represents the Euclidean distance.
[0040] When displacement deviation When the displacement deviation is not greater than the dynamic judgment threshold τd, the observation weight of the local observation object is maintained; when the displacement deviation is less than or equal to the dynamic judgment threshold τd, the observation weight of the local observation object is maintained. When the value exceeds the dynamic judgment threshold τd, the observation weight of the local observation object is reduced by a preset attenuation coefficient. ,in, This represents the observation weight of the a-th local observation object in the k-th frame. Let represent the updated observation weights, τd represent the dynamic judgment threshold, and λd represent the preset attenuation coefficient. The dynamic determination threshold τd is determined based on the displacement deviation statistics obtained by the mixed reality terminal through continuous sampling of the twin reference object before the task starts, and does not use a fixed empirical value.
[0041] The observation weight is determined based on the displacement deviation between the spatial displacement of the local observed object in consecutive frames and the terminal pose change corresponding to the terminal motion data. Specifically, the spatial displacement of the local observed object in consecutive frames is converted to the local coordinate system of the same mixed reality terminal to obtain the observed displacement of the object; the terminal pose change of the mixed reality terminal in the same time period is calculated based on the terminal motion data; the displacement deviation between the observed displacement of the object and the terminal pose change is calculated, and when the displacement deviation is greater than the dynamic judgment threshold, the observation weight of the local observed object is updated according to the preset attenuation coefficient, which is greater than 0 and less than 1.
[0042] The dynamic judgment threshold is determined based on the displacement deviation statistics obtained by the mixed reality terminal through continuous sampling of the twin reference object before the task starts. Specifically, before the task starts, the mixed reality terminal continuously samples the selected twin reference object to obtain the displacement deviation distribution of the stable object in multiple frames. The dynamic judgment threshold is determined based on the displacement deviation distribution. This threshold is derived from the actual observation fluctuations of the current mixed reality terminal and the current target work space. If the displacement deviation of the subsequent local observation object is greater than the dynamic judgment threshold, the change of the local observation object cannot be explained by the motion of the mixed reality terminal itself, indicating that it may be a person, a mobile tool, a temporary obstruction, or an object being disassembled or assembled. Therefore, its observation weight is reduced.
[0043] After obtaining the local observation objects and twin reference feature sets, the local observation objects in each mixed reality terminal are matched with the twin reference objects. The matching process includes candidate screening and confidence calculation. During candidate screening, candidate twin reference objects are determined based on the semantic categories of the local semantic categories and the twin reference objects. During confidence calculation, the geometric deviation between the local geometric features and the geometric reference features in the twin reference feature set is calculated, the topological deviation between the relative positional relationship between the local observation objects and the object topological relationship between the twin reference objects is calculated, and the matching confidence is generated by combining the observation weights.
[0044] Geometric deviation is determined by one or more of edge direction deviation, plane normal deviation, corner relative position deviation, and contour scale deviation. Topological deviation is determined by one or more of relative orientation deviation, difference in the number of adjacent objects, difference in the identification of connected objects, and difference in inclusion relationship. Matching confidence can be obtained by weighting the normalized geometric deviation, topological deviation, semantic category consistency results, and observation weights.
[0045] In this embodiment, the matching confidence score can be generated as follows.
[0046] Suppose that the a-th local observation object in the first mixed reality terminal is matched with the b-th twin reference object, and the geometric consistency is . The topological consistency is The semantic category consistency result is The observation weight is Then the matching confidence level for: ,in, This represents the matching confidence between the a-th local observation and the b-th twin reference object; Indicates geometric consistency; Indicates topological consistency; This indicates the result of semantic category consistency. The value is 1 when the semantic categories are consistent and 0 when the semantic categories are inconsistent. Indicates the observation weight of the local observation object; , , , These are the weight coefficients corresponding to geometric consistency, topological consistency, semantic category consistency results, and observation weights, respectively. , , , All are not less than 0. .
[0047] Geometric consistency The topological consistency is obtained by normalizing the geometric deviation. It is obtained by topological deviation normalization.
[0048] The smaller the geometric deviation, the higher the geometric consistency; the smaller the topological deviation, the higher the topological consistency.
[0049] Weighting coefficient , , , The weights can be determined based on the initial sampling results of the labeled twin reference objects before the task starts, or by taking the same weights when no separate calibration is performed.
[0050] Specifically, geometric deviation and topological deviation are converted into geometric consistency and topological consistency, respectively, with the values of geometric consistency and topological consistency ranging from 0 to 1. When the semantic categories are consistent, the semantic consistency result is 1, and when the semantic categories are inconsistent, the semantic consistency result is 0. Then, the geometric consistency, topological consistency, semantic consistency results and observation weights are weighted to generate the matching confidence score. Each weighting coefficient can be determined based on the initial sampling results of the labeled twin reference objects before the task starts, or the same weight can be used.
[0051] When the matching confidence is not less than the matching threshold, a matching relationship is generated between the local observation object and the twin reference object. The matching threshold is determined based on the lower limit of the matching confidence obtained by the mixed reality terminal in initial sampling of no less than three labeled twin reference objects before the task starts. During initial sampling, the labeled twin reference objects are provided by fixed objects with stable positions and clear semantic categories in the target work space, such as wall boundaries, fixed cabinets, fixed equipment shells, or fixed pipeline supports. The matching threshold adopts the lower limit of the matching confidence obtained by sampling the current scene and the current terminal, which can avoid using fixed empirical thresholds that do not match the accuracy of the current terminal and the conditions of the current scene.
[0052] Based on the generated matching relationship, the rotation and translation amounts from the local coordinate system to the twin coordinate system of the mixed reality terminal are calculated as the initial spatial transformation relationship of the mixed reality terminal. The rotation and translation amounts are calculated by spatial registration between the spatial feature points, spatial feature lines or spatial feature surfaces of the local observation object and the corresponding geometric reference features of the twin reference object. The initial spatial transformation relationship is used to transform the observation results in the local coordinate system of the mixed reality terminal to the twin coordinate system.
[0053] After multiple mixed reality terminals obtain the initial spatial transformation relationship, a shared spatial anchor constraint graph is constructed. The shared spatial anchor constraint graph is generated based on the initial spatial transformation relationship, matching relationship, matching confidence, observation weight, and object topological relationship corresponding to the multiple mixed reality terminals.
[0054] The shared space anchor point constraint graph includes terminal nodes, reference object nodes, terminal observation edges, and reference object topological edges. Terminal nodes correspond to mixed reality terminals, reference object nodes correspond to twin reference objects, terminal observation edges are formed by mixed reality terminals and their matched twin reference objects, and reference object topological edges are formed by the relative positional relationships, connection relationships, inclusion relationships, or adjacency relationships between twin reference objects.
[0055] Each terminal observation edge is assigned an edge weight.
[0056] In this embodiment, the edge weight of the terminal observation edge can be generated as follows: Let the terminal observation edge be formed between the i-th mixed reality terminal and the b-th twin reference object, with the edge weight being... ,but: ,in, Represents the terminal observation edge weight between the i-th mixed reality terminal and the b-th twin reference object; This indicates the matching confidence between the local observed object and the twin reference object; Indicates the observation weight of the local observation object; This represents the reprojection bias after normalization; This represents the normalized observation distance; This represents the normalized observation angle; Indicates the proportion of effective depth points; , , , , , These are the weight coefficients for the corresponding items, and none of them are less than 0. .
[0057] The normalized reprojection bias, normalized observation distance, and normalized observation angle range from 0 to 1.
[0058] The greater the reprojection bias, the greater the observation distance, or the greater the observation angle, the lower the corresponding edge weight; the higher the matching confidence, the higher the observation weight, or the higher the proportion of effective depth points, the higher the corresponding edge weight.
[0059] The weight coefficients of the above items can be determined based on the initial sampling results before the task starts, or the same weight can be used when no separate calibration is performed.
[0060] The edge weights are determined based on at least two of the following: matching confidence, observation weights, reprojection bias, observation distance, observation angle, and the proportion of effective points at depth.
[0061] Reprojection bias is the deviation between the position of the twin reference object and the corresponding local observation object after the twin reference object is projected onto the mixed reality terminal's field of view according to the current spatial transformation relationship.
[0062] The observation distance is the spatial distance between the mixed reality terminal and the local observed object. The observation angle is the angle between the line of sight of the mixed reality terminal and the main plane or contour direction of the local observed object. The effective depth point ratio is the ratio between the number of effective depth points in the local observed object area and the number of sampling points in that area. The edge weight can be obtained by weighting the normalized matching confidence, observation weight, reprojection bias, observation distance, observation angle and effective depth point ratio. Among them, the larger the reprojection bias, observation distance and observation angle, the lower the corresponding edge weight. The larger the matching confidence, observation weight and effective depth point ratio, the higher the corresponding edge weight.
[0063] After obtaining the shared space anchor point constraint diagram, the initial spatial transformation relationship of each mixed reality terminal is corrected based on the shared space anchor point constraint diagram.
[0064] To enable those skilled in the art to perform the correction process for spatial transformation relationships, this embodiment can employ the following optimization objective to calculate the unified spatial transformation relationship: Where Rᵢ represents the rotation of the i-th mixed reality terminal to the twin coordinate system, tᵢ represents the translation of the i-th mixed reality terminal to the twin coordinate system; Eo represents the set of terminal observation edges in the shared space anchor constraint graph; This represents the set of topological edges of the base object in the shared space plotted point constraint graph; Represents the terminal observation edge weight between the i-th mixed reality terminal and the b-th twin reference object; This represents the geometric datum feature point or geometric datum feature center point of the b-th twin datum object in the twin coordinate system; Let i represent the projection function of the i-th mixed reality terminal; This represents the observation position of the local observation object corresponding to the b-th twin reference object in the field of view of the i-th mixed reality terminal; This represents the topological relationship between the b-th twin reference object and the c-th twin reference object in the digital twin 3D model; This represents the topological relationship between the b-th twin reference object and the c-th twin reference object, calculated based on the observation results from the mixed reality terminal; μ represents the weight coefficient of the topological constraint term, and μ is not less than 0.
[0065] The first term in the above optimization objective is used to constrain the reprojection deviation of the same twin reference object in different mixed reality terminals, and the second term is used to constrain the topological deviation between twin reference objects.
[0066] By solving this optimization objective, a unified spatial transformation relationship from each mixed reality terminal to the twin coordinate system can be obtained.
[0067] The solution can be obtained by iterative least squares method, graph optimization method, or other spatial registration methods that can solve for rotation and translation.
[0068] During calibration, the reprojection deviation of the same twin reference object in different mixed reality terminals, as well as the topological deviation between twin reference objects, are used as the calibration objects. The edge weight of the terminal observation edge is used as the constraint strength of the terminal observation edge. For terminal observation edges with edge weights lower than the retention threshold, the weights are reduced or deleted. Based on the retained terminal observation edges and the topological edges of the reference object, the unified spatial transformation relationship from each mixed reality terminal to the twin coordinate system is recalculated. The retention threshold is determined based on the edge weight distribution of the terminal observation edges obtained by continuous sampling after the first consistency calibration. Specifically, after the first consistency calibration, the retained terminal observation edges are continuously sampled to obtain the edge weight distribution, and the retention threshold is determined based on the edge weight distribution to distinguish between stable observation constraints and unstable observation constraints.
[0069] After spatial alignment is completed, the task to be executed is divided into multiple task units. Each task unit is bound to a task identifier, an execution terminal identifier, a unique identifier of the target object, the pose of the target object, the state of the target object, and the task dependency relationship, generating a task object state graph. The task object state graph includes task nodes, target object nodes, and task dependency edges. Task nodes correspond to task units, target object nodes correspond to target objects in the twin object library, and task dependency edges correspond to task dependency relationships. The target object state includes occlusion state, pose state, disassembly / assembly state, and task execution state. The task dependency relationship includes one or more of the following: operation sequence relationship, target object connection relationship, spatial occlusion relationship, and safety distance relationship.
[0070] The mixed reality terminal overlays mixed reality guidance content into its display field of view based on a unified spatial transformation relationship and a task object state diagram. The mixed reality guidance content includes one or more of the following: task prompts, operation annotations, path guidance, and execution terminal location identifiers. When overlaying mixed reality guidance content, the occlusion relationship between the entity object and the mixed reality guidance content is determined based on the geometric information of the target object in the digital twin 3D model and the unified spatial transformation relationship. The mixed reality guidance content is then occluded and clipped according to this occlusion relationship. During occlusion and clipping, mixed reality guidance content located behind the entity object does not cover the front surface of the entity object, while mixed reality guidance content located in front of the entity object is displayed at the corresponding entity space position according to the unified spatial transformation relationship.
[0071] When any mixed reality terminal updates the state of a target object, the updated state of the target object is written into the task object state graph, and the affected task units are identified. When identifying the affected task units, the corresponding task units are first identified in the task object state graph based on the unique identifier of the target object whose state has changed. Then, the target objects that are associated with the target object whose state has changed are identified based on the task dependency relationship, and the task units bound to the associated target objects are identified as the affected task units. The target object state update includes occlusion state update, pose state update, disassembly and assembly state update, and task execution state update.
[0072] When the target object undergoing a state change is a twin reference object, or when the target object has an object topology relationship with at least one twin reference object, the terminal observation edge connected to the reference object node corresponding to the at least one twin reference object is determined as the relevant terminal observation edge, and the edge weight of the relevant terminal observation edge is adjusted according to the occlusion state, pose state or topology state in the target object state.
[0073] When the target object is in an occluded state, update the edge weights of the relevant terminal observation edges according to the occlusion attenuation coefficient: ,in, This represents the edge weight observed by the terminal after the occlusion state is updated. Let λo represent the terminal observation edge weights before the occlusion state update, and let λo represent the occlusion attenuation coefficient. .
[0074] When the target object's state changes, the reprojection deviation is recalculated based on the changed twin reference object's pose, and the edge weights of the corresponding terminal observation edges are recalculated based on the updated reprojection deviation.
[0075] When the target object's state changes due to a change in object topology, the corresponding reference object's topology edge is updated, and the topology deviation is recalculated based on the updated reference object's topology edge. If the recalculated topology deviation increases, the topology edge with the reference object is reduced.
[0076] Related terminal observation edge weights Among them, the occlusion state is used to reduce the edge weight of the terminal observation edge corresponding to the occluded twin reference object; the pose state is used to recalculate the reprojection deviation of the terminal observation edge corresponding to the twin reference object whose pose has changed, and update the edge weight according to the recalculated reprojection deviation; the topology state is used to update the reference object topology edge related to the twin reference object whose object topology relationship has changed, and recalculate the topology deviation according to the updated reference object topology edge.
[0077] When the adjusted edge weight is lower than the retention threshold, twin reference object replacement or local reregistration is triggered. When twin reference object replacement occurs, a twin reference object with an edge weight not lower than the retention threshold and meeting the twin reference object selection conditions is selected from the current field of view to replace the twin reference object with an edge weight lower than the retention threshold. The twin reference object selection conditions are: the pose change is not greater than the reference threshold and it does not have movement markers, disassembly markers, or occlusion markers. When local reregistration occurs, the unified spatial transformation relationship of the corresponding mixed reality terminal is recalculated based on the replaced twin reference object, and the corresponding terminal observation edge in the shared spatial anchor point constraint graph is updated.
[0078] In this embodiment, spatial alignment and task collaboration form a continuous processing relationship. The shared spatial anchor point constraint graph is used to correct the unified spatial transformation relationship from multiple mixed reality terminals to the twin coordinate system. The task object state graph is used to record the binding relationship between the task unit and the target object state. When the target object state changes and affects the twin reference object or its object topology relationship, the change is reversed through the task object state graph and applied to the relevant terminal observation edges in the shared spatial anchor point constraint graph. This causes the edge weights of the terminal observation edges, the replacement of the twin reference object, and the local re-registration to be updated as the task is executed. Thus, this embodiment does not only use the digital twin 3D model as the mixed reality display background, nor does it only share spatial anchor points or only synchronize task states. Instead, it establishes a closed-loop processing relationship between the target object state, spatial anchor point constraints, and multi-person mixed reality guidance.
[0079] Compared with the prior art, this embodiment objectifies the digital twin 3D model into a twin object library, and selects twin reference objects based on pose change records, movement markers, disassembly markers and occlusion markers, so that the spatial registration reference is determined from the twin objects that can be used for registration within the current task cycle, reducing the participation of dynamic objects, temporary occlusion objects and disassembly objects in spatial registration.
[0080] This embodiment uses joint matching of semantic categories, geometric features, and object topological relationships between local observed objects and twin reference objects, and constructs a shared spatial anchor point constraint graph by combining observation weights. This enables the spatial transformation relationships of multiple mixed reality terminals to be uniformly corrected based on the twin coordinate system, reducing the impact of mismatching of similar-shaped objects and single-terminal positioning drift on the spatial alignment of multiple users.
[0081] This embodiment links the task object state diagram with the shared spatial anchor point constraint diagram, so that the occlusion state, pose state, disassembly and assembly state or changes in the object topology relationship of the target object participate in the update of the edge weight of the relevant terminal observation. When the edge weight is lower than the retention threshold, it triggers the replacement of the twin reference object or local re-registration, thereby reducing the display offset and guidance error caused by the disconnect between the task state update and the maintenance of the spatial registration reference.
[0082] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0083] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for mixed reality multiplayer space alignment and task coordination based on digital twinning, characterized in that, Includes the following steps: S1. The digital twin 3D model of the target work space is objectified to generate a twin object library and a twin coordinate system, and a twin reference object is selected from the twin object library to form a twin reference feature set; S2. Multiple mixed reality terminals respectively collect environmental images, depth data and terminal motion data, extract the local observation objects corresponding to the entity objects, and generate the local geometric features, local semantic categories and observation weights of the local observation objects; S3. Based on local semantic categories, local geometric features, twin reference feature sets, object topological relationships, and observation weights, match local observed objects with twin reference objects to obtain matching relationships, matching confidence, and initial spatial transformation relationships. S4. Construct a shared spatial anchor point constraint graph based on the initial spatial transformation relationship, matching relationship, matching confidence, observation weight, and object topological relationship, and set edge weights for the terminal observation edges in the shared spatial anchor point constraint graph; S5. Correct the initial spatial transformation relationship based on the shared spatial anchor point constraint diagram to obtain the unified spatial transformation relationship from each mixed reality terminal to the twin coordinate system; S6. Divide the task to be executed into multiple task units, and bind the task units with the task identifier, execution terminal identifier, target object unique identifier, target object pose, target object state and task dependency relationship to generate a task object state diagram. S7. Based on the unified spatial transformation relationship and the task object state diagram, superimpose mixed reality guidance content on each mixed reality terminal; when the target object state is updated, update the task object state diagram and determine the affected task units; when the target object whose state changes is a twin reference object, or has an object topology relationship with at least one twin reference object, determine the terminal observation edge connected to the reference object node corresponding to the at least one twin reference object as the relevant terminal observation edge, adjust the edge weight of the relevant terminal observation edge according to the updated target object state, and trigger twin reference object replacement or local re-registration when the edge weight is lower than the retention threshold.
2. The method of claim 1, wherein the method is based on digital twinning. In step S1, the digital twin 3D model is generated from at least one of building information model, 3D point cloud model, equipment ledger data, and CAD spatial drawings; The twin objects in the twin object library have unique object identifiers, spatial geometric descriptions, object semantic categories, object poses, and object topological relationships. The objectification process includes mapping the physical components, equipment, pipelines, cabinets, platform boundaries, and passage boundaries in the target work space into twin objects.
3. The method for multi-person spatial alignment and task collaboration in mixed reality based on digital twins according to claim 2, characterized in that: In step S1, when selecting a twin reference object, based on the pose change record of the twin object and the movement marker, disassembly marker and occlusion marker corresponding to the twin object in the current task configuration data, a twin object whose pose change is not greater than the reference threshold and does not have movement marker, disassembly marker and occlusion marker is selected. The baseline threshold is determined based on model verification data before the task starts. The model verification data includes at least one of the following: pose difference between adjacent model versions of the same twin object, and registration residual between the on-site scan point cloud and the digital twin 3D model.
4. The method for multi-person spatial alignment and task collaboration in mixed reality based on digital twins according to claim 3, characterized in that: In step S1, the twin reference feature set includes the geometric reference features, semantic category, and topological relationship of the twin reference object; The geometric datum features are extracted from at least one of the three-dimensional contours, corners, edges, planes, and feature point clouds of the twin datum object, and the topological relationships include at least one of relative positional relationships, connectivity relationships, containment relationships, and adjacency relationships.
5. A mixed reality multi-person spatial alignment and task collaboration method based on digital twins according to claim 4, characterized in that: In step S2, when generating observation weights, the spatial displacement of the local observed object in consecutive frames is converted to the local coordinate system of the same mixed reality terminal, and then compared with the terminal pose change corresponding to the terminal motion data. When the displacement deviation between the two is greater than the dynamic judgment threshold, the observation weight of the local observation object is updated according to the preset attenuation coefficient. The preset attenuation coefficient is greater than 0 and less than 1. The dynamic judgment threshold is determined based on the displacement deviation statistics obtained by the mixed reality terminal through continuous sampling of the twin reference object before the task starts.
6. A method for multi-person spatial alignment and task collaboration in mixed reality based on digital twins according to claim 5, characterized in that: In step S3, when performing matching, candidate twin reference objects are first determined based on the semantic categories of the local semantic category and the twin reference object. Then, based on the geometric deviation between the local geometric features and the geometric reference features in the twin reference feature set, the topological deviation between the relative positional relationship between the local observation objects and the topological relationship between the twin reference objects, and the observation weight, a matching confidence score is generated. When the matching confidence is not less than the matching threshold, a matching relationship is generated between the local observation object and the twin reference object. Based on the matching relationship, the rotation and translation amounts from the local coordinate system to the twin coordinate system of the mixed reality terminal are calculated as the initial spatial transformation relationship.
7. A mixed reality multi-person spatial alignment and task collaboration method based on digital twins according to claim 6, characterized in that: In step S4, the shared space anchor point constraint graph includes terminal nodes, reference object nodes, terminal observation edges, and reference object topological edges; The terminal node corresponds to a mixed reality terminal, the reference object node corresponds to a twin reference object, the terminal observation edge is formed by the mixed reality terminal and its matched twin reference object, the reference object topological edge is formed by the relative positional relationship, connection relationship, inclusion relationship or adjacency relationship between twin reference objects, and the edge weight is determined based on at least two of the following: matching confidence, observation weight, reprojection bias, observation distance, observation angle and effective point ratio of depth.
8. A mixed reality multi-person spatial alignment and task collaboration method based on digital twins according to claim 7, characterized in that: In step S5, when correcting the initial spatial transformation relationship, the reprojection deviation of the same twin reference object in different mixed reality terminals and the topological deviation between twin reference objects are taken as the correction objects, and the edge weight of the terminal observation edge is taken as the constraint strength of the terminal observation edge. Terminal observation edges with weights below the retention threshold are downweighted or deleted, and the unified spatial transformation relationship is recalculated based on the retained terminal observation edges and the topological edges of the reference object; the retention threshold is determined based on the distribution of edge weights of terminal observation edges obtained by continuous sampling after the first completion of consistency correction.
9. A mixed reality multi-person spatial alignment and task collaboration method based on digital twins according to claim 8, characterized in that: In step S6, the task object state graph includes task nodes, target object nodes, and task dependency edges; the task node corresponds to a task unit, the target object node corresponds to a target object in the twin object library, and the task dependency edge corresponds to a task dependency relationship. The target object state includes occlusion state, pose state, disassembly / assembly state, and task execution state; The task dependency relationship includes at least one of operation sequence relationship, target object connection relationship, spatial occlusion relationship and safety distance relationship. When determining the affected task unit, the corresponding task unit is first determined in the task object state diagram according to the unique identifier of the target object whose state has changed. Then, the target object that is associated with the target object whose state has changed is determined according to the task dependency relationship, and the task unit bound to the associated target object is determined as the affected task unit.
10. A method for multi-person spatial alignment and task collaboration in mixed reality based on digital twins according to claim 9, characterized in that: In step S7, the mixed reality guidance content includes at least one of task prompts, operation annotations, path guidance, and execution terminal location identifiers; When overlaying mixed reality guidance content, the front and back occlusion relationship between the entity object and the mixed reality guidance content is determined based on the geometric information of the target object in the digital twin 3D model and the unified spatial transformation relationship, and the mixed reality guidance content is occluded and clipped according to the front and back occlusion relationship. The twin reference object replacement includes selecting twin reference objects with edge weights not lower than the retention threshold and meeting the twin reference object selection conditions in claim 3 from the current field of view, and replacing twin reference objects with edge weights lower than the retention threshold; the local re-registration includes recalculating the unified spatial transformation relationship of the corresponding mixed reality terminal based on the replaced twin reference object, and updating the corresponding terminal observation edge in the shared spatial anchor point constraint graph.