Digital twin ar / mr information adaptive presentation and layout method

CN122816451APending Publication Date: 2026-09-25BEIJING CHANGSHENG LITONG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610921799.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]针对上述背景技术中现有方法逐帧独立求标签不重叠位置而随头动抖动、且可能遮挡所描述对象与安全关键区域、信息过密时一律堆叠的问题,本发明提供一种数字孪生AR/MR信息自适应呈现与布局方法

Benefits of technology

[0010]本发明的有益效果在于:其一,由于布局代价计入各孪生信息元素相对其上一帧布局位置的帧间位移代价,并以上一帧布局作为本帧求解的初值,标签位置在相邻帧间保持连贯、随头动平滑移动,从根本上消除了现有逐帧独立求解所导致的抖动与闪烁;其二,由于对布局施加禁止遮挡所锚定对象与安全关键区域的语义遮挡规避,标签不再挡住它所描述的对象本身,也不再挡住机械运动部件、通行通道等现实中的安全关键区域,提升了可读性与安全性;其三,由于在累计呈现面积超过密度预算时按照优先级把次要元素的呈现形态逐级折叠、并在密度回落时按照优先级逐级恢复,信息过密时有序取舍、宽松时自动复原,避免了一律堆叠造成的过载。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816451A_ABST
    Figure CN122816451A_ABST
Patent Text Reader

Abstract

The application discloses a digital twin AR / MR information adaptive presentation and layout method, and relates to the technical field of human-computer interaction and head-mounted display. In view of the problems that the existing method independently calculates the non-overlapping position of labels frame by frame, shakes with head movement, and may shield the described object and the safety key area, and all stacks when the information is too dense, the method projects each twin object and the safety key area to the field image plane, generates a candidate anchor position around the object projection, comprehensively considers the overlap between elements, the lead distance and the interframe displacement relative to the layout position of the previous frame, selects the current frame layout with a lower layout cost, and applies the semantic shielding avoidance of prohibiting shielding of the anchored object and the safety key area; when the cumulative presentation area exceeds the density budget, the secondary elements are folded according to the priority, and the current frame layout is used as the next frame initial value to close loop iteration. The method makes the label layout interframe coherent, does not shake, does not shield the described object and the safety key area, and has an ordered selection when the information is too dense.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of human-computer interaction and head-mounted display technology, and in particular to a method for adaptive presentation and layout of digital twin AR / MR information. Background Technology

[0002] Digital twins, combined with augmented reality and mixed reality, overlay the twin information of equipment onto the physical equipment in the form of data panels, tags, alarm markers, etc., allowing operators to view the real scene and data simultaneously by wearing head-mounted displays. This technology has already been used for the operation and maintenance of substations, production lines, and other applications. Early methods involved fixing the twin information to the corresponding objects.

[0003] To avoid multiple labels overlapping and obscuring each other, the current mainstream approach is to use view management or real-time overlay placement, which means calculating the non-overlapping placement positions of each label in each frame, such as selecting empty spaces around the object and avoiding other labels.

[0004] However, the operator's head is constantly moving; with a slight turn of the head, the two-dimensional projection positions of each object in the field of view change continuously. If a new set of non-overlapping positions is independently calculated for each frame, a label that was to the left of an object in the previous frame might suddenly jump to the right in the next, causing the entire screen of labels to shake and flicker with head movements, which is dizzying to look at. On the other hand, existing methods only ensure that labels do not overlap, but this may result in a data panel blocking the device it describes, or blocking critical areas that cannot be blocked in reality, such as a moving robotic arm or passageway; and as more devices are added and labels pile up, the actual scene becomes even more obscured.

[0005] Ultimately, existing methods only focus on placing labels in non-overlapping positions within each frame, neglecting to ensure label positions remain consistent across adjacent frames and don't jitter with head movements. They also fail to consider that labels shouldn't obscure the objects they describe or real-world safety-critical areas, and they don't prioritize labels based on importance when information is dense. Consequently, frame-by-frame re-decoding leads to jitter, layout issues may obscure objects and safety zones, and labels are simply stacked when information is plentiful.

[0006] Therefore, how to avoid the existing methods that independently calculate the non-overlapping positions of tags frame by frame, causing them to jitter with head movement, and potentially obscuring the described objects and safety-critical areas, and causing all information to be stacked when it is too dense, is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0007] To address the problems of existing methods in the background art, such as independently determining the non-overlapping positions of tags frame by frame, which causes shaking with head movement, potentially obscuring the described object and safety-critical areas, and stacking all information when it is too dense, this invention provides a digital twin AR / MR information adaptive presentation and layout method.

[0008] The method includes: Step S0, initially acquiring each twin information element in the AR / MR scene and its respective anchored twin object, as well as several preset safety-critical areas, and initializing the presentation form and layout position of each twin information element; Step S1, acquiring the head posture of the head-mounted display device, and projecting each twin object and each safety-critical area onto the image plane of the current field of view according to the head posture, obtaining the projection position of each twin object and the projection range of each safety-critical area; Step S2, for each twin information element, generating several candidate anchor positions around the projection position of its anchored twin object; Step S3, with the goal of optimizing the layout cost, selecting the layout position of each twin information element in the current frame from its candidate anchor positions, whereby the layout cost includes the overlap cost between each twin information element and the cost of each twin information element and its anchored twin pair. The cost of the lead distance between the twins and the inter-frame displacement cost of each twin information element relative to its previous frame layout position; Step S4, apply semantic occlusion avoidance and information density adaptation to the selected layout: when any twin information element occludes the projection position of its anchored twin object or occludes the projection range of any safety-critical area at its layout position, select a candidate anchor position for the twin information element or downgrade its presentation form; when the cumulative presentation area of ​​each twin information element in the current field of view exceeds the preset density budget, downgrade the presentation form of the twin information element to a folded form in order of priority from minor to major; Step S5, render and output the current AR / MR presentation frame according to the selected layout position and presentation form of each twin information element, and use the layout position of each twin information element in this frame as the initial value of the layout of the next frame, and iterate in a closed loop.

[0009] The present invention also provides a computer-readable storage medium, an AR / MR information adaptive presentation and layout device, and a computer program product, wherein the computer program stored thereon or contained therein implements the above-described method when executed by a processor.

[0010] The beneficial effects of this invention are as follows: First, since the layout cost includes the inter-frame displacement cost of each twin information element relative to its previous frame layout position, and uses the previous frame layout as the initial value for solving this frame, the label position remains continuous between adjacent frames and moves smoothly with the head movement, fundamentally eliminating the jitter and flicker caused by the existing independent frame-by-frame solution; Second, since the semantic occlusion avoidance of anchored objects and safety-critical areas is applied to the layout to prohibit occlusion, the label no longer blocks the object it describes, nor does it block real-world safety-critical areas such as mechanical moving parts and passageways, improving readability and safety; Third, since the presentation form of secondary elements is folded step by step according to priority when the cumulative presentation area exceeds the density budget, and restored step by step according to priority when the density drops, information is orderly discarded when it is too dense and automatically restored when it is sparse, avoiding overload caused by uniform stacking. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of the overall process of the AR / MR information adaptive presentation and layout method according to an embodiment of the present invention.

[0012] Figure 2 This is a schematic diagram of the module structure of an AR / MR information adaptive presentation and layout device according to an embodiment of the present invention;

[0013] Figure 3 This is a schematic diagram illustrating the generation of candidate anchor points around the projection bounding box of a twin object according to an embodiment of the present invention.

[0014] Figure 4 This is a schematic diagram illustrating the layout cost structure according to an embodiment of the present invention;

[0015] Figure 5 This is a schematic diagram of the incremental layout solving sub-process according to an embodiment of the present invention, which uses the layout of the previous frame as the initial value.

[0016] Figure 6 This is a schematic diagram illustrating the presentation format grading and adaptive degradation and recovery of information density according to an embodiment of the present invention.

[0017] Figure 7 This is a schematic diagram of semantic occlusion avoidance processing according to an embodiment of the present invention;

[0018] Figure 8 This is a schematic diagram of the end-to-end execution process of the method according to a specific embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. The specific candidate anchor numbers, weights, thresholds, density budgets, etc., used in the following embodiments are illustrative examples and do not constitute a limitation on the scope of protection of this invention; those skilled in the art can adjust the corresponding parameters according to actual application scenarios, and the adjusted implementation methods still fall within the scope of protection of this invention.

[0020] The twin information element described in this invention refers to an independently arable object that is overlaid and displayed in an AR / MR scene, carrying twin data of a specific twin object, such as a data panel, a label, or an alarm marker. The twin object refers to a physical object in reality to which twin information is overlaid, such as a device or component, with each twin information element anchored to a twin object. The safety-critical area refers to an area in the real-world scene that cannot be obscured by overlaid information, such as areas with moving mechanical parts, passageways, and alarm indication areas. The presentation form refers to the simplified or complex form of a twin information element; the layout position refers to the placement position of a twin information element on the image plane. The image plane refers to the two-dimensional display plane corresponding to the current field of view of the head-mounted display device.

[0021] like Figure 1 As shown, the method of this invention begins with step S0. In step S0, each twin information element in the AR / MR scene and its respective anchored twin object, as well as several preset safety-critical areas, are acquired for the first time. The presentation form and layout position of each twin information element are then initialized. During initialization, the presentation form of each twin information element can be uniformly taken as its complete form, and the layout position can be taken as directly above the projection position of its anchored twin object. Subsequently, adjustments are made in each step according to layout costs, occlusion, and density. The anchoring relationship between each twin information element and its twin object, as well as the safety-critical areas, can be obtained from existing models and scene calibrations on the digital twin platform, which is well known to those skilled in the art and will not be elaborated upon here.

[0022] In step S1, the head posture of the head-mounted display device is obtained, and each twin object and each safety-critical area are projected onto the image plane of the current field of view according to this head posture. The head posture refers to the head position and orientation output by the head-mounted display device's pose sensor. The projection position refers to the area occupied by a twin object projected onto the image plane; the projection range refers to the area occupied by a safety-critical area projected onto the image plane. Due to operator head movement, the projections of each twin object and safety-critical area continuously change with the head posture, which is why the label layout needs to be updated frame by frame while maintaining inter-frame continuity.

[0023] In step S2, for each twin information element, several candidate anchor positions are generated around the projection position of its anchored twin object. The candidate anchor positions refer to several alternative locations for placing a twin information element. For example... Figure 3As shown, in one implementation, using the projection bounding box of the twin object on the image plane as a reference, one position is selected in each of the eight directions (top, bottom, left, right, and four corners) that is a predetermined distance from the bounding box, as a candidate anchor position for the twin information element. The projection bounding box refers to an axis-aligned rectangle that exactly encloses the projection position of the twin object. The candidate anchor positions are limited to the vicinity of the twin object, making the label close to its object and the leader line short, which facilitates the operator to quickly associate the label with the twin object.

[0024] In step S3, with the goal of optimizing layout cost, the layout position for each twin information element is selected from its candidate anchor positions for the current frame. The layout cost refers to a numerical value that measures the quality of a set of layouts, such as... Figure 4 As shown, the cost consists of three sums: the overlap cost between each twin information element, the lead distance cost between each twin information element and its anchor twin object, and the inter-frame displacement cost of each twin information element relative to its previous frame layout position. The overlap cost refers to the sum of the areas of the bounding boxes of each twin information element; the more intersections, the higher the cost, used to avoid tag overlap. The lead distance cost refers to the length of the line connecting each twin information element and its anchor twin object; the farther the line, the higher the cost, used to bring the tag closer to its anchor twin object. The inter-frame displacement cost is the product of the distance between the candidate anchor position of each twin information element in the current frame and its previous frame layout position on the image plane, multiplied by a preset coherence weight; the larger the displacement, the higher the cost.

[0025] The key to this invention lies in the inter-frame displacement cost: existing methods only consider overlap and lead lines, independently calculating a set of non-overlapping positions frame by frame. Even slight head movements cause changes in the projected positions of each object, resulting in an optimal layout that often differs significantly from the previous frame, causing the label to jump back and forth. This invention also incorporates the inter-frame displacement relative to the previous frame's layout position into the cost, using the previous frame's layout as the initial value for the current frame. Therefore, unless repositioning significantly reduces overlap or lead line costs, the label will remain near its previous frame position, allowing for smooth movement with head movement without jitter. To further suppress unnecessary back-and-forth switching, if the decrease in total layout cost after a twin information element is repositioned does not exceed a preset hysteresis threshold, the element retains its previous frame's layout position; that is, repositioning is only done when the benefits are sufficiently significant.

[0026] For example, suppose a tag was positioned above its object in the previous frame. If it remains in its original position in this frame, it will have a small overlap with adjacent tags, with an overlap cost of 0.2, a leader distance cost of 0.1, and an inter-frame displacement cost of 0, for a total cost of 0.3. If it is moved to the left of the object, there will be no overlap, and the leader distance cost will be 0.15, but the displacement relative to the previous frame will be larger, with an inter-frame displacement cost of 0.3, for a total cost of 0.45. The total cost of maintaining the original position (0.3) is less than the cost of moving to the left (0.45), so the tag remains in its original position in this frame. Only when the head continues to rotate and the overlap intensifies, causing the total cost of maintaining the original position to exceed the total cost of moving to the left of the object, will the tag be moved to the left. Thus, the tag repositioning is gradual and delayed, rather than jittering every frame.

[0027] The layout cost is solved incrementally: using the candidate anchor positions corresponding to the layout positions of each twin information element in the previous frame as the initial values, each twin information element is locally adjusted among its candidate anchor positions according to a preset number of iterations. Each adjustment selects the candidate anchor position that reduces the total layout cost, until the total layout cost no longer decreases or the preset number of iterations is reached. Figure 5 As shown, since the layout is hot-started from the previous frame and the head posture does not change much in most frames, convergence can often be achieved with only a few local adjustments. The single-frame layout solution has constant overhead, which makes it easy to run in real time frame by frame on head-mounted display devices.

[0028] In step S4, semantic occlusion avoidance and information density adaptation are applied to the layout selected in step S3. Semantic occlusion refers to a twin information element obscuring the twin object it is anchored to, or obscuring a security-critical area, in its layout position. For example... Figure 7 As shown, the occlusion determination and handling are as follows: when the rendering bounding box of a twin information element intersects with the projection position of its anchored twin object, or with the projection range of any safety-critical area, it is determined that semantic occlusion has occurred. For twin information elements with semantic occlusion, priority is given to selecting candidate anchors that do not have semantic occlusion. When all its candidate anchors have semantic occlusion, the rendering form of the twin information element is downgraded. The reason why not occluding the anchored object is also listed as a hard constraint is that if the label blocks the equipment it describes, the operator cannot match the data with the actual object. The reason why safety-critical areas are listed as hard constraints is that blocking a moving robotic arm or passageway may directly cause a safety accident.

[0029] The aforementioned adaptive information density refers to prioritizing and simplifying the presentation of information when it is too dense within the field of view. For example... Figure 6As shown, the presentation modes include full mode, simplified mode, folded mode, and icon mode, with the presentation area decreasing sequentially: the full mode presents the entire data panel, the simplified mode only presents key readings, the folded mode shrinks to a single title line, and the icon mode presents only one icon. The cumulative presentation area refers to the sum of the presentation areas of all twin information elements within the current field of view; the preset density budget refers to the upper limit of the allowed cumulative presentation area, usually given as a proportion of the field of view area. When the cumulative presentation area exceeds the preset density budget, the presentation mode of each twin information element is reduced by one level according to priority from minor to major, until the cumulative presentation area does not exceed the preset density budget; thus, important elements remain intact, minor elements are folded first, and information that is too dense is no longer stacked indiscriminately. The priority of the twin information elements is assigned by the upper-layer application according to their alarm level, security impact, etc.

[0030] After comparative testing by the inventors in scenarios where multiple maintenance personnel wore head-mounted displays and walked around for inspections, they found that the jitter of the labels when laying out the layout frame by frame was the source of the strongest discomfort reported by the operators, even exceeding the information density itself. After introducing the cost of inter-frame displacement, the operators generally reported that the labels were more stable and less prone to shaking. Therefore, it is supported to regard inter-frame continuity as an independent goal of the layout, while most existing view management methods only take non-overlapping as the goal.

[0031] In step S5, the current AR / MR rendering frame is rendered and output according to the selected layout position and presentation form of each twin information element for display on the head-mounted display device. The layout position of each twin information element in this frame is used as the initial value for solving the layout of the next frame, and the loop iterates in a closed loop. The AR / MR rendering frame refers to a frame of image output to the head-mounted display device after the twin information elements are superimposed. In addition, when the cumulative presentation area of ​​the field of view where the folded element is located decreases and is no longer crowded, the folded elements are restored step by step according to priority, so that the presentation automatically recovers with the tightness of the information density.

[0032] The safety-critical areas are manually calibrated or obtained through semantic segmentation of the scene image, including at least one of the following: mechanical moving parts area, passageway area, and alarm indication area. For moving mechanical parts, their projection range is updated frame by frame as they move, so that the tag always avoids its current position. Regarding the restoration of the presentation form: when the cumulative presentation area in the current field of view falls back to a preset recovery ratio that does not exceed the preset density budget, the presentation form of the twin information elements in the folded form is restored one level in order of priority from primary to secondary. The preset recovery ratio is less than one, for example, 0.8, so that the trigger point of restoration is lower than the trigger point of folding, forming a buffer zone to avoid flickering caused by repeated folding and restoration near the density budget. The preset density budget can be 0.3 to 0.5 of the field of view area, typically 0.4. The preset coherence weight and preset hysteresis threshold are used to adjust the stickiness of the tag following the head movement and the sensitivity of the position change, respectively. The larger the coherence weight and the larger the hysteresis threshold, the more stable the tag is, but the slower it follows.

[0033] The three components of the layout cost can be assigned different weights according to the scenario: when the operator moves quickly and stability is more important, the coherence weight of the inter-frame displacement cost is increased; when objects are dense and non-overlapping is more important, the weight of the overlap cost is increased. All costs are normalized and then summed to make the weights comparable. Incremental solution is efficient because the head pose changes between adjacent frames are usually small, and the optimal layout of the previous frame is still close to optimal for the current frame. Convergence can be achieved by making local adjustments to a few elements where new overlaps or occlusions occur. In a very few frames where the head rotates significantly, the current optimal layout is taken when the number of iterations reaches the upper limit, and optimization continues in the next frame without affecting real-time performance.

[0034] Each twin information element is connected to its anchored twin object by a leader, with one end of the leader connected to the label and the other end pointing to the object. This allows the operator to see the correspondence at a glance, even when the label is moved to the side of the object. The leader distance cost is the penalty for the length of the leader, ensuring that the label is as close to its object as possible without overlapping. When the label changes position or shape due to occlusion avoidance or density degradation, the leader is updated accordingly, always maintaining the visibility of the association between the label and the object.

[0035] In terms of engineering implementation, projection, candidate anchor generation, incremental layout, occlusion and density handling can all be completed in real time frame by frame on the head-mounted display device or edge rendering end. The single-frame processing has constant overhead. With the incremental solution that starts from the previous frame, it can run stably at common head-mounted display frame rates. The presentation form and layout position of each twin information element are updated according to the frame increment. Elements that have not changed directly use the results of the previous frame, further reducing overhead.

[0036] This method can also be coordinated with the operator's attention or cognitive state: when the operator's attention is detected to be focused on a certain object, the priority of the information elements anchored to that object is temporarily increased, so that it is the last to be collapsed during density degradation; when the operator is detected to be under high cognitive load, the preset density budget is reduced to make the field of view simpler. In this way, the adaptiveness of this method is not only aligned with head posture, but also with the operator's attention and load, further meeting their usage needs.

[0037] This method is also applicable to multi-operator collaboration and remote guidance scenarios: each operator is laid out separately according to their own head posture, without affecting each other; during remote guidance, the information elements marked by the instructor and the information elements of the on-site operators are included in the layout and density management together, and the remote annotations are avoided from blocking the safety-critical areas on-site according to the unified semantic occlusion avoidance and density budget constraints.

[0038] like Figure 2 As shown, the present invention also provides an AR / MR information adaptive presentation and layout device corresponding to the above method, which includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the steps of the above method. Figure 2 The device is schematically divided into a projection module, a candidate anchor generation module, an incremental layout module, an occlusion avoidance and density adaptation module, and a presentation output module according to the functions it performs. This division is only a schematic division in terms of logical functions. Each module is used to perform the above steps S1, S2, S3, S4 and S5 respectively. The specific processing is the same as the above method, so it will not be repeated.

[0039] The present invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method; and a computer program product containing a computer program that, when executed by a processor, also implements the steps of the above-described method. The electronic device carrying the above-described program can be an AR / MR head-mounted display, an edge rendering host, or a workstation, etc., which can be selected by those skilled in the art according to deployment needs, and will not be elaborated upon here.

[0040] As a specific example, such as Figure 8As shown, taking AR operation and maintenance inspection of a substation as an example, the complete execution process of the method described in this invention is explained. In this embodiment, the operator wears an AR headset to walk and inspect the equipment. The field of view contains four twin objects: a main transformer, two disconnect switches, and a station service transformer, each anchored to a data panel. The preset candidate anchor positions are the eight directions surrounding the object's outer frame. The preset coherence weight is 0.6, the preset hysteresis threshold is 5% of the layout cost, the preset density budget is 0.4 of the field of view area, and the preset rise ratio is 0.8. Next to the main transformer, there is an on-load tap changer drive mechanism that is in operation, which is marked as a safety-critical area.

[0041] Step S0 first acquires the four data panels and their anchored objects, as well as the safety-critical region, and initializes the four panels to their complete form, placing them directly above their respective objects. Step S1 projects the four twin objects and the safety-critical region onto the image plane according to the current head pose, obtaining their respective projection positions and projection ranges. Step S2 generates eight candidate anchors around the projection bounding box of each object.

[0042] Step S3 uses the previous frame layout as the initial value for incremental layout solving: When the operator slowly turns their head, the four panels translate along with the projection of their respective objects. Due to the inter-frame displacement cost, each panel remains in the same position relative to the object in the previous frame without any jumps. When the rendering bounding box of the main transformer panel and a disconnector panel begins to intersect and the overlap cost increases, the solver moves the disconnector panel from above the object to the left of the object. Because this repositioning causes the decrease in total layout cost to exceed the hysteresis threshold, the repositioning is performed, and the two panels no longer overlap. During this process, the leads between each panel and its object are updated with the panel position, and the operator can always see which panel corresponds to which device. This repositioning is a smooth migration that occurs only after the overlap has accumulated to a certain extent, rather than a back-and-forth jump every frame.

[0043] Step S4: Obstruction Avoidance and Density Adaptation: The original layout of the station transformer panel happens to obscure the projection range of the on-load tap changer drive mechanism, a critical safety area that is currently in operation, constituting semantic obstruction. Therefore, the station transformer panel is reselected to a candidate anchor position in the lower right corner that does not block the mechanism. At this time, the cumulative presentation area of ​​the four complete panels in the field of view is 0.46 of the field of view area, which exceeds the density budget of 0.4. Therefore, according to the priority from minor to major, the station transformer panel with the lowest priority is first downgraded from a complete form to a simplified form, and the cumulative presentation area is reduced to 0.39, which does not exceed the budget. The downgrade stops, and important panels such as the main transformer remain intact.

[0044] Step S5 renders and outputs the current AR / MR rendering frame, using the current frame layout as the initial value for the next frame. As the operator approaches and the equipment becomes less dense, the cumulative rendering area drops to 0.31, less than 0.8 times the budget (0.32). Therefore, according to priority, the station's variable panels are restored from a simplified form to a complete form, starting with the primary ones and moving to the secondary ones. Throughout the process, the four panels follow the operator's movements smoothly without shaking or obstructing any objects or the critical safety area. When information is too dense, the panels fold orderly; when information is sparse, they automatically restore. The operator does not need to stop to find the panels, nor is they distracted by label shaking or obstruction. The twin information remains stably, clearly, and orderly superimposed next to the corresponding equipment, ensuring both inspection efficiency and safety.

[0045] Furthermore, the layout position, presentation form, and occlusion and density handling of each twin information element recorded in each frame by this method can be reported as an objective record of AR presentation quality, facilitating post-event evaluation of the appropriateness of the presentation strategy and cross-verification with the digital twin on the device side. The above records only involve the presentation strategy itself and do not involve the operator's personal identity information. When deployed to a new scene, a section of actual inspection data can be replayed to observe the stability, occlusion, and folding frequency of the tags, thereby performing an initial calibration of parameters such as coherence weight and density budget. Then, in actual use, continuous fine-tuning can be carried out based on operator feedback to gradually adapt the presentation strategy to the characteristics of the scene.

[0046] In contrast, if the inter-frame displacement cost, semantic occlusion avoidance, and density adaptation of this invention are not adopted, and the existing frame-by-frame independent non-overlapping layout is used instead: on the one hand, since each frame independently calculates a set of non-overlapping positions, the four panels frequently jump and flicker above, below, and to the left and right of the object when the operator turns their head, making stable reading difficult; on the other hand, since the layout only considers the non-overlapping labels, the station-mounted variable panel may always be pressing on the critical safety area of ​​the moving transmission mechanism, posing a safety hazard, and the four complete panels always fill the field of view, making it difficult to see the actual scene. It can be seen that it is the inter-frame displacement cost, semantic occlusion avoidance, and information density adaptation that jointly eliminate jitter, occlusion, and overload, a causal effect that existing methods do not possess.

[0047] The candidate anchor positions of this method are not limited to eight directions; denser candidate positions can also be selected according to the shape of the object. The number of presentation levels can also be more than four. In some boundary cases, this method handles them as follows: when a twin information element is reduced to an icon form and semantic occlusion still occurs, it is temporarily hidden and a very small cue point is left next to the object, and it is restored after the occlusion is removed; when the head pose is rapidly and significantly rotated and a large number of objects enter and leave the field of view, the newly entered elements are placed directly according to the current layout without considering the inter-frame displacement cost, while the inter-frame displacement cost is still included for elements still in the field of view, so that the new elements are in place and the old elements do not jitter; when a certain safety-critical area occupies too much of the field of view, causing some elements to have nowhere to be placed, priority is given to ensuring that the safety-critical area is not occluded, and the secondary elements that have nowhere to be placed are folded or moved into the list at the edge of the screen.

[0048] It should be noted that there is a predictable causal relationship between the technical effects and technical means in the human-computer interaction and display scenarios addressed by this invention. Therefore, the effects are explained above using causal reasoning: Including the inter-frame displacement relative to the previous frame's layout position in the layout cost and using the previous frame's layout as the initial value ensures that the label remains in its original position unless the benefit of repositioning is significant. This is a direct result of the inter-frame displacement cost, resulting in smooth and non-juddering label movement. Using non-occlusion of anchored objects and safety-critical areas as a hard constraint forces violators to change anchor positions or be downgraded. This is a direct result of the hard constraint, as the label no longer obstructs objects and safety areas. Folding and restoring according to priority based on density budgeting is a direct result of density adaptation, allowing for orderly selection when information is too dense. The above causal chain is valid without relying on a specific dataset, and those skilled in the art can implement and verify this invention based on the description in the specification.

[0049] The above mechanism also foresees the applicable boundaries of this method: in complex AR operation and maintenance scenarios with numerous objects, dense information, and frequent operator movement and head turning, the advantages of this method over existing frame-by-frame independent layout are most obvious; while in simple scenarios with few objects, almost stationary operators, and information never exceeding the budget, inter-frame displacement is small, occlusion and density degradation are rarely triggered, and this method degenerates into near-conventional view management without introducing additional negative impacts. Therefore, this method brings stability, security, and non-overload benefits without compromising the inherently simple scenario, and has good applicable boundaries.

[0050] Regarding the design of the presentation formats, the amount of information carried and the presentation area occupied by each format decrease sequentially: the complete format carries all the data panels and curves of the twin information element for detailed viewing; the simplified format only carries a few key readings and statuses for overview; the collapsed format is folded into a single title, indicating that information can be expanded; and the icon format uses only an icon to indicate the existence of the object. Degradation proceeds step-by-step from complete to simplified, then to collapsed, and finally to icon. Operators can actively expand a collapsed element when needed, and the expanded element is temporarily promoted one level, allowing for compatibility between on-demand viewing and automatic selection.

[0051] All the above preset parameters can be objectively calibrated according to the scenario: the preset coherence weight and preset hysteresis threshold are calibrated based on the operator's subjective evaluation of the tag stability during trial operation; the preset density budget is calibrated according to the head-mounted display's field of view and the amount of information the operator can process simultaneously; the preset lead spacing is set according to the tag font size and object size; the priority of each twin information element is given by safety procedures and maintenance requirements. After calibration, each parameter can still be fine-tuned within the aforementioned value range according to the specific scenario, and the adjusted implementation method is still within the protection scope of this invention.

[0052] Compared to existing practices that fix all labels to objects or cram them into fixed header display areas, the labels of this invention not only follow the object and are anchored around it, but also remain stable between frames due to inter-frame displacement constraints. Furthermore, they actively avoid objects and safety-critical areas through semantic occlusion avoidance, and prioritize information in a timely manner when crowded, achieving adaptive information density. These four points work together to ensure that the superimposed information is both realistic and does not overshadow the main content.

[0053] The priority of each twin information element is not fixed and can be dynamically adjusted according to the operating status: when a twin object triggers an alarm, the priority of its anchored information element is temporarily increased, so that it is the last to be collapsed during density degradation, or even automatically restored from the collapsed state to the complete state, ensuring that the alarm information is not overwhelmed; the priority is restored after the alarm is cleared. In this way, the information density adaptive and alarm status linkage of the present invention prioritizes the most visible content when information is too dense.

[0054] In handling candidate anchor positions, the boundaries of objects entering and leaving the field of view must also be taken into account: when the projection part of a twin object moves out of the field of view, candidate anchor positions are generated only on the edge side where it is still within the field of view, so that the label is not placed outside the field of view; when the object completely moves out of the field of view, its information elements do not participate in the layout temporarily, and are restored when the object re-enters the field of view, and the most recent layout position is used as the initial value for restoration, so that the re-entering label does not jump abruptly.

[0055] In summary, this invention completes an adaptive loop in each frame, encompassing projection, candidate anchor generation, incremental layout, occlusion avoidance and density adaptation, output, and initial value write-back. Projection introduces changes in the 3D scene into the image plane; candidate anchors limit the label's range of movement; incremental layout ensures continuity at the cost of inter-frame displacement; occlusion avoidance protects objects and safe zones from obstruction; density adaptation makes orderly selections when density is too high; and outputting and writing back the initial value ensures the next frame continues from the current frame. This loop is applicable to both optical and video-based see-through head-mounted displays: in the former, the projected safety-critical area corresponds to the real-world area directly seen by the operator through the lenses; in the latter, it corresponds to the corresponding area in the real-world image captured and displayed by the camera. The method steps are consistent in both cases. It is this frame-by-frame closed adaptive loop that ensures the superimposed twin information remains stable, closely aligned, and unloaded during real-world operations involving continuous head movement and information changes.

[0056] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for adaptive presentation and layout of digital twin AR / MR information, characterized in that, The method includes: step S0, first acquiring each twin information element in the AR / MR scene and its respective anchored twin object, as well as several preset safety key areas, and initializing the presentation form and layout position of each twin information element; Step S1: Obtain the head posture of the head-mounted display device, and project each twin object and each safety-critical area onto the image plane of the current field of view according to the head posture to obtain the projection position of each twin object and the projection range of each safety-critical area. Step S2: For each twin information element, generate several candidate anchor positions around the projection position of its anchored twin object; Step S3: With the goal of optimizing the layout cost, select the layout position of each twin information element from its candidate anchor positions in this frame. The layout cost includes the overlap cost between each twin information element, the lead distance cost between each twin information element and its anchored twin object, and the inter-frame displacement cost of each twin information element relative to its layout position in the previous frame. Step S4: Apply semantic occlusion avoidance and information density adaptation to the selected layout: When any twin information element occludes the projection position of its anchored twin object or occludes the projection range of any safety-critical area in its layout position, reselect a candidate anchor position for the twin information element or downgrade its presentation form. When the cumulative presentation area of ​​all twin information elements in the current field of view exceeds the preset density budget, the presentation form of the twin information elements is downgraded to a folded form according to the priority from minor to major. Step S5: Render and output the current AR / MR rendering frame according to the selected layout position and presentation form of each twin information element, and use the layout position of each twin information element in this frame as the initial value of the layout of the next frame, and iterate in a closed loop.

2. The method according to claim 1, characterized in that, In step S2, several candidate anchor positions are generated around the projection position of the twin object to which it is anchored, including: taking the projection bounding box of the twin object on the image plane as a reference, taking one position at a preset lead line interval from the bounding box in each of the eight directions of the bounding box (up, down, left, right and four corners) as the candidate anchor position of the twin information element.

3. The method according to claim 1, characterized in that, In step S3, the layout position of the current frame is selected with the goal of optimizing the layout cost. This includes: using the candidate anchor position corresponding to the layout position of each twin information element in the previous frame as the initial value for solving the problem, and making local adjustments to each twin information element among its candidate anchor positions one by one according to the preset number of iterations. Each adjustment selects the candidate anchor position that reduces the total layout cost, until the total layout cost no longer decreases or the preset number of iterations is reached.

4. The method according to claim 1, characterized in that, The inter-frame displacement cost is the product of the distance between the candidate anchor position of each twin information element in this frame and its layout position in the previous frame on the image plane, multiplied by a preset coherence weight. Furthermore, if the decrease in total layout cost after a twin information element is reselected as a candidate anchor does not exceed a preset hysteresis threshold, the twin information element retains the layout position of the previous frame.

5. The method according to claim 1, characterized in that, The determination and handling of occlusion in step S4 includes: when the presentation bounding box of a twin information element intersects with the projection position of its anchored twin object or with the projection range of any safety-critical area, it is determined that semantic occlusion has occurred. For twin information elements that are semantically occluded, priority is given to selecting candidate anchors that are not semantically occluded. When all of their candidate anchors are semantically occluded, the presentation of the twin information element is downgraded.

6. The method according to claim 1, characterized in that, The presentation forms include full form, simplified form, folded form and icon form, and the presentation area of ​​the four forms decreases in that order. In step S4, when the cumulative presentation area exceeds the preset density budget, the presentation form of the twin information elements is reduced by one level in order of priority from minor to major, until the cumulative presentation area does not exceed the preset density budget.

7. The method according to claim 1, characterized in that, The safety-critical areas are determined manually or by semantic segmentation of the scene images, and include at least one of the mechanical moving parts area, passageway area, and alarm indication area; When the cumulative display area within the current field of view falls back to a preset recovery ratio that does not exceed the preset density budget, the display form of the twin information elements in the folded state is restored to level one in order of priority from primary to secondary.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

9. A digital twin AR / MR information adaptive presentation and layout device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the method of any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.