Virtual-real combined scene space interaction self-adaptive online processing method
By constructing a collaborative situation map and anchoring area units, the problem of stable display of virtual content in multi-user virtual-real combined scenarios was solved, achieving a unified spatial layout and taking into account individual perspectives in a multi-user environment, thus improving assembly efficiency and training effectiveness.
Patent Information
- Application Number
- CN202610080988.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-21
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2046-01-21
AI Technical Summary
Existing technologies struggle to achieve a unified spatial layout of virtual content in a shared virtual-real scenario involving multiple users and roles. This results in inconsistent virtual content placement among different users, virtual content obscuring real devices, and instability of virtual content when the environment changes, impacting assembly efficiency and training effectiveness.
By constructing a collaborative situation map, combining user posture, field of vision, and operation trajectory, anchoring zone units are generated. Based on the resource status of the anchoring zone and component anchoring requirements, global layout and personalized fine-tuning of group layer and role layer components are carried out, and strategy parameters are updated in real time to ensure the stable presentation of virtual content.
It achieves stable display of virtual content in multi-user virtual-real hybrid scenarios, reduces virtual content drift and occlusion issues, improves assembly efficiency and training effectiveness, and adapts to changes in environment and computing resources.
Smart Images

Figure CN121564296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, specifically to an adaptive online processing method for virtual-real combined scene space interaction. Background Technology
[0002] In virtual-real hybrid scenarios such as remote guidance at the manufacturing site, training on complex equipment assembly and adjustment, and multi-person safety inspections, personnel wearing head-mounted displays or mobile terminals overlay virtual markers, step prompts, hazard warnings, and virtual tool panels onto real workshops, machine rooms, or experimental areas. As production line scale and the number of trainees increase, a single real space often hosts multiple users, multiple roles, and various types of virtual content simultaneously. This necessitates continuously adding virtual content within limited field of view and available space, while meeting the viewing and operational needs of different users. Currently, in engineering practice, more and more enterprises are applying this to assembly, remote expert guidance across regions, and classroom training, leading to increasingly higher requirements for adaptive spatial layout. Existing virtual-real hybrid systems generally use synchronous positioning and mapping algorithms and 3D reconstruction technology to model the on-site environment and have a certain degree of automatic placement function for virtual content, such as attaching prompts to equipment or fixing interface panels in front of the user. Some systems also adaptively adjust parameters such as display clarity and refresh rate based on ambient lighting, tracking quality, or network status.
[0003] This type of method can reduce the difficulty of manually placing the interface and repeatedly adjusting its position when there is only one user, a small amount of virtual content, and a slow change in the environment. However, in real-world scenarios where multiple users and roles are working simultaneously, and there is a large variety and quantity of virtual content, this method can only adopt a local rule layout from a single perspective, lacking a unified modeling and scheduling mechanism that considers multiple users, multiple roles, and the overall collaborative relationships.
[0004] When multiple users share the same virtual-real hybrid scenario, several issues arise: First, the positions and relative relationships of virtual content vary among different users, making it difficult for collaborators to consistently refer to the same virtual markers or prompts, potentially leading to misunderstandings of instructions. Second, virtual prompts or panels may obscure important locations, safety signs, or actual operating areas of the real equipment, causing on-site personnel to switch their attention between virtual content and real targets. Third, when equipment vibrates, the environment obscures the view, or the network fluctuates, the tracking stability and computing resources in some areas change significantly. If the virtual content remains fixed in these areas, it can easily cause screen shake, virtual marker drift, or even sudden interface disappearance, interrupting operations. These problems are amplified in scenarios involving long periods, multiple batches, and multiple participants, affecting assembly efficiency and training effectiveness, and interfering with the transmission of safety-related information.
[0005] Therefore, the current technical problem to be solved is:
[0006] When multiple users and roles share the same virtual-real hybrid scenario, and environmental stability and computing resources change over time, existing technologies struggle to coordinate the spatial layout of a large number of virtual markers, prompts, and interface components under a unified model. They also struggle to avoid areas with unstable tracking and resource shortages in a timely manner, and simultaneously ensure consistency of different users' perspectives, unobstructed real key areas, and controllability of cognitive load for groups and individuals. Summary of the Invention
[0007] (a) Technical problems to be solved
[0008] To address the shortcomings of existing technologies, this invention provides an adaptive online processing method for virtual-real integrated scene spatial interaction. It generates a collaborative situation map by combining user posture, field of vision, line of sight, and operation trajectory. Based on this, it performs anchor zone scheduling according to the anchor zone resource status and component anchoring requirements. First, it completes the global layout of components at the group and role levels, and then performs personalized fine-tuning of components within each user's field of vision. During operation, it performs cross-layer adjustments based on anchor zone tracking confidence and user cognitive load, and updates strategy parameters online, achieving stable presentation of virtual content layout in multi-user virtual-real integrated scenes, thus solving the technical problems described in the background section.
[0009] (II) Technical Solution
[0010] To achieve the above objectives, the present invention provides the following technical solution:
[0011] The adaptive online processing method for virtual-real scene spatial interaction includes: collecting image and depth data of virtual-real scene through multiple sensors, 3D reconstruction and semantic segmentation, constructing scene semantic topology map containing adjacency, visibility and access relationships, and dividing the scene into anchoring area units with resource status according to time-series tracking confidence and anchorability level.
[0012] The virtual content is abstracted into hierarchical spatial interaction components that carry steps, interaction types and spatial levels. Anchoring requirements are set for the components, and a collaborative situation diagram with users, task objects and components as nodes is constructed based on user posture, field of vision and line of sight and operation trajectory.
[0013] Based on the resource status of the anchoring area unit and the anchoring requirements of the hierarchical spatial interaction components, the anchoring area scheduling is performed according to the adaptability. High-confidence anchoring areas are allocated to group layer and role layer components first, and group layer layout and individual layer fine-tuning are carried out in combination with multi-user perspectives within the anchoring area.
[0014] During operation, the tracking confidence level of the anchor zone unit and the user cognitive load in the collaboration status map are monitored. When the confidence level or cognitive load exceeds the threshold, the position and level of the deployed components are adjusted across layers, and the strategy parameters are updated according to the task completion status and user intervention records.
[0015] Furthermore, the multi-sensor system includes a color video camera, an infrared depth camera, an inertial measurement unit, and a laser ranging device. Simultaneous localization and mapping (SMR) algorithms are used to process environmental data, generating a 3D scene model that includes a 3D point cloud or mesh model and the attitude trajectory of the terminal device. A scene semantic topology map is then established based on the geometric proximity relationships, unobstructed line-of-sight relationships, and passability relationships between semantic objects.
[0016] Furthermore, the anchoring area resource status information includes the time-series tracking confidence sequence of the anchoring area unit within a preset time window, the visibility markers for different user viewpoints, the accessibility markers for different user bodies and hands, and the number and type information of the hierarchical spatial interaction components attached to the anchoring area unit. At the end of step one, the anchoring area resource status information is associated and stored with the scene semantic topology map.
[0017] Furthermore, each hierarchical spatial interaction component carries a step identifier, a type identifier, and an interaction type identifier, and records the semantic object or semantic region identifier to be associated. The anchoring requirements include the acceptable lower limit of temporal tracking confidence for the hierarchical spatial interaction component, the allowed range of anchoring zone levels, and the target field of view orientation and field of view depth range.
[0018] Furthermore, for each user, head posture, field of vision, and reachable space are maintained. The semantic object currently being gazed at and the hierarchical spatial interaction component being operated are associated with the user's role information. When constructing the collaborative situation map, the user, semantic object, and hierarchical spatial interaction component are used as nodes. The edges of co-viewing relationship, operation relationship, and guidance relationship are marked according to the line of sight and operation trajectory. Cognitive load and participation are recorded on the nodes.
[0019] Furthermore, when scheduling anchor areas for hierarchical spatial interaction components, the components are divided into group-related components and role-related components according to hierarchical attributes. The anchor area resource status is queried, and anchor area units with higher time-series tracking confidence and visibility are preferentially allocated to group-related components, followed by role-related components. When there are insufficient available anchor area units, multiple components are displayed in rotation on the same anchor area unit according to a preset time slice.
[0020] Furthermore, when fine-tuning the position and size of components within each user's field of vision, the position and size of the hierarchical spatial interaction components assigned to that user are adjusted based on each user's field of vision and accessible space.
[0021] Keep the anchor area scheduling result unchanged, so that components do not obscure semantic objects with occlusion importance weights, and limit the user's head rotation and arm extension to a preset range.
[0022] Furthermore, while monitoring the time-series tracking confidence and user cognitive load of the anchored area, it also monitors processing capacity and network conditions. When the processing capacity or network conditions are lower than the preset threshold, the presentation of the hierarchical spatial interactive components anchored on the corresponding anchored area is adjusted. Components presented as 3D models are simplified into icons or text labels, and components displayed in the group layer are adjusted to be displayed in the role layer or individual layer.
[0023] Furthermore, the layout and hierarchical control strategy parameters include the cognitive load threshold of each role, the resource priority weight of various anchoring area units, the cost weight for occlusion, viewpoint offset and anchoring area occupation in layout evaluation, and the sensitivity parameters for triggering hierarchical adjustments of layered spatial interaction components between the group layer, role layer and individual layer. The above parameters are corrected when the strategy parameters are updated.
[0024] Furthermore, the strategy parameter updates are based on recorded task performance data, which includes completion time, number of misoperations and redoes, and the number of times the user adjusts the layout. After the task is completed, the task performance data is summarized, and the cognitive load threshold, resource priority weight, and cost weight of each role are corrected. The corrected parameters are then written back to the layout and hierarchy control strategy parameters.
[0025] (III) Beneficial Effects
[0026] This invention provides an adaptive online processing method for virtual-real scene spatial interaction, which has the following beneficial effects:
[0027] By introducing temporal tracking confidence and anchoring zone unit resource status into the scene semantic topology graph, the spatial anchoring of virtual content is based on a unified characterization of environmental stability and spatial availability. In scenarios with occlusion, lighting changes, and slight device vibration, the virtual markers and operation panels are kept stably displayed, reducing interference caused by virtual content drift and frequent repositioning, and providing a reliable spatial benchmark for subsequent multi-user spatial layout.
[0028] This approach links task flows, role assignments, and spatial content within the same structure, distinguishing between group, role, and individual layers of information and reducing inconsistencies in how different users interpret the same device and virtual marker. Based on the anchoring requirements of layered spatial interaction components, high-quality anchoring area units are prioritized for group layer components and key role layer components. Then, their positions and sizes are fine-tuned within each user's field of view. This ensures the spatial layout is constrained by both the scene's semantic topology and the collaborative situational map, guaranteeing consistent placement of shared components while also considering readability and ease of operation from an individual perspective.
[0029] During operation, the actual tracking quality of anchor zone units, system resource status, and user cognitive load in the collaboration situation diagram are continuously monitored. When it is found that the stability of a certain anchor zone unit has decreased or a certain type of user is overburdened, the hierarchical spatial interaction component is triggered to migrate between different anchor zone units, making hierarchical adjustments between the group layer, role layer, and individual layer, and simplifying complex three-dimensional content into icons or text lists as needed, thereby maintaining the balance of the overall interaction rhythm when the environment and task intensity change.
[0030] Without changing the overall processing flow, the cognitive load thresholds for different roles, the resource priority of anchor units, and the weight configuration in layout evaluation are gradually adjusted so that the same method can adaptively converge across different workshops, teaching scenarios, and remote collaborative projects, reducing the workload of manual parameter tuning and scenario reconstruction. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the online processing method for adaptive interaction between virtual and real scene spaces according to the present invention. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Please see Figure 1 This invention provides an adaptive online processing method for virtual-real integrated scene space interaction, including:
[0034] Step 1: In spaces such as assembly workshops, computer rooms, or teaching and experimental areas, color video cameras, infrared depth cameras, inertial measurement units, and laser ranging devices are often installed in different locations and sample at different frequencies. If these raw data are used directly, it is difficult to form a unified description of the changes over time in the same ground area and the same equipment surface, and it is even more difficult to predict the tracking stability and anchoring capability for the same location in subsequent steps.
[0035] Therefore, it is necessary to first integrate multi-source observations into a set of basic scene units by unifying the spatial coordinate system. Subsequent updates can directly correct this set of basic units without gradually aligning from the raw data level. When selecting the workspace coordinate system, the installation positions and attitudes of all sensors are calibrated into this coordinate system all at once.
[0036] During operation, each frame of color image, depth image and attitude data of inertial measurement unit can be converted into a three-dimensional point set and pose in a unified coordinate system. According to the spatial sampling accuracy and business requirements, the workspace is divided into continuous voxel units or mesh units, and each unit corresponds to a scene basic unit.
[0037] When new sensor observations arrive, the corresponding pixel or point cloud sample is assigned to the basic unit based on spatial location parameters, and an additional observation is added to this basic unit. This allows for the accumulation of a large number of observations of the same physical area over time, rather than being scattered across different coordinate systems or data structures. The initial system deployment phase involves a complete inspection in the workshop or experimental area. For example, a color video camera on a head-mounted display terminal acquires multi-view images, while infrared depth cameras and laser ranging devices acquire depth data. This data is then mapped to the workspace coordinate system, and a regular grid is divided with the ground as a reference plane, forming a series of basic unit sets that closely resemble the real structure.
[0038] During device operation, when a sensor acquires a new image frame or depth frame, it can read the spatial coordinates of each point corresponding to the pose in the frame into this basic unit. At the same time, a record of descriptive quantities such as illumination, texture clarity, and motion blur in the observation history is added to this basic unit. Therefore, the observable situation in the same area can always be continuously repeated in the same basic unit, and the observation content acquired by different sensors can also be repeated in the same basic unit.
[0039] In application scenarios, each observable region has its own basic unit in a coordinate system, which can provide a spatial index for subsequent semantic segmentation, topological relationship establishment, and temporal confidence measurement. Multi-source observation data are integrated at the basic unit level, avoiding the need for subsequent algorithms to frequently perform registration on the original image or point cloud. The computational path of the scene semantic-topology graph construction process also becomes clear and controllable.
[0040] After constructing the basic units, the localization results of a single frame or a few frames are insufficient to determine whether a spatial region can be used for content stabilization in the future. It is necessary to use the observation history over a certain time window, incorporating occlusion rate, feature sharpness, device vibration, and other factors into the temporal confidence field to characterize the stability and predictability of each basic unit across different time ranges. An observation history sequence is maintained for each basic unit, including observation time, observation angle, lighting conditions, texture sharpness, and local motion. Observational attributes are statistically analyzed over a sliding time window to construct a comprehensive index function sensitive to stability, which is then used as a confidence level ranging from 0 to 1.
[0041] As time progresses, new observations continuously enter the window, while older observations are gradually weakened or discarded, thus causing the confidence surface to evolve slowly over time, rather than undergoing drastic jumps under single-frame fluctuations.
[0042] The temporal confidence function of a given basic unit is defined as follows: given the spatial location parameters of the basic unit... and current time parameter Next, calculate an observation stability index. Then, the confidence value is obtained through exponential mapping:
[0043]
[0044] Among them, the time series confidence function Spatial location parameters In time parameters The stability of time as an anchor candidate location is limited to a value between 0 and 1; observation stability index : A non-negative quantity calculated based on the number of occlusions, the magnitude of texture sharpness change, and the magnitude of local motion of the basic unit within a preset time window;
[0045] Attenuation coefficient : Positive real number, acting on the influence of the observation stability index on the confidence score; where, the observation stability index Occlusion frequency, texture blur, and other parameters can be quantized into a consistent value through summation. The number of occlusions is calculated by detecting whether the basic unit is occluded by other objects; texture sharpness is determined by image gradient or edge response; and the local motion amplitude is calculated by matching point clouds from adjacent frames to determine the average displacement, which is then summed to obtain the final value. .
[0046] Optional, observe stability index Defined as a linear combination of three normalized sub-indices:
[0047]
[0048] Among them: occlusion metric :Location Within the sliding time window, the base cell containing the location is marked as the percentage of frames occluded out of the total number of frames in that window, with the value... Texture variation metric Within the same window, perform inter-frame differencing on the gradient energy or local contrast of the image patch corresponding to the basic unit, and normalize it to the maximum possible difference. Motion measurement Within the same window, the average displacement of the center point of the basic unit is obtained based on the point cloud registration of adjacent frames, and then normalized to the preset maximum displacement. Weighting coefficient : A non-negative real number, set by the deployment personnel based on the scene's sensitivity to occlusion, texture, and motion.
[0049] When using it, a time-series confidence function is introduced based on the original unit observation history. This allows the semantic-topological graph of a scene to not only reflect the current geometric and semantic state, but also the expected stability over a future period. Any subsequent spatial layout requiring anchor points can then directly incorporate the temporal confidence function. This reduces the probability of virtual tag jitter drift.
[0050] Even after establishing basic units with temporal confidence attributes, not every basic unit can be directly used as a candidate anchor point for virtual content. If the area of a basic unit is too small, its surface tilt angle is too large, or the corresponding semantic object itself is not suitable for being covered, then even if these basic units have high confidence, they are not suitable as actual anchor locations.
[0051] Therefore, one or more semantic objects are selected from the scene semantic-topology graph as candidate anchoring carriers, such as flat surfaces of equipment shells, fixed walls in rooms, or the surface of workbenches. The surface of each candidate object is divided according to basic units, and adjacent basic units with similar geometric properties are grouped into several anchoring area units.
[0052] For each anchoring zone unit, calculations are performed on its area, the angle between the normal and gravity directions, and historical occlusion. An anchoring fitness coefficient is then derived from the relevant data of these factors to determine how many virtual components the anchoring zone can support and what type of virtual components they can support. Each anchoring zone can be numbered... And construct the anchoring fitness coefficient ,like:
[0053]
[0054] In the formula, the anchoring fitness coefficient Number is The anchoring area unit serves as the overall adaptability of the virtual content carrying area; area parameters : The actual area of the anchoring zone element on the candidate surface, used to reflect the available space size; included angle parameter The angle between the average normal of the anchorage area and the direction of gravity; the closer the angle is to a right angle, the stronger the anchorage. The closer to zero;
[0055] Obstruction history parameters The non-positive real number obtained from the statistics of historical occlusion events of all basic units in the anchoring area can be obtained by taking the negative value of the number of occlusions or the negative value of the occlusion frequency. This can be achieved by counting the number of occluded frames of the base unit within the anchoring area within the time window and taking the opposite number, or by taking the opposite number after linear normalization, to ensure that areas with more occlusion are protected. The value is smaller, thus in The middle is naturally suppressed; historical parameters will be obscured. Defined as the negative of the occlusion frequency:
[0056]
[0057] Where: Number of times occlusion occurs : Anchoring area within the most recent time window The number of frames marked as occluded within the basic unit, a non-negative integer; total number of observed frames. The number of frames within the same window where the anchored area has been successfully observed by at least one sensor, a positive integer. Weighting coefficient. , , These are non-negative real numbers representing the adjustment area, posture, and degree of influence of occlusion factors, respectively. The specific values can be set during the deployment phase based on the size of the workshop space and the characteristics of the display terminal.
[0058] In practice, larger-scale anchoring region units are constructed on top of the basic units. Each anchoring region simultaneously considers factors such as area, attitude, and occlusion history to obtain a precisely quantified anchoring fitness coefficient. In subsequent anchor zone scheduling and multi-user deployment, the anchor fitness coefficient can be directly adjusted. By sorting and filtering, virtual content is prioritized to fall in areas with suitable geometry and less occlusion, thereby improving space utilization and the reliability of the layout.
[0059] After dividing the candidate surface into several anchoring area units, judging the quality of the anchoring area from the perspective of a single user often overlooks the differences in visibility and accessibility of the same anchoring area for multiple users in different positions and standing postures.
[0060] It is necessary to incorporate multi-user perspectives and multi-role task requirements into the characterization of anchor zone resources, so that each anchor zone not only has geometrical anchor adaptability, but also has a resource state description oriented towards the group, thereby better supporting multi-user collaboration in subsequent layout.
[0061] Based on the above collaborative situational map, the observation and operation trajectories of multiple users in each anchorage area during typical task phases are statistically analyzed. Based on each user's field of vision cone, head posture, and body reachable space, the visibility index and reachability index of each user for a certain anchorage area are calculated. The visibility indices and reachability indices of multiple users are then added together to obtain the comprehensive usability index of the anchorage area, reflecting the average usage value of the anchorage area under group collaboration.
[0062] The total number of collaborative users can be set to [number]. Numbered 1 to Number each anchoring zone Define a multi-user comprehensive availability index for
[0063]
[0064] Among them, the multi-user comprehensive availability index Number is The average availability of the anchored area unit from the perspective of all participating users; User contribution index : No. User name relative to anchor zone number The overall visibility and accessibility index can be obtained by weighting and adding the proportion of the user's line of sight coverage and hand accessibility to the anchor area during the critical phase of the task. The value range is usually limited to between 0 and 1.
[0065] Total number of users parameter : is a positive integer, equal to the number of users participating simultaneously in the current collaboration scenario.
[0066] Among them, the user contribution index The calculation can rely on historical mission playback data or on-site data collection. Stable estimates can be obtained by statistically analyzing the spatial relationship between users and anchorage areas under several typical working conditions.
[0067] In a multi-student training scenario, instructors and multiple students are positioned at different locations around a single piece of equipment. For an anchoring zone directly in front of the equipment, despite its anchoring fitness coefficient... The overall usability index is relatively high, but if only the teacher stands directly in front, most students will have difficulty seeing the area directly from the side, resulting in a low multi-user overall usability index. The size will be too low, and when the system is setting up the group layer, it may prefer to choose another anchor area that is slightly smaller but can be directly seen by most trainees.
[0068] This approach achieves a unified characterization of geometric fitness and multi-user accessibility from the perspective of anchor zones, enabling the resource status of anchor zones to reflect physical conditions, collaboration needs, and anchor fitness coefficients. After combining, it will be easier to proceed according to the following processes. and Based on the combination of different scenarios, each anchoring zone is scheduled in a hierarchical manner to provide resource guarantees for the spatial structure of the virtual-real integrated scenario.
[0069] Step Two: In a hybrid virtual-real scenario, on-site operators, remote experts, and trainees often focus on different information on the same device. Some information guides the rotation direction of specific knobs, some highlights the risks of the current step, and others provides an overview of the progress for observers.
[0070] If we continue with the traditional interface approach, treating all virtual content as parallel two-dimensional windows superimposed on the viewer's field of vision, key operators' view will be dominated by non-critical information, while bystanders will find it difficult to determine which stage the process is currently at. Therefore, it is necessary to construct a task-oriented, hierarchical spatial interaction component description structure. When describing each piece of virtual content, its task stage, importance, time urgency, and expected service recipients should be simultaneously provided, thus laying the foundation for subsequent spatial arrangement and weight allocation.
[0071] At this point, for each business process, the process is first broken down into consecutive steps, each containing several operational sub-tasks. For each step, key actions and key prompts are extracted from the operating procedures, and these are mapped to several candidate spatial interaction component prototypes. Each candidate spatial interaction component prototype is assigned a unique number, establishing a one-to-one correspondence between it and the specific step, the device location, and the role it plays. This allows for direct instantiation based on this descriptive structure when generating actual virtual components later, without needing to reinterpret the business semantics.
[0072] Furthermore, during setup, each spatial interaction component is configured with its component number, associated task step, task type, importance level, time urgency level, set of roles, and expected associated device area.
[0073] For example, during the disassembly of a high-pressure pump housing, a virtual numbered component used to guide the bolt removal sequence will be set as the core component of this step, with a high importance level, a medium time urgency level, and the roles it plays being the on-site operators and trainees. The expected associated equipment area is the bolts around the pump housing.
[0074] During runtime, when the process is detected to have progressed to this step, the fields of the component are read from the description structure, a spatial interaction component entity containing arrow tips and text descriptions is instantiated in the scene, and a reference relationship is established between this entity and the anchoring area unit formed in step one, in order to prepare for subsequent spatial anchoring and display control.
[0075] In one specific embodiment, an operator wears a head-mounted display terminal on-site, a trainee watches via a tablet device, and an expert participates in the review remotely via a large screen in a conference room. In the spatial interaction component description structure generated in this scenario, the same disassembly sequence prompt can be marked as core information for the operator, important information for the trainee, and reference information for the expert. Subsequent layout phases determine the presentation method of the component on different terminals and its priority at the individual, role, or group level, without requiring separate logic to be written for each terminal. Thus, each virtual prompt in the scenario carries task and role markers from the outset and establishes a stable correspondence with the actual device area, laying a unified descriptive foundation for multi-user spatial interaction.
[0076] When in use, all virtual content is categorized from the outset into structured, hierarchical spatial interaction components. Each component has clearly defined fields for task sequence, role division, and device association, thus providing clear input for subsequent role-based spatial hierarchy decisions and cross-terminal display method selection.
[0077] After obtaining a structured description of the hierarchical spatial interaction components, the textual fields of importance, time urgency, and risk association are transformed into numerical values that can participate in subsequent decision-making, and based on this, it is determined whether each component is more suitable for the individual level, role level, or group level.
[0078] Without a unified numerical characterization, different project teams are prone to inconsistent scales when configuring components. The same meaning of importance can represent completely different strengths in different processes, which can lead to deviations in the sorting and selection of components in the subsequent layout stage.
[0079] Furthermore, three main task attribute metrics are set for each component, corresponding to task importance, time urgency, and risk relevance, respectively.
[0080] Task importance focuses on whether the component directly affects the success or failure of the task; time urgency indicates the time sensitivity of the component's information to failure; and risk correlation indicates the strength of the association between the component and security or quality risks.
[0081] The three types of indicators are normalized and assigned a set of weights in a linear combination. Components with high priority coefficients will be placed more often in the group layer or important role layer to ensure that all relevant personnel can see them. Components with low priority coefficients will be placed more often in the individual layer, and explanations will only be provided to individual users when needed.
[0082] The number for the spatial interaction component is set as follows. Task importance index is The time urgency index is The risk correlation index is Then generate the component hierarchy priority coefficient. :
[0083]
[0084] In the formula, the component hierarchy priority coefficient Number is The priority of spatial interaction components in spatial hierarchy division;
[0085] Task Importance Index This component represents a normalized value that influences task success. It can be configured based on whether the task needs rework or key quality indicators are affected if this component is missing; the value is typically set between 0 and 1. Time urgency index. The sensitivity of this component's information to time is also set between 0 and 1; Risk correlation index. The correlation between this component and safety and quality risks is set between 0 and 1; the weighting coefficient is also set. , , All are non-negative real numbers, used to balance the influence ratio of the three types of indicators;
[0086] When used, this ensures a clear ordering of hierarchical interactive components in subsequent layout and display decisions. Component hierarchy priority. Once determined, it can be repeatedly applied in different scenarios and deployments, serving as input for subsequent spatial layout along with the anchoring fitness coefficient of the anchoring area in step one and the multi-user comprehensive availability.
[0087] Given that the hierarchical spatial interaction components have been established, if there is a lack of quantitative description of the collaborative relationships among multiple users, it will be difficult to determine in the subsequent spatial layout stage which users need to carry out collaborative operations around the same device area at the same time and which users are more in the role of observation.
[0088] It is necessary to utilize the gaze sequences and operation sequences of multiple users to extract the collaboration strength between users and between users and task objects, and use this as the basis for the edge weights in the collaboration situation diagram, so that the subsequent layout can be adjusted around the real collaboration structure, rather than just statically dividing based on role names.
[0089] For each user, the system continuously records their head orientation, eye gaze direction, and hand gestures, mapping these records to basic and anchoring units in the scene semantic-topology graph. For any two users, the system calculates the percentage of time they spend jointly looking at the same device area or virtual component within the same time window, and the percentage of time one user operates while the other looks at the same area within adjacent time windows. The former reflects the shared gaze relationship between the two users at the information level, while the latter reflects the behavioral relationship of guidance, observation, or handover. Regarding the relationship between a user and a task object, the system calculates the percentage of time the user spends looking at a particular object and the number of times they operate near that object in each task stage to determine the strength of collaboration between the user and the object.
[0090] Among them, it is possible to set a total of in the current collaborative scenario. User ID, numbered 1 to For any two users and Define the total number of video views index and frequency index of collaborative operations And construct the strength weight of the cooperative relationship. for:
[0091]
[0092] Among them, the strength weight of the cooperative relationship :user With users The overall collaboration intensity surrounding the task object in the current scenario; the total number of video views index. :user With users The proportion of time spent jointly viewing the same device area or virtual component within the same time period, with a value between 0 and 1; [Common Viewing Index] It can be defined as the proportion of frames in the window where two users' gazes fall on the same device area or component within the sliding time window.
[0093] Collaborative operation frequency index :user With users The proportion of events in which the same object is operated on sequentially or by one person while another observes within a similar time period can be obtained by statistically analyzing the temporal adjacency of hand action events and gaze events, with values ranging from 0 to 1.
[0094] Weighting coefficient and : A non-negative real number used to balance the weights of shared-vision behavior and collaborative behavior in the strength of the cooperative relationship. The proportion of.
[0095] When used, the originally difficult-to-quantify collaboration process is transformed into one based on the strength weight of the collaboration relationship. The depicted graph structure transforms collaborative dynamics from a simple judgment of the presence or absence of a relationship into a refined description of the strength and weakness of those relationships. This allows for subsequent placement of group and role-level components based on the weighted strength of collaborative relationships. By aggregating and displaying information among closely collaborating users around the same device area, collaboration efficiency can be improved.
[0096] The strength of the collaboration relationship alone is not enough to reflect the differences in each user's state at the current task stage. Some users may be highly stressed and make frequent mistakes, while others may be distracted and not perform effective operations for a long time.
[0097] It is necessary to combine observable measures such as the strength of collaborative relationships, task response time, number of errors, and duration of stagnation to construct cognitive load index and participation index for each user, so that the nodes in the collaborative situation map not only have structural positions but also state attributes, providing a basis for subsequent differentiated information delivery to different users.
[0098] During each task execution, record each user's response time to various prompts, the number of operations performed on physical devices and virtual components, the number of erroneous operations, and the duration of inactivity within a given period. Compare these records with the aforementioned collaboration relationship strength weights. Combined, we can determine whether each user is currently under heavy load or participating in a low-volume manner.
[0099] To keep the cognitive load index within a finite range and to compress extreme cases, the system constructs the cognitive load function in exponential form, ensuring that the output approaches the upper bound as the input value increases, thus preventing the exponent from growing infinitely in extreme cases. Specifically, for the input value numbered... For users, the average response time is The error count is The duration of the pause is And select non-negative weighting coefficients. , , Constructing a cognitive load index for:
[0100]
[0101] Among them, cognitive load index Number is The user's workload level at the current task stage, with a value between 0 and 1; average response time. The average time between the appearance of several key prompts and the commencement of the corresponding operation by the user can be automatically recorded by the system during task execution; it is a non-negative real number; erroneous operation count. : The number of erroneous operations triggered by this user during the current task phase, a non-negative integer; Stall duration. This represents the cumulative time during which the user has not performed any effective actions within the set time window; it is a non-negative real number. (Weighting coefficient) , , This is used to adjust the contribution of three factors—reaction time, misoperation, and pause duration—to cognitive load.
[0102] Because the exponential function tends to 1 when the input is large, therefore when , , At the same time, when it is relatively large, the cognitive load index Approaching the upper limit helps identify high-load users earlier.
[0103] When using it, each user node in the collaborative situation diagram should have a cognitive load index. and by , , The resulting behavioral characteristics add a state dimension beyond the structural relationships. This is combined with the aforementioned collaborative relationship strength weights. The collaborative situation map can describe who is working with whom and to what extent each person is occupied by tasks, providing rich decision-making basis for the subsequent layout and adjustment of hierarchical spatial interaction components, and laying the data foundation for adopting differentiated layout strategies for high-load users and low-participation users in step three.
[0104] Step 3: If there is a lack of a fine connection between the anchor area resource status obtained in Step 1 and the hierarchical spatial interaction component attributes obtained in Step 2, it is easy to fall into a coarse processing method of placing nearby or allocating according to number order. This kind of method cannot simultaneously consider the timing tracking confidence, geometric fitness, multi-user comprehensive availability of the anchor area, as well as the task urgency and risk level of the component itself.
[0105] Therefore, a candidate relationship and comprehensive scoring structure are established between the anchor area side and the component side, so that each component has a set of filtered candidate anchor areas before entering the subsequent layout solution, and the adaptability of each candidate anchor area is numerically characterized.
[0106] First, read the number for each anchoring zone from step one. Calculated anchoring fitness coefficient and multi-user comprehensive availability index Simultaneously, the time-series tracking confidence function within the corresponding spatial range of the anchoring area is read. A representative statistic for the current time period, such as the lower bound or weighted average within a small time window.
[0107] Then, number each component as described in step two. Calculated component hierarchy priority coefficient In addition to the component's own anchoring requirements (such as the lowest anchorable level and the requirement to avoid obstructing safety signs), a set of anchoring areas that meets the hard constraints is selected from all anchoring areas as the anchoring area set for that component. A comprehensive score is calculated for each anchoring area in the set for scheduling options.
[0108] Among them, the number can be _____. Spatial interaction components and numbered Anchorage zone construction comprehensive scoring function Its form is as follows:
[0109]
[0110] In the formula, the comprehensive scoring function Spatial interaction components Anchored in the anchorage area The suitability of this combination is determined by comprehensively considering three dimensions: task priority, geometric adaptability, and multi-user overall availability; component-level priority coefficient. The result given in step two reflects the component's importance, time urgency, and risk relevance in the task chain, with a value between 0 and 1; anchor fitness coefficient. Step 1 defines the anchoring area. The degree of suitability in terms of area, orientation, and occlusion history; a non-negative real number; a multi-user comprehensive availability index. Also originating from step one, reflecting the anchoring area Average visibility and accessibility from the perspective of all participating users, ranging from 0 to 1; weighting coefficient. , , : A non-negative real number used to adjust task factors, geometric factors, and multi-user factors in the comprehensive scoring function during specific deployments. The proportion of.
[0111] When in use, a candidate relationship with a score is formed between each component and several anchor areas. This can prevent components from falling into positions that are unstable or difficult to observe, and also provides a clear numerical basis for the subsequent anchor area scheduling process. This is beneficial to ensure task priority while taking into account the observation and operation needs of multiple users.
[0112] Once each component has a set of candidate anchor areas and a comprehensive score, if the upper limit of anchor area occupancy is not considered and the anchor areas are allocated only according to the score from high to low, it is easy for a small number of anchor areas to be occupied by a large number of components, while the remaining anchor areas are idle for a long time. This is also not conducive to visual neatness and reduces the adjustable space for subsequent layout.
[0113] A balance needs to be struck between the overall scoring on the component side and the resource occupancy constraints on the anchoring area side. An initial anchoring area allocation scheme is generated through a deterministic allocation traversal process, and hierarchical adjustment suggestions are automatically generated in resource-scarce areas, providing a starting point that meets physical constraints for subsequent two-level layout solutions.
[0114] First, prioritize according to component hierarchy. All components are sorted from highest to lowest, with group-level and key role-level components ranked first. Then, for each component in the sorted sequence, its candidate anchor regions are evaluated according to a comprehensive scoring function. Arrange them from high to low, and check the resource usage and capacity limit of the corresponding anchor area in the current time period.
[0115] If a certain anchor area has surplus in terms of area, number of allowed components, and margin for occupancy risk, the component is temporarily assigned to that anchor area, and the occupancy status of that anchor area is updated. If all candidate anchor areas cannot receive the component under resource constraints, the system records a compression suggestion, indicating that the component needs to be considered for reducing its spatial hierarchy or using a time-switching method in subsequent steps.
[0116] Specifically, this can be achieved by iterating through each row, starting with the first component of the sorted sequence and matching the anchor regions one by one. Taking the aforementioned engine assembly training example, the safety-related group-layer components are first processed and assigned to the comprehensive scoring function. A high-level and resource-rich anchorage zone.
[0117] For example, a large wall area built around the engine, due to the anchoring fitness coefficient and multi-user comprehensive availability index Since the height is relatively high, priority will be given to supporting the global risk warning panel. As the usable area of this wall region gradually decreases, subsequent component processing will switch to using the side anchoring area around the engine. Furthermore, regarding the hierarchical priority coefficient... For lower-level personal components, if all candidate anchor areas cannot meet the resource consumption limit, the system will record in the component description that this level adjustment information will only be displayed locally on the personal terminal, and hand it over to the subsequent layout solving stage for processing.
[0118] In practice, without relying on complex mathematical solvers, an initial allocation scheme that balances task priority and anchor area resource constraints is constructed using sorting and sequential traversal engineering methods. Furthermore, hierarchical adjustment suggestions are pre-generated for resource-constrained components. This eliminates the need to deal with a completely blank allocation state for subsequent spatial layout solutions, improving the controllability of the scheme's implementation.
[0119] Furthermore, even after obtaining the initial anchoring zone allocation scheme and hierarchical adjustment suggestions, it is still impossible to guarantee the geometric relationship consistency of group layer and role layer components from the perspective of multiple users if we only consider whether each component has been anchored.
[0120] For example, even if a key tooltip is anchored in a highly adaptive area, if the tooltip is positioned completely differently in the field of vision for different users, pointing at the tooltip during collaboration can easily result in inaccuracies. It is necessary to incorporate user field-of-view geometry and collaboration relationship information at the anchoring area level, adjusting the precise positions and orientations of group and role-level components within the anchoring area to ensure that key components form a relatively stable and consistent geometric relationship from the perspective of most collaborating users.
[0121] Based on this, for each user, according to the scene semantic-topology map constructed in step one and the terminal posture, the user's line of sight incident direction and field of view coverage area for each anchorage area under the current posture are calculated. For group layer or key role layer components that have been assigned to a certain anchorage area, available layout sub-regions are divided on the surface of the anchorage area, and angle analysis is performed on different sub-regions to calculate the observation deflection angle, offset distance from the center of vision, and occlusion probability of each user in these sub-regions. At the same time, the perspective differences between users with close cooperative relationships in these sub-regions are estimated.
[0122] Refer to the strength weight of the cooperation relationship in step two. Prioritize users with close collaborative relationships to minimize their differing perspectives on key components.
[0123] Specifically, a view cost term can be constructed for each user and each candidate layout sub-region. The viewing angle, the distance to the edge of the field of view, and the probability of occlusion are aggregated into a geometric comfort index for a single user view according to a fixed weight. Then, the sub-region with the smaller aggregate cost is selected as the final position of the key component within the anchor area.
[0124] For example, in an engine training embodiment, the instructor and two lead trainees typically stand on the same side of the engine. When the system anchors a group-level overview panel on the wall above the engine, it prioritizes positioning the panel within the overlapping field of vision of these users. This ensures that the panel, when all three look up simultaneously, is positioned close to the center of the wall, rather than biased towards one person's side. For trainees observing from the other side, the system can later compensate for the difficulty in observing the wall panel from that perspective by replicating a simplified overview bar within their field of vision during the personalized layout phase.
[0125] When using it, add the user's view geometry and collaboration relationship strength weights to the anchor area. This allows the positions of group layer and role layer components to maintain a relatively fixed geometric relationship with the majority of collaborating users, while satisfying the resource constraints of the anchoring area, thereby reducing the inaccuracy of collaboration instructions and verbal communication.
[0126] After the group and role-level components are laid out within their anchored areas, if the differences in cognitive load among individual users and the cumulative effect of individual-level components are ignored, it's easy for some users to experience overload in their field of vision, while others experience insufficient information. This can be addressed by utilizing the cognitive load index constructed for each user in step two. By combining the number and importance of components in the user's personal layer and role layer, a personal information load index is constructed. Without disrupting the overall structure of the group layer, the layout of each user's personal layer is fine-tuned to keep the information volume and cognitive load in harmony.
[0127] Within the current task phase, for each user, the system calculates the set of components that truly require their attention within their field of vision. Components belonging to the group and role layers are considered the basic responsibility set, while personal-layer components accessible only to that user are considered the incremental set. For each component, the system directly reads its component hierarchy priority coefficient. This is considered the intensity of the component's occupation of the user's attention. A personal information load index is constructed for each user, summing the component hierarchy priority coefficients of all components in the user's current field of vision to obtain an information load characterization related to task importance. This characterization is then compared with the cognitive load index. These are combined to form a comprehensive indicator, which guides the fine-tuning of the position or display control of individual-level components.
[0128] Among them, it is possible to set the number as For users, their personal information load index is recorded as: The set of components that exist in an individual's field of vision is Then there is
[0129]
[0130] Among them, the personal information load index Number is The sum of the importance of all components in the user's current field of view; component collection : The set of component IDs currently displayed in the user's field of view; component hierarchy priority. The comprehensive value of the spatial interaction components defined in Step 2 in terms of task importance, time urgency, and risk relevance.
[0131] In obtaining personal information load index Afterwards, the system can construct a comprehensive load index. For example, let Among them, the comprehensive load index Indicates the number is In the current layout state, the user's overall burden level, considering both cognitive load and information load, is adjusted by the coefficient. It is a non-negative real number used to control the weight of information load in the comprehensive load index.
[0132] When a user's comprehensive load index When the pre-set safety threshold is exceeded, the system will proceed according to the component hierarchy priority coefficient. From low to high priority, select several individual or secondary role-level components and move them to the edge of the field of view or use a collapsed display method to reduce the user's current perceptual burden; for the overall load index For users with lower engagement levels, the system can appropriately introduce some additional personalized tips or practice guidance to increase their participation.
[0133] In the engine training implementation example, when the system detects that a trainee operator has difficulty recognizing the load index in several key steps... It is already close to the upper limit, and the personal information load index When the cognitive load is also at a high level, some of the more explanatory personal text descriptions will be temporarily collapsed into icons, and will only be expanded when the student pauses; while for the students standing in the back row who are observing, due to their cognitive load index Typically low, personal information load index It is also relatively small, and the system can add some structural diagrams or principle explanation components to its field of view, thereby improving the overall teaching effect without interfering with the main operator.
[0134] When applying, prioritize the component hierarchy. Cognitive Load Index and personal information load index By combining these approaches, a series of fine-tuned layouts are arranged at the individual level, with the amount of information corresponding to the individual's status. This ensures that users with high workloads are not disturbed while providing more opportunities for learning and exploration for users with low workloads. At the same time, a layer of delicate personalized adjustments is formed at the group level.
[0135] Step 4 and Step 3 result in a multi-user spatial layout based on the anchor fitness coefficient, multi-user comprehensive availability index and component hierarchy priority coefficient within a certain moment or short time window. However, the tracking quality and environmental occlusion in the virtual-real combined scenario often change slowly or abruptly over time.
[0136] Without a risk measure that comprehensively reflects the confidence function of time-series tracking and the resource occupancy of anchor areas at the anchor area level, it is difficult to promptly detect whether an anchor area is no longer suitable as a carrier location for key components in the future, which may cause jitter, drift, or frequent disappearance of group-level components.
[0137] Therefore, for each anchor region in the layout, the time-tracking confidence function will be applied. Used cyclically, numbered for each anchoring zone. Read the basic units contained therein, and use the time-tracking confidence function of these basic units in the current time window. By performing weighted integration or weighted summation, the risk of the anchor zone is mainly characterized by a decrease in confidence level.
[0138] Based on the number of spatial interaction components currently contained in the anchoring area and the component hierarchy priority coefficient of that component. A weighted term for the density of key components is constructed, and the weighted terms are superimposed and combined to obtain a cross-layer risk index: representing the risk level of the anchoring area at the current moment under the combined effects of spatial layout, tracking quality and mission importance;
[0139] Anchorage zone numbering can be set. The spatial region is A weighting function is defined to represent the contribution of the basic unit to the anchorage area, and an anchorage risk index is constructed at time parameter t. :
[0140]
[0141] In the formula, the anchorage risk index Number is The anchoring zone in time parameters The cumulative risk caused by the time-series tracking confidence function at any given time; time-series tracking confidence function The basic unit-level confidence level defined in step one ranges from 0 to 1. Approaching a certain point indicates that the tracking at that position is stable. A value close to zero indicates that position tracking is unreliable;
[0142] Integration area Number is The three-dimensional spatial region corresponding to the anchorage region; weighting function A non-negative function that assigns different weights to basic units located at different positions within the anchoring zone; A constant equal to the area of the basic unit can be used, or greater weight can be assigned to basic units near critical parts of the equipment to amplify the risk in critical areas.
[0143] If the anchorage region is divided into discrete basic units, the above integral can be approximated by a finite summation form, and the weighting function... This is reflected by the area or volume of the basic unit.
[0144] In practice, anchoring area Composed of a finite number of basic units, the above integral can be achieved by summation:
[0145]
[0146] Where: the set of basic units Anchorage area It contains the complete set of basic cell indexes; cell center location. Basic Unit Center coordinates in the workspace coordinate system; weight : You can take the area or volume of the basic unit, or you can take a larger value from the unit closer to the key part of the equipment.
[0147] When using it, the time-series tracking confidence function given in step one will be used. Unified as a risk index This allows the system to continuously track changes in tracking stability of each anchoring area throughout the entire layout process, facilitating cross-layer layout adjustments and avoiding localized decisions in a single frame.
[0148] In addition, the anchorage risk index After the construction is completed, if the anchorage risk index is not combined with the user-side comprehensive load index... If all components are migrated or downgraded based solely on thresholds, it will inevitably lead to excessive adjustments to critical components by heavy-load users or even disruption of key processes.
[0149] During the layout adjustment process, both spatial risks and user load need to be considered simultaneously. The adjustment intensity should reflect both the spatial instability of the anchor area and the responsibility of the components within the anchor area to various types of users. This allows for prioritizing the adjustment of low-priority components or components in the view of low-load users when necessary.
[0150] For each anchorage zone number Read the set of components anchored to the anchored area in the current layout and their component hierarchy priority coefficients. At the same time, combined with the comprehensive load index established for each user in step three. Construct a layout adjustment strength metric for the relationship between components, anchor areas, and users. This layout adjustment strength metric is used to determine which components need to be moved out of the current anchor area, whether to trigger hierarchy degradation, or simply move within individual user views.
[0151] Among them, it is possible to set the target number as The spatial interaction component is anchored at the numbered The anchorage area, and pay special attention to the numbered When users are constructing layouts, adjust the strength. for:
[0152]
[0153] Layout adjustment intensity In time parameters At any given moment, for components In the anchoring area A comprehensive measure of whether the layout needs adjustment and the extent of such adjustment; anchoring zone risk index. : A non-negative quantity constructed to reflect changes in the tracking stability of the anchoring zone; comprehensive load index Step three incorporates the cognitive load index. and personal information load index Constructed user-side load characterization; weighting coefficients and A non-negative real number, used to balance spatial risks and user load in the intensity of layout adjustments. The proportion of contribution.
[0154] You can set a layout adjustment threshold. ,when When, the component From the anchorage area Migrate to a lower-risk candidate anchoring area, or perform hierarchical degradation or simplified display on the component.
[0155] In practice, the risk index of a certain anchorage area near the channel. The load increases due to frequent occlusion, and the overview panel on this anchored area is a group-layer component, crucial for all students and teachers. Calculate the overall load index for the teacher node. It was found that the teacher's performance was still within a manageable range, so the intensity of the layout adjustment for this component in the teacher's view was adjusted. The risk index does not increase significantly; the system simply copies some explanatory content from the student's view onto another anchoring area on the wall and simplifies its display on the original anchoring area to mitigate the impact of passageway obstruction. The overall workload index of many trainees continues to rise. When the priority level also increases, the system begins to prioritize components according to their hierarchical coefficients. Gradually move less critical group-level components down to the role or individual level, retaining only the most essential overall process overview.
[0156] When using it, in the cross-layer risk index and comprehensive load index Based on the structural layout, adjust the intensity The document uses specific examples to illustrate how to take adjustments such as migration, downgrading, or simplification of displays under the dual constraints of rising spatial risks and changing user load, so as to improve the overall stability of the scene without interrupting critical tasks.
[0157] Once the aforementioned cross-layer layout adjustment mechanism is established, if all trigger thresholds, weight coefficients, and migration rules are fixed, the system's performance in different projects and different scenarios will heavily depend on the initial configuration. It will be impossible to change the strategy based on long-term accumulated task completion data and user feedback. The series of weight coefficients, thresholds, and hierarchical decision parameters required for the aforementioned layout adjustment process are abstracted into a strategy parameter vector. Furthermore, by aggregating performance indicators from multiple task processes, an objective function is constructed, and the strategy parameter vector is updated with gradients based on the objective function. Thus, while maintaining the overall decision structure, parameters that better reflect the actual situation are gradually constructed.
[0158] After each task process, a set of performance metrics is recorded, including but not limited to whether key steps were completed on the first attempt, number of rework attempts, total task duration, number of serious misoperation events, and quantitative results of user subjective evaluation. Furthermore, these performance metrics are correlated with the strategy parameter vector used in this process, accumulating performance data from multiple processes into a historical objective function. For this objective function, the system uses numerical differentiation or automatic differentiation tools to calculate the gradient approximation with respect to the strategy parameter vector, determining the direction and magnitude of adjustments for each parameter without disrupting the original structure.
[0159] Among them, the strategy parameter vector can be set as This includes the layout adjustment intensity weighting coefficient. , Multiple components, including risk threshold, comprehensive load threshold, and component-level priority coefficient combination weight;
[0160] Define the objective function This is used to characterize the overall performance of multiple task flows under a given policy parameter vector, such as a weighted sum of rework counts, serious error events, and timeouts. After each round of data accumulation, the policy parameter vector is updated as follows:
[0161]
[0162] Among them, the new round of policy parameter vector : Represents the set of strategy parameters after a single parameter adjustment, which will be used in subsequent task flows; the current strategy parameter vector. This is the set of parameters used before the execution of this round of tasks;
[0163] Learning step size coefficient : A positive real number whose value determines the magnitude of parameter adjustment during each update. Excessive size may cause fluctuations. If it is too small, the convergence will be slow;
[0164] objective function The larger the value of the task performance characterization function constructed above, the worse the task performance under the current policy parameter configuration; the long-term performance objective function within the task cycle can be defined as follows:
[0165]
[0166] Where: average number of rework cycles The average number of rework attempts per task across multiple tasks; the average number of serious errors. The average count of serious errors across multiple tasks; the average timeout duration exceeding... : Average duration exceeding the expected completion time; weight A non-negative real number used to balance the importance of different types of costs. Gradient It can be approximated by finite difference: for Each component is subjected to a small perturbation, and the comparison is performed. The changes before and after are used to obtain an approximate gradient value, and then the parameters are adjusted according to the above update formula.
[0167] gradient vector The combination of partial derivatives of the objective function with respect to the policy parameter vector can be approximated using finite difference or automatic differentiation, and is used to determine the direction in which each parameter should be adjusted. Common solvers such as quasi-Newton methods or conjugate gradient methods can be used, and this update rule can be embedded in the system background for periodic execution.
[0168] During the implementation of the assembly training line, a relatively high risk threshold and a relatively low overall load threshold were set based on experience during the initial deployment phase. As multiple training tasks were completed, the number of rework attempts and serious misoperations in each batch were gradually obtained and incorporated into the objective function. In the middle. Objective function When the gradient vector is higher in certain batches, A portion of the components points towards lowering the risk threshold and raising the overall load threshold, thereby adjusting the strategy parameter vector through rule updates. In subsequent training sessions, more training tasks will prioritize relocating elements from high-risk anchor areas earlier and appropriately lowering certain user load thresholds to achieve a more balanced strategy configuration for both experimental and production environments.
[0169] When using it, the strategy parameter vector should be adjusted without changing the overall decision-making framework. and objective function The gradient update utilizes long-term accumulated scenario performance data to coordinate and control multiple parameters of the system's layout and threshold adjustment, forming an experience-based strategy configuration for specific projects.
[0170] However, simply adjusting the weight coefficients in the strategy parameter vector is insufficient to segment and track the cognitive load limits of different groups and roles. The cognitive load index varies among individuals from different companies, positions, and training stages. and comprehensive load index The typical distribution of cognitive load varies significantly. If a uniform fixed threshold is still used, it is easy for some users to remain close to the threshold for extended periods, while other users remain far below the threshold most of the time. This leads to inconsistent sensitivity of layout adjustment decisions to different user groups. Therefore, it is necessary to construct user-oriented cognitive threshold parameters and apply them based on the cognitive load index. With comprehensive load index The historical trajectory is used to self-tune the threshold parameter, so that the system can form differentiated threshold configurations for different user groups while maintaining structural uniformity.
[0171] The system assigns a number to each user. Set a cognitive threshold parameter Initially, tasks can be assigned based on job category or training stage and experience. After each round of tasks is completed, the system calculates the user's cognitive load index for the current task stage. and comprehensive load index The time average or representative value is used to compare the relationship between the two. If it is found that the user's cognitive load index is high for most of the time... Far below the comprehensive load index This indicates that the current threshold setting is too conservative, and the threshold parameter can be appropriately increased. This is to prevent the system from adjusting the layout for the user too frequently; conversely, if the user's cognitive load index... For a long period of time, the load index is close to or even exceeds the comprehensive load index. Then the threshold parameter needs to be gradually reduced. This allows the system to trigger information simplification and layout migration earlier in subsequent tasks.
[0172] Each user can be assigned a unique ID. The cognitive threshold parameter is updated in the following form:
[0173]
[0174] Among them, the updated cognitive threshold parameter After the current task is completed, the task number will be... The new cognitive threshold set by the user will be used in the next round of tasks to determine whether the user needs to trigger view simplification and layout adjustment;
[0175] Current cognitive threshold parameter Threshold settings for the user before task execution; Cognitive load index The load profiles constructed in steps two and three based on reaction time, error count, and stagnation duration have values between 0 and 1; the comprehensive load index... Step 3: Combining cognitive load indices and personal information load index The obtained overall burden value also varies around 0 to 1; adjustment coefficient : A real number less than 1, used to control the magnitude of each threshold adjustment.
[0176] In an assembly line embodiment that includes both new trainees and experienced operators, different initial cognitive threshold parameters are assigned to the new trainees and the experienced operators respectively. After several training sessions, the system identified the cognitive load index of experienced operators. Significantly lower than the comprehensive load index for a long period of time This indicates that this type of user still has additional information storage space. According to the above rules, the cognitive threshold parameter... The cognitive load was gradually increased, reducing the number of times the system triggered view simplification in subsequent tasks, allowing more auxiliary information and complex components to be presented in the user's field of vision; at the same time, the system detected the cognitive load index of some new learners. Frequently exceeding the comprehensive load index The corresponding cognitive threshold parameter After several consecutive task cycles with slight reductions, the system pre-compacts the display area of non-critical components in the trainees' views, delaying the appearance of certain complex explanations, allowing them to focus on completing key steps. In actual observation, teachers can see that the virtual panel in front of new trainees is relatively simple, while experienced operators see more information. This differentiated effect is a direct manifestation of the self-tuning of the cognitive threshold parameter.
[0177] When in use, the thresholds for different user groups can be automatically adjusted over several task cycles, enabling different levels of information control for users with different roles and skill levels while maintaining the overall layout decision-making logic, and integrating with the aforementioned strategy parameter vector. The combination of gradient updates enables step four to simultaneously possess the ability to learn and adapt from two levels: cross-layer layout adjustment and long-term strategy evolution.
[0178] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0179] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0180] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0181] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0182] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A virtual-real integrated scene space interaction adaptive online processing method, characterized by: include, By collecting images and depth data of the virtual and real scene through multiple sensors, 3D reconstruction and semantic segmentation are performed to construct a scene semantic topology map containing adjacency, visibility and access relationships. The scene is then divided into anchoring area units with resource status according to the confidence level and anchorability level tracked in time. The virtual content is abstracted into hierarchical spatial interaction components that carry steps, interaction types and spatial levels. Anchoring requirements are set for the components, and a collaborative situation diagram with users, task objects and components as nodes is constructed based on user posture, field of vision and line of sight and operation trajectory. Based on the resource status of the anchoring area unit and the anchoring requirements of the hierarchical spatial interaction components, the anchoring area scheduling is performed according to the adaptability. High-confidence anchoring areas are allocated to group layer and role layer components first, and group layer layout and individual layer fine-tuning are carried out in combination with multi-user perspectives within the anchoring area. During operation, the tracking confidence level of the anchor zone unit and the user cognitive load in the collaboration status map are monitored. When the confidence level or cognitive load exceeds the threshold, the position and level of the deployed components are adjusted across layers, and the strategy parameters are updated according to the task completion status and user intervention records.
2. The adaptive online processing method for virtual-real combined scene space interaction according to claim 1, characterized in that: The multi-sensor system includes a color video camera, an infrared depth camera, an inertial measurement unit, and a laser ranging device. It uses simultaneous localization and mapping (SMR) algorithms to process environmental data and generate a 3D scene model that includes a 3D point cloud or mesh model and the attitude trajectory of the terminal device. It also establishes a scene semantic topology map based on the geometric proximity relationship, unobstructed line-of-sight relationship, and passability relationship between semantic objects.
3. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 2, characterized in that: The anchoring area resource status information includes the time-series tracking confidence sequence of the anchoring area unit within a preset time window, the visibility markers for different user viewpoints, the accessibility markers for different user bodies and hands, and the number and type information of the hierarchical spatial interaction components attached to the anchoring area unit. At the end of step one, the anchoring area resource status information is associated and stored with the scene semantic topology map.
4. The adaptive online processing method for virtual-real combined scene space interaction according to claim 3, characterized in that: Each hierarchical spatial interaction component carries a step identifier, a type identifier, and an interaction type identifier, and records the semantic object or semantic region identifier to be associated. Anchoring requirements include the acceptable lower limit of temporal tracking confidence for the hierarchical spatial interaction component, the allowed range of anchoring zone levels, and the target field of view orientation and field of view depth range.
5. The adaptive online processing method for virtual-real combined scene space interaction according to claim 4, characterized in that: For each user, maintain head posture, field of vision, and reachable space. Associate the currently gazed semantic object and the hierarchical spatial interaction component being operated with the user's role information. When constructing the collaborative situation map, use the user, semantic object, and hierarchical spatial interaction component as nodes. Mark the edges of co-viewing relationship, operation relationship, and guidance relationship according to the gaze and operation trajectory, and record the cognitive load and participation on the nodes.
6. The adaptive online processing method for virtual-real combined scene space interaction according to claim 5, characterized in that: When scheduling anchor areas for hierarchical spatial interaction components, the components are divided into group-related components and role-related components according to hierarchical attributes. The anchor area resource status is queried, and anchor area units with higher time-series tracking confidence and visibility are preferentially allocated to group-related components, and then allocated to role-related components. When there are insufficient available anchor area units, multiple components are displayed in rotation on the same anchor area unit according to a preset time slice.
7. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 6, characterized in that: When fine-tuning the position and size of components within each user's field of vision, the position and size of the hierarchical spatial interactive components assigned to that user are adjusted based on each user's field of vision and accessible space. Keep the anchor area scheduling result unchanged, so that components do not obscure semantic objects with occlusion importance weights, and limit the user's head rotation and arm extension to a preset range.
8. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 7, characterized in that: While monitoring the confidence level of time-series tracking in the anchoring area and the cognitive load of users, the system also monitors processing power and network conditions. When the processing power or network conditions are lower than the preset threshold, the presentation of the hierarchical spatial interactive components anchored in the corresponding anchoring area is adjusted. Components presented as 3D models are simplified into icons or text labels, and components displayed in the group layer are adjusted to be displayed in the role layer or individual layer.
9. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 8, characterized in that: The layout and hierarchy control strategy parameters include the cognitive load threshold of each role, the resource priority weight of various anchoring area units, the cost weight of occlusion, view offset and anchoring area occupation in the layout evaluation, and the sensitivity parameters for triggering the hierarchical adjustment of the hierarchical spatial interaction components between the group layer, role layer and individual layer. The above parameters are corrected when the strategy parameters are updated.
10. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 9, characterized in that: The strategy parameter update is based on the recorded task performance data, which includes completion time, number of misoperations and redo times, and the number of times the user adjusts the layout. After the task is completed, the task performance data is summarized, and the cognitive load threshold, resource priority weight, and cost weight of each role are corrected. The corrected parameters are then written back to the layout and hierarchy control strategy parameters.
Citation Information
Patent Citations
Indoor unfamiliar scene recognition system fusing knowledge graph and spatial semantic topological graph
CN114972938A
Autonomous lifelong SLAM method and system based on visual language model hidden space representation
CN120599495A
Information extraction processing method and system applied to data sharing service
CN120875010A
Clinical operation data analysis method based on virtual reality
CN121354088A
Scene flow digital twin method and system based on dynamic trajectory flow
WO2023207437A1