Mixed reality scene space interaction adaptive online processing method
By constructing scene semantic topology maps and collaborative situation maps, and combining multi-sensor data processing, the spatial layout of virtual content is dynamically adjusted, solving the problem of stability and consistency of virtual content in multi-user virtual-real integrated scenarios, and improving assembly and training efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to coordinate the spatial layout of numerous virtual markers, prompts, and interface components in a shared virtual-real scenario involving multiple users and roles. This leads to issues such as inconsistent virtual content placement for different users, obstruction of real devices, and unstable tracking, impacting assembly efficiency and training effectiveness.
By constructing a scene semantic topology map and generating a collaborative situation map, component anchoring area scheduling is performed based on user posture, field of vision and operation trajectory to achieve stable presentation and personalized fine-tuning of virtual content. Combined with changes in the multi-sensor data processing environment, the position and level of components are dynamically adjusted.
Maintaining stable display of virtual content in multi-user environments reduces occlusion and drift, ensuring consistency of perspective and ease of operation for different users, and improving assembly efficiency and training effectiveness.
Smart Images

Figure CN121564296B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, in particular to a virtual-real combined scene space interaction adaptive online processing method. BACKGROUND
[0002] In virtual-real combined scenes such as remote guidance in manufacturing sites, complex equipment installation and adjustment training, multi-person safety inspection, etc., personnel wearing head-mounted display terminals or mobile terminals superimpose virtual markers, step prompts, danger area reminders, virtual tool panels in real workshops, machine rooms or experimental areas. As the scale of the production line and the number of training personnel increase, a real space often has multiple users, multiple roles, and multiple types of virtual content at the same time, which requires constantly adding virtual content in limited field of view and limited available space, and meeting the viewing and operation needs of different users. In current engineering practice, more and more enterprises apply to assembly, cross-regional expert remote guidance and classroom training, etc., and the requirement for adaptive space layout is increasing. Existing virtual-real combined systems generally use synchronous positioning and mapping algorithms and three-dimensional reconstruction techniques to model the site environment, and have automatic placement functions for virtual content to a certain extent, such as attaching prompt information to equipment or fixing interface panels in front of the user. Some systems also adaptively adjust display clarity, refresh frequency and other parameters according to environmental lighting or tracking quality or network status.
[0003] Such methods can reduce the difficulty of manually placing interfaces and repeatedly adjusting positions when there is a single user, few virtual contents, and slow environmental changes. However, in real scenarios with multiple users and multiple roles working simultaneously, and with a large number of virtual content types and quantities, this method can still only adopt a single perspective-based local rule layout, lacking a unified modeling and scheduling mechanism from the perspective of multiple users, multiple roles and collaboration.
[0004] During the use of a virtual-real combined scene shared by multiple people, firstly, the positions and relative relationships of virtual contents corresponding to different users are inconsistent, making it difficult for collaborative personnel to uniformly refer to the same virtual marker or prompt, which can easily cause instruction understanding errors. Secondly, virtual prompts or virtual panels block important positions of real equipment, safety signs or actual operation areas, making on-site personnel switch attention between virtual content and real targets. Thirdly, when equipment shakes, the environment is blocked, or the network fluctuates, the tracking stability and computing resources of some areas change greatly, and if virtual content is still fixed in this area, it can easily cause picture shaking, virtual marker drifting, or even interface disappearance, interrupting operation. The above problems will be amplified in long-term, multi-batch and multi-person scenarios, affecting assembly efficiency and training effectiveness, and interfering with the communication of safety-related information.
[0005] Therefore, the technical problem to be solved at present is:
[0006] In the case that multiple users and multiple roles share the same virtual-real combined scene and the environmental stability and computing resources change over time, the prior art is difficult to coordinate the spatial layout of a large number of virtual markers, prompts and interface components under a unified model, difficult to avoid unstable tracking and resource-intensive areas in a timely manner, and difficult to simultaneously take into account the consistency of different user perspectives, the unobstructedness of real key areas and the controllability of group and individual cognitive load. SUMMARY
[0007] (I) Technical problems solved
[0008] In view of the deficiencies of the prior art, the virtual-real combined scene spatial interaction adaptive online processing method provided by the present application generates a cooperation situation diagram by combining user posture, field of view and line of sight and operation trajectory; on this basis, anchor area scheduling is performed according to anchor area resource state and component anchor demand, global layout of group layer and role layer components is completed first, and then individualized fine tuning of components within the field of view of each user is performed; during operation, cross-layer adjustment is performed according to anchor area tracking confidence and user cognitive load, and strategy parameters are updated online, realizing stable presentation of virtual content layout in a multi-user virtual-real combined scene, and solving the technical problems described in the background art.
[0009] (II) Technical solutions
[0010] To achieve the above object, the present application is implemented by the following technical solutions:
[0011] The virtual-real combined scene spatial interaction adaptive online processing method comprises: collecting image and depth data of the virtual-real combined scene through multiple sensors, three-dimensional reconstruction and semantic segmentation, constructing a scene semantic topology graph containing adjacency, visibility and passable relationship, and dividing the scene into anchor area units with resource states according to time sequence tracking confidence and anchorable level;
[0012] Virtual content is abstracted into hierarchical spatial interaction components carrying steps, interaction types and spatial levels, anchor demand is set for the components, and a cooperation situation diagram with users, task objects and components as nodes is constructed based on user posture, field of view and line of sight and operation trajectory;
[0013] According to the resource state of the anchor area unit and the anchor demand of the hierarchical spatial interaction component, anchor area scheduling is performed according to the adaptation degree, high-confidence anchor areas are preferentially allocated to group layer and role layer components, and group layer layout and individual layer fine tuning are performed in the anchor area in combination with the fields of view of multiple users;
[0014] During operation, the tracking confidence of the anchor area unit and the user cognitive load in the cooperation situation diagram are monitored, when the confidence or cognitive load exceeds the threshold, cross-layer adjustment is performed on the position and level of the arranged components, and strategy parameters are updated according to the task completion condition and user intervention record.
[0015] Further, the multi-sensor includes a color video camera, an infrared depth camera, an inertial measurement unit and a laser ranging device, the environment data is processed by a simultaneous localization and mapping algorithm to generate a scene three-dimensional model containing a three-dimensional point cloud or a grid model and a terminal device pose trajectory, and a scene semantic topology graph is established based on geometric proximity relationships, unobstructed line-of-sight relationships and passable relationships between semantic objects.
[0016] Further, the anchor area resource state information includes a time sequence tracking confidence sequence for the anchor area unit within a preset time window, visibility marks for different user viewpoints, accessibility marks for different user bodies and hands, and quantity and type information of the layered space interaction components connected to the anchor area unit, and the anchor area resource state information is stored in association with the scene semantic topology graph at the end of step one.
[0017] Further, each layered space interaction component carries a step identifier, a type identifier and an interaction type identifier, and records an expected associated semantic object or semantic region identifier, and the anchoring requirement includes a lower limit of the time sequence tracking confidence acceptable by the layered space interaction component, a permitted anchor area level range, and a target field of view direction and field of view depth range.
[0018] Further, the head pose, field of view range and reachable space are maintained for each user, the semantic object currently gazed at and the layered space interaction component currently operated are associated with the role information of the user, and in constructing the collaboration situation graph, the user, the semantic object and the layered space interaction component are taken as nodes, the edges of the co-visibility relationship, the operation relationship and the guidance relationship are marked according to the line of sight and the operation trajectory, and the cognitive load and the participation degree are recorded on the nodes.
[0019] Further, when scheduling the layered space interaction components in the anchor area, the components are divided into group-related components and role-related components according to the hierarchical attribute, the anchor area resource state is queried, the anchor area units with higher time sequence tracking confidence and visibility are preferentially allocated to the group-related components, and then allocated to the role-related components, and when the available anchor area units are insufficient, multiple components are displayed on the same anchor area unit in a preset time slice rotation manner.
[0020] Further, when adjusting the position and size of the components in the field of view of each user, the position and size of the layered space interaction components allocated to the user are adjusted according to the field of view range and the reachable space of each user.
[0021] The anchor area scheduling result is kept unchanged, the components do not block the semantic objects with blocking importance weights, and the user head rotation and arm stretching are limited within a preset range.
[0022] Further, while monitoring the anchor area timing tracking confidence and the user cognitive load, the processing capability and the network condition are also monitored, and when the processing capability or the network condition is lower than a preset threshold, the presentation of the layered space interaction component anchored on the corresponding anchor area is adjusted, the component presented in the three-dimensional model is simplified into an icon or a text label, and the component displayed in the group layer is adjusted to be displayed in the role layer or the individual layer.
[0023] Further, the layout and level control strategy parameters include the cognitive load threshold of each role, the resource priority weight of each anchor area unit, the cost weight of occlusion, perspective offset and anchor area occupation used in layout evaluation, and the sensitivity parameter for triggering the hierarchical adjustment of the layered space interaction component between the group layer, the role layer and the individual layer, and the above parameters are corrected when the strategy parameters are updated.
[0024] Further, the strategy parameter update is based on the recorded task performance data, the task performance data includes the completion time, the number of misoperations and the number of redos, and the number of times the user adjusts the layout, after completing the task, the task performance data is summarized, the cognitive load threshold of each role, the resource priority weight and the cost weight are corrected, and the corrected parameters are written back to the layout and level control strategy parameters.
[0025] (Three) beneficial effects
[0026] The present application provides a virtual-real combined scene space interaction adaptive online processing method, which has the following beneficial effects:
[0027] By introducing the timing tracking confidence and the anchor area unit resource state in the scene semantic topology graph, the spatial anchor of the virtual content is established on the unified description of environmental stability and spatial availability, and the stable display of the virtual marker and the operation panel is maintained in the scene with occlusion, light change and slight device jitter, reducing the interference caused by virtual content drift and frequent repositioning, and providing a reliable spatial reference for subsequent multi-user space layout.
[0028] The task flow, role division and space content are associated in the same structure, which can distinguish group layer, role layer and individual layer information, and reduce the situation that different users have different understandings of the same device and the same virtual marker. According to the anchor demand of the layered space interaction component, high-quality anchor area units are preferentially allocated to group layer components and key role layer components, and then the position and size are fine-tuned in each user's field of view, so that the space layout is constrained by the scene semantic topology graph and the cooperation situation graph at the same time, ensuring the consistency of the position of shared components while considering the readability and operation convenience of individual perspective.
[0029] The actual tracking quality of the anchor area unit, the system resource state and the cognitive load of the user in the cooperation situation map are continuously monitored during operation, and when it is found that the stability of a certain anchor area unit decreases or a certain type of user is overburdened, the hierarchical spatial interaction component is triggered to migrate between different anchor area units, to adjust the level between the group layer, the role layer and the individual layer, and to simplify the complex three-dimensional content into icons or text lists as needed, so as to maintain the balance of the overall interaction rhythm when the environment changes and the task intensity changes.
[0030] Without changing the overall processing flow, the cognitive load threshold of different roles, the resource priority of the anchor area unit and the weight configuration in the layout evaluation are gradually adjusted, so that the same method can adaptively converge between different workshops, teaching scenes and remote collaboration projects, reducing the workload of manual parameter adjustment and scene reconstruction. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 The figure is a schematic diagram of the virtual-real combined scene space interaction adaptive online processing method of the application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0033] Please refer to Figure 1 The application provides a virtual-real combined scene space interaction adaptive online processing method, comprising,
[0034] Step one, in a space such as an assembly workshop, a computer room or a teaching experiment area, color video cameras, infrared depth cameras, inertial measurement units and laser ranging devices are often installed at different positions and sampled at different frequencies. If these raw data are directly used separately, it is difficult to form a unified description of the changes over time of the same ground area and the same device surface, and it is more difficult to predict tracking stability and anchoring ability for the same position in subsequent steps.
[0035] Therefore, it is necessary to first integrate multiple source observations into a set of scene basic units through a unified space coordinate system, and each subsequent update can directly correct this set of basic units without gradually aligning from the raw data level. When selecting the workspace coordinate system, the installation position and attitude of all sensors are calibrated into the coordinate system at one time.
[0036] In the running process, each frame of color image, depth image and inertial measurement unit pose data can be converted into a three-dimensional point set and pose in a unified coordinate system. According to the spatial sampling accuracy and business requirements, the working space is divided into continuous voxel units or grid units, and each unit corresponds to a scene basic unit.
[0037] When new sensor observations arrive, the corresponding pixel or point cloud sample is classified into the basic unit according to the spatial position parameter, and an observation is added to the basic unit. In this way, the observations of the same physical area over time are accumulated rather than scattered in different coordinate systems or different data structures. Among them, the initial deployment stage of the system performs a complete inspection in the workshop or experimental area, for example, the color video camera on the head-mounted display terminal acquires multi-view images, the infrared depth camera and the laser ranging device acquire depth data, and then the data is mapped to the working space coordinate system. Divide the regular grid with the ground as the reference plane to form a basic unit set that fits the real structure layer by layer.
[0038] During the operation of the device, when a new image frame or depth frame is acquired by a sensor, the spatial coordinates of each point corresponding to the pose in the frame can be read into the basic unit, and a record about the description quantity of the observation history such as illumination, texture clarity and motion blur degree is added to the basic unit. Therefore, the observable conditions in the same area can be continuously repeated in the same basic unit, and the observation content obtained by different sensors can also be repeated in the same basic unit.
[0039] In application scenarios, each observable region has its own basic unit in a coordinate system, which can provide spatial indexing for subsequent semantic segmentation, topology relationship establishment and time series confidence calculation; multi-source observation data are fused at the basic unit level, avoiding the need for subsequent algorithms to frequently register on the original image or point cloud, and the calculation path of the scene semantic-topology graph construction process becomes clear and controllable.
[0040] After the construction of the basic unit is completed, only the positioning results of one frame or a small number of frames cannot determine whether a certain space region can be used to stabilize the content in the future period of time. It is necessary to use a certain time window of observation history to add occlusion rate, feature clarity, device vibration, etc. into the time series confidence field to depict the stability and predictability of each basic unit in different time ranges. A sequence of observation history is maintained for each basic unit, including observation time, observation angle, illumination condition, texture clarity, local motion, etc. The observation attributes are counted on the sliding time window to construct a comprehensive index function sensitive to stability, and then it is used as a confidence of 0 to 1.
[0041] As time goes on, new observations enter the window while old observations are gradually weakened or discarded, so that the confidence surface evolves slowly over time rather than jumping dramatically from frame to frame.
[0042] where the temporal confidence function of a certain base unit is defined as follows: given the spatial position parameter and the current time parameter , an observation stability indicator is calculated, and then the confidence value is obtained through an exponential mapping:
[0043]
[0044] where the temporal confidence function : spatial position parameter at time parameter is the stability degree of the anchor candidate position, whose value is limited in the range of 0 to 1; the observation stability indicator : a non-negative quantity calculated based on the number of occlusions, the texture definition variation amplitude and the local motion amplitude of the base unit within a preset time window;
[0045] the decay coefficient : a positive real number, which affects the influence of the observation stability indicator on the confidence value; where the observation stability indicator can be quantified into a consistent value by accumulation. The number of occlusions is calculated by detecting whether the base unit is occluded by other objects; the texture definition is determined by image gradient or edge response; the local motion amplitude is calculated by matching the point clouds of adjacent frames to obtain the average displacement, and then summed to obtain .
[0046] Alternatively, the observation stability indicator is defined as a linear combination of three normalized sub-indicators:
[0047]
[0048] where: the occlusion metric : the percentage of the number of frames in which the base unit is occluded in the total number of frames within the sliding time window in which the position is located, taking the value . The texture variation metric : the gradient energy or local contrast of the image block corresponding to the base unit within the same window is frame-differentiated, and normalized to ; the motion metric : the maximum possible difference. : In the same window, the average amplitude of the displacement of the center point of the basic unit is obtained according to the registration of the adjacent frame point clouds, and then normalized to ; the weight coefficient : a non-negative real number, which is set by the deployment personnel according to the sensitivity of the scene to occlusion, texture and motion.
[0049] When used, the time confidence function is introduced on the basis of the original unit observation history , so that the semantic-topological graph of the scene can not only reflect the current geometric and semantic state, but also reflect the stability expectation in the future period. Any subsequent spatial layout that needs an anchor point can directly introduce the time confidence function , reducing the probability of virtual marker jitter drift.
[0050] After having a basic unit with a time confidence attribute, not every basic unit can be directly used as a candidate anchor point for virtual content. If the basic unit area is too small, the surface inclination angle is too large, or the corresponding semantic object itself is not suitable to be covered, these basic units are not suitable as actual anchor positions even if the confidence is high.
[0051] Therefore, first select one or more types of semantic objects as candidate anchor carriers from the semantic-topological graph of the scene, such as flat surfaces of device enclosures, fixed walls in rooms, or surfaces on workbenches, etc., and divide each candidate object surface into basic units. Gather adjacent basic units with similar geometric properties into several anchor area units.
[0052] For each anchor area unit, the area, the angle between the normal and the direction of gravity, the history of occlusion, etc. are calculated, and the anchor fitness coefficient is obtained from the related data of these factors, so as to determine how many virtual components the anchor area can carry and what kind of virtual components it can carry. Each anchor area can be numbered and the anchor fitness coefficient is constructed, such as:
[0053]
[0054] In the formula, the anchor fitness coefficient : the comprehensive fitness of the anchor area unit numbered as a virtual content carrying area; the area parameter : the actual area of the anchor area unit on the candidate surface, used to reflect the size of the available space; the angle parameter : the average angle between the normal of the anchor area and the direction of gravity, the closer the angle is to a right angle, the closer to zero ;
[0055] the occlusion history parameter : Based on the non-positive real number obtained from the historical occlusion event statistics of all base units in the anchor area, the negative value of the number of occlusions or the negative value of the occlusion frequency can be obtained; By counting the number of occlusion frames of the base unit in the anchor area within a time window, and taking its opposite or taking its opposite after linear normalization, it is ensured that the area with more occlusions has a smaller value, so that in it is naturally depressed; the occlusion history parameter is defined as the opposite of the occlusion frequency:
[0056]
[0057] In the formula: the number of occlusions : In the latest time window, the number of frames in which the base unit in the anchor area is marked as occluded, a non-negative integer; the total number of observation frames : In the same window, the number of frames in which the anchor area is successfully observed by at least one sensor, a positive integer. The weight coefficient 、 、 、 are non-negative real numbers that adjust the influence degree of area, posture and occlusion factors, respectively. The specific values can be set in the deployment stage according to the size of the workshop space and the characteristics of the display terminal.
[0058] When used, a larger scale anchor area unit is constructed above the base unit, and each anchor area considers factors such as area, posture and occlusion history to obtain an accurate quantified anchor fitness coefficient , which can be directly sorted and selected in subsequent anchor area scheduling and multi-user layout to make virtual content preferentially fall in areas with suitable geometric conditions and not easily occluded, improving space utilization and layout reliability.
[0059] After dividing the candidate surface into anchor area units, if only a single user perspective is used to judge the pros and cons of the anchor area, the differences in visibility and accessibility of the same anchor area by multiple users in different positions and postures are often ignored.
[0060] It is necessary to introduce multi-user viewpoints and multi-role task requirements into the anchor area resource characterization, so that each anchor area not only has a geometrically meaningful anchor fitness, but also has a group-oriented resource state description, thereby better supporting multi-user collaboration in subsequent layout.
[0061] According to the above cooperation situation diagram, the observation and operation trajectory of multiple users in the typical task stage is counted in each anchor area. Based on the visual cone, head posture and body reachable space of each user, the visibility index and accessibility index of each user to a certain anchor area are calculated respectively, and the visibility index and accessibility of multiple users are added to obtain the comprehensive usability index of multiple users of the anchor area, reflecting the average use value of the anchor area in group cooperation.
[0062] wherein the total number of cooperation users can be set as , numbered 1 to , for each anchor area numbered , the multi-user comprehensive usability index is defined as
[0063]
[0064] wherein the multi-user comprehensive usability index : the average usability of the anchor area unit numbered under the perspective of all participating users; the user contribution index : the comprehensive visibility and accessibility index of the th user relative to the anchor area numbered , which can be obtained by weighting and adding the line-of-sight coverage ratio and hand accessibility ratio of the user to the anchor area in the critical task stage, and the numerical range is usually limited between 0 and 1;
[0065] the total number of users parameter : is a positive integer, equal to the number of users participating in the current cooperation scene at the same time.
[0066] wherein the calculation of the user contribution index can rely on historical task playback data or on-site collection data, and a stable estimate can be obtained by counting the spatial relationship between the user and the anchor area under several typical working conditions.
[0067] In a multi-student practical training scene, the teacher and multiple students are distributed in different directions around a device. For an anchor area close to the front of the device, although its anchor adaptability coefficient is high, if only the teacher stands in front, most students are difficult to see the area directly from the side, so the multi-user comprehensive usability index will be low, and the system may prefer to select another anchor area that is slightly smaller but directly visible to most students when laying out at the group level.
[0068] From the perspective of the anchor area, the geometric adaptability and the accessibility of multiple users are uniformly described, so that the resource state of the anchor area can reflect the physical conditions, cooperation needs, and anchor adaptability coefficient After the combination, it is convenient to schedule each anchor area according to the combination of With the combination, each anchor area is scheduled, and the spatial structure of the virtual-real combined scene is provided with resource guarantee.
[0069] Step two, in the virtual-real combined scene, the field operators, remote experts and practical students often pay attention to different information on the same device. Some content is used to guide the rotation direction of the specific knob, some content is used to remind the risk of the current step, and some content is used to provide progress overview for the bystander.
[0070] If all virtual contents are still regarded as parallel two-dimensional windows superimposed in front of the field of view according to the traditional interface method, the key operator's field of view will be occupied by non-key information, and the bystander observer will have difficulty in judging the current process to which link. Therefore, a task-oriented hierarchical spatial interaction component description structure needs to be constructed. When describing each virtual content, its task stage, importance, time urgency and expected service object are given at the same time, so as to provide a basis for subsequent spatial arrangement and weight allocation.
[0071] At this time, for each business process, the process is first divided into continuous steps, and each step contains a number of operation sub-tasks. For each step, key actions and key tips are extracted from the operation procedure, and these contents are mapped to a number of candidate spatial interaction component prototypes. And each candidate spatial interaction component prototype is assigned a unique number, and a one-to-one correspondence between it and the specific step, the device part and the role is established, so that when the actual virtual component is generated later, it can be directly instantiated according to the description structure without reinterpreting the business semantics.
[0072] Further, in the setting, the component number, the task step to which it belongs, the task type, the importance level, the time urgency level, the role set, and the expected associated device area are set in each spatial interaction component.
[0073] For example, in the process of disassembling the high-pressure pump shell, the virtual number component used to guide the bolt disassembly sequence will be set as the core component of the step, the importance level is high, the time urgency level is medium, the role set is the field operator and the practical student, and the expected associated device area is the bolt around the pump shell.
[0074] In runtime, when it is detected that the process advances to the step, the fields of the component are read from the description structure, a spatial interaction component entity containing arrow prompts and text instructions is instantiated in the scene, and the entity is associated with the anchor area unit formed in step one to establish a reference relationship, so as to prepare for subsequent spatial anchoring and display control.
[0075] In one embodiment, there is an operator wearing a head-mounted display terminal on the spot, a student watching through a tablet next to him, and an expert participating in the review through a conference room screen. In the spatial interaction component description structure generated in this scenario, the same disassembly sequence prompt can be marked as core information for the operator, important information for the student, and reference information for the expert. In the subsequent layout stage, the presentation of the component on different terminals and whether it should appear in the personal layer, role layer or group layer are determined accordingly, without the need to write separate logic for each terminal. Thus, each virtual prompt in the scenario is marked with task and role from the beginning, and a stable correspondence with the actual device area is established, laying a unified description foundation for multi-user spatial interaction.
[0076] When in use, all virtual content is summarized into structured hierarchical spatial interaction components from the beginning, each component having clear fields in terms of task sequence, role division and device association, thereby providing clear input for subsequent role-based spatial hierarchy decision and cross-terminal display mode selection.
[0077] After obtaining the structured hierarchical spatial interaction component description, the textual fields of importance, time urgency and risk association are converted into numerical values that can participate in subsequent decision-making, and accordingly it is determined whether each component is more suitable for the personal layer, role layer or group layer.
[0078] Without uniform numerical description, different project teams may have inconsistent scales when configuring components, and the same meaning of importance may represent different intensities in different processes, leading to deviations in subsequent layout stage sorting and filtering.
[0079] Further, three main task attribute indicators are set for each component, corresponding to task importance, time urgency and risk association.
[0080] Task importance focuses on whether the component directly affects the success of the task, time urgency represents the time sensitivity of the component information invalidation, and risk association represents the association strength of the component with safety risk or quality risk.
[0081] The three types of indicators are normalized and assigned a set of weights for linear combination. Components with high component hierarchy priority coefficients will be placed more in the group layer or important role layer, ensuring that all relevant personnel can see components with low component hierarchy priority coefficients placed more in the personal layer, providing instructions to individual users only when needed.
[0082] The component hierarchy priority coefficient is set to be , the task importance index is , the time urgency index is , and the risk association index is Then generate component layer priority coefficient :
[0083]
[0084] In the formula, the component layer priority coefficient represents the priority degree of the spatial interaction component numbered in the spatial layer division;
[0085] Task importance index : the normalized value of the role of the component in the success of the task, which can be configured according to whether the task needs to be reworked and whether it affects the key quality indicators when the component is missing, and the value range is usually set between 0 and 1; Time urgency index : the time sensitivity of the component information, and the value range is also set between 0 and 1; Risk correlation index : the correlation degree of the component with safety risk and quality risk, and the value range is also set between 0 and 1; Weight coefficient , , : all are non-negative real numbers, used to balance the influence proportion of the three types of indexes;
[0086] When used, the hierarchical spatial interaction component has a clear sorting rule in subsequent layout and display decisions. Component layer priority coefficient Once determined, it can be repeatedly applied in different scenarios and different deployments, and is used together with the anchoring adaptability coefficient and the multi-user comprehensive usability in step one to anchor the region as an input for subsequent spatial layout.
[0087] Under the premise that the hierarchical spatial interaction component has been established, if the quantitative description of the cooperation relationship between multiple users is missing, it is difficult to determine which users need to carry out collaborative operation around the same device area in the same time period and which users are more in the observation role in the subsequent spatial layout stage.
[0088] The line weight in the cooperation situation diagram is based on the cooperation strength between users and between users and task objects, which is extracted from the line-of-sight sequence and operation sequence of multiple users, so that the subsequent layout can be adjusted around the real cooperation structure, rather than just statically divided according to the role name.
[0089] For each user, continuously record his / her head orientation, eye gaze direction and hand motion events, and map these records to the basic units and anchor area units of the semantic-topological map. For any two users, the system calculates the proportion of time that both users gaze at the same device area or the same virtual component within the same time window, and the proportion of time that one user operates and the other user gazes at the same area within adjacent time windows. The former reflects the co-view relationship between the two users in the information layer, and the latter reflects the behavior relationship of guidance and observation or handover operation. For the relationship between the user and the task object, the system calculates the proportion of time that the user gazes at a certain object and the number of times that the user operates near the object in each task stage, to obtain the cooperation strength between the user and the object.
[0090] wherein the number of users commonly in the current cooperation scene can be set as , numbered 1 to , for any two users and , define the co-view times index and the cooperative operation frequency index , and construct the cooperation relationship strength weight as follows:
[0091]
[0092] wherein the cooperation relationship strength weight : the overall cooperation strength formed by user and user around the task object in the current scene; the co-view times index : the proportion of time that user and user co-gaze at the same device area or virtual component within the same period, with a value between 0 and 1; the co-view times index can be defined as the proportion of frames in which the lines of sight of the two users fall on the same device area or component within a sliding time window, to the number of frames in the window;
[0093] the cooperative operation frequency index : the proportion of events that user and user operate on the same object in sequence or one user operates and one user observes within adjacent time periods, which can be calculated by the time adjacency relationship between hand motion events and line of sight events, with a value between 0 and 1;
[0094] weight coefficients and : non-negative real numbers, used to balance the proportion of co-view behavior and cooperative operation behavior in the cooperation relationship strength weight .
[0095] When in use, the collaboration process that is originally difficult to quantify is converted into the collaboration relationship strength weight The graph structure is depicted, so that the collaboration situation is no longer a simple relationship judgment, but a fine strong and weak relationship description. Subsequently, when arranging the group layer and role layer components, the collaboration relationship strength weight The users who closely collaborate with each other are aggregated and displayed around the same device area, thereby improving the collaboration efficiency.
[0096] Only the collaboration relationship strength weight is not enough to reflect the state difference of each user in the current task stage. Some users may be highly nervous and make frequent mistakes, and some users may be inattentive and have not operated effectively for a long time.
[0097] It is necessary to combine the collaboration relationship strength, task reaction time, error number, and stagnation duration, to construct a cognitive load index and a participation index for each user, so that the nodes in the collaboration situation graph not only have a structural position, but also have a state attribute, providing a basis for subsequent differentiated information delivery for different users.
[0098] At each execution of the task, record the response time of each user to each prompt, the operation number of the physical device and virtual component, the number of caused misoperations, and the duration of no operation in a period of time. Combine these records with the previous collaboration relationship strength weight , to determine how much load each user is currently under, and whether the participation is large or small.
[0099] In order to keep the cognitive load index within a limited interval and be able to compress extreme cases, the system uses an exponential form to construct the cognitive load function, so that the output is closer to the upper limit when the input value is larger, thereby avoiding the exponential infinite growth in extreme cases. Among them, for the user numbered , the average reaction time is , the misoperation count is , and the stagnation duration is , and the non-negative weight coefficients , , are selected to construct the cognitive load index as follows:
[0100]
[0101] Among them, the cognitive load index : the load level of the user numbered in the current task stage, the value is between 0 and 1; the average reaction time : The average time between the appearance of several key cues and the start of the corresponding operation by the user, which can be automatically recorded by the system during the task execution, non-negative real number; Misoperation count : The number of error operations triggered by the user in the current task stage, which is a non-negative integer; Stagnation duration : The cumulative time of the user not performing effective operations within a set time window, which is a non-negative real number; Weight coefficient 、 、 Used to adjust the contribution of the three factors of reaction time, misoperation and stagnation duration to cognitive load.
[0102] Because the exponential function tends to 1 when the input is large, when 、 、 , the cognitive load index approaches the upper bound, which is conducive to early identification of high-load users.
[0103] When used, each user node of the collaboration situation diagram has a cognitive load index and behavior characteristics derived from 、 、 , thereby adding a state dimension to the structural relationship. Combined with the aforementioned collaboration relationship strength weight , the collaboration situation diagram can not only describe who is working together, but also describe how much each person is occupied by the task, providing rich decision-making basis for the placement and adjustment of subsequent hierarchical space interaction components, and laying a data foundation for the differentiated layout strategy for high-load users and low-participation users in step three.
[0104] Step three, if there is a lack of a fine connection chain between the anchor area resource state obtained in step one and the hierarchical space interaction component attributes obtained in step two, it is easy to fall into the rough processing mode of placing nearby or allocating in order, which cannot consider the time sequence tracking confidence, geometric adaptability, multi-user comprehensive availability of the anchor area, and the task urgency and risk level of the component itself.
[0105] Therefore, a set of candidate relationships and comprehensive scoring structures are established between the anchor area side and the component side, so that each component has a group of selected candidate anchor areas before entering the subsequent layout solution, and the adaptation degree of each candidate anchor area is numerically described in a comparable manner.
[0106] First, read the anchor adaptability coefficient and the multi-user comprehensive availability index calculated in step one for each anchor area Simultaneously read the time tracking confidence function in the spatial range corresponding to the anchor area The representative statistic of the current time period, such as the lower bound value or weighted average value in a small time window.
[0107] Then number each component according to step two The calculated component level priority coefficient And the anchor requirements of the component itself (such as the lowest anchorable level, the avoidance requirement of the safety identification shield) select a set that meets the hard constraint condition in all anchor areas as the anchor area set of the component. Calculate a comprehensive score for each anchor area in the set for scheduling options.
[0108] Among them, the spatial interaction component numbered and the anchor area numbered The comprehensive score function is constructed , which is in the form of:
[0109]
[0110] In the formula, the comprehensive score function : When the spatial interaction component is anchored in the anchor area , the degree of adaptation of the combination is comprehensively considered from the task priority, geometric adaptability, and multi-user comprehensive availability; The component level priority coefficient is given in step two, which reflects the comprehensive result of the importance of the component in the task chain, the time urgency, and the risk association, with a value between 0 and 1; The anchoring adaptability coefficient : Defined in step one, reflecting the suitability of the anchor area in area, pose, and occlusion history, which is a non-negative real number; The multi-user comprehensive availability index also comes from step one, reflecting the average visible and accessible degree of the anchor area from all participating users' perspectives, which is between 0 and 1; The weight coefficient , , : Non-negative real numbers, used to adjust the proportion of task factors, geometric factors, and multi-user factors in the comprehensive score function in specific deployment.
[0111] When used, a candidate relationship with a score is formed between each component and several anchor areas, which can avoid the component falling into a position with unstable tracking or difficult observation, and also provides a clear numerical basis for the subsequent anchor area scheduling process, which is beneficial to guarantee the task priority while considering the observation and operation needs of multiple users.
[0112] When each component has a set of candidate anchor areas and a comprehensive score, if the anchor area occupation upper limit is not considered, and only the score is allocated from high to low, it is easy to appear that a small number of anchor areas are occupied by a large number of components, and the remaining anchor areas are idle for a long time, which is also not conducive to the neatness of the field of view and reduces the adjustable space of subsequent layout.
[0113] It is necessary to balance between the component side comprehensive score and the resource occupation limit on the anchor area side. An initial anchor area allocation scheme is generated through a set of deterministic allocation traversal process, and hierarchical adjustment suggestions are automatically generated in resource-intensive areas, providing a starting point that meets the physical limit conditions for subsequent two-level layout solving.
[0114] First, according to the component level priority coefficient Sort all components from high to low, and arrange the group layer and key role layer components in the front row. Then, for each component in the sorted sequence, arrange the candidate anchor area set in the comprehensive score function From high to low, and check the resource occupation and capacity upper limit of the corresponding anchor area at the current time period.
[0115] If an anchor area still has a surplus in terms of area, allowed component number, and obstruction risk margin, the component is temporarily allocated to the anchor area, and the occupation state of the anchor area is updated; if all candidate anchor areas cannot receive the component under the resource occupation limit, the system records a compression suggestion, indicating that the component needs to consider reducing the space level or using time switching to display in subsequent steps.
[0116] Among them, the specific operation can be performed in a row-by-row traversal manner, starting from the first component in the sorted sequence and matching the anchor area one by one. Still taking the engine assembly training implementation example mentioned above as an example, first process the group layer components related to safety, and allocate them to the anchor area with a higher comprehensive score function And sufficient resources.
[0117] For example, a large wall area built around the engine, since the anchor fitness coefficient And the multi-user comprehensive availability index Are relatively high, so the global risk prompt panel will be given priority to be carried. After the available area of this wall area gradually decreases, subsequent component processing will switch to using the side anchor area around the engine. In addition, for personal layer components with a lower level priority coefficient When all candidate anchor areas cannot meet the resource occupation limit, the system will record in the component description that only the local display of this level adjustment information on the personal terminal, and hand over to the subsequent layout solving stage for processing.
[0118] When used, an initial allocation scheme that takes into account both task priority and resource constraints of anchor areas is constructed using sorting and sequential traversal of engineering methods without relying on complex mathematical solvers, and hierarchical adjustment suggestions are generated in advance for resource-constrained components. Thus, the subsequent spatial layout solution does not need to face a completely blank allocation state, improving the controllability of the scheme landing.
[0119] Further, after obtaining the initial anchor area allocation scheme and hierarchical adjustment suggestions, if only considering whether each component has been anchored, the geometric relationship consistency of group layer and role layer components in the multi-user perspective cannot be guaranteed.
[0120] For example, even if a certain key prompt panel is anchored in a high adaptability area, if the panel is in completely different directions in the field of view for different users, pointing to the location of the panel with fingers during collaboration is easy to deviate. It is necessary to introduce user visual field geometry and collaboration relationship information at the anchor area level to adjust the accurate position and orientation of group layer and role layer components within the anchor area, so that key components form a relatively stable consistent geometric relationship in the perspective of most collaborative users.
[0121] Based on this, for each user, according to the scene semantic-topological graph and terminal posture constructed in step one, the incident direction and field of view coverage area of each anchor area under the current posture of the user are calculated. For group layer or key role layer components that have been allocated to a certain anchor area, the available layout sub-regions are divided on the surface of the anchor area, and angle analysis is performed on different sub-regions, calculating the observation deflection angle, offset distance from the field of view center and occlusion probability of each user on these sub-regions, while estimating the visual angle difference between users with close collaboration relationship on these sub-regions.
[0122] Referring to the collaboration relationship strength weight in step two Set users with close collaboration relationship as the priority guarantee object, so that the visual angle difference of these users for key components is minimized.
[0123] Among them, the visual angle cost term can be constructed for each user and each candidate layout sub-region, and the line of sight deflection angle, field of view edge distance and occlusion probability are combined into a single-user view geometric comfort index according to fixed weights, and then the sub-region with smaller aggregated cost is selected as the final position of the key component within the anchor area.
[0124] For example, in an engine training scenario, a teacher and two main trainees usually stand on the same side of the engine. When the system anchors a group layer step overview panel on the wall above the engine, it will prefer to position the panel in the intersection area of the visual fields of these three users, so that the panel position in the central area of the wall is visible to all three of them at the same time, rather than being biased to one side for one of them. For the trainee standing on the other side, the system can compensate for the panel not being easily observable from this perspective by copying a simplified step overview bar into his visual field in the later personalized layout stage.
[0125] Using user visual field geometry and collaboration relationship strength weight in the anchoring area to make the positions of the group layer and role layer components remain relatively fixed in geometry with the majority of collaborating users under the premise of meeting the resource constraints of the anchoring area, thereby reducing the inaccuracy of collaboration instructions and verbal communication.
[0126] After the group layer and role layer components complete the layout in the anchoring area, if the cognitive load differences of individual users and the superimposed effects of personal layer components are ignored, it is easy to have the phenomenon of overload in the visual field of some users and insufficient information in the visual field of other users. Using the cognitive load index constructed for each user in step two , combined with the number and importance of the user's personal layer and role layer components, a personal information load index is constructed, and the personal layer layout of each user is fine-tuned without destroying the overall structure of the group layer, so that the amount of information and cognitive load are coordinated.
[0127] In the current task stage, for each user, the set of components that are truly needed to be focused on in the visual field is counted, the components belonging to the group layer and the role layer are considered as the basic responsibility set, and the personal layer components that are only open to the user are considered as the incremental set. For each component, the system directly reads the component level priority coefficient of the component , which is considered as the occupancy strength of the component on the user's attention. And for each user, a personal information load index is constructed, which is the accumulation of the component level priority coefficients of all components in the current visual field of the user, obtaining an information load characterization related to the importance of the task, and then superimposing this characterization with the cognitive load index to form a comprehensive index to guide the position fine-tuning or display control of the personal layer components.
[0128] Among them, for the user numbered , its personal information load index is denoted as , and the set of components in its visual field is denoted as , then
[0129]
[0130] wherein the personal information load index is the sum of the importance of all components in the current view of the user numbered ; the component set is the set of component numbers actually presented in the current view of the user; the component hierarchy priority coefficient is the comprehensive value of the spatial interaction component defined in step two in terms of task importance, time urgency and risk relevance.
[0131] After obtaining the personal information load index , the system can construct a comprehensive load index , for example, let wherein the comprehensive load index represents the overall burden level of the user numbered under the current layout state considering the cognitive load and information load comprehensively, and the adjustment coefficient is a non-negative real number used to control the weight of the information load in the comprehensive load index.
[0132] When the comprehensive load index of a certain user exceeds the safety threshold set in advance, several personal layers or secondary role layers components will be selected in order from low to high according to the component hierarchy priority coefficient and moved to the edge of the view or displayed in a folded manner to reduce the user's current perception burden; for users with a lower comprehensive load index , the system can introduce some additional personal layer prompts or exercise guidance to improve their participation.
[0133] In the engine training embodiment, when the system finds that the cognitive load index of a certain main operator trainee has approached the upper limit in several key steps, and the personal information load index is also at a high level, it will temporarily shrink some explanatory personal layer text instructions into icons and display them only when the trainee pauses; for the trainee standing in the back row, since his cognitive load index is usually low and the personal information load index is also small, the system can add some structural schematic or principle explanation components in his view to improve the overall teaching effect without interfering with the main operator.
[0134] In application, the component hierarchy priority coefficient , the cognitive load index and the personal information load index In combination, a round of fine-tuning layout for different users is arranged in the personal field of vision, the amount of information corresponds to the personal state, which not only ensures that high-load users are not disturbed but also provides more opportunities for low-load users to learn and explore, and at the same time, a layer of delicate personalized adjustment is formed at the group level layout.
[0135] Step four, the multi-user space layout obtained in step three is the result formed in a time window according to the anchor adaptability coefficient, the multi-user comprehensive availability index and the component level priority coefficient, and the tracking quality and the environment blocking condition in the virtual-real combined scene often change slowly or suddenly over time.
[0136] If there is no risk quantity that comprehensively reflects the timing tracking confidence function and the anchor area resource occupation at the anchor area level, it is difficult to discover that a certain anchor area is no longer suitable as a bearing position of a key component in the future time period, thereby causing the group layer component to vibrate, drift or frequently disappear.
[0137] Therefore, for each anchor area in the layout, the timing tracking confidence function is used in a loop, for each anchor area number , the basic unit contained therein is read, and the timing tracking confidence function of these basic units in the current time window is weighted and integrated or weighted and accumulated to construct the anchor area risk with confidence decline as the main factor.
[0138] Combined with the number of spatial interaction components currently contained in the anchor area and the component level priority coefficient of the component , a weighted term of key component density is constructed, and the weighted term is superimposed and combined to obtain a cross-layer risk index: which represents the risk level of the anchor area under the joint action of the spatial layout, the tracking quality and the task importance at the current time;
[0139] The space region where the anchor area number is located is , the weight function is defined to represent the contribution degree of the basic unit to the internal anchor area, and the anchor area risk index is constructed at the time parameter time:
[0140]
[0141] In the formula, the anchor area risk index : the risk accumulation amount of the anchor area number caused by the timing tracking confidence function at the time parameter ; the timing tracking confidence function : the basic unit level confidence defined in step one, the value range is between 0 and 1, and when Close to 1 means the position tracking is stable, when Close to 0 means the position tracking is unreliable;
[0142] Integral region : the three-dimensional space region corresponding to the anchor region numbered ; weight function : a non-negative function that gives different weights to the basic units at different positions within the anchor region; It can take a constant equal to the area of the basic unit, or give a larger weight to the basic unit near the key part of the device to amplify the risk of the key area.
[0143] If the anchor region is divided into discrete basic units, the above integral can be approximated by finite summation, and the weight function is represented by the area or volume of the basic unit.
[0144] In practice, the anchor region is composed of a finite number of basic units, and the above integral can be realized by summation:
[0145]
[0146] In the formula: the set of basic units : the anchor region contains all the basic unit index sets; the center position : the center coordinates of the basic unit in the workspace coordinate system; the weight : it can take the area or volume of the basic unit, or a larger value on the unit near the key part of the device.
[0147] When used, the time sequence tracking confidence function given in step one is unified as a risk index , which is convenient for the system to continuously track the changes in tracking stability of each anchor region during the entire layout process, convenient for cross-layer layout adjustment, and avoid making local decisions in single frame situation.
[0148] In addition, after the anchor region risk index is constructed, if it is not connected with the user-side comprehensive load index , and only the migration or degradation operation is performed on all components according to the threshold value, it will inevitably lead to important components being excessively adjusted or even interrupted by high-load users in the view.
[0149] The spatial side risk and the user side load should be considered simultaneously in the layout adjustment process, so that the adjustment intensity reflects not only the instability degree of the anchor area in space, but also the responsibility size of the components in the anchor area to various users, so that low-priority components or components in the view of low-load users are adjusted in priority when necessary.
[0150] For each anchor area number , read the component set anchored in the anchor area in the current layout and its component hierarchical priority coefficient , and combine the comprehensive load index established for each user in step three to construct a layout adjustment intensity measure for components-anchor areas-users. The layout adjustment intensity measure is used to determine which components need to be moved out of the current anchor area, whether to trigger hierarchical degradation, or only move in individual user views.
[0151] Wherein, when the spatial interaction component numbered is anchored in the anchor area numbered and focuses on the user numbered , the layout adjustment intensity is constructed as follows:
[0152]
[0153] The layout adjustment intensity : the comprehensive measure of whether the layout of component in anchor area needs to be adjusted and the adjustment amplitude at time parameter ; Anchor area risk index : a non-negative quantity constructed to reflect the stability change of the anchor area; Comprehensive load index : the user side load characterization constructed in step three by combining the cognitive load index and the personal information load index ; Weight coefficients and are non-negative real numbers used to balance the contribution proportion of spatial risk and user load in the layout adjustment intensity .
[0154] The layout adjustment threshold can be set, when , component is migrated from anchor area to a candidate anchor area with lower risk, or hierarchical degradation or simplified display is performed on the component
[0155] In actual implementation, the anchor area risk index The load increases due to frequent occlusion, and the overview panel on this anchored area is a group-layer component, crucial for all students and teachers. Calculate the overall load index for the teacher node. It was found that the teacher's performance was still within a manageable range, so the intensity of the layout adjustment for this component in the teacher's view was adjusted. The risk level does not increase significantly; the system simply copies some explanatory content from the student's view to another anchoring area on the wall and simplifies its display on the original anchoring area to mitigate the impact of passageway obstruction. The overall workload index of many trainees continues to rise. When the priority level also increases, the system begins to prioritize components according to their hierarchical coefficients. Gradually move less critical group-level components down to the role or individual level, retaining only the most essential overall process overview.
[0156] When using it, in the cross-layer risk index and comprehensive load index Based on the structural layout, adjust the intensity The document uses specific examples to illustrate how to improve the overall stability of the scenario without interrupting critical tasks by taking adjustments such as migration, downgrading, or simplifying the display under the dual constraints of rising spatial risks and changing user load.
[0157] Once the aforementioned cross-layer layout adjustment mechanism is established, if all trigger thresholds, weight coefficients, and migration rules are fixed, the system's performance in different projects and different scenarios will heavily depend on the initial configuration. It will be impossible to change the strategy based on long-term accumulated task completion data and user feedback. The series of weight coefficients, thresholds, and hierarchical decision parameters required for the aforementioned layout adjustment process are abstracted into a strategy parameter vector. Furthermore, by aggregating performance indicators from multiple task processes, an objective function is constructed, and the strategy parameter vector is updated with gradients based on the objective function. Thus, while maintaining the overall decision structure, parameters that better reflect the actual situation are gradually constructed.
[0158] After each task process, a set of performance metrics is recorded, including but not limited to whether key steps were completed on the first attempt, number of rework attempts, total task duration, number of serious misoperation events, and quantitative results of user subjective evaluation. Furthermore, these performance metrics are correlated with the strategy parameter vector used in this process, accumulating performance data from multiple processes into a historical objective function. For this objective function, the system uses numerical differentiation or automatic differentiation tools to calculate the gradient approximation with respect to the strategy parameter vector, determining the direction and magnitude of adjustments for each parameter without disrupting the original structure.
[0159] Among them, the strategy parameter vector can be set as This includes the layout adjustment intensity weighting coefficient. , Multiple components, including risk threshold, comprehensive load threshold, and component-level priority coefficient combination weight;
[0160] Define the objective function This is used to characterize the overall performance of multiple task flows under a given policy parameter vector, such as a weighted sum of rework counts, serious error events, and timeouts. After each round of data accumulation, the policy parameter vector is updated as follows:
[0161]
[0162] Among them, the new round of policy parameter vector : Represents the set of strategy parameters after a single parameter adjustment, which will be used in subsequent task flows; the current strategy parameter vector. This is the set of parameters used before the execution of this round of tasks;
[0163] Learning step size coefficient : A positive real number whose value determines the magnitude of parameter adjustment during each update. Excessive size may cause fluctuations. If it is too small, the convergence will be slow;
[0164] objective function The larger the value of the task performance characterization function constructed above, the worse the task performance under the current policy parameter configuration; the long-term performance objective function within the task cycle can be defined as follows:
[0165]
[0166] Where: average number of rework sessions The average number of rework attempts per task across multiple tasks; the average number of serious errors. The average count of serious errors across multiple tasks; the average timeout duration exceeding... : Average duration exceeding the expected completion time; weight A non-negative real number used to balance the importance of different types of costs. Gradient It can be approximated by finite difference: for Each component is subjected to a small perturbation, and the comparison is performed. The changes before and after are used to obtain an approximate gradient value, and then the parameters are adjusted according to the above update formula.
[0167] gradient vector The gradient of the objective function with respect to the policy parameter vector can be approximated by finite difference or automatic differentiation, which gives the direction in which each parameter should be adjusted. A common solver such as the quasi-Newton method or the conjugate gradient method can be used to embed the update rule in the system background and execute it regularly.
[0168] During the implementation of the assembly training line, a higher risk threshold and a lower comprehensive load threshold are set in the initial deployment stage. As multiple batches of training tasks are completed, the number of rework and serious misoperation counts are gradually obtained and added to the objective function The objective function When it is too high in some batches, part of the components of the gradient vector point to the direction of reducing the risk threshold and increasing the comprehensive load threshold. Thus, the policy parameter vector is adjusted through the update rule. In subsequent training, the components that are moved off the high-risk anchor area are selected earlier in more training tasks, and some user load thresholds are appropriately reduced, forming a balanced strategy configuration in experiments and production,
[0169] When in use, without changing the overall decision framework, the policy parameter vector and the gradient of the objective function are updated, and the multi-parameter layout and threshold adjustment in the system are coordinated and controlled using long-term accumulated scene performance data, forming an experienced strategy configuration for specific projects.
[0170] However, simply adjusting the weight coefficients in the policy parameter vector is not enough to subdivide and track the cognitive upper limit of different groups and different roles. The typical distribution of personnel in different enterprises, different positions, and different training stages on the cognitive load index and the comprehensive load index will be significantly different. If a unified fixed threshold is still used, it is easy for some users to be in a state close to the threshold for a long time, while other users are far below the threshold most of the time, leading to inconsistent sensitivity of layout adjustment decisions for different users. Therefore, a user-oriented cognitive threshold parameter needs to be constructed, and the threshold parameter is self-adjusted according to the historical trajectory of the cognitive load index and the comprehensive load index to make the system form differentiated threshold configurations for different user groups while maintaining structural unity.
[0171] The system sets a cognitive threshold parameter for each user number , which can be allocated according to the job category or training stage at the beginning. After each round of tasks is completed, the system statistics the cognitive load index and the cognitive load index of the user, and compare the relationship between the two. If it is found that the cognitive load index of the user is much lower than the comprehensive load index for most of the time, it means that the current threshold setting is conservative, and the threshold parameter can be moderately increased to prevent the system from adjusting the layout of the user too frequently; on the contrary, if the cognitive load index of the user is close to or even exceeds the comprehensive load index for a long time, the threshold parameter needs to be gradually reduced so that the system triggers information simplification and layout migration earlier in subsequent tasks.
[0172] wherein the cognitive threshold parameter of each user can be updated as follows:
[0173]
[0174] wherein the updated cognitive threshold parameter : the new cognitive threshold set for the user with the number after the current task is completed, which will be used to determine whether the user needs to trigger view simplification and layout adjustment in the next round of tasks;
[0175] current cognitive threshold parameter : the threshold setting of the user before the task is performed; cognitive load index : the load characterization constructed according to the reaction time, the number of maloperations, and the stagnation duration in steps two and three, with a value between 0 and 1; comprehensive load index : the overall burden value obtained by combining the cognitive load index and the personal information load index in step three, which also varies around 0 to 1; adjustment coefficient : a real number less than 1, used to control the magnitude of each threshold adjustment.
[0176] In an assembly line example containing new recruits and experienced operators, different initial cognitive threshold parameters are assigned to new recruits and experienced operators, respectively. After a number of consecutive training tasks, the system finds that the cognitive load index of experienced operators is significantly lower than the comprehensive load index for a long time, indicating that such users still have additional information carrying space. According to the above rules, the cognitive threshold parameter The threshold is gradually increased, so that the system reduces the number of view simplification triggers in subsequent tasks, allowing more auxiliary information and complex components to be presented in the field of view; at the same time, the cognitive load index of part of the new trainees Often higher than the comprehensive load index The corresponding cognitive threshold parameter Is slightly reduced in several consecutive task periods, and the system compresses the display area of non-key components in the view of these trainees in advance, and delays the appearance of some complex explanations, so that they can concentrate on completing key steps. In actual observation, teachers can see that the virtual panel in front of new trainees is relatively simple, while the information in front of experienced operators is more abundant, and this differentiated effect is a direct manifestation of the self-adjustment of the cognitive threshold parameter.
[0177] When used, the threshold of different user groups can be automatically adjusted in several task periods to achieve different granularity of information control for users of different roles and different proficiency while the overall layout decision logic remains unchanged, and combined with the gradient update of the aforementioned strategy parameter vector Step four can have the ability to learn and adapt from both cross-layer layout adjustment and long-term strategy evolution.
[0178] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0179] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0180] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0181] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment of the present application according to actual needs.
[0182] The above is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A virtual-real integrated scene space interaction adaptive online processing method, characterized by: include, By collecting images and depth data of the virtual and real scene through multiple sensors, 3D reconstruction and semantic segmentation are performed to construct a scene semantic topology map containing adjacency, visibility and access relationships. The scene is then divided into anchoring area units with resource status according to the confidence level and anchorability level tracked in time. The virtual content is abstracted into hierarchical spatial interaction components that carry steps, interaction types and spatial levels. Anchoring requirements are set for the components, and a collaborative situation diagram with users, task objects and components as nodes is constructed based on user posture, field of vision and line of sight and operation trajectory. Based on the resource status of the anchoring area unit and the anchoring requirements of the hierarchical spatial interaction components, the anchoring area scheduling is performed according to the adaptability. High-confidence anchoring areas are allocated to group layer and role layer components first, and group layer layout and individual layer fine-tuning are carried out in combination with multi-user perspectives within the anchoring area. During operation, the tracking confidence level of the anchoring area unit and the user cognitive load in the collaboration status map are monitored. When the confidence level or cognitive load exceeds the threshold, the position and level of the deployed components are adjusted across layers, and the strategy parameters are updated according to the task completion status and user intervention records. Each hierarchical spatial interaction component carries a step identifier, a type identifier, and an interaction type identifier, and records the semantic object or semantic region identifier to be associated. Anchoring requirements include the acceptable lower limit of temporal tracking confidence for the hierarchical spatial interaction component, the allowed range of anchoring zone levels, and the range of target field of view orientation and field of view depth. For each user, maintain head posture, field of vision, and reachable space. Associate the currently gazed semantic object and the hierarchical spatial interaction component being operated with the user's role information. When constructing the collaborative situation map, use the user, semantic object, and hierarchical spatial interaction component as nodes. Mark the edges of co-viewing relationship, operation relationship, and guidance relationship according to the gaze and operation trajectory, and record the cognitive load and participation on the nodes.
2. The adaptive online processing method for virtual-real combined scene space interaction according to claim 1, characterized in that: The multi-sensor system includes a color video camera, an infrared depth camera, an inertial measurement unit, and a laser ranging device. It uses simultaneous localization and mapping (SMR) algorithms to process environmental data and generate a 3D scene model that includes a 3D point cloud or mesh model and the attitude trajectory of the terminal device. It also establishes a scene semantic topology map based on the geometric proximity relationship, unobstructed line-of-sight relationship, and passability relationship between semantic objects.
3. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 2, characterized in that: The anchoring area resource status information includes the time-series tracking confidence sequence of the anchoring area unit within a preset time window, the visibility markers for different user viewpoints, the accessibility markers for different user bodies and hands, and the number and type information of the hierarchical spatial interaction components attached to the anchoring area unit. At the end of step one, the anchoring area resource status information is associated and stored with the scene semantic topology map.
4. The adaptive online processing method for virtual-real combined scene space interaction according to claim 3, characterized in that: When scheduling anchor areas for hierarchical spatial interaction components, the components are divided into group-related components and role-related components according to hierarchical attributes. The anchor area resource status is queried, and anchor area units with higher time-series tracking confidence and visibility are preferentially allocated to group-related components, and then allocated to role-related components. When there are insufficient available anchor area units, multiple components are displayed in rotation on the same anchor area unit according to a preset time slice.
5. The adaptive online processing method for virtual-real combined scene space interaction according to claim 4, characterized in that: When fine-tuning the position and size of components within each user's field of vision, the position and size of the hierarchical spatial interactive components assigned to that user are adjusted based on each user's field of vision and accessible space. Keep the anchor area scheduling result unchanged, so that components do not obscure semantic objects with occlusion importance weights, and limit the user's head rotation and arm extension to a preset range.
6. The adaptive online processing method for virtual-real combined scene space interaction according to claim 5, characterized in that: While monitoring the confidence level of time-series tracking in the anchored area and the cognitive load of users, the system also monitors processing capacity and network conditions. When the processing capacity or network conditions are lower than the preset threshold, the presentation of the hierarchical spatial interactive components anchored in the corresponding anchored area is adjusted. Components presented as 3D models are simplified into icons or text labels, and components displayed in the group layer are adjusted to be displayed in the role layer or individual layer.
7. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 6, characterized in that: The layout and hierarchy control strategy parameters include the cognitive load threshold of each role, the resource priority weight of various anchoring area units, the cost weight of occlusion, view offset and anchoring area occupation in the layout evaluation, and the sensitivity parameters for triggering the hierarchical adjustment of the hierarchical spatial interaction components between the group layer, role layer and individual layer. The above parameters are corrected when the strategy parameters are updated.
8. The adaptive online processing method for virtual-real combined scene spatial interaction according to claim 7, characterized in that: The strategy parameter update is based on the recorded task performance data, which includes completion time, number of misoperations and redo times, and the number of times the user adjusts the layout. After the task is completed, the task performance data is summarized, and the cognitive load threshold, resource priority weight, and cost weight of each role are corrected. The corrected parameters are then written back to the layout and hierarchy control strategy parameters.
Citation Information
Patent Citations
Indoor unfamiliar scene recognition system fusing knowledge graph and spatial semantic topological graph
CN114972938A
Autonomous lifelong SLAM method and system based on visual language model hidden space representation
CN120599495A