An asynchronous object pose stream frame alignment anchoring method and system for perspective blending

CN122820840APending Publication Date: 2026-09-25BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611139878.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

二者之间缺少一种面向真实物体锚定的中间机制:它既要知道位姿来自哪一帧图像,又要知道该帧图像采集时头戴设备中相机的世界位姿,还要根据感知质量决定是否把该位姿真正应用到混合现实锚点上

Benefits of technology

将外部异步六自由度物体位姿流转换为混合现实世界锚点时,使用图像采集时刻的参考相机世界位姿,而不是位姿消息到达时刻的头显位姿,以减少头显运动、网络传输和模型推理延迟引起的虚拟内容滑移,提高真实物体锚定的一致性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820840A_ABST
    Figure CN122820840A_ABST
Patent Text Reader

Abstract

The application discloses an asynchronous object pose stream frame alignment anchoring method for perspective mixing. The method comprises steps 1-10: step 1, collecting binocular perspective images and generating frame numbers; step 2, sending the binocular images and camera calibration information through a data surface; step 3, estimating depth after the external visual computing device receives the images; step 4, performing initial registration by the external visual computing device to obtain an initial six-degree-of-freedom pose of a target object in a camera coordinate system; step 5, generating a pose result message; step 6, inquiring a reference camera world pose corresponding to a collection time of the frame number; step 7, converting the camera coordinate pose; step 8, obtaining an anchor point pose of the target object in a mixed reality world coordinate system; step 9, determining whether to update the anchor point; and step 10, forming a closed loop through a state event and a command channel and outputting a mixed reality world anchor point. The method solves the problem of converting asynchronous six-degree-of-freedom object poses in a camera coordinate system into stable, world-consistent and recoverable real object anchor points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computers, and more specifically, to a method and system for asynchronous object pose stream frame alignment and anchoring oriented to perspective blending. Background Technology

[0002] In intelligent manufacturing and high-end equipment operation and maintenance scenarios, operators need to overlay operation prompts, status information, maintenance procedures, and risk markers onto real equipment, tools, or components. This information cannot simply be displayed in the field of vision; it must be stably attached to the real object. If slippage occurs between the virtual marker and the real object, the operator may misjudge the location of components, assembly sequence, or risk areas. For complex equipment, precision assembly, and emergency response, such deviations directly affect operational efficiency and safety. Therefore, stable mixed reality anchoring at the real object level is a problem that must be solved when related systems move from demonstration to field application.

[0003] Existing mixed reality devices typically rely on their own spatial positioning capabilities to generate planar anchor points, spatial anchor points, or scene anchor points. These anchor points are suitable for fixed walls, tabletops, or spatial areas, but not necessarily for real-world objects, especially those with complex shapes, that are movable, occluded, or require precise correspondence with a 3D model. Existing six-DOF object tracking methods can output the object's pose relative to the camera, but this pose itself is not an anchor point in the mixed reality world coordinate system. Transforming the camera coordinate pose into a stable world anchor point requires addressing issues such as acquisition timing, head-mounted display motion, coordinate system differences, network latency, and tracking failures.

[0004] In an external visual computing architecture, after a head-mounted device acquires an image, it needs to undergo image compression, network transmission, object segmentation, depth estimation, 3D registration, pose tracking, and result return. This process inherently involves latency. During this latency, the operator's head may have already moved significantly. If the system only reads the current head-mounted display pose and performs coordinate transformation when the pose result is returned, the camera pose used for the transformation will not be consistent with the camera pose at the time of image acquisition. This problem is more pronounced during rapid head turning, close-range observation, handheld objects, or low frame rate external inference.

[0005] Existing augmented reality (AR) or mixed reality (MVR) anchoring schemes often assume that perception, localization, and rendering occur within the same device runtime, or primarily focus on the preservation and restoration of static spatial points. Existing object pose estimation schemes, on the other hand, evaluate pose accuracy in the camera coordinate system. There is a lack of an intermediate mechanism for anchoring to real-world objects: this mechanism needs to know which frame the pose originates from, the world pose of the camera in the head-mounted device when that frame was captured, and, based on the perceived quality, whether to actually apply that pose to the MVR anchor.

[0006] In real-world scenarios, target objects are also affected by occlusion, lighting variations, reflections, motion blur, viewpoint changes, and depth estimation failures. If the system directly uses the pose of each frame to update virtual anchor points, a single error detection or pose jump could cause the virtual content to be instantly pasted into the wrong position. Using only ordinary smoothing filtering might slowly propagate the incorrect pose to the displayed results, or fail to provide an interpretable anchor point state when the target is briefly lost. Therefore, an anchor point control method is needed that can distinguish between "reliable update," "no update for now," "short-term hold," and "repositioning."

[0007] Therefore, it is necessary to propose an asynchronous six-DOF object pose flow frame alignment and anchoring method for perspective-based mixed reality. This method should link the image, the camera world pose at the acquisition time, the external visual pose result, and the mixed reality anchor point update, so that the externally calculated object pose can be transformed to the world coordinate system at the correct acquisition time, and maintain anchor point stability or trigger re-localization when perception is unreliable. Summary of the Invention

[0008] The purpose of this disclosure is to provide an asynchronous object pose flow frame alignment and anchoring method for perspective blending, which aims to solve the problems of how to convert the asynchronous six-degree-of-freedom object pose flow in the camera coordinate system into stable, world-consistent, and recoverable real object anchor points under the conditions of continuous motion of head-mounted perspective blending reality devices, computation and transmission delays of external vision servers, and low-frequency or intermittent failure of object pose results.

[0009] In one general aspect, an asynchronous object pose flow frame alignment and anchoring method for perspective blending is provided, comprising steps one through ten: Step 1: The head-mounted mixed reality device acquires binocular perspective images and generates frame numbers. Simultaneously record the frame number. Corresponding reference camera world pose ; Step two: The head-mounted mixed reality device sends the binocular images and camera calibration information to an external visual computing device via a data plane; Step 3: After receiving the image, the external visual computing device performs target segmentation, mask acquisition, depth estimation, and six-DOF pose estimation on the input image, and outputs the pose of the target object in the camera coordinate system. Segmentation mask Depth-aligned evidence and its reliability information; Step four, the external visual computing device will carry frame numbers. The pose result message is returned through the message plane. The pose result message includes at least the frame number, pose matrix, effective pose flag, reliability score, reliability sub-score and time information. Step 5, the head-mounted mixed reality device, based on the frame number... Query the reference camera world pose corresponding to the acquisition time from the frame pose history module. When the frame number is not hit The world anchor point update was rejected at this time; Step six: The head-mounted mixed reality device converts the camera coordinate system pose in the pose result message into the camera local coordinate system pose of the mixed reality engine, and combines it with the reference camera world pose to obtain the anchor point pose of the target object in the mixed reality world coordinate system. ; Step 7: The head-mounted mixed reality device determines whether to update the anchor point, maintain the anchor point, perform short-term battery life, or enter a repositioning state based on the reliability score, failure flag, and anchor point status. Step eight: The head-mounted mixed reality device receives commands to reset tracking, reacquire anchor points, pause, and resume via the command plane, and performs idempotent processing using request numbers to form closed-loop control. The anchor point pose satisfies: in, The world anchor point pose after frame alignment. To obtain the world pose of the camera at the time of data acquisition. This represents the target pose in the OpenCV camera coordinate system. For a fixed transformation from the Unity camera local coordinate system to the OpenCV camera coordinate system, This is a fixed transformation from the OpenCV camera coordinate system to the Unity camera local coordinate system.

[0010] Step 9: Determine whether to update the anchor point based on the reliability score and anchor point status; Step 10: Form a closed loop through state events and command channels, and output the mixed reality anchor point.

[0011] The data plane adopts a strategy of retaining only the latest data transmission, satisfying the following: Furthermore, the data plane employs a latest-value retention transmission strategy, satisfying the following: in, , , These are the sets of payload types for the data plane, message plane, and command plane, respectively. For binocular image payload, Calibrate the load for the camera. The load is the pose result. For state event payloads, For heartbeat load, To control the requested load, To control the response load, For the subject identifier in the data plane, For a moment The preceding belongs to the topic The set of data packets to be transmitted For the first in the set Data packets, For its arrival time, Theme At any moment The latest value is retained by the selected index. For at any time On the topic The actual number of data packets retained. and They are time points The actual consumption posture and heartbeat, A stream of state events preserved in the order of arrival.

[0012] The target segmentation includes image normalization, candidate region extraction, connected component filtering, edge consistency filtering, and area threshold filtering, and a target mask. The degree of preference satisfies: in, For the first Frame candidate mask set, For the first One candidate mask, The number of candidate masks, Candidate Mask Foreground area, For the set of indices of non-empty candidates in the foreground, To segment the raw candidate scores returned by the backend There is a marker for the candidate score. The scores of the candidates selected. For the selected single-target mask index, The threshold for mask binarization. This is the output single-target binary mask. For output size alignment operation, This is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise. This indicates that there is no valid mask output in the current frame.

[0013] The pose result message includes frame number, 4×4 object pose matrix in camera coordinate system, valid pose flag, pose source, reliability score, reliability flag, depth quality statistics and sending timestamp.

[0014] The specific method for querying the reference camera world pose at the acquisition time corresponding to the frame number from the frame pose history module is as follows: in, For frame pose history set, Define the frame number field for the historical set. To obtain accurate frame number lookup results, when At that time, the world coordinate anchoring update based on the current pose result is refused. This represents the maximum capacity of the frame attitude history cache.

[0015] The external visual computing device performs hierarchical scoring of pose quality, satisfying the following: in, For the current stage of the sensing pipeline, For a set of highly reliable stages, the components are... , , These correspond to the tracking, registration, and re-registration perception stages, respectively. For phased gating, To recently refuse to suppress scores, The total score for pose reliability is... For the overall quality score, For continuous high-quality frame confidence, For geometric subdivision, and These are the effective sub-fractions for color reprojection and depth alignment, respectively. For mask modulation, This is the lower limit coefficient of mask modulation. For masking subdivision, and Based on the weights, and The current effective weights, The lower bound of geometry, It is an interval cutoff function; As an effective indicator of the projected area ratio, and These are the observation mask area and the rendering projection area, respectively. The percentage of mask area. This is the threshold for area segmentation.

[0016] The effective subset of color reprojection and the effective subset of depth alignment satisfy the following: in, and These are the color reprojection score signal and the depth alignment score signal, respectively. To render depth status codes, This is a valid status code. To serve as a valid indicator for rendering depth state, and These are the depth in-point ratio and depth residual signal, respectively. A flag exists for rendering the depth signal. and These are valid indicators for the color and depth items, respectively. To determine the effective depth coverage within the mask, .

[0017] The recent rejection inhibition segment in the staged gating satisfies the following formula: in, To recently refuse to suppress scores, This is a recent rejection sign. To track rejection counts continuously in recent times, To suppress the lower limit, This is the score decay coefficient for a single rejection. This represents the maximum number of rejections that can be included in the suppression.

[0018] The local anchoring strategy module of the head-mounted mixed reality device includes a freely combinable motion model and an output strategy. The motion model forms control points for each received observation and provides prediction operators externally, satisfying the following relationship: As a specific implementation of a motion model, the constant velocity integral model satisfies: in, For the first This was the first time it was observed. For the first One control point, For the first Each control point timestamp For the first pose of control points and These are the translation and rotation components of the control point, respectively. For linear velocity, Angular velocity, For motion model parameters, To update the operator for control points, For the prediction operator, and These are the translation and attitude prediction components, respectively.

[0019] The output strategy includes three methods: zero-delay extrapolation fusion, delayed interpolation, and preservation of form. The zero-delay extrapolation fusion synthesizes the output by superimposing the seam residuals on the extrapolated poses of the control points, satisfying the following: in, Predict the extrapolated pose of the control point at the current moment. The seam residual is the seam residual, and the seam residual is attenuated by a factor in subsequent rendering frames. Exponential decay is performed; the delay interpolation satisfies: in, To subtract the interpolation target time after adaptive rendering latency, It is a cubic Hermite spline interpolation operator, and the interpolation operator independently limits the velocity tangent at the control point endpoints based on the translation and rotation chord length to eliminate overshoot during sudden stops. This is the final output pose.

[0020] The anchor point strategy module includes two types of control layers: static anchoring and low-resolution relocation. The static anchoring and low-resolution relocation satisfy the following: in, This is the cumulative duration of continuous static activity. For the static condition to be met, the threshold values ​​for the object's translational and rotational velocities are multiplied by an amplification factor adaptively determined by the head motion intensity. ; This is a local relocation trigger signal. For the current low score event The duration; To reacquire the trigger signal upstream, when local relocation is triggered and local geometric arbitration is performed. Below the geometrical disbelief threshold Time-triggered; the geometric arbitration division It is calculated by weighted log-geometric mean of reprojection score and depth validity score; This is the cumulative duration of continuous stillness in the previous moment. For the overall reliability score, The minimum score threshold for static conditions. and The translational and rotational velocities of the smoothed object. and As a reference, the threshold values ​​for translational and rotational velocities are used. For the observation update time interval, This is a relock suppression enable flag. This is the threshold for the duration of local relocation. To reduce repositioning cooling time, The last time a relocation was triggered. This indicates that a local motion model has been established.

[0021] The local control of the static anchoring consists of relock suppression, head stop freezing, dead zone, velocity escape, score-weighted cumulative sum, low-score release, absolute drift rope rental, missed lock creep, and seam attenuation, satisfying the following: in, The position residual after the dead zone, For CUSUM, accumulated and location evidence, its head stop freeze mark The residuals are iteratively accumulated over time. The updated locked output position is achieved through adaptive gain. The leaky lock creep mechanism is used to track the current observation position. Gradual convergence; To release the seam position residual after locking, in the release event Write the initial value of the seam residual relative to the current candidate output. And in subsequent rendering frames, according to the coefficient The final anchor point output pose is smoothly transitioned to the candidate output based on the seam residual decay after exponential decay. The current output position is locked. and These are the head motion tolerance and distance adaptive amplification factors, respectively. The width of the dead zone. This is based on the accumulation and location evidence from the previous moment. For accumulation and attenuation coefficients, For the observation update time interval, As a time normalization benchmark, This represents the total reliability score for current observations. This represents the residual position of the seam at the previous moment. Adaptive decay factor for rendering step size; release event Triggered when the accumulated evidence, absolute displacement drift, velocity escape duration, or low score duration exceeds the threshold.

[0022] For real objects, short-term occlusion and short-term motion do not always mean anchor point failure. The system sets a static anchoring layer, accumulates high-resolution, low-velocity, and low-angular-velocity observations into a continuous static duration based on the actual measurement time, and locks the system after reaching the dwell time threshold; if low-resolution observations persist, velocity escape, absolute drift, or accumulation and evidence continue to accumulate, the system releases the lock and enters a seam attenuation transition, and then returns to the candidate output.

[0023] In another general aspect, an asynchronous object pose flow frame alignment and anchoring system for perspective blending is provided, comprising eight modules: The head-mounted display acquisition module acquires left and right eye images and camera calibration information, and generates a frame number for each frame of image. The frame pose history module records the position and pose of the left eye image, right eye image, or center reference camera in the mixed reality world coordinate system when the image is acquired or transmitted. The external vision computing module outputs the six-degree-of-freedom pose of the target object in the visual algorithm coordinate system based on the binocular image, camera calibration, target segmentation results, depth information and target 3D model; The reliability assessment module generates a reliability score based on the target area, the effective depth ratio within the target area, pose changes between adjacent frames, reprojection consistency, and tracking failure status. The pose message module returns a pose result message, which includes at least the frame number, object pose matrix, whether a valid pose exists, pose source, reliability score, reliability flag, and time information. The frame alignment coordinate transformation module retrieves the reference camera world pose at the time of image acquisition according to the frame number in the pose result message, and converts the object pose into the anchor point pose in the mixed reality coordinate system. The anchor point strategy module outputs anchor point behaviors such as accepting updates, rejecting updates, maintaining anchor points, short-term prediction, tracking loss, or repositioning based on reliability scores, pose jumps, number of consecutive frames without valid poses, and server status. The command closed-loop module receives control commands such as reset tracking, reacquire anchor points, pause, and resume, and performs idempotent processing based on the request number to avoid duplicate commands causing multiple executions.

[0024] The technical effects to be achieved by the embodiments of the present invention are as follows: When converting external asynchronous six-DOF object pose streams into mixed reality anchors, the reference camera world pose at the moment of image acquisition is used instead of the head-mounted display pose at the moment the pose message arrives. This reduces virtual content slippage caused by head-mounted display motion, network transmission, and model inference latency, and improves the consistency of real object anchoring. High-bandwidth image data, low-bandwidth pose results, state events, and control commands are processed in different semantic channels. The latest data priority strategy is adopted for images and pose results, the state change process is preserved for state events, and control commands support request number deduplication and runtime sequential execution. This can reduce the backlog of old frames and improve the controllability of operations such as relocation, pause, and resume. By introducing perceptual reliability scoring and anchor point strategy control, the anchor point can be updated when the pose is reliable, remain stable when the pose is low or changes abruptly, make limited predictions when the pose is lost for a short time, and enter re-localization when the pose fails for a long time. This makes it more suitable for mixed reality use scenarios such as occlusion, rapid head movement, and objects briefly leaving the field of view. Attached Figure Description

[0025] The above and other objects and features of this disclosure will become clearer from the following description taken in conjunction with the accompanying drawings.

[0026] Figure 1 This is a schematic diagram illustrating a system architecture according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram illustrating the pose estimation flowchart of an external vision computing device according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram illustrating a reliability assessment and anchoring strategy state machine according to an embodiment of the present disclosure. Detailed Implementation

[0027] The following detailed embodiments are provided to aid the reader in gaining a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but may be changed as will become clear upon understanding this disclosure, except for operations that must occur in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.

[0028] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided only to illustrate some of the many feasible ways of implementing the methods, apparatus, and / or systems described herein, which will become clear upon understanding the disclosure of this application.

[0029] As used herein, the term “and / or” includes any one of the associated listed items and any combination of any two or more.

[0030] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, assemblies, regions, layers, or parts, these components, assemblies, regions, layers, or parts should not be limited by these terms. Rather, these terms are used only to distinguish one component, assembly, region, layer, or part from another. Thus, without departing from the teaching of the examples described herein, the first component, first assembly, first region, first layer, or first part referred to as the first component, first assembly, first region, first layer, or first part may also be referred to as the second component, second assembly, second region, second layer, or second part.

[0031] In the specification, when an element (such as a layer, region, or substrate) is described as being "on" another element, "connected to," or "bonded to" another element, the element may be directly "on" another element, directly "connected to," or "bonded to" the other element, or one or more other elements may be present in between. Conversely, when an element is described as being "directly on" another element, "directly connected to," or "directly bonded to" another element, no other elements may be present in between.

[0032] The terminology used herein is for the purpose of describing various examples only and is not intended to limit disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. The terms “comprising,” “including,” and “having” indicate the presence of the described features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0033] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains upon understanding this disclosure. Unless expressly defined herein, terms (such as those defined in a general dictionary) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and in this disclosure, and shall not be interpreted in an idealized or overly formalistic manner.

[0034] Furthermore, in the description of the examples, detailed descriptions of well-known related structures or functions will be omitted when it is believed that such detailed descriptions would lead to a vague interpretation of this disclosure.

[0035] Figure 1 This is a schematic diagram illustrating an asynchronous object pose stream frame alignment and anchoring method for perspective blending according to an embodiment of the present disclosure.

[0036] To achieve the aforementioned objectives, the present invention employs the following technical framework: Figure 1 As shown.

[0037] A frame alignment mixed reality anchoring method and system for asynchronous six-DOF object pose flow.

[0038] The method of the present invention includes the following steps: Step 1: The head-mounted mixed reality device acquires binocular perspective images and generates frame numbers. Simultaneously record the frame number. Corresponding reference camera world pose ; Step two: The head-mounted mixed reality device sends the binocular images and camera calibration information to an external visual computing device via a data plane; Step 3: After receiving the image, the external visual computing device performs target segmentation, mask acquisition, depth estimation, and six-DOF pose estimation on the input image, and outputs the pose of the target object in the camera coordinate system. Segmentation mask Depth-aligned evidence and its reliability information; Step four, the external visual computing device will carry frame numbers. The pose result message is returned through the message plane. The pose result message includes at least the frame number, pose matrix, effective pose flag, reliability score, reliability sub-score and time information. Step 5, the head-mounted mixed reality device, based on the frame number... Query the reference camera world pose corresponding to the acquisition time from the frame pose history module. When the frame number is not hit The world anchor point update was rejected at this time; Step six: The head-mounted mixed reality device converts the camera coordinate system pose in the pose result message into the camera local coordinate system pose of the mixed reality engine, and combines it with the reference camera world pose to obtain the anchor point pose of the target object in the mixed reality world coordinate system. ; Step 7: The head-mounted mixed reality device determines whether to update the anchor point, maintain the anchor point, perform short-term battery life, or enter a repositioning state based on the reliability score, failure flag, and anchor point status. Step eight: The head-mounted mixed reality device receives commands to reset tracking, reacquire anchor points, pause, and resume via the command plane, and performs idempotent processing using request numbers to form closed-loop control. The anchor point pose satisfies: in, The world anchor point pose after frame alignment. To obtain the world pose of the camera at the time of data acquisition. This represents the target pose in the OpenCV camera coordinate system. For a fixed transformation from the Unity camera local coordinate system to the OpenCV camera coordinate system, This is a fixed transformation from the OpenCV camera coordinate system to the Unity camera local coordinate system.

[0039] Step nine: The anchor point strategy module determines whether to update the anchor point based on the reliability score and anchor point status. When the reliability meets the threshold and the pose jump is within the allowable range, the system accepts the pose and updates the anchor point; when the reliability is low or an abnormal jump occurs, the system rejects the pose and maintains the last stable anchor point; when there is no pose for a short period but the historical state is reliable, the system can make a short-term prediction based on the most recent stable motion state; when there is no effective pose for a long time, the system enters the tracking loss or repositioning state and waits for the re-registration result, such as... Figure 2 As shown.

[0040] Step 10: The system forms a closed loop through status events and command channels. External visual computing devices issue status events such as "detecting," "tracking," "tracking lost," "paused," and "error." The head-mounted mixed reality device can send control commands such as "reset tracking," "reacquire anchor point," "pause," and "resume." Control commands carry a request number, and the system performs deduplication based on the request number to avoid duplicate commands causing multiple resets or repositioning. Specifically, as shown below... Figure 3 As shown.

[0041] The specific method for determining whether to update the anchor point is as follows: First, a perceptual reliability assessment is performed. Reliability scores and reliability flags are assigned based on input metrics such as mask area, effective depth ratio within the mask, pose jumps between adjacent frames, rendering reprojection consistency, and tracking failure status. When the reliability meets the threshold and the pose jump is within the allowable range, it is considered highly reliable and the pose jump is normal; therefore, the pose is accepted and the anchor point is updated. When the reliability is low or an abnormal jump occurs, the pose is rejected and the previous stable anchor point is maintained. When there is no pose for a short period but the historical state is reliable, a short-term prediction is made based on the most recent stable motion state. When there is no effective pose for a long period, the tracking is lost or repositioned, and the system waits for the re-registration result.

[0042] The specific method for forming a closed loop through state events and command channels is as follows: 1. Input left-eye image, right-eye image, camera calibration information, and target 3D model; 2. After intrinsic parameter mapping and image processing; 3. Perform target segmentation and mask acquisition; 4. Perform binocular depth estimation to obtain a depth map; 5. In the initial stage, perform registration, initial registration based on the mask, depth map, and 3D model to obtain a six-DOF pose, and perform continuous tracking in the tracking stage, performing pose tracking based on continuous frame images and depth information, and outputting the current frame pose; 6. Failure detection and reliability assessment, detecting state events such as tracking loss, pose jump, continuous mask loss, and low rendering reprojection consistency; 7. Determine whether the current pose is usable by detecting state events. If usable, generate a pose result message; if unusable, reset tracking and reacquire anchor points, returning to step 3, performing target segmentation and mask acquisition.

[0043] The data plane adopts a strategy of retaining only the latest data transmission, satisfying the following: Furthermore, the data plane employs a latest-value retention transmission strategy, satisfying the following: in, , , These are the sets of payload types for the data plane, message plane, and command plane, respectively. For binocular image payload, Calibrate the load for the camera. The load is the pose result. For state event payloads, For heartbeat load, To control the requested load, To control the response load, For the subject identifier in the data plane, For a moment The preceding belongs to the topic The set of data packets to be transmitted For the first in the set Data packets, For its arrival time, Theme At any moment The latest value is retained by the selected index. For at any time On the topic The actual number of data packets retained. and They are time points The actual consumption posture and heartbeat, A stream of state events preserved in the order of arrival.

[0044] The target segmentation includes image normalization, candidate region extraction, connected component filtering, edge consistency filtering, and area threshold filtering, and a target mask. The degree of preference satisfies: in, For the first Frame candidate mask set, For the first One candidate mask, The number of candidate masks, Candidate Mask Foreground area, For the set of indices of non-empty candidates in the foreground, To segment the raw candidate scores returned by the backend There is a marker for the candidate score. The scores of the candidates selected. For the selected single-target mask index, The threshold for mask binarization. This is the output single-target binary mask. For output size alignment operation, This is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise. This indicates that there is no valid mask output in the current frame.

[0045] The pose result message includes frame number, 4×4 object pose matrix in camera coordinate system, valid pose flag, pose source, reliability score, reliability flag, depth quality statistics and sending timestamp.

[0046] The specific method for querying the reference camera world pose at the acquisition time corresponding to the frame number from the frame pose history module is as follows: in, For frame pose history set, Define the frame number field for the historical set. To obtain accurate frame number lookup results, when At that time, the world coordinate anchoring update based on the current pose result is refused. This represents the maximum capacity of the frame attitude history cache.

[0047] The external visual computing device performs hierarchical scoring of pose quality, satisfying the following: in, For the current stage of the sensing pipeline, For a set of highly reliable stages, the components are... , , These correspond to the tracking, registration, and re-registration perception stages, respectively. For phased gating, To recently refuse to suppress scores, The total score for pose reliability is... For the overall quality score, For continuous high-quality frame confidence, For geometric subdivision, and These are the effective sub-fractions for color reprojection and depth alignment, respectively. For mask modulation, This is the lower limit coefficient of mask modulation. For masking subdivision, and Based on the weights, and The current effective weights, The lower bound of geometry, It is an interval cutoff function; As an effective indicator of the projected area ratio, and These are the observation mask area and the rendering projection area, respectively. The percentage of mask area. This is the threshold for area segmentation.

[0048] The effective subset of color reprojection and the effective subset of depth alignment satisfy the following: in, and These are the color reprojection score signal and the depth alignment score signal, respectively. To render depth status codes, This is a valid status code. To serve as a valid indicator for rendering depth state, and These are the depth in-point ratio and depth residual signal, respectively. A flag exists for rendering the depth signal. and These are valid indicators for the color and depth items, respectively. To determine the effective depth coverage within the mask, .

[0049] The recent rejection inhibition segment in the staged gating satisfies the following formula: in, To recently refuse to suppress scores, This is a recent rejection sign. To track rejection counts continuously in recent times, To suppress the lower limit, This is the score decay coefficient for a single rejection. This represents the maximum number of rejections that can be included in the suppression.

[0050] The local anchoring strategy module of the head-mounted mixed reality device includes a freely combinable motion model and an output strategy. The motion model forms control points for each received observation and provides prediction operators externally, satisfying the following relationship: As a specific implementation of a motion model, the constant velocity integral model satisfies: in, For the first This was the first time it was observed. For the first One control point, For the first Each control point timestamp For the first pose of control points and These are the translation and rotation components of the control point, respectively. For linear velocity, Angular velocity, For motion model parameters, To update the operator for control points, For the prediction operator, and These are the translation and attitude prediction components, respectively.

[0051] The output strategy includes three methods: zero-delay extrapolation fusion, delayed interpolation, and preservation of form. The zero-delay extrapolation fusion synthesizes the output by superimposing the seam residuals on the extrapolated poses of the control points, satisfying the following: in, Predict the extrapolated pose of the control point at the current moment. The seam residual is the seam residual, and the seam residual is attenuated by a factor in subsequent rendering frames. Exponential decay is performed; the delay interpolation satisfies: in, To subtract the interpolation target time after adaptive rendering latency, It is a cubic Hermite spline interpolation operator, and the interpolation operator independently limits the velocity tangent at the control point endpoints based on the translation and rotation chord length to eliminate overshoot during sudden stops. This is the final output pose.

[0052] The anchor point strategy module includes two types of control layers: static anchoring and low-resolution relocation. The static anchoring and low-resolution relocation satisfy the following: in, This is the cumulative duration of continuous static activity. For the static condition to be met, the threshold values ​​for the object's translational and rotational velocities are multiplied by an amplification factor adaptively determined by the head motion intensity. ; This is a local relocation trigger signal. For the current low score event The duration; To reacquire the trigger signal upstream, when local relocation is triggered and local geometric arbitration is performed. Below the geometrical disbelief threshold Time-triggered; the geometric arbitration division It is calculated by weighted log-geometric mean of reprojection score and depth validity score; This is the cumulative duration of continuous stillness in the previous moment. For the overall reliability score, The minimum score threshold for static conditions. and The translational and rotational velocities of the smoothed object. and As a reference, the threshold values ​​for translational and rotational velocities are used. For the observation update time interval, This is a relock suppression enable flag. This is the threshold for the duration of local relocation. To reduce repositioning cooling time, The last time a relocation was triggered. This indicates that a local motion model has been established.

[0053] The local control of the static anchoring consists of relock suppression, head stop freezing, dead zone, velocity escape, score-weighted cumulative sum, low-score release, absolute drift rope rental, missed lock creep, and seam attenuation, satisfying the following: in, The position residual after the dead zone, For CUSUM, accumulated and location evidence, its head stop freeze mark The residuals are iteratively accumulated over time. The updated locked output position is achieved through adaptive gain. The leaky lock creep mechanism is used to track the current observation position. Gradual convergence; To release the seam position residual after locking, in the release event Write the initial value of the seam residual relative to the current candidate output. And in subsequent rendering frames, according to the coefficient The final anchor point output pose is smoothly transitioned to the candidate output based on the seam residual decay after exponential decay. The current output position is locked. and These are the head motion tolerance and distance adaptive amplification factors, respectively. The width of the dead zone. This is based on the accumulation and location evidence from the previous moment. For accumulation and attenuation coefficients, For the observation update time interval, As a time normalization benchmark, This represents the total reliability score for current observations. This represents the residual position of the seam at the previous moment. Adaptive decay factor for rendering step size; release event Triggered when the accumulated evidence, absolute displacement drift, velocity escape duration, or low score duration exceeds the threshold.

[0054] For real objects, short-term occlusion and short-term motion do not always mean anchor point failure. The system sets a static anchoring layer, accumulates high-resolution, low-velocity, and low-angular-velocity observations into a continuous static duration based on the actual measurement time, and locks the system after reaching the dwell time threshold; if low-resolution observations persist, velocity escape, absolute drift, or accumulation and evidence continue to accumulate, the system releases the lock and enters a seam attenuation transition, and then returns to the candidate output.

[0055] The system includes the following 8 modules, such as Figure 2 As shown: The head-mounted display acquisition module is used to acquire left and right eye images and camera calibration information, and to generate a frame number for each frame of image. The frame pose history module is used to record the position and pose of the left, right, or center reference camera in the mixed reality coordinate system when the image is acquired or transmitted. The external vision computing module is used to output the six-degree-of-freedom pose of the target object in the visual algorithm coordinate system based on the binocular image, camera calibration, target segmentation results, depth information and target 3D model; The reliability assessment module is used to generate a reliability score based on the target area, the effective depth ratio within the target area, pose changes between adjacent frames, reprojection consistency, and tracking failure status. The pose message module is used to return pose result messages. These messages include at least the frame number, object pose matrix, whether a valid pose exists, pose source, reliability score, reliability flag, and time information. The frame alignment coordinate transformation module is used to look up the reference camera world pose at the time of image acquisition according to the frame number in the pose result message, and convert the object pose into the anchor point pose in the mixed reality coordinate system. The anchor point strategy module is used to output anchor point behaviors such as accept update, refuse update, maintain anchor point, short-term prediction, tracking loss or relocation based on reliability score, pose jump, number of consecutive frames without valid pose and server status. The command closed-loop module is used to receive control commands such as reset tracking, reacquire anchor points, pause and resume, and perform idempotent processing with request numbers to avoid repeated execution of duplicate commands.

[0056] While some embodiments of this disclosure have been shown and described, those skilled in the art will understand that modifications may be made to these embodiments without departing from the principles and spirit of this disclosure, which are defined by the claims and their equivalents.

Claims

1. A method for frame alignment and anchoring of asynchronous object pose streams oriented towards perspective blending, characterized in that, Including steps one through ten: Step 1: The head-mounted mixed reality device acquires binocular perspective images and generates frame numbers. Simultaneously record the frame number. Corresponding reference camera world pose ; Step two: The head-mounted mixed reality device sends the binocular images and camera calibration information to an external visual computing device via a data plane; Step 3: After receiving the image, the external visual computing device performs target segmentation, mask acquisition, depth estimation, and six-DOF pose estimation on the input image, and outputs the pose of the target object in the camera coordinate system. Segmentation mask Depth-aligned evidence and its reliability information; Step four, the external visual computing device will carry frame numbers. The pose result message is returned through the message plane. The pose result message includes at least the frame number, pose matrix, effective pose flag, reliability score, reliability sub-score and time information. Step 5, the head-mounted mixed reality device, based on the frame number... Query the reference camera world pose corresponding to the acquisition time from the frame pose history module. When the frame number is not hit The world anchor point update was rejected at this time; Step six: The head-mounted mixed reality device converts the camera coordinate system pose in the pose result message into the camera local coordinate system pose of the mixed reality engine, and combines it with the reference camera world pose to obtain the anchor point pose of the target object in the mixed reality world coordinate system. ; Step 7: The head-mounted mixed reality device determines whether to update the anchor point, maintain the anchor point, perform short-term battery life, or enter a repositioning state based on the reliability score, failure flag, and anchor point status. Step 8: The head-mounted mixed reality device receives commands to reset tracking, reacquire anchor points, pause, and resume via the command plane, and performs idempotent processing using request numbers to form closed-loop control; the anchor point pose satisfies: in, The world anchor point pose after frame alignment. To obtain the world pose of the camera at the time of data acquisition. This represents the target pose in the OpenCV camera coordinate system. For a fixed transformation from the Unity camera local coordinate system to the OpenCV camera coordinate system, This is a fixed transformation from the OpenCV camera coordinate system to the Unity camera local coordinate system; Step 9: Determine whether to update the anchor point based on the reliability score and anchor point status; Step 10: Form a closed loop through state events and command channels, and output the mixed reality anchor point.

2. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 1, characterized in that, The data plane adopts a strategy of retaining only the latest data transmission, satisfying the following: Furthermore, the data plane employs a latest-value retention transmission strategy, satisfying the following: in, , , These are the sets of payload types for the data plane, message plane, and command plane, respectively. For binocular image payload, Calibrate the load for the camera. The load is the pose result. For state event payloads, For heartbeat load, To control the requested load, To control the response load, For the subject identifier in the data plane, For a moment The preceding belongs to the topic The set of data packets to be transmitted For the first in the set Data packets, For its arrival time, Theme At any moment The latest value is retained by the selected index. For at any time On the topic The actual number of data packets retained. and They are time points The actual consumption posture and heartbeat, A stream of state events preserved in the order of arrival.

3. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 1, characterized in that, The target segmentation includes image normalization, candidate region extraction, connected component filtering, edge consistency filtering, and area threshold filtering, and a target mask. The degree of preference satisfies: in, For the first Frame candidate mask set, For the first One candidate mask, The number of candidate masks, Candidate Mask Foreground area, For the set of indices of non-empty candidates in the foreground, To segment the raw candidate scores returned by the backend There is a marker for the candidate score. The scores of the candidates selected. For the selected single-target mask index, The threshold for mask binarization. This is the output single-target binary mask. For output size alignment operation, This is an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise. This indicates that there is no valid mask output in the current frame.

4. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 1, characterized in that, The pose result message includes frame number, 4×4 object pose matrix in camera coordinate system, valid pose flag, pose source, reliability score, reliability flag, depth quality statistics and sending timestamp.

5. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 1, characterized in that, The specific method for querying the reference camera world pose at the acquisition time corresponding to the frame number from the frame pose history module is as follows: in, For frame pose history set, Define the frame number field for the historical set. To obtain accurate frame number lookup results, when At that time, the world coordinate anchoring update based on the current pose result is refused. This represents the maximum capacity of the frame attitude history cache.

6. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 1, characterized in that, The external visual computing device performs hierarchical scoring of pose quality, satisfying the following: in, For the current stage of the sensing pipeline, For a set of highly reliable stages, the components are... , , These correspond to the tracking, registration, and re-registration perception stages, respectively. For phased gating, To recently refuse to suppress scores, The total score for pose reliability is... For the overall quality score, For continuous high-quality frame confidence, For geometric subdivision, and These are the effective sub-fractions for color reprojection and depth alignment, respectively. For mask modulation, This is the lower limit coefficient of mask modulation. For masking subdivision, and Based on the weights, and The current effective weights, The lower bound of geometry, It is an interval cutoff function; As an effective indicator of the projected area ratio, and These are the observation mask area and the rendering projection area, respectively. The percentage of mask area. This is the threshold for area segmentation.

7. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 6, characterized in that, The effective subset of color reprojection and the effective subset of depth alignment satisfy the following: in, and These are the color reprojection score signal and the depth alignment score signal, respectively. To render depth status codes, This is a valid status code. To serve as a valid indicator for rendering depth state, and These are the depth in-point ratio and depth residual signal, respectively. A flag exists for rendering the depth signal. and These are valid indicators for the color and depth items, respectively. To determine the effective depth coverage within the mask, .

8. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 6, characterized in that, The recent rejection inhibition segment in the staged gating satisfies the following formula: in, To avoid suppressing the score in the near future, This is a recent rejection sign. To track rejection counts continuously in recent times, To suppress the lower limit, This is the score decay coefficient for a single rejection. This represents the maximum number of rejections that can be included in the suppression.

9. The asynchronous object pose flow frame alignment and anchoring method for perspective blending as described in claim 1, characterized in that, The local anchoring strategy module of the head-mounted mixed reality device includes a freely combinable motion model and an output strategy. The motion model forms control points for each received observation and provides prediction operators externally, satisfying the following relationship: As a specific implementation of the motion model, the constant velocity integral model satisfies: in, For the first This was the first time it was observed. For the first One control point, For the first Each control point timestamp For the first pose of control points and These are the translation and rotation components of the control point, respectively. Linear velocity, Angular velocity, For motion model parameters, To update the operator for control points, For the prediction operator, and These are the translation and attitude prediction components, respectively.

10. An asynchronous object pose flow frame alignment and anchoring system for perspective blending, employing the method described in any one of claims 1-9, comprising eight modules: The head-mounted display acquisition module acquires left and right eye images and camera calibration information, and generates a frame number for each frame of image. The frame pose history module records the position and pose of the left eye image, right eye image, or center reference camera in the mixed reality world coordinate system when the image is acquired or transmitted. The external vision computing module outputs the six-degree-of-freedom pose of the target object in the visual algorithm coordinate system based on the binocular image, camera calibration, target segmentation results, depth information and target 3D model; The reliability assessment module generates a reliability score based on the target area, the effective depth ratio within the target area, pose changes between adjacent frames, reprojection consistency, and tracking failure status. The pose message module returns a pose result message, which includes at least the frame number, object pose matrix, whether a valid pose exists, pose source, reliability score, reliability flag, and time information. The frame alignment coordinate transformation module retrieves the reference camera world pose at the time of image acquisition according to the frame number in the pose result message, and converts the object pose into the anchor point pose in the mixed reality coordinate system. The anchor point strategy module outputs anchor point behaviors such as accepting updates, rejecting updates, maintaining anchor points, short-term prediction, tracking loss, or repositioning based on reliability scores, pose jumps, number of consecutive frames without valid poses, and server status. The command closed-loop module receives control commands such as reset tracking, reacquire anchor points, pause, and resume, and performs idempotent processing based on the request number to avoid duplicate commands causing multiple executions.