Delay measurement method and system, calibration method and system, delay test method and system, electronic device, storage medium, and program product

By acquiring the visual pose sequence and motion pose sequence of the augmented reality device through a synchronous motion platform and a vision platform, the problem of delay measurement in continuous composite 3D motion of the augmented reality device is solved, achieving accurate delay measurement and improving user experience.

WO2026098360A1PCT designated stage Publication Date: 2026-05-15SUNNY OPTICAL ZHEJIANG RES INST CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SUNNY OPTICAL ZHEJIANG RES INST CO LTD
Filing Date
2025-10-31
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing extended reality devices have difficulty accurately measuring motion-to-light delay (MTP delay) in continuous complex three-dimensional human motion, resulting in a difference between visual motion perception and vestibular motion perception, which affects the user's immersion and causes dizziness.

Method used

A time delay measurement system and method are provided. By synchronizing the motion of a motion platform with a vision platform, the system acquires the visual pose sequence and motion pose sequence of the extended reality device under test. The system then uses a processing platform to perform time synchronization and coordinate system transformation to accurately measure the MTP delay.

Benefits of technology

It enables precise delay measurement of extended reality devices in continuous composite 3D motion, improving user immersion, reducing dizziness, and comprehensively evaluating the device's latency performance under different motion states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025131880_15052026_PF_FP_ABST
    Figure CN2025131880_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A delay measurement method and system, a calibration method and system, a delay test method and system, an electronic device, a storage medium, and a program product. The delay measurement system comprises a motion platform, a vision platform, and a processing platform. The motion platform is fixedly connected to an extended reality device to be measured and drives said extended reality device to move synchronously. The vision platform is fixedly connected to said extended reality device, and acquires a plurality of continuous images of calibration plates displayed by said extended reality device. The processing platform obtains a visual pose sequence of said extended reality device on the basis of the plurality of continuous images, and determines a delay of said extended reality device on the basis of the visual pose sequence and a motion pose sequence of the motion platform.
Need to check novelty before this filing date? Find Prior Art

Description

Delay measurement methods and systems, calibration methods and systems, delay test methods and systems, electronic devices, storage media, and software products.

[0001] Related applications

[0002] This application claims priority to Chinese patent applications filed on November 8, 2024, with application number 202411587590.2 entitled "Delay Measurement Method and System, Electronic Device, Storage Medium, and Program Product"; applications filed on November 8, 2024, with application number 202411587570.5 entitled "Calibration Method and System, Electronic Device, Storage Medium, and Program Product"; and applications filed on November 8, 2024, with application number 202411594876.3 entitled "Delay Test Method, Delay Detection Equipment, Electronic Device, and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of extended reality device technology, and in particular to a delay measurement method and system, a calibration method and system, a delay test method and system, electronic devices, storage media, and program products. Background Technology

[0004] With the development of science and technology, higher requirements have been placed on the performance of extended reality devices such as virtual reality devices, augmented reality devices, and mixed reality devices.

[0005] Taking head-mounted displays (HMDs) as an example, augmented reality devices need to respond to the user's head movements in real time. Although the user's head movements can be measured by various sensors, due to limitations in transmission and computation time, the HMD cannot immediately display the corresponding image to the user, resulting in image latency. This latency causes a difference between visual motion perception and vestibular motion perception, which not only reduces the sense of presence of the augmented reality device but also increases the burden on the brain and causes dizziness and nausea. Motion-to-photon (MTP) latency includes the time taken from the start of the user's movement to the corresponding image being displayed on the screen. It characterizes the delay time between the image seen by the user in an augmented reality device such as an HMD and the user's head movement, and can quantify the degree of matching between visual observation and user head movement. Therefore, for users, the smaller the MTP latency, the better the user's immersion; the larger the MTP latency, the stronger the user's dizziness.

[0006] Considering that human head movements are mostly continuous three-dimensional (three-axis) composite movements, for extended display devices that support motion prediction algorithms, analyzing the delay changes in continuous composite three-dimensional movements is a key indicator for evaluating their motion prediction algorithms.

[0007] Therefore, how to accurately measure the MTP latency of extended reality devices during continuous composite three-dimensional motion of the human body has become one of the research hotspots in the field of extended reality device technology. Summary of the Invention

[0008] According to various embodiments of this application, a delay measurement method and system, a calibration method and system, a delay test method and system, an electronic device, a storage medium, and a program product are provided.

[0009] In a first aspect, embodiments of this application provide a time delay measurement system, comprising: a motion platform fixedly connected to an extended reality device under test (ARD) and driving the ARD to move synchronously; a vision platform fixedly connected to the ARD and acquiring multiple consecutive images of a target displayed on the ARD; and a processing platform for obtaining a visual pose sequence of the ARD based on the multiple consecutive images, and for determining the time delay of the ARD based on the visual pose sequence and the motion pose sequence of the motion platform.

[0010] Secondly, embodiments of this application provide a delay measurement method, the method comprising: acquiring a motion pose sequence of a motion platform, wherein the motion platform is fixedly connected to an extended reality device under test and drives the extended reality device under test to move synchronously; acquiring multiple consecutive images of a target displayed on the extended reality device under test; obtaining a visual pose sequence of the extended reality device under test based on the multiple consecutive images; and determining the delay of the extended reality device under test based on the visual pose sequence and the motion pose sequence.

[0011] Thirdly, embodiments of this application provide a calibration system, which includes: a motion component fixedly connected to an extended reality device and driving the extended reality device to move synchronously; a camera component fixedly connected to the extended reality device and acquiring multiple images of a target displayed by the extended reality device during synchronous movement; and a processing component calibrating the extrinsic parameters of the camera component based on the multiple images and the motion pose sequence of the motion component.

[0012] Fourthly, embodiments of this application provide a calibration method, which includes: acquiring a motion pose sequence of a motion component, wherein the motion component is fixedly connected to an extended reality device and drives the extended reality device to move synchronously; using a camera component fixedly connected to the extended reality device, acquiring multiple images of a target displayed on the extended reality device during the synchronous movement; and calibrating the extrinsic parameters of the camera component based on the multiple images and the motion pose sequence of the motion component.

[0013] Fifthly, embodiments of this application provide a delay testing method, the method comprising: in response to receiving a first action sequence, obtaining a first motion trajectory of a simulated component within a preset time period based on the first action sequence, wherein the first action sequence includes first time information and first pose information; obtaining a second action sequence based on a display image of a display device during the movement of the simulated component according to the first motion trajectory, and determining a second motion trajectory within a preset time period based on the second action sequence, wherein the second action sequence includes second time information and second pose information; and determining a delay of the display device within the preset time period based on the first motion trajectory and the second motion trajectory.

[0014] Sixthly, embodiments of this application provide a delay detection device, which includes a simulation component, a drive motor, a camera module, and a controller. The simulation component is used to simulate user actions. The drive motor is used to drive the simulation component to move along a first motion trajectory, wherein the first motion trajectory is determined based on a first action sequence including first time information and first pose information. The camera module is used to acquire display images of the display device under test during the movement of the simulation component along the first motion trajectory. The controller is communicatively connected to the drive motor and the camera module, wherein the controller is configured to: determine a second motion trajectory for a preset time period based on the display images, and determine the delay of the display device under test during the preset time period based on the first motion trajectory and the second motion trajectory.

[0015] In a seventh aspect, embodiments of this application provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement, when executed, the delay measurement method of the second aspect, the calibration method of the fourth aspect, or the delay test method of the fifth aspect.

[0016] Eighthly, embodiments of this application provide a non-transient computer-readable storage medium storing computer instructions that enable a computer to perform, when executed, a delay measurement method as described in the second aspect, a calibration method as described in the fourth aspect, or a delay test method as described in the fifth aspect.

[0017] Ninthly, embodiments of this application provide a computer program product including a computer program, which, when executed by a processor, can implement the delay measurement method of the second aspect, the calibration method of the fourth aspect, or the delay test method of the fifth aspect.

[0018] It should be understood that the descriptions in this section are not intended to identify key or important features of the embodiments of this application, nor are they intended to limit the scope of this application. Other features of this application will become readily apparent from the descriptions below. Attached Figure Description

[0019] Other features, objects, and advantages relating to the embodiments of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. Wherein:

[0020] Figure 1 shows a schematic diagram of a delay measurement system according to an embodiment of this application.

[0021] Figure 2 shows a schematic diagram of an extended reality device under test according to an embodiment of this application.

[0022] Figure 3 shows a schematic diagram of a reference plate according to an embodiment of this application.

[0023] Figure 4 shows a schematic diagram of one of a plurality of consecutive images displayed by an extended reality device under test according to an embodiment of the present application.

[0024] Figure 5 shows a schematic diagram of a processing platform according to an embodiment of this application.

[0025] Figure 6 shows a schematic diagram of a processing platform according to an embodiment of this application.

[0026] Figure 7 illustrates a flowchart of the process for determining the delay of the extended reality device under test by segmenting the visual pose sequence and the motion pose sequence according to an embodiment of this application.

[0027] Figure 8 shows a schematic diagram of statistical results including a first statistical indicator and a second statistical indicator according to an embodiment of this application.

[0028] Figure 9 is a flowchart of a delay measurement method provided in an embodiment of this application.

[0029] Figure 10 shows a schematic diagram of a calibration system according to an embodiment of this application.

[0030] Figure 11 shows a schematic diagram of a motion component according to an embodiment of this application.

[0031] Figure 12 shows a schematic diagram of the structure of the processing component according to an embodiment of this application.

[0032] Figure 13 illustrates a schematic diagram of the process of a camera assembly acquiring multiple images according to an embodiment of this application.

[0033] Figure 14 shows a schematic diagram of the structure of the time correction unit according to an embodiment of this application.

[0034] Figure 15 shows a time correction flowchart according to an embodiment of this application.

[0035] Figure 16 shows a schematic diagram of aligning the first angular velocity and the second angular velocity using a cross-correlation method according to an embodiment of this application.

[0036] Figure 17 shows a schematic flowchart of camera component extrinsic parameter calibration according to an embodiment of this application.

[0037] Figure 18 shows a flowchart illustrating the calibration of extrinsic parameters of a camera component based on multiple images and a motion pose sequence of a motion component according to an embodiment of this application.

[0038] Figure 19 shows a schematic diagram of motion information of the motion component and the calibrated camera component in the same coordinate system according to an embodiment of this application.

[0039] Figure 20 is a flowchart of a calibration method provided in an embodiment of this application.

[0040] Figure 21 illustrates an exemplary system framework that can be applied to the delay testing method according to this application.

[0041] Figure 22 illustrates another exemplary system framework that can be applied to the delay testing method according to this application.

[0042] Figure 23 shows a schematic flowchart of a delay test method according to an exemplary embodiment of this application.

[0043] Figure 24 shows a schematic diagram of the attitude angle according to an exemplary embodiment of this application.

[0044] Figure 25 shows a schematic diagram of the distribution of test points according to an exemplary embodiment of this application.

[0045] Figure 26 shows a schematic diagram of a first motion trajectory and a second motion trajectory according to an exemplary embodiment of this application.

[0046] Figure 27 shows a schematic diagram of linear fitting of the first motion trajectory within region G in Figure 6.

[0047] Figure 28 shows a schematic diagram of linear fitting of the second motion trajectory within region G in Figure 6.

[0048] Figure 29 shows a schematic diagram of the fitted first motion trajectory and the second motion trajectory according to an exemplary embodiment of this application.

[0049] Figure 30 shows a schematic diagram of the delay curve of a display device according to an exemplary embodiment of this application.

[0050] Figure 31 illustrates an interactive schematic diagram of a display device and a delay detection device according to an exemplary embodiment of this application.

[0051] Figure 32 illustrates an interaction diagram of a display device, a delay detection device, and a server according to an exemplary embodiment of this application.

[0052] Figure 33 shows a schematic diagram of the structure of an electronic device according to an exemplary embodiment of this application. Detailed Implementation

[0053] To better understand this application, various aspects of this application will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely illustrative of embodiments of this application and are not intended to limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0054] In this specification, the terms "first," "second," etc., are used only to distinguish one feature from another and do not imply any limitation on the features. Therefore, without departing from the teachings of this application, the first computational subunit discussed below may also be referred to as the second computational subunit, and the first time-pose curve may also be referred to as the second time-pose curve.

[0055] Expressions such as “comprising,” “including,” “having,” “containing,” and / or “comprising” are open-ended rather than closed expressions in this specification, indicating the presence of the stated features, elements, and / or components, but not excluding the presence of one or more other features, elements, components, and / or combinations thereof. Additionally, the word “exemplarily” is intended to refer to an example or illustration.

[0056] Unless otherwise specified, all terms used in this application (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. Terms (e.g., those defined in common dictionaries) shall be understood to have the meaning consistent with their meaning in the context of the relevant art, and shall not be interpreted in an idealized or overly formal sense, unless expressly so specified in this application.

[0057] Unless otherwise specified, the embodiments and features described in this application can be combined with each other. Furthermore, unless explicitly limited or contradicted by the context, the specific steps included in the methods described in this application are not limited to the order in which they are described, but can be performed in any order or in parallel. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0058] Figure 1 shows a schematic diagram of a delay measurement system 1000 according to an embodiment of the present application. Figure 2 shows a schematic diagram of an extended reality device under test 400 according to an embodiment of the present application. Figure 3 shows a schematic diagram of a reference panel 410 according to an embodiment of the present application. Figure 4 shows a schematic diagram of one of a plurality of consecutive images displayed by the extended reality device under test 400 according to an embodiment of the present application.

[0059] As shown in Figures 1-4, this application embodiment provides a time delay measurement system 1000, which may include a motion platform 100, a vision platform 200, and a processing platform 300. The motion platform 100 is fixedly connected to the extended reality device under test (ARD) 400 and drives the ARD 400 to move synchronously. The vision platform 200 is also fixedly connected to the ARD 400 and acquires multiple consecutive images of a target 410 displayed by the ARD 400, such as image 420 shown in Figure 4. The processing platform 300 obtains the visual pose sequence of the ARD 400 based on the multiple consecutive images, and determines the time delay of the ARD 400 based on the visual pose sequence and the motion pose sequence of the motion platform 100.

[0060] According to at least one embodiment of the latency measurement system provided in this application, a motion platform that moves synchronously with the extended reality device under test (ARD) can simulate continuous composite three-dimensional motion of the human body. Therefore, it can comprehensively evaluate the latency performance of the ARD under different motion states and observe the dynamic changes of the ARD in continuous composite three-dimensional motion. Using a vision platform, multiple consecutive images of a target displayed by the ARD under test can be acquired, and the visual pose sequence of the ARD under test can be determined. By comparing the difference between the visual pose sequence of the ARD under test and the motion pose sequence of the motion platform, the MTP latency of the ARD in continuous composite three-dimensional motion of the human body can be accurately measured.

[0061] Specifically, in some embodiments of this application, the extended reality device 400 under test may include a virtual reality device, an augmented reality device, or a mixed reality device. It should be understood that the delay measurement system 1000 can also be applied to delay measurement scenarios of other display devices, and this is not limited here.

[0062] The motion platform 100 is fixedly connected to the extended reality device under test (AMD) 400, for example, through a rigid connection. Furthermore, the motion platform 100 can drive the AMD 400 to move synchronously. The motion platform 100 can include any suitable structural components and can perform various translational and rotational movements in the world coordinate system. Taking the AMD 400 as an example, the motion platform 100, fixedly connected to the AMD, can simulate horizontal movement, horizontal rotation, tilting motion, etc., of the human body.

[0063] The vision platform 200 can be used to simulate a scenario where the human eye receives a series of images displayed by the extended reality device 400 under test. Optionally, the vision platform 200 may include at least one of a camera and a video camera, such as an IR camera (infrared camera), an RGB camera (red, green, and blue camera), or a monochrome camera. Furthermore, the vision platform 200 may also include other structures such as a photoelectric sensor; this application does not limit the specific internal structure of the vision platform 200.

[0064] Optionally, the calibration board 410 may be a virtual calibration board located within the extended reality device 400 under test. The virtual calibration board may include visual marker calibration boards from a visual benchmark library, such as ApriTag, ArUco, or other QR code calibration boards. Using a virtual calibration board for delay measurement simplifies the structure of the delay measurement system and reduces the difficulty of implementing delay measurements.

[0065] The processing platform 300 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processing platform 300 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The processing platform 300 obtains a visual pose sequence of the extended reality device under test 400 based on multiple consecutive images, and determines the delay of the extended reality device under test 400 based on the visual pose sequence and the motion pose sequence of the motion platform 100.

[0066] Optionally, the processing platform 300 may include various suitable communication buses and input / output interfaces for connecting to the motion platform 100 and the vision platform 200.

[0067] Referring again to FIG1, in some embodiments of this application, the delay measurement system 1000 further includes a time platform 500, which may include a clock source, wherein the clock source is used to time synchronize the motion platform 100 and the vision platform 200 before and during the synchronized motion.

[0068] Specifically, the clock source can provide unified time synchronization to the motion platform 100 and the vision platform 200 via the PTP (Precision Time Protocol). Using a high-precision and high-accuracy clock source as the master clock, timestamps are sent to the motion platform 100 and the vision platform 200, which act as slave clocks. The slave clocks compare the timestamps with their local time and adjust the time accordingly, thereby ensuring that the motion pose sequence obtained from the motion platform 100 is synchronized with the visual pose sequence obtained from the vision platform 200.

[0069] Optionally, referring to Figures 1 and 2, the visual pose sequence may include multiple first time-pose information items arranged in chronological order. The first time-pose information includes first time information and first attitude angles corresponding to the first time information, wherein the first attitude angles include a first roll angle, a first yaw angle, and a first pitch angle. For example, the form of the first time-pose information may include: timestamp_roll_pitch_yaw, where 000_10_20_30 can represent the roll, pitch, and yaw angles at the start time as (10°, 20°, 30°), respectively. In the three-dimensional right-handed Cartesian coordinate system, roll is the roll angle around the z-axis; pitch is the pitch angle around the x-axis; and yaw is the yaw angle around the y-axis.

[0070] The motion pose sequence may include multiple second time-pose information items arranged in chronological order. The second time-pose information may include second time information and corresponding second attitude angles, where the second attitude angles include a second roll angle, a second yaw angle, and a second pitch angle. Similarly, the second time-pose information may also be in the form of a timestamp: `_roll_pitch_yaw`.

[0071] To compare motion posture changes in the real world with those in the virtual world of the extended reality device 400 under test, these two types of motion posture changes need to be unified under the same coordinate system. Motion changes in the virtual world can be determined in a first coordinate system (e.g., camera coordinate system) with reference to the visual platform 200, while motion changes in the real world can be determined in a second coordinate system (e.g., world coordinate system) with reference to the motion platform 100. Therefore, coordinate system transformation calibration of the motion platform 100 and the visual platform 200 is required.

[0072] Because different extended reality devices under test have different eye-fitting distances, the position of the visual platform 200 needs to be adjusted to adapt to the extended reality device 400 before measurement. This causes a change in the relative position of the visual platform 200 and the motion platform 100, rendering the original conversion relationship inapplicable and requiring recalibration. To solve this problem, a feasible solution is to calibrate the extrinsic parameters of the camera or video camera in the visual platform 200 by photographing the virtual target displayed by the extended reality device under test using the visual platform 200.

[0073] Figure 5 shows a schematic diagram of a processing platform 300 according to an embodiment of this application.

[0074] Specifically, referring to Figures 1 and 5, in some embodiments of this application, the processing platform 300 may include a first conversion unit 320, a second conversion unit 330, and a calibration unit 350. The first conversion unit 320 can extract the pixel coordinates of feature points of the target plate 410 (as shown in Figure 3) based on multiple consecutive images acquired by the vision platform 200. The second conversion unit 330 can convert the pixel coordinates into a preliminary pose sequence in a first coordinate system referenced to the vision platform 200 according to preset spatial geometric constraints between feature points. The calibration unit 350 converts the preliminary pose sequence into a visual pose sequence in a second coordinate system referenced to the motion platform 100.

[0075] Optionally, the processing platform 300 may further include an image processing unit 310. The image processing unit 310 can process multiple consecutive images before extracting the pixel coordinates of feature points of the target plate 410. This processing includes at least one of downsampling and Gaussian blurring. Since moiré patterns can lead to inaccurate or even impossible extraction of 2D feature points when capturing the screen of the extended reality device 400 under test, it is necessary to process the acquired multiple consecutive images. Downsampling and Gaussian blurring, among other methods, can reduce the impact of moiré patterns on the images.

[0076] Furthermore, the processing platform 300 may also include a time correction unit 340, which can perform time correction on the initial pose sequence before converting the initial pose sequence into a visual pose sequence. Although the timestamps of the images of the target and the trajectory timestamps of the motion platform are from the same clock source, the image capture process has a certain time delay due to the presence of MTP, thus requiring time correction.

[0077] For clarity, Figure 5 shows the processing platform 300, which includes, in optional cases, an image processing unit 310 and a time correction unit 340. It can be distinguished that the optional image processing unit 310 or time correction unit 340 is shown with a dashed box, and dashed connecting lines are used to show the connections between the image processing unit 310 or time correction unit 340 and other units when the image processing unit 310 or time correction unit 340 is included.

[0078] Optionally, the fixed connection between the motion platform 100 and the extended reality device under test 400 can be considered as a rigid connection of rigid bodies, and the fixed connection between the vision platform 200 and the extended reality device under test 400 can also be considered as a rigid connection of rigid bodies. Based on the principle of rotational invariance of rigid bodies, time correction can be performed on the initial pose sequence before converting it into a visual pose sequence. For example, after filtering the angular velocity of the pose, the time delay of the image capture process caused by the presence of MTP can be obtained based on cross-correlation, and the relevant data can be time-corrected and aligned.

[0079] Figure 6 shows a schematic diagram of the processing platform 300 according to an embodiment of this application. Figure 7 shows a schematic flowchart of segmented processing of visual pose sequences and motion pose sequences to determine the delay of the extended reality device under test according to an embodiment of this application. Figures 26-30 show schematic flowcharts of generating a first time-pose curve and a second time-pose curve according to an embodiment of this application.

[0080] As shown in Figures 1, 6, 7, and 26-30, in some embodiments of this application, the processing platform 300 further includes a splitting unit 370, a fitting unit 380, and a delay unit 390. The splitting unit 370 splits the visual pose sequence and the motion pose sequence into multiple motion segments based on the motion characteristic parameters of the synchronous motion between the motion platform 100 and the extended reality device under test 400 (as shown in Figure 2). The fitting unit 380 fits the multiple motion segments of the visual pose sequence into multiple first time-pose curves and the multiple motion segments of the motion pose sequence into multiple second time-pose curves. The delay unit 390 determines the delay of the extended reality device under test 400 based on the difference between the first time-pose curves and the second time-pose curves that correspond to the first time-pose curves in time.

[0081] Specifically, relevant MTP latency measurements can only obtain the latency at the start or end of a motion, and cannot measure the latency during the motion process. In other words, relevant MTP measurements can be divided into manually calculated latency and automatically calculated latency. Manually calculated latency uses human-estimated values ​​to calculate the latency of the extended reality device, resulting in significant errors and failing to objectively reflect the true latency of the extended display device. Automated latency calculation measures the latency of the extended reality device through measuring equipment. The key is to capture changes in the motion state of the extended reality device and monitor changes in the image displayed on the screen. By comparing the delays of these two aspects, the latency of the extended reality device is determined. However, automated latency calculation is limited to measuring the latency generated at a single motion node, such as the start or end of a motion. In actual use cases, the motion of the extended reality device and the response of the virtual reality scene within it are continuous processes. The latency generated during this process is a set of dynamic and continuous data. Therefore, measuring only the latency generated at a single motion node cannot reflect the dynamic changes in latency of the extended reality device during continuous motion. Especially considering that human movement is mostly continuous three-dimensional composite motion, for extended display devices that support motion prediction algorithms, analyzing the delay changes in continuous composite three-dimensional motion is a key indicator for evaluating their motion prediction algorithms.

[0082] According to at least one embodiment of the latency measurement system provided in this application, the motion platform, which moves synchronously with the extended reality device under test (ARD), can simulate continuous composite three-dimensional motion of the human body. Therefore, it can comprehensively evaluate the latency performance of the ARD under different motion states and observe the dynamic changes of the ARD in continuous composite three-dimensional motion. Using the vision platform, multiple consecutive images of a target displayed by the ARD under test can be acquired, and the visual pose sequence of the ARD under test can be determined. By comparing the difference between the visual pose sequence of the ARD under test and the motion pose sequence of the motion platform, the MTP latency of the ARD in continuous composite three-dimensional motion of the human body can be accurately measured.

[0083] To enhance the above effects and improve the accuracy and precision of MTP delay measurement, the visual pose sequence and motion pose sequence can be segmented separately. Using motion segment similarity analysis, the visual pose sequence and motion pose sequence can be divided into multiple motion segments. By analyzing each of the multiple motion segments, fitting line segments one by one, and statistically analyzing the "gap" mean of the fitted first time-pose curve and second time-pose curve, the MTP delay of the extended reality device in continuous composite three-dimensional motion of the human body can be accurately measured.

[0084] Optionally, during motion segment similarity analysis, the motion characteristic parameters of the synchronous motion may include at least one of the velocity, acceleration, and derivative of the acceleration of the synchronous motion. Therefore, appropriate synchronous motion characteristic parameters can be selected based on the type of extended reality device under test and actual needs, or based on the characteristics of the synchronous motion, to accurately measure the MTP delay of the extended reality device during continuous composite three-dimensional motion of the human body. This application does not limit the specific content of the motion characteristic parameters of the synchronous motion.

[0085] Furthermore, in the process of using motion segment similarity analysis to decompose the visual pose sequence and the motion pose sequence into multiple motion segments, the motion feature parameters of the synchronous motion between the motion platform 100 and the extended reality device under test 400 can be used to decompose the visual pose sequence and the motion pose sequence into multiple motion segments. For example, at least one of the velocity, acceleration, and derivative of the synchronous motion can be regarded as fluctuation reference data, and a predetermined threshold can be set. If at least one of the velocity, acceleration, and derivative of the synchronous motion is greater than its corresponding predetermined threshold, the fluctuation of this part of the data can be considered too large. Based on this, in the visual pose sequence or the motion pose sequence, the previous time-pose information can be identified as the end point of the previous motion segment, and the next time-pose information can be identified as the starting point of the next motion segment.

[0086] Optionally, as shown in Figure 6, the processing platform 300 may further include a preprocessing unit 360. The preprocessing unit 360 may perform zero-bias preprocessing on the visual pose sequence and the motion pose sequence respectively before fitting multiple motion segments of the visual pose sequence and multiple motion segments of the motion pose sequence (as shown in the "Preprocessing" flow in Figure 7). Considering error factors such as equipment assembly errors and jitter errors, performing zero-bias preprocessing on the visual pose sequence and the motion pose sequence separately can improve the accuracy and precision of MTP delay measurement.

[0087] For clarity, Figure 6 shows the processing platform 300, which includes the preprocessing unit 360 in an optional configuration. It can be distinguished that the optional preprocessing unit 360 is shown with a dashed box, and dashed connecting lines are used to show the connections between the preprocessing unit 360 and other units when the preprocessing unit 360 is included.

[0088] In the process of segment-by-segment analysis and line segment fitting of multiple motion segments, fitting can be based on straight lines or quadratic curves. This application does not limit the method of line segment fitting. Subsequently, the "gap" mean can be statistically analyzed on the fitted first time-pose curve and second time-pose curve to accurately measure the MTP delay of the extended reality device in continuous composite three-dimensional motion of the human body.

[0089] Figures 26-30 illustrate the process of generating a first time-pose curve and a second time-pose curve according to embodiments of this application. In Figure 26, the first motion trajectory represents the visual pose sequence, and the second motion trajectory represents the motion pose sequence. Therefore, Figure 26 shows the attitude angle changes of the visual pose sequence and the motion pose sequence over time. Using motion segment similarity analysis, the visual pose sequence and the motion pose sequence can be divided into multiple motion segments based on the motion characteristic parameters of the synchronous motion of the motion platform and the extended reality device under test. For example, Figure 26 shows a motion segment of the visual pose sequence and a motion segment of the motion pose sequence that have a temporal correspondence, both indicated by dashed box A. Figures 27 and 28 respectively illustrate the process of linear fitting of the two motion segments within dashed box A in Figure 26, wherein the first motion trajectory and the second motion trajectory in Figures 27 and 28 respectively represent a motion segment of the visual pose sequence and a motion segment of the motion pose sequence that temporally corresponds to a motion segment of the visual pose sequence. In Figure 29, the fitted first and second motion trajectories respectively represent the attitude angle changes of a motion segment in the visual pose sequences of Figures 27 and 28 over time after fitting, as well as the attitude angle changes of a motion segment in the motion pose sequences of the aforementioned figures over time after fitting. Alternatively, the fitted first and second motion trajectories in Figure 29 represent the fitted first and second time-pose curves, respectively. By performing "gap" averaging on the first and second time-pose curves, the delay curve shown in Figure 30 can be obtained. This delay curve accurately represents the MTP delay of the extended reality device during the entire process of continuous composite three-dimensional human motion.

[0090] In addition, to further improve the accuracy and precision of MTP delay measurement, expand the application areas of delay measurement systems, and optimize application scenarios, multi-dimensional delay evaluation indicators can be used to characterize the results of MTP delay measurement.

[0091] Figure 8 shows a schematic diagram of statistical results including a first statistical indicator and a second statistical indicator according to an embodiment of this application.

[0092] Optionally, referring to Figures 6 and 8, the delay unit 390 may further include a first calculation subunit 391. The first calculation subunit 391 determines a first statistical index based on the difference between a first time-pose curve and a second time-pose curve that corresponds to the first time-pose curve in time. The first statistical index may include at least one of the following: delay extreme value, delay average value, delay median, and delay standard deviation of the extended reality device under test 400 (as shown in Figure 2).

[0093] Furthermore, the delay unit 390 may also include a second calculation subunit 392. The second calculation subunit 392 may determine a second statistical index based on at least one of the mean and standard deviation of a first statistical index obtained from multiple measurements.

[0094] The latency evaluation index characterizing the results of MTP latency measurement may include at least one of a first statistical index and a second statistical index. The first statistical index includes latency extremes, average latency, median latency, and standard deviation, which can be used to measure the latency of the extended reality device under test (ARTD) in a single measurement process. The second statistical index is at least one of the mean and standard deviation of the first statistical index from multiple measurements, and can be used to measure the stability of the latency of the ARTD from a macroscopic or overall perspective. Figure 8 shows the statistical results of six MTP latency measurements performed on an ARTD device 400 according to at least one embodiment of this application. Users can select appropriate latency evaluation indexes from the above multi-dimensional latency evaluation indexes to characterize the results of MTP latency measurement based on the type of their ARTD device and its application field.

[0095] Therefore, according to at least one embodiment of the delay measurement system provided in this application, the motion platform moving synchronously with the extended reality device under test (ARD) can simulate continuous composite three-dimensional motion of the human body, thus enabling a comprehensive evaluation of the delay performance of the ARD under different motion states and allowing observation of the dynamic changes of the ARD in continuous composite three-dimensional motion. Multiple consecutive images of a target displayed by the ARD under test can be acquired using the vision platform, thereby determining the visual pose sequence of the ARD under test. By comparing the difference between the visual pose sequence of the ARD under test and the motion pose sequence of the motion platform, the MTP delay of the ARD in continuous composite three-dimensional motion of the human body can be accurately measured.

[0096] Figure 9 is a flowchart of a delay measurement method 2000 provided in an embodiment of this application, wherein the delay measurement method 2000 may include the following steps:

[0097] S1, acquire the motion pose sequence of the motion platform, wherein the motion platform is fixedly connected to the extended reality device under test and drives the extended reality device under test to move synchronously.

[0098] S2, acquire multiple consecutive images of a target displayed via the extended reality device under test.

[0099] S3, obtain the visual pose sequence of the extended reality device under test based on multiple consecutive images.

[0100] S4, determine the delay of the extended reality device under test based on the visual pose sequence and the motion pose sequence.

[0101] The following will describe in detail each step of the above-described delay measurement method 2000 with reference to the accompanying drawings.

[0102] Step S1

[0103] Referring to Figures 1, 2, and 9, in some embodiments of this application, the motion platform 100 is fixedly connected to the extended reality device under test (AMD) 400, for example, through a rigid connection. Furthermore, the motion platform 100 can drive the AMD 400 to move synchronously. The motion platform 100 can include any suitable structural components and can perform various translational and rotational movements in a world coordinate system. Taking the AMD 400 as an example, the motion platform 100 fixedly connected to the AMD can simulate horizontal movement, horizontal rotation, tilting motion, etc., of the human body.

[0104] Optionally, the vision platform 200 acquires multiple consecutive images of the target plate 410 (as shown in FIG3) displayed by the extended reality device under test 400. The delay measurement method 2000 further includes: time synchronization of the motion platform 100 and the vision platform 200 before and during the synchronous motion.

[0105] For example, a clock source is used to synchronize the motion platform 100 and the vision platform 200. During the time synchronization process, the motion platform 100 and the vision platform 200 can be uniformly timed via the PTP protocol. A high-precision and high-accuracy clock source is used as the master clock, sending timestamps to the motion platform 100 and the vision platform 200, which act as slave clocks. The slave clocks compare the timestamps with their local time and adjust the time according to the difference, thereby ensuring that the motion pose sequence obtained based on the motion platform 100 is synchronized with the visual pose sequence obtained based on the vision platform 200.

[0106] Step S2

[0107] Referring to Figures 1-4 and Figure 9, in some embodiments of this application, a vision platform 200 can be used to acquire multiple consecutive images of a target 410 displayed via an extended reality device under test 400, for example, Figure 4 shows one of the multiple consecutive images, image 420.

[0108] The vision platform 200 can be used to simulate a scenario where the human eye receives a series of images displayed by the extended reality device 400 under test. Optionally, the vision platform 200 may include at least one of a camera and a video camera, such as an IR camera (infrared camera), an RGB camera (red-green-blue camera), or a monochrome camera. Furthermore, the vision platform 200 may also include other structures such as a photoelectric sensor; this application does not limit the specific internal structure of the vision platform 200.

[0109] Optionally, the calibration board 410 may be a virtual calibration board located within the extended reality device 400 under test. The virtual calibration board may include visual marker calibration boards from a visual benchmark library, such as ApriTag, ArUco, or other QR code calibration boards. Using a virtual calibration board for delay measurement simplifies the structure of the delay measurement system and reduces the difficulty of implementing delay measurements.

[0110] Step S3

[0111] Referring to Figures 1, 3, and 4, in some embodiments of this application, step S3, obtaining the visual pose sequence of the extended reality device under test based on multiple consecutive images, may include, for example,: extracting the pixel coordinates of feature points of the target plate 410 based on multiple consecutive images; converting the pixel coordinates into a preliminary pose sequence in a first coordinate system with the visual platform 200 that captures multiple consecutive images as a reference, according to the preset spatial geometric constraints between the feature points; and converting the preliminary pose sequence into a visual pose sequence in a second coordinate system with the motion platform 100 as a reference.

[0112] Specifically, the processing platform 300 can be used to obtain a visual pose sequence of the extended reality device under test based on multiple consecutive images. The processing platform 300 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processing platform 300 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc.

[0113] The processing platform 300 obtains the visual pose sequence of the extended reality device under test 400 based on multiple consecutive images, and determines the delay of the extended reality device under test 400 based on the visual pose sequence and the motion pose sequence of the motion platform 100. Optionally, the processing platform 300 may include various suitable communication buses and input / output interfaces for connecting to the motion platform 100 and the vision platform 200.

[0114] Furthermore, before extracting the pixel coordinates of feature points of the target plate 410 based on multiple consecutive images, the multiple consecutive images can be processed. This processing may include at least one of downsampling processing and Gaussian blur processing. Since the presence of moiré patterns when photographing the screen of the extended reality device 400 under test can lead to inaccurate or even impossible extraction of 2D feature points, it is necessary to process the acquired multiple consecutive images. Downsampling processing and Gaussian blur processing, among other methods, can reduce the impact of moiré patterns on the images.

[0115] Furthermore, in order to compare the changes in motion posture in the real world with the changes in motion posture in the virtual world of the extended reality device 400 under test, it is necessary to unify these two types of motion posture changes under the same coordinate system. The motion changes in the virtual world can be determined in a first coordinate system (e.g., the camera coordinate system) with the visual platform 200 as the reference, while the motion changes in the real world can be determined in a second coordinate system (e.g., the world coordinate system) with the motion platform 100 as the reference. Therefore, it is necessary to calibrate the motion platform 100 and the visual platform 200 by coordinate system transformation.

[0116] Since the eye-fitting distances of different extended reality devices under test vary, the position of the visual platform 200 needs to be adjusted to adapt to the extended reality device 400 under test before measurement. This further leads to a change in the relative position of the visual platform 200 and the motion platform 100, causing the original conversion relationship to no longer apply and requiring recalibration.

[0117] Optionally, the initial pose sequence can be time-corrected before being converted into a visual pose sequence. Although the timestamps of the image of the target and the trajectory timestamps of the motion platform are from the same clock source, the image acquisition process has a certain time delay due to the presence of MTP, thus requiring time correction.

[0118] For example, the fixed connection between the motion platform 100 and the extended reality device under test 400 can be considered a rigid connection of rigid bodies, and the fixed connection between the vision platform 200 and the extended reality device under test 400 can also be considered a rigid connection of rigid bodies. Based on the principle of rigid body rotation invariance, time correction can be performed on the initial pose sequence before converting it into a visual pose sequence. For example, after filtering the angular velocity of the pose, the time delay of the image capture process caused by the presence of MTP is obtained based on cross-correlation, and the relevant data is time-corrected and aligned.

[0119] Step S4

[0120] Referring to Figures 1, 2, 7, and 26-30, step S4, which determines the delay of the extended reality device under test based on the visual pose sequence and the motion pose sequence, may include, for example, the following: dividing the visual pose sequence and the motion pose sequence into multiple motion segments according to the synchronous motion feature parameters; fitting the multiple motion segments of the visual pose sequence into multiple first time-pose curves, and fitting the multiple motion segments of the motion pose sequence into multiple second time-pose curves; and determining the delay of the extended reality device under test based on the difference between the first time-pose curves and the second time-pose curves that correspond to the first time-pose curves in time.

[0121] Optionally, the visual pose sequence may include multiple first time-pose information items arranged in chronological order. The first time-pose information includes first time information and first attitude angles corresponding to the first time information, wherein the first attitude angles include a first roll angle, a first yaw angle, and a first pitch angle. For example, the form of the first time-pose information may include: timestamp_roll_pitch_yaw, where 000_10_20_30 can represent the roll, pitch, and yaw angles at the start time as (10°, 20°, 30°), respectively. In a three-dimensional right-handed Cartesian coordinate system, roll is the roll angle around the z-axis; pitch is the pitch angle around the x-axis; and yaw is the yaw angle around the y-axis.

[0122] The motion pose sequence may include multiple second time-pose information items arranged in chronological order. The second time-pose information may include second time information and corresponding second attitude angles, where the second attitude angles include a second roll angle, a second yaw angle, and a second pitch angle. Similarly, the second time-pose information may also be in the form of a timestamp: `_roll_pitch_yaw`.

[0123] Relevant MTP latency measurements can only obtain the latency at the start or end of a motion, and cannot measure the latency during the motion process. In other words, relevant MTP measurements can be divided into manually calculated latency and automatically calculated latency. Manually calculated latency uses human-estimated values ​​to calculate the latency of the extended reality device, resulting in significant errors and failing to objectively reflect the true latency of the extended display device. Automated latency calculation measures the latency of the extended reality device using measuring equipment. Its key is to capture changes in the motion state of the extended reality device and monitor changes in the image displayed on the screen. By comparing the delays of these two measurements, the latency of the extended reality device is determined. However, automated latency calculation is limited to measuring the latency generated at a single motion node, such as the start or end of a motion. In actual use cases, the motion of the extended reality device and the response of the virtual reality scene within it are continuous processes. The latency generated during this process is a set of dynamic and continuous data. Therefore, measuring only the latency generated at a single motion node cannot reflect the dynamic changes in latency of the extended reality device during continuous motion. Especially considering that human movement is mostly continuous three-dimensional composite motion, for extended display devices that support motion prediction algorithms, analyzing the delay changes in continuous composite three-dimensional motion is a key indicator for evaluating their motion prediction algorithms.

[0124] According to at least one embodiment of the latency measurement method provided in this application, a motion platform that moves synchronously with the extended reality device under test (ARD) can simulate continuous composite three-dimensional motion of the human body. Therefore, it is possible to comprehensively evaluate the latency performance of the ARD under different motion states and observe the dynamic changes of the ARD in continuous composite three-dimensional motion. Based on multiple consecutive images of a target displayed by the ARD under test, the visual pose sequence of the ARD under test can be determined. By comparing the difference between the visual pose sequence of the ARD under test and the motion pose sequence of the motion platform, the MTP latency of the ARD in continuous composite three-dimensional motion of the human body can be accurately measured.

[0125] To enhance the above effects and improve the accuracy and precision of MTP delay measurement, the visual pose sequence and motion pose sequence can be segmented separately. Using motion segment similarity analysis, the visual pose sequence and motion pose sequence can be divided into multiple motion segments. By analyzing each of the multiple motion segments, fitting line segments one by one, and statistically analyzing the "gap" mean of the fitted first time-pose curve and second time-pose curve, the MTP delay of the extended reality device in continuous composite three-dimensional motion of the human body can be accurately measured.

[0126] Optionally, during motion segment similarity analysis, the motion characteristic parameters of the synchronous motion may include at least one of the velocity, acceleration, and derivative of the acceleration of the synchronous motion. Therefore, appropriate synchronous motion characteristic parameters can be selected based on the type of extended reality device under test and actual needs, or based on the characteristics of the synchronous motion, to accurately measure the MTP delay of the extended reality device during continuous composite three-dimensional motion of the human body. This application does not limit the specific content of the motion characteristic parameters of the synchronous motion.

[0127] Furthermore, in the process of using motion segment similarity analysis to decompose the visual pose sequence and the motion pose sequence into multiple motion segments, the motion feature parameters of the synchronous motion between the motion platform 100 and the extended reality device under test 400 can be used to decompose the visual pose sequence and the motion pose sequence into multiple motion segments. For example, at least one of the velocity, acceleration, and derivative of the synchronous motion can be regarded as fluctuation reference data, and a predetermined threshold can be set. If at least one of the velocity, acceleration, and derivative of the synchronous motion is greater than its corresponding predetermined threshold, the fluctuation of this part of the data can be considered too large. Based on this, in the visual pose sequence or the motion pose sequence, the previous time-pose information can be identified as the end point of the previous motion segment, and the next time-pose information can be identified as the starting point of the next motion segment.

[0128] Furthermore, in some embodiments of this application, considering error factors such as equipment assembly error and jitter error, zero-bias preprocessing can be performed on the visual pose sequence and the motion pose sequence before fitting multiple motion segments of the visual pose sequence and multiple motion segments of the motion pose sequence, respectively.

[0129] In the process of segment-by-segment analysis and line segment fitting of multiple motion segments, fitting can be based on straight lines or quadratic curves. This application does not limit the method of line segment fitting. Subsequently, the "gap" mean can be statistically analyzed on the fitted first time-pose curve and second time-pose curve to accurately measure the MTP delay of the extended reality device during the entire motion process of continuous composite three-dimensional human motion.

[0130] To further improve the accuracy and precision of MTP delay measurement, expand the application areas of delay measurement systems, and optimize application scenarios, multi-dimensional delay evaluation indicators can be used to characterize the results of MTP delay measurement.

[0131] The latency evaluation index characterizing the results of MTP latency measurement may include at least one of a first statistical index and a second statistical index. The first statistical index includes latency extremes, average latency, median latency, and standard deviation, which can be used to measure the latency of the extended reality device under test (ARTD) in a single measurement process. The second statistical index is at least one of the mean and standard deviation of the first statistical index from multiple measurements, and can be used to measure the stability of the latency of the ARTD from a macroscopic or overall perspective. Figure 8 shows the statistical results of six MTP latency measurements performed on an ARTD device 400 according to at least one embodiment of this application. Users can select appropriate latency evaluation indexes from the above multi-dimensional latency evaluation indexes to characterize the results of MTP latency measurement based on the type of their ARTD device and its application field.

[0132] Therefore, according to the delay measurement method provided in at least one embodiment of this application, the motion platform moving synchronously with the extended reality device under test (ARD) can simulate continuous composite three-dimensional motion of the human body, comprehensively evaluate the delay performance of the ARD under different motion states, and observe the dynamic changes of the ARD in continuous composite three-dimensional motion. Based on multiple consecutive images of the target displayed by the ARD under test, the visual pose sequence of the ARD under test can be determined. By comparing the difference between the visual pose sequence of the ARD under test and the motion pose sequence of the motion platform, the MTP delay of the ARD in continuous composite three-dimensional motion of the human body can be accurately measured.

[0133] Figure 10 shows a schematic diagram of a calibration system 3000 according to an embodiment of the present application. Figure 2 shows a schematic diagram of an extended reality device 400 according to an embodiment of the present application. Figure 3 shows a schematic diagram of a pointer 410 according to an embodiment of the present application. Figure 4 shows a schematic diagram of one of a plurality of images displayed by the extended reality device 400 according to an embodiment of the present application, namely image 420.

[0134] As shown in Figures 10, 2, 3, and 4, this application embodiment provides a calibration system 3000, which may include a motion component 600, a camera component 700, and a processing component 800. The motion component 600 is fixedly connected to an extended reality device 400 and drives the extended reality device 400 to move synchronously. The camera component 700 is also fixedly connected to the extended reality device 400 and, during synchronous movement, acquires multiple images of a target board 410 displayed via the extended reality device 400; for example, Figure 4 shows one of the multiple images, image 420. The processing component 800 determines the extrinsic parameters of the camera component 700 based on the multiple images and the motion pose sequence of the motion component 600. The extrinsic parameters of the camera component 700 can be understood as parameters describing the position and orientation of the camera component 700 in a world coordinate system referenced to the motion component 600, and may include rotation matrices and translation vectors.

[0135] According to at least one embodiment of the calibration system provided in this application, a motion component that moves synchronously with an augmented reality device can simulate continuous composite three-dimensional motion of the human body. Furthermore, a camera component can acquire multiple images of a target displayed via the augmented reality device. Therefore, by using the multiple images of the target displayed via the augmented reality device acquired by the camera component and the motion pose sequence of the motion component, the extrinsic parameters of the camera component can be obtained. This unifies the motion information of the camera component and the motion component under the same coordinate system, providing a coordinate system reference for measuring the delay of the augmented reality device in continuous three-dimensional composite motion.

[0136] Specifically, in some embodiments of this application, the extended reality device 400 may include a virtual reality device, an augmented reality device, or a mixed reality device. It should be understood that the calibration system 3000 may also be applied to calibration scenarios of other display devices, and this is not limited here.

[0137] The motion component 600 is fixedly connected to the extended reality device 400, for example, through a rigid connection. Furthermore, the motion component 600 can drive the extended reality device 400 to move synchronously. The motion component 600 can include any suitable structural components; for example, it can include a robotic arm and a controller to control the robotic arm's movement. The motion component 600 can perform various translational and rotational movements in a world coordinate system. Taking the extended reality device 400 as an HMD as an example, the motion component 600, fixedly connected to the HMD, can simulate the rotational movements of the human head in the sagittal, coronal, and horizontal planes, etc.

[0138] The camera assembly 700 can be used to simulate a scenario where the human eye receives an image displayed by the extended reality device 400. Optionally, the camera assembly 700 may include at least one of a camera and a video camera. The camera assembly may include an IR camera (Infrared Camera), an RGB camera (Red Green Blue Camera), or a monochrome camera, etc. In addition, the camera assembly 700 may also include other structures such as a photoelectric sensor. This application does not limit the specific internal structure of the camera assembly 700.

[0139] Optionally, the calibration board 410 may be a virtual calibration board located within the extended reality device 400. The virtual calibration board may include visual marker calibration boards from a visual benchmark library, such as ApriTag, ArUco, and other QR code calibration boards. Using a virtual calibration board simplifies the structure of the calibration system, reduces the difficulty of calibration implementation, and decreases resource consumption.

[0140] Processing component 800 can be a variety of general-purpose and / or special-purpose processing units with processing and computing capabilities. Some examples of processing component 800 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processing component 800 determines the extrinsic parameters of camera component 700 based on multiple images acquired by camera component 700 and the motion pose sequence of motion component 600.

[0141] Optionally, the processing component 800 may include various suitable communication buses and input / output interfaces for connecting to the motion component 600 and the camera component 700.

[0142] Referring again to FIG10, in some embodiments of this application, the calibration system 3000 further includes a time component 900, which may include a clock source for synchronizing the motion component 600 and the camera component 700 before and during the synchronized motion.

[0143] Specifically, the clock source can provide unified time synchronization to the motion component 600 and the camera component 700 via the PTP (Precision Time Protocol). Using a high-precision and high-accuracy clock source as the master clock, timestamps are sent to the motion component 600 and the camera component 700, which act as slave clocks. The slave clocks compare the timestamps with their local time and adjust the time accordingly, thereby ensuring that the motion pose sequence obtained subsequently based on the motion component 600 is synchronized with the time of the multiple images acquired by the camera component 700.

[0144] Figure 11 shows a schematic diagram of a motion component 600 according to an embodiment of the present application. Figure 12 shows a structural schematic diagram of a processing component 800 according to an embodiment of the present application.

[0145] Referring to Figures 10, 2-4, 11, and 12, in some embodiments of this application, the processing component 800 may include: a first conversion unit 320, a second conversion unit 330, and a calibration unit 350. The first conversion unit 320 can extract the pixel coordinates of feature points of the target plate 410 based on multiple images of the target plate 410 displayed via the extended reality device 400, acquired by the camera component 700 during synchronous motion. The second conversion unit 330 converts the pixel coordinates into a visual pose sequence according to preset spatial geometric constraints between feature points. The calibration unit 350 calibrates the extrinsic parameters of the camera component 700 based on the visual pose sequence and the motion pose sequence.

[0146] Optionally, the processing component 800 may further include an image processing unit 310. The image processing unit 310 may process the multiple images before extracting the pixel coordinates of feature points of the target 410 based on the multiple images. The processing may include at least one of downsampling and Gaussian blurring. Since moiré patterns can lead to inaccurate or even impossible extraction of 2D feature points when the extended reality device 400 is photographed using the camera component 700, the acquired multiple images need to be processed. Downsampling and Gaussian blurring, among other methods, can reduce the impact of moiré patterns on the images.

[0147] Furthermore, the processing component 800 may also include an adjustment unit 360, which can adjust the positional relationship between the camera component 700 and the extended reality device 400 so that multiple feature points specified in the target plate 410 are displayed in multiple images. This can improve the generation quality of multiple images and ensure high stability and high accuracy of the calibration results.

[0148] As shown in Figure 5, as an option, the motion component 600 may include a robotic arm, which can be fixedly connected to the augmented reality device 400 and drive the augmented reality device 400 to move synchronously. A predetermined motion trajectory can be designed and encoded into the robotic arm, so that the robotic arm moves the augmented reality device 400 in various directions according to the predetermined motion trajectory. Furthermore, by optimizing the aforementioned predetermined motion trajectory, while ensuring the activation of various rotations and accelerations of the augmented reality device 400 in various directions, the blurring of the calibration plate image caused by excessively high movement speed can be reduced. This ensures high stability and high accuracy of the calibration results.

[0149] Figure 13 shows a schematic diagram of the process of a camera assembly 700 acquiring multiple images according to an embodiment of this application.

[0150] As shown in Figures 10, 3, and 13, the camera assembly 700 can acquire multiple images of the target 410 displayed via the extended reality device 400 using various acquisition modes. At the start of data acquisition for multiple images, the target (e.g., a virtual target) 410 located in the extended reality device 400 can be opened. Then, a data acquisition mode can be selected. Alternatively, the camera assembly 700 can acquire images of the target 410 displayed via the extended reality device 400 only when the extended reality device 400 passes a waypoint during synchronous motion. Alternatively, the camera assembly 700 can continuously acquire images of the target 410 displayed via the extended reality device 400 throughout the entire synchronous motion. The multiple image data acquired by the camera assembly 700, along with the triaxial encoder data of the motion component 600, can then be stored as a motion pose sequence of the motion component 600.

[0151] A waypoint can be understood as a series of pre-set key locations during the movement of a robotic arm, which the arm needs to pass through sequentially to complete the task. In other words, multiple key locations can be selected as waypoints when a predetermined motion trajectory is set. Multiple waypoints can fully activate the movement of the extended reality device 400 in various directions, including various rotational and acceleration motion information. Therefore, as an option, during synchronous movement, images of the target board 410 displayed by the extended reality device 400 can be captured only when the extended reality device 400 passes through a waypoint.

[0152] In this scenario, synchronous motion is discrete motion between different waypoints. The camera component 700 is stationary at a waypoint, capturing images. For example, synchronous motion might remain stationary at a waypoint for several seconds before moving to the next waypoint. Therefore, the moment the camera component 700 captures the image is the same as the moment the extended reality device 400 (or motion component 600) synchronously moves to the corresponding waypoint, and no correction is needed for the acquisition time.

[0153] In another alternative, the camera assembly 700 can continuously capture images of the target 410 displayed via the extended reality device 400 throughout the entire synchronous motion, where the synchronous motion is continuous. Therefore, although the timestamps of the images captured by the camera assembly 700 and the trajectory timestamps of the motion assembly 600 (or the extended reality device) are from the same clock source, the image capture process has a certain time delay due to the presence of the MTP, thus requiring time correction.

[0154] Optionally, referring again to FIG12, the processing component 800 may further include: a time correction unit 340, which can perform time correction on the visual pose sequence converted from pixel coordinates by the second conversion unit 330 according to the preset spatial geometric constraint relationship between feature points.

[0155] For clarity, Figure 12 shows the processing assembly 800, which optionally includes the image processing unit 310, time correction unit 340, and adjustment unit 360. It can be distinguished that the optional image processing unit 310, time correction unit 340, or adjustment unit 360 is shown with dashed boxes, and dashed connecting lines are used to show the connections of the image processing unit 310, time correction unit 340, or adjustment unit 360 to other units when they are included.

[0156] Figure 14 shows a schematic diagram of the time correction unit 340 according to an embodiment of the present application. Figure 15 shows a flowchart of the time correction process according to an embodiment of the present application.

[0157] As shown in Figures 10 and 14, in some embodiments of this application, the time correction unit 340 may include: an angular velocity calculation subunit 343, a cross-correlation subunit 344, and a correction subunit 345. The angular velocity calculation subunit 343 may determine a first angular velocity based on the motion pose sequence of the motion component 600, and determine a second angular velocity based on a visual pose sequence converted from multiple images acquired by the camera component 700. The cross-correlation subunit 343 may align the first and second angular velocities based on a cross-correlation method to determine the time offset of the visual pose sequence. The correction subunit 345 may perform time correction on the visual pose sequence based on the time offset.

[0158] Referring to Figures 10, 2, 14, and 15, the visual pose sequence may include multiple first-time-pose information items arranged in chronological order. Each first-time-pose information item includes first-time information and corresponding first attitude angles, where the first attitude angles include a first roll angle, a first yaw angle, and a first pitch angle. For example, the form of the first-time-pose information may include: timestamp_roll_pitch_yaw, where 000_10_20_30 can represent the initial roll, pitch, and yaw angles as (10°, 20°, 30°), respectively. In a three-dimensional right-handed Cartesian coordinate system, roll is the roll angle around the z-axis; pitch is the pitch angle around the x-axis; and yaw is the yaw angle around the y-axis.

[0159] The motion pose sequence may include multiple second time-pose information items arranged in chronological order. The second time-pose information may include second time information and corresponding second attitude angles, where the second attitude angles include a second roll angle, a second yaw angle, and a second pitch angle. Similarly, the second time-pose information may also be in the form of a timestamp: `_roll_pitch_yaw`.

[0160] Optionally, the fixed connection between the motion component 600 and the extended reality device 400 can be considered as a rigid connection of a rigid body, and the fixed connection between the camera component 700 and the extended reality device 400 can also be considered as a rigid connection of a rigid body. Based on the principle of the rotational invariance of a rigid body, the time offset can be obtained by aligning the first angular velocity of the rotation of the motion platform 100 with the second angular velocity estimated from the attitude of the camera component 700.

[0161] Figure 16 shows a schematic diagram of aligning the first angular velocity and the second angular velocity using a cross-correlation method according to an embodiment of this application.

[0162] As shown in Figure 16, in the original signal diagram, the first signal can characterize the first angular velocity determined based on the motion pose sequence of the motion component 600, and the second signal can characterize the second angular velocity determined based on the visual pose sequence converted from multiple images acquired by the camera component 700. Therefore, the first diagram of Figure 16 shows the amplitude of the first angular velocity and the second angular velocity at different times. From the first diagram, it can be observed that the first angular velocity and the second angular velocity have different times at the same amplitude, or in other words, the amplitudes of the first angular velocity and the second angular velocity corresponding to the same time are different.

[0163] The second figure in Figure 16 illustrates the cross-correlation function. Here, the cross-correlation method is used to align two angular velocities. The accuracy of the classic cross-correlation estimation method depends on the sampling frequency of the two angular velocity sequences. Therefore, to improve the accuracy and precision of aligning two angular velocities using the cross-correlation method, interpolation can be performed on the second angular velocity determined based on the visual pose sequence converted from multiple images acquired by the camera component 700.

[0164] Referring again to Figure 14, the time correction unit 340 may further include an interpolation subunit 342, which performs interpolation processing on the visual pose sequence before determining the second angular velocity, so that the sampling frequencies of the motion pose sequence and the visual pose sequence are the same.

[0165] Furthermore, since the second angular velocity determined based on camera component sampling is estimated from a target (e.g., a virtual target) in the extended reality device, and the quality of the target image is affected by the accuracy of the extended reality device's own localization and tracking (SLAM) system, filtering can be performed before interpolation. Then, the visual pose sequence is interpolated to the same sampling frequency as the motion pose sequence using the aforementioned interpolation process. The temporal bias of the visual pose sequence is then determined based on a cross-correlation method, and the relevant data is time-compensated and aligned. For example, referring to Figure 14, the time correction unit 340 may further include a filtering subunit 341, which filters the visual pose sequence before performing interpolation.

[0166] For clarity, Figure 14 shows the time correction unit 340, which includes, in an optional configuration, both a filtering subunit 341 and an interpolation subunit 342. Distinguishingly, the optional filtering subunit 341 or interpolation subunit 342 is shown with dashed boxes, and dashed connecting lines further illustrate the connections between the filtering subunit 341 or interpolation subunit 342 and other subunits when included.

[0167] Referring again to Figure 16, in the aligned signal graph, the first and second signals are processed using a cross-correlation function to obtain the aligned first and second signals. The aligned first signal represents the aligned first angular velocity, and the aligned second signal represents the aligned second angular velocity. The third figure in Figure 16 shows the amplitudes of the aligned first and second angular velocities at different times after aligning the first and second angular velocities using the cross-correlation method. From the third figure, it can be observed that the aligned first and second angular velocities at the same amplitude have approximately the same time, or in other words, the amplitudes of the aligned first and second angular velocities at the same time are approximately the same.

[0168] Figure 17 shows a schematic flowchart of camera component extrinsic parameter calibration according to an embodiment of this application.

[0169] Referring to Figures 10, 12 and 17, the calibration unit 350 can calibrate the extrinsic parameters of the camera component 700 based on the acquired visual pose sequence and motion pose sequence; or it can calibrate the extrinsic parameters of the camera component 700 based on the visual pose sequence and motion pose sequence modified by at least one of the above embodiments.

[0170] Specifically, at the start of calibration, a target board (e.g., the virtual target board shown in Figure 3) in the extended reality device 400 can be opened. The robotic arm of the motion component 600 can be fixedly connected to the extended reality device 400 and drive the extended reality device 400 to move synchronously. By pre-setting a predetermined motion trajectory and encoding it into the robotic arm, the robotic arm can be made to move the extended reality device 400 in various directions according to the predetermined motion trajectory.

[0171] The camera assembly 700 can acquire multiple images of a target displayed via the extended reality device 400 during synchronous motion. Data acquisition of these multiple images by the camera assembly 700 can include various data acquisition modes. Alternatively, during synchronous motion, the camera assembly 700 can acquire images of the target 410 displayed via the extended reality device 400 only when the extended reality device 400 passes a waypoint. Alternatively, the camera assembly 700 can continuously acquire images of the target 410 displayed via the extended reality device 400 throughout the entire synchronous motion. The multiple image data acquired by the camera assembly 700, along with the triaxial encoder data of the motion assembly 600, can then be stored as a motion pose sequence of the motion assembly 600.

[0172] In another alternative scenario, although the timestamps of the image acquired by the camera assembly 700 and the trajectory timestamps of the motion assembly 600 (or, the extended reality device) are from the same clock source, the image capture process has a certain time delay due to the presence of the MTP. Therefore, time correction is required to align the time of the motion assembly 600 and the camera assembly 700. After acquiring the visual pose sequence and motion pose sequence, the calibration unit 350 can calibrate the extrinsic parameters of the camera assembly 700 based on the visual pose sequence and motion pose sequence.

[0173] In the real world, motion changes are referenced to the motion component's coordinate system, while in the virtual world, motion changes are referenced to the camera component's coordinate system. Therefore, coordinate system transformation calibration is required for both the camera component and the motion component, which is known as hand-eye calibration. Calibrating the extrinsic parameters of the camera component (700) can be understood as hand-eye calibration.

[0174] Optionally, the extrinsic parameters of the camera assembly 700 can be obtained by solving the hand-eye calibration equations AX = XB, where A is the attitude transformation matrix of the camera assembly, B is the attitude transformation matrix of the end effector of the motion component (e.g., a robotic arm), and X is the hand-eye matrix to be solved. By solving the hand-eye calibration equations, the coordinate transformation relationship between the camera assembly 700 and the end effector of the motion component 600 can be obtained.

[0175] Figure 18 shows a flowchart illustrating the calibration of extrinsic parameters of a camera component based on multiple images and a motion pose sequence of a motion component according to an embodiment of this application.

[0176] Referring to Figures 10, 12, and 18, in some embodiments of this application, the first conversion unit 320 of the processing component 800 can extract the pixel coordinates of feature points of the target 410 (2D coordinate extraction of target feature points) based on multiple images of the target 410 displayed via the extended reality device 400, acquired by the camera component 700 during synchronous motion. The second conversion unit 330 of the processing component 800 can convert the pixel coordinates into a visual pose sequence according to preset spatial geometric constraints between feature points, such as preset target feature point ID resolution, where ID can be understood as the spatial position number of the target feature point. An ID-3D mapping table can be established through preset target feature point ID resolution. By looking up the ID-3D mapping table, the 3D spatial position of the target feature points can be obtained, thereby establishing a 2D-3D point pair relationship and converting the pixel coordinates into a visual pose sequence. The calibration unit 350 of the processing component 800 can calibrate the extrinsic parameters of the camera component 700 based on the visual pose sequence and the motion pose sequence.

[0177] Optionally, the calibration unit 350 may include a linear transformation subunit 351 and a nonlinear transformation subunit 352. The linear transformation subunit 351 may obtain the initial values ​​of the extrinsic parameters of the camera component 700 based on the visual pose sequence and the motion pose sequence through a direct linear transformation method. The nonlinear transformation subunit 352 may optimize the initial values ​​of the extrinsic parameters of the camera component 700 based on minimizing the reprojection error through a nonlinear transformation method, thereby obtaining the extrinsic parameters of the camera component 700.

[0178] Specifically, the direct linear transformation method can be divided into two steps. First, the projection matrix M is calculated, which represents the projection relationship between 2D and 3D point pairs. As shown in formula (1), u and v represent the coordinates of 2D points, X, Y and Z represent the coordinates of 3D points, K represents the intrinsic parameters of camera component 700, and R and T represent the extrinsic parameters of camera component 700 and the target plate in the world coordinate system (the second coordinate system with the motion component as the reference). Based on a series of point pairs, the QM equation system is constructed (as shown in formula (2)), and the matrix M is obtained through singular value decomposition. Since the extrinsic parameter R is an orthogonal matrix, the matrix M can be decomposed by QR to obtain the intrinsic parameters K and extrinsic parameters R and T as shown in formula (3). M = K(R|T)

[0179] As shown in formula (4), given the initial intrinsic parameters K, extrinsic parameters R and T, the distortion D can be preset to 0, and nonlinear joint optimization can be performed based on minimizing the reprojection error (as shown in formula (2)) to obtain the extrinsic parameters of the camera component 700.

[0180] To compare motion posture changes in the real world with those in the virtual world of the extended reality device 400 (as shown in Figure 11), it is necessary to unify these two types of motion posture changes under the same coordinate system. Motion changes in the virtual world can be determined in a first coordinate system (e.g., the camera component coordinate system) with reference to the camera component 700, while motion changes in the real world can be determined in a second coordinate system (e.g., the world coordinate system) with reference to the motion component 600. Therefore, it is necessary to calibrate the motion component 600 and the camera component 700 by coordinate system transformation.

[0181] Because different extended reality devices have different viewing distances, the positional relationship between the camera assembly 700 and the extended reality device 400 needs to be adjusted before measurement, which causes a change in the positional relationship between the motion component 600 and the camera assembly 700. To solve this problem, a feasible solution is to calibrate the extrinsic parameters of the camera assembly 700 by taking pictures of the virtual target displayed on the extended reality device using the camera assembly 700.

[0182] In the process of calibrating the extrinsic parameters of camera component 700 by solving the hand-eye calibration equations AX = XB, the extrinsic parameters can be solved using traditional hand-eye calibration algorithms such as the T-SAI method. However, this application does not limit the algorithm for solving the extrinsic parameters. In the solution process, the pose transformation matrix of the camera component represented by A can be understood as the visual pose sequence obtained in any of the above embodiments, and the pose transformation matrix of the motion component end effector (e.g., robotic arm) represented by B can be understood as the motion pose sequence obtained in any of the above embodiments. The extrinsic parameters of camera component 700 can be obtained by solving A and B. In other words, by solving the hand-eye calibration equations, the coordinate transformation relationship between camera component 700 and the end effector of motion component 600 can be obtained.

[0183] Figure 19 shows a schematic diagram of motion information of the motion component and the calibrated camera component in the same coordinate system according to an embodiment of this application.

[0184] As shown in Figure 19, the motion information of the moving components in the same coordinate system is a compound three-dimensional motion (3DOF), which can be characterized by yaw motion, pitch motion, and roll motion. In a right-handed Cartesian coordinate system in three-dimensional space, roll is the roll angle around the z-axis; pitch is the pitch angle around the x-axis; and yaw is the yaw angle around the y-axis. Similarly, the motion information of the calibrated camera components in the same coordinate system is also a compound three-dimensional motion, which can be characterized by yaw photon, pitch photon, and roll photon motion. Likewise, in a right-handed Cartesian coordinate system in three-dimensional space, roll is the roll angle around the z-axis; pitch is the pitch angle around the x-axis; and yaw is the yaw angle around the y-axis.

[0185] As can be observed from Figure 19, after time-delay correction of the motion component and the calibrated camera component, the motion information of the time-delay corrected motion component in the same coordinate system and the motion information of the calibrated camera component in the same coordinate system are approximately the same in degree at different times in each direction (roll, pitch and yaw).

[0186] Therefore, according to at least one embodiment of the calibration system provided in this application, the motion component moving synchronously with the extended reality device can simulate continuous composite three-dimensional motion of the human body. Furthermore, the camera component can acquire multiple images of the target displayed via the extended reality device. Thus, by using the multiple images of the target displayed via the extended reality device acquired by the camera component and the motion pose sequence of the motion component, the extrinsic parameters of the camera component can be obtained, thereby unifying the motion information of the camera component and the motion component in the same coordinate system, providing a coordinate system reference for measuring the delay of the extended reality device in continuous three-dimensional composite motion.

[0187] Figure 20 is a flowchart of a calibration method 4000 provided in an embodiment of this application, wherein the calibration method 4000 may include the following steps:

[0188] S5, obtain the motion pose sequence of the motion component, wherein the motion component is fixedly connected to the extended reality device and drives the extended reality device to move synchronously.

[0189] S6 uses a camera assembly that is fixedly connected to the augmented reality device to capture multiple images of a target displayed via the augmented reality device during synchronized motion.

[0190] S7 calibrates the extrinsic parameters of the camera components based on multiple images and the motion pose sequence of the motion components.

[0191] The following will describe in detail each step of the above calibration method 4000 with reference to the accompanying drawings.

[0192] Step S5

[0193] Referring to Figures 10, 11, and 20, in some embodiments of this application, the motion component 600 is fixedly connected to the extended reality device 400, for example, through a rigid connection. Furthermore, the motion component 600 can drive the extended reality device 400 to move synchronously. The motion component 600 can include any suitable structural components and can perform various translational and rotational movements in the world coordinate system. Taking the extended reality device 400 as an HMD as an example, the motion component 600 fixedly connected to the HMD can simulate the rotational movements of the human head in the sagittal, coronal, and horizontal planes, etc.

[0194] Optionally, the calibration method 4000 further includes acquiring multiple images of the target 410 (as shown in FIG3) displayed via the extended reality device 400 using the camera component 700, and performing time synchronization of the motion component 600 and the camera component 700 before and during the synchronized motion.

[0195] For example, a clock source is used to synchronize the motion component 600 and the camera component 700. During the time synchronization process, the motion component 600 and the camera component 700 can be uniformly timed via the PTP protocol. A high-precision and high-accuracy clock source is used as the master clock to send timestamps to the motion component 600 and the camera component 700, which act as slave clocks. The slave clocks compare the timestamps with their own local time and adjust the time according to the difference, thereby ensuring that the motion pose sequence obtained based on the motion component 600 is synchronized with the visual pose sequence obtained based on the camera component 700.

[0196] Step S6

[0197] Referring to Figures 10, 2-4, 11 and 20, in some embodiments of this application, a camera assembly 700 fixedly connected to the extended reality device 400 may be used to capture multiple images of a target 410 displayed via the extended reality device 400 during synchronous motion, for example, Figure 4 shows one of the multiple images, image 420.

[0198] The camera assembly 700 can be used to simulate a scenario where the human eye receives an image displayed by the extended reality device 400. Optionally, the camera assembly 700 may include at least one of a camera and a video camera, and the camera assembly may include an IR camera assembly (Infrared Camera), an RGB camera assembly (Red-Green-Blue Camera), or a monochrome camera assembly, etc. In addition, the camera assembly 700 may also include other structures such as a photoelectric sensor, and this application does not limit the specific internal structure of the camera assembly 700.

[0199] Optionally, the calibration board 410 may be a virtual calibration board located within the extended reality device 400. The virtual calibration board may include visual marker calibration boards from a visual reference library, such as ApriTag, ArUco, and other QR code calibration boards. Using a virtual calibration board simplifies the structure of the calibration system and reduces the difficulty of calibration implementation.

[0200] Step S7

[0201] In some embodiments of this application, step S7, which calibrates the extrinsic parameters of the camera component based on multiple images and the motion pose sequence of the motion component, may include, for example,: extracting the pixel coordinates of feature points of the target plate based on multiple images; converting the pixel coordinates into a visual pose sequence according to the preset spatial geometric constraint relationship between the feature points; and calibrating the extrinsic parameters of the camera component based on the visual pose sequence and the motion pose sequence.

[0202] Specifically, referring to Figures 10, 11, and 20, the processing component 800 can be used to obtain a visual pose sequence of an extended reality device based on multiple images. The processing component 800 can be a variety of general-purpose and / or dedicated processing units with processing and computing capabilities. Some examples of the processing component 800 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc.

[0203] The processing component 800 obtains a visual pose sequence of the extended reality device 400 based on multiple images, and determines the delay of the extended reality device 400 based on the visual pose sequence and the motion pose sequence of the motion component 600. Optionally, the processing component 800 may include various suitable communication buses and input / output interfaces for connecting to the motion component 600 and the camera component 700.

[0204] Furthermore, before extracting the pixel coordinates of feature points of the target plate 410 (as shown in Figure 3) based on multiple images, the multiple images can be processed. This processing may include at least one of downsampling processing and Gaussian blur processing. Since the presence of moiré patterns when capturing the screen of the extended reality device 400 can lead to inaccurate or even impossible extraction of 2D feature points, it is necessary to process the acquired multiple images. Downsampling processing and Gaussian blur processing, among others, can reduce the impact of moiré patterns on the images.

[0205] Optionally, before acquiring multiple images of the target 410 displayed via the extended reality device 400, the positional relationship between the camera assembly 700 and the extended reality device 400 can be adjusted so that multiple feature points specified in the target 410 are shown in the multiple images. This improves the quality of the generated multiple images and ensures high stability and high accuracy of the calibration results.

[0206] Furthermore, in order to compare the changes in motion posture in the real world with the changes in motion posture in the virtual world of the extended reality device 400, it is necessary to unify these two types of motion posture changes under the same coordinate system. The motion changes in the virtual world can be determined in a first coordinate system (e.g., the camera component coordinate system) with reference to the camera component 700, while the motion changes in the real world can be determined in a second coordinate system (e.g., the world coordinate system) with reference to the motion component 600. Therefore, it is necessary to calibrate the motion component 600 and the camera component 700 by coordinate system transformation.

[0207] Because different extended reality devices have different viewing distances, the positional relationship between the camera assembly 700 and the extended reality device 400 needs to be adjusted before measurement, which causes a change in the positional relationship between the motion component 600 and the camera assembly 700. To solve this problem, a feasible solution is to calibrate the extrinsic parameters of the camera assembly 700 by taking pictures of the virtual target displayed on the extended reality device using the camera assembly 700.

[0208] Optionally, the camera assembly 700 can acquire multiple images of the target displayed via the extended reality device 400 during synchronous motion. Data acquisition of multiple images by the camera assembly 700 can include various data acquisition modes. Alternatively, the camera assembly 700 can acquire images of the target 410 displayed via the extended reality device 400 only when the extended reality device 400 passes a waypoint during synchronous motion. Alternatively, the camera assembly 700 can continuously acquire images of the target 410 displayed via the extended reality device 400 throughout the entire synchronous motion. The multiple image data acquired by the camera assembly 700 and the three-axis encoder data of the motion assembly 600 can then be stored, wherein the three-axis encoder data of the motion assembly 100 can serve as a motion pose sequence of the motion assembly 600.

[0209] In another alternative scenario, although the timestamps of the image acquired by the camera component 700 and the trajectory timestamps of the motion component 600 (or, the extended reality device) are from the same clock source, the image capture process has a certain time delay due to the presence of the MTP. Therefore, time correction is required for the visual pose sequence to align the time of the motion component 600 and the camera component 700. Optionally, time correction is performed on the visual pose sequence based on the time offset of the visual pose sequence.

[0210] As shown in Figure 2, the visual pose sequence may include multiple first-time-pose information items arranged in chronological order. The first-time-pose information includes first-time information and first attitude angles corresponding to the first-time information, wherein the first attitude angles include a first roll angle, a first yaw angle, and a first pitch angle. For example, the form of the first-time-pose information may include: timestamp_roll_pitch_yaw, where 000_10_20_30 can represent the roll, pitch, and yaw angles at the start time as (10°, 20°, 30°), respectively. In the three-dimensional right-handed Cartesian coordinate system, roll is the roll angle around the z-axis; pitch is the pitch angle around the x-axis; and yaw is the yaw angle around the y-axis.

[0211] The motion pose sequence may include multiple second time-pose information items arranged in chronological order. The second time-pose information may include second time information and corresponding second attitude angles, where the second attitude angles include a second roll angle, a second yaw angle, and a second pitch angle. Similarly, the second time-pose information may also be in the form of a timestamp: `_roll_pitch_yaw`.

[0212] Referring to Figures 10, 11 and 20, in some embodiments of this application, the fixed connection between the motion component 600 and the extended reality device 400 can be considered as a rigid connection of a rigid body, and the fixed connection between the camera component 700 and the extended reality device 400 can also be considered as a rigid connection of a rigid body. Based on the principle of rigid body rotation invariance, the time offset can be obtained by aligning the first angular velocity of the rotation of the motion platform 100 with the second angular velocity estimated from the attitude of the camera component 700.

[0213] For example, a first angular velocity is determined based on the motion pose sequence of the motion component 600, and a second angular velocity is determined based on the visual pose sequence converted from multiple images acquired by the camera component 700; the first and second angular velocities are aligned based on a cross-correlation method to determine the temporal offset of the visual pose sequence; and the visual pose sequence is temporally corrected based on the temporal offset.

[0214] In addition, time correction of the visual pose sequence may also include interpolating the visual pose sequence before determining the second angular velocity to make the sampling frequency of the motion pose sequence and the visual pose sequence the same.

[0215] Optionally, since the second angular velocity determined based on the sampling of the camera components is estimated by a target (e.g., a virtual target) in the extended reality device, and the quality of the target image is affected by the accuracy of the extended reality device's own localization and tracking (SLAM) system, filtering can be performed before interpolation. Then, the visual pose sequence is interpolated to the same sampling frequency as the motion pose sequence through the above interpolation process. The temporal offset of the visual pose sequence is then determined based on the cross-correlation method, and the relevant data is time-compensated and aligned.

[0216] In some embodiments of this application, the extrinsic parameters of the camera component 700 can be calibrated based on the obtained visual pose sequence and motion pose sequence; or the extrinsic parameters of the camera component 700 can be calibrated based on the visual pose sequence and motion pose sequence modified by at least one of the above embodiments.

[0217] Specifically, at the start of calibration, a target board (e.g., the virtual target board shown in Figure 3) in the extended reality device 400 can be opened. The robotic arm of the motion component 600 can be fixedly connected to the extended reality device 400 and drive the extended reality device 400 to move synchronously. By pre-setting a predetermined motion trajectory and encoding it into the robotic arm, the robotic arm can be made to move the extended reality device 400 in various directions according to the predetermined motion trajectory.

[0218] The camera assembly 700 can acquire multiple images of a target displayed via the extended reality device 400 during synchronous motion. Data acquisition of these multiple images by the camera assembly 700 can include various data acquisition modes. Alternatively, during synchronous motion, the camera assembly 700 can acquire images of the target 410 displayed via the extended reality device 400 only when the extended reality device 400 passes a waypoint. Alternatively, the camera assembly 700 can continuously acquire images of the target 410 displayed via the extended reality device 400 throughout the entire synchronous motion. The multiple image data acquired by the camera assembly 700, along with the triaxial encoder data of the motion assembly 600, can then be stored as a motion pose sequence of the motion assembly 600.

[0219] In another alternative scenario, although the timestamps of the image acquired by the camera component 700 and the trajectory timestamps of the motion component 600 (or, the extended reality device) are from the same clock source, the image capture process has a certain time delay due to the presence of the MTP. Therefore, time correction is required to align the time of the motion component 600 and the camera component 700. After acquiring the visual pose sequence and motion pose sequence, the extrinsic parameters of the camera component 700 can be calibrated based on the visual pose sequence and motion pose sequence.

[0220] In the real world, motion changes are referenced to the motion component's coordinate system, while in the virtual world, motion changes are referenced to the camera component's coordinate system. Therefore, coordinate system transformation calibration is required for both the camera component and the motion component, which is known as hand-eye calibration. Calibrating the extrinsic parameters of the camera component (700) can be understood as hand-eye calibration.

[0221] Optionally, the extrinsic parameters of the camera assembly 700 can be obtained by solving the hand-eye calibration equations AX = XB, where A is the attitude transformation matrix of the camera assembly, B is the attitude transformation matrix of the end effector of the motion component (e.g., a robotic arm), and X is the hand-eye matrix to be solved. By solving the hand-eye calibration equations, the coordinate transformation relationship between the camera assembly 700 and the end effector of the motion component 600 can be obtained.

[0222] In other words, in the process of calibrating the extrinsic parameters of the camera component 700 by solving the hand-eye calibration equations AX = XB, the extrinsic parameters can be solved using traditional hand-eye calibration algorithms such as the T-SAI method. However, this application does not limit the algorithm for solving the extrinsic parameters. In the solution process, the pose transformation matrix of the camera component represented by A can be understood as the visual pose sequence obtained in any of the above embodiments, and the pose transformation matrix of the motion component end effector (e.g., a robotic arm) represented by B can be understood as the motion pose sequence obtained in any of the above embodiments. The extrinsic parameters of the camera component 700 can be obtained by solving A and B. In other words, by solving the hand-eye calibration equations, the coordinate transformation relationship between the camera component 700 and the end effector of the motion component 600 can be obtained.

[0223] Optionally, the initial values ​​of the extrinsic parameters of the camera component 700 can be obtained based on the visual pose sequence and the motion pose sequence using a direct linear transformation method; then, the initial values ​​of the extrinsic parameters of the camera component 700 can be optimized based on minimizing the reprojection error using a nonlinear transformation method to obtain the extrinsic parameters of the camera component 700.

[0224] Therefore, according to the calibration method provided in at least one embodiment of this application, the motion component that moves synchronously with the extended reality device can simulate continuous composite three-dimensional motion of the human body, and the camera component can acquire multiple images of the target displayed by the extended reality device. By acquiring multiple images of the target displayed by the extended reality device and the motion pose sequence of the motion component from the multiple images acquired by the camera component, the extrinsic parameters of the camera component can be obtained, thereby unifying the motion information of the camera component and the motion component in the same coordinate system, providing a coordinate system reference for measuring the delay of the extended reality device in continuous three-dimensional composite motion.

[0225] In related technologies, MTP latency testing is a necessary step before display devices leave the factory. The accuracy of MTP latency testing is crucial, as it guides R&D personnel in improving display devices and reducing their MTP latency. However, the accuracy of MTP latency measurements obtained using relevant testing methods is relatively low.

[0226] Figures 21 and 22 respectively illustrate an exemplary system architecture 5000 of a delay testing method applicable to some embodiments of this application.

[0227] In an exemplary embodiment, as shown in FIG21, the system architecture 5000 may include a display device 110 and a delay detection device 120. The display device 110 may include one or more of a virtual reality device, an augmented reality device, and a mixed reality device. The delay detection device 120 may include a simulation component 121, a controller 122, a drive motor 123, and a camera module 124. The simulation component 121 can be used to simulate user actions. The drive motor 123 can be used to drive the movement of the simulation component 121. The camera module 124 can be used to capture display images of the display device 110 during the movement of the simulation component 121. The controller 122 may be communicatively connected to the drive motor 123 and the camera module 124, for example. The controller 122 may also be communicatively connected to the display device 110, for example.

[0228] The delay test method of this application embodiment can be executed by the controller 122 of the delay detection device 120.

[0229] In an embodiment, as shown in FIG22, the system architecture 5000 may include a display device 110, a delay detection device 120, and a server 130. The display device 110 may include one or more of a virtual reality device, an augmented reality device, and a mixed reality device. The delay detection device 120 may include a simulation component 121, a controller 122, a drive motor 123, and a camera module 124. The simulation component 121 can be used to simulate user actions. The drive motor 123 can be used to drive the movement of the simulation component 121. The camera module 124 can be used to capture display images of the display device 110 during the movement of the simulation component 121. The controller 122 may be communicatively connected to the drive motor 123 and the camera module 124, for example. The server 130 may be communicatively connected to the display device 110 and the controller 122, for example. The server 130 can provide various services based on the display device 110 and the delay detection device 120. For example, the server 130 may analyze and process a received first action sequence and generate a processing result (e.g., obtaining a first motion trajectory of the simulation component 121 based on the first action sequence).

[0230] Server 130 can be either hardware or software. When server 130 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 130 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide distributed services), or as a single software program or software module. No specific limitations are made here. Furthermore, the latency testing method of this embodiment can be executed by server 130.

[0231] Figure 23 is a schematic flowchart of a delay test method 6000 according to an embodiment of this application. The delay test method 6000 can be executed, for example, by an electronic device such as the controller 122 of the delay detection device 120 (as shown in Figure 21) or a server 130 (as shown in Figure 22). It should be understood that the delay test method 600 may also include additional steps not shown and / or the steps shown may be omitted; the scope of this application is not limited in this respect. The delay test method 200 can be applied to application scenarios such as delay testing of a display device 110, wherein the display device 110 may include a virtual reality device, an augmented reality device, or a mixed reality device. It should be understood that the delay test method 6000 can also be applied to delay testing application scenarios of other display devices, and this is not limited here.

[0232] As shown in Figure 23, the delay test method 6000 may include the following steps:

[0233] S210, in response to receiving the first action sequence, the simulation component obtains the first motion trajectory of the simulated component in a preset time period according to the first action sequence, wherein the first action sequence includes first time information and first pose information.

[0234] S220, a second action sequence is obtained from the display image of the display device during the movement of the simulated component according to the first motion trajectory, and a second motion trajectory for a preset time period is determined according to the second action sequence, wherein the second action sequence includes second time information and second pose information.

[0235] S230, determine the delay of the display device in a preset time period based on the first motion trajectory and the second motion trajectory.

[0236] The method provided in the embodiments of this application obtains the first motion trajectory of a simulated component within a preset time period through a first action sequence. This allows the simulated component to simulate the actions of different users and makes the actions of the simulated component more realistic. Then, while the simulated component moves according to the first motion trajectory, a second motion trajectory within the preset time period is determined based on a second action sequence obtained from the display image of the display device. The delay of the display device within the preset time period can be determined based on the first and second motion trajectories, thus realizing delay testing within the preset time period and improving the accuracy of delay testing of the display device.

[0237] The steps S210 to S230 of the embodiments of this application will be described in detail below.

[0238] In step S210, in response to receiving the first action sequence, the first motion trajectory of the simulated component in a preset time period is obtained according to the first action sequence, wherein the first action sequence may include first time information and first pose information.

[0239] In this embodiment, the first action sequence can be an action sequence corresponding to the action to be tested. For example, the first action sequence can be an action sequence constructed from action parameters at different time points during the user's performance of the action to be tested. The number of action parameters in the first action sequence can be determined by the action execution time and sampling time interval of the action to be tested. The sampling time interval between two adjacent action parameters in the first action sequence can be the same or different. For example, if the action execution time of the action to be tested is 1 second, the sampling time interval is 0.01 seconds, and the total number of samples from the start time to the end time of the action to be tested is 101, then the number of action parameters in the first action sequence is 101.

[0240] The action parameters within the first action sequence can be obtained via "t". m _roll m _pitch m _yaw m The expression is in the form of "t", where 1 ≤ m ≤ a, a is the number of action parameters in the first action sequence, and t is the number of action parameters in the first action sequence. m For the time corresponding to the m-th action parameter in the first action sequence, rollm _pitch m _yaw m This refers to the attitude angle corresponding to the m-th action parameter in the first action sequence. As shown in Figure 24, in the world coordinate system, roll is the angle of rotation around the Z-axis (i.e., roll angle), pitch is the angle of rotation around the X-axis (i.e., pitch angle), and yaw is the angle of rotation around the Y-axis (i.e., yaw angle). The aforementioned t... m For the first time information of the first action sequence, the above-mentioned roll m _pitch m _yaw m This is the first pose information of the first action sequence.

[0241] For example, the number of motion parameters in the first motion sequence can be 101. The first motion parameter in the first motion sequence can be "t1_roll1_pitch1_yaw1", where t1 is the start time of the action to be performed, and roll1_pitch1_yaw1 is the attitude angle corresponding to the start time. "t1_roll1_pitch1_yaw1" can be, for example, "000_10_20_30", where "000" represents the start time, and "10_20_30" represents that the roll angle is 10°, the pitch angle is 20°, and the yaw angle is 30° at the start time. The 101st motion parameter in the first motion sequence can be "t 101 _roll 101 _pitch 101 _yaw 101 ” t 101 To determine the termination time of the action to be tested, roll 101 _pitch 101 _yaw 101 The attitude angle corresponding to the termination time.

[0242] In this embodiment, after receiving the first action sequence, it is determined whether a target action sequence corresponding to the first action sequence exists in the database. If a target action sequence corresponding to the first action sequence exists in the database, the target motion trajectory corresponding to the target action sequence is determined as the first motion trajectory of the simulation component 121. The database in this document may, for example, be a pre-configured mapping relationship between action sequences and the motion trajectories of the simulation component 121. After determining that a target action sequence corresponding to the first action sequence exists in the database, the first motion trajectory of the simulation component 121 can be determined according to the aforementioned mapping relationship.

[0243] When the delay test method is implemented by the controller 122 of the delay detection device 120, the database can be stored in the controller 122. When the delay test method is implemented by the server 130, the database can be stored in the server 130.

[0244] In this embodiment, if the database does not contain a target action sequence corresponding to the first action sequence, the pose information of the reference point at different time points during the user's action corresponding to the first action sequence is obtained. For example, after determining that the database does not contain a target action sequence corresponding to the first action sequence, a request is sent to the display device 110 to obtain the pose information of the user's head at different time points during the user's action corresponding to the first action sequence. Then, the display device 110 uses its own positioning and tracking system to obtain the pose information of the user's head at different time points according to a preset sampling time interval. After receiving the pose information of the user's head at different time points sent by the display device 110, the pose information of the reference point at different time points can be determined based on the pose information of the user's head at different time points.

[0245] In other embodiments, if the database does not contain a target action sequence corresponding to the first action sequence, then the pose information of multiple test points at different time points during the user's action corresponding to the first action sequence is acquired. Then, the pose information of a reference point at different time points is determined based on the pose information of the multiple test points at different time points, where the reference point is the center point of the three-dimensional space formed by the multiple test points. Specifically, after determining that the database does not contain a target action sequence corresponding to the first action sequence, a request is sent to the motion capture system (not shown) to acquire the pose information of multiple test points at different time points during the user's action according to the test action. From the start time to the end time of the user's action according to the test action, the motion capture system can acquire the pose information of multiple test points at different time points according to a preset sampling time interval. For example, as shown in Figure 25, the number of test points can be six, which may include test point A located at the user's left ear, test point B located at the user's left forehead, test point C located below the user's left eye, test point D located at the user's right ear, test point E located at the user's right forehead, and test point F located below the user's right eye. The center point O of the three-dimensional space formed by the above six test points can be used as the reference point. After receiving the pose information of multiple test points at different time points sent by the motion capture system, the pose information of the reference point (e.g., the center point O) at different time points can be determined based on the pose information of multiple test points at different time points.

[0246] After determining the pose information of the reference point at different time points, a first motion trajectory of the simulation component 121 within a preset time period can be determined based on the pose information of the reference point at different time points. For example, the pose information of the simulation component 121 at different time points is determined based on the pose information of the reference point at different time points. Then, the motion parameters of the simulation component 121 at different time points are determined based on the pose information of the simulation component 121 at different time points. The motion parameters of the simulation component 121 may include, but are not limited to, one or more of the rotation angle, velocity, and acceleration of the simulation component 121. The rotation angle of the simulation component 121 can be determined based on the pose information of the simulation component 121 at two adjacent time points. The velocity and acceleration of the simulation component 121 can be determined based on the rotation angle of the simulation component 121, the time interval between two adjacent time points, and the action execution time. Finally, a third motion trajectory is generated based on the pose information and motion parameters of the simulation component 121 at different time points, and the third motion trajectory is determined as the first motion trajectory of the simulation component 121 within the preset time period.

[0247] In this embodiment, the user's head motion sequence is collected, and the motion planning and control algorithm maps the user's head motion sequence to a simulation component. The simulation component can be, for example, a simulated head model, to meet the needs of personalized testing and make the motion of the simulation component more realistic.

[0248] In this embodiment, the pose of the simulation component 121 can be controlled by three drive motors, which may include a first drive motor, a second drive motor, and a third drive motor. The first drive motor can be used to control the roll angle of the simulation component 121, the second drive motor can be used to control the pitch angle of the simulation component 121, and the third drive motor can be used to control the yaw angle of the simulation component 121. Accordingly, the rotation angle of the simulation component 121 may include the rotation angles of the roll angle, pitch angle, and yaw angle of the simulation component 121.

[0249] In this embodiment, when generating the third motion trajectory based on the pose information and motion parameters of the simulation component 121 at different time points, it is determined whether the actual rotation angle of the simulation component 121 corresponds to the target rotation angle. If the actual rotation angle of the simulation component 121 does not correspond to the target rotation angle, the control signal of the drive motor is adjusted so that the actual rotation angle of the simulation component 121 corresponds to the target rotation angle, thereby making the third motion trajectory smoother.

[0250] In this embodiment, when the movement of the drive motor (e.g., the first drive motor, the second drive motor, and the third drive motor) exceeds the safe range, an alarm is triggered, and the drive motor is automatically stopped. Software or hardware limits can also be added to prevent the drive motor's movement from exceeding the safe range, thereby improving safety.

[0251] In this embodiment, the first motion trajectory of the simulation component 121 within a preset time period is obtained through the first action sequence. This allows the simulation component 121 to simulate the actions of different users, making its actions more realistic. This facilitates personalized testing for different users, provides researchers with a wider range of test data, and better improves the display device, reducing its latency. Simultaneously, the third motion trajectory generated based on the pose information and action parameters of the simulation component 121 is smoother and more realistic.

[0252] In step S220, a second motion sequence is obtained based on the displayed image of the display device during the movement of the simulated component according to the first motion trajectory, and a second motion trajectory for a preset time period is determined based on the second motion sequence, wherein the second motion sequence includes second time information and second pose information.

[0253] Specifically, after obtaining the first motion trajectory of the simulation component 121 within a preset time period, the simulation component 121 can move according to the first motion trajectory. From the start time to the end time of the movement of the simulation component 121 according to the first motion trajectory, the display images of the display device 110 at different time points are acquired, wherein the display images correspond to the movement of the simulation component 121.

[0254] Because moiré patterns exist in the displayed image, these patterns can affect the accuracy of feature points extracted from the displayed image. Therefore, after acquiring the displayed image, filtering processing, such as Gaussian blurring, is performed on the displayed image to reduce the impact of moiré patterns on the accuracy of feature points extracted from the displayed image.

[0255] After processing the displayed image using Gaussian blur and other methods, feature extraction is performed on the displayed image of the display device 110 at different time points to obtain feature points, and the two-dimensional coordinates of the feature points at different time points are determined. The number of feature points can be, for example, multiple. Since each feature point has a unique identifier, and there is a predefined mapping relationship between the identifier and the three-dimensional coordinates, the three-dimensional coordinates of the feature point at different time points can be determined based on the identifier and the aforementioned mapping relationship, and two-dimensional coordinate-three-dimensional coordinate matching point pairs of the feature point at different time points are constructed. Then, a second action sequence can be determined based on the two-dimensional coordinate-three-dimensional coordinate matching point pairs of the feature point at different time points. For example, pose calculation is performed on the two-dimensional coordinate-three-dimensional coordinate matching point pairs of the feature point at different time points to obtain the rotation matrix of the displayed image relative to the world coordinate system at different time points; then, the second action sequence is determined based on the rotation matrix of the displayed image relative to the world coordinate system at different time points. Specifically, the rotation matrix of the displayed image relative to the world coordinate system at different time points can be obtained using the PNP (Perspective-n-Points) algorithm. After determining the second action sequence, a second motion trajectory for a preset time period can be determined based on the second action sequence.

[0256] The action parameters within the second action sequence can be obtained via "t". n _roll n _pitch n _yaw n The expression is in the form of ", where 1≤n≤b, b is the number of action parameters in the second action sequence, and t n For the time corresponding to the nth action parameter in the second action sequence, roll n _pitch n _yaw n This refers to the attitude angle corresponding to the nth action parameter within the second action sequence. The aforementioned t... n For the second time information of the second action sequence, the above-mentioned roll n _pitch n _yaw n This refers to the second pose information of the second action sequence. The number of motion parameters in the second action sequence can be the same as the number of motion parameters in the first action sequence. The sampling time interval between adjacent motion parameters in the second action sequence can be the same as the sampling time interval between adjacent motion parameters in the first action sequence.

[0257] In this embodiment, preprocessing the displayed image, such as Gaussian blurring, can reduce moiré patterns and improve the accuracy of feature points extracted from the displayed image. Simultaneously, matching pairs of two-dimensional and three-dimensional coordinates of the feature points at different time points are constructed. Based on these matching pairs, the pose changes of the virtual world at different time points are calculated to obtain a second action sequence. Finally, a second motion trajectory for a preset time period is determined based on the second action sequence.

[0258] In step S230, the delay of the display device in a preset time period is determined based on the first motion trajectory and the second motion trajectory.

[0259] In related technologies, when testing the latency of a display device 110, only the latency of the start and / or end times of a user's action can usually be obtained, but the latency of other times during the user's action besides the start and end times cannot be obtained. However, the latency of other times during the user's action is also quite important, as it still affects the user's immersion.

[0260] In order to perform a time delay test on the display device 110 for a preset time period, a first motion trajectory and a second motion trajectory for a preset time period can be obtained. After obtaining the first motion trajectory and the second motion trajectory, linear fitting and / or quadratic curve fitting are performed on the first motion trajectory and the second motion trajectory. For example, the linear interval and non-linear interval of the first motion trajectory and the second motion trajectory are identified, linear fitting is performed on the linear interval, and quadratic curve fitting is performed on the non-linear interval.

[0261] Figure 26 is a schematic diagram of the first and second motion trajectories according to an embodiment of this application. Figure 27 is a schematic diagram of linear fitting of the first motion trajectory within region G in Figure 26. The formula for linear fitting of the first motion trajectory within region G is y1 = -6.198*x1 + 3.098, where x1 is the time corresponding to the first motion trajectory and y1 is the attitude angle corresponding to the first motion trajectory. Figure 28 is a schematic diagram of linear fitting of the second motion trajectory within region G in Figure 26. The formula for linear fitting of the second motion trajectory within region G is y2 = -6.21*x2 + 3.043, where x2 is the time corresponding to the second motion trajectory and y2 is the attitude angle corresponding to the second motion trajectory. It should be understood that the formulas for linear fitting of the first and second motion trajectories within region G are merely exemplary, and this application does not impose specific limitations on them.

[0262] As shown in Figures 26, 27, and 28, linear fitting is performed on the first and second motion trajectories within region G, resulting in the fitted first and second motion trajectories. The fitted first and second motion trajectories can be referenced in Figure 29, which shows the fitted first and second motion trajectories within the time frame of 0.3 seconds to 0.7 seconds.

[0263] By performing linear fitting on the linear intervals of the first and second motion trajectories and quadratic curve fitting on the nonlinear intervals, the fitting effect of the first and second motion trajectories can be improved, which is beneficial to improving the testing accuracy of the display device's latency. In other examples, zero-bias preprocessing can be performed on the first and second motion trajectories before fitting them.

[0264] After fitting the first and second motion trajectories, the delay of the display device 110 within a preset time period can be determined based on the fitted first and second motion trajectories. As shown in Figure 26, both the first and second motion trajectories are time-pose curves, with time on the horizontal axis and posture (e.g., posture angle) on the vertical axis. Both the first and second motion trajectories can represent pose changes at different time points. The time difference between the second and first motion trajectories under the same pose is the delay of the display device 110. Therefore, the delay curve of the display device 110 can be obtained based on the time difference between the fitted second and first motion trajectories under the same pose. Then, the delay of the display device 110 within a preset time period can be determined based on the delay curve of the display device 110. By fitting the first and second motion trajectories, the delay test of the display device within a preset time period can be achieved, and the test accuracy of the display device's delay can be improved.

[0265] In other embodiments, as shown in Figures 26, 29 and 30, the time-pose curve can be a time-attitude angle curve. The delay curve of the display device 110 can be obtained based on the time difference between the second motion trajectory and the first motion trajectory fitted under the same attitude angle. The attitude angle may include roll angle, pitch angle and yaw angle.

[0266] The time-attitude curve may include a time-roll angle curve, a time-pitch angle curve, and a time-yaw angle curve. Correspondingly, the first motion trajectory may include time-roll angle curve I, time-pitch angle curve I, and time-yaw angle curve I, and the second motion trajectory may include time-roll angle curve II, time-pitch angle curve II, and time-yaw angle curve II. The delay curves of the display device 110 may include roll angle delay curves, pitch angle delay curves, and yaw angle delay curves.

[0267] The roll angle delay curve of display device 110 can be obtained by calculating the time difference between the fitted time-roll angle curve II and the time-roll angle curve I under the same roll angle. Similarly, the pitch angle delay curve of display device 110 can be obtained by calculating the time difference between the fitted time-pitch angle curve II and the time-pitch angle curve I under the same pitch angle. Likewise, the yaw angle delay curve of display device 110 can be obtained by calculating the time difference between the fitted time-yaw angle curve II and the time-yaw angle curve I under the same yaw angle.

[0268] In this embodiment, by performing linear fitting on the linear intervals of the first and second motion trajectories and quadratic curve fitting on the nonlinear intervals, the fitting effect of the first and second motion trajectories can be improved. Then, by using the delay curves obtained from the fitted first and second motion trajectories, a delay test of the display device for a preset time period can be achieved, and the delay test accuracy of the display device can be improved.

[0269] Figure 31 is a schematic diagram of the interaction between a display device and a delay detection device according to an embodiment of this application. The specific test process of the delay of the display device will be described below with reference to Figure 31.

[0270] S31, the display device 110 sends the first action sequence to the controller 122 of the delay detection device 120.

[0271] S32, after receiving the first action sequence, the controller 122 determines whether a target action sequence corresponding to the first action sequence exists in the database. In response to the existence of a target action sequence corresponding to the first action sequence in the database, the controller determines the target motion trajectory corresponding to the target action sequence as the first motion trajectory of the simulation component.

[0272] S33, in response to the absence of a target action sequence corresponding to the first action sequence in the database, the controller 122 sends a request to the display device 110 to obtain the pose information of the user's head at different time points during the user's action corresponding to the first action sequence.

[0273] S34, after receiving the above-mentioned request to obtain pose information, the display device 110 sends the pose information of the user's head at different time points during the user's actions corresponding to the first action sequence to the controller 122.

[0274] S35, after receiving the above-mentioned acquired pose information, the controller 122 determines the pose information of the reference point at different time points based on the pose information of the user's head at different time points, and determines the first motion trajectory of the simulation component in the preset time period based on the pose information of the reference point at different time points.

[0275] S36, after determining the first motion trajectory according to step S32 or S35, the controller 122 sends a request to the camera module 124 of the delay detection device 120 to acquire the display image of the display device during the movement of the analog component according to the first motion trajectory.

[0276] S37, after receiving the above-mentioned request to acquire the display image, the camera module 124 sends the display image of the display device during the movement of the analog component according to the first motion trajectory to the controller 122.

[0277] S38, after receiving the above-mentioned display image, the controller 122 obtains the second action sequence according to the display image, and determines the second motion trajectory for a preset time period according to the second action sequence.

[0278] S39, after determining the second motion trajectory, the controller 122 determines the delay of the display device within a preset time period based on the first motion trajectory and the second motion trajectory.

[0279] In other embodiments, the delay detection device 120 may include a data acquisition module (not shown). The data acquisition module can be used to acquire a first action sequence and send the first action sequence to the controller 122 of the delay detection device 120. The data acquisition module can also be used to acquire pose information of multiple test points at different time points during the user's actions corresponding to the first action sequence, and send the pose information to the controller 122 of the delay detection device 120.

[0280] Figure 32 is a schematic diagram illustrating the interaction between the display device, the latency detection device, and the server 130 according to an embodiment of this application. The specific testing process for the latency of the display device will be described below with reference to Figure 32.

[0281] S41, the display device 110 sends the first action sequence to the server 130.

[0282] S42, after receiving the first action sequence, the server 130 determines whether a target action sequence corresponding to the first action sequence exists in the database. In response to the existence of a target action sequence corresponding to the first action sequence in the database, the server determines the target motion trajectory corresponding to the target action sequence as the first motion trajectory of the simulation component.

[0283] S43, in response to the absence of a target action sequence corresponding to the first action sequence in the database, the server 130 sends a request to the display device 110 to obtain the pose information of the user's head at different time points during the user's action corresponding to the first action sequence.

[0284] S44, after receiving the above request to obtain pose information, the display device 110 sends the pose information of the user's head at different time points during the user's actions corresponding to the first action sequence to the server 130.

[0285] S45, after receiving the above-mentioned acquired pose information, the server 130 determines the pose information of the reference point at different time points based on the pose information of the user's head at different time points, and determines the first motion trajectory of the simulation component in the preset time period based on the pose information of the reference point at different time points.

[0286] S46, after determining the first motion trajectory according to step S42 or S45, the server 130 sends a request to the controller 122 of the delay detection device 120 to obtain the display image of the display device during the movement of the simulated component according to the first motion trajectory.

[0287] S47, after receiving the above-mentioned request to obtain the display image, the controller 122 sends the display image of the display device during the movement of the analog component according to the first motion trajectory to the server 130.

[0288] S48, after receiving the above-mentioned display image, the server 130 obtains the second action sequence based on the display image, and determines the second motion trajectory for a preset time period based on the second action sequence.

[0289] S49, after determining the second motion trajectory, the server 130 determines the delay of the display device within a preset time period based on the first and second motion trajectories.

[0290] In other embodiments, the delay detection device 120 may have a data acquisition module (not shown). The data acquisition module can be used to acquire a first action sequence and send the first action sequence to the server 130. The data acquisition module can also be used to acquire pose information of multiple test points at different time points during the user's action corresponding to the first action sequence, and send the pose information to the server 130.

[0291] One embodiment of this application also provides a delay detection device, which may include a simulation component, a drive motor, a camera module, and a controller. The simulation component can be used to simulate user actions. The simulation component may be, for example, a simulated head model. The drive motor can be used to drive the simulation component to move, for example, driving the simulation component to move along a first motion trajectory, wherein the first motion trajectory may be determined based on a first action sequence, the first action sequence may include first time information and first pose information. The camera module can be used to acquire images, for example, acquiring display images of the display device under test during the movement of the simulation component along the first motion trajectory. The controller may be communicatively connected to the drive motor and the camera module, for example. The controller may be configured to: determine a second motion trajectory for a preset time period based on the display images, and determine the delay of the display device under test within the preset time period based on the first and second motion trajectories.

[0292] In other examples, the second motion trajectory may be determined based on a second action sequence obtained from the displayed image, which may include second timing information and second pose information.

[0293] In other examples, the controller can obtain a first motion trajectory for a preset time period based on a first action sequence, and control the drive motor to drive the analog component to move along the first motion trajectory.

[0294] The controller of the delay detection equipment can execute the aforementioned delay test method 6000. All the aforementioned units, modules, and components can be implemented by corresponding processors or other hardware devices.

[0295] Figure 33 illustrates a schematic block diagram of an example electronic device 7000 that can be used to implement embodiments of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0296] As shown in Figure 33, the electronic device 7000 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0297] Multiple components in electronic device 7000 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 7000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0298] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as delay measurement methods, calibration methods, or delay testing methods. For example, in some embodiments, the delay measurement method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 7000 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the delay measurement method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the delay measurement method by any other suitable means (e.g., by means of firmware).

[0299] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0300] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0301] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0302] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0303] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0304] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0305] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0306] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A time delay measurement system, characterized by, include: A motion platform is fixedly connected to the extended reality device under test and drives the extended reality device under test to move synchronously. A vision platform is fixedly connected to the extended reality device under test and acquires multiple consecutive images of a target displayed by the extended reality device under test; as well as The processing platform obtains the visual pose sequence of the extended reality device under test based on the multiple consecutive images, and determines the delay of the extended reality device under test based on the visual pose sequence and the motion pose sequence of the motion platform. The processing platform also includes: The splitting unit, based on the motion feature parameters of the synchronized motion, splits the visual pose sequence and the motion pose sequence into multiple motion segments respectively; The fitting unit fits multiple motion segments of the visual pose sequence into multiple first time-pose curves, and fits multiple motion segments of the motion pose sequence into multiple second time-pose curves; and The delay unit determines the delay of the extended reality device under test based on the difference between the first time-pose curve and the second time-pose curve that corresponds to the first time-pose curve in time.

2. The system of claim 1, wherein, The motion characteristic parameters of the synchronous motion include at least one of the velocity of the synchronous motion, the acceleration of the synchronous motion, and the derivative of the acceleration of the synchronous motion.

3. The system of claim 1, wherein, The delay unit further includes: The first calculation subunit determines a first statistical index based on the difference between the first time-pose curve and the second time-pose curve that corresponds to the first time-pose curve in time. The first statistical index includes at least one of the following: the delay extreme value, the delay average value, the delay median, and the delay standard deviation of the extended reality device under test.

4. The system of claim 3, wherein, The delay unit further includes: The second calculation subunit determines the second statistical index based on at least one of the mean and standard deviation of the first statistical index obtained from multiple measurements.

5. The system of claim 1, wherein, The processing platform also includes: The preprocessing unit performs zero-bias preprocessing on the visual pose sequence and the motion pose sequence before fitting multiple motion segments of the visual pose sequence and multiple motion segments of the motion pose sequence, respectively.

6. The system of claim 1, wherein, The system also includes: A time platform, including a clock source, wherein the clock source is used to synchronize the motion platform and the vision platform in time before and during the synchronized motion.

7. The system of claim 1, wherein, The processing platform also includes: The first conversion unit extracts the pixel coordinates of the feature points of the target plate based on the multiple consecutive images; The second conversion unit converts the pixel coordinates into a preliminary pose sequence in a first coordinate system referenced by the visual platform, based on the preset spatial geometric constraints between the feature points; and The calibration unit converts the initial pose sequence into the visual pose sequence in a second coordinate system with the motion platform as a reference.

8. The system of claim 7, wherein, The processing platform also includes: The time correction unit performs time correction on the initial pose sequence before converting it into the visual pose sequence.

9. The system of claim 7, wherein, The processing platform also includes: The image processing unit processes the plurality of consecutive images before extracting the pixel coordinates of feature points of the target plate based on the plurality of consecutive images. The processing includes at least one of downsampling processing and Gaussian blur processing.

10. The system of claim 1, wherein, The target is a virtual target, which includes a visual marker calibration board in the visual reference library.

11. The system of claim 1, wherein, The visual pose sequence includes multiple first time-pose information items arranged in chronological order. Each first time-pose information item includes first time information and a first attitude angle corresponding to the first time information. The first attitude angle includes a first roll angle, a first yaw angle, and a first pitch angle. The motion pose sequence includes multiple second time-pose information arranged in chronological order. The second time-pose information includes second time information and a second attitude angle corresponding to the second time information. The second attitude angle includes a second roll angle, a second yaw angle, and a second pitch angle.

12. A method of measuring delay, characterized by, include: Acquire the motion pose sequence of the motion platform, wherein the motion platform is fixedly connected to the extended reality device under test and drives the extended reality device under test to move synchronously; Acquire multiple consecutive images of a target displayed via the extended reality device under test; The visual pose sequence of the extended reality device under test is obtained based on the multiple consecutive images. as well as The delay of the extended reality device under test is determined based on the visual pose sequence and the motion pose sequence. Determining the latency of the extended reality device under test based on the visual pose sequence and the motion pose sequence includes: Based on the motion characteristic parameters of the synchronized motion, the visual pose sequence and the motion pose sequence are respectively divided into multiple motion segments; The multiple motion segments of the visual pose sequence are respectively fitted into multiple first time-pose curves, and the multiple motion segments of the motion pose sequence are respectively fitted into multiple second time-pose curves; and The delay of the extended reality device under test is determined based on the difference between the first time-pose curve and the second time-pose curve that corresponds to the first time-pose curve in time.

13. The method of claim 12, wherein, The motion characteristic parameters of the synchronous motion include at least one of the velocity of the synchronous motion, the acceleration of the synchronous motion, and the derivative of the acceleration of the synchronous motion.

14. The method of claim 12, wherein, Based on the difference between the first time-pose curve and the second time-pose curve that corresponds to the first time-pose curve in time, the delay of the extended reality device under test is determined by: Based on the difference between the first time-pose curve and the second time-pose curve that corresponds to the first time-pose curve in time, a first statistical index is determined, wherein the first statistical index includes at least one of the following: the delay extreme value, the delay average value, the delay median, and the delay standard deviation of the extended reality device under test.

15. The method of claim 14, wherein, Determining the latency of the extended reality device under test based on the difference between the first time-pose curve and the second time-pose curve that corresponds to the first time-pose curve in time further includes: The second statistical indicator is determined based on at least one of the mean and standard deviation of the first statistical indicator obtained from multiple measurements.

16. The method of claim 12, wherein, Before fitting multiple motion segments of the visual pose sequence and multiple motion segments of the motion pose sequence, the method further includes: Zero-bias preprocessing is performed on the visual pose sequence and the motion pose sequence respectively.

17. The method according to claim 12, wherein, The method of acquiring the multiple consecutive images using a vision platform further includes: performing time synchronization between the motion platform and the vision platform before and during the synchronized motion.

18. The method according to claim 12, wherein, Obtaining the visual pose sequence of the extended reality device under test based on the multiple consecutive images includes: Based on the multiple consecutive images, the pixel coordinates of the feature points of the target plate are extracted; Based on the preset spatial geometric constraints between the feature points, the pixel coordinates are converted into a preliminary pose sequence in a first coordinate system referenced to the visual platform that captured the multiple consecutive images; and The initial pose sequence is converted into the visual pose sequence in a second coordinate system with the motion platform as a reference.

19. The method according to claim 18, wherein, The method further includes: performing time correction on the initial pose sequence before converting the initial pose sequence into the visual pose sequence.

20. The method according to claim 18, wherein, The method further includes: processing the plurality of consecutive images before extracting the pixel coordinates of the feature points of the target based on the plurality of consecutive images, the processing including at least one of downsampling processing and Gaussian blur processing.

21. The method according to claim 12, wherein, The target is a virtual target, which includes a visual marker calibration board in the visual reference library.

22. The method according to claim 12, wherein, The visual pose sequence includes multiple first time-pose information items arranged in chronological order. Each first time-pose information item includes first time information and a first attitude angle corresponding to the first time information. The first attitude angle includes a first roll angle, a first yaw angle, and a first pitch angle. The motion pose sequence includes multiple second time-pose information arranged in chronological order. The second time-pose information includes second time information and a second attitude angle corresponding to the second time information. The second attitude angle includes a second roll angle, a second yaw angle, and a second pitch angle.

23. A calibration system, characterized in that, include: A motion component is fixedly connected to the extended reality device and drives the extended reality device to move synchronously. A camera assembly, fixedly connected to the extended reality device, acquires multiple images of a target displayed via the extended reality device during the synchronized motion. as well as The processing component calibrates the extrinsic parameters of the camera component based on the plurality of images and the motion pose sequence of the motion component.

24. The system according to claim 23, wherein, The processing component includes: The first conversion unit extracts the pixel coordinates of the feature points of the target plate based on the multiple images; The second conversion unit converts the pixel coordinates into a visual pose sequence based on the preset spatial geometric constraints between the feature points; and The calibration unit calibrates the extrinsic parameters of the camera component based on the visual pose sequence and the motion pose sequence.

25. The system according to claim 24, wherein, The calibration unit includes: The linear transformation subunit, through a direct linear transformation method, obtains the initial values ​​of the extrinsic parameters of the camera component based on the visual pose sequence and the motion pose sequence; and The nonlinear transformation subunit optimizes the initial values ​​of the extrinsic parameters of the camera component based on minimizing the reprojection error using a nonlinear transformation method, thereby obtaining the extrinsic parameters of the camera component.

26. The system according to claim 24, wherein, The visual pose sequence includes multiple first time-pose information items arranged in chronological order. Each first time-pose information item includes first time information and a first attitude angle corresponding to the first time information. The first attitude angle includes a first roll angle, a first yaw angle, and a first pitch angle. The motion pose sequence includes multiple second time-pose information arranged in chronological order. The second time-pose information includes second time information and a second attitude angle corresponding to the second time information. The second attitude angle includes a second roll angle, a second yaw angle, and a second pitch angle.

27. The system according to claim 23, wherein, The system also includes: A timing component includes a clock source, wherein the clock source is used to time-synchronize the motion component and the camera component before and during the synchronized motion.

28. The system according to claim 23, wherein, During the synchronized motion, the camera assembly captures images of the markers displayed on the augmented reality device at the moment the augmented reality device passes a waypoint.

29. The system according to claim 24, wherein, During the synchronized motion, the camera assembly continuously captures images of the target displayed via the extended reality device.

30. The system according to claim 29, wherein, The processing component further includes: The time correction unit performs time correction on the visual pose sequence.

31. The system according to claim 30, wherein, The time correction unit includes: An angular velocity calculation subunit determines a first angular velocity based on the motion pose sequence and a second angular velocity based on the visual pose sequence. The cross-correlation subunit, based on the cross-correlation method, aligns the first angular velocity and the second angular velocity to determine the temporal bias of the visual pose sequence; and The correction subunit performs time correction on the visual pose sequence based on the time offset.

32. The system according to claim 31, wherein, The time correction unit further includes: The interpolation subunit performs interpolation processing on the visual pose sequence before determining the second angular velocity, so that the sampling frequency of the motion pose sequence and the visual pose sequence are the same.

33. The system according to claim 31, wherein, The time correction unit further includes: The filtering subunit performs filtering processing on the visual pose sequence before interpolating it.

34. The system according to claim 24, wherein, The processing component further includes: The image processing unit processes the plurality of images before extracting the pixel coordinates of feature points of the target plate based on the plurality of images. The processing includes at least one of downsampling processing and Gaussian blur processing.

35. The system according to claim 23, wherein, The target is a virtual target, which includes a visual marker calibration board in the visual reference library.

36. The system according to claim 23, wherein, The processing component further includes: The adjustment unit adjusts the positional relationship between the camera assembly and the extended reality device so that multiple feature points specified in the target plate are displayed in the multiple images.

37. A calibration method, characterized in that, include: Acquire the motion pose sequence of the motion component, wherein the motion component is fixedly connected to the extended reality device and drives the extended reality device to move synchronously; Using a camera assembly fixedly connected to the extended reality device, multiple images of a target displayed via the extended reality device are captured during the synchronized motion. as well as Based on the multiple images and the motion pose sequence of the motion component, the extrinsic parameters of the camera component are calibrated.

38. The calibration method according to claim 37, wherein, Based on the multiple images and the motion pose sequence of the motion component, the extrinsic parameters of the camera component are calibrated as follows: Based on the multiple images, the pixel coordinates of the feature points of the target plate are extracted; Based on the preset spatial geometric constraints between the feature points, the pixel coordinates are converted into a visual pose sequence; and Based on the visual pose sequence and the motion pose sequence, the extrinsic parameters of the camera component are calibrated.

39. The calibration method according to claim 38, wherein, Based on the visual pose sequence and the motion pose sequence, the extrinsic parameters of the camera component are calibrated as follows: The initial values ​​of the extrinsic parameters of the camera component are obtained using a direct linear transformation method based on the visual pose sequence and the motion pose sequence; and The initial values ​​of the extrinsic parameters of the camera component are optimized by using a nonlinear transformation method based on minimizing the reprojection error, thereby obtaining the extrinsic parameters of the camera component.

40. The calibration method according to claim 38, wherein, The visual pose sequence includes multiple first time-pose information items arranged in chronological order. Each first time-pose information item includes first time information and a first attitude angle corresponding to the first time information. The first attitude angle includes a first roll angle, a first yaw angle, and a first pitch angle. The motion pose sequence includes multiple second time-pose information arranged in chronological order. The second time-pose information includes second time information and a second attitude angle corresponding to the second time information. The second attitude angle includes a second roll angle, a second yaw angle, and a second pitch angle.

41. The calibration method according to claim 37, wherein, The calibration method further includes: Using a clock source, the motion component and the camera component are time-synchronized before and during the synchronized motion.

42. The calibration method according to claim 37, wherein, During the synchronized motion, acquiring multiple images of the target displayed via the extended reality device includes: During the synchronized motion, the camera assembly captures images of the markers displayed on the augmented reality device at the moment the augmented reality device passes a waypoint.

43. The calibration method according to claim 38, wherein, During the synchronized motion, acquiring multiple images of the target displayed via the extended reality device includes: During the synchronized motion, the camera assembly continuously captures images of the target displayed via the extended reality device.

44. The calibration method according to claim 43, wherein, The calibration method further includes: performing time correction on the visual pose sequence.

45. The calibration method according to claim 44, wherein, The time correction of the visual pose sequence includes: The first angular velocity is determined based on the motion pose sequence, and the second angular velocity is determined based on the visual pose sequence; Based on a cross-correlation method, the first angular velocity and the second angular velocity are aligned to determine the temporal bias of the visual pose sequence; and The visual pose sequence is time-corrected based on the time offset.

46. ​​The calibration method according to claim 45, wherein, The time correction of the visual pose sequence also includes: Before determining the second angular velocity, the visual pose sequence is interpolated to make the sampling frequency of the motion pose sequence and the visual pose sequence the same.

47. The calibration method according to claim 45, wherein, The time correction of the visual pose sequence also includes: Before interpolating the visual pose sequence, the visual pose sequence is filtered.

48. The calibration method according to claim 38, wherein, The calibration method further includes: processing the plurality of images before extracting the pixel coordinates of the feature points of the target plate based on the plurality of images, wherein the processing includes at least one of downsampling processing and Gaussian blur processing.

49. The calibration method according to claim 37, wherein, The target is a virtual target, which includes a visual marker calibration board in the visual reference library.

50. The calibration method according to claim 37, wherein, The calibration method further includes: Before acquiring multiple images of a target displayed via the extended reality device, the positional relationship between the camera assembly and the extended reality device is adjusted so that multiple feature points specified in the target are shown in the multiple images.

51. A delay testing method, characterized in that, include: In response to receiving a first action sequence, the simulation component obtains a first motion trajectory within a preset time period based on the first action sequence, wherein the first action sequence includes first time information and first pose information; A second motion sequence is obtained based on the displayed image of the simulated component moving according to the first motion trajectory, and a second motion trajectory for the preset time period is determined based on the second motion sequence, wherein the second motion sequence includes second time information and second pose information; and The delay of the display device in the preset time period is determined based on the first motion trajectory and the second motion trajectory.

52. The method according to claim 51, wherein, The first motion trajectory of the simulated component within a preset time period is obtained based on the first action sequence, including: In response to the existence of a target action sequence in the database that corresponds to the first action sequence, the target motion trajectory corresponding to the target action sequence is determined as the first motion trajectory.

53. The method according to claim 51, wherein, The first motion trajectory of the simulated component within a preset time period is obtained based on the first action sequence, including: In response to the absence of a target action sequence corresponding to the first action sequence in the database, the pose information of the reference point at different time points during the user's action corresponding to the first action sequence is obtained; and The first motion trajectory is determined based on the pose information of the reference point at different time points.

54. The method according to claim 53, wherein, Acquire the pose information of the reference point at different time points during the user's actions corresponding to the first action sequence, including: Acquire pose information of multiple test points at different time points during the user's action corresponding to the first action sequence; and The pose information of the reference point at different time points is determined based on the pose information of multiple test points at different time points, wherein the reference point is the center point of the three-dimensional space formed by the multiple test points.

55. The method according to claim 53, wherein, Determining the first motion trajectory based on the pose information of the reference point at different time points includes: The pose information of the simulation component at different time points is determined based on the pose information of the reference point at different time points. The motion parameters of the simulated component at different time points are determined based on the pose information of the simulated component at different time points, wherein the motion parameters include one or more of the following: rotation angle, velocity, and acceleration of the simulated component; and A third motion trajectory is generated based on the pose information and motion parameters of the simulation component at different time points, and the third motion trajectory is determined as the first motion trajectory.

56. The method according to any one of claims 51 to 55, wherein, A second action sequence is obtained based on the displayed image of the simulated component moving along the first motion trajectory, including: Feature extraction is performed on the displayed image to obtain feature points, and the two-dimensional coordinates of the feature points at different time points are determined; Based on the identity of the feature point, determine the three-dimensional coordinates of the feature point at different time points, and construct two-dimensional coordinate-three-dimensional coordinate matching point pairs of the feature point at different time points; and The second action sequence is determined based on the two-dimensional coordinate-three-dimensional coordinate matching point pairs of the feature points at different time points.

57. The method according to claim 56, wherein, The second action sequence is determined based on the two-dimensional coordinate-three-dimensional coordinate matching point pairs of the feature points at different time points, including: Pose calculations are performed on the 2D-3D coordinate matching point pairs of the feature points at different time points to obtain the rotation matrix of the displayed image relative to the world coordinate system at different time points; and The second action sequence is determined based on the rotation matrix at different time points.

58. The method according to claim 56, wherein, Before performing feature extraction on the displayed image to obtain feature points, the method further includes: The displayed image is subjected to Gaussian blur processing.

59. The method according to any one of claims 51 to 55, wherein, Determining the delay of the display device in the preset time period based on the first motion trajectory and the second motion trajectory includes: Linear fitting and / or quadratic curve fitting are performed on the first motion trajectory and the second motion trajectory; and The delay of the display device in the preset time period is determined based on the fitted first motion trajectory and the second motion trajectory.

60. The method according to claim 59, wherein, Performing linear fitting and / or quadratic curve fitting on the first motion trajectory and the second motion trajectory includes: Identify the linear and nonlinear intervals of the first and second motion trajectories; and Linear fitting is performed on the linear interval, and quadratic curve fitting is performed on the nonlinear interval.

61. The method according to claim 59, wherein, The first motion trajectory and the second motion trajectory are time-pose curves. Determining the delay of the display device within the preset time period based on the fitted first motion trajectory and the second motion trajectory includes: The delay curve of the display device is obtained based on the time difference between the fitted second motion trajectory and the first motion trajectory under the same pose; and The delay of the display device in the preset time period is determined based on the delay curve of the display device.

62. The method according to claim 61, wherein, The time-pose curve is a time-attitude angle curve. The delay curve of the display device is obtained based on the time difference between the fitted second motion trajectory and the first motion trajectory under the same pose, including: The delay curve of the display device is obtained by the time difference between the second motion trajectory and the first motion trajectory fitted under the same attitude angle, wherein the attitude angle includes roll angle, yaw angle and pitch angle.

63. A delay detection device, characterized in that, include: Simulation components are used to simulate user actions; A drive motor is used to drive the simulated component to move according to a first motion trajectory, wherein the first motion trajectory is determined based on a first action sequence including first time information and first pose information; A camera module is used to acquire display images of the display device under test during the movement of the simulated component according to the first motion trajectory; and The controller is communicatively connected to the drive motor and the camera module. The controller is configured to: determine a second motion trajectory for a preset time period based on the displayed image, and determine the delay of the display device under test during the preset time period based on the first motion trajectory and the second motion trajectory, wherein the second motion trajectory is determined based on a second action sequence obtained from the displayed image, and the second action sequence includes second time information and second pose information.

64. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, causes the at least one processor to perform the delay measurement method of any one of claims 12-22, the calibration method of any one of claims 37-50, or the delay test method of any one of claims 51-62.

65. A computer-readable storage medium storing a computer program, characterized in that, The computer instructions are used to cause the computer to execute the delay measurement method of any one of claims 12-22, the calibration method of any one of claims 37-50, or the delay test method of any one of claims 51-62.

66. A computer program product comprising a computer program that, when executed by a processor, implements the delay measurement method according to any one of claims 12-22, the calibration method according to any one of claims 37-50, or the delay test method according to any one of claims 51-62.