Calibration method, device, mixed reality device, system, storage medium and product
Through an optical see-through head-mounted device with a monocular camera and a fixed depth target, a preset calibration algorithm is used to calculate the transformation matrix between the eyeball and the monocular camera and the target, which solves the problem of cumbersome and high cost of mixed reality device calibration, and achieves the effect of simplifying the calibration process and reducing costs.
Patent Information
- Application Number
- CN202310597192.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-05-22
AI Technical Summary
Existing calibration methods for mixed reality devices are cumbersome and costly, relying on high-cost equipment such as infrared eye trackers and RGB-D cameras, and the system design is complex.
An optical see-through head-mounted device using a monocular camera and a fixed depth target obtains the target imaging point of the target point and uses a preset calibration algorithm to calculate the transformation matrix between the eyeball, monocular camera and target, simplifying the calibration process and reducing costs.
A simplified calibration process for mixed reality devices is achieved, which reduces device costs and improves calibration efficiency and simplicity.
Smart Images

Figure CN116612197B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mixed reality technology, and in particular to a calibration method, apparatus, mixed reality device, system, storage medium and product. Background Art
[0002] Mixed Reality (MR) technology is usually used to achieve deep integration of physical space and virtual objects.
[0003] Currently, mixed reality devices using mixed reality technology utilize relatively expensive auxiliary equipment, such as infrared eye trackers, RGB-D cameras, and depth cameras. Furthermore, these devices often require complex system designs, employing complex algorithms such as visual inertial navigation fusion and binocular vision positioning. Furthermore, for mixed reality devices without eye trackers, calibration is often required to achieve optimal display quality.
[0004] However, the traditional calibration method is complicated and costly. Summary of the Invention
[0005] Based on this, it is necessary to provide a calibration method, device, mixed reality device, system, computer-readable storage medium and computer program product that can reduce the cost of mixed reality equipment and simplify the calibration process to address the above technical problems.
[0006] In a first aspect, the present application provides a calibration method. This method is applied to a mixed reality device, wherein the mixed reality device includes a monocular camera; the method comprises:
[0007] Obtain target imaging points corresponding to target points on the mixed reality device; generate multiple sets of target point pairs based on the target imaging points corresponding to the target points; input the multiple sets of target point pairs into a preset calibration algorithm for calibration to obtain a target extrinsic parameter matrix of the eyeball; the target extrinsic parameter matrix of the eyeball is used to represent the transformation matrix from the eyeball to the target;
[0008] When multiple target point pairs are aligned, the target extrinsic parameter matrix of the monocular camera is determined based on the actual position information of each target point on the target and the target position information of each target point in the aligned target image taken by the monocular camera. The target extrinsic parameter matrix of the monocular camera is used to represent the transformation matrix from the monocular camera to the target.
[0009] According to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball, the target transformation matrix between the monocular camera and the eyeball is obtained.
[0010] In one embodiment, obtaining target imaging points corresponding to target points on a mixed reality device includes:
[0011] Obtaining an initial intrinsic parameter matrix of the eyeball, an initial extrinsic parameter matrix of the monocular camera, and an initial transformation matrix between the monocular camera and the eyeball; the initial extrinsic parameter matrix of the monocular camera is determined based on an initial target image taken by the monocular camera before alignment;
[0012] Based on the initial intrinsic parameter matrix of the eyeball, the initial extrinsic parameter matrix of the monocular camera, and the initial transformation matrix between the monocular camera and the eyeball, the coordinates of each target point on the target are transformed to generate the target imaging points corresponding to each target point on the mixed reality device.
[0013] In one embodiment, obtaining an initial intrinsic parameter matrix of the eye, an initial extrinsic parameter matrix of the monocular camera, and an initial transformation matrix between the monocular camera and the eye includes:
[0014] Obtaining the lens resolution, pixel size, preset pupil distance, and preset exit pupil distance of the mixed reality device; determining an initial intrinsic parameter matrix of the eyeball based on the lens resolution, pixel size, preset pupil distance, and preset exit pupil distance of the mixed reality device;
[0015] Controlling the monocular camera to photograph the target to obtain an initial target image, and determining the initial position information of each target point in the initial target image; generating an initial extrinsic parameter matrix of the monocular camera based on the initial position information of each target point in the initial target image, the actual position information of each target point, and the intrinsic parameter matrix of the monocular camera;
[0016] Obtain a preset relative position relationship between the monocular camera and the eyeball; based on the preset relative position relationship, determine the initial transformation matrix between the monocular camera and the eyeball.
[0017] In one embodiment, multiple sets of target point pairs are input into a preset calibration algorithm for calibration to generate a target extrinsic parameter matrix of the eyeball, including:
[0018] Input multiple target point pairs into a preset calibration algorithm to generate a calibration equation group, solve the calibration equation group, and obtain a perspective transformation matrix; the perspective transformation matrix is composed of the target intrinsic parameter matrix of the eyeball and the target extrinsic parameter matrix of the eyeball;
[0019] The perspective transformation matrix is decomposed to generate the target intrinsic parameter matrix and the target extrinsic parameter matrix of the eyeball.
[0020] In one embodiment, the method further comprises:
[0021] Control the monocular camera to obtain real-time target images of the target;
[0022] According to the actual position information of each target point on the target and the real-time position information of each target point in the real-time target image, the real-time external parameter matrix of the monocular camera is generated.
[0023] In one embodiment, the method further comprises:
[0024] Get the projection transformation matrix of the virtual object;
[0025] Perform coordinate conversion on the virtual object's position information based on the virtual object's projection transformation matrix, the eye's target intrinsic parameter matrix, and the target transformation matrix between the monocular camera and the eye, to determine the virtual object's position information on the mixed reality device;
[0026] Based on the position information of the virtual object on the mixed reality device, the virtual object is imaged on the mixed reality device.
[0027] In one embodiment, if the projection transformation matrix of the virtual object is the transformation matrix between the virtual object and the target, coordinate transformation is performed on the position information of the virtual object according to the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball to determine the position information of the virtual object on the mixed reality device, including:
[0028] According to the transformation matrix between the virtual object and the target, the target intrinsic parameter matrix of the eyeball, the target transformation matrix between the monocular camera and the eyeball, and the real-time extrinsic parameter matrix of the monocular camera, the position information of the virtual object is transformed into coordinates to determine the position information of the virtual object on the mixed reality device.
[0029] In one embodiment, if the projection transformation matrix of the virtual object is the transformation matrix between the virtual object and the monocular camera, coordinate transformation is performed on the position information of the virtual object according to the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball to determine the position information of the virtual object on the mixed reality device, including:
[0030] According to the transformation matrix between the virtual object and the monocular camera, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball, the position information of the virtual object is transformed to determine the position information of the virtual object on the mixed reality device.
[0031] In one embodiment, a target extrinsic parameter matrix of the monocular camera is generated based on the actual position information of each target point on the target and the target position information of each target point in the aligned target image captured by the monocular camera, including:
[0032] When each target point is aligned with the target imaging point pair corresponding to each target point, the monocular camera is controlled to shoot the target to obtain a target image;
[0033] Obtaining target position information of each target point in the target image;
[0034] The target extrinsic parameter matrix of the monocular camera is generated according to the target position information of each target point in the target image, the actual position information of each target point on the target and the intrinsic parameter matrix of the monocular camera.
[0035] In a second aspect, the present application further provides a calibration device. This device is applied to a mixed reality device, wherein the mixed reality device includes a monocular camera; the device includes:
[0036] The first generation module is used to obtain target imaging points corresponding to target points on the mixed reality device; based on the target imaging points corresponding to the target points, multiple groups of target point pairs are generated, and the multiple groups of target point pairs are input into a preset calibration algorithm for calibration to obtain a target extrinsic parameter matrix of the eyeball; the target extrinsic parameter matrix of the eyeball is used to represent the transformation matrix from the eyeball to the target;
[0037] The second generation module is used to determine the target extrinsic parameter matrix of the monocular camera based on the actual position information of each target point on the target and the target position information of each target point in the aligned target image captured by the monocular camera when multiple groups of target point pairs are aligned; the target extrinsic parameter matrix of the monocular camera is used to represent the transformation matrix from the monocular camera to the target;
[0038] The third generation module is used to generate a target transformation matrix between the monocular camera and the eyeball according to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball.
[0039] In a third aspect, the present application further provides a mixed reality device. The mixed reality device includes a monocular camera, a memory, and a processor. The memory stores a computer program. The monocular camera is configured to track and photograph a target to obtain an image of the target. The processor, when executing the computer program, implements the steps of the calibration method described in the first aspect.
[0040] In a fourth aspect, the present application further provides a calibration system, comprising a target and the mixed reality device of the third aspect; at least six target points are set on the target;
[0041] The mixed reality device is used to obtain the actual position information of each target point on the target, and implement the steps of the calibration method in the first aspect based on the actual position information of each target point.
[0042] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the calibration method in the first aspect.
[0043] In a sixth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the calibration method in the first aspect.
[0044] The above-mentioned calibration method, apparatus, mixed reality device, system, storage medium and computer program product obtain target imaging points on the mixed reality device corresponding to each target point on the target; generate multiple groups of target point pairs based on each target point and the target imaging point corresponding to each target point, input the multiple groups of target point pairs into a preset calibration algorithm for calibration, and obtain the target extrinsic parameter matrix of the eyeball, that is, the transformation matrix from the eyeball to the target; and, when all the multiple groups of target point pairs are aligned, determine the target extrinsic parameter matrix of the monocular camera, that is, the transformation matrix from the monocular camera to the target, based on the actual position information of each target point on the target and the target position information of each target point in the aligned target image captured by the monocular camera; and then, based on the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball, generate the target transformation matrix between the monocular camera and the eyeball. When using the calibrating method of the present application to calibrate a mixed reality device, the transformation matrix from the monocular camera to the target can be determined with the help of the monocular camera on the mixed reality device, and then the transformation matrix from the eyeball to the target can be obtained by using a preset calibration algorithm through multiple sets of target points and imaging points; finally, the target transformation matrix from the monocular camera to the eyeball can be obtained through the transformation matrix from the eyeball to the target and the transformation matrix from the monocular camera to the target, thereby realizing the calibration of the mixed reality device; compared with traditional methods, only a monocular camera can be used to realize device calibration, the calibration cost is greatly reduced, and the calibration process is relatively simple. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A diagram illustrating an application environment of a calibration method according to an embodiment;
[0046] Figure 2 Schematic diagram of a flow chart of a calibration method in one embodiment;
[0047] Figure 3 A schematic diagram of a coordinate system of a camera, a target, and a head-mounted device in one embodiment;
[0048] Figure 4 is a schematic flow chart of a calibration method in another embodiment;
[0049] Figure 5 Schematic diagram of object-image point pair matching in one embodiment;
[0050] Figure 6 is a schematic flow chart of a calibration method in another embodiment;
[0051] Figure 7 is a schematic flow chart of a calibration method in another embodiment;
[0052] Figure 8 is a schematic flow chart of a calibration method in another embodiment;
[0053] Figure 9 Schematic diagram of a coordinate system for anchoring and following a virtual object in one embodiment;
[0054] Figure 10 A schematic diagram of a specific flow chart of a calibration method in one embodiment;
[0055] Figure 11 is a structural block diagram of a calibration device in one embodiment;
[0056] Figure 12 2 is a diagram of the internal structure of a mixed reality device in one embodiment. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0058] Mixed Reality (MR) technology is typically used to achieve a deep fusion of physical space and virtual objects. Through deep interaction between virtual objects and the real physical space, virtual objects can mimic the way the human eye observes real objects, forming images on MR lenses. This creates an indistinguishable "blending" of virtual and real. Existing MR devices are highly component-dependent, relying on costly devices such as infrared eye trackers, RGB-D cameras, and binocular cameras to achieve MR effects. Furthermore, the calibration process for existing MR devices is complex.
[0059] For a single application scenario, such as a surgical operation scenario, the equipment cost is high when using existing mixed reality equipment to achieve mixed reality effects, and the process of calibrating the mixed reality equipment is relatively complicated.
[0060] Based on this, this application proposes a calibration method that can achieve mixed reality effects by relying only on a monocular camera, an optical perspective head-mounted device with a fixed depth target and corresponding calibration operations, which can reduce equipment costs; in addition, this application simplifies the operational process of the calibration scheme, which not only makes the calibration process simpler, but also shortens the calibration time, thereby improving calibration efficiency.
[0061] The calibration method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, a user wears a mixed reality device 120 and uses a monocular camera 122 on the mixed reality device 120 to capture a depth target 140 in real space for device calibration. The mixed reality device 120 may be a head-mounted mixed reality device.
[0062] In one embodiment, Figure 2 As shown, a calibration method is provided, which is applied to Figure 1 Taking the mixed reality device 120 in FIG. 1 as an example, the method includes the following steps:
[0063] Step 201: Obtain target imaging points corresponding to target points on the mixed reality device; generate multiple sets of target point pairs based on the target imaging points corresponding to the target points, input the multiple sets of target point pairs into a preset calibration algorithm for calibration, and obtain the target extrinsic parameter matrix of the eyeball.
[0064] Among them, the target extrinsic parameter matrix of the eyeball is used to represent the transformation matrix from the eyeball to the target. A target for calibration of the mixed reality device is set in the real physical space, and multiple target points are set on the target. When the user wears the mixed reality device to observe the target, each target point on the target will be projected onto the lens of the mixed reality device, and an imaging point corresponding to each target point will be obtained. The target can be a three-dimensional depth target, with at least 6 target points set on the target, and at least two target points are located on different planes. The coordinates of each target point can be a three-dimensional spatial coordinate in the real physical space.
[0065] For example, during calibration, the mixed reality device can randomly generate multiple imaging points, and the coordinates of each imaging point can be the two-dimensional plane coordinates in the coordinate system of the mixed reality device lens; then, each imaging point is output and displayed on the lens of the mixed reality device, and the user wears the mixed reality device and aligns each imaging point on the lens with each target point on the target through translation, rotation and other movement operations; when it is determined that each imaging point and each target point are aligned, a mapping relationship between each target point and each imaging point is generated to obtain the target imaging point corresponding to each target point.
[0066] Then, based on this mapping relationship, the target point and the corresponding target imaging point can be used as a set of target point pairs, thereby obtaining multiple sets of target point pairs. Furthermore, based on these multiple sets of target point pairs, the transformation matrix between the eye and the target, i.e., the target extrinsic parameter matrix of the eye, can be determined. For example, the multiple sets of target point pairs can be input into a preset calibration algorithm, which calculates and outputs the target extrinsic parameter matrix of the eye.
[0067] For example, when obtaining the target imaging points corresponding to each target point on the target on the mixed reality device, it is also possible to calculate the imaging point on the lens of the mixed reality device corresponding to the target point based on the actual physical coordinates of each target point using the imaging principle, thereby obtaining the target imaging point corresponding to each target point, and generating multiple groups of target point pairs based on each target point and the target imaging point corresponding to each target point; further, based on the multiple groups of target point pairs, the target extrinsic parameter matrix of the eyeball is calculated.
[0068] Step 202 : When all target point pairs are aligned, determine the target extrinsic parameter matrix of the monocular camera based on the actual position information of each target point on the target and the target position information of each target point in the aligned target image captured by the monocular camera.
[0069] Among them, the target extrinsic parameter matrix of the monocular camera is used to represent the transformation matrix from the monocular camera to the target.
[0070] For example, after the user wears the mixed reality device, he can align each target point with the target imaging point displayed on the lens by moving. When each target point and the target imaging point pair corresponding to each target point are aligned, the monocular camera on the mixed reality device can be triggered to shoot the target, thereby obtaining a target image of the target in its state; then, the mixed reality device can perform image analysis on the target image to obtain the target position information of each target point in the target image in the target image; the target position information can be the two-dimensional plane coordinates of the target point in the target image in the imaging coordinate system of the monocular camera.
[0071] At this time, when the target position information in the target image taken by the monocular camera corresponding to each target point in the real physical space is obtained, the actual position information of each target point in the real physical space and the target position information corresponding to each target point can be input into the above-mentioned preset calibration algorithm for calibration, so that the transformation matrix between the monocular camera and the target can be obtained, that is, the target extrinsic parameter matrix of the monocular camera.
[0072] In one implementation, after determining the target position information in the target image captured by the monocular camera corresponding to each target point in the real physical space, the target extrinsic parameter matrix of the monocular camera can also be generated based on the target position information of each target point in the target image, the actual position information of each target point on the target, and the intrinsic parameter matrix of the monocular camera. The intrinsic parameter matrix of the monocular camera can be determined based on the inherent properties of the monocular camera itself. Based on the imaging principle, the conversion relationship between the physical point in space and the imaging point can be expressed by formula (1) or a variation of formula (1).
[0073] P i=K·T·P w (1)
[0074] Among them, P i represents an imaging point, and in this embodiment, can be used to represent the target position information of the target point in the target image; P w represents a physical point, and in this embodiment, can be used to represent the actual position information of the target point in the real physical space; K represents the intrinsic parameter matrix of the camera, and T represents the extrinsic parameter matrix of the camera.
[0075] Given the internal parameter matrix K of the monocular camera and the actual position information P of multiple sets of target points w And the target position information P of each target point in the target image i In the case of , the target external parameter matrix T of the monocular camera can be calculated by the above formula (1).
[0076] Step 203 : Generate a target transformation matrix between the monocular camera and the eyeball according to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball.
[0077] refer to Figure 3 As shown, it shows the coordinate transformation relationship between the imaging coordinate system of the mixed reality device, the target coordinate system and the camera coordinate system of the monocular camera on the mixed reality device, wherein, Represents the transformation matrix from camera to target, that is, the target extrinsic parameter matrix of the camera, Represents the transformation matrix from camera to eyeball, Represents the transformation matrix from the eyeball to the target, that is, the target extrinsic parameter matrix of the eyeball.
[0078] Transformation matrix from camera to target Camera to eye transformation matrix and the transformation matrix from eyeball to target The following relationship exists:
[0079]
[0080] Therefore, the camera-to-eye transformation matrix is It can be expressed as:
[0081]
[0082] That is to say, after obtaining the target external parameter matrix of the monocular camera and the target extrinsic matrix of the eyeball In the case of , the target transformation matrix between the monocular camera and the eyeball can be calculated according to the above formula (3):
[0083] In the above calibration method, target imaging points corresponding to target points on the target are obtained on the mixed reality device; multiple groups of target point pairs are generated based on each target point and the target imaging point corresponding to each target point, and the multiple groups of target point pairs are input into a preset calibration algorithm for calibration to obtain the target extrinsic parameter matrix of the eyeball, that is, the transformation matrix from the eyeball to the target; and, when the multiple groups of target point pairs are aligned, the target extrinsic parameter matrix of the monocular camera, that is, the transformation matrix from the monocular camera to the target, is determined according to the actual position information of each target point on the target and the target position information of each target point in the aligned target image taken by the monocular camera; and then, according to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball, the target transformation matrix between the monocular camera and the eyeball is generated. When using the calibrating method of the present application to calibrate a mixed reality device, the transformation matrix from the monocular camera to the target can be determined with the help of the monocular camera on the mixed reality device, and then the transformation matrix from the eyeball to the target can be obtained by using a preset calibration algorithm through multiple sets of target points and imaging points; finally, the target transformation matrix from the monocular camera to the eyeball can be obtained through the transformation matrix from the eyeball to the target and the transformation matrix from the monocular camera to the target, thereby realizing the calibration of the mixed reality device; compared with traditional methods, only a monocular camera can be used to realize device calibration, the calibration cost is greatly reduced, and the calibration process is relatively simple.
[0084] In one embodiment, when the mixed reality device is calibrated, it can randomly generate multiple imaging points so that the user can align the randomly generated multiple imaging points with the target points on the target; it can also adopt the imaging principle to calculate the imaging points obtained after imaging on the lens of the mixed reality device based on the actual physical coordinates of the target point; illustratively, when using the imaging principle to calculate the imaging points corresponding to the target points, the imaging points of each target point on the target on the lens can be calculated based on the initial system parameters, the initial system parameters or the initial intrinsic parameter matrix of the eyeball generated by pre-calibration, the initial transformation matrix from the camera to the target, and the initial transformation matrix from the camera to the eyeball; then, the user can manually align the target points on the target and the imaging points on the lens to obtain the target imaging points corresponding to each target point in the aligned state. When this method is used to obtain the target imaging points corresponding to each target point, since the imaging points on the lens are multiple imaging points determined based on the eye intrinsic parameter matrix, the camera-to-target transformation matrix, and the camera-to-eye transformation matrix determined by preset system parameters, the error between the multiple imaging points and the actual target points is small, and it will be easier for the user to align each target point and the imaging point when performing manual alignment operations; compared with randomly generated imaging points, it can greatly improve the simplicity of the user's alignment operation and the efficiency of the alignment processing, which is conducive to achieving rapid calibration.
[0085] Figure 4This embodiment is a flow chart of a calibration method in another embodiment. This embodiment involves an optional implementation process of generating multiple imaging points by initial system parameters to obtain target imaging points corresponding to each target point on the mixed reality device. Based on the above embodiment, Figure 4 As shown, the above step 201 includes:
[0086] Step 401: Obtain an initial intrinsic parameter matrix of the eyeball, an initial extrinsic parameter matrix of the monocular camera, and an initial transformation matrix between the monocular camera and the eyeball.
[0087] The initial extrinsic parameter matrix of the monocular camera is determined based on the initial target image captured by the monocular camera before alignment. In other words, at the beginning of calibration, or before imaging points are set on the lens of the mixed reality device, the monocular camera is controlled to capture the target in real physical space to obtain the initial target image.
[0088] For example, after obtaining the initial target image, the initial target image can be analyzed to determine the initial position information of each target point in the initial target image; then, the mixed reality device can adopt the method for determining the target extrinsic parameter matrix of the monocular camera described in the above step 202 to generate the initial extrinsic parameter matrix of the monocular camera based on the initial position information of each target point in the initial target image, the actual position information of each target point in the real physical space and the intrinsic parameter matrix of the monocular camera.
[0089] In addition, for the initial intrinsic parameter matrix of the eyeball, the lens resolution, pixel size, preset pupil distance and preset exit pupil distance of the mixed reality device can be obtained, and the initial intrinsic parameter matrix of the eyeball can be determined based on the lens resolution, pixel size, preset pupil distance and preset exit pupil distance of the mixed reality device; wherein, the lens resolution and pixel size can be determined based on the attribute information of the mixed reality device, and the preset pupil distance and preset exit pupil distance can be determined based on the user's eye measurement information, or based on the pupil distance and exit pupil distance of multiple users; optionally, the mixed reality device can record the pupil distance and exit pupil distance of different users. When different users wear the mixed reality device, the preset pupil distance and preset exit pupil distance corresponding to the user identifier can be determined from the pre-stored different user information based on the user identifier, thereby improving the convenience and flexibility of different users using the same mixed reality device.
[0090] Furthermore, the initial transformation matrix between the monocular camera and the eyeball can be determined by obtaining a preset relative position relationship between the monocular camera and the eyeball, and based on this preset relative position relationship. The preset relative position relationship between the monocular camera and the eyeball can be determined by measuring the position information of the user's eye and the position information of the monocular camera when the user wears the mixed reality device; or it can be determined comprehensively based on the relative position relationship between the position information of the eyes of multiple users and the position information of the monocular camera, thereby obtaining an average parameter applicable to most users as the preset relative position relationship.
[0091] In step 402 , based on the initial intrinsic parameter matrix of the eyeball, the initial extrinsic parameter matrix of the monocular camera, and the initial transformation matrix between the monocular camera and the eyeball, coordinate transformation is performed on each target point on the target to generate target imaging points corresponding to each target point on the mixed reality device.
[0092] Referring to the above formula (1), the conversion relationship between the lens imaging point and the target point for the mixed reality device can be obtained, which can be expressed as:
[0093]
[0094] Among them, K ini Represented as the initial internal parameter matrix of the eyeball, It is expressed as the initial external parameter matrix of the eyeball, that is, the initial transformation matrix from the eyeball to the target.
[0095] Combined with the above formula (2), the formula (4) can be expressed as:
[0096]
[0097] in, Expressed as the initial external parameter matrix of the monocular camera, that is, the initial transformation matrix from the monocular camera to the target, Represented as the initial transformation matrix between the monocular camera and the eyeball.
[0098] Through the above formula (5), we can get the initial internal parameter matrix K based on the eyeball ini , the initial external parameter matrix of the monocular camera And the initial transformation matrix between the monocular camera and the eyeball For each target point P on the target w Perform coordinate transformation to generate target imaging points P corresponding to each target point on the mixed reality device i .
[0099] In addition, when using the above formula (5) to calculate the imaging point corresponding to each target point, multi-point synchronous acquisition can also be achieved, that is, the imaging points corresponding to multiple target points are calculated simultaneously to obtain multiple target point pairs; for example, for P in formula (5) w , which can be a coordinate matrix P including multiple target point coordinates w , through the coordinate transformation of formula (5), the coordinate matrix P including the coordinates of multiple imaging points can be obtained i , thereby generating imaging points corresponding to multiple target points at one time, realizing multi-point synchronous acquisition; for example: Figure 5 As shown, 9 target points can be set on the target, including 8 vertices and a center point. When the 9 target points on the target are projected onto the lens of the mixed reality device, 9 imaging points correspond to the lens. The coordinates of the imaging points are the two-dimensional plane coordinates in the lens uv coordinate system. i (u',v'), can also be expressed as P i (u', v', 1), the coordinates of the target point are the three-dimensional space coordinates of the target xyz coordinate system, through the coordinate P w (X w ,Y w ,Z w ) can also be expressed as P w (X w ,Y w ,Z w ,1); When using multi-point synchronous acquisition, the coordinates of the 9 target points can form a 9*4 coordinate matrix P w By performing coordinate transformation using the above formula (5), 9 imaging points corresponding to 9 target points can be obtained simultaneously, that is, a 9*2 coordinate matrix Pi is obtained, which includes the two-dimensional coordinates of the 9 imaging points.
[0100] Furthermore, when multiple target point pairs are obtained at one time, a preset calibration algorithm can be used to calculate the perspective transformation matrix of the virtual object in the physical space entering the human eye based on the multiple target point pairs, and then the target extrinsic parameter matrix of the eyeball can be obtained based on the perspective transformation matrix; that is, the user can directly obtain the target extrinsic parameter matrix of the eyeball through a one-step alignment operation, which not only simplifies the calibration process, but also improves the efficiency of calibration.
[0101] In this embodiment, an initial intrinsic parameter matrix of the eye, an initial extrinsic parameter matrix of the monocular camera, and an initial transformation matrix between the monocular camera and the eye are obtained. Then, based on the initial intrinsic parameter matrix of the eye, the initial extrinsic parameter matrix of the monocular camera, and the initial transformation matrix between the monocular camera and the eye, coordinate transformation is performed on each target point on the target to generate a target imaging point corresponding to each target point on the mixed reality device. The initial extrinsic parameter matrix of the monocular camera is determined based on an initial target image captured by the monocular camera before alignment. Because each target imaging point in the target imaging point determination method of this embodiment is determined based on the initial intrinsic parameter matrix of the eye, the initial camera-to-target transformation matrix, and the initial camera-to-eye transformation matrix, the error between the multiple target imaging points and the actual target points is small, making it easier for a user to align each target point with the target imaging point during manual alignment. Compared to randomly generated imaging points, this greatly improves the simplicity of user alignment operations and the efficiency of alignment processing, facilitating rapid calibration.
[0102] Figure 6 This embodiment involves inputting multiple sets of target point pairs into a preset calibration algorithm for calibration, and generating an optional implementation process of the target external parameter matrix of the eyeball. Figure 6 As shown, the above step 201 includes:
[0103] In step 601, multiple target point pairs are input into a preset calibration algorithm to generate a calibration equation group, and the calibration equation group is solved to obtain a perspective transformation matrix.
[0104] Among them, the perspective transformation matrix is composed of the target intrinsic parameter matrix of the eyeball and the target extrinsic parameter matrix of the eyeball.
[0105] Referring to the above formula (1), it can be seen that based on the imaging principle, by multiplying the target intrinsic parameter matrix of the eyeball and the target extrinsic parameter matrix of the eyeball, the conversion relationship between the physical point in space and the imaging point on the lens can be obtained. Based on this, assuming that the product of the target intrinsic parameter matrix of the eyeball and the target extrinsic parameter matrix of the eyeball is represented by the perspective transformation matrix G, the above formula (1) can be adjusted to:
[0106] P i =G·P w (6)
[0107] The coordinates of the imaging point are expressed as P i (u',v',1), the coordinates of the physical point are expressed as P w (X w ,Y w ,Z wIn the case of ,1), the perspective transformation matrix G should be a 3*4 transformation matrix. Therefore, there are 12 unknowns in the perspective transformation matrix G. A calibration equation can be constructed through a set of target point pairs, and two unknowns can be solved. Therefore, to solve the 12 unknowns, at least six sets of target point pairs are required.
[0108] For example, six groups of target point pairs can be input into the above formula (6) to generate a calibration equation group, and the perspective transformation matrix G can be calculated by solving the calibration equation group.
[0109] Step 602: Decompose the perspective transformation matrix to generate an eyeball target intrinsic parameter matrix and an eyeball target extrinsic parameter matrix.
[0110] Since the perspective transformation matrix G represents the product of the target intrinsic parameter matrix of the eyeball and the target extrinsic parameter matrix of the eyeball, after calculating the perspective transformation matrix G, the perspective transformation matrix G can be decomposed to obtain the target intrinsic parameter matrix of the eyeball and the target extrinsic parameter matrix of the eyeball.
[0111] In this embodiment, based on the pinhole camera imaging principle, multiple groups of target point pairs are input into a preset calibration algorithm to generate a calibration equation group, and the perspective transformation matrix is obtained by solving the calibration equation group; then, the perspective transformation matrix is decomposed to generate the target intrinsic parameter matrix and the target extrinsic parameter matrix of the eyeball; thereby realizing the calibration of the mixed reality device and obtaining accurate eyeball intrinsic and extrinsic parameters.
[0112] Figure 7 This embodiment is an optional implementation process of updating camera external parameters in real time. Figure 7 As shown, the above method also includes:
[0113] Step 701: Control a monocular camera to acquire a real-time target image of a target.
[0114] For example, after calibration is completed, when the user wears the mixed reality device to operate and experience it, the user's head wearing the mixed reality device will move. Then, compared with the relative position between the target and the mixed reality device during calibration, the relative position between the target and the mixed reality device will change in real time during the movement of the user's head. If the projection imaging is still based on the calibrated transformation matrices at this time, imaging misalignment may occur. Therefore, in order to ensure the accuracy of imaging when the mixed reality device moves, it is necessary to obtain in real time the transformation matrix between the user's eyeball and the target, that is, the real-time extrinsic parameter matrix of the eyeball, and the transformation matrix between the camera and the target, that is, the real-time extrinsic parameter matrix of the camera.
[0115] According to the above formula (2), the extrinsic parameter matrix of the eyeball can be determined based on the transformation matrix from the camera to the target, that is, the extrinsic parameter matrix of the camera, and the transformation matrix from the camera to the eyeball; when the user moves while wearing the mixed reality device, it can be assumed that there is no relative movement between the monocular camera on the mixed reality device and the user's head, that is, the target transformation matrix between the monocular camera and the user's eyeball remains unchanged by default; therefore, when obtaining the real-time extrinsic parameter matrix of the eyeball, it is necessary to first obtain the real-time extrinsic parameter matrix of the monocular camera, and then the real-time extrinsic parameter matrix of the eyeball can be determined based on the real-time extrinsic parameter matrix of the monocular camera and the target transformation matrix from the monocular camera to the eyeball.
[0116] Based on the relevant content description in the above step 202, the external parameter matrix of the monocular camera can be determined by the target image captured by the monocular camera on the mixed reality device tracking the target; therefore, after the calibration is completed, when performing projection imaging, the monocular camera can be controlled to obtain a real-time target image of the target.
[0117] Step 702 : Generate a real-time extrinsic parameter matrix of the monocular camera based on the actual position information of each target point on the target and the real-time position information of each target point in the real-time target image.
[0118] When a real-time target image of the target is acquired, the real-time position information of each target point in the real-time target image can be determined. Then, based on the real-time position information of each target point in the real-time target image and the actual position information of each target point in real physical space, a real-time extrinsic parameter matrix of the monocular camera is obtained. The specific implementation of determining the real-time extrinsic parameter matrix of the monocular camera based on the real-time position information of each target point in the real-time target image and the actual position information of each target point in real physical space can be referred to the relevant description in step 202 above and will not be repeated here.
[0119] Furthermore, the target extrinsic parameter matrix of the eyeball can be updated according to the real-time extrinsic parameter matrix of the monocular camera and the target transformation matrix between the monocular camera and the eyeball to generate the real-time extrinsic parameter matrix of the eyeball.
[0120] Referring to the above formula (2), the real-time extrinsic parameter matrix of the monocular camera is multiplied by the inverse of the target transformation matrix between the monocular camera and the eyeball to obtain the real-time extrinsic parameter matrix of the eyeball.
[0121] In this embodiment, during the real-time imaging process after calibration is completed, the monocular camera is controlled to acquire a real-time target image of the target. A real-time extrinsic parameter matrix of the monocular camera is generated based on the actual position information of each target point on the target and the real-time position information of each target point in the real-time target image. This allows for projected imaging based on the real-time extrinsic parameter matrix of the monocular camera and the target transformation matrix between the monocular camera and the eyeball when the relative position between the mixed reality device and the target changes, thereby improving imaging accuracy. Furthermore, the target extrinsic parameter matrix of the eyeball can be updated based on the real-time extrinsic parameter matrix of the monocular camera and the target transformation matrix between the monocular camera and the eyeball, thereby generating a real-time extrinsic parameter matrix of the eyeball. This allows for projected imaging based on the real-time extrinsic parameter matrix of the eyeball and the intrinsic parameter matrix of the eyeball, thereby improving imaging accuracy.
[0122] Figure 8 This embodiment relates to an optional implementation process of projection imaging of a mixed reality device. Based on the above embodiment, Figure 8 As shown, the above method also includes:
[0123] Step 801: Obtain the projection transformation matrix of the virtual object.
[0124] The virtual object can be projected at a fixed position in the real physical space, or it can be displayed at a fixed position on the lens according to the movement of the spatial scene.
[0125] In one case, assuming that the virtual object is fixedly projected at a certain position in the real physical space, such as: a virtual object is fixedly displayed at a preset position in the real physical space, at this time, the position of the virtual object in the real physical space can be represented by the position of the fixed target, that is, the projection transformation matrix of the virtual object can be expressed as the transformation matrix between the virtual object and the target For example, the transformation matrix between the virtual object and the target can be determined based on the position information of the virtual object in the real physical space and the position information of the fixed target in the real physical space. Get the projection transformation matrix of the virtual object.
[0126] In another case, it is assumed that the virtual object is fixedly displayed at a certain position on the lens of the mixed reality device, that is, the virtual object moves with the movement of the head, and the virtual object moves in the virtual space; for example, a virtual menu is projected and displayed at a preset position on the lens, and the virtual menu moves with the movement of the head; at this time, the imaging position of the virtual object on the lens can be represented by the position of the monocular camera, that is, the projection transformation matrix of the virtual object can be expressed as the transformation matrix between the virtual object and the monocular camera For example, the transformation matrix between the virtual object and the monocular camera can be determined based on the position information of the virtual object on the lens and the position information of the monocular camera. Get the projection transformation matrix of the virtual object.
[0127] Step 802 : Based on the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball, coordinate transformation is performed on the position information of the virtual object to determine the position information of the virtual object on the mixed reality device.
[0128] For example, refer to Figure 9 As shown, the virtual object can be a virtual axis-sagittal-crown three-view (MPR) displayed at a fixed position in the real physical space. In this case, the projection transformation matrix of the virtual object is the transformation matrix between the virtual object and the target, that is, In this example, since the virtual object is a virtual axis sagittal crown three-view (MPR), It can also be expressed as In this case, the position information of the virtual object can be converted into coordinates based on the transformation matrix between the virtual object and the target, the target intrinsic parameter matrix of the eyeball, the target transformation matrix between the monocular camera and the eyeball, and the real-time extrinsic parameter matrix of the monocular camera to determine the position information of the virtual object on the mixed reality device.
[0129] Assume that the point coordinates of the virtual axis-sagittal-crown three-view in the real physical space are P v , then the coordinate P of the target point in the real physical space is w It can be expressed as: in, Represents the transformation matrix between the virtual axis sagittal crown three-view image and the target.
[0130] Based on this, the virtual object point P corresponding to the virtual axis sagittal crown three-view v The projection on the lenses of a mixed reality device can be described as:
[0131]
[0132] Among them, K represents the target intrinsic parameter matrix of the eyeball, represents the real-time extrinsic parameter matrix of the monocular camera, Represents the target transformation matrix between the monocular camera and the eyeball.
[0133] For example, refer to Figure 9As shown in FIG, the virtual object can also be a virtual menu (VM) fixedly displayed on the lens, that is, the virtual menu moves in the virtual space along with the movement of the user's head, that is, the virtual menu is displayed with the movement; in this case, the projection transformation matrix of the virtual object is the transformation matrix between the virtual object and the monocular camera, that is, In this case, the position information of the virtual object can be converted into coordinates based on the transformation matrix between the virtual object and the monocular camera, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball to determine the position information of the virtual object on the mixed reality device.
[0134] Here, the polar coordinate system is used to determine the position coordinates of the virtual menu in the real physical space, that is, the transformation matrix between the virtual object and the monocular camera is used. To indicate that, at this time, the virtual object point P corresponding to the virtual menu v The projection on the lenses of a mixed reality device can be described as:
[0135]
[0136] Among them, K represents the target intrinsic parameter matrix of the eyeball, represents the target transformation matrix between the monocular camera and the eyeball, Represents the transformation matrix between the virtual object and the monocular camera.
[0137] Step 803: Based on the position information of the virtual object on the mixed reality device, image the virtual object on the mixed reality device.
[0138] In this embodiment, when imaging a virtual object, the virtual object's projection transformation matrix is first obtained. Next, based on the virtual object's projection transformation matrix, the eye's target intrinsic parameter matrix, and the target transformation matrix between the monocular camera and the eye, coordinate transformation is performed on the virtual object's position information to determine the virtual object's position information on the mixed reality device. Finally, based on the virtual object's position information on the mixed reality device, the virtual object is imaged on the mixed reality device. In this embodiment, the virtual object can be fixedly displayed in real physical space, achieving spatial anchoring, or it can move with the user's head movement, appearing as a fixed display at a preset position on the lens, achieving virtual object tracking. This meets diverse imaging needs and enhances imaging diversity and flexibility.
[0139] In one embodiment, a specific implementation process of a calibration method is provided. Figure 10 As shown, the calibration method includes the following steps:
[0140] Step 1: After the user puts on the head-mounted mixed reality device, the calibration process begins. The mixed reality device uses the initial intrinsic parameter matrix K of the eyeball ini , the initial external parameter matrix of the monocular camera And the initial transformation matrix between the monocular camera and the eyeball Calculate each target point P in the real physical space separately w Image point P on the lens i ,Right now Thus, each target point P is obtained w (X w ,Y w ,Z w ,1) The corresponding target imaging point P i (u,v,1), multiple target point pairs are obtained; then, each imaging point P i The output shows that multiple calibration points are obtained on the lens.
[0141] Optionally, a calibration frame can be formed by multiple calibration points, refer to the above Figure 5 shown.
[0142] Step 2: The user moves autonomously to align the calibration points (or calibration frames) on the lens with the target points (or targets) in the real physical space. When the alignment is confirmed, the user can manually confirm the alignment operation and trigger the monocular camera to obtain the target image in the aligned state. The monocular camera tracks and identifies the target to obtain the current position of the monocular camera, that is, the target extrinsic parameter matrix of the monocular camera.
[0143] Step 3: Input multiple sets of target point pairs into the preset calibration algorithm to determine the target intrinsic parameter matrix K and target extrinsic parameter matrix of the eyeball
[0144] Referring to the above description, when calculating the target internal parameter matrix K and target external parameter matrix of the eyeball , it is necessary to first calculate the perspective transformation matrix G, which contains 12 unknowns. However, each pair of image-object points can only provide one linear equation. Therefore, at least six point pairs are required to solve the intrinsic and extrinsic parameter matrices. In addition, considering sampling noise, the number of point pairs selected is usually more than six. In this case, the optimal intrinsic and extrinsic parameter matrices can be obtained through least squares optimization. At the same time, to prevent feature degradation, the selected target points cannot be completely coplanar, that is, at least one depth target point must exist.
[0145] Step 4: According to the target external parameter matrix of the monocular camera obtained in step 2 And the target external parameter matrix of the eyeball obtained in step 3 The above formula (3) is used to calculate the transformation matrix from camera to eyeball: Right now
[0146] It should be noted that, by executing the above steps 2-4, the target intrinsic parameter matrix of the first eye and the transformation matrix from the camera to the first eye can be determined.
[0147] Step 5: Repeat steps 2-4 above to obtain the target intrinsic parameter matrix of the second eye and the transformation matrix from the camera to the second eyeball.
[0148] Step 6: When the user moves while wearing the mixed reality device, the transformation matrix from the camera to the eyeball is assumed to be the same as the transformation matrix from the camera to the eyeball, assuming that there is no relative movement between the user's head and the mixed reality device. The target internal parameter matrix of the eyeball will not change, and the external parameter matrix of the eyeball will not change. At this time, the external parameter matrix of the eyeball needs to be updated in real time. Since the external parameter matrix of the eyeball can be transformed from the camera to the target matrix And the transformation matrix from camera to eyeball Therefore, when the user is wearing a helmet and performing activities, he only needs to track the head frame target according to the monocular camera to determine the real-time transformation matrix from the camera to the target. Then, the real-time transformation matrix from camera to target can be and the camera-to-eye target transformation matrix Update the external parameter matrix of the eyeball to obtain the real-time external parameter matrix of the eyeball
[0149] In step 7, considering the various projection requirements in real scenes, including spatial anchoring imaging of images such as axial-sagittal-coronal three-view (MPR) and holograms, and follow-up imaging of virtual menus, two projection methods can be provided in this embodiment; spatial anchoring imaging and follow-up imaging will be explained separately below.
[0150] Method 1: Spatial anchor imaging; obtain the mapping from virtual space to real space. The spatial coordinates of virtual objects such as virtual three-view image (MPR) or virtual hologram can be defined by a fixed head frame target, that is, the transformation matrix between the virtual object and the target is determined. Let the virtual object point be P v , then the target point P in the real physical space w It can be expressed as: The complete process of an anchored point in virtual space appearing on the head-mounted device lens can be described as follows:
[0151]
[0152] Method 2: Follow-up imaging; define the spatial coordinates of the virtual menu through the polar coordinate system and determine the transformation matrix between the virtual object and the monocular camera So the complete process of the virtual menu points appearing on the head-mounted device lenses can be described as:
[0153]
[0154] It should be noted that the above imaging process is an imaging process for a single eye. When imaging the left eye, the anchored and following virtual object imaging is achieved based on the target intrinsic parameter matrix of the left eye, the transformation matrix from the camera to the target, the transformation matrix from the camera to the left eye, the transformation matrix from the virtual object to the target, and the transformation matrix from the virtual object to the camera; when imaging the right eye, the anchored and following virtual object imaging is achieved based on the target intrinsic parameter matrix of the right eye, the transformation matrix from the camera to the target, the transformation matrix from the camera to the right eye, the transformation matrix from the virtual object to the target, and the transformation matrix from the virtual object to the camera.
[0155] The calibration method in this embodiment has the following advantages:
[0156] (1) Low equipment dependence
[0157] Compared with traditional solutions that rely on expensive equipment such as eye trackers and depth cameras to obtain eye parameters or assist in calibration, the calibration solution provided in this embodiment only requires a monocular camera and a fixed target to achieve calibration and imaging of a certain range of mixed reality devices.
[0158] (2) Low system complexity and easy to operate
[0159] Compared with the traditional single-point acquisition solution (i.e., aligning a group of imaging points and target points at a time) or the manually adjusted calibration target alignment solution, the calibration solution provided in this embodiment supports multi-point synchronous acquisition (i.e., determining multiple groups of target point pairs at a time to form a calibration frame). The user can obtain the perspective transformation matrix of the virtual object in the physical space entering the human eye with a one-step alignment operation.
[0160] (3) Optimize manual calibration scheme to improve accuracy and shorten operation time
[0161] Compared with the traditional manual calibration object-image point alignment process, which is slow or difficult to align, resulting in reduced accuracy, the calibration solution provided in this embodiment automatically initializes and generates virtual calibration points / calibration frames through the monocular camera and target system and initializes system parameters, thereby improving calibration accuracy while greatly shortening the calibration time.
[0162] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0163] Based on the same inventive concept, the present application also provides a calibration device for implementing the aforementioned calibration method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more calibration device embodiments provided below can be found in the above-mentioned limitations of the calibration method and will not be repeated here.
[0164] In one embodiment, Figure 11 As shown, a calibration device is provided, including: a first generation module 1101, a second generation module 1102 and a third generation module 1103, wherein:
[0165] The first generation module 1101 is configured to obtain target imaging points corresponding to target points on the mixed reality device; generate multiple sets of target point pairs based on the target imaging points corresponding to the target points; input the multiple sets of target point pairs into a preset calibration algorithm for calibration to obtain a target extrinsic parameter matrix of the eye; the target extrinsic parameter matrix of the eye is used to represent the transformation matrix from the eye to the target;
[0166] The second generation module 1102 is configured to determine a target extrinsic parameter matrix of the monocular camera based on actual position information of each target point on the target and target position information of each target point in the aligned target image captured by the monocular camera when all target point pairs are aligned. The target extrinsic parameter matrix of the monocular camera is used to represent a transformation matrix from the monocular camera to the target.
[0167] The third generating module 1103 is configured to obtain a target transformation matrix between the monocular camera and the eyeball according to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball.
[0168] In one embodiment, the first generation module 1101 includes a first acquisition submodule and a first generation submodule; the first acquisition submodule is used to obtain the initial intrinsic parameter matrix of the eyeball, the initial extrinsic parameter matrix of the monocular camera, and the initial transformation matrix between the monocular camera and the eyeball; the initial extrinsic parameter matrix of the monocular camera is determined based on the initial target image before alignment taken by the monocular camera; the first generation submodule is used to perform coordinate transformation on each target point on the target based on the initial intrinsic parameter matrix of the eyeball, the initial extrinsic parameter matrix of the monocular camera, and the initial transformation matrix between the monocular camera and the eyeball, and generate target imaging points corresponding to each target point on the mixed reality device.
[0169] In one embodiment, the first acquisition submodule includes a first determination unit, a second determination unit and a third determination unit; the first determination unit is used to obtain the lens resolution, pixel size, preset pupil distance and preset exit pupil distance of the mixed reality device; determine the initial intrinsic parameter matrix of the eyeball according to the lens resolution, pixel size, preset pupil distance and preset exit pupil distance of the mixed reality device; the second determination unit is used to control the monocular camera to shoot the target to obtain an initial target image, and determine the initial position information of each target point in the initial target image; generate the initial extrinsic parameter matrix of the monocular camera according to the initial position information of each target point in the initial target image, the actual position information of each target point and the intrinsic parameter matrix of the monocular camera; the third determination unit is used to obtain the preset relative position relationship between the monocular camera and the eyeball; based on the preset relative position relationship, determine the initial transformation matrix between the monocular camera and the eyeball.
[0170] In one embodiment, the first generation module 1101 also includes a first determination submodule and a second generation submodule; wherein the first determination submodule is used to input multiple groups of target point pairs into a preset calibration algorithm to generate a calibration equation group, solve the calibration equation group, and obtain a perspective transformation matrix; the perspective transformation matrix is composed of a target intrinsic parameter matrix of the eyeball and a target extrinsic parameter matrix of the eyeball; the second generation submodule is used to decompose the perspective transformation matrix to generate a target intrinsic parameter matrix of the eyeball and a target extrinsic parameter matrix of the eyeball.
[0171] In one embodiment, the device further includes: a control module and a fourth generation module; wherein the control module is used to control the monocular camera to obtain a real-time target image of the target; the fourth generation module is used to generate a real-time extrinsic parameter matrix of the monocular camera based on the actual position information of each target point on the target and the real-time position information of each target point in the real-time target image.
[0172] In one embodiment, the device also includes: an acquisition module, a determination module and an imaging module; wherein the acquisition module is used to obtain the projection transformation matrix of the virtual object; the determination module is used to perform coordinate conversion on the position information of the virtual object according to the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball, and determine the position information of the virtual object on the mixed reality device; the imaging module is used to image the virtual object on the mixed reality device based on the position information of the virtual object on the mixed reality device.
[0173] In one embodiment, if the projection transformation matrix of the virtual object is the transformation matrix between the virtual object and the target, the above-mentioned determination module is specifically used to perform coordinate conversion on the position information of the virtual object according to the transformation matrix between the virtual object and the target, the target intrinsic parameter matrix of the eyeball, the target transformation matrix between the monocular camera and the eyeball, and the real-time extrinsic parameter matrix of the monocular camera, so as to determine the position information of the virtual object on the mixed reality device.
[0174] In one embodiment, if the projection transformation matrix of the virtual object is the transformation matrix between the virtual object and the monocular camera, the above-mentioned determination module is specifically used to perform coordinate conversion on the position information of the virtual object according to the transformation matrix between the virtual object and the monocular camera, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball, so as to determine the position information of the virtual object on the mixed reality device.
[0175] In one embodiment, the second generation module 1102 includes a control submodule, a second acquisition submodule, and a third generation submodule; wherein the control submodule is used to control the monocular camera to shoot the target to obtain a target image when each target point and the target imaging point pair corresponding to each target point are aligned; the second acquisition submodule is used to obtain the target position information of each target point in the target image; the third generation submodule is used to generate a target extrinsic parameter matrix of the monocular camera based on the target position information of each target point in the target image, the actual position information of each target point on the target, and the intrinsic parameter matrix of the monocular camera.
[0176] Each module in the calibration device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0177] In one embodiment, a mixed reality device is provided, referring to Figure 1As shown, the mixed reality device includes a monocular camera, a memory and a processor, the memory stores a computer program, and the monocular camera is used to track and shoot the target to obtain the target image; when the processor executes the computer program, the steps of the calibration method in any of the above embodiments are implemented.
[0178] In one embodiment, a mixed reality device is provided, whose internal structure diagram can be as follows: Figure 12 As shown. The mixed reality device includes a processor, a memory, a communication interface, a display unit and an input device connected via a system bus; the display unit can be a display lens. The processor of the mixed reality device is used to provide computing and control capabilities. The memory of the mixed reality device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface of the mixed reality device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a calibration method is implemented. The display lens of the mixed reality device can be a liquid crystal display lens or an electronic ink display lens. The input device of the mixed reality device can be a touch layer covering the display lens, or a button, trackball or touchpad provided on the housing of the mixed reality device, or an external keyboard, touchpad or mouse.
[0179] Those skilled in the art will understand that Figure 12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the mixed reality device to which the solution of the present application is applied. The specific mixed reality device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0180] In one embodiment, a calibration system is provided, referring to Figure 1 As shown, the calibration system includes a target and the above-mentioned mixed reality device, wherein a monocular camera is provided on the mixed reality device; at least 6 target points are provided on the target, and the mixed reality device is used to obtain the actual position information of each target point on the target, and implement the steps of the calibration method in any of the above-mentioned embodiments based on the actual position information of each target point.
[0181] For example, when obtaining the actual physical coordinates of each target point on a target, the mixed reality device can capture the target using a monocular camera on the mixed reality device, thereby obtaining a target image. Subsequently, the mixed reality device can perform image recognition on the target in the target image to obtain the spatial physical coordinates of each target point on the target, thereby obtaining the actual physical coordinates of each target point. Furthermore, when the mixed reality device implements device calibration based on the actual position information of each target point, its specific implementation process can be referred to the relevant description of the various embodiments above and will not be repeated here.
[0182] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the calibration method in each of the above embodiments are implemented.
[0183] In one embodiment, a computer program product is provided, including a computer program, which implements the steps of the calibration method in each of the above embodiments when executed by a processor.
[0184] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0185] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0186] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0187] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A calibration method, characterized in that: Applied to a mixed reality device, the mixed reality device including a monocular camera; the method comprising: Obtaining target imaging points corresponding to target points on the mixed reality device; generating multiple sets of target point pairs based on the target points and the target imaging points corresponding to the target points; inputting the multiple sets of target point pairs into a preset calibration algorithm for calibration to obtain a target extrinsic parameter matrix of the eyeball; the target extrinsic parameter matrix of the eyeball is used to represent the transformation matrix from the eyeball to the target; When the multiple sets of target point pairs are aligned, determining a target extrinsic parameter matrix of the monocular camera according to actual position information of each target point on the target and target position information of each target point in the aligned target image taken by the monocular camera; the target extrinsic parameter matrix of the monocular camera is used to represent a transformation matrix from the monocular camera to the target; A target transformation matrix between the monocular camera and the eyeball is obtained according to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball.
2. The method according to claim 1, characterized in that The acquiring target imaging points corresponding to target points on the mixed reality device includes: Obtaining an initial intrinsic parameter matrix of the eyeball, an initial extrinsic parameter matrix of the monocular camera, and an initial transformation matrix between the monocular camera and the eyeball; the initial extrinsic parameter matrix of the monocular camera is determined based on an initial target image before alignment taken by the monocular camera; Based on the initial intrinsic parameter matrix of the eyeball, the initial extrinsic parameter matrix of the monocular camera and the initial transformation matrix between the monocular camera and the eyeball, coordinate transformation is performed on each target point on the target to generate a target imaging point corresponding to each target point on the mixed reality device.
3. The method according to claim 2, characterized in that The obtaining of the initial intrinsic parameter matrix of the eyeball, the initial extrinsic parameter matrix of the monocular camera, and the initial transformation matrix between the monocular camera and the eyeball includes: Obtaining a lens resolution, a pixel size, a preset pupil distance, and a preset exit pupil distance of the mixed reality device; determining an initial intrinsic parameter matrix of the eyeball according to the lens resolution, the pixel size, the preset pupil distance, and the preset exit pupil distance of the mixed reality device; Controlling the monocular camera to photograph the target to obtain an initial target image, and determining initial position information of each target point in the initial target image; generating an initial extrinsic parameter matrix of the monocular camera according to the initial position information of each target point in the initial target image, the actual position information of each target point, and the intrinsic parameter matrix of the monocular camera; Acquire a preset relative position relationship between the monocular camera and the eyeball; and determine an initial transformation matrix between the monocular camera and the eyeball based on the preset relative position relationship.
4. The method according to claim 1, wherein The plurality of target point pairs are input into a preset calibration algorithm for calibration to generate a target extrinsic parameter matrix of the eyeball, including: Inputting the plurality of target point pairs into a preset calibration algorithm to generate a calibration equation group, solving the calibration equation group to obtain a perspective transformation matrix; the perspective transformation matrix is composed of a target intrinsic parameter matrix of the eyeball and a target extrinsic parameter matrix of the eyeball; The perspective transformation matrix is decomposed to generate a target intrinsic parameter matrix of the eyeball and a target extrinsic parameter matrix of the eyeball.
5. The method according to claim 4, characterized in that The method further comprises: Controlling the monocular camera to acquire a real-time target image of the target; A real-time extrinsic parameter matrix of the monocular camera is generated according to the actual position information of each target point on the target and the real-time position information of each target point in the real-time target image.
6. The method according to claim 5, characterized in that The method further comprises: Get the projection transformation matrix of the virtual object; Performing coordinate transformation on the position information of the virtual object according to the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball to determine the position information of the virtual object on the mixed reality device; Based on the position information of the virtual object on the mixed reality device, the virtual object is imaged on the mixed reality device.
7. The method according to claim 6, characterized in that If the projection transformation matrix of the virtual object is a transformation matrix between the virtual object and the target, performing coordinate transformation on the position information of the virtual object according to the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball to determine the position information of the virtual object on the mixed reality device includes: According to the transformation matrix between the virtual object and the target, the target intrinsic parameter matrix of the eyeball, the target transformation matrix between the monocular camera and the eyeball, and the real-time extrinsic parameter matrix of the monocular camera, the position information of the virtual object is coordinate transformed to determine the position information of the virtual object on the mixed reality device.
8. The method according to claim 6, characterized in that If the projection transformation matrix of the virtual object is a transformation matrix between the virtual object and the monocular camera, performing coordinate transformation on the position information of the virtual object according to the projection transformation matrix of the virtual object, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball to determine the position information of the virtual object on the mixed reality device, including: According to the transformation matrix between the virtual object and the monocular camera, the target intrinsic parameter matrix of the eyeball, and the target transformation matrix between the monocular camera and the eyeball, the position information of the virtual object is converted into coordinates to determine the position information of the virtual object on the mixed reality device.
9. A calibration device, characterized in that: Applied to a mixed reality device, the mixed reality device includes a monocular camera; the device includes: A first generation module is configured to obtain target imaging points corresponding to target points on the mixed reality device; generate multiple sets of target point pairs based on the target points and the target imaging points corresponding to the target points; input the multiple sets of target point pairs into a preset calibration algorithm for calibration to obtain a target extrinsic parameter matrix of the eye; the target extrinsic parameter matrix of the eye is used to represent a transformation matrix from the eye to the target; a second generating module, configured to determine, when all the groups of target point pairs are aligned, a target extrinsic parameter matrix of the monocular camera based on actual position information of each target point on the target and target position information of each target point in the aligned target image captured by the monocular camera; the target extrinsic parameter matrix of the monocular camera is used to represent a transformation matrix from the monocular camera to the target; The third generating module is used to generate a target transformation matrix between the monocular camera and the eyeball according to the target extrinsic parameter matrix of the monocular camera and the target extrinsic parameter matrix of the eyeball.
10. A mixed reality device comprising a monocular camera, a memory, and a processor, wherein the memory stores a computer program, characterized in that: The monocular camera is used to track and shoot the target to obtain the target image; When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A calibration system, characterized in that: The mixed reality device comprises a target and the mixed reality device according to claim 10; the target is provided with at least 6 target points; The mixed reality device is used to obtain actual position information of each target point on the target, and implement the steps of the calibration method according to any one of claims 1 to 8 based on the actual position information of each target point.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Monocular visual error measurement system for cooperative target and error limit quantification method
CN104729534A
Calibration method and equipment for an optical perspective augmented reality display
CN109615664A