Virtual object positioning system and method for mixed reality technology

By using cameras and image recognition modules to identify and calculate the three-dimensional coordinates of spatial anchor points in mixed reality technology, the stability and accuracy of virtual object positioning in the prior art are solved, and high-precision and stable virtual object positioning are achieved, improving user experience.

CN120070838APending Publication Date: 2025-05-30SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411913213.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing spatial anchoring technology has stability and accuracy problems when achieving the precise placement of virtual objects.

Method used

The camera module is used to obtain real-world images, identify spatial anchor points through the image recognition module, and use the coordinate calculation module to calculate three-dimensional coordinates in real time based on the identified identifiers, and bind them to the virtual object to achieve real-time positioning.

Benefits of technology

It significantly improves the positioning accuracy and system stability of spatial anchor points, reduces the impact of environmental factors on positioning accuracy, and is applicable to multiple fields and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070838A_ABST
    Figure CN120070838A_ABST
Patent Text Reader

Abstract

The invention provides a virtual object positioning system and method for a mixed reality technology, and belongs to the technical field of electric digital data processing. The virtual object positioning system comprises a camera module, an image recognition module, a coordinate calculation module and a virtual object positioning module. The camera module is used for acquiring images of the real world. The image recognition module is used for recognizing a preset space anchor point identifier in the image. And the coordinate calculation module is used for calculating the three-dimensional coordinates of the space anchor point identifier in the real world in real time according to the identified space anchor point identifier. And the virtual object positioning module is used for binding the calculated three-dimensional coordinates with the virtual object so as to realize real-time positioning of the virtual object. The technical scheme of the invention has the technical effects of high positioning precision, high system stability, wide application scene and optimization of user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electrical digital data processing, and in particular to the technical field of realizing virtual object positioning in mixed reality by using spatial anchor technology. Background Art

[0002] With the development of mixed reality technology, how to accurately and stably place virtual objects in the real world has become a key issue. The spatial anchor technology provides a solution to achieve the accurate placement of virtual objects by setting fixed reference points, namely spatial anchors, in the real world. However, the existing spatial anchor technology uses gyroscopes for spatial positioning, but there are still problems in terms of stability and accuracy in implementation. Summary of the Invention

[0003] In order to overcome the above technical defects, a first aspect of the present invention provides a virtual object positioning system for mixed reality technology, which includes:

[0004] A camera module for acquiring images of the real world;

[0005] An image recognition module for recognizing preset spatial anchor identifiers in the images;

[0006] A coordinate calculation module for calculating the three-dimensional coordinates of the spatial anchor identifier in the real world in real time according to the recognized spatial anchor identifier;

[0007] A virtual object positioning module for binding the calculated three-dimensional coordinates to the virtual object, thereby realizing real-time positioning of the virtual object.

[0008] Further, the coordinate calculation module calculates the three-dimensional coordinates of the spatial anchor identifier in the real world by using the following formula:

[0009] P world = K -1 ·P image ·D

[0010] Where P world is the three-dimensional coordinates of the spatial anchor in the real world, K is the internal parameter matrix of the camera, K -1 is the inverse matrix of the internal parameter matrix, P image is the two-dimensional coordinates of the spatial anchor identifier in the image, D is the depth information obtained by a depth camera or other sensors, and "·" represents matrix multiplication.

[0011] Further, the internal parameter matrix is a 3x3 matrix or a 4x4 matrix.

[0012] Further, the content of the depth information includes: the distance from a point to the camera, three-dimensional coordinate information, and surface geometry information.

[0013] Further, the method by which the virtual object positioning module binds the calculated three-dimensional coordinates to the virtual object includes:

[0014] (1) Coordinate calculation: Use a sensor to calculate the three-dimensional coordinates (Xr, Yr, Zr) of the target object in the real world;

[0015] (2) Coordinate transformation: Transform the coordinates in the real world to the coordinate system of the virtual world through a coordinate transformation matrix. The coordinate transformation matrix defines the mapping relationship between the real world coordinate system and the virtual world coordinate system. Among them, the formula of the coordinate transformation matrix is:

[0016] XvYvZv = [RT]XrYrZr

[0017] where R is the rotation matrix, T is the translation vector, and (Xv, Yv, Zv) are the transformed virtual world coordinates;

[0018] (3) Binding implementation: Bind the virtual world coordinates to the virtual object by updating the transformation attributes of the virtual object.

[0019] It should be noted that the camera module and the image recognition module in this application can adopt any conventional settings commonly used in the MR field. This application does not particularly limit them. For example, the image recognition module adopts OpenCV, etc. The preset method of the spatial anchor identifier can adopt any conventional settings commonly used in the art. This application does not particularly limit it.

[0020] Exemplarily, R is a 3x3 rotation matrix, and T is a 3x1 translation vector.

[0021] The second aspect of this application provides a virtual object positioning method for mixed reality technology using the above-mentioned virtual object positioning system for mixed reality technology, which includes:

[0022] Step S1: Obtain an image of the real world;

[0023] Step S2: Identify the preset spatial anchor identifier in the image;

[0024] Step S3: According to the identified spatial anchor identifier, calculate the three-dimensional coordinates of the spatial anchor identifier in the real world in real time;

[0025] Step S4: Bind the calculated three-dimensional coordinates to the virtual object, thereby realizing real-time positioning of the virtual object.

[0026] Further, in step S3, the three-dimensional coordinates of the spatial anchor identifier in the real world are calculated using the following formula:

[0027] P world = K -1 ·P image ·D

[0028] where P world is the three-dimensional coordinate of the spatial anchor in the real world, K is the internal parameter matrix of the camera, and K -1 is the inverse matrix of the internal parameter matrix. P image is the two-dimensional coordinate of the spatial anchor identifier in the image, D is the depth information obtained by a depth camera or other sensors, and "·" represents matrix multiplication.

[0029] Further, in step S4, the method of binding the calculated three-dimensional coordinates to the virtual object includes:

[0030] (1) Coordinate calculation: Use a sensor to calculate the three-dimensional coordinates (Xr, Yr, Zr) of the target object in the real world;

[0031] (2) Coordinate transformation: Transform the coordinates in the real world to the coordinate system of the virtual world through a coordinate transformation matrix. The coordinate transformation matrix defines the mapping relationship between the real world coordinate system and the virtual world coordinate system. The formula for the coordinate transformation matrix is:

[0032] XvYvZv=[RT]XrYrZr

[0033] where R is the rotation matrix, T is the translation vector, and (Xv, Yv, Zv) are the transformed virtual world coordinates;

[0034] (3) Binding implementation: Bind the virtual world coordinates to the virtual object by updating the transformation attributes of the virtual object.

[0035] Further, the virtual object positioning method is applicable to the medical field, the education and training field, the architectural design field, the remote collaboration field, and the game field.

[0036] After adopting the above technical solutions, compared with the prior art, the following beneficial effects are achieved:

[0037] The present invention provides a virtual object positioning system and a virtual object positioning method for mixed reality technology through an improved spatial anchor technology. Tests in actual applications show that the present invention has the following beneficial effects:

[0038] (1)Improve positioning accuracy: The present invention significantly improves the positioning accuracy of spatial anchor points by introducing depth information (D) and combining it with the internal parameter matrix (K) of the camera to perform precise three-dimensional coordinate calculations. This method reduces the influence of environmental factors (such as light changes, angle changes, etc.) on the positioning accuracy.

[0039] (2)Improve system stability: By accurately identifying spatial anchor point identifiers through image recognition technology, the system of the present invention can maintain a high degree of stability in complex environments and accurately identify and locate spatial anchor points even in dynamically changing scenarios.

[0040] (3)Wide range of application scenarios: The spatial anchor point technology of the present invention is not only applicable to MR games, but also can be applied to multiple fields such as education, training, and architectural design preview, realizing the seamless integration of virtual content and the real world. This technology has broad application prospects and great commercial value, and is expected to promote the further development and popularization of mixed reality technology.

[0041] (4)Optimize user experience: By improving positioning accuracy and stability, the present invention significantly enhances the immersion and interactivity of users in the mixed reality experience, thus optimizing the user experience. Brief Description of the Drawings

[0042] Figure 1 It is a module structure diagram of the virtual object positioning system for mixed reality technology in this application;

[0043] Figure 2 It is a flowchart of the virtual object positioning method for mixed reality technology in this application. Detailed Embodiments

[0044] The advantages of the present invention are further elaborated below in conjunction with the drawings and specific embodiments. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.

[0045] Example 1 Application Example in the Medical Field

[0046] During the surgical operation, the virtual object positioning method for mixed reality technology includes the following steps:

[0047] Step S1.1: Obtain an image of the real world;

[0048] Step S1.2: Identify the preset spatial anchor point identifier in the image;

[0049] Step S1.3: According to the identified spatial anchor point identifier, calculate the three-dimensional coordinates of the spatial anchor point identifier in the real world in real time;

[0050] The three-dimensional coordinates of the spatial anchor identifier in the real world are calculated using the following formula:

[0051] P world = K -1 ·P image ·D

[0052] where P world is the three-dimensional coordinate of the spatial anchor in the real world, K is the internal parameter matrix of the camera, and K -1 is the inverse matrix of the internal parameter matrix. P image is the two-dimensional coordinate of the spatial anchor identifier in the image, D is the depth information obtained through a depth camera or other sensors, and "·" represents matrix multiplication.

[0053] This formula is used to convert the point P image in the image back to the point P world in the three-dimensional world coordinates. This formula is derived based on the imaging principle of the camera. During the camera imaging process, points in the three-dimensional world are mapped onto a two-dimensional image through the internal and external parameters of the camera (not involved here). This formula is the inverse process of the imaging process, that is, given a point on the image, depth information, and the internal parameters of the camera, the corresponding point in the three-dimensional world can be calculated.

[0054] The Intrinsic Matrix is an important part of camera parameters. It contains parameters related to the characteristics of the camera itself and is mainly used to describe the internal optical and geometric characteristics of the camera. The intrinsic matrix is usually a 3x3 matrix, but in practical applications, to handle homogeneous coordinates, it is sometimes extended to a 4x4 matrix (the last row is usually [0 0 0 1]). The intrinsic matrix mainly includes the following: (1) Focal Length: fx, fy: respectively represent the focal lengths of the camera in the x-axis and y-axis directions, with the unit of pixels. The focal length is the distance from the optical center of the camera lens to the image plane. However, in the intrinsic matrix, the focal length is measured in pixels because the camera ultimately outputs an image composed of pixels. Example: For a camera, fx in its intrinsic matrix may be 800 pixels, and fy may be 600 pixels, indicating that the focal lengths of the camera in the x-axis and y-axis directions are different, which is usually due to the design of the camera sensor and lens. (2) Principal Point Coordinates: cx, cy: The principal point is the intersection of the camera optical axis and the image plane. In the intrinsic matrix, it represents the coordinates of the image center (ideally) in the pixel coordinate system. Example: The principal point coordinate cx may be half of the image width (assuming the image width is 1920 pixels, then cx is 960), and cy may be half of the image height (assuming the image height is 1080 pixels, then cy is 540). However, in practical applications, due to errors in the camera manufacturing and installation processes, the principal point coordinates may deviate slightly. (3) Axis Skew Parameter: s (sometimes also called alpha): Ideally, the x-axis and y-axis of the image coordinate system should be orthogonal. However, in reality, due to the design of the camera sensor and lens, there may be a certain tilt angle between the x-axis and y-axis, which is represented by the s parameter in the intrinsic matrix. However, in many camera models, s is usually assumed to be 0, that is, the x-axis and y-axis are considered orthogonal.

[0055] Example of an ideal intrinsic matrix in the ideal case:

[0056] Without considering the axis skew parameter s, an ideal intrinsic matrix K can be expressed as:

[0057]

[0058] This matrix projects points in the three-dimensional camera coordinate system onto the two-dimensional image coordinate system and takes into account the focal length and principal point offset of the camera. The intrinsic matrix is a parameter that is fixed when the camera leaves the factory (unless the camera hardware changes, such as replacing the lens or sensor). Therefore, in related applications of computer vision and robotics, it is usually necessary to calibrate the camera first to obtain an accurate intrinsic matrix.

[0059] The content of depth information includes: (1) Distance from a point to the camera: This is the most basic and intuitive content of depth information, that is, the straight-line distance from each point in the scene to the camera. In machine vision, depth images are usually used to represent this distance information, where each pixel value in the image corresponds to the distance from the corresponding point in the scene to the camera. (2) Three-dimensional coordinate information: In addition to simple distance information, depth information can be further extended to the three-dimensional coordinate information of each point in the scene. This means that in addition to knowing the distance from a point to the camera, we can also know the specific position of these points in three-dimensional space (including the values on the X, Y, and Z axes). (3) Surface geometry information: Depth images also describe the geometric information of the three-dimensional scene surface, such as the shape, size, and surface concavity of objects. These information are crucial for tasks such as three-dimensional reconstruction and scene understanding. Example: Application of depth images: In the field of autonomous driving, vehicles obtain depth images of the surrounding environment through mounted cameras and sensors, so as to be able to perceive the distance and position information of objects such as roads, pedestrians, and vehicles, and then make accurate driving decisions. Three-dimensional reconstruction: In the three-dimensional reconstruction task, first, depth images of the scene are collected from multiple perspectives, and then the three-dimensional coordinate information in these depth images is used to reconstruct the three-dimensional model of the scene. This model can be used in multiple fields such as virtual reality, game development, and architectural design. Augmented reality: In augmented reality applications, depth information is used to accurately place virtual objects in the appropriate positions in the real scene.

[0060] Step S1.4: Bind the calculated three-dimensional coordinates to the virtual object, so as to realize the real-time positioning of the virtual object.

[0061] The methods for binding the calculated three-dimensional coordinates to the virtual object include:

[0062] (1) Coordinate calculation

[0063] First of all, the system needs to calculate the three-dimensional coordinates of the target object in the real world. This is usually achieved through some form of sensor or tracking system, such as using a depth camera, lidar (LiDAR), or optical tracker, etc.

[0064] Suppose we already have the three-dimensional coordinates (Xr, Yr, Zr) of the target in the real world.

[0065] (2) Coordinate transformation

[0066] Next, it is necessary to transform the coordinates in the real world into the coordinate system of the virtual world. This usually involves a coordinate transformation matrix, which defines the mapping relationship between the real-world coordinate system and the virtual-world coordinate system.

[0067] The coordinate transformation can be expressed as:

[0068] XvYvZv = [RT]XrYrZr

[0069] Preferably, R is a 3x3 rotation matrix, T is a 3x1 translation vector, and (Xv, Yv, Zv) are the transformed virtual world coordinates.

[0070] (3) Binding implementation

[0071] Once the coordinates in the virtual world are available, they can be bound to virtual objects. This is typically achieved by updating the transformation properties (such as position, rotation) of the virtual objects.

[0072] For example, in a 3D game engine such as Unity, the coordinates can be bound by setting the position property of the virtual object's Transform component: virtualObject.transform.position = new Vector3(X_v, Y_v, Z_v).

[0073] Exemplarily, the application of the technical solution of this application in the surgical operation process includes the following stages:

[0074] 1. Spatial scanning and three-dimensional positioning of the MR helmet

[0075] During the operation, the doctor wears an MR helmet, which has powerful spatial scanning and three-dimensional modeling functions. Through the high-precision sensors built into the helmet, the system will scan the surgical operation area in real time, generate an accurate three-dimensional model, and combine it with the pre-acquired imaging data (such as CT, MRI) to create a precise virtual image. This virtual image can be perfectly aligned with the patient's anatomical structure, helping the doctor to clearly observe the target area, such as the key areas for resection and suture, during the operation.

[0076] This real-time spatial scanning and positioning technology greatly reduces the frequent adjustment requirements of relying on intraoperative imaging equipment in traditional surgeries. Through the MR helmet, the doctor can immediately obtain a clear view at any operation step, without the need to pause the operation for repositioning or inspection.

[0077] 2. Precise registration of virtual projection and real anatomical structure

[0078] Another core technology of the MR helmet is the precise registration function of virtual projection and real anatomical structure. Before the operation starts, the doctor can upload the pre-prepared imaging data into the MR system to generate an individualized virtual model of the patient. These virtual models include important anatomical structures, such as the positions of blood vessels and nerves.

[0079] During the operation, the MR helmet uses real-time spatial scanning and manual adjustment functions to accurately align the virtual model with the patient's actual anatomical structure one-to-one. Doctors can not only observe the virtual model from a fixed perspective, but also switch between multiple perspectives by moving the helmet to view the correspondence between the virtual image and the real organ at different angles. This free observation and operation capability enables doctors to always maintain accurate spatial perception in complex surgical scenarios.

[0080] For example, during a surgical operation, the MR helmet directly superimposes the virtual images of the surgical site and surrounding tissues (such as blood vessels and nerves) on the actual operation area, ensuring that the doctor can avoid important anatomical structures and accurately treat the target site. This precise registration reduces the risk of misoperation and reduces damage to surrounding normal tissues.

[0081] 3. Dynamic adjustment and free movement during surgery

[0082] The introduction of MR technology makes dynamic adjustments during surgery simpler and more flexible. In traditional surgery, the doctor's operating angle is limited to fixed cameras or displays, and assistants are usually required to adjust the equipment to obtain different viewing angles. MR helmets break this limitation, and doctors can adjust the projection position and viewing angle of virtual images at any time by simply moving their heads.

[0083] This dynamic adjustment function allows doctors to move freely during surgery. Whether looking down, sideways or from more complex angles, doctors can clearly see the matching effect between virtual images and actual anatomical structures. In this way, doctors no longer rely on the fixed perspective of external devices, and can quickly make accurate judgments and operations. This is especially important for complex surgeries involving multiple anatomical levels, because doctors can continuously observe and fine-tune the operation path at different angles to ensure high precision and low risk of surgery.

[0084] 4. Real-time update and interaction of virtual images

[0085] MR helmets not only provide static virtual images, but also support real-time image updates. During the operation, as the doctor operates, the system automatically adjusts the display content of the virtual image. For example, when the doctor starts to cut tissue or suture, the virtual image will change in real time to reflect the changes in the anatomical structure after the operation. This real-time interactive capability allows doctors to obtain the latest progress of the operation at any time and make further operation plans based on the current situation.

[0086] This interactivity is particularly important in minimally invasive surgery. Through real-time virtual image feedback, doctors can confirm the accuracy of the surgery, the damage to surrounding tissues, and the next operation direction. This feedback mechanism ensures the accuracy of each operation step, reduces the uncertainty during the surgery, and improves the safety and success rate of the surgery.

[0087] Embodiment 2: Using MR for teaching

[0088] In the field of education and training, teachers can use spatial anchor technology to display 3D models in the classroom, enabling students to more intuitively understand complex concepts or structures.

[0089] For example, when conducting MR teaching, the teacher needs to display a virtual solid geometry model from multiple angles in the real world. Based on the virtual object positioning system and its virtual object positioning method for mixed reality technology provided in this application, the following steps are included:

[0090] Step 2.1: Capture images of the real world through the camera of the MR device;

[0091] Step 2.2: Identify the pre-set spatial anchor markers in the image through image recognition technology (such as OpenCV, etc.);

[0092] Step 2.3: According to the identified spatial anchor markers, the coordinate calculation module uses the following formula to calculate the three-dimensional coordinates of the spatial anchor markers in the real world in real time;

[0093] P world = K -1 ·P image ·D

[0094] Where P world is the three-dimensional coordinate of the point of the spatial anchor in the real world; K is the internal parameter matrix of the camera. Exemplarily, it is a 3×3 matrix that contains internal parameters such as the focal length and optical center of the camera; K -1 is the inverse matrix of the internal parameter matrix; P image is the two-dimensional coordinate of the point of the spatial anchor marker in the image, represented here in homogeneous coordinates (i.e., adding one dimension, usually 1) for matrix multiplication; D is the depth information of the point obtained through a depth camera or other sensors, usually a scalar representing the distance of the point from the camera; "·" represents matrix multiplication.

[0095] This formula is used to convert the point P image in the image back to the point P world in the three-dimensional world coordinates。This formula is derived based on the imaging principle of the camera. During the camera imaging process, points in the three-dimensional world are mapped to the two-dimensional image through the camera's intrinsic and extrinsic parameters (not involved here). This formula is the inverse process of the imaging process, that is, given a point on the image, depth information, and the camera's intrinsic parameters, the corresponding point in the three-dimensional world can be calculated.

[0096] The intrinsic matrix is an important part of the camera parameters. It contains parameters related to the characteristics of the camera itself and is mainly used to describe the internal optical and geometric characteristics of the camera. The intrinsic matrix is usually a 3x3 matrix, but in practical applications, to handle homogeneous coordinates, it is sometimes extended to a 4x4 matrix (the last row is usually [0 0 0 1]). The intrinsic matrix mainly includes the following: (1) Focal Length: fx, fy: respectively represent the focal lengths of the camera in the x-axis and y-axis directions, and the unit is pixels. The focal length is the distance from the optical center of the camera lens to the image plane. However, in the intrinsic matrix, the focal length is measured in pixels because the camera finally outputs an image, and the image is composed of pixels. Example: For a camera, fx in its intrinsic matrix may be 800 pixels, and fy may be 600 pixels, indicating that the focal lengths of the camera in the x-axis and y-axis directions are different, which is usually due to the design of the camera sensor and lens. (2) Principal Point Coordinates: cx, cy: The principal point is the intersection of the camera optical axis and the image plane. In the intrinsic matrix, it represents the coordinates of the image center (ideally) in the pixel coordinate system. Example: The principal point coordinate cx may be half of the image width (assuming the image width is 1920 pixels, then cx is 960), and cy may be half of the image height (assuming the image height is 1080 pixels, then cy is 540). However, in practical applications, due to errors in the camera manufacturing and installation processes, the principal point coordinates may deviate slightly. (3) Axis Skew Parameter: s (sometimes also called alpha): Ideally, the x-axis and y-axis of the image coordinate system should be orthogonal. However, in reality, due to the design of the camera sensor and lens, there may be a certain tilt angle between the x-axis and y-axis, and this tilt angle is represented by the s parameter in the intrinsic matrix. However, in many camera models, s is usually assumed to be 0, that is, the x-axis and y-axis are considered orthogonal.

[0097] Example of the intrinsic matrix in the ideal case:

[0098] Without considering the axis skew parameter s, an ideal intrinsic matrix K can be expressed as:

[0099]

[0100] This matrix projects points in the three-dimensional camera coordinate system onto the two-dimensional image coordinate system, taking into account the camera's focal length and principal point offset. The intrinsic matrix is a parameter that is fixed when the camera leaves the factory (unless there are changes to the camera hardware, such as changing the lens or sensor). Therefore, when conducting relevant applications in computer vision and robotics, it is usually necessary to calibrate the camera first to obtain the accurate intrinsic matrix.

[0101] The content of depth information includes: (1) Distance from the point to the camera: This is the most basic and intuitive content of depth information, that is, the straight-line distance from each point in the scene to the camera. In machine vision, depth images are usually used to represent this distance information, where each pixel value of the image corresponds to the distance from the corresponding point in the scene to the camera. (2) Three-dimensional coordinate information: In addition to simple distance information, depth information can be further extended to the three-dimensional coordinate information of each point in the scene. This means that in addition to knowing the distance from the point to the camera, one can also know the specific position of these points in three-dimensional space (including the values on the X, Y, and Z coordinate axes). (3) Surface geometry information: The depth image also describes the surface geometry information of the three-dimensional scene, such as the shape, size, and surface concavity of objects. This information is crucial for tasks such as three-dimensional reconstruction and scene understanding. Examples: Applications of depth images: In the field of autonomous driving, vehicles obtain depth images of the surrounding environment through onboard cameras and sensors, enabling them to perceive the distance and position information of objects such as roads, pedestrians, and vehicles, and then make accurate driving decisions. Three-dimensional reconstruction: In the three-dimensional reconstruction task, depth images of the scene are first collected from multiple perspectives, and then the three-dimensional coordinate information in these depth images is used to reconstruct the three-dimensional model of the scene. Such models can be used in multiple fields such as virtual reality, game development, and architectural design. Augmented reality: In augmented reality applications, depth information is used to accurately place virtual objects in the appropriate positions in the real scene. For example, in an AR game, players can see scenes where virtual characters or items are seamlessly integrated with the real environment, thanks to the accurate perception of the distance of each point in the scene by depth information.

[0102] Step 2.4: The virtual object positioning module binds the calculated three-dimensional coordinates to the virtual solid geometry model, thereby achieving the precise positioning of the virtual solid geometry model in the real world.

[0103] The methods by which the virtual object positioning module binds the calculated three-dimensional coordinates to the virtual object (virtual solid geometry model) include:

[0104] (1) Coordinate calculation

[0105] First, the system needs to calculate the three-dimensional coordinates of the target object in the real world. This is usually achieved through some form of sensor or tracking system, such as using a depth camera, lidar (LiDAR), or optical tracker, etc.

[0106] Suppose we already have the three-dimensional coordinates (Xr, Yr, Zr) of the target in the real world.

[0107] (2) Coordinate transformation

[0108] Next, it is necessary to transform the coordinates in the real world into the coordinate system of the virtual world. This usually involves a coordinate transformation matrix, which defines the mapping relationship between the real-world coordinate system and the virtual-world coordinate system.

[0109] The coordinate transformation can be expressed as:

[0110] XvYvZv = [RT]XrYrZr

[0111] Preferably, R is a 3x3 rotation matrix, T is a 3x1 translation vector, and (Xv, Yv, Zv) are the transformed virtual-world coordinates.

[0112] (3) Binding implementation

[0113] Once the coordinates in the virtual world are available, these coordinates can be bound to virtual objects. This is usually achieved by updating the transformation properties (such as position, rotation) of the virtual objects.

[0114] For example, in a 3D game engine such as Unity, the coordinates can be bound by setting the position property of the virtual object's Transform component: virtualObject.transform.position = new Vector3(X_v, Y_v, Z_v).

[0115] Example 3 Using MR for architectural design preview

[0116] Architectural design preview: Architects can use the technology of the present invention to preview virtual architectural models in a real environment, so as to better evaluate the feasibility and visual effects of the design scheme. Based on the virtual object positioning system and its virtual object positioning method for mixed reality technology provided in this application, the following steps are included:

[0117] Step 3.1: Capture an image of the real world through the camera of the MR device;

[0118] Step 3.2: Identify the pre-set spatial anchor point identifiers in the image through image recognition technology (such as OpenCV, etc.);

[0119] Step 3.3: According to the identified spatial anchor point identifiers, the coordinate calculation module uses the following formula to calculate the three-dimensional coordinates of the spatial anchor point identifiers in the real world in real time;

[0120] P world = K -1 ·P image ·D

[0121] Where P world is the three - dimensional coordinates of the point of the spatial anchor in the real world; K is the internal parameter matrix of the camera. By way of example, it is a 3×3 matrix that contains internal parameters such as the focal length and optical center of the camera; K -1 is the inverse matrix of the internal parameter matrix; P image is the two - dimensional coordinates of the point of the spatial anchor identification in the image, represented here in homogeneous coordinates (i.e., adding one dimension, usually 1) for matrix multiplication; D is the depth information of the point obtained through a depth camera or other sensors, usually a scalar representing the distance of the point from the camera; “·” represents matrix multiplication.

[0122] This formula is used to convert the point P image in the image back to the point P world in the three - dimensional world coordinates. This formula is derived based on the imaging principle of the camera. During the camera imaging process, points in the three - dimensional world are mapped to a two - dimensional image through the internal and external parameters of the camera (not involved here). This formula is the inverse process of the imaging process, that is, given a point on the image, depth information, and the internal parameters of the camera, the corresponding point in the three - dimensional world can be calculated.

[0123] The intrinsic matrix is an important part of camera parameters. It contains parameters related to the characteristics of the camera itself and is mainly used to describe the internal optical and geometric characteristics of the camera. The intrinsic matrix is usually a 3x3 matrix, but in practical applications, to handle homogeneous coordinates, it is sometimes extended to a 4x4 matrix (the last row is usually [0 0 0 1]). The intrinsic matrix mainly includes the following: (1) Focal length: fx, fy: respectively represent the focal lengths of the camera in the x-axis and y-axis directions, with the unit of pixels. The focal length is the distance from the optical center of the camera lens to the image plane. However, in the intrinsic matrix, the focal length is measured in pixels because the camera ultimately outputs an image composed of pixels. Example: For a camera, fx in its intrinsic matrix may be 800 pixels and fy may be 600 pixels, indicating that the focal lengths of the camera in the x-axis and y-axis directions are different, usually due to the design of the camera sensor and lens. (2) Principal point coordinates: cx, cy: The principal point is the intersection of the camera optical axis and the image plane. In the intrinsic matrix, it represents the coordinates of the image center (ideally) in the pixel coordinate system. Example: The principal point coordinate cx may be half of the image width (assuming the image width is 1920 pixels, then cx is 960), and cy may be half of the image height (assuming the image height is 1080 pixels, then cy is 540). However, in practical applications, due to errors in the camera manufacturing and installation processes, the principal point coordinates may deviate slightly. (3) Axis skew parameter: s (sometimes also called alpha): Ideally, the x-axis and y-axis of the image coordinate system should be orthogonal. However, in reality, due to the design of the camera sensor and lens, there may be a certain tilt angle between the x-axis and y-axis, which is represented by the s parameter in the intrinsic matrix. However, in many camera models, s is usually assumed to be 0, that is, the x-axis and y-axis are considered orthogonal.

[0124] Example of an ideal intrinsic matrix:

[0125] Without considering the axis skew parameter s, an ideal intrinsic matrix K can be expressed as:

[0126]

[0127] This matrix projects points in the three-dimensional camera coordinate system onto the two-dimensional image coordinate system and takes into account the focal length and principal point offset of the camera. The intrinsic matrix is a parameter that is fixed when the camera leaves the factory (unless the camera hardware changes, such as replacing the lens or sensor). Therefore, in related applications of computer vision and robotics, it is usually necessary to calibrate the camera first to obtain the accurate intrinsic matrix.

[0128] The content of depth information includes: (1) Distance from a point to the camera: This is the most basic and intuitive content of depth information, which is the straight-line distance from each point in the scene to the camera. In machine vision, depth images are usually used to represent this distance information, where each pixel value in the image corresponds to the distance from the corresponding point in the scene to the camera. (2) Three-dimensional coordinate information: In addition to simple distance information, depth information can be further extended to the three-dimensional coordinate information of each point in the scene. This means that in addition to knowing the distance from a point to the camera, we can also know the specific position of these points in three-dimensional space (including the values on the X, Y, and Z axes). (3) Surface geometry information: The depth image also describes the geometric information of the three-dimensional scene surface, such as the shape, size, and surface concavity of objects. These information are crucial for tasks such as three-dimensional reconstruction and scene understanding. Examples: Applications of depth images: In the field of autonomous driving, vehicles obtain depth images of the surrounding environment through mounted cameras and sensors, enabling them to perceive the distance and position information of objects such as roads, pedestrians, and vehicles, and then make accurate driving decisions. Three-dimensional reconstruction: In the three-dimensional reconstruction task, depth images of the scene are first collected from multiple perspectives, and then the three-dimensional coordinate information in these depth images is used to reconstruct the three-dimensional model of the scene. Such models can be used in multiple fields such as virtual reality, game development, and architectural design. Augmented reality: In augmented reality applications, depth information is used to accurately place virtual objects in the appropriate positions in the real scene. For example, in an AR game, players can see scenes where virtual characters or items seamlessly blend with the real environment, which benefits from the accurate perception of the distance of each point in the scene by depth information.

[0129] Step 3.4: The virtual object positioning module binds the calculated three-dimensional coordinates to the virtual building model, thereby achieving the precise positioning of the virtual building model in the real world.

[0130] The methods for the virtual object positioning module to bind the calculated three-dimensional coordinates to the virtual object (virtual building model) include:

[0131] (1) Coordinate calculation

[0132] First, the system needs to calculate the three-dimensional coordinates of the target object in the real world. This is usually achieved through some form of sensor or tracking system, such as using a depth camera, lidar (LiDAR), or an optical tracker, etc.

[0133] Suppose we already have the three-dimensional coordinates (Xr, Yr, Zr) of the target in the real world.

[0134] (2) Coordinate transformation

[0135] Next, it is necessary to convert the coordinates in the real world to the coordinate system of the virtual world. This usually involves a coordinate transformation matrix that defines the mapping relationship between the real-world coordinate system and the virtual-world coordinate system.

[0136] The coordinate transformation can be expressed as:

[0137] XvYvZv = [RT]XrYrZr

[0138] Preferably, R is a 3x3 rotation matrix, T is a 3x1 translation vector, and (Xv, Yv, Zv) are the transformed virtual-world coordinates.

[0139] (3) Binding implementation

[0140] Once the coordinates in the virtual world are obtained, these coordinates can be bound to virtual objects. This is usually achieved by updating the transformation properties (such as position and rotation) of the virtual objects.

[0141] For example, in a 3D game engine such as Unity, the coordinates can be bound by setting the position property of the virtual object's Transform component: virtualObject.transform.position = new Vector3(X_v, Y_v, Z_v).

[0142] Example 4: Using MR for remote collaboration

[0143] Remote collaboration: In a remote work scenario, team members can use spatial anchors to share the positions and states of virtual objects, thereby improving collaboration efficiency. Based on the virtual object positioning system and its virtual object positioning method for mixed reality technology provided in this application, it includes the following steps:

[0144] Step 4.1: Capture an image of the real world through the camera of the MR device;

[0145] Step 4.2: Identify the pre-set spatial anchor identifier in the image through image recognition technology (such as OpenCV, etc.);

[0146] Step 4.3: According to the identified spatial anchor identifier, the coordinate calculation module uses the following formula to calculate the three-dimensional coordinates of the spatial anchor identifier in the real world in real time;

[0147] P world = K -1 ·P image ·D

[0148] where, P worldare the three-dimensional coordinates of the points of the spatial anchor in the real world; K is the internal parameter matrix of the camera. By way of example, it is a 3×3 matrix that contains internal parameters such as the focal length and optical center of the camera; K -1 is the inverse matrix of the internal parameter matrix; P image are the two-dimensional coordinates of the points of the spatial anchor identifier in the image, represented here in homogeneous coordinates (i.e., adding one dimension, usually 1) for matrix multiplication; D is the depth information of the points obtained through a depth camera or other sensors, usually a scalar representing the distance of the points from the camera; “·” represents matrix multiplication.

[0149] This formula is used to transform the point P image in the image back to the point P world in the three-dimensional world coordinates. This formula is derived based on the imaging principle of the camera. During the camera imaging process, the points in the three-dimensional world are mapped onto the two-dimensional image through the internal and external parameters of the camera (not involved here). This formula is the inverse process of the imaging process, that is, given the points on the image, the depth information, and the internal parameters of the camera, the corresponding points in the three-dimensional world can be calculated.

[0150] The Intrinsic Matrix is an important part of camera parameters. It contains parameters related to the characteristics of the camera itself and is mainly used to describe the internal optical and geometric characteristics of the camera. The intrinsic matrix is usually a 3x3 matrix, but in practical applications, in order to handle homogeneous coordinates, it is sometimes extended to a 4x4 matrix (the last row is usually [0 0 0 1]). The intrinsic matrix mainly includes the following: (1) Focal Length: fx, fy: respectively represent the focal lengths of the camera in the x-axis and y-axis directions, with the unit of pixels. The focal length is the distance from the optical center of the camera lens to the image plane. However, in the intrinsic matrix, the focal length is measured in pixels because the final output of the camera is an image, which is composed of pixels. Example: For a camera, fx in its intrinsic matrix may be 800 pixels, and fy may be 600 pixels, indicating that the focal lengths of the camera in the x-axis and y-axis directions are different, which is usually due to the design of the camera sensor and lens. (2) Principal Point Coordinates: cx, cy: The principal point is the intersection of the camera optical axis and the image plane. In the intrinsic matrix, it represents the coordinates of the image center (ideally) in the pixel coordinate system. Example: The principal point coordinate cx may be half of the image width (assuming the image width is 1920 pixels, then cx is 960), and cy may be half of the image height (assuming the image height is 1080 pixels, then cy is 540). However, in practical applications, due to errors in the camera manufacturing and installation processes, the principal point coordinates may deviate slightly. (3) Axis Skew Parameter: s (sometimes also called alpha): Ideally, the x-axis and y-axis of the image coordinate system should be orthogonal. However, in reality, due to the design of the camera sensor and lens, there may be a certain tilt angle between the x-axis and y-axis, which is represented by the s parameter in the intrinsic matrix. However, in many camera models, s is usually assumed to be 0, that is, the x-axis and y-axis are considered orthogonal.

[0151] Example of an ideal intrinsic matrix in the ideal case:

[0152] Without considering the axis skew parameter s, an ideal intrinsic matrix K can be expressed as:

[0153]

[0154] This matrix projects points in the three-dimensional camera coordinate system onto the two-dimensional image coordinate system and takes into account the focal length and principal point offset of the camera. The intrinsic matrix is a parameter that is fixed when the camera leaves the factory (unless the camera hardware changes, such as replacing the lens or sensor). Therefore, when conducting related applications in computer vision and robotics, it is usually necessary to calibrate the camera first to obtain an accurate intrinsic matrix.

[0155] The content of depth information includes: (1) Distance from a point to the camera: This is the most basic and intuitive content of depth information, which is the straight-line distance from each point in the scene to the camera. In machine vision, depth images are usually used to represent this distance information, where each pixel value of the image corresponds to the distance from the corresponding point in the scene to the camera. (2) Three-dimensional coordinate information: In addition to simple distance information, depth information can be further extended to the three-dimensional coordinate information of each point in the scene. This means that in addition to knowing the distance from a point to the camera, we can also know the specific position of these points in three-dimensional space (including the values on the X, Y, and Z axes). (3) Surface geometry information: The depth image also describes the geometric information of the three-dimensional scene surface, such as the shape, size, and surface concavity and convexity of objects. These information are crucial for tasks such as three-dimensional reconstruction and scene understanding. Example: Application of depth images: In the field of autonomous driving, vehicles obtain depth images of the surrounding environment through mounted cameras and sensors, thereby being able to perceive the distance and position information of objects such as roads, pedestrians, and vehicles, and then make accurate driving decisions. Three-dimensional reconstruction: In the three-dimensional reconstruction task, first, depth images of the scene are collected from multiple perspectives, and then the three-dimensional coordinate information in these depth images is used to reconstruct the three-dimensional model of the scene. This model can be used in multiple fields such as virtual reality, game development, and architectural design. Augmented reality: In augmented reality applications, depth information is used to accurately place virtual objects in appropriate positions in the real scene. For example, in an AR game, players can see scenes where virtual characters or items are seamlessly integrated with the real environment, which benefits from the accurate perception of the distance of each point in the scene by depth information.

[0156] Step 4.4: The virtual object positioning module binds the calculated three-dimensional coordinates to the virtual object, thereby achieving precise positioning of the virtual object in the real world.

[0157] The methods by which the virtual object positioning module binds the calculated three-dimensional coordinates to the virtual object include:

[0158] (1) Coordinate calculation

[0159] First of all, the system needs to calculate the three-dimensional coordinates of the target object in the real world. This is usually achieved through some form of sensor or tracking system, such as using a depth camera, lidar (LiDAR), or an optical tracker, etc.

[0160] Suppose we already have the three-dimensional coordinates (Xr, Yr, Zr) of the target in the real world.

[0161] (2) Coordinate transformation

[0162] Next, it is necessary to convert the coordinates in the real world to the coordinate system of the virtual world. This usually involves a coordinate transformation matrix that defines the mapping relationship between the real-world coordinate system and the virtual-world coordinate system.

[0163] The coordinate transformation can be expressed as:

[0164] XvYvZv = [RT]XrYrZr

[0165] Preferably, R is a 3x3 rotation matrix, T is a 3x1 translation vector, and (Xv, Yv, Zv) are the transformed virtual-world coordinates.

[0166] (3) Binding implementation

[0167] Once the coordinates in the virtual world are available, these coordinates can be bound to virtual objects. This is usually achieved by updating the transformation properties (such as position, rotation) of the virtual objects.

[0168] For example, in a 3D game engine such as Unity, the coordinates can be bound by setting the position property of the virtual object's Transform component: virtualObject.transform.position = new Vector3(X_v, Y_v, Z_v).

[0169] Example 5: Using MR for gaming

[0170] In an MR game, players can enhance the interactivity and realism of the game by precisely placing virtual characters, props, or buildings in the real world by setting spatial anchors.

[0171] For example, in an MR game, a player needs to place a virtual treasure chest in the real world. Based on the virtual object positioning system and its virtual object positioning method for mixed reality technology provided in this application, the following steps are included:

[0172] Step 5.1: Capture an image of the real world through the camera of the MR device;

[0173] Step 5.2: Identify the pre-set spatial anchor identifier in the image through image recognition technology (such as OpenCV, etc.);

[0174] Step 5.3: According to the identified spatial anchor identifier, the coordinate calculation module uses the following formula to calculate the three-dimensional coordinates of the spatial anchor identifier in the real world in real time.

[0175] P world = K -1 ·P image ·D

[0176] Among them, P world is the three-dimensional coordinates of the point of the spatial anchor in the real world; K is the internal parameter matrix of the camera. By way of example, it is a 3×3 matrix that contains internal parameters such as the focal length and optical center of the camera; K -1 is the inverse matrix of the internal parameter matrix; P image is the two-dimensional coordinates of the point of the spatial anchor identification in the image, which is represented here in homogeneous coordinates (that is, adding one dimension, usually 1) for matrix multiplication; D is the depth information of the point obtained through a depth camera or other sensors, and is usually a scalar representing the distance of the point from the camera; "·" represents matrix multiplication.

[0177] This formula is used to convert the point P image in the image back to the point P world in the three-dimensional world coordinates. This formula is derived based on the imaging principle of the camera. During the camera imaging process, the points in the three-dimensional world are mapped onto the two-dimensional image through the internal and external parameters of the camera (not involved here). This formula is the inverse process of the imaging process, that is, given the points on the image, the depth information, and the internal parameters of the camera, the corresponding points in the three-dimensional world can be calculated.

[0178] The intrinsic matrix is an important part of camera parameters. It contains parameters related to the characteristics of the camera itself and is mainly used to describe the internal optical and geometric characteristics of the camera. The intrinsic matrix is usually a 3x3 matrix, but in practical applications, to handle homogeneous coordinates, it is sometimes extended to a 4x4 matrix (the last row is usually [0 0 0 1]). The intrinsic matrix mainly includes the following: (1) Focal length: fx, fy: respectively represent the focal lengths of the camera in the x-axis and y-axis directions, with the unit of pixels. The focal length is the distance from the optical center of the camera lens to the image plane. However, in the intrinsic matrix, the focal length is measured in pixels because the final output of the camera is an image, which is composed of pixels. Example: For a camera, fx in its intrinsic matrix may be 800 pixels, and fy may be 600 pixels, indicating that the focal lengths of the camera in the x-axis and y-axis directions are different, which is usually due to the design of the camera sensor and lens. (2) Principal point coordinates: cx, cy: The principal point is the intersection of the camera optical axis and the image plane. In the intrinsic matrix, it represents the coordinates of the image center (ideally) in the pixel coordinate system. Example: The principal point coordinate cx may be half of the image width (assuming the image width is 1920 pixels, then cx is 960), and cy may be half of the image height (assuming the image height is 1080 pixels, then cy is 540). However, in practical applications, due to errors in the camera manufacturing and installation processes, the principal point coordinates may deviate slightly. (3) Axis skew parameter: s (sometimes also called alpha): Ideally, the x-axis and y-axis of the image coordinate system should be orthogonal. However, in reality, due to the design of the camera sensor and lens, there may be a certain tilt angle between the x-axis and y-axis, which is represented by the s parameter in the intrinsic matrix. However, in many camera models, s is usually assumed to be 0, that is, the x-axis and y-axis are considered orthogonal.

[0179] Example of an ideal intrinsic matrix under ideal conditions:

[0180] Without considering the axis skew parameter s, an ideal intrinsic matrix K can be expressed as:

[0181]

[0182] This matrix projects points in the three-dimensional camera coordinate system onto the two-dimensional image coordinate system and takes into account the focal length and principal point offset of the camera. The intrinsic matrix is a parameter that is fixed when the camera leaves the factory (unless the camera hardware changes, such as replacing the lens or sensor). Therefore, in related applications of computer vision and robotics, it is usually necessary to calibrate the camera first to obtain an accurate intrinsic matrix.

[0183] The content of depth information includes: (1) Distance from a point to the camera: This is the most basic and intuitive content of depth information, which is the straight-line distance from each point in the scene to the camera. In machine vision, depth images are usually used to represent this distance information, where each pixel value of the image corresponds to the distance from the corresponding point in the scene to the camera. (2) Three-dimensional coordinate information: In addition to simple distance information, depth information can be further extended to the three-dimensional coordinate information of each point in the scene. This means that in addition to knowing the distance from a point to the camera, we can also know the specific position of these points in three-dimensional space (including the values on the X, Y, and Z axes). (3) Surface geometry information: Depth images also describe the geometric information of the three-dimensional scene surface, such as the shape, size, and surface concavity and convexity of objects. These information are crucial for tasks such as three-dimensional reconstruction and scene understanding. Example: Application of depth images: In the field of autonomous driving, vehicles obtain depth images of the surrounding environment through mounted cameras and sensors, so as to be able to perceive the distance and position information of objects such as roads, pedestrians, and vehicles, and then make accurate driving decisions. Three-dimensional reconstruction: In the three-dimensional reconstruction task, first, depth images of the scene are collected from multiple perspectives, and then the three-dimensional coordinate information in these depth images is used to reconstruct the three-dimensional model of the scene. Such models can be used in multiple fields such as virtual reality, game development, and architectural design. Augmented reality: In augmented reality applications, depth information is used to accurately place virtual objects in the appropriate positions in the real scene. For example, in an AR game, players can see scenes where virtual characters or items are seamlessly integrated with the real environment, which benefits from the accurate perception of the distance of each point in the scene by depth information.

[0184] Step 5.4: The virtual object positioning module binds the calculated three-dimensional coordinates to the virtual treasure chest, thereby achieving the precise positioning of the virtual treasure chest in the real world.

[0185] The methods by which the virtual object positioning module binds the calculated three-dimensional coordinates to the virtual object include:

[0186] (1) Coordinate calculation

[0187] First of all, the system needs to calculate the three-dimensional coordinates of the target object in the real world. This is usually achieved through some form of sensor or tracking system, such as using a depth camera, lidar (LiDAR), or an optical tracker, etc.

[0188] Suppose we already have the three-dimensional coordinates (Xr, Yr, Zr) of the target in the real world.

[0189] (2) Coordinate transformation

[0190] Next, it is necessary to convert the coordinates in the real world into the coordinate system of the virtual world. This usually involves a coordinate transformation matrix, which defines the mapping relationship between the real-world coordinate system and the virtual-world coordinate system.

[0191] The coordinate transformation can be expressed as:

[0192] XvYvZv = [RT]XrYrZr

[0193] Preferably, R is a 3x3 rotation matrix, T is a 3x1 translation vector, and (Xv, Yv, Zv) are the transformed virtual-world coordinates.

[0194] (3)Binding implementation

[0195] Once the coordinates in the virtual world are obtained, these coordinates can be bound to virtual objects. This is usually achieved by updating the transformation properties (such as position, rotation) of the virtual objects.

[0196] For example, in a 3D game engine such as Unity, the coordinates can be bound by setting the position property of the virtual object's Transform component: virtualObject.transform.position = new Vector3(X_v, Y_v, Z_v).

[0197] Comparative Example 1

[0198] Through tests in actual applications, the comparison of the performance indicators between the technical solution of this application and the traditional technical solution is shown in Table 1.

[0199] Table 1 Comparison of performance indicators between the technical solution of this application and the traditional technical solution

[0200] Performance Index Traditional Technical Solution Technical Solution of This Application Positioning Accuracy (m) Average: ±0.20, Maximum: ±0.50 Average: ±0.05, Maximum: ±0.10 Stability (%) After continuous operation for 72 hours, the probability of failure is 5% After continuous operation for 72 hours, the probability of failure is 1% Anchor Creation Time (s) Average: 5.0, Maximum: 10.0 Average: 1.0, Maximum: 2.0 Anchor Update Frequency (Hz) 10 50 Temperature Stability (°C) In the temperature range from -10°C to 50°C, the positioning error variation can reach ±0.15 m In the temperature range from -10°C to 50°C, the positioning error variation does not exceed ±0.02 m Anchor Persistence (days) The average time that the anchor can continuously exist and remain stable in the environment is 30 days The time that the anchor can continuously exist and remain stable in the environment exceeds 90 days Environmental Adaptability Sensitive to environmental factors such as light and occlusion, which may affect positioning accuracy and stability Has strong adaptability to environmental factors such as light and occlusion, and can maintain high accuracy and stability in complex environments

[0201] The tests in the above actual applications show that the present invention has the following beneficial effects:

[0202] (1)Improve positioning accuracy: The present invention significantly improves the positioning accuracy of spatial anchors by introducing depth information (D) and combining it with the internal parameter matrix (K) of the camera for precise three-dimensional coordinate calculation. This method reduces the influence of environmental factors (such as light changes, angle changes, etc.) on the positioning accuracy.

[0203] (2)Improve system stability: By accurately identifying the spatial anchor identifiers through image recognition technology, the system of the present invention can maintain a high degree of stability in complex environments and can accurately identify and locate spatial anchors even in dynamically changing scenarios.

[0204] (3)Wide range of application scenarios: The spatial anchor technology of the present invention is not only applicable to MR games, but also can be applied to multiple fields such as education, training, architectural design preview, etc., to achieve seamless integration of virtual content and the real world. This technology has broad application prospects and great commercial value, and is expected to promote the further development and popularization of mixed reality technology.

[0205] (4)User experience optimization: By improving the positioning accuracy and stability, the present invention significantly enhances the immersion and interactivity of users in the mixed reality experience, thereby optimizing the user experience.

[0206] It should be noted that the embodiments of the present invention have good implementability and do not impose any form of limitation on the present invention. Any person skilled in the art may use the disclosed technical content to modify or transform it into equivalent effective embodiments. However, as long as it does not depart from the technical solution of the present invention, any modification, equivalent change or modification made to the above embodiments based on the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A virtual object positioning system for mixed reality technology, characterized in that: include: A camera module, wherein the camera module is used to obtain images of the real world; An image recognition module, wherein the image recognition module is used to recognize a preset spatial anchor point mark in an image; A coordinate calculation module, the coordinate calculation module is used to calculate the three-dimensional coordinates of the spatial anchor point mark in the real world in real time according to the identified spatial anchor point mark; A virtual object positioning module is used to bind the calculated three-dimensional coordinates to the virtual object, thereby achieving real-time positioning of the virtual object.

2. The virtual object positioning system for mixed reality technology according to claim 1, characterized in that: The coordinate calculation module uses the following formula to calculate the three-dimensional coordinates of the spatial anchor point marker in the real world: P world = K -1 ·P image ·D Among them, P world is the three-dimensional coordinate of the spatial anchor point in the real world, K is the intrinsic parameter matrix of the camera, K -1 is the inverse matrix of the internal parameter matrix, P image is the two-dimensional coordinate of the spatial anchor point in the image, D is the depth information obtained by the depth camera or other sensors, and "·" represents matrix multiplication.

3. The virtual object positioning system for mixed reality technology according to claim 2, characterized in that: The internal parameter matrix is ​​a 3x3 matrix or a 4x4 matrix.

4. The virtual object positioning system for mixed reality technology according to claim 2, characterized in that: The depth information includes: the distance from the point to the camera, three-dimensional coordinate information and surface geometry information.

5. The virtual object positioning system for mixed reality technology according to claim 1, characterized in that: The method in which the virtual object positioning module binds the calculated three-dimensional coordinates to the virtual object includes: (1) Coordinate calculation: Use sensors to calculate the three-dimensional coordinates (Xr, Yr, Zr) of the target object in the real world; (2) Coordinate transformation: The coordinates in the real world are transformed into the coordinate system of the virtual world through the coordinate transformation matrix. The coordinate transformation matrix defines the mapping relationship between the coordinate system of the real world and the coordinate system of the virtual world. The formula of the coordinate transformation matrix is: XvYvZv=[RT]XrYrZr Where R is the rotation matrix, T is the translation vector, and (Xv, Yv, Zv) is the converted virtual world coordinate; (3) Binding implementation: Bind the virtual world coordinates to the virtual object by updating the transformation properties of the virtual object.

6. The virtual object positioning system for mixed reality technology according to claim 5, characterized in that: R is a 3x3 rotation matrix and T is a 3x1 translation vector.

7. A method for positioning a virtual object for mixed reality technology using the virtual object positioning system for mixed reality technology according to any one of claims 1 to 6, characterized in that: include: Step S1: Acquire an image of the real world; Step S2: Identify preset spatial anchor point markers in the image; Step S3: Calculate the three-dimensional coordinates of the spatial anchor point marker in the real world in real time according to the identified spatial anchor point marker; Step S4: Binding the calculated three-dimensional coordinates to the virtual object, thereby achieving real-time positioning of the virtual object.

8. The virtual object positioning method according to claim 7, characterized in that: In step S3, the three-dimensional coordinates of the spatial anchor point marker in the real world are calculated using the following formula: P world = K -1 ·P image ·D Among them, P world is the three-dimensional coordinate of the spatial anchor point in the real world, K is the intrinsic parameter matrix of the camera, K -1 is the inverse matrix of the internal parameter matrix, P image is the two-dimensional coordinate of the spatial anchor point in the image, D is the depth information obtained by the depth camera or other sensors, and "·" represents matrix multiplication.

9. The virtual object positioning method according to claim 7, characterized in that: In step S4, the method of binding the calculated three-dimensional coordinates to the virtual object includes: (1) Coordinate calculation: Use sensors to calculate the three-dimensional coordinates (Xr, Yr, Zr) of the target object in the real world; (2) Coordinate transformation: The coordinates in the real world are transformed into the coordinate system of the virtual world through the coordinate transformation matrix. The coordinate transformation matrix defines the mapping relationship between the coordinate system of the real world and the coordinate system of the virtual world. The formula of the coordinate transformation matrix is: XvYvZv=[RT]XrYrZr Where R is the rotation matrix, T is the translation vector, and (Xv, Yv, Zv) is the converted virtual world coordinate; (3) Binding implementation: Bind the virtual world coordinates to the virtual object by updating the transformation properties of the virtual object.

10. The virtual object positioning method according to any one of claims 7 to 9, characterized in that: The virtual object positioning method is applicable to the medical field, the education and training field, the architectural design field, the remote collaboration field and the game field.