Virtual reality positioning method of head-mounted device and related device
By combining a visual sensing module and a laser positioning module, and dynamically evaluating and fusing positioning data, the problem of positioning instability in virtual reality devices in complex environments is solved, achieving a more stable and consistent positioning effect and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PIMAX TECH (SHANGHAI) CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing positioning technologies for virtual reality headsets struggle to balance stability and scene adaptability in complex environments. Laser positioning is susceptible to occlusion, while visual positioning is limited by lighting conditions, resulting in instability and poor consistency.
By combining a visual sensing module and a laser positioning module, the availability of the two types of positioning data is assessed through environmental perception information, and the positioning mode is dynamically selected or fused to obtain stable pose information.
It improves the positioning stability and consistency of virtual reality devices in complex scenarios, and optimizes the user experience.
Smart Images

Figure CN121829480A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of virtual reality, in particular to a virtual reality positioning method of a head-mounted device and related equipment. BACKGROUND
[0002] In the application of virtual reality head-mounted devices, the reliability of positioning technology directly affects the user experience. Laser positioning systems can provide high-precision positioning capabilities and have strong anti-environmental light interference characteristics, but their implementation relies on the deployment of external base stations, and they are easily blocked by physical obstacles during device movement, resulting in unstable or even interrupted positioning signals. Camera-based visual simultaneous localization and mapping technology does not require additional hardware support, can flexibly adapt to various scenarios and support augmented reality and other extended functions, but its performance is severely limited by environmental lighting conditions and scene texture characteristics, and the positioning accuracy will decrease significantly in low-light or texture-lacking areas. In complex real-world scenarios, environmental lighting changes and physical obstacle distribution are often random and variable, and the inherent characteristics of different positioning technologies make it difficult for them to adapt to diverse environments alone, thereby affecting the stability and consistency of virtual reality device positioning performance. SUMMARY
[0003] The present application provides a virtual reality positioning method of a head-mounted device and related equipment to improve the stability and environmental adaptability of head-mounted device positioning, thereby optimizing the user experience.
[0004] In a first aspect, the present application provides a virtual reality positioning method for a head-mounted device, the head-mounted device comprising a visual sensing module and a laser positioning module, the method comprising: obtaining environmental perception information of a target scene; obtaining first positioning data and second positioning data representing the pose of the head-mounted device in the target scene, wherein the first positioning data is generated by the visual sensing module based on the collected images, and the second positioning data is provided by the laser positioning module; based on the environmental perception information, evaluating the available state of the first positioning data and the second positioning data; determining a target positioning mode according to the result of the available state evaluation; based on the target positioning mode, processing the first positioning data and / or the second positioning data to determine the pose information for constructing a target virtual reality scene.
[0005] In some embodiments, the result of the available state evaluation includes a first confidence score of the first positioning data and a second confidence score of the second positioning data; the available state evaluation of the first positioning data and the second positioning data comprises: calculate the first confidence score according to the light condition information in the environment perception information and the visual feature information extracted based on the image collected by the visual sensing module; calculate the second confidence score according to the spatial structure information in the environment perception information and the laser signal parameter provided by the laser positioning module.
[0006] In some embodiments, the determining the target positioning mode according to the result of the available state evaluation comprises: if the first confidence score is greater than a first threshold and the second confidence score is less than a second threshold, determining the target positioning mode as a first mode, and in the first mode, the pose information is determined according to the first positioning data; if the second confidence score is greater than the second threshold, determining the target positioning mode as a second mode, and in the second mode, the pose information is determined according to the second positioning data.
[0007] In some embodiments, the determining the target positioning mode according to the result of the available state evaluation further comprises: if the first confidence score is greater than a first threshold and the second confidence score is greater than a second threshold, determining the target positioning mode as a third mode; the processing the first positioning data and / or the second positioning data based on the target positioning mode to determine the pose information for constructing the target virtual reality scene comprises: in the third mode, converting the first positioning data and the second positioning data to the same coordinate system through a preset conversion matrix to obtain converted first positioning data and converted second positioning data; determining a first fusion weight of the converted first positioning data and a second fusion weight of the converted second positioning data according to the first confidence score and the second confidence score; performing weighted fusion on the converted first positioning data and the converted second positioning data according to the first fusion weight and the second fusion weight to determine the pose information.
[0008] In some embodiments, the method further comprises: obtaining perspective image information based on the image collected by the visual sensing module; constructing the target virtual reality scene based on the perspective image information and the pose information determined in the second mode to realize the perspective function in the target virtual reality scene.
[0009] In some embodiments, the method further comprises: determine to support interaction using the second positioning data when a first handle access signal associated with the laser positioning module is detected; and / or determine to support interaction using the first positioning data when a second handle access signal associated with the visual sensing module is detected.
[0010] In some embodiments, the visual sensing module comprises a processing unit for implementing visual simultaneous localization and mapping function, and the laser positioning module is a laser positioning unit detachably connected to the head-mounted device.
[0011] In a second aspect, the present application provides a virtual reality positioning apparatus for a head-mounted device, the head-mounted device comprising a visual sensing module and a laser positioning module, the apparatus comprising: an environment perception module configured to acquire environment perception information of a target scene; a data acquisition module configured to acquire first positioning data and second positioning data representing a pose of the head-mounted device in the target scene, wherein the first positioning data is generated by the visual sensing module based on an image acquired, and the second positioning data is provided by the laser positioning module; a state evaluation module configured to evaluate an available state of the first positioning data and the second positioning data based on the environment perception information; a mode determination module configured to determine a target positioning mode according to a result of the available state evaluation; a pose determination module configured to process the first positioning data and / or the second positioning data based on the target positioning mode, and determine pose information for constructing a target virtual reality scene.
[0012] In a third aspect, the present application provides a head-mounted device, the head-mounted device comprising a visual sensing module and a laser positioning module, the head-mounted device further comprising: one or more processors; and a memory associated with the one or more processors, the memory being configured to store program instructions, which when executed by the one or more processors, perform the steps of the method of any one of the first aspect.
[0013] In a fourth aspect, the present application provides a computer program product comprising a computer program, which when executed by a processor, implements the steps of the method of any one of the first aspect.
[0014] According to the specific embodiments provided by the present application, the following technical effects are disclosed: The virtual reality positioning method and related equipment for head-mounted devices in this application acquire environmental perception information of the target scene and two types of positioning data generated by the visual sensing module and the laser positioning module, respectively. Based on the environmental perception information, the usability of the two types of positioning data is evaluated to determine the target positioning mode. Then, the corresponding positioning data is processed according to the target positioning mode to obtain pose information as the virtual reality positioning result. It is understood that this application can combine the advantages of visual positioning and laser positioning in the virtual reality positioning process, improve the consistency and reliability of virtual reality positioning results, solve the problem of unstable positioning of head-mounted devices in complex scenes, provide effective support for the stable construction of the target virtual reality scene, and thus optimize the user experience.
[0015] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the above drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the modules of the head-mounted device provided in an embodiment of this application.
[0018] Figure 2 A flowchart illustrating the steps of a virtual reality positioning method for a head-mounted device provided in an embodiment of this application.
[0019] Figure 3 The flowchart illustrates the steps of determining the target positioning mode in the virtual reality positioning method for a head-mounted device provided in this application embodiment.
[0020] Figure 4 The flowchart illustrates the steps of determining pose information for constructing a target virtual reality scene in the virtual reality positioning method for a head-mounted device provided in this application embodiment.
[0021] Figure 5 A flowchart illustrating the steps for implementing the perspective function in the virtual reality positioning method for a head-mounted device provided in this application embodiment.
[0022] Figure 6 This is a schematic diagram of a virtual reality positioning device module for a head-mounted device provided in an embodiment of this application.
[0023] Figure 7 A schematic block diagram of a head-mounted device provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0025] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0026] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0027] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0028] In virtual reality headsets, the reliability of positioning technology directly impacts user experience. Laser positioning systems offer high-precision positioning and strong resistance to ambient light interference; however, their implementation relies on external base station deployments and is susceptible to occlusion by physical obstacles during device movement, leading to unstable or even interrupted positioning signals. Camera-based real-time visual positioning and mapping technology requires no additional hardware support, can flexibly adapt to various scenarios, and supports extended functions such as augmented reality; however, its performance is severely limited by ambient lighting conditions and scene texture features, with positioning accuracy significantly decreasing in low-light or texture-deficient areas. In complex real-world scenarios, factors such as ambient lighting changes and the distribution of physical obstacles are often random and variable. The inherent characteristics of different positioning technologies make it difficult for them to adapt to diverse usage environments independently, thus affecting the stability and consistency of virtual reality device positioning performance.
[0029] Furthermore, traditional virtual reality headsets often employ a single positioning method, making it difficult to balance positioning stability with scene adaptability. For example, while laser positioning technology offers high accuracy and resistance to light interference, it requires additional base stations and is susceptible to occlusion. On the other hand, camera-based SLAM positioning, although requiring no external equipment and supporting various extended functions, is significantly affected by lighting conditions, performing poorly in low-light or texture-deficient scenes.
[0030] Based on this, this application provides a virtual reality positioning method and related equipment for head-mounted devices, which can improve the stability and environmental adaptability of head-mounted device positioning, thereby optimizing the user experience.
[0031] The present application will be further described below with reference to the accompanying drawings.
[0032] refer to Figure 1 , Figure 1 This is a schematic diagram of the modules of a head-mounted device provided in an embodiment of this application. In some embodiments, the head-mounted device 100 used in this application can be a PCVR device, which can integrate a visual sensing module 110 and a laser positioning module 120.
[0033] The head-mounted device can be worn on a user's head to provide virtual reality or augmented reality experiences. It can integrate components such as a display screen, sensors, and processors, enabling it to sense the user and their surrounding environment and generate corresponding virtual scenes. The visual sensing module is a component within the head-mounted device used to acquire image or video information. This module can include one or more cameras to obtain visual features of the target scene, providing a data foundation for vision-based positioning. The laser positioning module is a component within the head-mounted device used to emit and receive laser signals for positioning. This module can calculate the spatial position and attitude of the head-mounted device by measuring the time-of-flight or angular information of the laser signal, combined with external base station or internal inertial measurement unit information.
[0034] It is understood that the visual sensing module corresponds to SLAM positioning technology, and the laser positioning module corresponds to Lighthouse positioning technology. This application can simultaneously acquire positioning data through these two modules, and can determine the corresponding positioning mode based on the actual environmental conditions. Then, based on one or more of the positioning data, the device's pose information can be obtained, thereby realizing virtual reality functions.
[0035] Understandably, PCVR is a virtual reality headset that needs to be connected to a computer, and can rely on the high performance of the computer to provide a high-quality, low-latency VR experience; SLAM is a real-time localization and mapping technology that uses visual sensors to collect environmental images and can simultaneously complete its own localization and environmental map construction without relying on external auxiliary equipment; Lighthouse is a laser-based positioning technology that uses external base stations to emit laser signals, which are received and calculated by sensors on the device to achieve high-precision spatial tracking.
[0036] refer to Figure 2 , Figure 2 This is a flowchart illustrating the steps of a virtual reality positioning method for a head-mounted device provided in an embodiment of this application. In some embodiments, this application proposes a virtual reality positioning method for a head-mounted device, the method being applied to, for example... Figure 1 The head-mounted device 100 shown includes a visual sensing module 110 and a laser positioning module 120. The method may include at least the following steps: Step 210: Obtain environmental perception information of the target scene; Step 220: Obtain first positioning data and second positioning data representing the pose of the head-mounted device in the target scene. The first positioning data is generated by the visual sensing module based on the acquired images, and the second positioning data is provided by the laser positioning module. Step 230: Based on environmental perception information, assess the usability of the first and second positioning data; Step 240: Determine the target positioning mode based on the results of the availability assessment; Step 250: Based on the target localization mode, process the first localization data and / or the second localization data to determine the pose information used to construct the target virtual reality scene.
[0037] The environmental perception information can be data acquired through internal or external sensors of the head-mounted device, used to describe the characteristics of the target scene. This information may include lighting conditions, spatial structure, texture richness, obstacle distribution, etc., and is used to evaluate the applicability of different positioning methods. The first positioning data can be position and attitude data generated by the visual sensing module based on the acquired images. This data can be calculated using visual real-time positioning and mapping algorithms or other visual odometry methods, reflecting the relative or absolute pose of the head-mounted device in the scene. The second positioning data can be position and attitude data provided by the laser positioning module. This data can be acquired through technologies such as lidar scanning, laser triangulation, or laser inertial navigation, and has high accuracy and anti-interference capabilities. The usability assessment is the process of judging the reliability or applicability of the first and second positioning data based on the environmental perception information. The assessment aims to determine which positioning data is more valuable under the current scene conditions, or the confidence level of each of the two data.
[0038] In some embodiments, the target localization mode can be a specific strategy for subsequent pose information calculation, determined based on the results of an assessment of available states. The mode can indicate the use of visual localization alone, laser localization alone, or a fusion of both localization data. Pose information can be data describing the position and orientation of the head-mounted device in three-dimensional space. Pose information can include the device's three-dimensional coordinates (X, Y, Z) in a coordinate system and its rotation angles (pitch, yaw, roll) relative to the coordinate axes, and is the foundation for constructing virtual reality scenes.
[0039] It is understood that, based on the virtual reality positioning method for head-mounted devices provided in this embodiment, this application can execute the following steps in sequence: First, environmental perception information of the target scene is acquired. This information can be obtained through various sensors within the head-mounted device. For example, a light sensor can be used to obtain the brightness information of the current scene, an inertial measurement unit (IMU) can be used to obtain the device's motion state information, or a pre-set scene database can be used to obtain the type information of the current scene. Alternatively, environmental perception information can also be obtained through an external sensor network or through manual user input, such as allowing the user to manually select the current scene as "bright indoors" or "dim outdoor."
[0040] Next, first and second positioning data, representing the pose of the head-mounted device in the target scene, are acquired. The first positioning data is generated by the visual sensing module based on acquired images. For example, the visual sensing module can continuously acquire scene images and calculate the relative motion of the device in the image sequence using image processing algorithms (such as feature point matching and optical flow tracking), thereby deducing the first positioning data. Alternatively, the first positioning data can also be obtained by directly acquiring the absolute pose of the device by recognizing preset visual markers (such as QR codes or AR markers) in the scene. The second positioning data is provided by the laser positioning module. For example, the laser positioning module can periodically emit laser beams and receive reflected signals, construct an environmental point cloud by measuring the time of flight or phase difference, and perform matching or registration based on the point cloud data to calculate the second positioning data. Alternatively, the laser positioning module can also cooperate with an external laser base station to triangulate the device's pose by measuring the time difference of the laser signal arriving at the base station.
[0041] Subsequently, based on environmental perception information, the usability of the first and second positioning data is assessed. The assessment aims to determine the reliability or applicability of the two types of positioning data in the current environment. For example, based on the light intensity in the environmental perception information, it can be determined whether visual positioning data might be affected; or based on the obstacle distribution in the environmental perception information, it can be determined whether laser positioning data might be obstructed. As one implementation, the assessment can be a simple binary judgment, determining whether each type of data is "available" or "unavailable." Alternatively, the assessment can be performed using an expert system or rule engine, outputting the applicability level of each type of positioning data based on a preset set of rules and environmental parameters.
[0042] Then, based on the results of the availability assessment, a target localization mode is determined. The target localization mode indicates the strategy for subsequent pose information determination. For example, if the assessment results show that visual localization data performs well in the current environment, while laser localization data may have problems, the target localization mode can be determined to prioritize the use of visual localization data. If the assessment results show that laser localization data performs excellently, the target localization mode can be determined to prioritize the use of laser localization data. As one implementation approach, mode determination can be based on simple priority rules; for example, when visual localization data is assessed as available, the visual localization mode is selected; otherwise, if laser localization data is assessed as available, the laser localization mode is selected.
[0043] Furthermore, based on the target localization mode, the first localization data and / or the second localization data are processed to determine the pose information used to construct the target virtual reality scene. For example, if the target localization mode indicates that the first localization data should be used first, then the first localization data is directly used as the final pose information. If the target localization mode indicates that the second localization data should be used first, then the second localization data is directly used as the final pose information. As one implementation approach, when both types of localization data are evaluated as available, one type of data can be simply selected as the pose information, for example, always selecting the data with higher accuracy, or selecting the data with lower computational overhead.
[0044] In summary, this embodiment introduces a visual sensing module and a laser positioning module, and combines environmental perception information to assess the usability of both types of positioning data. This allows for adaptive selection or processing of positioning data based on different scene conditions. Therefore, it overcomes the problem of traditional single virtual reality positioning methods struggling to balance positioning stability and scene adaptability in complex and changing scenarios. The method effectively addresses challenges such as changes in lighting, missing textures, or occlusion, ensuring that the head-mounted device obtains stable and reliable pose information in various virtual reality applications, thereby improving the user experience.
[0045] In some embodiments, this application further proposes that the results of the availability assessment include a first confidence score of the first positioning data and a second confidence score of the second positioning data. The availability assessment of the first and second positioning data may include at least the following steps: calculating the first confidence score based on illumination condition information in the environmental perception information and visual feature information extracted from images acquired by the visual sensing module; and calculating the second confidence score based on spatial structure information in the environmental perception information and laser signal parameters provided by the laser positioning module.
[0046] The first confidence score and the second confidence score are indicators used to quantify the reliability or availability of the first and second location data. Higher scores indicate better quality and greater reliability of the corresponding location data, providing a foundation for subsequently determining the target location pattern.
[0047] Specifically, when calculating the first confidence score, the system comprehensively considers both the lighting condition information from the environmental perception information and the visual feature information extracted from the images acquired by the visual sensing module. Lighting condition information can be obtained by analyzing the brightness, contrast, and histogram distribution of the images acquired by the visual sensing module. For example, in excessively bright or dark environments, image quality degrades, visual features are difficult to extract, or are easily affected by noise; in this case, the lighting condition information will indicate a lower confidence score. Visual feature information refers to the feature points or feature lines extracted from the image for visual localization. It can assess the number of feature points, their uniformity of distribution, the uniqueness of descriptors, and the stability of feature point tracking. For example, when the number of feature points in the image is sparse, their distribution is uneven, or the feature point tracking loss rate is high, the confidence score for visual localization will decrease. The calculation of the first confidence score can be performed using a pre-trained machine learning model or a rule-based expert system, taking the lighting condition information and visual feature information as inputs for comprehensive judgment.
[0048] Simultaneously, when calculating the second confidence score, the system considers spatial structure information from the environmental perception information and laser signal parameters provided by the laser positioning module. Spatial structure information can be analyzed using point cloud data obtained from LiDAR scanning. For example, it assesses whether there are sufficient geometric features such as planes, edges, or corners in the scene, or whether there are large open areas or overly complex structures. For instance, in open environments lacking reflectors, the accuracy of laser positioning decreases, and spatial structure information indicates a lower confidence score. Laser signal parameters can be the signal characteristics generated by the laser positioning module during operation, such as laser signal intensity, signal-to-noise ratio, reflectivity, density of the scanned point cloud, and the number of effective points within the scanning range. For instance, when the laser signal intensity is too low, the signal-to-noise ratio is poor, or the number of effective point clouds is insufficient, the accuracy and stability of laser positioning are affected, and the laser signal parameters will indicate a lower confidence score. The calculation of the second confidence score can also employ machine learning models or rule-based systems, using spatial structure information and laser signal parameters as inputs for comprehensive judgment.
[0049] refer to Figure 3 , Figure 3 A flowchart illustrating the steps for determining a target positioning mode in the virtual reality positioning method for a head-mounted device provided in this application embodiment; as follows: Figure 3 As shown, in some embodiments, this application further proposes a specific method for determining the target positioning mode based on the results of the availability status assessment, which may include at least the following steps: Step 310: If the first confidence score is greater than the first threshold and the second confidence score is less than the second threshold, then the target positioning mode is determined to be the first mode, and the pose information in the first mode is determined based on the first positioning data. Step 320: If the second confidence score is greater than the second threshold, the target positioning mode is determined to be the second mode, and the pose information in the second mode is determined based on the second positioning data.
[0050] The first threshold is a preset value used to determine whether the first confidence score meets the minimum reliability standard sufficient for pose determination based solely on the first positioning data. When the first confidence score exceeds this threshold, it indicates that the visual positioning data has high availability in the current environment. The first threshold can be calibrated and adjusted according to the actual application scenario, the performance of the visual sensing module, and the requirements for positioning accuracy.
[0051] The second threshold is a preset value used to determine whether the second confidence score meets the minimum reliability standard sufficient to determine pose independently based on the second positioning data. When the second confidence score exceeds this threshold, it indicates that the laser positioning data has high availability in the current environment. Similar to the first threshold, the second threshold can also be calibrated and adjusted according to the actual application scenario, the performance of the laser positioning module, and the requirements for positioning accuracy.
[0052] Understandably, the first mode is the positioning strategy selected when the system determines that the visual positioning data is sufficiently reliable while the laser positioning data is unreliable. In this mode, the pose information of the head-mounted device will be determined entirely based on the first positioning data (i.e., the visual positioning data). This is suitable for scenarios such as well-lit, textured indoor environments, but where the laser signal may be interfered with or there may be insufficient reflective surfaces.
[0053] The second mode is the positioning strategy selected when the system determines that the laser positioning data is sufficiently reliable. In this mode, the pose information of the head-mounted device will be determined entirely based on the second positioning data (i.e., the laser positioning data). This is suitable for environments, for example, with poor lighting conditions, sparse visual features, but good laser reflection conditions.
[0054] In summary, by comparing the first confidence score with the first threshold, and the second confidence score with the second threshold, the system intelligently selects the most suitable positioning mode for the current environment. This dynamic selection mechanism ensures that the system can flexibly switch positioning strategies according to real-time environmental changes, thereby maximizing positioning accuracy and stability.
[0055] refer to Figure 4 , Figure 4 A flowchart illustrating the steps for determining pose information for constructing a target virtual reality scene in the virtual reality positioning method for a head-mounted device provided in this application embodiment; as shown... Figure 4As shown, in some embodiments, this application further proposes determining the target positioning mode as a third mode when the first confidence score is greater than a first threshold and the second confidence score is greater than a second threshold; based on the target positioning mode, processing the first positioning data and / or the second positioning data to determine the pose information used to construct the target virtual reality scene may include at least the following steps: Step 410: In the third mode, the first positioning data and the second positioning data are transformed to the same coordinate system using a preset transformation matrix to obtain the transformed first positioning data and the transformed second positioning data. Step 420: Determine the first fusion weight of the converted first location data and the second fusion weight of the converted second location data based on the first confidence score and the second confidence score. Step 430: Based on the first fusion weight and the second fusion weight, perform weighted fusion on the transformed first positioning data and the transformed second positioning data to determine the pose information.
[0056] Understandably, when the first confidence score of the first positioning data generated by the visual sensing module is greater than a preset first threshold, and the second confidence score of the second positioning data provided by the laser positioning module is also greater than a preset second threshold, it indicates that both types of positioning data have high reliability and usability. At this point, the system determines the target positioning mode as the third mode. The third mode aims to fully utilize these two high-quality positioning data to obtain more accurate and stable pose information through fusion, rather than relying solely on a single data source.
[0057] Since the visual sensing module and the laser positioning module can have their own independent coordinate systems, their output first and second positioning data may differ spatially. To effectively fuse these two types of data, they need to be unified into a single reference coordinate system. A pre-defined transformation matrix can be obtained through offline calibration or online calibration. For example, by placing specific markers at known locations, simultaneously acquiring visual images and laser data, and then calculating the rotation and translation parameters required to transform the data from the visual coordinate system to the laser coordinate system (or vice versa, or to a unified world coordinate system). The transformation matrix ensures the spatial alignment of data from different sensors, laying the foundation for subsequent fusion operations.
[0058] Fusion weights are used to quantify the contribution of different data sources in the fusion process. When both types of location data are available, their respective confidence scores can be used to determine the fusion weights. For example, normalized confidence scores can be used as weights, or a function proportional to the confidence scores can be used to calculate the weights. Specifically, the higher the confidence score of a data source, the greater its corresponding fusion weight, indicating that the data source is more reliable in the current environment and should have a larger proportion in the fusion result. This dynamic weight allocation mechanism based on confidence scores enables the system to intelligently adjust the influence of different data sources, thereby optimizing the fusion effect.
[0059] Weighted fusion is a method that combines information from multiple data sources by assigning a weight to each data source and then combining the weighted data. In the third mode, the first and second positioning data, after coordinate system transformation, are processed by weighted averaging or more complex fusion algorithms according to their respective determined first and second fusion weights. For example, pose information can be represented as position and orientation, and position and orientation can be weighted and fused separately. This fusion method effectively combines the texture richness advantage of visual positioning with the distance accuracy advantage of laser positioning. With both types of data being of high quality, the complementary enhancement reduces the noise or local errors that may exist in a single sensor, thereby obtaining more accurate, stable, and robust head-mounted device pose information than a single data source.
[0060] Through the above technical solution, when both the visual sensing module and the laser positioning module can provide high-confidence positioning data, the system can unify the two types of positioning data into the same coordinate system through a preset transformation matrix, eliminating spatial differences between the sensors. Subsequently, the fusion weights are dynamically determined based on their respective confidence scores, enabling the system to rationally allocate the contributions of the two data sources in the final pose determination according to the current environment and sensor performance. Finally, by weighted fusion of the transformed positioning data, the rich texture information of visual positioning and the high-precision distance information of laser positioning are effectively combined, thereby improving the accuracy, stability, and robustness of the head-mounted device's pose information.
[0061] refer to Figure 5 , Figure 5 A flowchart illustrating the steps for implementing the see-through function in the virtual reality positioning method for a head-mounted device provided in this application embodiment; as follows: Figure 5 As shown, in some embodiments, this application further proposes that the method may include at least the following steps: Step 510: Obtain perspective image information based on the image acquired by the visual sensing module; Step 520: Based on the perspective image information and the pose information determined in the second mode, construct the target virtual reality scene to realize the perspective function in the target virtual reality scene.
[0062] Specifically, the visual sensing module may include one or more cameras to capture real-time images of the user's surrounding environment. These cameras can be configured as wide-angle or fisheye lenses to provide a wider field of view. After acquiring the raw image data, the system performs a series of preprocessing steps, such as distortion correction, color balancing, and exposure adjustment, to generate high-quality perspective image information suitable for direct display to the user. This step is fundamental to enabling the user to perceive the real world, ensuring that subsequent virtual reality scene construction accurately reflects the physical environment.
[0063] In the second mode, the pose information of the head-mounted device is primarily provided by a laser positioning module, offering high accuracy and stability. When constructing the target virtual reality scene, the system utilizes this precise pose information to accurately map previously acquired perspective image information into the virtual space. Specifically, the perspective image information is rendered onto one or more virtual planes in the virtual scene. The position and orientation of these virtual planes are adjusted in real time based on the current pose information of the head-mounted device, ensuring that the real-world image seen by the user in the virtual scene remains consistent with their actual physical position and viewpoint. This mapping process is crucial for achieving seamless integration between the real and virtual worlds.
[0064] The see-through function allows users of head-mounted devices to clearly see their surrounding physical environment while wearing the device. By combining see-through image information with pose information provided by the laser positioning module to construct a virtual reality scene, users can not only perceive virtual content, but also simultaneously see and understand their real environment.
[0065] In some embodiments, this application further proposes that the method further includes: determining that interaction using second positioning data is supported when a first handle access signal associated with the laser positioning module is detected; and / or determining that interaction using the first positioning data is supported when a second handle access signal associated with the vision sensing module is detected.
[0066] When the system detects a first handle access signal associated with the laser positioning module, it signifies that an input device designed or optimized for the laser positioning system has been connected and is ready. This association can be pre-configured within the system; for example, a specific handle model may be identified as working in conjunction with the laser positioning module, or the handle itself may integrate sensors or markers required for laser positioning. The detection methods for the access signal can include, but are not limited to: establishing communication via a wired connection (such as USB), completing pairing via a wireless connection (such as Bluetooth or Wi-Fi), or the handle entering the effective working range of the head-mounted device and being recognized by the system. Once this access signal is detected, the system explicitly determines that for interactions generated by the first handle, the second positioning data provided by the laser positioning module should be used preferentially or primarily to determine its position and orientation in the virtual environment.
[0067] Furthermore, when the system detects a second handle access signal associated with the visual sensing module, it indicates that an input device designed or optimized for the visual positioning system has been connected. This association can also be pre-configured; for example, the handle may have specific visual markers so that the visual sensing module can accurately track it. The access signal is detected in a similar manner to the first handle. Once this access signal is detected, the system explicitly determines that for interactions generated by the second handle, the position and orientation in the virtual environment should be determined primarily or preferentially using the first positioning data provided by the visual sensing module.
[0068] In some embodiments, this application further proposes that the visual sensing module includes a processing unit for realizing real-time visual positioning and map building functions, and the laser positioning module is a laser positioning unit detachably connected to the head-mounted device.
[0069] The visual sensing module, which includes a processing unit for real-time visual localization and mapping (LAMR) functionality, is the core component of the module. Its main function is to execute LAMR algorithms. LAMR technology allows head-mounted devices to simultaneously estimate their own position and orientation, and construct a 3D map of the environment, in unknown environments, solely using image data acquired by visual sensors.
[0070] Furthermore, the laser positioning module is a laser positioning unit that is detachably connected to the head-mounted device. Detachable connection means that the laser positioning unit is not a fixed component of the head-mounted device, but can be flexibly installed or removed according to user needs or application scenarios. For example, the laser positioning module can be set in the positioning components of the head-mounted device, such as the face mask, and fixed via magnetic interfaces, USB interfaces, snap-fit connections, or screws. When connected, it can communicate with the main system of the head-mounted device, providing high-precision secondary positioning data. The laser positioning unit itself can contain a laser emitter and / or receiver for interacting with an external laser base station, or it can emit a laser and receive reflected signals for positioning. Its detachability allows users to install it in scenarios requiring high-precision, wide-area positioning, and remove it when not needed, thereby reducing device weight, lowering power consumption, or simplifying the device's form factor.
[0071] refer to Figure 6 , Figure 6 This is a schematic diagram of a virtual reality positioning device module for a head-mounted device provided in an embodiment of this application. Figure 6 As shown, this application provides a virtual reality positioning device 600 for a head-mounted device. The head-mounted device includes a visual sensing module and a laser positioning module. The device includes: The environmental perception module 610 is used to acquire environmental perception information of the target scene; The data acquisition module 620 is used to acquire first positioning data and second positioning data that characterize the pose of the head-mounted device in the target scene. The first positioning data is generated by the visual sensing module based on the acquired image, and the second positioning data is provided by the laser positioning module. The status assessment module 630 is used to assess the usability status of the first positioning data and the second positioning data based on environmental perception information. The pattern determination module 640 is used to determine the target positioning pattern based on the results of the availability status assessment. The pose determination module 650 is used to process first positioning data and / or second positioning data based on the target positioning mode to determine pose information for constructing the target virtual reality scene.
[0072] refer to Figure 7 , Figure 7 A schematic block diagram of a head-mounted device provided in an embodiment of this application; in some embodiments, this application provides a computer program product including a computer program that, when executed by a processor, implements the steps of any of the methods in the above embodiments.
[0073] in, Figure 7 The architecture of the head-mounted device provided in the embodiments of this application is illustrated by way of example. Figure 7As shown, the head-mounted device 100 may include a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, and a memory 720. The processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, and memory 720 can communicate with each other via a communication bus 730.
[0074] The processor 710 can be implemented using a general-purpose CPU, microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs and implement the technical solution provided in this application.
[0075] The memory 720 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 720 can store the operating system 721 for controlling the terminal's operation, and the basic input / output system (BIOS) 722 for controlling the terminal's low-level operations. Additionally, it may include storage for a web browser 723, a data storage management system 724, and a virtual reality positioning device 500 or 600 for a head-mounted device, etc. The aforementioned control device can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when implementing the technical solution provided in this application through software or firmware, the relevant program code is stored in the memory 720 and executed by the processor 710.
[0076] Input / output interface 713 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0077] Network interface 714 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0078] Bus 730 includes a pathway for transmitting information between multiple components of the device, such as processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, and memory 720.
[0079] It should be noted that although the above-described device only shows the processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, memory 720, bus 730, etc., in specific implementations, the head-mounted device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the solution of this application, and does not necessarily include all the components shown in the figures.
[0080] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer program product. This computer program product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of multiple embodiments or some parts of the embodiments of this application.
[0081] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A virtual reality positioning method for a head-mounted device, characterized in that, The head-mounted device includes a visual sensing module and a laser positioning module, and the method includes: Acquire environmental perception information of the target scene; Acquire first positioning data and second positioning data that characterize the pose of the head-mounted device in the target scene, wherein the first positioning data is generated by the visual sensing module based on the acquired image, and the second positioning data is provided by the laser positioning module; Based on the environmental perception information, the availability status of the first positioning data and the second positioning data is assessed. Based on the results of the availability assessment, the target positioning mode is determined; Based on the target positioning mode, the first positioning data and / or the second positioning data are processed to determine the pose information used to construct the target virtual reality scene.
2. The method according to claim 1, characterized in that, The results of the availability assessment include a first confidence score of the first location data and a second confidence score of the second location data; The assessment of the availability status of the first location data and the second location data includes: The first confidence score is calculated based on the illumination condition information in the environmental perception information and the visual feature information extracted from the image acquired by the visual sensing module. The second confidence score is calculated based on the spatial structure information in the environmental perception information and the laser signal parameters provided by the laser positioning module.
3. The method according to claim 2, characterized in that, Determining the target positioning mode based on the results of the availability status assessment includes: If the first confidence score is greater than the first threshold and the second confidence score is less than the second threshold, then the target positioning mode is determined to be the first mode, and the pose information in the first mode is determined based on the first positioning data; If the second confidence score is greater than the second threshold, the target positioning mode is determined to be the second mode, and the pose information in the second mode is determined based on the second positioning data.
4. The method according to claim 2, characterized in that, The step of determining the target positioning mode based on the result of the availability status assessment further includes: if the first confidence score is greater than the first threshold and the second confidence score is greater than the second threshold, then the target positioning mode is determined to be the third mode; The step of processing the first positioning data and / or the second positioning data based on the target positioning mode to determine the pose information for constructing the target virtual reality scene includes: In the third mode, the first positioning data and the second positioning data are transformed to the same coordinate system using a preset transformation matrix to obtain the transformed first positioning data and the transformed second positioning data. Based on the first confidence score and the second confidence score, determine the first fusion weight of the converted first location data and the second fusion weight of the converted second location data; The transformed first positioning data and the transformed second positioning data are weighted and fused according to the first fusion weight and the second fusion weight to determine the pose information.
5. The method according to claim 3, characterized in that, The method further includes: Based on the images acquired by the visual sensing module, perspective image information is obtained; Based on the perspective image information and the pose information determined in the second mode, the target virtual reality scene is constructed to realize the perspective function in the target virtual reality scene.
6. The method according to claim 1, characterized in that, The method further includes: When a first handle access signal associated with the laser positioning module is detected, it is determined that interaction using the second positioning data is supported; and / or When a second handle access signal associated with the visual sensing module is detected, it is determined that interaction using the first positioning data is supported.
7. The method according to any one of claims 1 to 6, characterized in that, The visual sensing module includes a processing unit for realizing real-time visual positioning and map building functions, and the laser positioning module is a laser positioning unit that is detachably connected to the head-mounted device.
8. A virtual reality positioning device for a head-mounted device, characterized in that, The head-mounted device includes a visual sensing module and a laser positioning module. The device includes: The environmental perception module is used to acquire environmental perception information of the target scene; The data acquisition module is used to acquire first positioning data and second positioning data that characterize the pose of the head-mounted device in the target scene, wherein the first positioning data is generated by the visual sensing module based on the acquired image, and the second positioning data is provided by the laser positioning module. The status assessment module is used to assess the usability status of the first positioning data and the second positioning data based on the environmental perception information. The pattern determination module is used to determine the target positioning pattern based on the results of the available state assessment. The pose determination module is used to process the first positioning data and / or the second positioning data based on the target positioning mode to determine the pose information for constructing the target virtual reality scene.
9. A head-mounted device, characterized in that, The head-mounted device includes a visual sensing module and a laser positioning module, and the head-mounted device also includes: One or more processors; and A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.