A large-scene meta-universe space superposition method and device
By combining visual images, IMU data, geomagnetic data, and GPS data, a 3D AR digital space is constructed and server-side visual feature matching is performed, solving the problem of poor integration between virtual and real spaces in large scenes and achieving efficient spatial overlay and a natural user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2026-04-14
AI Technical Summary
In large-scale scenes, due to the lack of visual features, traditional visual positioning methods cannot effectively identify real-world scenes, resulting in poor superposition effects of the metaverse space and affecting the fusion effect of virtual space and real space.
By combining visual images, IMU data, geomagnetic data, and GPS data, the geographical coordinates and altitude of the terminal device are obtained, a three-dimensional AR digital space is constructed, and the server is used to perform visual feature matching and scene matching to determine the position of the terminal device in the metaverse virtual space. Finally, the virtual space is integrated into the AR digital space.
It improves the positioning accuracy and speed of terminal devices in large scenes, ensures the natural integration of virtual and real spaces, and enhances the user's viewing comfort and the smoothness of scene transitions.
Smart Images

Figure CN115731370B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of metaverse technology and also to the field of spatial positioning technology, specifically to a method, device, electronic device, computer-readable storage medium, and computer program product for large-scale metaverse spatial overlay. Background Technology
[0002] The metaverse is an open and shared online platform that integrates information technology, communication technology, AR, VR, and other virtual technologies. It is a vast and evolving virtual universe. Although currently limited by the level of technological development, its applications mainly include games, social networking, advertising and marketing, and virtual offices, these applications will expand over time and with technological advancements, attracting more and more people to the metaverse. The metaverse space itself is the fundamental carrier of this virtual universe, and the quality of its construction directly impacts the various scenarios within the metaverse and the user experience.
[0003] Based on the connection between the metaverse and the real world, they can be roughly divided into three types: One is a digital world completely detached from the real world, such as some online games created by game companies, where players are placed in a space completely different from reality. Another is a fused space resulting from the superposition and integration of virtual and real spaces. For example, AR technology currently creates a digital space that completely overlaps with reality based on the current environment, and augments reality by adding virtual markers and 3D models within this space. The third is a virtual space that replicates and simulates real space at a 1:1 scale, such as the virtual space in a virtual office setting. Of course, there are also intersections and fusions of these various situations, such as superimposing a 1:1 virtual space onto the real space, merging it seamlessly with it.
[0004] To integrate digital content in virtual space with real space, AR devices typically employ technologies such as visual recognition and V-SLAM to create a digital space overlaid with reality, and then fuse the digital content into this digital space. When overlaying virtual space onto the digital space, it's necessary to identify the real-world scene and accurately locate the device's position in the real-world space to determine its position in the virtual space. In typical scenarios, identifying the structural features of entities in the real-world scene allows for relatively good scene recognition and localization. However, for large scenes with fewer spatial entities, such as open beaches, ocean surfaces, deserts, or large plazas, traditional visual localization methods cannot effectively identify the real-world scene due to the limited visual features of spatial entities. This affects the accuracy of localization, resulting in a poor overlay effect between virtual and real space, ultimately impacting the fusion of digital content and real-world space. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for overlaying large-scene metaverse spaces, to solve the technical problem that the overlay effect of metaverse spaces is poor in large scenes due to the inability to effectively identify the current scene.
[0006] According to one aspect of the present invention, a method for overlaying virtual spaces in a large-scale metaverse is provided, comprising the following steps:
[0007] The system acquires real-time visual images of the current large-scale scene environment, sensor data from the terminal device, and GPS data from the terminal device. The sensor data from the terminal device includes at least IMU data and geomagnetic data.
[0008] Visual features are extracted from the visual images of the current large scene environment; the IMU data is processed to obtain the pose data of the terminal device; the geographical location planar coordinates of the terminal device are obtained based on GPS data and geomagnetic data;
[0009] A 3D AR digital space that overlaps with the current environment is constructed based on visual images of the current environment and pose data of the terminal device.
[0010] Based on the GPS data or geographic location plane coordinates of the terminal device, a 1:1 metaverse virtual space and a virtual 3D space model of the current large scene are obtained from the metaverse system for overlay.
[0011] Visual positioning is performed based on extracted visual features, geographic location planar coordinates, and a virtual 3D spatial model of the current large scene to obtain the 3D coordinates of the real space;
[0012] The virtual spatial position of the terminal device in the metaverse virtual space is determined based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and
[0013] The metaverse virtual space is integrated into the AR digital space according to the virtual space location.
[0014] According to another aspect of the present invention, a method for overlaying a large-scale metaverse virtual space is provided, the method being applied to a terminal device, comprising the following steps:
[0015] The system acquires real-time visual images of the current large-scale scene environment, sensor data from the terminal device, and GPS data from the terminal device. The sensor data from the terminal device includes at least IMU data and geomagnetic data.
[0016] Visual features are extracted from the visual images of the current large scene environment; the IMU data is processed to obtain the pose data of the terminal device; the geographical location planar coordinates of the terminal device are obtained based on GPS data and geomagnetic data;
[0017] A 3D AR digital space that overlaps with the current environment is constructed based on visual images of the current environment and pose data of the terminal device.
[0018] Send the terminal device's GPS data or geographic location plane coordinates and requests to the server to obtain a 1:1 metaverse virtual space for overlay;
[0019] Send the visual image of the current large scene environment or the visual feature information extracted based on the visual image of the current large scene environment to the server, as well as the geographical location planar coordinates of the terminal device, and receive one or more first estimated three-dimensional coordinates from the server.
[0020] When multiple first estimated three-dimensional coordinates are obtained, air pressure data is obtained from the sensor data, and the altitude of the terminal device is calculated based on the air pressure data.
[0021] Calculate the difference between the altitude and the height coordinate in each first estimated three-dimensional coordinate, and take the first estimated three-dimensional coordinate with the smallest difference as the real space three-dimensional coordinate;
[0022] The virtual spatial position of the terminal device in the metaverse virtual space is determined based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and
[0023] The metaverse virtual space is integrated into the AR digital space according to the virtual space location.
[0024] According to another aspect of the present invention, a method for overlaying a large-scale metaverse virtual space is provided, the method being applied to a terminal device, comprising the following steps:
[0025] The system acquires real-time visual images of the current large-scale scene environment, sensor data from the terminal device, and GPS data from the terminal device. The sensor data from the terminal device includes IMU data, geomagnetic data, and air pressure data.
[0026] The IMU data is processed to obtain the pose data of the terminal device; the geographical location plane coordinates of the terminal device are obtained based on GPS data and geomagnetic data; the altitude of the terminal device is calculated based on the air pressure data.
[0027] A 3D AR digital space that overlaps with the current environment is constructed based on visual images of the current environment and pose data of the terminal device.
[0028] Send the terminal device's GPS data or geographic location plane coordinates and requests to the server to obtain a 1:1 metaverse virtual space for overlay;
[0029] Send the visual image of the current large scene environment, the geographical location planar coordinates of the terminal device, and the altitude of the terminal device to the server, and receive the three-dimensional coordinates of the terminal device in the real space of the current large scene from the server;
[0030] The virtual spatial position of the terminal device in the metaverse virtual space is determined based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and
[0031] The metaverse virtual space is integrated into the AR digital space according to the virtual space location.
[0032] According to another aspect of the present invention, a method for overlaying virtual spaces in a large-scale metaverse is provided, the method being applied to a server and comprising the following steps:
[0033] When responding to a location request sent by a terminal device, the device obtains its geographic location plane coordinates and visual feature information of the current large scene environment from the location request.
[0034] Based on the virtual 3D space model of the current large scene, the extracted visual features are matched to obtain a predicted position array. The predicted position array includes multiple predicted 3D coordinates that each correspond to a predicted scene that matches the visual features.
[0035] Based on the geographic location planar coordinates, one or more first estimated three-dimensional coordinates are selected from the estimated location array and sent to the terminal device; and
[0036] In response to receiving a metaverse virtual space request from a terminal device, the system returns the corresponding metaverse virtual space data to the terminal device.
[0037] According to another aspect of the present invention, a method for overlaying virtual spaces in a large-scale metaverse is provided, the method being applied to a server and comprising the following steps:
[0038] When responding to a location request sent by a terminal device, the device obtains the terminal device's geographic location plane coordinates, the visual image of the current large scene environment, and the terminal device's altitude from the location request.
[0039] Visual features are extracted from the visual images of the current large-scale scene environment;
[0040] The matching range in the virtual 3D space model of the current large scene is determined based on the altitude.
[0041] Scene matching is performed on the extracted visual features within the matching range to obtain a predicted location array, which includes multiple predicted three-dimensional coordinates corresponding to a predicted scene that matches the visual features.
[0042] Based on the geographic location planar coordinates, an estimated three-dimensional coordinate is selected from the estimated location array as the actual spatial three-dimensional coordinate, and then sent to the terminal device; and
[0043] In response to receiving a metaverse virtual space request from a terminal device, the system returns the corresponding metaverse virtual space data to the terminal device.
[0044] According to another aspect of the present invention, a large-scale metaverse virtual space overlay device is provided, the device being applied to a terminal device and comprising the following modules:
[0045] The client communication module is configured to transmit data with the server.
[0046] The data acquisition module is configured to acquire in real time visual images of the current large scene environment, terminal device sensor data, and terminal device GPS data, wherein the terminal device sensor data includes at least IMU data and geomagnetic data.
[0047] The data preprocessing module, connected to the data acquisition module, is configured to extract visual features based on the visual image of the current large scene environment; process the IMU data to obtain the pose data of the terminal device; and obtain the geographical location planar coordinates of the terminal device based on GPS data and geomagnetic data.
[0048] A digital space construction module, which is connected to the data acquisition module and the data preprocessing module, is configured to construct a three-dimensional AR digital space that overlaps with the current environment based on the visual images of the current environment and the pose data of the terminal device.
[0049] The request module, which is connected to the data acquisition module, the data preprocessing module, and the client communication module, is configured to send a location request to the server, the location request including at least visual feature information and geographic location planar coordinates; and to send a metaverse virtual space request to the server.
[0050] The first positioning module is connected to the client communication module and the data acquisition module, and is configured to determine one of the multiple first estimated three-dimensional coordinates corresponding to an estimated scene from the positioning data returned by the server as the real space three-dimensional coordinate.
[0051] The second positioning module, connected to the first positioning module, is configured to determine the virtual spatial position of the terminal device in the metaverse virtual space based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and
[0052] The spatial fusion module, which is connected to the digital space construction module, the client communication module, and the second positioning module, is configured to fuse the metaverse virtual space returned by the server into the three-dimensional AR digital space according to the virtual space location for enhanced display.
[0053] According to another aspect of the present invention, the present invention provides a large-scale metaverse virtual space overlay device, the device being applied to a server and comprising the following modules:
[0054] The server-side communication module is configured to transmit data with the terminal device, wherein it receives a location request and a metaverse virtual space request sent by the terminal device. The location request includes the geographic location plane coordinates and the visual feature information of the current large scene environment.
[0055] A visual positioning module, which is connected to the server communication module, is configured to perform scene matching on visual features based on the virtual three-dimensional space model of the current large scene based on the positioning request to obtain a predicted position array. The predicted position array includes multiple predicted three-dimensional coordinates that each correspond to a predicted scene that matches the visual features.
[0056] A location data filtering module, connected to the visual positioning module and a server-side communication module, is configured to filter one or more first estimated three-dimensional coordinates from the estimated location array based on geographic location planar coordinates, and send them to the terminal device via the server-side communication module; and
[0057] The resource processing module, which is connected to the server communication module, is configured to return the corresponding metaverse virtual space data to the terminal device when a metaverse virtual space request is sent by the terminal device.
[0058] According to another aspect of the present invention, an electronic device is provided, including a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the aforementioned method for overlaying large-scene metaverse virtual space in a terminal device; or, the processor, when executing the computer program instructions, implements the aforementioned method for overlaying large-scene metaverse virtual space in a server.
[0059] According to another aspect of the present invention, a computer-readable storage medium is provided, characterized in that the computer storage medium stores computer program instructions, which, when executed by a processor, implement the aforementioned method for overlaying large-scene metaverse virtual space in a terminal device; or, implement the aforementioned method for overlaying large-scene metaverse virtual space in a server.
[0060] According to another aspect of the present invention, the present invention provides a computer program product, characterized in that it includes computer program instructions, which, when executed by a processor, implement the aforementioned method for overlaying large-scene metaverse virtual space in a terminal device; or, implement the aforementioned method for overlaying large-scene metaverse virtual space in a server.
[0061] In situations where the terminal device is located in a large scene with few visual features and a monotonous environment, this invention can quickly determine the spatial location of the terminal device by using geographic location coordinates. This reduces the computational load of visual positioning and improves the accuracy and speed of positioning. As a result, it can respond to the movement speed of the terminal user in a large scene, making the fused scene seen by the terminal user natural, and the scene transitions and changes smooth, thereby improving the viewing comfort of the terminal user. Attached Figure Description
[0062] To more clearly illustrate the implementation of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below.
[0063] Figure 1 This is a schematic diagram of an AR system architecture based on a server and a terminal device according to an embodiment of the present invention.
[0064] Figure 2 This is a schematic diagram of a virtual-real image fusion method for AR navigation using a mobile app.
[0065] Figure 3 This is a flowchart of a method for overlaying a large-scene metaverse virtual space according to an embodiment of the present invention.
[0066] Figure 4 This is a flowchart illustrating the positioning initialization process achieved through interaction between a terminal device and a server in a large-scene metaverse virtual space overlay method according to an embodiment of the present invention.
[0067] Figure 5 This is a schematic diagram of visual computing principle according to an embodiment of the present invention.
[0068] Figure 6 This is a flowchart of the server positioning process in a large-scene metaverse virtual space overlay method according to an embodiment of the present invention.
[0069] Figure 7 This is a flowchart of the positioning process of a terminal device in a large-scene metaverse virtual space overlay method according to another embodiment of the present invention.
[0070] Figure 8 Is with Figure 7 The positioning processing flow diagram of the server corresponding to the positioning processing flow of the terminal device described in the document.
[0071] Figure 9 This is a schematic diagram of a large-scene metaverse virtual space overlay device applied to a terminal device according to an embodiment of the present invention.
[0072] Figure 10 This is a schematic diagram of a large-scale metaverse virtual space overlay device applied to a server according to an embodiment of the present invention.
[0073] Figure 11 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention.
[0074] Figure 12 This is a schematic diagram of the software structure of an exemplary terminal device according to an embodiment of the present invention. Detailed Implementation
[0075] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided to make the principles and spirit of the present invention clearer and more thorough, enabling those skilled in the art to better understand and implement the principles and spirit of the present invention. The exemplary embodiments provided herein are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments described herein without inventive effort are within the scope of protection of the present invention.
[0076] Those skilled in the art will recognize that embodiments of the present invention can be implemented as a system, apparatus, device, method, computer-readable storage medium, or computer program product. Therefore, the present invention can be specifically implemented in at least one of the following forms: entirely hardware, entirely software, or a combination of hardware and software. Based on a specific embodiment of the present invention, protection is claimed for one such embodiment.
[0077] In this document, terms such as first, second, and third are used only to distinguish one entity (or operation) from another, and are not intended to require or imply any order or relationship between these entities (or operations).
[0078] Embodiments of the present invention can be applied to servers and terminal devices. Please refer to... Figure 1This diagram schematically illustrates an AR system architecture based on a server and terminal devices. The AR system architecture includes a server 10 and several terminal devices 20. In some examples, the terminal devices 20 are AR devices, which can be dedicated AR devices such as head-mounted displays (HMDs), smart gloves, clothing, and other smart wearable electronic devices. In other examples, the terminal devices 20 can be general-purpose AR devices, such as mobile phones, laptops, tablets, virtual reality (VR) devices, in-vehicle devices, navigation devices, gaming devices, etc.
[0079] Taking AR helmets or AR glasses as an example, a head-mounted display, machine vision system, and mobile computer can be integrated into a wearable device. This device has a display resembling glasses and is worn on the user's head. It transmits augmented reality information to the display or projects it onto the user's eyes, enhancing the user's visual immersion. In some examples, AR devices also include cameras, which can be wide-angle, telephoto, or structured light cameras (also known as point cloud depth cameras, 3D structured light cameras, or depth cameras). Structured light cameras, based on 3D vision technology, can acquire the planar and depth information of objects. A structured light camera projects light with specific structural features onto the object being photographed using a near-infrared laser. The reflected light is then collected by an infrared camera and processed by a processor chip. The calculation principle involves calculating the object's position and depth information based on changes in the light signal caused by the object, presenting a 3D image. Typical terminal devices, such as mobile phones, display two-dimensional images and cannot show the depth of different locations within the image. Structured light cameras can capture and acquire 3D image information, obtaining not only color and other information at different locations but also depth information, which can be used for AR ranging. Of course, ordinary terminal devices can also acquire 2D images using optical cameras and combine this with deep learning algorithms to obtain depth information, ultimately displaying 3D images as well.
[0080] In some examples, terminal device 20 has AR-enabled software or an application (APP) installed. Server 10 can be a management server or application server for this software or APP. Server 10 can be a single server, a server cluster consisting of multiple servers, or a cloud server, etc. Terminal device 20 integrates modules with networking capabilities, such as Wireless-Fidelity (Wi-Fi) modules, Bluetooth modules, 2G / 3G / 4G / 5G communication modules, etc., to connect to server 10 via a network.
[0081] For example, users can log in to their user accounts through an app installed on their mobile phones, or through software installed on AR glasses.
[0082] Taking an AR navigation app as an example, the app can possess capabilities such as high-precision map navigation, environmental understanding, and virtual-real fusion rendering. The app can report its current geographical location information to the server 10 through the terminal device 20, and the server 10 provides AR navigation services to the user based on the real-time geographical location information. For example, if the terminal device 20 is a mobile phone, in response to the user's operation of launching the app, the mobile phone can activate its camera to capture images of the real environment. Then, the system performs AR enhancement on the real environment images captured by the camera, integrating or overlaying rendered AR effects (such as navigation route signs, road names, merchant information, advertising displays, etc.) into the real environment images, and displaying the virtual-real fusion image on the mobile phone screen.
[0083] Figure 2 The illustration schematically shows a virtual-real fusion image for AR navigation using a mobile app, where the AR navigation pointer arrows are superimposed on the real road surface and space in the image, and the electronic promotional materials of merchants float in the space in the form of parachutes carrying gift boxes at designated locations.
[0084] Embodiments of the present invention relate to terminal devices and / or servers. The principles and spirit of the present invention will be explained in detail below through several exemplary embodiments or representative implementations.
[0085] Figure 3 This is a flowchart of a method for overlaying a large-scene metaverse virtual space according to an embodiment of the present invention. In this embodiment, the method is applied to an AR application, such as an exhibition held in the metaverse virtual space, a created building or building complex, or a manufactured landscape. When an end user launches the AR application to view it, the metaverse virtual space needs to be enhanced and displayed in the current scene. Specifically, it includes the following steps:
[0086] Step S11: The terminal device collects data and sends a request for the metaverse virtual space. The terminal device collects visual images at a certain frequency, and some sensors within the terminal device collect data at a certain frequency. These sensors include, for example, an IMU unit that collects the terminal device's three-axis acceleration and three-axis angular velocity; a geomagnetic sensor for collecting geomagnetic data; and an optional barometric pressure sensor for collecting barometric pressure data at the current altitude. The GPS module in the terminal device obtains GPS data.
[0087] Step S12 involves data processing to obtain the corresponding dataset. This includes processing the acquired visual images to obtain the visual features of the current space. These visual features include information such as points, lines, surfaces, colors, textures, and depth that constitute the spatial structure, thus forming a visual feature set C. Ti Here, Ti corresponds to the capture time of the visual image. For example, the terminal device first filters the camera based on the environmental visual image, such as using median filtering or Gaussian filtering algorithms to remove noise and reduce interference. Then, calibration is performed, which includes calculating the intrinsic and extrinsic parameters and distortion parameters of the camera. This includes determining the pixel coordinate system, image coordinate system, camera coordinate system, and world coordinate system based on the acquired image and lens data, and calculating the overall relationship between the four coordinate systems to obtain the intrinsic parameter matrices of the left and right cameras. The left and right distortion coefficients are calculated based on the lens data of the cameras, thereby obtaining the radial distortion and tangential distortion, which constitute the distortion matrix. Based on the Kruppa equation and Zhang Zhengyou calibration method, the rotation and translation matrices of one camera relative to another are obtained. The distortion matrix, rotation matrix, and translation matrix are collectively referred to as the extrinsic parameter matrix. Then, based on the camera's intrinsic and extrinsic parameters and distortion parameters, stereo correction is performed on the real-time acquired environmental visual image, including correcting distortion errors, changing the viewing angle, and horizontal objects, thereby obtaining a distortion-free and corrected image located on the same plane. Then, the terminal device performs visual feature recognition based on the acquired visual images of the current environment. For example, a pre-trained 3D convolutional network model can be used to detect and extract features from each frame of the image, obtaining multiple visual features, including information such as points, lines, surfaces, colors, textures, and depth that constitute the spatial structure. For the large scene space in this invention, only a very limited number of visual features can usually be obtained. For example, in an open square, the visual features obtained might include the ground, buildings at the edge of the square, and other visual features.
[0088] Regarding the acceleration and angular velocity acquired by the IMU unit, since the acquisition frequency of IMU data is greater than that of visual images, in one embodiment, when processing IMU data, the acquisition period of visual images is used as the processing time period of IMU data. After pre-integration, integration and other processing, the angular offset and position offset of the terminal device corresponding to two frames of visual images are obtained, thereby obtaining the current pose data.
[0089] The geographical location plane coordinates of the terminal device are obtained based on GPS and geomagnetic data. Since GPS data can accurately represent the location of a moving terminal device, while geomagnetic data provides a more accurate location when the device is stationary, in one embodiment, the motion state of the terminal device is first determined based on IMU data. For example, if the triaxial acceleration value in the IMU data is 0 for a given period, the terminal device is determined to be stationary; if the triaxial acceleration value is not 0, the terminal device is determined to be moving. When the terminal device is stationary, its geographical location plane coordinates are determined based on geomagnetic data; when the terminal device is moving, its geographical location plane coordinates are determined based on GPS data. Specifically, geomagnetic sensors detect geomagnetic information in the current scene, and after preprocessing, geomagnetic positioning data such as total magnetic field strength, vector strength, magnetic inclination, magnetic declination, and intensity gradient are obtained. Then, based on the current geomagnetic positioning data, a pre-set geomagnetic positioning information table for the current site is retrieved to obtain the geographical location plane coordinates of the terminal device. In addition, to maintain the accuracy of geomagnetic data and reduce information errors caused by interference, geomagnetic data is calibrated based on GPS data when the terminal device is in motion.
[0090] The altitude of the terminal device is calculated based on the air pressure data collected by the barometric pressure sensor. Under normal circumstances, this altitude is approximately equal to the height of the terminal device above the ground in the current scene.
[0091] Step S13: Determine the location of the terminal device. This includes the location initialization process and the real-time location process, see details below. Figure 4 And its explanation. Step S13 determines the position of the terminal device in the real space, and then determines its virtual position in the metaverse virtual space. The virtual position includes the three-dimensional coordinates and pose of a point in space. The pose is specifically the angular offset relative to the three axes of the spatial coordinate system. Therefore, the virtual position can not only express the position of the terminal device at a specific point in space, but also determine the specific orientation it is facing, which corresponds to a specific spatial scene.
[0092] Step S14: Construct a 3D AR digital space that overlaps with the current environment based on the visual images of the current environment and the pose data of the terminal device.
[0093] Step S15: After obtaining the virtual position of the terminal device in the metaverse virtual space, the metaverse virtual space is merged into the current 3D AR digital space according to this position. This means that the metaverse virtual space can be added to the terminal device for display. The metaverse virtual space can be, for example, a virtual scene of the current real-world scene, plus virtual buildings, landscapes, exhibitions, etc., added within that virtual scene. The current real-world scene and its virtual scene are in a 1:1 ratio, and the virtual buildings, landscapes, exhibitions, etc., added within the virtual scene are also in a 1:1 ratio with the real scene. Therefore, when the metaverse virtual space is merged into the current 3D AR digital space according to the terminal device's virtual position in the metaverse virtual space, the terminal user can have an immersive experience. For example, when the metaverse virtual space includes a group of virtual buildings, through positioning, the terminal user sees a specific location within the merged virtual building group through the terminal device, and the building group seen corresponds to their position in the scene. Furthermore, since the positioning method provided in step S13 of this invention is fast and accurate, it can quickly respond to the movement speed of the end user in the scene, thereby ensuring that the scene seen by the end user through the terminal device is natural, and the scene transitions and changes are smooth, thus providing the end user with viewing comfort.
[0094] Step S16: Determine whether the application has ended. If it has ended, stop the overlay and positioning processes. If not, return to step S13 to continue positioning, obtain the real-time position, and in step S15, adjust the display angle of the metaverse virtual space in the three-dimensional AR digital space according to the real-time position.
[0095] Figure 4 This is a flowchart illustrating the positioning initialization process in a large-scene metaverse virtual space overlay method according to an embodiment of the present invention, where a terminal device interacts with a server.
[0096] See Figure 4 In step S131, the terminal device sends a location request to the server. The location request includes a visual image of the current large-scale scene environment or visual feature information extracted from the visual image of the current large-scale scene environment, as well as the geographical location planar coordinates of the terminal device. In one embodiment, the terminal device obtains a visual feature set C after processing an image. Ti The processed images are then sent to the server. In another embodiment, the terminal device processes multiple images and obtains corresponding visual feature sets C. Ti Then, multiple visual feature sets CTi are simultaneously sent to the server. Alternatively, the terminal device does not perform image processing but directly sends the captured visual image of the current large scene environment to the server, which then performs visual feature recognition based on the current environment visual image to obtain the visual feature set C. Ti .
[0097] Step S231: After receiving the location request sent by the terminal device, the server obtains the visual feature set C from it. Ti And the geographical location plane coordinates of the terminal equipment.
[0098] Step S232 involves retrieving and matching scenes in the database based on the 3D virtual scene model of the current location to obtain an estimated location array. Specifically, the server retrieves the corresponding 3D virtual scene model from the database based on the geographic location's planar coordinates in the location request. Since the scene where the terminal device is located in this invention is a large scene with relatively few spatial features, the aforementioned visual feature set C... Ti The set of visual features is relatively small, so when performing scene retrieval and matching, multiple sets of visual features C will be obtained. Ti The matched scene, referred to here as the estimated scene, has position coordinates calibrated for the 3D virtual scene model in the database. Therefore, after determining the estimated scene, the corresponding estimated 3D coordinates can be obtained, thus forming an estimated position array. In determining the estimated scene, in one embodiment, the matched scenes are sorted in descending order of matching degree, and a preset number S of scenes at the top of the list are selected. i As a predicted scenario. In another embodiment, a matching threshold is set, and scenarios S with a matching degree greater than the matching threshold are considered as predicted scenarios. i As a predicted scene. When multiple predicted scenes are obtained, the corresponding position coordinates (x, y, y) are acquired. 0i ,y 0i ,z 0i ) is used as the estimated 3D coordinates, at which point the estimated position array is obtained, which includes multiple estimated 3D coordinates, such as [(x 01 ,y 01 ,z 01 ),(x 02 ,y 02 ,z 02 )……(x 0i ,y 0i ,z 0i )).
[0099] Step S233: Based on the geographic location planar coordinates, one or more first estimated three-dimensional coordinates are selected from the estimated location array. Specifically, the geographic location planar coordinates (x... p ,y p ) and the three-dimensional coordinates (x) in the estimated position array 0i ,y 0i ,z 0i Perform intersection calculations. Determine if the estimated location array contains any coordinates (x, y) that correspond to the geographic location's planar coordinates. p ,y p If the three-dimensional coordinates are the same, then they will be compared with the geographic location plane coordinates (x, y). p ,yp The same three-dimensional coordinates are used as the first estimated three-dimensional coordinates. If not, the geographic location plane coordinates (x, y) are calculated. p ,y p The difference between the x and y coordinates in each estimated 3D coordinate system and the estimated 3D coordinate system is used as the first estimated 3D coordinate system. In another embodiment, the estimated 3D coordinate system is determined according to the geographical location plane coordinates (x, y, ...). p ,y p The estimated 3D coordinates are sorted by the difference between their x-coordinates and y-coordinates in each estimated 3D coordinate system, and the estimated 3D coordinates with the smallest preset number of differences are selected as the first estimated 3D coordinates. In another embodiment, the x-coordinates and y-coordinates of the multiple first estimated 3D coordinates can be calibrated, for example, by calculating the difference between the x-coordinate in the determined first estimated 3D coordinates and the x-coordinate in the geographic location plane coordinates. p The average value of the coordinates is used as the x-coordinate of the first estimated three-dimensional coordinates, and the y-coordinate of the first estimated three-dimensional coordinates is obtained in the same way. This ensures that the first estimated three-dimensional coordinates sent to the terminal device are calibrated coordinates, thereby improving the positioning accuracy.
[0100] Step S234: Return visual positioning data to the terminal device. In one embodiment, the server sends the estimated location array of the scene corresponding to each shooting time to the terminal device, or, if processing time allows, the server sends the estimated location arrays of the scene corresponding to multiple shooting times in batches to the terminal device.
[0101] In step S132, after receiving the estimated location data returned by the server, the terminal device determines one of the multiple first estimated three-dimensional coordinates as the actual three-dimensional coordinate based on the altitude obtained from the air pressure data. In one embodiment, the difference between the altitude and the height coordinate in each of the first estimated three-dimensional coordinates is calculated, and the first estimated three-dimensional coordinate with the smallest difference is used as the actual three-dimensional coordinate. In another embodiment, the height coordinate of the first estimated three-dimensional coordinate with the smallest difference is calibrated. One calibration method is to calculate the average value of the altitude and the height coordinate of the current first estimated three-dimensional coordinate, and use the average value as the actual three-dimensional coordinate.
[0102] Step S133: Use the currently determined three-dimensional coordinates of the real space as the initial positioning coordinates, and obtain the angle offset and position offset from the IMU data of the current scene as the initial pose.
[0103] Step S134: Based on the current initial positioning position, calibrate the spatial position coordinates corresponding to the IMU accumulated position.
[0104] In the aforementioned positioning initialization process, the server can quickly filter out more reliable location coordinates from multiple estimated 3D coordinates using the assistance of geographic location coordinates. This reduces processing load, speeds up the initial positioning process for the terminal device, and improves the accuracy of the initial positioning location. When the terminal device is located in a large scene with few or monotonous visual features, making it impossible to obtain an accurate location, there is no need to repeatedly acquire visual images and extract visual features.
[0105] In another embodiment of the above-described positioning initialization, after calculating the difference between the altitude and the height coordinate in each of the first estimated three-dimensional coordinates, the difference is further compared with a threshold. When the difference is greater than or equal to the threshold, a height of the terminal device relative to the ground is obtained through visual calculation, and then the distance is used to replace the height coordinate in the first estimated three-dimensional coordinates. The first estimated three-dimensional coordinates with the replaced height coordinates are then used as the real-world three-dimensional coordinates. The visual calculation process, for example, first identifies the ground plane using multiple current visual images, and then calculates a visual depth map. When calculating the visual depth, the distance of the terminal device's camera relative to the current ground can be obtained, such as... Figure 5 The hypotenuses l1 and l2 of the two triangles in the diagram can be determined, and the distance d between the two hypotenuses can be obtained. Then, by solving the triangles, the distance H from the terminal device's camera to the ground can be obtained. In this embodiment, since the altitude obtained from air pressure data may be inaccurate, a relatively accurate distance from the terminal device's camera to the ground can be obtained through plane detection, depth calculation, and triangle solving, thereby improving the accuracy of the obtained three-dimensional coordinates of the real space.
[0106] Then, a real-time positioning process is performed based on the current initial position. This real-time positioning process includes the same procedures as the initial positioning process described above, plus pose calculation and cumulative error correction. In some embodiments, when processing the height coordinates of the first estimated three-dimensional coordinates sent by the server, position data from the pose data can also be referenced. For example, based on the three-axis acceleration data in the current IMU data, the current three-axis position offset is obtained after integration. Then, based on the position data from the previous pose data, the current spatial position, i.e., the three-dimensional spatial coordinates (x, y, x) can be calculated. I ,y I ,z I ). Using the aforementioned three-dimensional spatial coordinates (x... I ,y I ,z IThe first estimated 3D coordinates are calibrated to obtain the actual 3D coordinates in space. This calibration may involve calculating the average value of the coordinates and using this average value as the final actual 3D coordinates in space, or keeping the first estimated 3D coordinates unchanged when the difference between the two is small. By calibrating various data during the real-time positioning process, consistency of various data types is maintained, and accumulated errors in some data (such as IMU data and geomagnetic data) are eliminated. This not only improves positioning accuracy but also positioning speed and responsiveness to the actual movement speed of the terminal.
[0107] Figure 6 This is a flowchart illustrating the server's positioning process in a large-scene metaverse virtual space overlay method according to an embodiment of the present invention. In this embodiment, the terminal device and the server jointly implement the large-scene metaverse virtual space overlay method, and the overlay process of the metaverse virtual space is the same as described above. Figure 3 The process is the same and will not be repeated here. The terminal positioning process is as follows: Figure 4 As shown, the process will not be repeated here. The following description will explain the server's location processing flow.
[0108] In step S231a, after receiving the location request from the terminal device, the server obtains the geographical location planar coordinates of the terminal device and multiple visual images from the location request, and processes the visual images to obtain the visual feature set C corresponding to each visual image. Ti .
[0109] Step S232a: Perform visual calculations based on multiple visual images to obtain the distance H from the terminal device's camera to the ground.
[0110] Step S233a: In the database, the distance H determines the first scene matching range for scene matching of the three-dimensional virtual scene model at the current location.
[0111] Step S234a: Determine a second matching range within the first scene matching range based on the geographic location planar coordinates;
[0112] Step S235a: Perform scene matching on the extracted visual features within the second matching range to obtain a predicted position array, wherein the predicted position array includes one or more first predicted three-dimensional coordinates.
[0113] In this embodiment, when performing visual positioning, the server determines the matching range by using the distance H from the terminal device's camera to the ground and the geographical location's planar coordinates obtained through visual calculation, which greatly reduces the amount of matching calculation and thus effectively improves the positioning speed.
[0114] Figure 7 This is a flowchart of the positioning process of a terminal device in a large-scene metaverse virtual space overlay method according to another embodiment of the present invention. Figure 8 This is the corresponding server location processing flowchart. In this embodiment, the terminal device and the server jointly implement the large-scene metaverse virtual space overlay method. The metaverse virtual space overlay process is the same as described above. Figure 3 The same applies, so I won't repeat it here.
[0115] In step S131b, the terminal device sends a location request to the server. The location request includes a visual image of the current large scene environment or visual feature information extracted from the visual image of the current large scene environment, the geographical location plane coordinates of the terminal device, and the altitude of the terminal device calculated based on air pressure data.
[0116] Step S132b: Receive the positioning data returned from the server and obtain the three-dimensional coordinates from it.
[0117] Step S133b: Calculate the difference between the three-dimensional coordinates and the position coordinates obtained when calculating the pose.
[0118] Step S134b: Determine whether the difference is greater than or equal to a threshold. If it is greater than or equal to the threshold, in step S135b, the position coordinates obtained during pose calculation are used to calibrate the current three-dimensional coordinates of the real space, and in step S136b, the calibrated coordinates are used as the three-dimensional coordinates of the real space. If the difference is less than the threshold, in step S137b, the coordinates currently received from the server are used as the three-dimensional coordinates of the real space.
[0119] See Figure 8 In step S231b, a positioning request sent by the terminal device is received, and the geographical location plane coordinates of the terminal device, the visual image of the current large scene environment, and the altitude of the terminal device are obtained from the positioning request.
[0120] Step S232b: Extract visual features from the visual image of the current large scene environment.
[0121] Step S233b: Determine the matching range in the virtual 3D space model of the current large scene based on the altitude.
[0122] Step S234b: The extracted visual features are matched with the scene within the matching range to obtain a predicted position array. The predicted position array includes multiple predicted three-dimensional coordinates that each correspond to a predicted scene that matches the visual features.
[0123] Step S235b: Based on the geographic location planar coordinates, select an estimated three-dimensional coordinate from the estimated location array as the real space three-dimensional coordinate, and send it to the terminal device.
[0124] In this embodiment, the altitude of the terminal device, calculated based on air pressure data, is fully utilized. By using the altitude of the terminal device, the scene matching range is narrowed, thereby reducing the amount of calculation and improving the positioning speed.
[0125] It should be noted that, for clarity, all embodiments of the present invention are described as a combination of a series of actions or processes. Those skilled in the art should understand that the implementation process is not limited by the order of the described actions or processes, and some steps in the embodiments of the present invention may be processed in other orders or simultaneously.
[0126] Figure 9 This is a schematic diagram of a large-scene metaverse virtual space overlay device applied to a terminal device according to an embodiment of the present invention, including a client communication module 11, a data acquisition module 12, a data preprocessing module 13, a digital space construction module 14, a request module 15, a first positioning module 16, a second positioning module 17, and a spatial fusion module 18.
[0127] The client communication module 11 is used for data transmission with the server. The data acquisition module 12 is used to acquire in real time the visual images of the current large-scale scene environment, the sensor data of the terminal device, and the GPS data of the terminal device. The sensor data of the terminal device includes at least IMU data and geomagnetic data. For example, the data acquisition module 12 controls the camera of the terminal device to capture images at a certain shooting frequency, reads the sensor data of each sensor at a certain frequency, and reads GPS data at a certain frequency. This data is then sent to the data preprocessing module 13.
[0128] The data preprocessing module 13 is connected to the data acquisition module 12 and is used to process various types of data. For example, it performs visual feature extraction based on the visual image of the current large scene environment; processes the IMU data to obtain the pose data of the terminal device; obtains the geographical location plane coordinates of the terminal device based on GPS data and geomagnetic data; and obtains the altitude based on air pressure data.
[0129] The digital space construction module 14 is connected to the data acquisition module 12 and the data preprocessing module 13. Based on the visual images of the current environment and the pose data of the terminal device, it constructs a three-dimensional AR digital space that overlaps with the current environment. For example, a spatial map is constructed based on SLAM technology and displayed in real time, thereby forming a three-dimensional AR digital space that overlaps with the current real-world scene.
[0130] The request module 15 is connected to the data acquisition module 12, the data preprocessing module 13 and the client communication module 11, and is used to send a location request to the server, wherein the location request includes at least visual feature information and geographic location plane coordinates; at the same time, it sends a metaverse virtual space request to the server.
[0131] The first positioning module 16 communicates with the client module 11 and the data acquisition module 12, and determines one of the multiple estimated three-dimensional coordinates corresponding to a predicted scene from the positioning data returned by the server based on the altitude of the terminal device as the real-world three-dimensional coordinate. Alternatively, it determines one of the multiple estimated three-dimensional coordinates corresponding to a predicted scene from the positioning data returned by the server based on the height obtained through visual calculation as the real-world three-dimensional coordinate.
[0132] The second positioning module 17 is connected to the first positioning module 16, and determines the virtual space position of the terminal device in the metaverse virtual space based on the three-dimensional coordinates of the real space and the current pose data of the terminal device.
[0133] The spatial fusion module 18 is connected to the digital space construction module 14, the client communication module 11, and the second positioning module 17, and is used to fuse the metaverse virtual space returned by the server into the three-dimensional AR digital space according to the virtual space position for enhanced display.
[0134] Figure 10 This is a schematic diagram of a large-scene metaverse virtual space overlay device applied to a server according to an embodiment of the present invention. In this embodiment, the device includes a server-side communication module 21, a visual positioning module 22, a positioning data filtering module 23, and a resource processing module 24. The server-side communication module 21 is used for data transmission with a terminal device, receiving positioning requests and metaverse virtual space requests sent by the terminal device. The positioning request includes geographic location planar coordinates and visual feature information of the current large-scene environment. The visual positioning module 22 is connected to the server-side communication module 21. Based on the positioning request, it performs scene matching on the visual features based on the virtual three-dimensional space model of the current large scene to obtain a predicted position array. The predicted position array includes multiple predicted three-dimensional coordinates, each corresponding to a predicted scene that matches the visual features.
[0135] The positioning data filtering module 23 is connected to the visual positioning module 22 and the server communication module 21. Based on the geographic location plane coordinates, it filters one or more first estimated three-dimensional coordinates from the estimated location array and sends them to the terminal device via the server communication module.
[0136] The resource processing module 24 is connected to the server communication module 21. When it receives a metaverse virtual space request sent by the terminal device, it returns the metaverse virtual space data corresponding to the request to the terminal device.
[0137] In an optional embodiment, the device further includes a visual computing module 25 connected to the server communication module 21. This module calculates the height H of the terminal device relative to the ground based on multiple current visual images and sends the distance H to the visual positioning module 22. When performing scene matching, the visual positioning module 22 first determines the matching range within the virtual 3D spatial model of the current large scene based on the height of the terminal device relative to the ground. Then, it performs scene matching on the extracted visual features within this matching range to obtain an estimated position array. Additionally, when the positioning request includes the altitude of the terminal device, the visual positioning module 22, when performing scene matching, first determines the matching range within the virtual 3D spatial model of the current large scene based on the height of the terminal device relative to the ground, and then performs scene matching on the extracted visual features within this matching range to obtain an estimated position array.
[0138] Those skilled in the art will understand that the embodiments described herein are preferred embodiments, and the actions, steps, modules, or units involved are not necessarily essential to the embodiments of the present invention. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0139] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 60 includes a processor 61, a memory 62, and a communication bus for connecting the processor 61 and the memory 62. The memory 62 stores a computer program that can run on the processor 61. When the processor 61 runs the computer program, it can execute or implement the steps of the methods in the various embodiments of the present invention. The electronic device 60 also includes a communication interface for receiving and sending data. The electronic device 60 can be a server in the embodiments of the present invention, or it can be a cloud server. The electronic device 60 can also be a terminal device or an AR device in the embodiments of the present invention. Where appropriate, the electronic device can also be referred to as a computing device.
[0140] In some embodiments, processor 61 may be a central processing unit (CPU), graphics processing unit (GPU), application processor (AP), modem processor, image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, neural-network processing unit (NPU), etc. Processor 61 may also be other general-purpose processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors may be microprocessors or any conventional processor. The neural network processor (NPU), by drawing inspiration from biological neural network structures, can rapidly process input information and continuously learn itself. The NPU electronic device 60 can realize applications such as intelligent cognition, including image recognition, face recognition, semantic recognition, speech recognition, and text understanding.
[0141] In some embodiments, memory 62 may be an internal storage unit of electronic device 60, such as a hard disk or memory of electronic device 60; memory 62 may also be an external storage device of electronic device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on electronic device 60. Memory 62 may include both internal storage units and external storage devices of electronic device 60. Memory 62 can be used to store operating system, application programs, bootloader, data, and other programs, such as program code of computer programs. Memory 62 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). Memory 62 is used to store program code executed by electronic device 60 and data transmitted. Memory 62 can also be used to temporarily store data that has been output or will be output.
[0142] Those skilled in the art will understand that Figure 11 This is merely an example of electronic device 60 and does not constitute a limitation on electronic device 60. Electronic device 60 may include more or fewer components than shown, or combine certain components, or include different components, such as input / output devices, network access devices, etc.
[0143] Figure 12 This is a schematic diagram of the software structure of a terminal device according to an embodiment of the present invention. Taking the Android operating system as an example, in some embodiments, the Android system is divided into four layers: the application layer, the application framework layer (FWK), the system layer, and the hardware abstraction layer. The layers communicate with each other through software interfaces.
[0144] First, the application layer can include multiple application packages, which can be various application apps such as calling, camera, video, navigation, weather, instant messaging, education, etc., or application apps based on AR technology.
[0145] Second, the Application Framework Layer (FWK) provides application programming interfaces (APIs) and programming frameworks for applications within the application layer. The application framework layer can include predefined functions, such as functions for receiving events sent by the application framework layer.
[0146] The application framework layer may include a window manager, a resource manager, and a notification manager, among others.
[0147] The window manager manages the windowed applications. It can determine the screen size, the presence of a status bar, lock the screen, and capture screenshots. The content provider stores and retrieves data, making it accessible to applications. This data can include videos, images, audio, made and received phone calls, browsing history and bookmarks, and phonebook entries.
[0148] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0149] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0150] In addition, the application framework layer may include a view system, which includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build the application. The display interface can consist of one or more views; for example, the display interface of a text notification icon may include a view for displaying text and a view for displaying images.
[0151] Third, the system layer can include multiple functional modules, such as sensor service modules, physical state recognition modules, 3D graphics processing libraries (e.g., OpenGLES), and so on.
[0152] The sensor service module monitors sensor data uploaded by various sensors at the hardware layer to determine the physical state of the phone; the physical state recognition module analyzes and recognizes user gestures, faces, etc.; and the 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0153] In addition, the system layer may include a surface manager and a media library. The surface manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The media library supports playback and recording of various common audio and video formats, as well as still image files.
[0154] Finally, the hardware abstraction layer is the layer between hardware and software. The hardware abstraction layer can include display drivers, camera drivers, sensor drivers, etc., used to drive the relevant hardware in the hardware layer, such as displays, cameras, and sensors.
[0155] This invention also provides a computer-readable storage medium storing a computer program or instructions that, when executed, implement the steps in the large-scene metaverse virtual space overlay method described in the above embodiments.
[0156] This invention also provides a computer program product, including a computer program or instructions, which, when executed, implement the steps in the large-scene metaverse virtual space overlay method described in the above embodiments. For example, the computer program product may be a software installation package.
[0157] Those skilled in the art should understand that the functions of the methods, steps, or related modules / units described in the embodiments of the present invention can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product, or by a processor executing computer program instructions. The computer program product includes at least one computer program instruction, which can be composed of corresponding software modules. These software modules can be stored in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, portable hard disk, read-only optical disk (CD-ROM), or any other form of storage medium known in the art. The computer program instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (e.g., SSD).
[0158] Regarding the various devices / products described in the above embodiments, the modules / units included can be software modules / units, hardware modules / units, or a combination of both. For example, for devices / products applied to or integrated into a chip, all of its modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs running on a processor integrated within the chip, while the remaining modules / units can be implemented using hardware methods such as circuits. Similarly, for devices / products applied to or integrated into a terminal, all of its modules / units can be implemented using hardware methods such as circuits, or at least some modules / units can be implemented using software programs running on a processor integrated within the terminal, while the remaining modules / units can be implemented using hardware methods such as circuits.
[0159] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A method for overlaying virtual spaces in a large-scale metaverse, characterized in that, include: The system acquires real-time visual images of the current large-scale scene environment, sensor data from the terminal device, and GPS data from the terminal device. The sensor data from the terminal device includes at least IMU data and geomagnetic data. Visual features are extracted from the visual images of the current large scene environment; the IMU data is processed to obtain the pose data of the terminal device; the geographical location planar coordinates of the terminal device are obtained based on GPS data and geomagnetic data; A 3D AR digital space that overlaps with the current environment is constructed based on visual images of the current environment and pose data of the terminal device. Based on the GPS data or geographic location plane coordinates of the terminal device, a 1:1 metaverse virtual space and a virtual 3D space model of the current large scene are obtained from the metaverse system for overlay. Visual positioning is performed based on extracted visual features, geographic location planar coordinates, and a virtual 3D spatial model of the current large scene to obtain the 3D coordinates of the real space; The virtual spatial position of the terminal device in the metaverse virtual space is determined based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and The metaverse virtual space is integrated into the AR digital space according to the virtual space location; The steps for visual localization based on extracted visual features, geographic location planar coordinates, and a virtual 3D spatial model of the current large scene include: Based on the virtual 3D space model of the current large scene, scene matching is performed on the extracted visual features to obtain a predicted location array. The predicted location array includes multiple predicted 3D coordinates of predicted scenes that are respectively matched with the visual features. One or more first predicted 3D coordinates are selected from the predicted location array based on the geographical location plane coordinates. The difference between the altitude of the terminal device and the height coordinate in each first predicted 3D coordinate is calculated, and the first predicted 3D coordinate with the smallest difference is taken as the real space 3D coordinate. Alternatively, the matching range in the virtual 3D space model of the current large scene is determined based on the altitude of the terminal device; scene matching is performed on the extracted visual features within the matching range to obtain an estimated position array, the estimated position array including multiple estimated 3D coordinates corresponding to an estimated scene that matches the visual features; and an estimated 3D coordinate is selected from the estimated position array based on the geographic location plane coordinates as the real space 3D coordinate.
2. The method according to claim 1, characterized in that, The steps to obtain the geographic location plane coordinates of the terminal device based on GPS data and geomagnetic data include: Determine the motion status of the terminal device based on IMU data; In response to the terminal device being in a stationary state, the geographical location plane coordinates of the terminal device are determined based on geomagnetic data; In response to the terminal device being in motion, the geographical location planar coordinates of the terminal device are determined based on GPS data; and Geomagnetic data is calibrated based on GPS data when the terminal device is in motion.
3. The method according to claim 1, characterized in that, The step of filtering the first estimated three-dimensional coordinates from the estimated position array includes: Calculate the difference between the planar coordinates of the geographic location and the planar coordinates of each estimated 3D coordinate in the estimated location array; and One or more estimated 3D coordinates with the smallest difference are selected as the first estimated 3D coordinates.
4. The method according to claim 1, characterized in that, After calculating the difference between the altitude of the terminal device and the altitude coordinate in each first estimated three-dimensional coordinate, the method further includes: comparing the difference with a threshold. In response to the difference being greater than or equal to a threshold, acquire the current multiple visual images; Ground plane recognition is performed based on the multiple visual images, and a visual depth map is obtained; The height of the terminal device relative to the ground is calculated based on the visual depth map; and The height coordinates in the first estimated three-dimensional coordinates are replaced by the height coordinates, and the first estimated three-dimensional coordinates with the replaced height coordinates are used as the actual three-dimensional coordinates in space.
5. The method according to claim 1, characterized in that, The steps for visual localization based on extracted visual features, geographic location planar coordinates, and a virtual 3D spatial model of the current large scene include: Air pressure data is obtained from the sensor data, and the altitude of the device is calculated based on the air pressure data.
6. The method according to claim 1, characterized in that, The steps for visual localization based on extracted visual features, geographic location planar coordinates, and the current large-scale virtual 3D spatial model include: Ground plane recognition is performed based on multiple current visual images, and a visual depth map is obtained; The height of the terminal device relative to the ground is calculated based on the visual depth map; The matching range in the current large-scale virtual 3D space model is determined based on the height of the terminal device relative to the ground.
7. The method according to claim 1, characterized in that, Further includes: The cumulative error of position and angle in pose data is corrected based on one or more of the following: real-world three-dimensional coordinates, GPS data, and geomagnetic data.
8. A method for overlaying virtual spaces in a large-scale metaverse, characterized in that, The method is applied to a terminal device and includes: The system acquires real-time visual images of the current large-scale scene environment, sensor data from the terminal device, and GPS data from the terminal device. The sensor data from the terminal device includes at least IMU data and geomagnetic data. Visual features are extracted from the visual images of the current large scene environment; the IMU data is processed to obtain the pose data of the terminal device; the geographical location planar coordinates of the terminal device are obtained based on GPS data and geomagnetic data; A 3D AR digital space that overlaps with the current environment is constructed based on visual images of the current environment and pose data of the terminal device. Send the terminal device's GPS data or geographic location plane coordinates and requests to the server to obtain a 1:1 metaverse virtual space for overlay; The system sends a visual image of the current large-scale scene environment or visual feature information extracted from the visual image of the current large-scale scene environment, as well as the geographical location plane coordinates of the terminal device to the server, and receives one or more first estimated three-dimensional coordinates from the server. The one or more first estimated three-dimensional coordinates are selected from the estimated location array based on the geographical location plane coordinates. The estimated location array includes multiple estimated three-dimensional coordinates that each correspond to an estimated scene that matches the visual features. The estimated location array is obtained by scene matching of the extracted visual features based on the virtual three-dimensional space model of the current large-scale scene. When multiple first estimated three-dimensional coordinates are obtained, air pressure data is obtained from the sensor data, and the altitude of the terminal device is calculated based on the air pressure data. Calculate the difference between the altitude and the height coordinate in each first estimated three-dimensional coordinate, and take the first estimated three-dimensional coordinate with the smallest difference as the real space three-dimensional coordinate; The virtual spatial position of the terminal device in the metaverse virtual space is determined based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and The metaverse virtual space is integrated into the AR digital space according to the virtual space location.
9. The method according to claim 8, characterized in that, After calculating the difference between the altitude and the altitude coordinate in each first estimated three-dimensional coordinate system, the method further includes: The difference is compared with a threshold. In response to the difference being greater than or equal to a threshold, acquire the current multiple visual images; Ground plane recognition is performed based on the multiple visual images, and a visual depth map is obtained; The height of the terminal device relative to the ground is calculated based on the visual depth map; and The height coordinates in the first estimated three-dimensional coordinates are replaced by the height coordinates, and the first estimated three-dimensional coordinates with the replaced height coordinates are used as the actual three-dimensional coordinates in space.
10. The method according to claim 8, characterized in that, Further includes: The cumulative error of position and angle in pose data is corrected based on one or more of the following: real-world three-dimensional coordinates, GPS data, and geomagnetic data.
11. A method for overlaying virtual spaces in a large-scale metaverse, characterized in that, The method is applied to a terminal device and includes: The system acquires real-time visual images of the current large-scale scene environment, sensor data from the terminal device, and GPS data from the terminal device. The sensor data from the terminal device includes IMU data, geomagnetic data, and air pressure data. The IMU data is processed to obtain the pose data of the terminal device; the geographical location plane coordinates of the terminal device are obtained based on GPS data and geomagnetic data; the altitude of the terminal device is calculated based on the air pressure data. A 3D AR digital space that overlaps with the current environment is constructed based on visual images of the current environment and pose data of the terminal device. Send the terminal device's GPS data or geographic location plane coordinates and requests to the server to obtain a 1:1 metaverse virtual space for overlay; The system sends a visual image of the current large-scale scene environment, the geographical location planar coordinates of the terminal device, and the altitude of the terminal device to the server, and receives the real-world three-dimensional coordinates of the terminal device in the current large-scale scene from the server. The real-world three-dimensional coordinates are a predicted three-dimensional coordinate selected from a predicted position array based on the geographical location planar coordinates. The predicted position array is obtained by scene matching of extracted visual features within a matching range. The predicted position array includes multiple predicted three-dimensional coordinates, each corresponding to a predicted scene that matches a visual feature. The matching range is determined based on the altitude within the virtual three-dimensional space model of the current large-scale scene. The visual features are extracted from the visual image of the current large-scale scene environment. The virtual spatial position of the terminal device in the metaverse virtual space is determined based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and The metaverse virtual space is integrated into the AR digital space according to the virtual space location.
12. A method for overlaying virtual spaces in a large-scale metaverse, characterized in that, The method is applied to a server and includes: When responding to a location request sent by a terminal device, the device obtains its geographic location plane coordinates and visual feature information of the current large scene environment from the location request. Based on the virtual 3D space model of the current large scene, the extracted visual features are matched to obtain a predicted position array. The predicted position array includes multiple predicted 3D coordinates that each correspond to a predicted scene that matches the visual features. Based on the geographic location planar coordinates, one or more first estimated three-dimensional coordinates are selected from the estimated location array and sent to the terminal device. This allows the terminal device to obtain air pressure data from its sensor data when it receives multiple first estimated three-dimensional coordinates, calculate the terminal device's altitude based on the air pressure data, calculate the difference between the altitude and the height coordinate in each first estimated three-dimensional coordinate, and use the first estimated three-dimensional coordinate with the smallest difference as the real-world three-dimensional coordinate. Based on the real-world three-dimensional coordinates and the terminal device's current pose data, the virtual spatial position of the terminal device in the metaverse virtual space is determined; and the metaverse virtual space is integrated into the AR digital space according to the virtual spatial position. In response to receiving a metaverse virtual space request from a terminal device, the system returns the corresponding metaverse virtual space data to the terminal device.
13. The method according to claim 12, characterized in that, Further includes: The next step before scene matching includes: Ground plane recognition is performed based on multiple current visual images, and a visual depth map is obtained; The height of the terminal device relative to the ground is calculated based on the visual depth map; The matching range in the virtual 3D space model of the current large scene is determined based on the height of the terminal device relative to the ground; and Scene matching is performed on the extracted visual features within the matching range to obtain an array of estimated locations.
14. A method for overlaying virtual spaces in a large-scale metaverse, characterized in that, The method is applied to a server and includes: When responding to a location request sent by a terminal device, the device obtains the terminal device's geographic location plane coordinates, the visual image of the current large scene environment, and the terminal device's altitude from the location request. Visual features are extracted from the visual images of the current large-scale scene environment; The matching range in the virtual 3D space model of the current large scene is determined based on the altitude. Scene matching is performed on the extracted visual features within the matching range to obtain a predicted location array, which includes multiple predicted three-dimensional coordinates corresponding to a predicted scene that matches the visual features. Based on the geographic location planar coordinates, an estimated 3D coordinate is selected from the estimated location array as the real-world 3D coordinate, and sent to the terminal device. This allows the terminal device to determine its virtual spatial position in the metaverse virtual space based on the real-world 3D coordinates and its current pose data. The metaverse virtual space is then integrated into the AR digital space according to the virtual spatial position. In response to receiving a metaverse virtual space request from a terminal device, the system returns the corresponding metaverse virtual space data to the terminal device.
15. A large-scale metaverse virtual space overlay device, characterized in that, The device is applied to a terminal equipment and includes: The client communication module is configured to transmit data with the server. The data acquisition module is configured to acquire in real time visual images of the current large scene environment, terminal device sensor data, and terminal device GPS data, wherein the terminal device sensor data includes at least IMU data and geomagnetic data. The data preprocessing module, connected to the data acquisition module, is configured to extract visual features based on the visual image of the current large scene environment; process the IMU data to obtain the pose data of the terminal device; and obtain the geographical location planar coordinates of the terminal device based on GPS data and geomagnetic data. A digital space construction module, which is connected to the data acquisition module and the data preprocessing module, is configured to construct a three-dimensional AR digital space that overlaps with the current environment based on the visual images of the current environment and the pose data of the terminal device. The request module, which is connected to the data acquisition module, the data preprocessing module, and the client communication module, is configured to send a location request to the server, the location request including at least visual feature information and geographic location planar coordinates; and to send a metaverse virtual space request to the server. The first positioning module is connected to the client communication module and the data acquisition module. It is configured to determine one of the multiple estimated three-dimensional coordinates corresponding to a predicted scene from the positioning data returned by the server as the real space three-dimensional coordinate. The first estimated three-dimensional coordinate is selected from the estimated position array based on the geographic location plane coordinate. The estimated position array includes multiple estimated three-dimensional coordinates corresponding to a predicted scene that matches the visual features. The estimated position array is obtained by scene matching of the extracted visual features based on the virtual three-dimensional space model of the current large scene. The second positioning module, connected to the first positioning module, is configured to determine the virtual spatial position of the terminal device in the metaverse virtual space based on the three-dimensional coordinates of the real space and the current pose data of the terminal device; and The spatial fusion module, which is connected to the digital space construction module, the client communication module, and the second positioning module, is configured to fuse the metaverse virtual space returned by the server into the three-dimensional AR digital space according to the virtual space location for enhanced display.
16. A large-scale metaverse virtual space overlay device, characterized in that, The device is used in a server and includes: The server-side communication module is configured to transmit data with the terminal device, wherein it receives a location request and a metaverse virtual space request sent by the terminal device. The location request includes the geographic location plane coordinates and the visual feature information of the current large scene environment. A visual positioning module, which is connected to the server communication module, is configured to perform scene matching on visual features based on the virtual three-dimensional space model of the current large scene based on the positioning request to obtain a predicted position array. The predicted position array includes multiple predicted three-dimensional coordinates that each correspond to a predicted scene that matches the visual features. A positioning data filtering module, connected to the visual positioning module and a server-side communication module, is configured to filter one or more first estimated three-dimensional coordinates from the estimated location array based on geographic location planar coordinates. These first estimated three-dimensional coordinates are then sent to the terminal device via the server-side communication module. This allows the terminal device to obtain air pressure data from its sensor data when it receives multiple first estimated three-dimensional coordinates, calculate its altitude based on the air pressure data, calculate the difference between the altitude and the height coordinate in each first estimated three-dimensional coordinate, and use the first estimated three-dimensional coordinate with the smallest difference as the real-world three-dimensional coordinate. Based on the real-world three-dimensional coordinates and the terminal device's current pose data, the module determines the terminal device's virtual spatial position in the metaverse virtual space. Finally, the module integrates the metaverse virtual space into the AR digital space according to the virtual spatial position. The resource processing module, which is connected to the server communication module, is configured to return the corresponding metaverse virtual space data to the terminal device when a metaverse virtual space request is sent by the terminal device.
17. The apparatus according to claim 16, characterized in that, The location request includes a visual image of the current large-scale scene environment, and the device further includes: A visual computing module, which is connected to the server communication module, is configured to calculate the height of the terminal device relative to the ground using multiple current visual images; Correspondingly, when performing scene matching, the visual positioning module first determines the matching range in the virtual three-dimensional space model of the current large scene based on the height of the terminal device relative to the ground, and then performs scene matching on the extracted visual features within the matching range to obtain the estimated position array.
18. The apparatus according to claim 16, characterized in that, When the location request includes the altitude of the terminal device, the visual positioning module first determines the matching range in the virtual three-dimensional space model of the current large scene based on the height of the terminal device relative to the ground when performing scene matching. Then, it performs scene matching on the extracted visual features within the matching range to obtain the estimated position array.
19. An electronic device, characterized in that, It includes a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the large-scene metaverse virtual space overlay method applied in a terminal device as described in any one of claims 8-11; or, when the processor executes the computer program instructions, it implements the large-scene metaverse virtual space overlay method applied in a server as described in any one of claims 12-14.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the large-scene metaverse virtual space overlay method applied in a terminal device as described in any one of claims 8-11; or implement the large-scene metaverse virtual space overlay method applied in a server as described in any one of claims 12-14.
21. A computer program product, characterized in that, It includes computer program instructions that, when executed by a processor, implement the large-scene metaverse virtual space overlay method as described in any one of claims 8-11 in a terminal device; or implement the large-scene metaverse virtual space overlay method as described in any one of claims 12-14 in a server.
Citation Information
Patent Citations
Augmented reality positioning method and device based on environment visual feature point recognition technology
CN110533719A
Intelligent shopping mall service transaction platform based on meta universe
CN114493785A