Position acquisition method, device, electronic device, storage medium and program product
By constructing a voxel feature map and performing feature matching, the problem of low posture accuracy of image acquisition equipment in unknown environments is solved, and efficient and accurate posture acquisition is achieved.
Patent Information
- Application Number
- CN202310146798.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-02-08
AI Technical Summary
In existing technologies, image acquisition devices cannot accurately obtain the scale of the real world in unknown environments, resulting in low pose (position and attitude) accuracy and low efficiency.
By obtaining a three-dimensional map image of the target space area and multiple two-dimensional real-scene images, the target voxels corresponding to the two-dimensional real-scene images are selected, feature assignment is performed to construct a voxel feature map, and the real-scene image features are matched with the voxel feature map to determine the position and posture of the image acquisition device in the target space area.
The accuracy and efficiency of the pose are improved, the dependence on three-dimensional map images is reduced, and the target pose can still be accurately obtained even when the three-dimensional map accuracy is low.
Smart Images

Figure CN116503474B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a posture acquisition method, device, electronic device, storage medium and program product. Background Art
[0002] With the rapid development of augmented reality navigation technology, efficient and accurate acquisition of the device's position in the spatial area affects navigation accuracy.
[0003] In related technologies, an image acquisition device typically starts from an unknown location in an unknown environment and repeatedly observes map features during movement to determine its own position and posture. Because the image acquisition device loses the depth of the scene during capture, it cannot resolve the real-world scale and can only recover relative coordinates, which is not suitable for real-world positioning. This results in low accuracy in the determined posture (position and posture). Furthermore, the time-consuming repeated observation of map features leads to low posture acquisition efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a posture acquisition method, device, electronic device, computer-readable storage medium and computer program product, which can effectively improve the efficiency and accuracy of posture acquisition.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present invention provides a method for obtaining a posture, including:
[0007] Acquire a three-dimensional map image of a target spatial area and a plurality of two-dimensional real scene images of the target spatial area;
[0008] Acquiring the poses corresponding to the two-dimensional real scene images, and selecting target voxels corresponding to the two-dimensional real scene images from a plurality of voxels of the three-dimensional map image based on the poses;
[0009] Acquiring image features of each of the two-dimensional real scene images, and assigning features to each of the target voxels based on the image features to obtain a voxel feature map of the target space area;
[0010] Acquiring real-scene image features of the real-scene captured image acquired by an image acquisition device, and matching the real-scene image features with voxel features of the voxel feature map to obtain a matching result;
[0011] Based on the matching result, a target posture of the image acquisition device in the target space area is determined.
[0012] The present invention provides a posture acquisition device, comprising:
[0013] An acquisition module, configured to acquire a three-dimensional map image of a target space area and a plurality of two-dimensional real scene images of the target space area;
[0014] a selection module, configured to obtain the poses corresponding to the two-dimensional real scene images, and select, based on the poses, target voxels corresponding to the two-dimensional real scene images from a plurality of voxels of the three-dimensional map image;
[0015] an assignment module, configured to obtain image features of each of the two-dimensional real scene images, and perform feature assignment on each of the target voxels based on the image features to obtain a voxel feature map of the target space area;
[0016] a matching module, configured to obtain real-scene image features of the real-scene photograph captured by the image acquisition device, and match the real-scene image features with voxel features of the voxel feature map to obtain a matching result;
[0017] A determination module is used to determine the target posture of the image acquisition device in the target space area based on the matching result.
[0018] In some embodiments, the selection module is further used to generate at least one ray corresponding to each of the two-dimensional real scene images in the three-dimensional map image based on the posture corresponding to each of the two-dimensional real scene images; obtain at least one intersecting voxel between each of the rays and the three-dimensional map image; determine, for each of the rays, the distance between each of the intersecting voxels of the ray and the starting point of the ray, and determine the intersecting voxel with the smallest distance as the target voxel.
[0019] In some embodiments, the posture is used to indicate the target position and attitude angle of the real scene acquisition device that acquires the two-dimensional real scene image in the target space area; the above-mentioned selection module is also used to perform the following processing for each of the two-dimensional real scene images: based on the target position, determine the target map coordinates corresponding to the target position in the three-dimensional map image; based on the attitude angle and the viewing angle of the real scene acquisition device, determine the ray angle range of the ray in the three-dimensional map image; in the three-dimensional map image, use the target map coordinates as the starting point of the ray and generate at least one ray within the ray angle range.
[0020] In some embodiments, the above-mentioned selection module is also used to obtain a position-coordinate mapping file, which is used to record the mapping relationship between each position in the target space area and the corresponding map coordinates in the three-dimensional map image, and the position in the target space area corresponds one to one with the map coordinates in the three-dimensional map image; in the position-coordinate mapping file, the target mapping relationship including the target position is queried, and the map coordinates in the target mapping relationship are determined as the target map coordinates.
[0021] In some embodiments, the above-mentioned selection module is also used to obtain a reference viewing angle, and twice the size of the reference viewing angle is equal to the viewing angle size of the real scene acquisition device; subtracting the attitude angle from the reference viewing angle to obtain the minimum angle value of the ray angle range, and adding the attitude angle and the reference viewing angle to obtain the maximum angle value of the ray angle range; the angle range between the minimum angle value and the maximum angle value is determined as the ray angle range.
[0022] In some embodiments, the image features include pixel features of each pixel point in the two-dimensional real scene image; the above-mentioned assignment module is also used to obtain at least one associated pixel point associated with the corresponding target voxel from the multiple pixel points included in each of the two-dimensional real scene images; combining the pixel features of each of the associated pixel points, determining the target voxel features of the target voxels corresponding to each of the two-dimensional real scene images; in the three-dimensional map image, based on the target voxel features of each of the target voxels, feature assignment is performed on each of the target voxels to obtain a voxel feature map of the target space area.
[0023] In some embodiments, the above-mentioned assignment module is also used to perform the following processing on the target voxels corresponding to each of the two-dimensional real-scene images: when the number of the associated pixel points is one, the pixel features of the associated pixel points are determined as the target voxel features of the target voxel; when the number of the associated pixel points is multiple, the pixel features of each of the associated pixel points are weightedly summed to obtain the target voxel features of the target voxel.
[0024] In some embodiments, the voxel features include target voxel features of each target voxel in the voxel feature map; the above-mentioned matching module is also used to perform feature matching on the real-scene image features with each target voxel feature respectively to obtain feature matching results corresponding to each target voxel feature; when the feature matching result indicates that the real-scene image features and the target voxel features are successfully matched, the real-scene captured image and the target voxel corresponding to the target voxel feature are combined into a captured image-voxel pair; each captured image-voxel pair obtained by the combination is determined as the matching result.
[0025] In some embodiments, the above-mentioned matching module is also used to perform the following processing for each target voxel feature: determine the feature distance between the target voxel feature and the real-scene image feature; when the feature distance is greater than or equal to the distance threshold, determine the feature matching result corresponding to the target voxel feature as a first matching result, and the first matching result is used to indicate that the real-scene image feature and the target voxel feature are successfully matched; when the feature distance is less than the distance threshold, determine the feature matching result corresponding to the target voxel feature as a second matching result, and the second matching result is used to indicate that the real-scene image feature and the target voxel feature fail to match.
[0026] In some embodiments, the matching result includes at least one photographic image-voxel pair, and the photographic image-voxel pair is used to indicate the mapping relationship between the real-scene photographic image and the target voxel; the above-mentioned determination module is further used to select a target photographic image-voxel pair from the at least one photographic image-voxel pair; wherein the target photographic image-voxel pair is the photographic image-voxel pair with the highest matching degree among the at least one photographic image-voxel pair, and the matching degree is used to indicate the degree of matching between the real-scene photographic image and the target voxel; based on the target photographic image-voxel pair, the target posture of the image acquisition device in the target space area is determined.
[0027] In some embodiments, the above-mentioned determination module is also used to obtain the three-dimensional position coordinates of the target voxel in the target shot image-voxel pair in the three-dimensional map image, and obtain at least one associated pixel point associated with the target voxel from the real-scene shot image; determine the two-dimensional position coordinates of each of the associated pixel points in the real-scene shot image; based on the three-dimensional position coordinates and the two-dimensional position coordinates, perform posture prediction on the image acquisition device to obtain the target posture of the image acquisition device in the target space area.
[0028] In some embodiments, the posture prediction is achieved through a posture prediction model, which includes a parameter transformation layer and a posture estimation layer; the above-mentioned determination module is also used to call the parameter transformation layer to perform parameter transformation on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain a parameter transformation matrix; call the posture estimation layer to perform posture estimation on the image acquisition device based on the parameter transformation matrix to obtain the target posture of the image acquisition device in the target space area.
[0029] In some embodiments, the above-mentioned posture acquisition device also includes: a map module, which is used to receive a posture acquisition request sent by the image acquisition device during the navigation process; in response to the posture acquisition request, the target posture is sent to the image acquisition device; wherein, the target posture is used by the image acquisition device to combine the target posture and the real-scene shooting image to render a real-scene navigation map of the target space area.
[0030] An embodiment of the present application provides an electronic device, including:
[0031] a memory for storing computer-executable instructions or computer programs;
[0032] The processor is used to implement the posture acquisition method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0033] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute and implement the posture acquisition method provided in the embodiment of the present application.
[0034] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the posture acquisition method described in the present invention.
[0035] The embodiments of the present application have the following beneficial effects:
[0036] By assigning features to each target voxel in the 3D map image, a voxel feature map is obtained, and the real-life image features are matched with the voxel features of the voxel feature map. The target pose is determined based on the matching results. This effectively reduces the reliance on the 3D map image in the target pose determination process (the reliance on the 3D map image is converted to reliance on the voxel feature map). Since the accuracy of the 3D map image often depends on the acquisition device used to collect the 3D map image, low accuracy of the acquisition device will directly result in the 3D map image failing to accurately reflect the features of the target spatial area. Therefore, the pose is obtained by determining the voxel feature map. Even if the 3D map image has low accuracy, the target pose can still be accurately obtained, thereby effectively ensuring the accuracy of the determined target pose. By assigning features to some voxels (target voxels) of the 3D map image based on image features, rather than all voxels, the resulting voxel feature map is smaller in size. In the process of matching using the voxel feature map, the matching computation amount can be effectively reduced, thereby effectively improving the efficiency of pose acquisition. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 Schematic diagram of the structure of the posture acquisition system provided in the embodiment of the present application;
[0038] Figure 2 Schematic diagram of the structure of an electronic device for obtaining posture provided in an embodiment of the present application;
[0039] Figures 3 to 7 Schematic diagram of the process of obtaining the posture provided by the embodiment of the present application;
[0040] Figure 8 This is a schematic diagram of the effect of the target space area provided by the embodiment of the present application;
[0041] Figure 9 This is a schematic diagram of the structure of a three-dimensional map acquisition device and a real scene acquisition device provided in an embodiment of the present application;
[0042] Figure 10 Schematic diagram of the principle of the posture acquisition method provided in the embodiment of the present application;
[0043] Figure 11 This is a schematic diagram of the display interface of the image acquisition device provided in an embodiment of the present application;
[0044] Figure 12 This is a schematic diagram of the interface of the image acquisition device provided in an embodiment of the present application;
[0045] Figure 13 Schematic diagram of the process of obtaining the posture provided by the embodiment of the present application;
[0046] Figure 14This is a schematic diagram of the effect of the point cloud map provided by the embodiment of the present application;
[0047] Figure 15 It is a schematic diagram of the principle of the posture acquisition method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0049] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0050] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0051] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in the embodiments of this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0052] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0053] 1) Pose: This describes the position and orientation of an object (e.g., coordinates) within a specified coordinate system. Pose describes the position and orientation of an object within a spatial coordinate system. Position indicates the coordinates of an object within the spatial coordinate system, while attitude describes the orientation of the object within the spatial coordinate system.
[0054] 2) Voxel: Voxel stands for Volume Pixel. A volume containing voxels can be represented through volume rendering or by extracting polygonal isosurfaces with a given threshold outline. As the name suggests, voxels are the smallest unit of digital data used in three-dimensional space segmentation. Voxels are used in fields such as 3D imaging, scientific data, and medical imaging. Conceptually, they are similar to pixels, the smallest unit of two-dimensional space, used in image data in two-dimensional computer graphics.
[0055] 3) Target spatial area: refers to a specific spatial area in the real world, for example, the spatial area corresponding to a shopping mall in a city, the spatial area corresponding to a scenic spot, etc.
[0056] 4) Augmented Reality (AR): This technology cleverly integrates virtual information with the real world. It utilizes a wide range of technologies, including multimedia, 3D modeling, real-time registration, intelligent interaction, and sensing. It simulates computer-generated virtual information, such as text, images, 3D models, music, and video, and then applies it to the real world. The two types of information complement each other, thereby "enhancing" the real world. Also known as augmented reality, AR is a relatively new technology that integrates real-world and virtual-world information. It simulates physical information, previously difficult to experience within the spatial confines of the real world, using computer science and other scientific technologies. This overlay allows virtual information to be effectively applied in the real world, and this process is perceived by human senses, creating a sensory experience beyond reality. The real environment and virtual objects overlap, allowing them to coexist in the same image and space.
[0057] 5) Real Map: A real map is a map that displays a real street view. Based on the electronic map basemap architecture, secondary development is completed according to customer needs to display and navigate the locations of customer branches and outlets. This not only allows for rapid location positioning and bus route queries, but also integrates a vast amount of multimedia information into the map points, allowing customers to view the information integrated into the punctuation points while retrieving the specific geographic locations of branches and outlets. By viewing the 360-degree real scenes at the punctuation points, customers can gain an immersive understanding of the internal and external environment without leaving their homes. Real Map innovatively combines three-dimensional real scenes with electronic maps to provide real-view map search services for the general public.
[0058] 6) The PnP (Perspective-n-Point) problem is a method for solving the motion of 3D-to-2D point pairs, aiming to determine the pose of the camera coordinate system relative to the world coordinate system. It describes the process of calculating the corresponding perspective projection relationship, given the coordinates of n 3D points (relative to the world coordinate system) and their pixel coordinates, to obtain the camera pose (also called camera pose) or object pose (also called object pose). The solvepnp function provided by OpenCV can be used to solve the PnP problem. This function can be used to measure the camera pose or object pose, as well as for spatial positioning.
[0059] 7) Camera Calibration: In image measurement and machine vision applications, to determine the relationship between the 3D geometric position of a point on a spatial object's surface and its corresponding point in the image, a geometric model of the camera's imaging, known as the camera model, must be established. These camera model parameters are known as the camera parameters. The process of determining these parameters is called camera calibration. Camera parameters include intrinsic and extrinsic parameters. Intrinsic parameters are determined by the camera itself and do not change due to external factors.
[0060] 8) Camera Model: This describes the process of mapping coordinate points in a 3D world coordinate system to a 2D image plane. It serves as the link between points in 3D space and points on the 2D plane. Camera models include at least the pinhole camera model and the fisheye camera model. For example, the pinhole camera model contains four coordinate systems: the 3D world coordinate system, the 3D camera coordinate system, the 2D image physical coordinate system, and the 2D image pixel coordinate system.
[0061] 9) Camera coordinate system: It is a three-dimensional rectangular coordinate system, also known as the three-dimensional camera coordinate system. The optical center of the camera is the origin O of the coordinate system. The x-axis and y-axis are parallel to the two perpendicular sides of the camera imaging plane, that is, the x and y directions parallel to the image are the x-axis and y-axis, and the optical axis of the camera is the z-axis (or the z-axis is parallel to the optical axis). x, y, and z are perpendicular to each other, and the unit is the length unit.
[0062] 10) 3D Camera Coordinate System: Points in the 3D camera coordinate system are transformed into corresponding points in the 2D image coordinate system through perspective projection. Perspective projection uses a central projection method to project an object onto a projection surface, resulting in a single-sided projection that closely resembles the visual effect (similar to a shadow puppet). Perspective projection conforms to human psychology, where objects closer to the viewpoint appear larger, objects farther away appear smaller, and parallel lines that are not parallel to the imaging plane intersect at a vanishing point.
[0063] 11) Image physical coordinate system: The image physical coordinate system (also called the two-dimensional image coordinate system) takes the intersection of the camera's optical axis and the physical imaging plane (also called the image plane) as the coordinate origin O'. The x' axis and y' axis are parallel to the two perpendicular sides of the image plane, respectively. The unit is the unit of length.
[0064] 12) Image Pixel Coordinate System: The image pixel coordinate system (also known as the pixel coordinate system) uses the image vertex as the coordinate origin, Opixel, with the u and v directions parallel to the x' and y' axes, and is measured in pixels. In practical applications, images captured by a camera are first converted into standard electrical signals and then into digital images through analog-to-digital conversion. Each image is stored as an M×N array, with the value of each element in the M rows and N columns representing the grayscale of the image point. Each such element is called a pixel, and the pixel coordinate system is an image coordinate system based on pixels.
[0065] During the implementation of the embodiments of this application, the applicant discovered that the related technology has the following problems:
[0066] In related technologies, an image capture device typically starts from an unknown location in an unknown environment and uses repeated observations of map features during movement to determine its own position and posture. Because the image capture device loses the depth of the scene during its capture, it cannot determine the real-world scale and can only recover relative coordinates, which is not suitable for real-world positioning, resulting in inaccurate determined pose (position and posture).
[0067] The embodiments of the present application provide a posture acquisition method, device, electronic device, computer-readable storage medium and computer program product, which can effectively improve the efficiency and accuracy of posture acquisition. The following describes an exemplary application of the posture acquisition system provided by the embodiments of the present application.
[0068] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the posture acquisition system 100 provided in an embodiment of the present application. The terminal (terminal 400 is shown as an example) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0069] The terminal 400 is used for the user to use the client 410 and display the real scene captured image on the graphical interface 410-1 (graphic interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0070] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a car terminal, etc., but is not limited to this. The electronic device provided in the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application.
[0071] In some embodiments, the server 200 obtains a three-dimensional map image and multiple two-dimensional real-scene images of the target space area, obtains the posture corresponding to each two-dimensional real-scene image, and based on the posture, selects the target voxel corresponding to the two-dimensional real-scene image, obtains the image features of each two-dimensional real-scene image, and assigns features to the target voxel based on the image features to obtain a voxel feature map; obtains the real-scene image features of the real-scene geophotograph collected by the image acquisition device, matches the real-scene image features with the voxel feature map to obtain a matching result, determines the target posture of the image acquisition device in the target space area based on the matching result, and sends the target posture to the terminal 400 corresponding to the image acquisition device.
[0072] In other embodiments, the terminal 400 corresponding to the image acquisition device obtains a three-dimensional map image and multiple two-dimensional real-scene images of the target space area, obtains the posture corresponding to each two-dimensional real-scene image, and based on the posture, selects the target voxel corresponding to the two-dimensional real-scene image, obtains the image features of each two-dimensional real-scene image, and assigns features to the target voxels based on the image features to obtain a voxel feature map; obtains the real-scene image features of the real-scene geophotograph collected by the image acquisition device, matches the real-scene image features with the voxel feature map to obtain a matching result, and determines the target posture of the image acquisition device in the target space area based on the matching result.
[0073] In other embodiments, the embodiments of the present application can be implemented with the help of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or local area network to realize data calculation, storage, processing, and sharing.
[0074] Cloud technology is a general term for network, information, integration, management platform, and application technologies used in the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a key support. The backend services of technical network systems require a large amount of computing and storage resources.
[0075] See also Figure 2 , Figure 2 is a structural diagram of an electronic device 500 for obtaining posture provided in an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1 The server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 410, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0076] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0077] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0078] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0079] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0080] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0081] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0082] In some embodiments, the posture acquisition device provided in the embodiments of the present application can be implemented in software. Figure 2 A posture acquisition device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 4551, a selection module 4552, an assignment module 4553, a matching module 4554, and a determination module 4555. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0083] In other embodiments, the posture acquisition device provided in the embodiments of the present application can be implemented in hardware. As an example, the posture acquisition device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the posture acquisition method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.
[0084] In some embodiments, the terminal or server can implement the posture acquisition method provided in the embodiments of the present application by running a computer program or computer executable instructions. For example, the computer program can be a native program in the operating system (for example, a dedicated posture acquisition program) or a software module, for example, a posture acquisition module that can be embedded in any program (such as an instant messaging client, a photo album program, an electronic map client, a navigation client); for example, it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be any form of application, module or plug-in.
[0085] The posture acquisition method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.
[0086] See also Figure 3 , Figure 3 This is a flow chart of the posture acquisition method provided in the embodiment of the present application, which will be combined with Figure 3 Steps 101 to 107 are shown for illustration. The posture acquisition method provided in the embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will be illustrated by taking the server alone as an example.
[0087] In step 101, a three-dimensional map image of a target space area and a plurality of two-dimensional real scene images of the target space area are obtained.
[0088] In some embodiments, the target spatial area refers to a specific spatial area in the real world, for example, the spatial area corresponding to a shopping mall in a city, the spatial area corresponding to a scenic spot, etc.
[0089] For example, see Figure 8 , Figure 8 This is a schematic diagram of the effect of the target space area provided in the embodiment of the present application. The target space area can be a real-world space area such as the "8000 Service Store", "221 Training Classroom F2 Floor", "F2 East Elevator Hall F2 Floor", "F2 West Elevator Hall F2 Floor", "F2 South Elevator Hall F2 Floor", and "233 Conference Room F2 Floor" in a shopping mall.
[0090] In some embodiments, a three-dimensional map image, also known as a three-dimensional electronic map, is a three-dimensional, abstract depiction of one or more aspects of the real world or a portion thereof, at a specific scale, based on a three-dimensional electronic map database. Online three-dimensional electronic maps not only provide users with map search capabilities such as map query and travel navigation through intuitive geographical simulations, but also integrate a range of services including lifestyle information, e-government, e-commerce, virtual communities, and travel navigation.
[0091] In some embodiments, the two-dimensional real scene image is used to reflect the real scene of the target space area at a certain angle from a two-dimensional perspective.
[0092] In some embodiments, the above step 101 can be implemented as follows: performing three-dimensional map acquisition on the target space area through a three-dimensional map acquisition device to obtain a three-dimensional map image of the target space area; performing multiple two-dimensional image acquisition on the target space area through a real scene acquisition device to obtain multiple two-dimensional real scene images of the target space area.
[0093] For example, see Figure 9 , Figure 9 This is a structural schematic diagram of the three-dimensional map acquisition device and the real scene acquisition device provided in the embodiment of the present application. The three-dimensional map acquisition device 1 is used to acquire a three-dimensional map of the target space area to obtain a three-dimensional map image of the target space area; the real scene acquisition device 2 is used to acquire multiple two-dimensional images of the target space area to obtain multiple two-dimensional real scene images of the target space area.
[0094] In this way, by obtaining a three-dimensional map image of the target space area and multiple two-dimensional real-scene images of the target space area, it is convenient to subsequently construct a voxel feature map of the target space area based on the three-dimensional map image and the two-dimensional real-scene images, providing reliable data guarantee for the subsequent determination of the target posture.
[0095] In step 102, the posture corresponding to each two-dimensional real scene image is obtained.
[0096] In some embodiments, the posture corresponding to the two-dimensional real scene image is used to indicate the target position and posture angle in the target space area when the real scene acquisition device that acquires the two-dimensional real scene image acquires the two-dimensional real scene image.
[0097] In some embodiments, the attitude angle includes pitch angle, yaw angle and roll angle. When the real scene acquisition device acquires two-dimensional real scene images, it adopts different postures. The image content of the acquired two-dimensional real scene images is different, and the postures corresponding to the two-dimensional real scene images with different image contents are different.
[0098] In step 103 , based on the position and posture, target voxels corresponding to the respective two-dimensional real scene images are selected from the plurality of voxels of the three-dimensional map image.
[0099] In some embodiments, voxel stands for volume pixel. A volume containing voxels can be represented by volume rendering or by extracting polygonal isosurfaces with a given threshold outline. As the name suggests, voxels are the smallest unit of digital data used in three-dimensional space segmentation. Voxels are used in fields such as 3D imaging, scientific data, and medical imaging. Conceptually, they are similar to pixels, the smallest unit of two-dimensional space, used in image data of two-dimensional computer graphics.
[0100] In some embodiments, the number of pixels in a two-dimensional real-scene image depends on the device parameters of the real-scene acquisition device used to capture the two-dimensional real-scene image. The number of voxels in a three-dimensional map image depends on the device parameters of the three-dimensional map acquisition device used to capture the three-dimensional map image. The number of voxels in a three-dimensional map image is greater than the number of pixels in the two-dimensional real-scene image. Each pixel in the two-dimensional real-scene image corresponds to a target voxel, and each target voxel corresponds to a pixel in at least one two-dimensional real-scene image.
[0101] In some embodiments, the content indicated by the target voxel in the target space region includes the content indicated by each pixel corresponding to the target voxel in the target space region. The content indicated by the target voxel in the target space region may be a real object actually existing in the target space region.
[0102] In some embodiments, see Figure 4 , Figure 4 is a flow chart of the posture acquisition method provided in the embodiment of the present application, Figure 3 Step 103 shown may be performed by executing Figure 4 Steps 1031 to 1033 are shown to be implemented.
[0103] In step 1031 , based on the postures corresponding to the two-dimensional real scene images, at least one ray corresponding to each of the two-dimensional real scene images is generated in the three-dimensional map image.
[0104] In some embodiments, the posture is used to indicate the target position and posture angle of a real scene acquisition device that acquires a two-dimensional real scene image in a target space area.
[0105] In some embodiments, the starting point of the ray may be a target map coordinate corresponding to the target position in the three-dimensional map image.
[0106] For example, see Figure 10 , Figure 10 Schematic diagram of the principle of the posture acquisition method provided in the embodiment of the present application. Based on the posture T corresponding to the two-dimensional real scene image, the rays R1, R2, and R3 corresponding to the two-dimensional real scene image are generated in the three-dimensional map image.
[0107] In some embodiments, the above-mentioned step 1031 can be implemented as follows: the following processing is performed for each two-dimensional real scene image respectively: based on the target position, the target map coordinates corresponding to the target position are determined in the three-dimensional map image; based on the attitude angle and the viewing angle of the real scene acquisition device, the ray angle range of the ray in the three-dimensional map image is determined; in the three-dimensional map image, with the target map coordinates as the starting point of the ray, at least one ray is generated within the ray angle range.
[0108] In some embodiments, the viewing angle of the real scene acquisition device refers to the maximum angle range that can be captured by the real scene acquisition device.
[0109] In some embodiments, the above-mentioned determination of the target map coordinates corresponding to the target position in the three-dimensional map image based on the target position can be achieved as follows: based on the target position of the real scene acquisition device that acquires the two-dimensional real scene image in the target space area, the target map coordinates corresponding to the target position are determined in the three-dimensional map image.
[0110] In some embodiments, the target position of the real scene acquisition device uses the image physical coordinate system corresponding to the target space area as the reference coordinate system, and the target map coordinates use the three-dimensional camera coordinate system corresponding to the three-dimensional map image as the reference coordinate system. The image physical coordinate system corresponding to the target position can be converted into a coordinate system to obtain the coordinates corresponding to the target position in the three-dimensional camera coordinate system (i.e., the target map coordinates).
[0111] In some embodiments, the above-mentioned determination of the target map coordinates corresponding to the target position in the three-dimensional map image based on the target position can be achieved as follows: obtaining a position-coordinate mapping file, the position-coordinate mapping file being used to record the mapping relationship between each position in the target spatial area and the corresponding map coordinates in the three-dimensional map image, and the positions in the target spatial area have a one-to-one correspondence with the map coordinates in the three-dimensional map image; in the position-coordinate mapping file, querying the target mapping relationship including the target position, and determining the map coordinates in the target mapping relationship as the target map coordinates.
[0112] In some embodiments, the above-mentioned position-coordinate mapping file is used to record the mapping relationship between each position in the target space area and the corresponding map coordinates in the three-dimensional map image. It can be understood that the reference coordinate systems corresponding to the positions in the target space area and the map coordinates in the three-dimensional map image are different. The reference coordinate system corresponding to the positions in the target space area is the image physical coordinate system of the target space area, and the reference coordinate system corresponding to the map coordinates in the three-dimensional map image is the three-dimensional camera coordinate system of the three-dimensional map image. The mapping relationship between each position in the target space area and the corresponding map coordinates in the three-dimensional map image recorded in the position-coordinate mapping file can indicate the conversion relationship between the above-mentioned image physical coordinate system and the three-dimensional camera coordinate system from the perspective of the coordinate system.
[0113] In some embodiments, the above-mentioned determination of the ray angle range of the ray in the three-dimensional map image based on the attitude angle and the viewing angle of the real scene acquisition device can be achieved in the following way: obtain a reference viewing angle, twice the size of the reference viewing angle is equal to the viewing angle size of the real scene acquisition device; subtract the attitude angle from the reference viewing angle to obtain the minimum angle value of the ray angle range, and add the attitude angle and the reference viewing angle to obtain the maximum angle value of the ray angle range; determine the angle range between the minimum angle value and the maximum angle value as the ray angle range.
[0114] In step 1032 , at least one intersecting voxel between each ray and the three-dimensional map image is obtained.
[0115] For example, see Figure 10 , obtain the intersection voxel L1 between the ray R1 and the three-dimensional map image, the intersection voxels L2 and L4 between the ray R2 and the three-dimensional map image, and the intersection voxel L3 between the ray R3 and the three-dimensional map image.
[0116] In step 1033 , for each ray, the distance between each intersecting voxel of the ray and the starting point of the ray is determined, and the intersecting voxel with the smallest distance is determined as the target voxel.
[0117] For example, see Figure 10 When the number of intersecting voxels is one, for ray R1, the intersecting voxel L1 of ray R1 is determined as the target voxel. When the number of intersecting voxels is multiple, the distances between the intersecting voxels L2 and L4 of ray R2 and the starting point T of the ray are determined, and the intersecting voxel with the smallest distance is determined as the target voxel.
[0118] In this way, based on the pose corresponding to each 2D real-world image, at least one ray is generated in the 3D map image corresponding to each 2D real-world image. The intersecting voxel with the smallest distance from the ray's starting point is determined as the target voxel, facilitating subsequent assignment of values to the target voxels and resulting in a more accurate voxel feature map. By selectively selecting some voxels in the 3D map image as target voxels, the subsequent assignment process does not require assigning values to all voxels in the 3D map image, effectively reducing the algorithm's computational complexity and improving computational efficiency, thereby effectively improving the efficiency of acquiring poses and achieving efficient pose acquisition.
[0119] In step 104 , image features of each two-dimensional real scene image are obtained.
[0120] In some embodiments, the above step 104 may be implemented in the following manner: performing image feature extraction on each two-dimensional real scene image to obtain image features of each two-dimensional real scene image.
[0121] In some embodiments, the above-mentioned image feature extraction can be implemented through an image coding model.
[0122] In step 105 , feature values are assigned to each target voxel based on the image features to obtain a voxel feature map of the target spatial region.
[0123] In some embodiments, the feature assignment refers to a process of assigning corresponding features to target voxels.
[0124] In some embodiments, the image features include pixel features of each pixel in the two-dimensional real scene image.
[0125] In some embodiments, see Figure 5 , Figure 5 is a flow chart of the posture acquisition method provided in the embodiment of the present application, Figure 3 Step 105 shown may be performed by executing Figure 5 Steps 1051 to 1053 are shown to be implemented.
[0126] In step 1051 , at least one associated pixel point associated with the corresponding target voxel is obtained from a plurality of pixel points included in each two-dimensional real scene image.
[0127] In some embodiments, the number of pixels in a two-dimensional real-scene image depends on the device parameters of the real-scene acquisition device used to capture the two-dimensional real-scene image. The number of voxels in a three-dimensional map image depends on the device parameters of the three-dimensional map acquisition device used to capture the three-dimensional map image. The number of voxels in a three-dimensional map image is greater than the number of pixels in the two-dimensional real-scene image. Each pixel in the two-dimensional real-scene image corresponds to a target voxel, and each target voxel corresponds to at least one associated pixel. When a target voxel corresponds to multiple associated pixels, the multiple associated pixels corresponding to the target voxel can be from the same two-dimensional real-scene image or from different two-dimensional real-scene images.
[0128] In some embodiments, the content indicated by the target voxel in the target space area includes: the content indicated by each associated pixel point corresponding to the target voxel in the target space area, the content indicated by the target voxel in the target space area may be the real scene content actually existing in the target space area, and the content indicated by the associated pixel point in the target space area may be the real scene content actually existing in the target space area.
[0129] As an example, the real scene content actually existing in the target space area may be an object, a person, etc. in the target space area.
[0130] In step 1052 , the target voxel features of the target voxels corresponding to each two-dimensional real scene image are determined by combining the pixel features of each associated pixel point.
[0131] In some embodiments, the two-dimensional real scene image includes multiple pixel points, the multiple pixel points include associated pixel points, and the image features of the two-dimensional real scene image include pixel features of the associated pixel points. The pixel features of the above-mentioned associated pixel points can be extracted from the image features of the two-dimensional real scene image.
[0132] In some embodiments, the target voxel feature of the target voxel may be determined by pixel features of each associated pixel point associated with the target voxel.
[0133] In some embodiments, the above-mentioned step 1052 can be implemented as follows: the following processing is performed on the target voxels corresponding to each two-dimensional real-scene image: when the number of associated pixel points is one, the pixel features of the associated pixel points are determined as the target voxel features of the target voxel; when the number of associated pixel points is multiple, the pixel features of each associated pixel point are weighted and summed to obtain the target voxel features of the target voxel.
[0134] In some embodiments, the above-mentioned weighted summation of the pixel features of each associated pixel point to obtain the target voxel feature of the target voxel can be determined as follows: obtain the weight of the pixel feature of each associated pixel point, and perform weighted summation of the pixel features of each associated pixel point according to the corresponding weight to obtain the target voxel feature of the target voxel.
[0135] In some embodiments, the weights of the pixel features of the above-mentioned associated pixel points can be determined in the following manner: the following processing is performed for each associated pixel point respectively: the target distance between the ray corresponding to the associated pixel point and the center point of the corresponding target voxel is obtained, and based on the target distance, the weights of the pixel features of the associated pixel point are determined, wherein the above-mentioned target distance is inversely proportional to the value of the weight, the larger the target distance, the smaller the weight, and the smaller the target distance, the greater the weight.
[0136] In step 1053 , in the three-dimensional map image, based on the target voxel features of each target voxel, feature assignment is performed on each target voxel to obtain a voxel feature map of the target space area.
[0137] In some embodiments, the above step 1053 can be implemented as follows: in the three-dimensional map image, the following processing is performed for each target voxel: the initial voxel feature of the target voxel is obtained, and the initial voxel feature of the target voxel is replaced with the target voxel feature of the target voxel to obtain a voxel feature map of the target space area.
[0138] In some embodiments, the initial voxel feature of the target voxel may be determined during the image acquisition phase of the three-dimensional map image. The initial voxel feature of the target voxel may be missing, that is, the initial voxel feature of the target voxel may be empty.
[0139] In this way, by assigning features to each target voxel based on image features, a voxel feature map of the target spatial area is obtained. In the process of determining the voxel feature map, some voxels of the three-dimensional map image (i.e., target voxels) are assigned values, and there is no need to assign values to all voxels in the three-dimensional map image, thereby effectively reducing the amount of calculation for feature assignment and eliminating the need to assign values to all voxels in the three-dimensional map image. This effectively reduces the amount of algorithm calculation and improves computational efficiency, thereby effectively improving the efficiency of obtaining posture and achieving efficient posture acquisition.
[0140] In step 106, the real scene image features of the real scene photographed image captured by the image capture device are obtained.
[0141] In some embodiments, the image acquisition device may be a mobile terminal with an image acquisition function (eg, a mobile phone with a camera, etc.) or other dedicated image acquisition devices with display and communication functions.
[0142] In some embodiments, the above-mentioned step 106 can capture the real-scene shooting image of the target space area through the image acquisition device, and send the real-scene shooting image to the server, so that the server obtains the real-scene shooting image captured by the image acquisition device, and extracts image features of the real-scene shooting image to obtain the real-scene image features of the real-scene shooting image.
[0143] For example, see Figure 11 , Figure 11 Schematic diagram of the display interface of the image acquisition device provided in the embodiment of the present application. Figure 11 The display interface of the image acquisition device shown, the real-scene captured image displayed in the display interface of the image acquisition device, the image acquisition device sends the real-scene captured image to the server, so that the server obtains the real-scene captured image captured by the image acquisition device, and extracts image features of the real-scene captured image to obtain the real-scene image features of the real-scene captured image, so as to subsequently determine the target posture of the image acquisition device and send the target posture to the image acquisition device.
[0144] In step 107 , the real scene image features are matched with the voxel features of the voxel feature map to obtain a matching result.
[0145] In some embodiments, the above matching refers to a process of matching the real scene image features with the target voxel features of each target voxel in the voxel feature map.
[0146] In some embodiments, the matching result is used to indicate a mapping relationship between the real scene captured image and the target voxel.
[0147] In some embodiments, the voxel features include target voxel features for each target voxel in the voxel feature map, see Figure 6 , Figure 6 is a flow chart of the posture acquisition method provided in the embodiment of the present application, Figure 3 Step 107 shown may be performed by executing Figure 6 Steps 1071 to 1073 are shown to be implemented.
[0148] In step 1071 , feature matching is performed on the real scene image features and the target voxel features respectively to obtain feature matching results corresponding to the target voxel features.
[0149] In some embodiments, the feature matching can be achieved by determining the feature distance between the real scene image feature and the target voxel feature. The feature matching result indicates whether the real scene image feature matches the target voxel feature successfully.
[0150] In some embodiments, the above-mentioned step 1071 can be implemented by performing the following processing for each target voxel feature: determining the feature distance between the target voxel feature and the real-scene image feature; when the feature distance is greater than or equal to the distance threshold, determining the feature matching result corresponding to the target voxel feature as the first matching result, and the first matching result is used to indicate that the real-scene image feature and the target voxel feature are successfully matched; when the feature distance is less than the distance threshold, determining the feature matching result corresponding to the target voxel feature as the second matching result, and the second matching result is used to indicate that the real-scene image feature and the target voxel feature fail to match.
[0151] In some embodiments, the characteristic distance between the target voxel feature and the real-scene image feature may refer to the Manhattan distance, Euclidean distance, or Bischoff distance between the target voxel feature and the real-scene image feature. The specific expression of the characteristic distance does not constitute a limitation on the embodiments of the present application.
[0152] In some embodiments, the above-mentioned feature distance is used to indicate the similarity between the target voxel feature and the real-scene image feature. The similarity between the target voxel feature and the real-scene image feature is proportional to the feature distance between the target voxel feature and the real-scene image feature. That is, the greater the similarity between the target voxel feature and the real-scene image feature, the greater the feature distance between the target voxel feature and the real-scene image feature, and the smaller the similarity between the target voxel feature and the real-scene image feature, the smaller the feature distance between the target voxel feature and the real-scene image feature.
[0153] In step 1072 , when the feature matching result indicates that the real scene image feature matches the target voxel feature successfully, the real scene captured image and the target voxel corresponding to the target voxel feature are combined into a captured image-voxel pair.
[0154] In some embodiments, a photographic image-voxel pair is used to indicate a mapping relationship between the real-scene photographic image and the target voxel. By determining the photographic image-voxel pair, it is convenient to subsequently determine the target position of the image acquisition device in the target spatial region based on the target voxel in the photographic image-voxel pair and the real-scene photographic image.
[0155] In step 1073 , each combined captured image-voxel pair is determined as a matching result.
[0156] In some embodiments, the matching result includes at least one photographic image-voxel pair, and the matching result is used to indicate a mapping relationship between the real scene photographic image and the target voxel.
[0157] In this way, by matching the real-scene image features with the target voxel features respectively, the feature matching results corresponding to the target voxel features are obtained. Therefore, in the feature matching process, there is no need to match all the voxels in the three-dimensional map image, which effectively reduces the computational complexity of the feature matching process and improves the computational efficiency, thereby effectively improving the efficiency of obtaining the posture and achieving efficient posture acquisition.
[0158] In step 108 , based on the matching result, the target posture of the image acquisition device in the target space area is determined.
[0159] In some embodiments, the matching result includes at least one photographic image-voxel pair, where the photographic image-voxel pair is used to indicate a mapping relationship between the real scene photographic image and the target voxel.
[0160] In some embodiments, the target posture of the image acquisition device in the target space area is used to indicate the position and posture of the image acquisition device in the target space area.
[0161] In some embodiments, see Figure 7 , Figure 7 is a flow chart of the posture acquisition method provided in the embodiment of the present application, Figure 3 Step 108 shown may be performed by executing Figure 7 Steps 1081 to 1082 are shown to be implemented.
[0162] In step 1081 , a target shot-voxel pair is selected from at least one shot-voxel pair.
[0163] In some embodiments, the above step 1081 can be implemented as follows: when the number of the photographic image-voxel pair is one, the photographic image-voxel pair is determined as the target photographic image-voxel pair; when the number of the photographic image-voxel pairs is multiple, the target photographic image-voxel pair is selected from the multiple photographic image-voxel pairs.
[0164] In some embodiments, the above-mentioned selection of the target shot-voxel pair from multiple shot-voxel pairs can be achieved by: selecting a candidate shot-voxel pair from multiple shot-voxel pairs through a random sampling consensus algorithm (RANSAC); and selecting a target shot-voxel pair from the candidate shot-voxel pairs through a voting matching algorithm (Voting-Based PoseEstimation for Robotic Assembly Using a 3D Sensor).
[0165] In this way, by selecting a target shot image-voxel pair from at least one shot image-voxel pair, when there are multiple shot image-voxel pairs, the shot image-voxel pairs are screened to obtain the target shot image-voxel pair, thereby effectively reducing the number of shot image-voxel pairs for subsequent determination of the target pose, thereby effectively reducing the time for determining the target pose, improving the computational efficiency, and thus effectively improving the efficiency of obtaining the pose.
[0166] In this way, by selecting a target shot image-voxel pair from at least one shot image-voxel pair, when there are multiple shot image-voxel pairs, the shot image-voxel pairs are screened to obtain a target shot image-voxel pair. Since the obtained target shot image-voxel pair has high validity, the accuracy of the target pose subsequently determined based on the target shot image-voxel pair is effectively improved.
[0167] In step 1082 , the target pose of the image acquisition device in the target space region is determined based on the target shot image-voxel pairs.
[0168] In some embodiments, the target photographic image-voxel pair is a photographic image-voxel pair with the highest matching degree among the at least one photographic image-voxel pair, and the matching degree is used to indicate the degree of matching between the real scene photographic image and the target voxel.
[0169] In some embodiments, the target position of the image acquisition device in the target space region can be determined by the target voxels in the target shot image-voxel pair and the real scene shot image in the target shot image-voxel pair.
[0170] In some embodiments, the above-mentioned step 1082 can be implemented as follows: for the target voxel in the target shot image-voxel pair, obtain the three-dimensional position coordinates of the target voxel in the three-dimensional map image, and obtain at least one associated pixel point associated with the target voxel from the real-scene shot image; determine the two-dimensional position coordinates of each associated pixel point in the real-scene shot image; based on the three-dimensional position coordinates and the two-dimensional position coordinates, predict the posture of the image acquisition device to obtain the target posture of the image acquisition device in the target space area.
[0171] In some embodiments, the target voxel in the three-dimensional map image is a three-dimensional position coordinate, which is the coordinate position with respect to the coordinate origin of the three-dimensional coordinate system corresponding to the three-dimensional map image; the associated pixel point in the real-scene photograph is a two-dimensional position coordinate with respect to the coordinate origin of the two-dimensional coordinate system corresponding to the real-scene photograph.
[0172] In some embodiments, the content indicated by the target voxel in the target space region includes: the content indicated by each associated pixel point corresponding to the target voxel in the target space region.
[0173] In some embodiments, at least one associated pixel associated with the target voxel may be from different real-scene images. For example, at least one associated pixel associated with the target voxel includes: associated pixel 1 and associated pixel 2, associated pixel 1 and associated pixel 2, associated pixel 1 is a real-scene image. Figure 1 The pixel point in the image, the associated pixel point 2 is the real scene shooting Figure 2 Pixels in .
[0174] In some embodiments, the above-mentioned posture prediction of the image acquisition device based on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain the target posture of the image acquisition device in the target space area can be achieved as follows: the three-dimensional position coordinates and the two-dimensional position coordinates are used as input, and the corresponding algorithm for three-dimensional posture estimation, such as the Solvepnp algorithm, is input to obtain the target posture of the image acquisition device in the target space area.
[0175] In actual implementation, the server can use the 3D and 2D position coordinates, combined with the image acquisition device's device parameters, to predict the image acquisition device's posture, thereby obtaining the target position of the image acquisition device in the target spatial region, i.e., the posture of the image acquisition device in the target spatial region. The posture here indicates the position of each device point of the image acquisition device in the target spatial region. It should be noted that the posture can be represented by the rotation matrix and translation vector of all device points from the target spatial region to the 3D map image.
[0176] In some embodiments, the above-mentioned posture prediction is achieved through a posture prediction model, which includes a parameter transformation layer and a posture estimation layer.
[0177] In some embodiments, the above-mentioned posture prediction of the image acquisition device based on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain the target posture of the image acquisition device in the target space area can be achieved in the following ways: calling the parameter transformation layer to perform parameter transformation on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain the parameter transformation matrix; calling the posture estimation layer to perform posture estimation on the image acquisition device based on the parameter transformation matrix to obtain the target posture of the image acquisition device in the target space area.
[0178] In some embodiments, the parameter transformation layer is used to perform parameter transformation (matrix transformation) on three-dimensional position coordinates and two-dimensional position coordinates to obtain a parameter transformation matrix, which is used to estimate the posture of the image acquisition device.
[0179] In some embodiments, the pose estimation layer may estimate the pose of the image acquisition device by using the least squares method to find an approximate optimal solution, thereby obtaining the target pose of the image acquisition device in the target space area.
[0180] In this way, for the target voxel in the target shot image-voxel pair, the posture of the image acquisition device is predicted through the three-dimensional position coordinates of the target voxel in the three-dimensional map image and the two-dimensional position coordinates of each associated pixel point in the real-scene shot image, and the target posture of the image acquisition device in the target space area is obtained. Since the target voxel in the target shot image-voxel pair can accurately reflect the mapping relationship between the indicated real-scene shot image and the three-dimensional map image, the target posture determined based on the target shot image-voxel pair is more accurate.
[0181] In some embodiments, Figure 3 After step 108 shown, real-scene navigation can be performed in the following manner: receiving a posture acquisition request sent by the image acquisition device during the navigation process; in response to the posture acquisition request, sending the target posture to the image acquisition device; wherein, the target posture is used by the image acquisition device to combine the target posture and the real-scene shooting image to render a real-scene navigation map of the target space area.
[0182] In some embodiments, see Figure 12 , Figure 12 1 is a schematic diagram of an interface of an image acquisition device provided in an embodiment of the present application. In the display interface of the image acquisition device, a real-life shot image 21 is displayed. The content displayed in the real-life shot image 21 is the real-life content of the target spatial area (current floor: F2, target floor: F2) acquired by the image acquisition device. In response to a navigation trigger operation for the real-life shot image in the display interface of the image acquisition device, the image acquisition device sends a pose acquisition request to the server. After receiving the pose acquisition request sent by the image acquisition device during the navigation process, the server sends the calculated target pose to the image acquisition device. After receiving the target pose, the image acquisition device combines the target pose and the real-life shot image to render a real-life navigation map 22 of the target spatial area. In the real-life navigation map 22, corresponding navigation icons 221 are rendered at corresponding positions in the real-life shot image. The navigation icons 221 are used to guide the user using the image acquisition device to navigate.
[0183] In some embodiments, Figure 3 After step 108 shown, the augmented reality image can be rendered in the following manner: receiving a posture acquisition request sent by the image acquisition device during the navigation process; in response to the posture acquisition request, sending the target posture to the image acquisition device; wherein, the target posture is used by the image acquisition device to combine the target posture and the real-scene shooting image to render an augmented reality image of the target space area.
[0184] In this way, by assigning features to each target voxel in the 3D map image, a voxel feature map is obtained, and the real-life image features are matched with the voxel features of the voxel feature map. The target pose is determined based on the matching results. This effectively reduces the reliance on the 3D map image in the process of determining the target pose (converting the reliance on the 3D map image to the reliance on the voxel feature map). Since the accuracy of the 3D map image often depends on the acquisition device used to collect the 3D map image, low accuracy of the acquisition device will directly result in the 3D map image not accurately reflecting the characteristics of the target spatial area. Therefore, the pose is obtained by determining the voxel feature map. Even if the 3D map image has low accuracy, the target pose can still be accurately obtained, thereby effectively ensuring the accuracy of the determined target pose. By assigning features to some voxels (target voxels) of the 3D map image based on image features, rather than all voxels, the resulting voxel feature map is smaller in size. In the process of matching using the voxel feature map, the matching computation amount can be effectively reduced, thereby effectively improving the efficiency of acquiring the pose.
[0185] The following describes an exemplary application of the embodiment of the present application in an actual real-scene navigation application scenario.
[0186] A real-life map is a map that shows real street scenes. Based on the electronic map base map architecture, secondary development is completed according to customer needs to realize the display and navigation of customer branches and outlets. It can not only realize the rapid positioning of geographical locations and bus route inquiries, but also integrate massive multimedia information on the map points, allowing users to view the information integrated on the punctuation points while retrieving the specific geographical locations of branches and outlets.
[0187] The embodiment of the present application proposes a posture acquisition method that uses an image to add features to a laser point cloud map, constructs a laser voxel feature map, and uses the map for visual positioning. It is mainly used for mobile camera positioning. The invention is divided into two parts: laser point cloud feature assignment and visual positioning, which specifically include the following aspects: Extract image features: Use a deep neural network model to extract the descriptor feature map of each pixel in the image. Laser point cloud feature assignment: Use the posture of the image to weight and bind the features corresponding to each pixel of the image to the corresponding point cloud voxel. Visual and map matching positioning: Match each pixel of the query image to the point cloud voxel, use geometric constraints to filter out erroneous matches, and finally calculate the position and posture of the camera.
[0188] This embodiment of the application proposes a novel algorithm for visual positioning using laser point cloud maps. This algorithm directly utilizes high-precision laser point cloud maps, avoiding the use of low-precision and algorithmically complex methods such as 3D reconstruction and SLAM to recover visual feature point maps. Furthermore, to address the problem of matching image pixels to laser point clouds, a deep neural network is used to assign feature values to the laser point clouds. Finally, a camera pose recovery algorithm based on a voting algorithm is proposed.
[0189] In some embodiments, see Figure 13 , Figure 13 The posture acquisition method provided by the embodiment of the present application can be obtained by Figure 13 The steps 301 to 308 shown are implemented, and the following will be combined with Figure 13 Steps 301 to 308 are described.
[0190] In step 301 , radar data is acquired.
[0191] In some embodiments, the above radar data can be obtained by radar acquisition equipment through radar acquisition of the target space area.
[0192] In step 302 , a point cloud map is constructed based on the acquired radar data.
[0193] In some embodiments, laser point cloud mapping can be used to obtain a point cloud map (i.e., a three-dimensional map image of the target spatial area described above) by using algorithms such as laser SLAM algorithm and FAST-LIO algorithm. After constructing the point cloud map, the laser point cloud map is saved at the same time.
[0194] For example, see Figure 14 , Figure 14 This is a schematic diagram of the effect of the point cloud map provided by the embodiment of the present application. Based on the acquired radar data, the following is constructed: Figure 14 The point cloud map shown. A point cloud map, also known as a laser point cloud, is a collection of scanned points. A LiDAR system scans the ground to obtain the 3D coordinates of ground reflection points. Each ground reflection point is distributed in 3D space as a point according to the 3D coordinates, called a scan point. The 3D map image described above includes a laser point cloud.
[0195] In step 303, first visual data is acquired.
[0196] In some embodiments, the first visual data can be acquired through a real scene acquisition device, and the first visual data is the multiple two-dimensional real scene images of the target space area described above.
[0197] In step 304 , image feature extraction is performed based on the first visual data.
[0198] In some embodiments, performing image feature extraction based on the first visual data refers to performing image feature extraction (image encoding) on the first visual data to obtain image features corresponding to the first visual data.
[0199] In some embodiments, a deep learning network S2Dnet based on a CNN (Convolutional Neural Network) can be used to extract a D-dimensional feature map F from each first visual data. i ∈R W*H*D The feature map is L2 normalized on the channel to improve versatility. The length of the feature map is consistent with the image. After obtaining the image features corresponding to the first visual data, the image features are saved.
[0200] In step 305 , a feature voxel map is determined based on the point cloud map and the image features.
[0201] In some embodiments, the above step 305 can be implemented as follows: Obtaining the image frame pose: The acquisition device is pre-calibrated, and the external parameters between the laser radar and the camera are obtained. By aligning the timestamp of the laser radar trajectory, the timestamp of the image frame and the laser point cloud collected by the laser radar, the image frame pose (x, y, z, pitch angle, yaw angle, roll angle) is obtained. Laser point cloud feature assignment: After obtaining the image frame pose, refer to Figure 15 , Figure 15 This is a schematic diagram of the principle of the pose acquisition method provided in the embodiment of the present application, where a ray is drawn from the camera optical center o to the image pixel. To ensure sparsity, each projection is separated by n pixels, where n depends on the image resolution. The next step is to bind the feature values to the point cloud. Using the octree, we retrieve all voxels that the ray passes through, find the voxel center of the first intersecting voxel, and bind the D-dimensional feature corresponding to the pixel to that voxel. When different pixels correspond to the same voxel, a weighted average is performed. The "weight" of each ray is inversely proportional to the "distance from the voxel center to the intersection of the ray and the voxel." The closer the distance, the greater the weight, and the sum of the weights is 1.
[0202] In step 306 , second visual data is acquired.
[0203] In some embodiments, the second visual data is the real scene image captured by the image capture device as described above.
[0204] In step 307, image feature extraction is performed based on the second visual data.
[0205] In some embodiments, the process of extracting image features in step 307 is similar to the process of extracting image features in step 304 .
[0206] In step 308, based on the image features and the feature voxel map of the second visual data, voting matching optimization is performed to calculate the positioning result.
[0207] In some embodiments, the above step 308 can be implemented as follows: after the voxel feature map of the scene is established offline, S2D-net is also used to extract features for the query map during positioning, and the correspondence between the 2D features of the image and the 3D points of the voxel centers in the map is established through feature value matching. After obtaining enough 2D-3D Matches, RANSAC is used to screen multiple 2D-3D Matches to obtain candidate 2D-3D Matches and PnP algorithm based on the target 2D-3DMatches to calculate the camera pose (the pose of the camera that shoots the query map relative to the origin of the map), and the position and pose of the camera's six degrees of freedom (x, y, z, pitch angle, yaw angle, and roll angle) can be obtained. 2D-3D voting matching is performed to screen the candidate 2D-3DMatches to obtain the target 2D-3D After extracting features from the query graph using S2Dnet, a voting matching algorithm is used to first obtain one-to-many matching relationships for each feature. These matching relationships are then passed through a series of geometric filters to eliminate most erroneous matches. A voting shape is then constructed. If prior location information (such as GPS location) is available, the voting shape can be further constrained. A voting shape is constructed for each 2D-3D matching relationship. Similar to Hough voting, a vote is performed on each region, and the pose represented by the region with the most votes is considered the true camera pose. The image pose is calculated, and matching and verification are performed to obtain matching inliers. The final pose (the target pose described above) is determined using the RANSAC and PnP algorithms. This target pose is then sent to the client, which renders the navigation screen based on the pose.
[0208] In this way, the pose acquisition method provided by the embodiment of the present application can obtain the target pose more quickly and accurately. Moreover, for the camera, a pixel corresponds to a real-world point rather than a 3D point, and the use of voxels is more in line with the actual situation of taking pictures. By using feature matching positioning instead of triangulated visual feature point cloud set matching, the accuracy and real-time requirements for the terminal's pose calculation are reduced, which is more suitable for mobile products with low computing power, thereby effectively improving the efficiency of pose acquisition.
[0209] In some embodiments, when binding eigenvalues, pixel points with better eigenvalue properties can be projected in a targeted manner rather than randomly selected. When voting, restrictions on the direction of gravity can be added to help quickly determine the area where the query graph is located.
[0210] In this way, by assigning features to each target voxel in the 3D map image, a voxel feature map is obtained, and the real-life image features are matched with the voxel features of the voxel feature map. The target pose is determined based on the matching results. This effectively reduces the reliance on the 3D map image in the process of determining the target pose (converting the reliance on the 3D map image to the reliance on the voxel feature map). Since the accuracy of the 3D map image often depends on the acquisition device used to collect the 3D map image, low accuracy of the acquisition device will directly result in the 3D map image not accurately reflecting the characteristics of the target spatial area. Therefore, the pose is obtained by determining the voxel feature map. Even if the 3D map image has low accuracy, the target pose can still be accurately obtained, thereby effectively ensuring the accuracy of the determined target pose. By assigning features to some voxels (target voxels) of the 3D map image based on image features, rather than all voxels, the resulting voxel feature map is smaller in size. In the process of matching using the voxel feature map, the matching computation amount can be effectively reduced, thereby effectively improving the efficiency of acquiring the pose.
[0211] It is understandable that in the embodiments of the present application, related data such as three-dimensional map images and two-dimensional real-scene images are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0212] The following continues to describe the exemplary structure of the posture acquisition device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2As shown, the software modules stored in the posture acquisition device 455 of the memory 450 may include: an acquisition module 4551, used to acquire a three-dimensional map image of the target space area, and multiple two-dimensional real-scene images of the target space area; a selection module 4552, used to acquire the posture corresponding to each two-dimensional real-scene image, and based on the posture, select the target voxel corresponding to each two-dimensional real-scene image from multiple voxels of the three-dimensional map image; an assignment module 4553, used to acquire the image features of each two-dimensional real-scene image, and based on the image features, assign features to each target voxel to obtain a voxel feature map of the target space area; a matching module 4554, used to acquire the real-scene image features of the real-scene shooting image captured by the image acquisition device, and match the real-scene image features with the voxel features of the voxel feature map to obtain a matching result; a determination module 4555, used to determine the target posture of the image acquisition device in the target space area based on the matching result.
[0213] In some embodiments, the selection module 4552 is further used to generate at least one ray corresponding to each two-dimensional real scene image in the three-dimensional map image based on the posture corresponding to each two-dimensional real scene image; obtain at least one intersecting voxel of each ray and the three-dimensional map image; for each ray, determine the distance between each intersecting voxel of the ray and the starting point of the ray, and determine the intersecting voxel with the smallest distance as the target voxel.
[0214] In some embodiments, the posture is used to indicate the target position and attitude angle of the real scene acquisition device that acquires the two-dimensional real scene image in the target space area; the above-mentioned selection module 4552 is also used to perform the following processing for each two-dimensional real scene image: based on the target position, determine the target map coordinates corresponding to the target position in the three-dimensional map image; based on the attitude angle and the viewing angle of the real scene acquisition device, determine the ray angle range of the ray in the three-dimensional map image; in the three-dimensional map image, use the target map coordinates as the starting point of the ray and generate at least one ray within the ray angle range.
[0215] In some embodiments, the above-mentioned selection module 4552 is also used to obtain a position-coordinate mapping file, which is used to record the mapping relationship between each position in the target space area and the corresponding map coordinates in the three-dimensional map image. The positions in the target space area correspond one-to-one to the map coordinates in the three-dimensional map image; in the position-coordinate mapping file, the target mapping relationship including the target position is queried, and the map coordinates in the target mapping relationship are determined as the target map coordinates.
[0216] In some embodiments, the above-mentioned selection module 4552 is also used to obtain a reference viewing angle, and twice the size of the reference viewing angle is equal to the viewing angle size of the real scene acquisition device; the attitude angle and the reference viewing angle are subtracted to obtain the minimum angle value of the ray angle range, and the attitude angle and the reference viewing angle are added to obtain the maximum angle value of the ray angle range; the angle range between the minimum angle value and the maximum angle value is determined as the ray angle range.
[0217] In some embodiments, the image features include pixel features of each pixel point in the two-dimensional real scene image; the above-mentioned assignment module 4553 is also used to obtain at least one associated pixel point associated with the corresponding target voxel from the multiple pixel points included in each two-dimensional real scene image; combining the pixel features of each associated pixel point, determining the target voxel features of the target voxel corresponding to each two-dimensional real scene image; in the three-dimensional map image, based on the target voxel features of each target voxel, feature assignment is performed on each target voxel to obtain a voxel feature map of the target space area.
[0218] In some embodiments, the above-mentioned assignment module 4553 is also used to perform the following processing on the target voxels corresponding to each two-dimensional real-scene image: when the number of associated pixel points is one, the pixel features of the associated pixel points are determined as the target voxel features of the target voxel; when the number of associated pixel points is multiple, the pixel features of each associated pixel point are weighted and summed to obtain the target voxel features of the target voxel.
[0219] In some embodiments, the voxel features include target voxel features of each target voxel in the voxel feature map; the above-mentioned matching module 4554 is also used to perform feature matching on the real-scene image features with each target voxel feature respectively to obtain feature matching results corresponding to each target voxel feature; when the feature matching result indicates that the real-scene image features and the target voxel features are successfully matched, the real-scene captured image and the target voxel corresponding to the target voxel feature are combined into a captured image-voxel pair; each captured image-voxel pair obtained by the combination is determined as a matching result.
[0220] In some embodiments, the above-mentioned matching module 4554 is also used to perform the following processing for each target voxel feature: determine the feature distance between the target voxel feature and the real-scene image feature; when the feature distance is greater than or equal to the distance threshold, determine the feature matching result corresponding to the target voxel feature as the first matching result, and the first matching result is used to indicate that the real-scene image feature and the target voxel feature are successfully matched; when the feature distance is less than the distance threshold, determine the feature matching result corresponding to the target voxel feature as the second matching result, and the second matching result is used to indicate that the real-scene image feature and the target voxel feature fail to match.
[0221] In some embodiments, the matching result includes at least one photographic image-voxel pair, which is used to indicate the mapping relationship between the real-scene photographic image and the target voxel; the above-mentioned determination module 4555 is also used to select a target photographic image-voxel pair from the at least one photographic image-voxel pair; wherein the target photographic image-voxel pair is the photographic image-voxel pair with the highest matching degree among the at least one photographic image-voxel pair, and the matching degree is used to indicate the degree of matching between the real-scene photographic image and the target voxel; based on the target photographic image-voxel pair, the target position of the image acquisition device in the target space area is determined.
[0222] In some embodiments, the above-mentioned determination module 4555 is also used to obtain the three-dimensional position coordinates of the target voxel in the target shot image-voxel pair in the three-dimensional map image, and obtain at least one associated pixel point associated with the target voxel from the real-scene shot image; determine the two-dimensional position coordinates of each associated pixel point in the real-scene shot image; based on the three-dimensional position coordinates and the two-dimensional position coordinates, predict the posture of the image acquisition device to obtain the target posture of the image acquisition device in the target space area.
[0223] In some embodiments, posture prediction is achieved through a posture prediction model, which includes a parameter transformation layer and a posture estimation layer; the above-mentioned determination module 4555 is also used to call the parameter transformation layer to perform parameter transformation on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain a parameter transformation matrix; call the posture estimation layer to perform posture estimation on the image acquisition device based on the parameter transformation matrix to obtain the target posture of the image acquisition device in the target space area.
[0224] In some embodiments, the above-mentioned posture acquisition device 455 also includes: a map module, which is used to receive a posture acquisition request sent by the image acquisition device during the navigation process; in response to the posture acquisition request, the target posture is sent to the image acquisition device; wherein, the target posture is used by the image acquisition device to combine the target posture and the real-scene shooting image to render a real-scene navigation map of the target space area.
[0225] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the posture acquisition method described in the present invention.
[0226] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the posture acquisition method provided by the embodiment of the present application, for example, Figure 3 The pose acquisition method shown.
[0227] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various electronic devices including one or any combination of the above memories.
[0228] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0229] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0230] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0231] In summary, the embodiments of the present application have the following beneficial effects:
[0232] (1) By assigning features to each target voxel in the three-dimensional map image, a voxel feature map is obtained, and the features of the real scene image and the voxel features of the voxel feature map are matched, and the target pose is determined based on the matching results. In this way, in the process of determining the target pose, the dependence on the three-dimensional map image is effectively reduced (the dependence on the three-dimensional map image is converted into dependence on the voxel feature map). Since the accuracy of the three-dimensional map image often depends on the acquisition device that acquires the three-dimensional map image, the low accuracy of the acquisition device will directly lead to the three-dimensional map image being unable to accurately reflect the features of the target spatial area. Therefore, the pose is obtained by determining the voxel feature map. Even if the accuracy of the three-dimensional map image is low, the target pose can still be accurately obtained, thereby effectively ensuring the accuracy of the determined target pose. By assigning features to some voxels (target voxels) of the three-dimensional map image based on image features, rather than all voxels, the volume of the obtained voxel feature map is smaller. In the process of matching using the voxel feature map, the matching operation amount can be effectively reduced, thereby effectively improving the efficiency of acquiring the pose.
[0233] (2) By obtaining a three-dimensional map image of the target space area and multiple two-dimensional real-scene images of the target space area, it is convenient to subsequently construct a voxel feature map of the target space area based on the three-dimensional map image and the two-dimensional real-scene image, providing reliable data guarantee for the subsequent determination of the target position.
[0234] (3) Based on the postures corresponding to the two-dimensional real scene images, at least one ray corresponding to each two-dimensional real scene image is generated in the three-dimensional map image, and the intersecting voxel with the smallest distance between each intersecting voxel of the ray and the starting point of the ray is determined as the target voxel, so as to facilitate the subsequent assignment of values to the target voxels, so that the obtained voxel feature map is more accurate. By selectively selecting some voxels in the three-dimensional map image as target voxels, the subsequent assignment process does not need to assign values to all voxels in the three-dimensional map image, thereby effectively reducing the amount of algorithm calculation and improving the operation efficiency, thereby effectively improving the efficiency of obtaining the posture and achieving efficient posture acquisition.
[0235] (4) By assigning features to each target voxel based on image features, a voxel feature map of the target spatial area is obtained. In the process of determining the voxel feature map, some voxels (i.e., target voxels) of the three-dimensional map image are assigned values, and there is no need to assign values to all voxels in the three-dimensional map image, thereby effectively reducing the amount of calculation for feature assignment and there is no need to assign values to all voxels in the three-dimensional map image, thereby effectively reducing the amount of algorithm calculation and improving the operation efficiency, thereby effectively improving the efficiency of obtaining the posture and achieving efficient posture acquisition.
[0236] (5) By matching the features of the real scene image with the features of each target voxel respectively, the feature matching results corresponding to the features of each target voxel are obtained. Therefore, in the process of feature matching, there is no need to match all the voxels of the three-dimensional map image, which effectively reduces the computational complexity of the feature matching process and improves the computational efficiency, thereby effectively improving the efficiency of obtaining the posture.
[0237] (6) By selecting a target shot-voxel pair from at least one shot-voxel pair, when there are multiple shot-voxel pairs, the shot-voxel pairs are screened to obtain a target shot-voxel pair, thereby effectively reducing the number of shot-voxel pairs for subsequently determining the target posture, thereby effectively reducing the time for determining the target posture, improving the computational efficiency, and thus effectively improving the efficiency of obtaining the posture.
[0238] (7) By selecting a target photographic image-voxel pair from at least one photographic image-voxel pair, when there are multiple photographic image-voxel pairs, the photographic image-voxel pairs are screened to obtain a target photographic image-voxel pair. Since the obtained target photographic image-voxel pair has high validity, the accuracy of the target position determined subsequently based on the target photographic image-voxel pair is effectively improved.
[0239] (8) For the target voxels in the target shot image-voxel pair, the posture of the image acquisition device is predicted through the three-dimensional position coordinates of the target voxels in the three-dimensional map image and the two-dimensional position coordinates of each associated pixel point in the real scene shot image, and the target posture of the image acquisition device in the target space area is obtained. Since the target voxels in the target shot image-voxel pair can accurately reflect the mapping relationship between the indicated real scene shot image and the three-dimensional map image, the target posture determined based on the target shot image-voxel pair is more accurate.
[0240] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A posture acquisition method, characterized in that: The method comprises: Acquire a three-dimensional map image of a target spatial area and a plurality of two-dimensional real scene images of the target spatial area; Acquiring the poses corresponding to the two-dimensional real scene images, and selecting target voxels corresponding to the two-dimensional real scene images from a plurality of voxels of the three-dimensional map image based on the poses; Acquiring image features of each of the two-dimensional real scene images, and assigning features to each of the target voxels based on the image features to obtain a voxel feature map of the target space area; Acquiring real-scene image features of the real-scene captured image acquired by an image acquisition device, and matching the real-scene image features with voxel features of the voxel feature map to obtain a matching result; Based on the matching result, a target posture of the image acquisition device in the target space area is determined.
2. The method according to claim 1, characterized in that The selecting, based on the posture, target voxels corresponding to the two-dimensional real scene images from a plurality of voxels of the three-dimensional map image includes: Based on the postures corresponding to the two-dimensional real scene images, generating at least one ray corresponding to each of the two-dimensional real scene images in the three-dimensional map image; respectively acquiring at least one intersecting voxel between each of the rays and the three-dimensional map image; For each of the rays, the distance between each of the intersecting voxels of the ray and the starting point of the ray is determined, and the intersecting voxel with the smallest distance is determined as the target voxel.
3. The method according to claim 2, characterized in that The posture is used to indicate the target position and posture angle of the real scene acquisition device that acquires the two-dimensional real scene image in the target space area; The generating, in the three-dimensional map image, at least one ray corresponding to each of the two-dimensional real scene images based on the postures corresponding to each of the two-dimensional real scene images respectively includes: The following processing is performed on each of the two-dimensional real scene images: Based on the target position, determining target map coordinates corresponding to the target position in the three-dimensional map image; determining a ray angle range of the ray in the three-dimensional map image based on the attitude angle and the viewing angle of the real scene acquisition device; In the three-dimensional map image, the at least one ray is generated within the ray angle range with the target map coordinates as the starting point of the ray.
4. The method according to claim 3, characterized in that The determining, based on the target position, target map coordinates corresponding to the target position in the three-dimensional map image includes: Obtaining a position-coordinate mapping file, wherein the position-coordinate mapping file is used to record a mapping relationship between each position in the target spatial area and the corresponding map coordinates in the three-dimensional map image, wherein the positions in the target spatial area correspond one to one with the map coordinates in the three-dimensional map image; In the position-coordinate mapping file, a target mapping relationship including the target position is searched, and the map coordinates in the target mapping relationship are determined as the target map coordinates.
5. The method according to claim 3, characterized in that The determining, based on the attitude angle and the viewing angle of the real scene acquisition device, a ray angle range of the ray in the three-dimensional map image includes: Acquire a reference viewing angle, where twice the reference viewing angle is equal to the viewing angle of the real scene acquisition device; Subtracting the attitude angle from the reference viewing angle to obtain a minimum angle value of the ray angle range, and adding the attitude angle to the reference viewing angle to obtain a maximum angle value of the ray angle range; The angle range between the minimum angle value and the maximum angle value is determined as the ray angle range.
6. The method according to claim 1, characterized in that The image features include pixel features of each pixel in the two-dimensional real scene image; The step of assigning a feature value to each target voxel based on the image feature to obtain a voxel feature map of the target spatial region includes: Acquire at least one associated pixel point associated with the corresponding target voxel from a plurality of pixel points included in each of the two-dimensional real scene images; Determining target voxel features of target voxels corresponding to each of the two-dimensional real scene images by combining pixel features of each of the associated pixel points; In the three-dimensional map image, based on the target voxel features of each target voxel, feature assignment is performed on each target voxel to obtain a voxel feature map of the target space area.
7. The method according to claim 6, characterized in that Determining the target voxel features of the target voxels corresponding to the two-dimensional real scene images by combining the pixel features of the associated pixel points includes: The following processing is performed on the target voxels corresponding to each of the two-dimensional real scene images: When the number of the associated pixel points is one, determining the pixel feature of the associated pixel point as the target voxel feature of the target voxel; When there are multiple associated pixel points, pixel features of each associated pixel point are weighted and summed to obtain the target voxel feature of the target voxel.
8. The method according to claim 1, characterized in that The voxel features include target voxel features of each target voxel in the voxel feature map; matching the real scene image features with the voxel features of the voxel feature map to obtain a matching result, including: Perform feature matching on the real scene image features and the target voxel features respectively to obtain feature matching results corresponding to the target voxel features; When the feature matching result indicates that the real scene image feature matches the target voxel feature successfully, combining the real scene captured image and the target voxel corresponding to the target voxel feature into a captured image-voxel pair; Each of the combined captured image-voxel pairs is determined as the matching result.
9. The method according to claim 8, characterized in that The performing feature matching on the real scene image features and the target voxel features respectively to obtain feature matching results corresponding to the target voxel features includes: The following processing is performed for each target voxel feature: Determining a feature distance between the target voxel feature and the real scene image feature; When the feature distance is greater than or equal to a distance threshold, determining the feature matching result corresponding to the target voxel feature as a first matching result, wherein the first matching result is used to indicate that the real scene image feature is successfully matched with the target voxel feature; When the feature distance is less than the distance threshold, the feature matching result corresponding to the target voxel feature is determined as a second matching result, and the second matching result is used to indicate that the real scene image feature fails to match the target voxel feature.
10. The method according to claim 1, characterized in that The matching result includes at least one photographic image-voxel pair, wherein the photographic image-voxel pair is used to indicate a mapping relationship between the real scene photographic image and the target voxel; Determining the target posture of the image acquisition device in the target space area based on the matching result includes: Selecting a target shot-voxel pair from the at least one shot-voxel pair; The target photographic image-voxel pair is a photographic image-voxel pair with the highest matching degree among the at least one photographic image-voxel pair, and the matching degree is used to indicate the degree of matching between the real scene photographic image and the target voxel; Based on the target shot image-voxel pairs, a target posture of the image acquisition device in the target space area is determined.
11. The method according to claim 10, characterized in that The determining the target pose of the image acquisition device in the target space area based on the target shot image-voxel pair includes: For the target voxel in the target captured image-voxel pair, obtaining the three-dimensional position coordinates of the target voxel in the three-dimensional map image, and obtaining at least one associated pixel point associated with the target voxel from the real scene captured image; Determining the two-dimensional position coordinates of each of the associated pixel points in the real scene photograph; Based on the three-dimensional position coordinates and the two-dimensional position coordinates, the posture of the image acquisition device is predicted to obtain the target posture of the image acquisition device in the target space area.
12. The method according to claim 11, characterized in that The posture prediction is achieved by a posture prediction model, which includes a parameter transformation layer and a posture estimation layer; The performing posture prediction on the image acquisition device based on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain the target posture of the image acquisition device in the target space area includes: Calling the parameter transformation layer to perform parameter transformation on the three-dimensional position coordinates and the two-dimensional position coordinates to obtain a parameter transformation matrix; The pose estimation layer is called to perform pose estimation on the image acquisition device based on the parameter transformation matrix to obtain the target pose of the image acquisition device in the target space area.
13. The method according to claim 1, wherein After determining the target posture of the image acquisition device in the target space area based on the matching result, the method further includes: receiving a pose acquisition request sent by the image acquisition device during navigation; In response to the pose acquisition request, sending the target pose to the image acquisition device; The target posture is used by the image acquisition device to combine the target posture and the real-scene photograph to render a real-scene navigation map of the target space area.
14. A posture acquisition device, characterized in that: The device comprises: An acquisition module, configured to acquire a three-dimensional map image of a target space area and a plurality of two-dimensional real scene images of the target space area; a selection module, configured to obtain the poses corresponding to the two-dimensional real scene images, and select, based on the poses, target voxels corresponding to the two-dimensional real scene images from a plurality of voxels of the three-dimensional map image; an assignment module, configured to obtain image features of each of the two-dimensional real scene images, and perform feature assignment on each of the target voxels based on the image features to obtain a voxel feature map of the target space area; a matching module, configured to obtain real-scene image features of the real-scene photograph captured by the image acquisition device, and match the real-scene image features with voxel features of the voxel feature map to obtain a matching result; A determination module is used to determine the target posture of the image acquisition device in the target space area based on the matching result.
15. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the posture acquisition method according to any one of claims 1 to 13 when executing the computer executable instructions or computer program stored in the memory.
16. A computer-readable storage medium storing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the posture acquisition method according to any one of claims 1 to 13 is implemented.
17. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the posture acquisition method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Voxel map construction method and device, computer readable medium and electronic equipment
CN112927363A
Method, device and system for obtaining pose information
CN113838129A