Camera positioning method and apparatus, electronic device, and computer-readable storage medium

By establishing a target tracking connection and performing pose transformation during the camera localization process, the problems of high computational load and low efficiency in existing technologies are solved, achieving efficient and low-cost camera localization.

CN115311364BActive Publication Date: 2026-02-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110497270.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-07
Publication Date
2026-02-03
Estimated Expiration
2041-05-07

AI Technical Summary

Technical Problem

In existing technologies, the camera positioning process involves a large amount of computation, is inefficient, cannot achieve real-time positioning, and has stringent hardware requirements and high costs.

Method used

By establishing a target tracking connection between the first and second images, and utilizing the fixed invariance of the mapping region on the positioning space surface, combined with the pose transformation processing of the world coordinate system and camera coordinate system corresponding to the mapping region, the pose of the camera when acquiring the second image is determined.

Benefits of technology

This reduces the computational load during camera positioning, improves positioning efficiency and accuracy, and lowers the demand for computing resources from electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115311364B_ABST
    Figure CN115311364B_ABST
Patent Text Reader

Abstract

The application provides a camera positioning method and device, electronic equipment and computer readable storage medium. The method comprises: creating a first tracking area in a first image collected by a camera; mapping the first tracking area to a surface of a positioning space, and determining a pose of a mapping area corresponding to a world coordinate system; performing target tracking processing on a second image collected by the camera according to the first tracking area, to obtain a second tracking area in the second image; determining a pose of the mapping area corresponding to a camera coordinate system of the camera when collecting the second image according to the second tracking area; and performing pose transformation processing on the pose of the mapping area corresponding to the world coordinate system and the pose of the mapping area corresponding to the camera coordinate system, to obtain a pose of the camera corresponding to the world coordinate system when collecting the second image. Through the application, the amount of calculation in the camera positioning process can be reduced, and the efficiency of camera positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to image technology, and more particularly to a camera positioning method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] A camera is a device that uses optical imaging principles to form images. It is widely used in daily life, such as digital cameras and camera components built into mobile devices (like smartphones). Because there are usually various shooting needs when using a camera, the camera is often in motion, resulting in a non-fixed camera position.

[0003] To address this, solutions in related technologies typically involve reconstructing a point cloud containing tens of thousands of points around the camera based on images captured by the camera, and then estimating the camera's pose based on the reconstructed point cloud. However, this approach is computationally intensive, consuming significant computing resources of electronic devices, and suffers from low camera localization efficiency. Summary of the Invention

[0004] This application provides a camera positioning method, apparatus, electronic device, and computer-readable storage medium, which can reduce the amount of computation in the camera positioning process and improve the efficiency of camera positioning.

[0005] This application provides a camera positioning method, including:

[0006] A first tracking region is created in the first image acquired by the camera;

[0007] The first tracking area is mapped onto the surface of the positioning space including the camera, and the pose of the mapped area in the world coordinate system is determined.

[0008] Based on the first tracking region, target tracking processing is performed on the second image acquired by the camera to obtain the second tracking region in the second image;

[0009] Based on the second tracking region, determine the pose of the camera coordinate system corresponding to the mapping region when the camera is acquiring the second image;

[0010] Based on the pose of the mapped region corresponding to the world coordinate system and the pose of the mapped region corresponding to the camera coordinate system, pose transformation processing is performed to obtain the pose of the camera in the world coordinate system when acquiring the second image.

[0011] This application provides a camera positioning device, including:

[0012] The region creation module is used to create a first tracking region in the first image acquired by the camera;

[0013] The mapping module is used to map the first tracking area onto a surface of the positioning space including the camera, and to determine the pose of the mapped area in the world coordinate system.

[0014] A region tracking module is used to perform target tracking processing on a second image acquired by the camera based on the first tracking region, to obtain a second tracking region in the second image;

[0015] The region tracking module is further configured to determine, based on the second tracking region, the pose of the camera coordinate system corresponding to the mapping region when the camera is acquiring the second image;

[0016] The pose transformation module is used to perform pose transformation processing based on the pose of the mapped region corresponding to the world coordinate system and the pose of the mapped region corresponding to the camera coordinate system, so as to obtain the pose of the camera in the world coordinate system when acquiring the second image.

[0017] This application provides an electronic device, including:

[0018] Memory, used to store executable instructions;

[0019] The processor, when executing executable instructions stored in the memory, implements the camera positioning method provided in the embodiments of this application.

[0020] This application provides a computer-readable storage medium storing executable instructions for inducing a processor to execute and implement the camera positioning method provided in this application.

[0021] The embodiments of this application have the following beneficial effects:

[0022] Establishing a connection between the first and second images through target tracking essentially achieves target tracking from the first tracking area to the second tracking area. Simultaneously, since the mapped area of ​​the positioning space surface remains fixed during camera movement, the pose of the camera coordinate system (the camera coordinate system when acquiring the second image) corresponding to the mapped area can be determined based on the second tracking area. Furthermore, pose transformation processing is performed based on the pose of the world coordinate system corresponding to the mapped area and the pose of the camera coordinate system corresponding to the mapped area to obtain the pose of the camera in the world coordinate system when acquiring the second image. In summary, this embodiment reduces the computational load during camera positioning while also improving the accuracy of camera positioning. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the architecture of the camera positioning system provided in the embodiments of this application;

[0024] Figure 2 This is a schematic diagram of the architecture of the terminal device provided in the embodiments of this application;

[0025] Figure 3A This is a schematic flowchart of the camera positioning method provided in the embodiments of this application;

[0026] Figure 3B This is a schematic flowchart of the camera positioning method provided in the embodiments of this application;

[0027] Figure 3C This is a schematic flowchart of the camera positioning method provided in the embodiments of this application;

[0028] Figure 3D This is a schematic flowchart of the camera positioning method provided in the embodiments of this application;

[0029] Figure 4 This is a schematic diagram of the camera interface provided in an embodiment of this application;

[0030] Figure 5 This is a schematic diagram of the camera interface provided in an embodiment of this application;

[0031] Figure 6 This is a schematic diagram illustrating the changes in the virtual scene provided in the embodiments of this application;

[0032] Figure 7 This is a schematic flowchart of the camera positioning method provided in the embodiments of this application;

[0033] Figure 8 This is a schematic diagram of camera imaging provided in an embodiment of this application. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0036] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein. In the following description, the term "multiple" means at least two.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0038] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0039] 1) Camera: refers to a device that forms an image using the principle of optical imaging. In the embodiments of this application, the camera can be a standalone device, such as a data camera, or a camera component built into or external to an electronic device (such as a terminal device).

[0040] 2) Tracking region: The tracking region is a region in the image captured by the camera. The camera pose is calibrated by tracking the target in the tracking region. The shape of the tracking region is not limited in this embodiment; for example, it can be rectangular or circular.

[0041] 3) Positioning Space: This refers to the created three-dimensional space that includes the camera and is used for camera positioning. The shape of the positioning space is not limited in this embodiment; for example, it can be a sphere, cuboid, triangular pyramid, or square pyramid. In this embodiment, camera positioning can be achieved through fixed mapping areas (also called anchor blocks) on the surface of the positioning space.

[0042] 4) Pose: Includes attitude (also known as rotational attitude) and position (also known as displacement), used to describe an object in three-dimensional space.

[0043] 5) Inertial Measurement Unit (IMU): Also known as an inertial measurement unit or inertial sensor, the inertial measurement unit can collect its own attitude in the corresponding world coordinate system during operation.

[0044] 6) Artificial Intelligence (AI): Theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Computer Vision (CV) is an important branch of AI, primarily researching related theories and technologies, and attempting to establish AI systems capable of extracting information from images or multidimensional data. In the embodiments of this application, camera localization can be achieved using computer vision technology, for example, target tracking processing can be performed using computer vision technology.

[0045] 7) Virtual Scene: A virtual scene is a scene output by an electronic device that differs from the real world. Visual perception of the virtual scene can be formed with the naked eye or with device assistance. Examples include two-dimensional images output through a display screen, and three-dimensional images output through stereoscopic display technologies such as stereoscopic projection, virtual reality, and augmented reality. Furthermore, various possible hardware can be used to create auditory, tactile, olfactory, and motion perceptions, simulating the real world. A virtual scene can be a simulation of the real world, a semi-simulated / semi-fictional virtual environment, or a purely fictional virtual environment. A virtual scene can be any of a two-dimensional, 2.5-dimensional, or three-dimensional virtual scene; this application does not limit the dimension of the virtual scene.

[0046] Given the frequent movement of cameras, related technologies primarily employ two methods for camera localization: Simultaneous Localization and Mapping (SLAM) and Visual Inertial Odometry (VIO). However, these technologies suffer from at least the following problems: 1) Due to limitations in algorithmic theory, a point cloud around the camera needs to be created based on the parallax during camera translation. This requires users to spend several seconds or even longer translating the camera initially. Furthermore, the camera pose can only be estimated after the point cloud has been created, meaning the preparation phase for camera localization is lengthy and real-time localization is not possible; 2) The solutions provided by these technologies require the integration of accelerometers in the inertial measurement unit (IMU) to calculate accurate displacement. Therefore, high-precision IMUs are needed to achieve relatively accurate camera localization, resulting in stringent hardware requirements and high implementation costs; 3) In these solutions, camera localization can only be performed after reconstructing a point cloud containing tens of thousands of points, leading to a large computational load and high demands on the computing performance of electronic devices.

[0047] This application provides a camera positioning method, apparatus, electronic device, and computer-readable storage medium, which can reduce the computational load in the camera positioning process and improve the efficiency of camera positioning. The following describes exemplary applications of the electronic device provided in this application. The electronic device provided in this application can be implemented as various types of terminal devices or as a server.

[0048] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the camera positioning system 100 provided in the embodiment of this application. The terminal device 400 is connected to the server 200 through the network 300, and the server 200 is connected to the database 500. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0049] In some embodiments, taking the electronic device as a terminal device as an example, the camera positioning method provided in this application embodiment can be implemented by the terminal device. For example, the terminal device 400 creates a first tracking region in a first image acquired by the camera, wherein the camera can be built into or external to the terminal device 400, and the terminal device 400 can send an image acquisition command to the camera to enable the camera to acquire the image; in addition, the camera can also be independent of the terminal device 400, for example, the terminal device 400 can obtain the image acquired by the camera (e.g., from the distributed file system of server 200, database 500, or blockchain, etc.), but there is no direct or indirect connection between the terminal device 400 and the camera. The terminal device 400 maps the first tracking region to the surface of the positioning space including the camera, and determines the pose of the mapped region in the world coordinate system. For a second image acquired by the camera, the terminal device 400 performs target tracking processing on the second image according to the first tracking region to obtain a second tracking region in the second image. Then, the terminal device 400 determines the pose of the camera coordinate system corresponding to the second tracking area when the camera is acquiring the second image based on the second tracking area, and performs pose transformation processing based on the pose of the world coordinate system corresponding to the mapping area and the pose of the camera coordinate system corresponding to the mapping area to obtain the pose of the camera in the world coordinate system when acquiring the second image.

[0050] In some embodiments, taking the electronic device as a server as an example, the camera positioning method provided in this application embodiment can also be implemented by a server. For example, the server 200 can obtain the first image and the second image acquired by the camera from the server 200's distributed file system, database 500, or blockchain, and after a series of processing, obtain the pose of the camera in the world coordinate system corresponding to the acquisition of the second image.

[0051] In some embodiments, the camera positioning method provided in this application can also be implemented collaboratively by a terminal device and a server. For example, the terminal device 400 can send a first image and a second image acquired by the camera to the server 200. After performing a series of processing on the received first and second images, the server 200 sends the resulting pose of the camera in the world coordinate system corresponding to the second image acquisition to the terminal device 400.

[0052] In some embodiments, when any image is captured by a camera, the virtual scene can be processed based on the observation pose, and the virtual scene obtained through the observation processing can be presented; wherein, the observation pose represents the pose of the camera in the world coordinate system when capturing any image. Figure 1 As shown, terminal device 400 can determine the virtual scene to be presented based on the observed pose, and load, parse, and render display data (referring to the display data related to the virtual scene), ultimately outputting the virtual scene through graphics output hardware (such as a screen). Alternatively, server 200 can determine the virtual scene to be presented based on the observed pose and send the display data related to the virtual scene to terminal device 400. Terminal device 400 can then rely on graphics computing hardware (such as a graphics processor) to load, parse, and render the display data, and output the virtual scene through graphics output hardware to form a visual perception.

[0053] In some embodiments, various results involved in the camera positioning process (such as the first image, the second image, and the pose of the camera in the world coordinate system when acquiring the second image) can be stored in a blockchain. Because the blockchain is immutable, the accuracy of the data in the blockchain can be guaranteed. When needed, the electronic device can send a query request to the blockchain to retrieve the data stored therein. For example, the terminal device 400 can query the pose of the camera in the world coordinate system when acquiring the second image, stored in the blockchain, to determine the virtual scene to be presented based on that pose.

[0054] In some embodiments, the terminal device 400 or the server 200 can implement the camera positioning method provided in this application embodiment by running a computer program, such as... Figure 1The client 410 shown could be, for example, a camera client or a short video client. For instance, the computer program could be a native program or software module within an operating system; it could be a native application (APP), i.e., a program that needs to be installed on the operating system to run; it could also be a small program, i.e., a program that only needs to be downloaded to a browser environment to run; or it could be a small program that can be embedded into any APP, and this small program component can be controlled by the user to run or close. In short, the above-mentioned computer program can be any form of application, module, or plugin.

[0055] In some embodiments, server 200 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The cloud service may be a camera positioning service, which can be invoked by terminal device 400. Terminal device 400 may be a smartphone, tablet, laptop, desktop computer, smart TV, smartwatch, etc., but is not limited to these. Terminal device 400 and server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0056] In some embodiments, the database 500 and the server 200 can be set up independently. In some embodiments, the database 500 and the server 200 can also be integrated together, that is, the database 500 can be regarded as existing inside the server 200 and integrated with the server 200, and the server 200 can provide the data management functions of the database 500.

[0057] Taking the example of a terminal device provided in this application embodiment, it can be understood that in the case where the electronic device is a server, Figure 2 Some parts of the structure shown (such as the user interface, presentation module, and input processing module) can be omitted. See also Figure 2 , Figure 2 This is a schematic diagram of the structure of the terminal device 400 provided in the embodiments of this application. Figure 2The terminal device 400 shown includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0058] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0059] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0060] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0061] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0062] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0063] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0064] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0065] Presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with user interface 430;

[0066] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0067] In some embodiments, the camera positioning device provided in this application can be implemented in software. Figure 2 A camera positioning device 455 stored in memory 450 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a region creation module 4551, a mapping module 4552, a region tracking module 4553, and a pose transformation module 4554. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0068] The camera positioning method provided in this application will be described in conjunction with exemplary applications and implementations of the electronic devices provided in the embodiments of this application.

[0069] See Figure 3A , Figure 3A This is a flowchart illustrating the camera positioning method provided in the embodiments of this application, which will be combined with... Figure 3A The steps shown are explained.

[0070] In step 101, a first tracking region is created in the first image acquired by the camera.

[0071] Here, the first image can be any frame captured by the camera, and determining the first image also means starting camera localization. For example, the first image can be the first frame captured after the camera is started; or, for another example, the latest image captured by the camera can be used as the first image when a camera localization request is received.

[0072] For the acquired first image, a tracking region is created in the first image. For ease of distinction, the tracking region created here is named the first tracking region. This embodiment of the application does not limit the shape of the tracking region; for example, it can be a rectangle, a circle, or an irregular shape.

[0073] In some embodiments, the creation of a first tracking region in a first image acquired by a camera can be achieved by weighting the size of the first image to obtain a weighted size; and creating a first tracking region in the first image based on the weighted size, with the optical center of the first image as the center.

[0074] For example, a weighted size can be obtained by weighting the dimensions of the first image using a weighting parameter. This weighting parameter is a number greater than 0 and less than 1, and can be set according to the specific application scenario, such as 0.7. Here, the resulting weighted size is smaller than the dimensions of the first image. Then, a first tracking region is created in the first image based on the obtained weighted size. The center of this first tracking region is the optical center of the first image. The optical center of the first image is the projection point of the origin of the camera coordinate system (referring to the camera coordinate system when acquiring the first image) onto the first image. If the first image is not distorted, the optical center of the first image is its center. This provides an effective way to create the first tracking region, facilitating subsequent target tracking processing.

[0075] In some embodiments, the creation of a first tracking region in a first image acquired by a camera can be achieved by performing target detection processing on the first image according to a target detection model, and using the resulting region including the set target as the created first tracking region.

[0076] This application also provides another method for creating the first tracking region: performing object detection processing on the first image using an object detection model, and using any region that includes the specified target as the first tracking region. This application does not limit the type of object detection model; for example, it can be a region-convolutional neural network (R-CNN) model or a You Only LookOnce (YOLO) model. Similarly, it does not limit the specified target; the specified target can be at least one of several targets that the object detection model can detect, such as faces, cats, and dogs. By using the above method, the first tracking region includes the specified target, which facilitates improved accuracy in subsequent object tracking processing.

[0077] In step 102, the first tracking region is mapped onto the surface of the positioning space including the camera, and the pose of the mapped region in the world coordinate system is determined.

[0078] Here, a positioning space including the camera is created. This embodiment does not limit the shape of the positioning space; for example, it can be a sphere, cuboid, triangular pyramid, or square pyramid. The first tracking area created in step 101 is mapped (projected) onto the surface of the positioning space to obtain a mapped area. The pose of the mapped area in the world coordinate system is then determined. The world coordinate system is an absolute coordinate system used for camera positioning. This embodiment does not limit the coordinate axis directions or origin of the world coordinate system. For example, the world coordinate system can be created with the camera's real-time position as the origin when acquiring the first image. It is worth noting that in this embodiment, the pose of the mapped area in the world coordinate system refers to the pose of the mapped area relative to the world coordinate system; the same applies below.

[0079] In this embodiment of the application, the mapping region is fixed and the pose of the world coordinate system corresponding to the mapping region is also fixed. Therefore, in subsequent steps, the pose of the camera coordinate system corresponding to the world coordinate system can be determined by determining the pose of the camera coordinate system corresponding to the mapping region.

[0080] In step 103, target tracking processing is performed on the second image acquired by the camera based on the first tracking region to obtain the second tracking region in the second image.

[0081] Here, a second image acquired by the camera is obtained, wherein the second image can be any frame acquired later than the first image. Target tracking processing is performed on the second image acquired by the camera based on the first tracking region to obtain a second tracking region in the second image. That is, the first tracking region and the second tracking region include the same target. This application embodiment does not limit the method of target tracking processing; for example, target tracking processing can be implemented using an object tracking algorithm or an object detection algorithm.

[0082] In some embodiments, after step 103, the method further includes: when the second tracking region meets the reconstruction conditions, using the second image as a new first image, and creating a new first tracking region in the new first image; wherein the reconstruction conditions include at least one of the following: the second tracking region extends beyond the boundary of the second image; the size of the second tracking region is less than a size threshold; mapping the new first tracking region to the surface of the positioning space, and determining the pose of the new mapped region corresponding to the world coordinate system; performing target tracking processing on the new second image acquired by the camera based on the new first tracking region to obtain a new second tracking region in the new second image.

[0083] During target tracking, when the obtained second tracking region meets the reconstruction conditions, it proves that accurate camera positioning cannot be achieved based on this second tracking region. Therefore, the second image can be used as a new first image to restart camera positioning based on the new first image. The reconstruction conditions include at least one of the following: the second tracking region extends beyond the boundary of the second image; the size of the second tracking region is smaller than a size threshold, where the size threshold is smaller than the size of the second image, and can be specifically set according to the actual application scenario.

[0084] For ease of understanding, let's take the new first image P1 as an example. During the restart of camera localization, the dimensions of the new first image P1 can be weighted according to weighting parameters to obtain a weighted dimension. Then, centered on the optical center of the new first image P1, a new first tracking region is created within P1 based on this weighted dimension. Of course, the method for creating a new first tracking region is not limited to this; for example, it can also be created using a target detection model. Next, the new first tracking region is mapped onto the surface of the localization space to obtain a new mapped region, and the pose of the new mapped region in the world coordinate system is determined. Based on the new first tracking region, target tracking processing is performed on the new second image P2 acquired by the camera to obtain a new second tracking region within P2. Here, P2 can be any frame acquired later than P1. Similarly, in subsequent steps 104 and 105, "second tracking region" can be replaced with "new second tracking region," "mapped region" can be replaced with "new mapped region," and "second image" can be replaced with "new second image." This method improves the success rate and accuracy of camera localization.

[0085] In step 104, the pose of the camera coordinate system corresponding to the mapping area when the camera is acquiring the second image is determined based on the second tracking area.

[0086] To facilitate differentiation, the camera coordinate system used when the camera is capturing the first image is named the first camera coordinate system, and the camera coordinate system used when the camera is capturing the second image is named the second camera coordinate system.

[0087] Based on the second tracking region obtained from the target tracking process, the pose of the mapping region corresponding to the second camera coordinate system can be determined. For example, the depth of the mapping region corresponding to the second camera coordinate system can be determined based on the second tracking region, and then the pose of the mapping region corresponding to the second camera coordinate system can be determined based on the depth, which will be explained later.

[0088] In step 105, pose transformation is performed based on the pose of the world coordinate system corresponding to the mapped region and the pose of the camera coordinate system corresponding to the mapped region to obtain the pose of the camera in the world coordinate system when acquiring the second image.

[0089] Since the mapping region remains fixed, given the pose of the mapping region in the world coordinate system and the pose of the mapping region in the second camera coordinate system, the pose of the camera in the world coordinate system at the time of acquiring the second image can be obtained through pose transformation, thus achieving pose calibration of the camera at the moment of acquiring the second image. For example, the inverse matrix of the pose of the mapping region in the world coordinate system and the pose of the mapping region in the second camera coordinate system can be multiplied to obtain the pose of the camera in the world coordinate system at the time of acquiring the second image.

[0090] Similarly, camera localization can be performed based on images acquired after the second image. For example, target tracking processing can be performed on the third image acquired by the camera based on the second tracking region to obtain the third tracking region in the third image. Based on the third tracking region, the pose of the camera in the camera coordinate system corresponding to the mapping region when the third image was acquired is determined. Based on the pose of the world coordinate system corresponding to the mapping region and the pose of the camera coordinate system corresponding to the mapping region, pose transformation processing is performed to obtain the pose of the camera in the world coordinate system corresponding to the acquisition of the third image. Here, the third image is any frame acquired after the second image.

[0091] In some embodiments, between any steps, the method further includes: when any image is captured by the camera, performing observation processing on the virtual scene based on the observation pose, and presenting the virtual scene obtained through the observation processing; wherein, the observation pose represents the pose of the camera in the world coordinate system when capturing any image.

[0092] This application can be applied to Augmented Reality (AR) scenarios. For example, when any image is captured by a camera, the entire virtual scene is observed and processed based on the observation pose. That is, the entire virtual scene is observed and processed based on the viewpoint corresponding to the observation pose, and the virtual scene obtained through the observation processing is presented. Here, the observation pose refers to the pose of the camera in the world coordinate system when capturing any image. For example, it could be the pose of the camera in the world coordinate system when capturing the first image, or the pose of the camera in the world coordinate system when capturing the second image. The virtual scene obtained through the observation processing can be the entire virtual scene or a part of the entire virtual scene. This application does not limit the entire virtual scene; it could be, for example, virtual effects or virtual space. Through the above method, the accuracy of the presented virtual scene can be improved, that is, regardless of how the camera moves, the corresponding virtual scene can be presented, thus improving the human-computer interaction experience.

[0093] In some embodiments, the above-described presentation of the virtual scene obtained through observation processing can be achieved by performing any of the following processes: replacing any image to be presented with the virtual scene obtained through observation processing; overlaying the virtual scene obtained through observation processing with any image, and presenting the image obtained through overlay processing.

[0094] This application provides two presentation methods for virtual scenes obtained through observation processing. The first method replaces any image (the one that triggers the observation processing) with the virtual scene obtained through the observation processing, i.e., only the virtual scene obtained through observation processing is presented. This method is suitable for situations where the entire virtual scene is a virtual space (such as a game virtual space). The second method overlays the virtual scene obtained through observation processing with any image (again, the image that triggers the observation processing) and presents the resulting image. This method is suitable for situations where the entire virtual scene is a virtual effect (such as a magic pendant effect). These methods improve the flexibility of presentation, allowing any presentation method to be applied according to the needs of the actual application scenario.

[0095] like Figure 3A As shown, the embodiments of this application can achieve camera positioning with just two frames of images, which can greatly reduce the amount of computation in the camera positioning process and improve the efficiency of camera positioning.

[0096] In some embodiments, see Figure 3B , Figure 3B This is a schematic flowchart of the camera positioning method provided in the embodiments of this application. Figure 3A Step 102 shown can be implemented through steps 201 to 205, which will be explained in conjunction with each step.

[0097] In step 201, the first tracking area is mapped onto the surface of the positioning space including the camera to obtain the mapped area.

[0098] In step 202, the pose of the world coordinate system corresponding to the mapped region is determined based on the pose of the local mapped coordinate system corresponding to the mapped region.

[0099] Here, a local coordinate system for the mapping region is created. For ease of distinction, this local coordinate system is named the mapping coordinate system. By performing attitude transformation based on the attitude of the mapping region in the mapping coordinate system and the attitude of the mapping coordinate system in the world coordinate system, the attitude of the mapping region in the world coordinate system can be obtained.

[0100] In some embodiments, before step 202, the method further includes: creating a mapping coordinate system for the mapping region with the center of the mapping region as the origin and according to the coordinate axis direction of the world coordinate system; the above-mentioned determination of the orientation of the mapping region corresponding to the world coordinate system based on the orientation of the mapping region corresponding to the local mapping coordinate system can be achieved in this way: determining the orientation of the mapping region corresponding to the mapping coordinate system as the orientation of the mapping region corresponding to the world coordinate system.

[0101] Here, the center of the mapping region can be used as the origin, and a local mapping coordinate system for the mapping region can be created according to the coordinate axis direction of the world coordinate system. The center of the mapping region is the mapping point (projection point) of the center of the first tracking region on the surface of the positioning space.

[0102] Since the coordinate axes of the created mapped coordinate system are aligned with those of the world coordinate system, the pose of the mapped region in the mapped coordinate system can be directly used as the pose of the mapped region in the world coordinate system. In this case, the pose of the mapped region in the world coordinate system can be described by an identity matrix. Using this method, the pose of the mapped region in the world coordinate system can be determined quickly and accurately.

[0103] In step 203, the position of the center of the mapping area corresponding to the first camera coordinate system is determined, which is used as the position of the mapping area corresponding to the first camera coordinate system; wherein, the first camera coordinate system represents the camera coordinate system when the camera is acquiring the first image.

[0104] Here, since the center of the mapping region can represent the entire mapping region to a certain extent, the position (displacement) of the center of the mapping region corresponding to the first camera coordinate system can be directly used as the position of the mapping region corresponding to the first camera coordinate system.

[0105] In some embodiments, before step 102, the method further includes: creating a positioning space centered on the camera; the determination of the position of the center of the mapping area corresponding to the first camera coordinate system can be achieved in such a way as: determining the distance between the center of the mapping area and the camera as the position of the center of the mapping area corresponding to the first camera coordinate system.

[0106] For example, when capturing the first image using a camera, a positioning space centered on the camera's real-time position can be created. After mapping the area, the distance between the center of the mapped area and the camera's real-time position (the position when capturing the first image) can be determined, and this distance can be used as the position of the center of the mapped area corresponding to the first camera coordinate system.

[0107] For example, the positioning space can be a spherical positioning space with radius r centered on the real-time position of the camera. Then, the distance r between the center of the mapping area and the real-time position of the camera can be determined. Using this method, the position of the center of the mapping area corresponding to the first camera coordinate system can be determined quickly and accurately.

[0108] In step 204, position transformation is performed based on the pose of the camera in the world coordinate system when acquiring the first image and the position of the mapped area in the first camera coordinate system to obtain the position of the mapped area in the world coordinate system.

[0109] Here, the pose of the camera in the world coordinate system when acquiring the first image is determined, and a position transformation process is performed based on the pose of the camera in the world coordinate system when acquiring the first image and the position of the mapped area in the first camera coordinate system to obtain the position of the mapped area in the world coordinate system.

[0110] In some embodiments, before step 204, the method further includes: when the camera acquires the first image, creating a world coordinate system with the camera as the origin, and using the attitude acquired by the inertial measurement unit as the attitude of the inertial measurement unit corresponding to the world coordinate system; performing attitude transformation processing based on the attitude of the inertial measurement unit corresponding to the world coordinate system and the attitude of the camera corresponding to the coordinate system of the inertial measurement unit to obtain the attitude of the camera in the world coordinate system when acquiring the first image; using the zero matrix as the position of the camera in the world coordinate system when acquiring the first image; and combining the attitude of the camera in the world coordinate system when acquiring the first image and the position of the camera in the world coordinate system when acquiring the first image to obtain the pose of the camera in the world coordinate system when acquiring the first image.

[0111] For example, both the camera and the inertial measurement unit (IMU) are built-in components of electronic devices (such as mobile phones). Typically, the camera and IMU are fixed to the circuit board of the electronic device at the time of manufacture. Therefore, the attitude of the coordinate system corresponding to the camera and the IMU can be obtained in advance, for example, from the manufacturer of the electronic device. Of course, in this embodiment, the camera and IMU do not necessarily need to be deployed in the same electronic device, as long as the attitude of the coordinate system corresponding to the camera and the IMU can be obtained. For example, the camera and IMU can be deployed in different electronic devices, or they can both be deployed as independent devices.

[0112] When the camera acquires the first image, its real-time position can be used as the origin, and a world coordinate system can be created based on the set coordinate axis directions. Simultaneously, the attitude acquired by the inertial measurement unit (IMU) is obtained, and this attitude is used as the attitude of the IMU in the corresponding world coordinate system. Then, attitude transformation is performed based on the attitude of the IMU in the corresponding world coordinate system and the known attitude of the camera in the IMU's coordinate system to obtain the camera's attitude in the world coordinate system at the time of acquiring the first image.

[0113] Since the world coordinate system is constructed with the camera's real-time position as the origin, the zero matrix is ​​used as the camera's position in the world coordinate system when the first image is acquired. Then, by combining the camera's pose in the world coordinate system at the time of acquiring the first image with its position in the same coordinate system, the camera's pose in the world coordinate system at that time can be obtained. Using this method, camera pose initialization can be achieved with just one image frame, significantly improving camera localization efficiency compared to point cloud creation methods in related technologies.

[0114] In step 205, the pose of the world coordinate system corresponding to the mapped region and the position of the world coordinate system corresponding to the mapped region are combined to obtain the pose of the world coordinate system corresponding to the mapped region.

[0115] Given the pose of the mapped region in the world coordinate system and its position in the world coordinate system, the two can be combined to obtain the pose of the mapped region in the world coordinate system.

[0116] like Figure 3B As shown, the embodiments of this application can initialize the pose of the camera in the world coordinate system when acquiring the first image, and then determine the pose of the mapped area in the world coordinate system based on the pose, which can improve the efficiency of camera positioning, that is, it can achieve real-time positioning without going through a long preparation stage.

[0117] In some embodiments, see Figure 3C , Figure 3C This is a schematic flowchart of the camera positioning method provided in the embodiments of this application. Figure 3A Step 104 shown can be implemented through steps 301 to 303, which will be explained in conjunction with each step.

[0118] In step 301, the depth of the mapping region corresponding to the second camera coordinate system is determined based on the second tracking region; wherein, the second camera coordinate system represents the camera coordinate system when the camera is acquiring the second image.

[0119] Here, the depth of the mapped region corresponding to the second camera coordinate system is determined based on the second tracking region. This depth refers to the depth along the vertical axis of the second camera coordinate system. For example, if the world coordinate system is established with due south as the x-axis, due west as the y-axis, and the vertically downward direction as the z-axis, then the depth can refer to the depth along the z-axis.

[0120] In some embodiments, before step 102, the method further includes: creating a positioning space centered on the camera; the above-mentioned determination of the depth of the mapping area corresponding to the second camera coordinate system based on the second tracking area can be achieved in such a way that the depth of the mapping area corresponding to the second camera coordinate system is determined based on the size of the first tracking area, the size of the second tracking area, and the distance between the center of the mapping area and the camera.

[0121] When acquiring the first image using the camera, a positioning space can be created centered on the camera's real-time position. In this case, the depth of the mapping area corresponding to the second camera coordinate system can be determined based on the size of the first tracking area, the size of the second tracking area, and the distance between the center of the mapping area and the camera.

[0122] For example, the distance can be multiplied by the width of the first tracking region (e.g., the width on the x-axis of the image plane coordinate system of the first image), and the result of the multiplication can be divided by the width of the second tracking region (e.g., the width on the x-axis of the image plane coordinate system of the second image) to obtain the depth of the mapped region in the second camera coordinate system. In this way, the depth of the mapped region in the second camera coordinate system can be accurately and effectively determined.

[0123] In step 302, the depth of the mapped region in the second camera coordinate system, the camera's intrinsic parameter matrix, and the center position of the second image are fused to obtain the position of the mapped region in the second camera coordinate system.

[0124] For example, the depth of the mapped region corresponding to the second camera coordinate system, the inverse of the camera's intrinsic parameter matrix, and the center position of the second image (such as the center position expressed in homogeneous coordinates) can be multiplied together to obtain the position of the mapped region corresponding to the second camera coordinate system.

[0125] The camera's intrinsic parameter matrix can be obtained in advance or through real-time calibration.

[0126] In step 303, the pose of the camera in the world coordinate system when acquiring the second image and the position of the mapped region in the second camera coordinate system are combined to obtain the pose of the mapped region in the second camera coordinate system.

[0127] Here, the pose of the camera in the world coordinate system when acquiring the second image can be used as the pose of the mapped region in the second camera coordinate system. Then, the pose of the mapped region in the second camera coordinate system and its position in the second camera coordinate system can be combined to obtain the pose of the mapped region in the second camera coordinate system.

[0128] In some embodiments, before step 303, the method further includes: when the camera acquires the second image, taking the attitude acquired by the inertial measurement component as the attitude of the inertial measurement component in the world coordinate system; performing attitude transformation processing based on the attitude of the inertial measurement component in the world coordinate system and the attitude of the camera in the coordinate system of the inertial measurement component to obtain the attitude of the camera in the world coordinate system when acquiring the second image.

[0129] Similarly, when the camera and inertial measurement unit (IMU) are deployed in the same electronic device, the attitude acquired by the IMU during the acquisition of the second image can be used as the attitude of the IMU in the corresponding world coordinate system. Then, based on the attitude of the IMU in the corresponding world coordinate system and the known attitude of the camera in the IMU's coordinate system, an attitude transformation is performed to obtain the camera's attitude in the corresponding world coordinate system during the acquisition of the second image. In this way, the camera's attitude in the corresponding world coordinate system during the acquisition of the second image can be accurately determined by combining the IMU with the data from the camera.

[0130] like Figure 3C As shown, the embodiments of this application can quickly and accurately determine the pose of the second camera coordinate system corresponding to the mapping area based on the second tracking area, thereby improving the accuracy of camera positioning.

[0131] In some embodiments, see Figure 3D , Figure 3D This is a schematic flowchart of the camera positioning method provided in the embodiments of this application. Figure 3A Step 103 shown can be implemented by any one of steps 401 to 403, and will be explained in conjunction with each step.

[0132] In step 401, a region detection model is trained based on the first tracking region in the first image and the region that is different from the first tracking region, and the second image is processed by the trained region detection model to obtain the second tracking region in the second image.

[0133] This application provides three methods for target tracking processing, which will be described in detail below.

[0134] In the first approach, a region detection model is trained based on a first tracking region in the first image and regions in the first image that are distinct from the first tracking region. For example, the first tracking region in the first image can be used as a positive sample, and regions in the first image that are distinct from the first tracking region can be used as negative samples. The region detection model is trained based on the positive and negative samples. The type of region detection model is not limited; for example, it can be a filter model. Similarly, the shape of the regions distinct from the first tracking region is not limited; for example, they can have the same shape as the first tracking region, such as both being rectangles.

[0135] After the region detection model is trained, the second image is processed by region detection based on the trained region detection model. The region obtained by the region detection process is used as the second tracking region in the second image, thereby realizing target tracking processing.

[0136] In step 402, forward point tracking and reverse point tracking are performed based on multiple initial points in the first tracking region. Based on the positional differences between the multiple initial points and the multiple points obtained through reverse point tracking, the multiple points obtained through forward point tracking are filtered, and the second tracking region in the second image is determined based on the filtered points. Herein, forward point tracking refers to point tracking from the first image to the second image; reverse point tracking refers to point tracking from the second image to the first image.

[0137] In the second approach, target tracking is performed on a point-by-point basis, based on forward-backward consistency. For example, multiple initial points are uniformly generated in the first tracking region of the first image, and forward point tracking is performed on each initial point to obtain the corresponding point in the second image (named the forward tracking point for easy distinction). The forward point tracking can be implemented based on the pixel movement from the first image to the second image. For ease of understanding, taking an initial point 1 in the first tracking region of the first image as an example, after forward point tracking, the corresponding forward tracking point 2 in the second image can be obtained.

[0138] Then, reverse point tracking processing is performed on each forward tracking point in the second image to obtain the corresponding point in the first image (named the reverse tracking point for easy distinction). For example, reverse point tracking processing is performed on the forward tracking point point2 in the second image to obtain the corresponding reverse tracking point point3 in the first image.

[0139] Then, based on the positional differences between multiple initial points and multiple reverse tracking points obtained through reverse point tracking, the multiple forward tracking points obtained through forward point tracking are filtered. For example, for initial point 1, the positional difference between initial point 1 and reverse tracking point 3 can be determined. The positional difference corresponding to each initial point can be determined in the same way. Then, the forward tracking points corresponding to the initial points with the smallest positional differences are selected as the filtered forward tracking points. The number of filtered forward tracking points can be preset, such as being set to half the total number of initial points.

[0140] Based on the selected positive tracking points, the second tracking region in the second image can be determined. For example, the smallest region in the second image that includes all the selected positive tracking points can be taken as the second tracking region; however, the method is not limited to this.

[0141] In step 403, the target to be tracked is obtained by performing target detection processing on the first tracking region according to the target detection model, and the target to be tracked is obtained by performing target detection processing on the second image according to the target detection model, and the region including the target to be tracked is taken as the second tracking region.

[0142] The first and second methods described above both use target tracking algorithms. In the third method provided in this application embodiment, target tracking can also be achieved through target detection algorithms.

[0143] For example, a trained object detection model can be used to perform object detection processing on the first tracking region, and the target obtained from the object detection processing (i.e., the target included in the first tracking region) can be used as the target to be tracked, such as a face. Then, the same object detection model can be used to perform object detection processing on the second image, and the region in the detected second image that includes the target to be tracked can be used as the second tracking region, that is, object tracking is achieved based on the detected target to be tracked. This application does not limit the type of object detection model; for example, it can be an R-CNN model or a YOLO model.

[0144] like Figure 3D As shown, the embodiments of this application can improve the flexibility of target tracking processing, and any of the above methods can be applied according to the needs of actual application scenarios.

[0145] The following describes an exemplary application of the embodiments of this application in AR application scenarios. In the embodiments of this application, when images are captured by the camera in an electronic device (such as a smartphone), the entire virtual scene can be observed and processed according to the pose of the camera in the world coordinate system (i.e., the entire virtual scene can be observed according to the viewpoint corresponding to the pose), and the virtual scene obtained through observation and processing can be presented in the user interface (such as the camera interface) provided by the electronic device. AR interaction can be realized based on the presented virtual scene, thereby improving the human-computer interaction experience. The observation and processing and the presentation of the virtual scene can be implemented by a specific rendering engine in the electronic device.

[0146] For a virtual scene to be presented, the virtual scene can be overlaid on an image captured by a camera and then presented. As an example, embodiments of this application provide... Figure 4 The diagram shows a camera interface where virtual scene 41 is a magic pendant effect superimposed on image 42, which is an image captured by the camera of a keyboard in the real world. As an example, embodiments of this application also provide... Figure 5 The diagram shows a camera interface where a virtual scene 51 is overlaid on an image 52, which is an image captured by the camera of a room in the real world. It is worth noting that in this embodiment, for the virtual scene to be presented, the image captured by the camera can be replaced with the virtual scene itself, i.e., only the virtual scene is presented.

[0147] During the presentation of a virtual scene, the presented virtual scene can change accordingly as the camera moves. As an example, embodiments of this application provide... Figure 6 The diagram illustrates the virtual scene. At the first real-time moment, virtual scene 61 is displayed on the camera interface. At the second real-time moment, because the camera's pose in the world coordinate system changes compared to the first moment, virtual scene 62 is displayed on the camera interface, distinct from virtual scene 61. At the third real-time moment, because the camera's pose in the world coordinate system changes compared to the second moment, virtual scene 63 is displayed on the camera interface, distinct from virtual scene 62. The second moment is later than the first moment, and the third moment is later than the second moment. Thus, by moving the camera, users can observe the entire virtual scene from different perspectives on the camera interface, enabling them to roam within the entire virtual scene and enhancing the user experience.

[0148] Next, the camera positioning scheme of this application embodiment will be described from the perspective of low-level implementation. As an example, the following is provided: Figure 7 The flowchart shown will be combined with Figure 7Please provide an explanation.

[0149] Step 1) Initialization.

[0150] like Figure 7 As shown, the initialization process involves sub-steps such as image acquisition, obtaining IMU values, and determining the camera's pose in the world coordinate system when acquiring the first frame image. For example, when the first frame image (e.g., the first image mentioned above) is acquired by the camera in the electronic device, a world coordinate system W is created with the camera's real-time position as the origin. The direction of the coordinate axes of the world coordinate system W is not limited; for example, due south can be used as the x-axis, due west as the y-axis, and the vertically downward direction as the z-axis. Simultaneously, the attitude R acquired by the inertial measurement unit (IMU) of the electronic device is obtained. imu1 R imu1 That is, the value obtained from the IMU, R imu1 It can be a 3x3 matrix. According to R... imu1 The pose R of the camera in the world coordinate system W when acquiring the first frame can be determined. c1 Formula R c1 =R ic -1 *R imu1 *R ic Among them, R ic This indicates the orientation of the camera (or camera coordinate system) relative to the inertial measurement unit's coordinate system when the first frame of image is acquired. Since the inertial measurement unit and camera are typically fixed to the circuit board of an electronic device, R can be known in advance. ic The value of R ic It can be a 3x3 matrix. (R) c1 This represents the pose of the camera (or camera coordinate system) in the world coordinate system W when the first frame of the image is acquired. c1 It can also be a 3x3 matrix.

[0151] It is worth noting that, when the electronic device is a smartphone, and the smartphone is placed horizontally with the screen facing upwards and the line connecting the charging port and the earpiece pointing north, the orientation of the camera coordinate system coincides with the orientation of the world coordinate system W. In this case, the camera's orientation corresponds to the world coordinate system W. That is, the identity matrix.

[0152] When acquiring the first frame, since the world coordinate system W is created with the camera's real-time position as the origin, the camera's displacement (position) P corresponding to the world coordinate system W at this time is... c1 It is 0, that is

[0153] R c1 and P c1By combining these steps, we can obtain the pose T of the camera in the world coordinate system W when the first frame of the image is acquired. W,c1 ,Right now

[0154] Step 2) Create an anchor block (corresponding to the mapping area above) and determine the pose of the anchor block in the world coordinate system.

[0155] When acquiring the first frame image, a positioning space is created centered on the camera's real-time position. This embodiment does not limit the shape of the positioning space; it can be a sphere, cuboid, triangular pyramid, or square pyramid, etc. For ease of understanding, a sphere is used as an example. The created spherical positioning space is named the Sky Sphere, and its radius is s. The value of s can be set according to the actual application scenario, for example, it can be 1.5 meters.

[0156] For a camera, its intrinsic parameter matrix K can be obtained. K can be a 3x3 matrix, for example... Where f represents the camera's focal length, which can be in pixels; u represents the horizontal coordinate of the optical center in the current frame image (referring to any frame image captured by the camera); and v represents the vertical coordinate of the optical center in the current frame image.

[0157] For ease of understanding, the embodiments of this application provide, as follows: Figure 8 The diagram shown illustrates the camera imaging process. Figure 8 In the diagram, coordinate system Oxyz represents the camera coordinate system, O represents the optical center of the camera coordinate system, and coordinate system O'-x'-y' represents the coordinate system of the pixel plane (also known as the physical imaging plane or image plane), with O' representing the optical center of the pixel plane. Then, u can represent the x-coordinate of the optical center O' in the current frame image, and v can represent the y-coordinate of the optical center O' in the current frame image. It is possible to treat the x-coordinate and y-coordinate of any point corresponding to any corner of the current frame image as 0 to facilitate determining the x-coordinate and y-coordinate of the optical center O'.

[0158] Let `width` represent the width of the current frame image captured by the camera (i.e., the width along the x' axis), and `height` represent the height of the current frame image (i.e., the height along the y' axis). If there is no distortion in the current frame image, then `width = 2 * u` and `height = 2 * v`. A tracking region B (corresponding to the first tracking region mentioned above) can be created centered on the optical center p(u,v) of the current frame image. This embodiment does not limit the shape of the tracking region; it can be rectangular, circular, elliptical, or irregular. For ease of understanding, a rectangle will be used as an example below. The width of the tracking region B is `ratio * width`, and its height is `ratio * height`. Here, `p(u,v)` is... Figure 8In the O'; ratio*width and ratio*height correspond to the weighted dimensions mentioned above. ratio is a weighted parameter that can be set according to the actual application scenario, such as setting it to 0.7.

[0159] Mapping the created tracking region B onto the surface of the sky sphere yields the anchor block (corresponding to the mapped region above). The 3D center of this anchor block (corresponding to the center of the mapped region above) is P. an P an This corresponds to p(u,v) in the current frame image. Then, based on the coordinate axis directions of the world coordinate system W, an anchor block coordinate system (corresponding to the mapped coordinate system mentioned above) is created for the anchor block itself, for example, with the 3D center P of the anchor block as the coordinate system. an An anchor block coordinate system is created with the origin at south, the x-axis at due west, and the z-axis at vertically downward. Therefore, the pose of the anchor block coordinate system is the same as the pose of the world coordinate system W; that is, the anchor block (or anchor block coordinate system) corresponds to the pose of the world coordinate system W. That is, R W,Anchor It can be a 3x3 identity matrix.

[0160] Due to the three-dimensional center P an On the surface of a sky sphere with radius s centered at the camera, and P an The projection point (mapping point) in the current frame image is the optical center p(u,v). Therefore, P can be... an The position of the anchor block in the corresponding camera coordinate system (referring to the camera coordinate system used when acquiring the current frame image) is taken as the position t of the anchor block in the corresponding camera coordinate system. c,Anchor ,Right now

[0161] When acquiring the current frame image, the pose T of the camera in the corresponding world coordinate system W can be determined. Wc ,For example Among them, R W,c The t represents the camera's pose in the world coordinate system W. W,c This represents the camera's position in the world coordinate system W. If the current frame is the first frame, then T... W,c =T W,c1 .

[0162] According to T W,c and t c,Anchor By performing a position transformation, we can obtain the position t of the anchor block in the world coordinate system W. W,Anchor The formula for position transformation processing is as follows: t W,Anchor =R W,c ×t c,Anchor +t W,c .

[0163] R W,Anchor and t W,Anchor By combining the data, we can obtain the pose T of the anchor block in the world coordinate system W. W,Anchor ,Right now Since the anchor block is fixed in place, regardless of how the camera moves subsequently, T W,Anchor It is fixed and unchanging.

[0164] Step 3) Target tracking processing.

[0165] like Figure 7 As shown, for each new image acquired after the current frame, target tracking is performed by a tracker. The tracker can be implemented using an object tracking algorithm or an object detection algorithm. Object tracking algorithms include Kernelized Correlation Filter (KCF), Tracking-Learning-Detection (TLD), or Median Flow Tracker, etc. Object detection algorithms include Region-Convolutional Neural Networks (R-CNN) and You Only Look Once (YOLO) algorithms, etc. The implementation method of the tracker is not limited in this embodiment.

[0166] For ease of understanding, let's take the current frame image as the first frame image captured by the camera as an example. Then, in the second frame image captured by the camera (corresponding to the second image mentioned above), the center of the tracking region B' (corresponding to the second tracking region mentioned above) obtained through target tracking processing is p2(u2,v2), and the width of the tracking region B' is w2. For ease of calculation, the position of p2 can be converted into homogeneous coordinates, i.e., p2'(u2,v2,1).

[0167] like Figure 7 As shown, based on the obtained tracking region B', the anchor block can be tracked. For example, in the camera coordinate system used to acquire the second frame image (for ease of distinction, it is named the second camera coordinate system), the three-dimensional center P of the anchor block can be determined. an The depth 'depth2' corresponding to the second camera coordinate system is given by the formula as follows: Where depth2 represents the three-dimensional center P an The value of the z-axis corresponding to the second camera coordinate system.

[0168] Furthermore, P can be determined an The three-dimensional coordinates t corresponding to the second camera coordinate system c2,Anchor (i.e., position), the formula is as follows

[0169] While acquiring the second frame image, the attitude R acquired by the inertial measurement unit (IMU) of the electronic device is obtained. imu2 R imu2 This is the new value of the IMU obtained. According to R... imu2 Determine the pose R of the camera in the world coordinate system W when acquiring the second frame image. c2 Formula R c2 =R ic -1 *R imu2 *R ic .

[0170] R c2 and t c2,Anchor By performing combined processing, the pose T of the anchor block corresponding to the second camera coordinate system can be obtained. c2,Anchor ,Right now

[0171] Since the fixed pose T of the anchor block in the world coordinate system W has been obtained in the above steps, W,Anchor Therefore, it can be based on T W,Anchor and T c2,Anchor Determine the pose T of the camera in the world coordinate system W when acquiring the second frame image. W,c2 That is, T W,c2 =T W,Anchor ×(T c2,Anchor ) -1 .

[0172] In the same way, the pose T of each frame acquired after the first frame can be calibrated. W,cn , where n is an integer greater than 1.

[0173] It is worth noting that, such as Figure 7 As shown, in determining the three-dimensional center P of the anchor block anBefore the depth2 corresponding to the second camera coordinate system, it can be determined whether to recreate the tracking region and the corresponding anchor based on the tracking region B'. For the newly created anchor, the pose of the new anchor in the world coordinate system W is also determined to facilitate tracking of the new anchor. For example, when the tracking region (such as tracking region B') obtained during target tracking meets the reconstruction conditions, the tracking region and the corresponding anchor are recreated according to step 2). The reconstruction conditions include at least one of the following: ① The tracking region exceeds the boundary (edge) of the image, for example, any vertex of the tracking box moves to or exceeds the boundary of the image; ② The size of the tracking region is less than a size threshold, for example, the width of the tracking region is less than ratio_min*width, where ratio_min is less than ratio. The specific value can be set according to the actual application scenario, such as setting ratio_min to 0.3. This ensures the accuracy and effectiveness of target tracking processing.

[0174] like Figure 7 As shown, when it is necessary to create a new tracking region and a new anchor, target tracking processing can be performed based on the new tracking region to track the new anchor, thereby determining the pose of the camera in the world coordinate system W. When it is not necessary to create a new tracking region and a new anchor, the tracking region obtained from the target tracking processing (such as tracking region B') can be used to track the already created anchor, thereby determining the pose of the camera in the world coordinate system W.

[0175] The embodiments of this application have at least the following technical advantages: 1) It does not require a long time to translate the camera to create a point cloud; the embodiments of this application can initialize the pose of the camera in the world coordinate system when acquiring the first frame image; 2) It does not require a high-precision inertial measurement unit; a high camera positioning accuracy can still be achieved with a common inertial measurement unit, and it can be applied to more types of electronic devices; 3) The amount of computation in the camera positioning process is small, that is, less computing resources are consumed, and the performance requirements of the central processing unit (CPU) are low, so it can be applied to more types of electronic devices; 4) The camera positioning accuracy is high; the camera can be accurately positioned even when the user moves the camera in any direction; 5) The corresponding virtual scene is presented according to the pose of the camera in the world coordinate system, which can achieve a realistic AR effect and can be widely used in scenarios such as mobile AR and AR glasses.

[0176] The following continues to describe the exemplary structure of the camera positioning device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the camera positioning device 455 in the memory 450 may include: a region creation module 4551, used to create a first tracking region in a first image acquired by the camera; a mapping module 4552, used to map the first tracking region onto a surface of the positioning space including the camera, and determine the pose of the mapped region in the world coordinate system; a region tracking module 4553, used to perform target tracking processing on a second image acquired by the camera based on the first tracking region, to obtain a second tracking region in the second image; the region tracking module 4553 is also used to determine the pose of the mapped region in the camera coordinate system when the camera acquires the second image based on the second tracking region; and a pose transformation module 4554, used to perform pose transformation processing based on the pose of the mapped region in the world coordinate system and the pose of the mapped region in the camera coordinate system, to obtain the pose of the camera in the world coordinate system when acquiring the second image.

[0177] In some embodiments, the mapping module 4552 is further configured to: determine the pose of the mapping region corresponding to the world coordinate system based on the pose of the mapping region corresponding to the local mapping coordinate system; determine the position of the center of the mapping region corresponding to the first camera coordinate system as the position of the mapping region corresponding to the first camera coordinate system; wherein the first camera coordinate system represents the camera coordinate system when the camera is acquiring the first image; perform position transformation processing based on the pose of the camera corresponding to the world coordinate system when acquiring the first image and the position of the mapping region corresponding to the first camera coordinate system to obtain the position of the mapping region corresponding to the world coordinate system; and combine the pose of the mapping region corresponding to the world coordinate system and the position of the mapping region corresponding to the world coordinate system to obtain the pose of the mapping region corresponding to the world coordinate system.

[0178] In some embodiments, the mapping module 4552 is further configured to: when acquiring a first image through a camera, create a world coordinate system with the camera as the origin, and use the attitude acquired by the inertial measurement unit as the attitude of the inertial measurement unit corresponding to the world coordinate system; perform attitude transformation processing based on the attitude of the inertial measurement unit corresponding to the world coordinate system and the attitude of the camera corresponding to the coordinate system of the inertial measurement unit to obtain the attitude of the camera in the world coordinate system when acquiring the first image; use the zero matrix as the position of the camera in the world coordinate system when acquiring the first image; and combine the attitude of the camera in the world coordinate system when acquiring the first image and the position of the camera in the world coordinate system when acquiring the first image to obtain the pose of the camera in the world coordinate system when acquiring the first image.

[0179] In some embodiments, the mapping module 4552 is further configured to: create a mapping coordinate system for the mapping region with the center of the mapping region as the origin and according to the coordinate axis direction of the world coordinate system; and determine the orientation of the mapping coordinate system corresponding to the mapping region as the orientation of the world coordinate system corresponding to the mapping region.

[0180] In some embodiments, the mapping module 4552 is further configured to: create a positioning space centered on the camera; determine the distance between the center of the mapping area and the camera, so as to use the position of the center of the mapping area corresponding to the first camera coordinate system.

[0181] In some embodiments, the region tracking module 4553 is further configured to: determine the depth of the mapping region corresponding to the second camera coordinate system based on the second tracking region; wherein the second camera coordinate system represents the camera coordinate system when the camera acquires the second image; fuse the depth of the mapping region corresponding to the second camera coordinate system, the camera's intrinsic parameter matrix, and the center position of the second image to obtain the position of the mapping region corresponding to the second camera coordinate system; and combine the pose of the camera in the world coordinate system when acquiring the second image and the position of the mapping region corresponding to the second camera coordinate system to obtain the pose of the mapping region corresponding to the second camera coordinate system.

[0182] In some embodiments, the mapping module 4552 is further configured to: create a positioning space centered on the camera; the region tracking module 4553 is further configured to: determine the depth of the mapping region corresponding to the second camera coordinate system based on the size of the first tracking region, the size of the second tracking region, and the distance between the center of the mapping region and the camera.

[0183] In some embodiments, the region tracking module 4553 is further configured to: perform any of the following processes: train a region detection model based on a first tracking region in a first image and a region distinct from the first tracking region, and perform region detection processing on a second image based on the trained region detection model to obtain a second tracking region in the second image; perform forward point tracking processing and reverse point tracking processing based on multiple initial points in the first tracking region, filter multiple points obtained through forward point tracking processing based on the positional differences between the multiple initial points and the multiple points obtained through reverse point tracking processing, and determine the second tracking region in the second image based on the filtered points; wherein, forward point tracking processing represents point tracking processing from the first image to the second image; and reverse point tracking processing represents point tracking processing from the second image to the first image; perform target detection processing on the first tracking region based on the target detection model to obtain a target to be tracked, and perform target detection processing on the second image based on the target detection model, and take the obtained region including the target to be tracked as the second tracking region.

[0184] In some embodiments, the region creation module 4551 is further configured to perform any of the following processes: weighting the size of the first image to obtain a weighted size, and creating a first tracking region in the first image based on the weighted size with the optical center of the first image as the center; performing target detection processing on the first image according to the target detection model, and using the obtained region including the set target as the created first tracking region.

[0185] In some embodiments, the region tracking module 4553 is further configured to: when the second tracking region meets the reconstruction conditions, use the second image as a new first image, and create a new first tracking region in the new first image; wherein the reconstruction conditions include at least one of the following: the second tracking region extends beyond the boundary of the second image; the size of the second tracking region is less than a size threshold; map the new first tracking region onto the surface of the positioning space, and determine the pose of the new mapped region corresponding to the world coordinate system; perform target tracking processing on the new second image acquired by the camera based on the new first tracking region to obtain a new second tracking region in the new second image.

[0186] In some embodiments, the camera positioning device 455 further includes a presentation module, which is used to perform observation processing on the virtual scene according to the observation pose when any image is captured by the camera, and present the virtual scene obtained through the observation processing; wherein, the observation pose represents the pose of the camera in the world coordinate system when capturing any image.

[0187] In some embodiments, the presentation module is further configured to perform any of the following processes: replacing any presented image with a virtual scene obtained through observation processing; overlaying the virtual scene obtained through observation processing with any image, and presenting the image obtained through overlay processing.

[0188] This application provides a computer program product or computer program that includes computer instructions (i.e., executable instructions) stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the camera positioning method described above in this application.

[0189] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to perform the method provided in this application, for example... Figure 3A , Figure 3B , Figure 3C and Figure 3D The camera positioning method is shown.

[0190] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0191] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0192] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0193] As an example, executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0194] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A camera positioning method, characterized in that, The method includes: A first tracking region is created in the first image acquired by the camera; The first tracking area is mapped onto the surface of the positioning space including the camera, and the pose of the mapped area corresponding to the world coordinate system is determined. The positioning space is a three-dimensional space created that includes the camera and is used for camera positioning. The shape of the positioning space includes a sphere, a cuboid, a triangular pyramid, or a square pyramid. During the movement of the camera, the mapped area on the surface of the positioning space and the pose of the world coordinate system corresponding to the mapped area remain unchanged. Based on the first tracking region, target tracking processing is performed on the second image acquired by the camera to obtain the second tracking region in the second image; Based on the second tracking region, determine the pose of the camera coordinate system corresponding to the mapping region when the camera is acquiring the second image; Based on the pose of the mapped region corresponding to the world coordinate system and the pose of the mapped region corresponding to the camera coordinate system, pose transformation processing is performed to obtain the pose of the camera in the world coordinate system when acquiring the second image.

2. The method according to claim 1, characterized in that, The determination of the pose in the world coordinate system corresponding to the mapped region obtained by mapping includes: The attitude of the mapping region corresponding to the world coordinate system is determined based on the attitude of the mapping region corresponding to the local mapping coordinate system. The position of the center of the mapping area corresponding to the first camera coordinate system is determined, and is used as the position of the mapping area corresponding to the first camera coordinate system; wherein, the first camera coordinate system represents the camera coordinate system when the camera is acquiring the first image; Based on the pose of the camera in the world coordinate system when acquiring the first image and the position of the mapped region in the first camera coordinate system, a position transformation process is performed to obtain the position of the mapped region in the world coordinate system. The pose of the mapped region in the world coordinate system and the position of the mapped region in the world coordinate system are combined to obtain the pose of the mapped region in the world coordinate system.

3. The method according to claim 2, characterized in that, Before performing position transformation processing based on the pose of the camera in the world coordinate system when acquiring the first image and the position of the mapped region in the first camera coordinate system, the method further includes: When the first image is captured by the camera, the world coordinate system is created with the camera as the origin, and the attitude acquired by the inertial measurement unit is taken as the attitude of the inertial measurement unit corresponding to the world coordinate system. An attitude transformation process is performed based on the attitude of the inertial measurement unit corresponding to the world coordinate system and the attitude of the camera corresponding to the coordinate system of the inertial measurement unit to obtain the attitude of the camera corresponding to the world coordinate system when acquiring the first image. The zero matrix is ​​used as the position of the camera in the world coordinate system when acquiring the first image; The pose of the camera in the world coordinate system when acquiring the first image and the position of the camera in the world coordinate system when acquiring the first image are combined to obtain the pose of the camera in the world coordinate system when acquiring the first image.

4. The method according to claim 2, characterized in that, Before determining the pose of the world coordinate system corresponding to the mapped region based on the pose of the local mapped coordinate system corresponding to the mapped region, the method further includes: A mapping coordinate system for the mapping region is created with the center of the mapping region as the origin, based on the coordinate axis directions of the world coordinate system. Determining the pose of the mapped region corresponding to the world coordinate system based on the pose of the mapped region in the local mapped coordinate system includes: The pose of the mapped region corresponding to the mapped coordinate system is determined, and is used as the pose of the mapped region corresponding to the world coordinate system.

5. The method according to claim 2, characterized in that, Before mapping the first tracking region onto a surface including the positioning space of the camera, the method further includes: Create a positioning space centered on the camera; Determining the position of the center of the mapping region corresponding to the first camera coordinate system includes: The distance between the center of the mapping area and the camera is determined as the position of the center of the mapping area corresponding to the first camera coordinate system.

6. The method according to claim 1, characterized in that, The step of determining the pose of the camera coordinate system corresponding to the mapping region when the camera acquires the second image, based on the second tracking region, includes: The depth of the mapping region corresponding to the second camera coordinate system is determined based on the second tracking region; wherein, the second camera coordinate system represents the camera coordinate system when the camera is acquiring the second image; The depth of the mapped region corresponding to the second camera coordinate system, the intrinsic parameter matrix of the camera, and the center position of the second image are fused to obtain the position of the mapped region corresponding to the second camera coordinate system. The pose of the camera in the world coordinate system when acquiring the second image and the position of the mapped region in the second camera coordinate system are combined to obtain the pose of the mapped region in the second camera coordinate system.

7. The method according to claim 6, characterized in that, Before mapping the first tracking region onto a surface including the positioning space of the camera, the method further includes: Create a positioning space centered on the camera; Determining the depth of the mapping region corresponding to the second camera coordinate system based on the second tracking region includes: The depth of the mapping region corresponding to the second camera coordinate system is determined based on the size of the first tracking region, the size of the second tracking region, and the distance between the center of the mapping region and the camera.

8. The method according to claim 1, characterized in that, The step of performing target tracking processing on the second image acquired by the camera based on the first tracking region to obtain the second tracking region in the second image includes: Perform any of the following processes: A region detection model is trained based on the first tracking region in the first image and regions different from the first tracking region, and the second image is processed by region detection based on the trained region detection model to obtain the second tracking region in the second image. Forward point tracking and reverse point tracking are performed based on multiple initial points in the first tracking region. Based on the positional differences between the multiple initial points and the multiple points obtained through the reverse point tracking, the multiple points obtained through the forward point tracking are filtered, and the second tracking region in the second image is determined based on the filtered points. Wherein, the forward point tracking processing refers to point tracking processing from the first image to the second image; the reverse point tracking processing refers to point tracking processing from the second image to the first image; The target to be tracked is obtained by performing target detection processing on the first tracking region according to the target detection model, and the target to be tracked is obtained by performing target detection processing on the second image according to the target detection model, and the region including the target to be tracked is taken as the second tracking region.

9. The method according to claim 1, characterized in that, Creating a first tracking region in the first image acquired by the camera includes: Perform any of the following processes: The size of the first image is weighted to obtain a weighted size, and a first tracking region is created in the first image with the optical center of the first image as the center, based on the weighted size. The first image is processed by object detection model, and the region including the target is used as the first tracking region.

10. The method according to claim 1, characterized in that, After performing target tracking processing on the second image acquired by the camera based on the first tracking region to obtain the second tracking region in the second image, the method further includes: When the second tracking region meets the reconstruction conditions, the second image is used as the new first image, and a new first tracking region is created in the new first image; The reconstruction conditions include at least one of the following: the second tracking region extends beyond the boundary of the second image; the size of the second tracking region is smaller than a size threshold. The new first tracking region is mapped onto the surface of the positioning space, and the pose of the new mapped region corresponding to the world coordinate system is determined. Based on the new first tracking region, target tracking processing is performed on the new second image acquired by the camera to obtain a new second tracking region in the new second image.

11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: When any image is captured by the camera, the virtual scene is observed and processed according to the observation pose, and the virtual scene obtained through the observation and processing is presented. Wherein, the observation pose refers to the pose of the camera in the world coordinate system when acquiring any one of the images.

12. The method according to claim 11, characterized in that, The presentation of the virtual scene obtained through the observation processing includes: Perform any of the following processes: Replace any of the presented images with a virtual scene obtained through the observation process; The virtual scene obtained through the observation process is overlaid with any one of the images, and the image obtained through the overlay process is presented.

13. A camera positioning device, characterized in that, The device includes: The region creation module is used to create a first tracking region in the first image acquired by the camera; The mapping module is used to map the first tracking area onto the surface of the positioning space including the camera, and to determine the pose of the mapped area in the world coordinate system. The positioning space is a three-dimensional space created that includes the camera and is used for camera positioning. The shape of the positioning space includes a sphere, a cuboid, a triangular pyramid, or a square pyramid. During the movement of the camera, the mapped area on the surface of the positioning space and the pose of the world coordinate system corresponding to the mapped area remain unchanged. A region tracking module is used to perform target tracking processing on a second image acquired by the camera based on the first tracking region, to obtain a second tracking region in the second image; The region tracking module is further configured to determine, based on the second tracking region, the pose of the camera coordinate system corresponding to the mapping region when the camera is acquiring the second image; The pose transformation module is used to perform pose transformation processing based on the pose of the mapped region corresponding to the world coordinate system and the pose of the mapped region corresponding to the camera coordinate system, so as to obtain the pose of the camera in the world coordinate system when acquiring the second image.

14. An electronic device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the camera positioning method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, It stores executable instructions for implementing the camera positioning method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Patent Citations

  • Pose determination method and device, equipment and storage medium

    CN111768454A