Target detection method of AR device and AR device
By optimizing the pose of the target object using target detection algorithms and camera pose changes in AR devices, the problems of high cost and high power consumption of LiDAR on portable devices are solved, achieving high-precision target detection and interactive effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HISENSE ELECTRONICS TECH SHENZHEN CO LTD
- Filing Date
- 2022-11-07
- Publication Date
- 2026-08-04
AI Technical Summary
In existing technologies, 3D target detection technology based on lidar is costly and consumes a lot of power on portable devices, resulting in low target detection accuracy and precision.
By using target detection algorithms in AR devices to obtain the scale of the target object and the pose of the camera in the world coordinate system, and combining the change in the camera pose to optimize the initial pose of the target object, the accurate positioning of the target object in the camera coordinate system can be achieved.
Without using LiDAR, the accuracy of target detection was improved, the impact of position jitter on interaction was mitigated, and precise target positioning was achieved.
Smart Images

Figure CN116092071B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology for virtual reality devices, and more particularly to a target detection method for AR devices and an AR device. Background Technology
[0002] Augmented Reality (AR) technology provides users with rich interactive experiences between the virtual and real worlds. Users can interact with virtual objects in real-world scenes using AR devices. Currently, this interaction primarily involves users generating virtual objects in real-world scenes using AR devices and determining the virtual objects' positions in the real world using information from the AR device, such as the camera's pose. Furthermore, users increasingly desire virtual objects to interact with real-world objects, for example... Figure 1 Virtual items can be placed around real-world objects. Or, for example... Figure 2 The movement of virtual objects can create collision effects with real-world objects. In the interaction between virtual and real-world objects, 3D (3D) object detection technology plays a crucial role.
[0003] Existing 3D target detection technologies primarily employ lidar-based solutions. While generating 3D point cloud data using lidar can yield high-precision 3D target positions, the high cost and power consumption of lidar and similar acquisition devices make them unsuitable for direct application in portable devices. However, existing target detection technologies that do not utilize lidar suffer from low detection accuracy, resulting in a low overall target detection rate. Summary of the Invention
[0004] This application provides a target detection method and an AR device for AR devices, which can achieve accurate target positioning when using acquisition devices such as LiDAR, thereby improving the accuracy of target detection.
[0005] In a first aspect, embodiments of this application provide a target detection method for an AR device, comprising:
[0006] For an image containing a target object captured by the camera in an AR device at any given time, a target detection algorithm is used to detect the target in the image to obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width, and height of the target object; and,
[0007] Based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, the pose of the camera in the target coordinate system at the current moment is obtained, wherein the target coordinate system is the coordinate system corresponding to the target object;
[0008] Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, the initial pose of the target object in the camera coordinate system is obtained at the current moment, wherein the local pose of the target is obtained based on the scale of the target;
[0009] The initial pose of the target object at the current moment is optimized by the pose change of the camera to obtain the target pose of the target object in the camera coordinate system at the current moment. The pose change is the change of the camera pose in the world coordinate system between the current moment and the previous moment.
[0010] A second aspect of this application provides an AR device, including a processor and a memory, wherein the processor and the memory are connected via a bus;
[0011] The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program:
[0012] For an image containing a target object captured by the camera in an AR device at any given time, a target detection algorithm is used to detect the target in the image to obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width, and height of the target object; and,
[0013] Based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, the pose of the camera in the target coordinate system at the current moment is obtained, wherein the target coordinate system is the coordinate system corresponding to the target object;
[0014] Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, the initial pose of the target object in the camera coordinate system is obtained at the current moment, wherein the local pose of the target is obtained based on the scale of the target;
[0015] The initial pose of the target object at the current moment is optimized by the pose change of the camera to obtain the target pose of the target object in the camera coordinate system at the current moment. The pose change is the change of the camera pose in the world coordinate system between the current moment and the previous moment.
[0016] According to a third aspect of the present invention, a computer storage medium is provided, the computer storage medium storing a computer program for performing the method as described in the first aspect.
[0017] In the above embodiments of this application, the camera pose in the target coordinate system is obtained based on the scale of the detected target object and the camera pose in the world coordinate system. Then, based on the camera pose in the target coordinate system and the local pose of the target object, the initial pose of the target object in the camera coordinate system is obtained. Finally, the initial pose of the target object is optimized by the camera pose change to obtain the target pose of the target object in the camera coordinate system. In this embodiment, optimizing the initial pose of the target object by the camera pose change can effectively smooth the 3D position of the target object relative to the camera in each frame of the video, mitigating the interactive effects of positional jitter on the target object. Thus, accurate target positioning is achieved without using acquisition equipment such as LiDAR, improving the accuracy of target detection. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 An exemplary illustration shows one of the application scenarios provided in the embodiments of this application;
[0020] Figure 2 The second example illustration shows a schematic diagram of an application scenario provided in an embodiment of this application;
[0021] Figure 3 This document exemplarily illustrates the third application scenario diagram provided in the embodiments of this application;
[0022] Figure 4 One of the flowcharts of the target detection method for AR devices provided in this application is illustrated by way of example;
[0023] Figure 5 An exemplary schematic diagram of the process for determining the second pose change provided in an embodiment of this application is shown;
[0024] Figure 6 An exemplary illustration shows a flowchart of determining the target pose of the target object in the camera coordinate system at the current moment, provided by an embodiment of this application.
[0025] Figure 7 An exemplary schematic diagram illustrates the process of optimizing the scale of a target object according to an embodiment of this application;
[0026] Figure 8An exemplary illustration shows a flowchart of determining the target scale of the target object at the current moment, provided in an embodiment of this application.
[0027] Figure 9 A second flowchart of the target detection method for AR devices provided in this application is illustrated by way of example;
[0028] Figure 10 An exemplary schematic diagram of the target detection device for an AR device provided in an embodiment of this application is shown;
[0029] Figure 11 An exemplary hardware structure diagram of an AR device provided in an embodiment of this application is shown. Detailed Implementation
[0030] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.
[0031] Based on the exemplary embodiments described in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the appended claims. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation on its own.
[0032] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0033] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to be omnipresent but not exclusive; for example, a product or device comprising a series of components is not necessarily limited to those explicitly listed, but may include other components not explicitly listed or inherent to such product or device.
[0034] As used in this application, the term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.
[0035] The following is an overview of the ideas behind the embodiments of this application.
[0036] Current 3D target detection technologies primarily employ LiDAR-based solutions. While generating 3D point cloud data using LiDAR can yield high-precision 3D target positions, the high cost and power consumption of LiDAR and similar acquisition devices make them difficult to apply directly to portable devices. However, performing 3D target detection without LiDAR or similar acquisition devices results in lower detection accuracy and reduced target detection precision.
[0037] Addressing the issue of low target detection accuracy in existing technologies without the use of acquisition devices such as LiDAR, this application provides a target detection method for AR devices. The method involves obtaining the camera's pose in the target coordinate system based on the detected target object's scale and the camera's pose in the world coordinate system. Then, based on the camera's pose in the target coordinate system and the target object's local pose, the initial pose of the target object in the camera coordinate system is obtained. Finally, the initial pose of the target object is optimized using the camera's pose change, resulting in the target pose of the target object in the camera coordinate system. This embodiment optimizes the initial pose of the target object using the camera's pose change, effectively smoothing the 3D position of the target object relative to the camera in each frame of the video, mitigating the interactive effects of positional jitter on the target object. Therefore, it achieves accurate target localization without the use of acquisition devices such as LiDAR, improving the accuracy of target detection.
[0038] The embodiments of this application are described in detail below with reference to the accompanying drawings.
[0039] Figure 1 An exemplary illustration shows a schematic diagram of an application scenario provided by an embodiment of this application; such as Figure 1 As shown, this application scenario uses an AR device as the server as an example. This application scenario includes an AR device 110 and a server 120. Server 120 can be implemented using a single server or multiple servers. Server 120 can be implemented using a physical server or a virtual server.
[0040] In one possible application scenario, AR device 110 sends an image containing a target object to server 120. Server 120 receives the image containing the target object captured by the camera in AR device 110 at any given time, performs target detection on the image using a target detection algorithm, and obtains the scale of the target object at the current time. The scale of the target object at the current time includes the length, width, and height of the target object. Based on the scale of the target object at the current time and the pose of the camera in the world coordinate system at the current time, server 120 obtains the pose of the camera in the target coordinate system at the current time. The target coordinate system is the coordinate system corresponding to the target object. Based on the pose of the camera in the target coordinate system at the current time and the local pose of the target object at the current time, server 120 obtains the initial pose of the target object in the camera coordinate system at the current time. The local pose of the target object is obtained based on the scale of the target object. Finally, server 120 optimizes the initial pose of the target object at the current moment by the pose change of the camera, and obtains the target pose of the target object in the camera coordinate system at the current moment. The pose change is the change of the camera pose in the world coordinate system between the current moment and the previous moment.
[0041] like Figure 2 The diagram illustrates another application scenario, which includes an AR device 110, a server 120, and a storage device 130. The AR device 110 stores images containing the target object captured in the storage device 130. The server 120 retrieves an image containing the target object at any given moment from the storage device 130 and performs target detection on the image using a target detection algorithm to obtain the scale of the target object at the current moment. The scale of the target object at the current moment includes the length, width, and height of the target object. Based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, the server 120 obtains the pose of the camera in the target coordinate system at the current moment, where the target coordinate system is the coordinate system corresponding to the target object. Furthermore, based on the pose of the camera in the target coordinate system at the current moment and the local pose of the target object at the current moment, the server 120 obtains the initial pose of the target object in the camera coordinate system at the current moment, where the local pose of the target object is obtained based on the target's scale. Finally, server 120 optimizes the initial pose of the target object at the current moment by the pose change of the camera, and obtains the target pose of the target object in the camera coordinate system at the current moment. The pose change is the change of the camera pose in the world coordinate system between the current moment and the previous moment.
[0042] like Figure 3The diagram illustrates another application scenario, which includes an AR device 110 and a memory 130. The AR device 110, for an image containing a target object captured at any given time, performs target detection using a target detection algorithm to obtain the scale of the target object. The scale of the target object at the current time includes its length, width, and height. Based on the target object's scale at the current time and the camera's pose in the world coordinate system at the current time, the AR device 110 obtains the camera's pose in the target coordinate system at the current time, where the target coordinate system is the coordinate system corresponding to the target object. Furthermore, based on the camera's pose in the target coordinate system at the current time and the target object's local pose at the current time, the initial pose of the target object in the camera coordinate system at the current time is obtained. The target's local pose is obtained based on the target's scale. Finally, the AR device 110 optimizes the initial pose of the target object at the current moment by the pose change of the camera, and obtains the target pose of the target object in the camera coordinate system at the current moment. The pose change is the change of the camera's pose in the world coordinate system between the current moment and the previous moment.
[0043] The description in this application focuses on a single AR device 110, a single server 120, and a single memory 130. However, those skilled in the art should understand that the illustrated AR device 110, server 120, and memory 130 are intended to illustrate the operation of the AR device 110, server 120, and memory 130 involved in the technical solutions of this application, and do not imply any limitation on the number, type, or location of the AR device 110, server 120, and memory 130. It should be noted that adding additional modules to or removing individual modules from the illustrated environment will not change the underlying concept of the exemplary embodiments of this application.
[0044] It should be noted that the target detection method for AR devices proposed in this application is not only applicable to… Figure 1 , Figure 2 and Figure 3 The application scenarios shown can also be applied to any target detection device with AR equipment.
[0045] The target detection method of the AR device of this application, in conjunction with the application scenarios described above and with reference to the accompanying drawings, is described below as an exemplary embodiment. It should be noted that the above application scenarios are only shown to facilitate understanding of the methods and principles of this application, and the implementation of this application is not limited in any way in this respect.
[0046] like Figure 4 The diagram shown illustrates a target detection method for AR devices, which may include the following steps:
[0047] Step 401: For an image containing a target object captured by the camera in the AR device at any time, perform target detection on the image using a target detection algorithm to obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width and height of the target object;
[0048] In this embodiment, before performing target detection on the image using the target detection algorithm, preprocessing is required. This preprocessing includes, but is not limited to, dynamically filling, scaling, and normalizing the image to a single-precision type of a specified size. In this embodiment, the single-precision type is float32, and the specified size is 512×512×1. However, this embodiment does not limit the image type or size; the image type and size can be set according to actual circumstances.
[0049] It should be noted that the target detection algorithm in this embodiment can be set according to the actual situation, and this embodiment does not limit the target detection algorithm.
[0050] Step 402: Based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, obtain the pose of the camera in the target coordinate system at the current moment, wherein the target coordinate system is the coordinate system corresponding to the target object;
[0051] Since AR devices are equipped with camera pose tracking algorithms, they can directly obtain the camera's pose in the world coordinate system at every moment through these algorithms.
[0052] In one embodiment, step 402 can be implemented as follows: inputting the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment into the pnp (Perspective-n-Point) algorithm to obtain the pose of the camera in the target coordinate system at the current moment.
[0053] It should be noted that the target coordinate system in this embodiment is a spatial rectangular coordinate system established with the centroid of the target object as the origin and the lines corresponding to the length, width, and height of the target object as the x-axis, y-axis, and z-axis, respectively. However, the specific directions of the coordinate axes can be set according to the actual situation, and this embodiment does not limit them.
[0054] Step 403: Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, obtain the initial pose of the target object in the camera coordinate system at the current moment, wherein the local pose of the target is obtained based on the scale of the target;
[0055] It should be noted that the local pose of the target is determined by defining the length of the target's dimensions as the x-coordinate, the width of the target's dimensions as the y-coordinate, and the height of the target's dimensions as the y-coordinate. The direction of the target is a preset initial direction. Therefore, the position coordinates of the local pose are obtained using the x-coordinate, y-coordinate, y-coordinate, and initial direction.
[0056] In one embodiment, step 403 can be implemented as follows: multiplying the current pose of the camera in the target coordinate system with the current local pose of the target object to obtain the current initial pose of the target object in the camera coordinate system. The initial pose of the target object in the camera coordinate system can be obtained using formula (1):
[0057]
[0058] in, Let i be the initial pose of the target object in the camera coordinate system at the current time. Let i be the pose of the camera in the target coordinate system at the current time. Let i be the local pose of the target object at the current time i.
[0059] Step 404: Optimize the initial pose of the target object at the current moment by using the pose change of the camera to obtain the target pose of the target object in the camera coordinate system at the current moment, wherein the pose change is the change of the camera pose in the world coordinate system between the current moment and the previous moment.
[0060] The pose change includes a first pose change and a second pose change. The first pose change is obtained based on the pose of the camera in the world coordinate system, and the second pose change is obtained based on the pose of the camera in the target coordinate system.
[0061] Below, we will first provide a detailed introduction to the methods for determining the first pose change and the second pose change:
[0062] First pose change: The first pose change is obtained by subtracting the current pose of the camera in the world coordinate system from the pose of the camera in the world coordinate system at the previous moment.
[0063] Second pose change: such as Figure 5 The diagram shown illustrates the process for determining the second pose change, including the following steps:
[0064] Step 501: Based on the current pose of the camera in the target coordinate system and the current pose of the camera in the world coordinate system, convert the local pose of the target object into the local pose of the target object in the world coordinate system.
[0065] In one embodiment, step 501 can be implemented as follows: dividing the current pose of the camera in the target coordinate system by the current pose of the camera in the world coordinate system to obtain the local pose of the target object in the world coordinate system. The local pose of the target object in the world coordinate system can be obtained through formula (2):
[0066]
[0067] Among them, T O2W The local pose of the target object in the world coordinate system. Let i be the pose of the camera in the target coordinate system at the current time. Let i be the pose of the camera in the world coordinate system at the current time.
[0068] Step 502: Based on the local pose of the target object in the world coordinate system and the pose of the camera in the target coordinate system at the previous moment, obtain the pose of the camera in the world coordinate system at the previous moment;
[0069] Since the local pose of the target object in the world coordinate system is invariant, the pose of the camera in the world coordinate system at the previous moment can be determined by the local pose of the target object in the world coordinate system.
[0070] In one embodiment, step 502 can be implemented as follows: multiplying the local pose of the target object in the world coordinate system with the pose of the camera in the target coordinate system at the previous moment to obtain the pose of the camera in the world coordinate system at the previous moment. The pose of the camera in the world coordinate system at the previous moment can be obtained through formula (3):
[0071]
[0072] in, Let T be the pose of the camera in the world coordinate system at the previous time i-1. O2W The local pose of the target object in the world coordinate system. The pose of the camera in the target coordinate system at the previous time i-1.
[0073] Step 503: Subtract the pose of the camera in the world coordinate system at the previous moment from the pose of the camera in the world coordinate system at the current moment to obtain the second pose change.
[0074] After introducing the first pose change and how to determine it, the following describes the specific method for determining the target pose of the target object in the camera coordinate system at the current moment, such as... Figure 6 The diagram illustrates the process of determining the target pose of the target object in the camera coordinate system at the current moment, including the following steps:
[0075] Step 601: Based on the intermediate weight matrix of the target object at the previous time step, obtain the initial weight matrix of the target object at the current time step;
[0076] In one embodiment, step 601 can be implemented as follows: multiplying the intermediate weight matrix with a preset state transition matrix to obtain a first intermediate state transition matrix; multiplying the first intermediate state transition matrix with the inverse of the preset state transition matrix to obtain a second intermediate state transition matrix; and adding the second intermediate state transition matrix with a preset first noise matrix to obtain the initial weight matrix of the target object at the current time. The initial weight matrix of the target object at the current time can be obtained using formula (4):
[0077]
[0078] in, Let A be the initial weight matrix of the target object at the current time i, and let P be the preset state transition matrix. i-1 Let A' be the intermediate weight matrix corresponding to the target object at the previous time i-1, let Q be the inverse matrix of the preset state transition matrix, and let Q be the preset first noise matrix.
[0079] It should be noted that if the current time is the first time, then the intermediate weight matrix corresponding to the target object in the previous time is the preset initial intermediate weight matrix.
[0080] Step 602: Using the initial weight matrix of the target object at the current time, obtain the target weight matrix of the target object at the current time; wherein, the target weight matrix of the target object at the current time can be obtained through formula (5):
[0081]
[0082] Among them, K i Let H be the target weight matrix of the target object at the current time i, H be the preset second noise matrix, H′ be the inverse matrix of the second noise matrix, and R be the preset third noise matrix.
[0083] After obtaining the target weight matrix, the intermediate weight matrix at the current time can be updated using formula (6):
[0084]
[0085] Among them, P i Let I be the intermediate weight matrix at the current time i, and let I be the preset identity matrix.
[0086] Step 603: Obtain the target pose change using the target weight matrix, the first pose change, and the second pose change; wherein the target pose change can be obtained using formula (7):
[0087]
[0088] Where, ΔT i The change in the target pose. This represents the second pose change. This represents the change in the first pose.
[0089] Step 604: Based on the target pose change and the initial pose, obtain the target pose of the target object in the camera coordinate system at the current moment.
[0090] In one embodiment, step 604 can be implemented as follows: multiplying the target pose change by the initial pose to obtain the target pose. The target pose can be obtained using formula (8):
[0091]
[0092] in, Let ΔT be the target pose of the target object in the camera coordinate system at the current time i. i The change in the target pose. Let i be the initial pose of the target object in the camera coordinate system at the current time i.
[0093] To further improve the accuracy of target detection, in one embodiment, after performing step 404, as follows: Figure 7 The diagram shows a flowchart for optimizing the scale of a target object, including the following steps:
[0094] Step 701: Optimize the scale of the target object at the current moment using the target pose of the target object at the current moment to obtain the target scale of the target object at the current moment;
[0095] like Figure 8 The diagram illustrates the process of determining the target scale of the target object at the current moment, including the following steps:
[0096] Step 801: Reproject the target pose of the target object at the current moment onto the corresponding image to obtain the actual position of the target object in the image;
[0097] The reprojection method in this embodiment can be set according to the actual situation, and this embodiment does not limit the reprojection method.
[0098] Step 802: Input the target pose of the target object into a pre-trained neural network to obtain the predicted position of the target object in the image;
[0099] It should be noted that this embodiment does not limit the neural network; the neural network in this embodiment can be set according to the actual situation.
[0100] Step 803: Obtain the adjustment weight based on the actual position and the predicted position;
[0101] In one embodiment, step 803 can be implemented as follows: multiply the initial adjustment weight by the actual position to obtain the adjusted actual position; subtract the adjusted actual position from the predicted position to obtain an error value; if the error value is not less than a specified threshold, adjust the initial adjustment weight using a preset method to obtain the adjusted initial adjustment weight; determine the adjusted initial adjustment weight as the initial adjustment weight; then return to the step of multiplying the initial adjustment weight by the actual position to obtain the adjusted actual position, until the error value is less than the specified threshold, then determine the initial adjustment weight as the adjustment weight.
[0102] The preset method involves increasing or decreasing the initial adjustment weight by a preset value each time to obtain the adjusted initial adjustment weight. However, this embodiment does not limit the preset method; the preset method in this embodiment can be set according to the actual situation.
[0103] Step 804: Adjust the scale of the target object at the current moment based on the adjustment weight to obtain the target scale of the target object at the current moment.
[0104] In one embodiment, the adjustment weight is multiplied by the scale of the target object at the current moment to obtain the target scale of the target object at the current moment. The target scale of the target object at the current moment can be obtained through formula (9):
[0105]
[0106] in, The target scale of the target object at the current moment. Let α0 be the scale of the target object at the current moment, and let α0 be the adjustment weight.
[0107] Step 702: Input the target scale of the target object at the current moment and the scale of the target object in the image at the next moment into the Kalman filter algorithm to smooth the scale of the target object in the image at the next moment, and obtain the smoothed scale of the target object in the image at the next moment. Then, determine the smoothed scale as the scale of the target object in the image at the next moment, so as to obtain the target pose of the target object in the camera coordinate system at the next moment based on the scale of the target object in the image at the next moment.
[0108] In this embodiment, the Kalman filter algorithm is used to smooth the scale of the image to obtain the smoothed scale of the target object in the image at the next moment, which is the method in the prior art, and will not be described in detail here.
[0109] To further improve the accuracy of target detection, in one embodiment, before performing step 402, the scale of the target object at the current moment and the target scale of the target object in the image at the previous moment are input into the Kalman filter algorithm to smooth the scale of the target object in the image at the current moment, so as to obtain the smoothed scale of the target object in the image at the current moment, and the smoothed scale of the target object in the image at the current moment is determined as the scale of the target object at the current moment.
[0110] Once the target pose of the target object in the camera coordinate system is obtained, a virtual target can be placed on the target object, or a collision between the virtual target and the target object can be achieved based on this pose. Since these two methods are existing technologies, the specific implementation details will not be elaborated upon in this embodiment.
[0111] To further connect the technical solutions in this application, the following is combined with... Figure 9 A detailed explanation may include the following steps:
[0112] Step 901: For an image containing a target object captured by the camera in the AR device at any time, perform target detection on the image using a target detection algorithm to obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width and height of the target object;
[0113] Step 902: Input the scale of the target object at the current moment and the target scale of the target object in the image at the previous moment into the Kalman filter algorithm to smooth the scale of the target object in the image at the current moment, and obtain the smoothed scale of the target object in the image at the current moment, and determine the smoothed scale of the target object in the image at the current moment as the scale of the target object at the current moment;
[0114] Step 903: Input the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment into the PNP algorithm to obtain the pose of the camera in the target coordinate system at the current moment;
[0115] Step 904: Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, obtain the initial pose of the target object in the camera coordinate system at the current moment, wherein the local pose of the target is obtained based on the scale of the target;
[0116] Step 905: Based on the intermediate weight matrix of the target object at the previous time step, obtain the initial weight matrix of the target object at the current time step;
[0117] Step 906: Using the initial weight matrix of the target object at the current time, obtain the target weight matrix of the target object at the current time;
[0118] Step 907: Obtain the target pose change using the target weight matrix, the first pose change, and the second pose change; wherein the first pose change is obtained based on the camera's pose in the world coordinate system, and the second pose change is obtained based on the camera's pose in the target coordinate system.
[0119] Step 908: Based on the target pose change and the initial pose, obtain the target pose of the target object in the camera coordinate system at the current moment;
[0120] Step 909: Optimize the scale of the target object at the current moment using the target pose of the target object at the current moment to obtain the target scale of the target object at the current moment;
[0121] Step 910: Input the target scale of the target object at the current moment and the scale of the target object in the image at the next moment into the Kalman filter algorithm to smooth the scale of the target object in the image at the next moment, and obtain the smoothed scale of the target object in the image at the next moment. Then, determine the smoothed scale as the scale of the target object in the image at the next moment, so as to obtain the target pose of the target object in the camera coordinate system at the next moment based on the scale of the target object in the image at the next moment.
[0122] Based on the same inventive concept, the target detection method for AR devices disclosed above can also be implemented by a target detection device for AR devices. The effect of this target detection device for AR devices is similar to that of the aforementioned method, and will not be described in detail here.
[0123] Figure 10 This is a schematic diagram of the structure of a target detection device for an AR device according to an embodiment of the present disclosure.
[0124] like Figure 10 As shown, the target detection device 1000 of the AR device disclosed herein may include a target detection module 1010, a camera pose determination module 1020, an initial pose determination module 1030, and a target pose determination module 1040.
[0125] The target detection module 1010 is used to perform target detection on an image containing a target object captured by the camera in the AR device at any time using a target detection algorithm, and obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width and height of the target object;
[0126] The camera pose determination module 1020 is used to obtain the pose of the camera in the target coordinate system at the current moment based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, wherein the target coordinate system is the coordinate system corresponding to the target object;
[0127] The initial pose determination module 1030 is used to obtain the initial pose of the target object in the camera coordinate system at the current moment based on the pose of the camera in the target coordinate system at the current moment and the local pose of the target object at the current moment, wherein the local pose of the target is obtained based on the scale of the target.
[0128] The target pose determination module 1040 is used to optimize the initial pose of the target object at the current moment by the pose change of the camera, so as to obtain the target pose of the target object in the camera coordinate system at the current moment, wherein the pose change is the change of the camera pose in the world coordinate system between the current moment and the previous moment.
[0129] In one embodiment, the apparatus further includes:
[0130] The scale optimization module 1050 is used to optimize the initial pose of the target object at the current moment by the pose change of the camera, and after obtaining the target pose of the target object in the camera coordinate system at the current moment, optimize the scale of the target object at the current moment by the target pose of the target object at the current moment, and obtain the target scale of the target object at the current moment.
[0131] The target scale of the target object at the current moment and the scale of the target object in the image at the next moment are input into the Kalman filter algorithm to smooth the scale of the target object in the image at the next moment, so as to obtain the smoothed scale of the target object in the image at the next moment, and the smoothed scale is determined as the scale of the target object in the image at the next moment, so as to obtain the target pose of the target object in the camera coordinate system at the next moment based on the scale of the target object in the image at the next moment.
[0132] In one embodiment, the scale optimization module 1050 performs the optimization of the scale of the target object at the current moment using the target pose of the target object at the current moment to obtain the target scale of the target object at the current moment, specifically for:
[0133] The target pose of the target object at the current moment is reprojected onto the corresponding image to obtain the actual position of the target object in the image; and...
[0134] The target pose of the target object is input into a pre-trained neural network to obtain the predicted position of the target object in the image;
[0135] The adjustment weights are obtained based on the actual location and the predicted location;
[0136] The scale of the target object at the current moment is adjusted based on the adjustment weight to obtain the target scale of the target object at the current moment.
[0137] In one embodiment, the camera pose determination module 1020 is specifically used for:
[0138] The scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment are input into the PNP algorithm to obtain the pose of the camera in the target coordinate system at the current moment.
[0139] The initial pose determination module 1030 is specifically used for:
[0140] Multiply the current pose of the camera in the target coordinate system by the current local pose of the target object to obtain the current initial pose of the target object in the camera coordinate system.
[0141] In one embodiment, the pose change includes a first pose change and a second pose change, wherein the first pose change is obtained based on the pose of the camera in the world coordinate system, and the second pose change is obtained based on the pose of the camera in the target coordinate system.
[0142] The target pose determination module 1040 is specifically used for:
[0143] Based on the intermediate weight matrix of the target object at the previous time step, the initial weight matrix of the target object at the current time step is obtained.
[0144] Using the initial weight matrix of the target object at the current time, the target weight matrix of the target object at the current time is obtained;
[0145] The target pose change is obtained by using the target weight matrix, the first pose change, and the second pose change.
[0146] Based on the target pose change and the initial pose, the target pose of the target object in the camera coordinate system at the current moment is obtained.
[0147] In one embodiment, the target pose determination module 1040 executes the process of obtaining the initial weight matrix of the target object at the current time based on the intermediate weight matrix corresponding to the target object at the previous time step, specifically for:
[0148] Multiply the intermediate weight matrix by the preset state transition matrix to obtain the first intermediate state transition matrix;
[0149] Multiply the first intermediate state transition matrix by the inverse of the preset state transition matrix to obtain the second intermediate state transition matrix;
[0150] The second intermediate state transition matrix is added to the preset first noise matrix to obtain the initial weight matrix of the target object at the current time.
[0151] In one embodiment, the target pose determination module 1040 is further configured to:
[0152] The initial weight matrix is obtained using the following formula:
[0153]
[0154] in, Let A be the initial weight matrix of the target object at the current time i, and let P be the preset state transition matrix. i-1 Let A' be the intermediate weight matrix corresponding to the target object at the previous time i-1, let Q be the inverse matrix of the preset state transition matrix, and let Q be the preset first noise matrix.
[0155] In one embodiment, the target pose determination module 1040 performs the step of obtaining the target weight matrix of the target object at the current time using the initial weight matrix of the target object at the current time, specifically for:
[0156] The target weight matrix of the target object at the current time can be obtained using the following formula:
[0157]
[0158] Among them, K i Let H be the target weight matrix of the target object at the current time i, H be the preset second noise matrix, H′ be the inverse matrix of the second noise matrix, and R be the preset third noise matrix.
[0159] The target pose determination module 1040 performs the step of obtaining the target pose change amount through the target weight matrix, the first pose change amount, and the second pose change amount, specifically for:
[0160] The target pose change is obtained using the following formula:
[0161]
[0162] Where, ΔT i The change in the target pose. This represents the second pose change. This represents the change in the first pose.
[0163] In one embodiment, the target pose determination module 1040 performs the step of obtaining the target pose of the target object in the camera coordinate system at the current moment based on the target pose change and the initial pose, specifically for:
[0164] The target pose change is obtained by multiplying the target pose change by the initial pose.
[0165] After introducing a target detection method and apparatus for an AR device according to an exemplary embodiment of the present invention, the following describes an AR device according to another exemplary embodiment of the present invention.
[0166] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as "circuit", "module", or "system".
[0167] In some possible implementations, the AR device according to the present invention may include at least one processor and at least one computer storage medium. The computer storage medium stores program code that, when executed by the processor, causes the processor to perform the steps in the target detection method of the AR device according to various exemplary embodiments of the present invention described above. For example, the processor may perform actions such as... Figure 4 Steps 401-404 are shown in the diagram.
[0168] The following reference Figure 11 To describe an AR device 1100 according to this embodiment of the present invention. Figure 11 The AR device 1100 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0169] like Figure 11 As shown, AR device 1100 is presented in the form of a general AR device. The components of AR device 1100 may include, but are not limited to: at least one processor 1101, at least one computer storage medium 1102, and a bus 1103 connecting different system components (including computer storage medium 1102 and processor 1101).
[0170] Bus 1103 represents one or more of several bus structures, including a computer storage media bus or computer storage media controller, peripheral bus, processor, or local bus using any of the various bus structures.
[0171] Computer storage medium 1102 may include readable media in the form of volatile computer storage media, such as random access computer storage medium (RAM) 1121 and / or cache storage medium 1122, and may further include read-only computer storage medium (ROM) 1123.
[0172] The computer storage medium 1102 may also include a program / utility 1125 having a set (at least one) of program modules 1124, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0173] AR device 1100 can also communicate with one or more external devices 1104 (e.g., keyboard, pointing device, etc.), one or more devices that enable a user to interact with AR device 1100, and / or any device that enables AR device 1100 to communicate with one or more other AR devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1105. Furthermore, AR device 1100 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1106. As shown, network adapter 1106 communicates with other modules used for AR device 1100 via bus 1103. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with AR device 1100, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0174] In some possible implementations, various aspects of the target detection method for an AR device provided by the present invention can also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to cause the computer device to perform the steps in the target detection method for an AR device according to various exemplary embodiments of the present invention described above.
[0175] The program product may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access computer storage media (RAM), read-only computer storage media (ROM), erasable programmable read-only computer storage media (EPROM or flash memory), optical fibers, portable compact disk read-only computer storage media (CD-ROM), optical computer storage media, magnetic computer storage media, or any suitable combination thereof.
[0176] The target detection program product of the AR device according to embodiments of the present invention can be a portable compact disc read-only computer storage medium (CD-ROM) and include program code, and can run on the AR device. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0177] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0178] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0179] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as "C" or similar languages. The program code can execute entirely on the user's AR device, partially on the user's device, as a standalone software package, partially on the user's AR device and partially on a remote AR device, or entirely on a remote AR device or server. In cases involving remote AR devices, the remote AR device can be connected to the user's AR device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external AR device (e.g., via the Internet using an Internet service provider).
[0180] It should be noted that although several modules of the apparatus have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0181] Furthermore, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0182] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk computer storage media, CD-ROMs, optical computer storage media, etc.) containing computer-usable program code.
[0183] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0184] These computer program instructions may also be stored in a computer-readable computer storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable computer storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0185] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0186] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A target detection method for an AR device, characterized in that, The method includes: For an image containing a target object captured by the camera in an AR device at any given time, a target detection algorithm is used to detect the target in the image to obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width, and height of the target object; and, Based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, the pose of the camera in the target coordinate system at the current moment is obtained, wherein the target coordinate system is the coordinate system corresponding to the target object; Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, the initial pose of the target object in the camera coordinate system is obtained at the current moment, wherein the local pose of the target is obtained based on the scale of the target; Based on the intermediate weight matrix of the target object at the previous time step, the initial weight matrix of the target object at the current time step is obtained. Using the initial weight matrix of the target object at the current time, the target weight matrix of the target object at the current time is obtained; The target pose change is obtained by using the target weight matrix, the first pose change, and the second pose change. Based on the target pose change and the initial pose, the target pose of the target object in the camera coordinate system at the current moment is obtained, wherein the first pose change is obtained based on the pose of the camera in the world coordinate system, and the second pose change is obtained based on the pose of the camera in the target coordinate system.
2. The method according to claim 1, characterized in that, After obtaining the target pose of the target object in the camera coordinate system at the current moment based on the target pose change and the initial pose, the method further includes: The target size of the target object at the current moment is optimized by using the target pose of the target object at the current moment; The target scale of the target object at the current moment and the scale of the target object in the image at the next moment are input into the Kalman filter algorithm to smooth the scale of the target object in the image at the next moment, so as to obtain the smoothed scale of the target object in the image at the next moment, and determine the smoothed scale as the scale of the target object in the image at the next moment, so as to obtain the target pose of the target object in the camera coordinate system at the next moment based on the scale of the target object in the image at the next moment.
3. The method according to claim 2, characterized in that, The step of optimizing the scale of the target object at the current moment using the target pose of the target object at the current moment to obtain the target scale of the target object at the current moment includes: The target pose of the target object at the current moment is reprojected onto the corresponding image to obtain the actual position of the target object in the image; and... The target pose of the target object is input into a pre-trained neural network to obtain the predicted position of the target object in the image; The adjustment weights are obtained based on the actual location and the predicted location; The scale of the target object at the current moment is adjusted based on the adjustment weight to obtain the target scale of the target object at the current moment.
4. The method according to claim 1, characterized in that, The step of obtaining the camera's pose in the target coordinate system at the current moment based on the target object's scale at the current moment and the camera's pose in the world coordinate system at the current moment includes: The scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment are input into the PNP algorithm to obtain the pose of the camera in the target coordinate system at the current moment. Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, the initial pose of the target object in the camera coordinate system at the current moment is obtained, including: Multiply the current pose of the camera in the target coordinate system by the current local pose of the target object to obtain the current initial pose of the target object in the camera coordinate system.
5. The method according to claim 1, characterized in that, The step of obtaining the initial weight matrix of the target object at the current time based on the intermediate weight matrix corresponding to the target object at the previous time step includes: Multiply the intermediate weight matrix by the preset state transition matrix to obtain the first intermediate state transition matrix; Multiply the first intermediate state transition matrix by the inverse of the preset state transition matrix to obtain the second intermediate state transition matrix; The second intermediate state transition matrix is added to the preset first noise matrix to obtain the initial weight matrix of the target object at the current time.
6. The method according to claim 5, characterized in that, The initial weight matrix is obtained using the following formula: ; in, Let i be the initial weight matrix of the target object at the current time i. The preset state transition matrix, For the target object at the previous moment The corresponding intermediate weight matrix, The inverse of the preset state transition matrix. The preset first noise matrix.
7. The method according to claim 1, characterized in that, The step of obtaining the target weight matrix of the target object at the current time using the initial weight matrix of the target object at the current time includes: The target weight matrix of the target object at the current time can be obtained using the following formula: ; in, Let i be the target weight matrix of the target object at the current time i. The second noise matrix is preset. This is the inverse of the second noise matrix. This is the preset third noise matrix; The step of obtaining the target pose change using the target weight matrix, the first pose change, and the second pose change includes: The target pose change is obtained using the following formula: ; in, The change in the target pose. This represents the second pose change. This represents the change in the first pose.
8. The method according to claim 1, characterized in that, The step of obtaining the target pose of the target object in the camera coordinate system at the current moment based on the target pose change and the initial pose includes: The target pose change is obtained by multiplying the target pose change by the initial pose.
9. An AR device, characterized in that, It includes a processor and a memory, which are connected via a bus; The memory stores a computer program, and the processor is configured to perform the following operations based on the computer program: For an image containing a target object captured by the camera in an AR device at any given time, a target detection algorithm is used to detect the target in the image to obtain the scale of the target object at the current time, wherein the scale of the target object at the current time includes the length, width, and height of the target object; and, Based on the scale of the target object at the current moment and the pose of the camera in the world coordinate system at the current moment, the pose of the camera in the target coordinate system at the current moment is obtained, wherein the target coordinate system is the coordinate system corresponding to the target object; Based on the current pose of the camera in the target coordinate system and the current local pose of the target object, the initial pose of the target object in the camera coordinate system is obtained at the current moment, wherein the local pose of the target is obtained based on the scale of the target; Based on the intermediate weight matrix of the target object at the previous time step, the initial weight matrix of the target object at the current time step is obtained. Using the initial weight matrix of the target object at the current time, the target weight matrix of the target object at the current time is obtained; The target pose change is obtained by using the target weight matrix, the first pose change, and the second pose change. Based on the target pose change and the initial pose, the target pose of the target object in the camera coordinate system at the current moment is obtained, wherein the first pose change is obtained based on the pose of the camera in the world coordinate system, and the second pose change is obtained based on the pose of the camera in the target coordinate system.