A data processing method, apparatus, device, and storage medium
By transforming and processing the first and auxiliary frames of the image acquisition device and applying depth coefficients, the accuracy problem of visual inertial odometry initialization is solved, and the trajectory estimation accuracy in augmented reality, virtual reality and extended reality applications is improved.
Patent Information
- Application Number
- CN202310745416.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-06-21
AI Technical Summary
In existing technologies for augmented reality, virtual reality, and extended reality applications, the initial accuracy of visual inertial odometry is insufficient, which affects the accuracy of subsequent trajectory estimation.
By transforming the first and auxiliary frames of the object captured by the image acquisition device, and combining the depth coefficient and visual reprojection, the spatial position of the object in the acquisition coordinate system is determined, and then the depth of the first frame and the depth coefficient are processed to improve the initialization accuracy.
It improves the accuracy and success rate of visual inertial odometry initialization and enhances the accuracy of subsequent trajectory estimation.
Smart Images

Figure CN116777967B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and more particularly to the field of visual positioning and navigation technology, which can be used in scenarios such as augmented reality, virtual reality, or extended reality. Background Technology
[0002] Applications such as Augmented Reality (AR), Virtual Reality (VR), and Extended Reality (XR) require real-time calculation of the 6 degrees of freedom (DOF) pose of mobile devices (such as mobile phones / headsets).
[0003] The related technology, Visual-Inertial Odometry (VIO), uses algorithms to fuse raw data from two sensors in a mobile device (phone / head-mounted display): camera and inertial measurement unit (IMU), to determine the 6DOF pose and calculate the device's movement trajectory over a period of time.
[0004] The VOI algorithm needs to be initialized when determining the 6dof pose, that is, to determine the d6dof pose at the initial time. The accuracy of the initialization is of great significance to the subsequent trajectory estimation of VIO. Summary of the Invention
[0005] This disclosure provides a data processing method, apparatus, device, and medium.
[0006] According to one aspect of this disclosure, a data processing method is provided, the method comprising:
[0007] The first frame image of the object acquired by the image acquisition device is transformed to obtain the first frame observation position of the object, and the auxiliary frame image of the object acquired by the image acquisition device is transformed to obtain the auxiliary observation position of the object.
[0008] The spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted, to obtain the first spatial position to be predicted;
[0009] Visual reprojection is performed using the first frame observation position and the first frame depth to be predicted to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0010] Based on the first spatial position to be predicted and the second spatial position to be predicted, the first frame depth to be predicted and the depth coefficient to be predicted are processed to obtain the processing result.
[0011] According to another aspect of this disclosure, a data processing apparatus is provided, the apparatus comprising:
[0012] The observation position determination module is used to transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and to transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0013] The first spatial position determination module is used to predict the spatial position of the object in the second acquisition coordinate system using the auxiliary frame image, the auxiliary observation position and the depth coefficient to be predicted, so as to obtain the first spatial position to be predicted.
[0014] The second spatial position determination module is used to perform visual reprojection using the first frame observation position and the first frame depth to be predicted, and to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0015] The processing result determination module is used to process the first frame depth to be predicted and the depth coefficient to be predicted based on the first spatial position to be predicted and the second spatial position to be predicted, and obtain the processing result.
[0016] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the data processing method described in any embodiment of this disclosure.
[0020] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the data processing method described in any embodiment of this disclosure.
[0021] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the data processing method according to any embodiment of this disclosure.
[0022] According to the technology disclosed herein, the accuracy of VIO initialization can be improved.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0024] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0025] Figure 1 This is a flowchart of a data processing method provided according to an embodiment of the present disclosure;
[0026] Figure 2 This is a flowchart of another data processing method provided according to an embodiment of the present disclosure;
[0027] Figure 3 This is a flowchart of yet another data processing method provided according to an embodiment of the present disclosure;
[0028] Figure 4 This is a flowchart of yet another data processing method provided according to an embodiment of the present disclosure;
[0029] Figure 5 This is a flowchart of yet another data processing method provided according to an embodiment of the present disclosure;
[0030] Figure 6 This is a schematic diagram of the structure of a data processing apparatus provided according to an embodiment of the present disclosure;
[0031] Figure 7 This is a block diagram of an electronic device used to implement the data processing method of the embodiments of this disclosure. Detailed Implementation
[0032] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Furthermore, it should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of the first frame image, auxiliary frame image, IMU data, and other related data involved in the technical solution of the present invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0035] Figure 1 This is a flowchart of a data processing method provided according to an embodiment of this disclosure. The method is applicable to how to initialize VIO in AR, VR, or XR scenarios. The method can be executed by a data processing device, which can be implemented in software and / or hardware and can be integrated into an electronic device carrying data processing functions, such as an image acquisition device. Optionally, the image acquisition device can be a camera or a head-mounted display device (head-mounted display); further, the image acquisition device includes an IMU. Figure 1 As shown, the data processing method in this embodiment may include:
[0036] S101, transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0037] In this embodiment, the first frame image refers to the image in which the object first appears. The auxiliary frame images refer to images containing the object that are not included in the first frame image. It should be noted that in this embodiment, images of objects in the environment are acquired continuously over a period of time; the first image to be observed containing the object is the first frame image; images containing the object observed after the first frame image are auxiliary frame images; the number of auxiliary frame images can be one or more.
[0038] The so-called acquisition coordinate system can be the acquisition coordinate system of the image acquisition device, that is, a three-dimensional spatial coordinate system, such as the camera coordinate system. The so-called first acquisition coordinate system refers to the acquisition coordinate system corresponding to the first frame image, that is, the three-dimensional spatial coordinate system; the so-called second acquisition coordinate system refers to the acquisition coordinate system corresponding to the auxiliary frame images, that is, the three-dimensional spatial coordinate system.
[0039] The first-frame observation position refers to the observation position of the object in the first acquisition coordinate system corresponding to the first frame image. The auxiliary observation position refers to the observation position of the object in the second acquisition coordinate system of the auxiliary frame image pair. It should be noted that the observation position refers to the transformation of the two-dimensional pixel position of the object in the image coordinate system to the corresponding three-dimensional plane position in the acquisition coordinate system, that is, the z-coordinate of this three-dimensional plane position, which is the depth, is 1.
[0040] Specifically, objects in the environment can be continuously observed using an image acquisition device to obtain the first frame image and auxiliary frame image of the object. Then, based on the transformation relationship between the first acquisition coordinate system and the image coordinate system, the first frame image of the object can be transformed to obtain the first frame observation position of the object. And based on the transformation relationship between the second acquisition coordinate system and the image coordinate system, the auxiliary frame image of the object can be transformed to obtain the auxiliary observation position of the object.
[0041] S102, using auxiliary frame images, auxiliary observation positions, and depth coefficients to be predicted, the spatial position of the object in the second acquisition coordinate system is predicted to obtain the first spatial position to be predicted.
[0042] In this embodiment, the depth coefficient is used to characterize the difference between the actual depth and the predicted depth of the object in the second acquisition coordinate system; the depth coefficient is unknown, that is, it needs to be predicted.
[0043] The so-called first spatial position refers to the spatial position of an object in the second acquisition coordinate system, which is determined by auxiliary observation positions, etc. It can be represented by (x, y, z), where z represents the depth of the object in the second acquisition coordinate system.
[0044] Specifically, based on a spatial position prediction model, auxiliary frame images, auxiliary observation positions, and depth coefficients to be predicted can be used to predict the spatial position of an object in the second acquisition coordinate system, thus obtaining the object's first predicted spatial position. The spatial position prediction model can be determined based on the relationship between the image coordinate system and the second acquisition coordinate system.
[0045] S103 uses the first frame observation position and the first frame depth to be predicted for visual reprojection to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0046] In this embodiment, the depth of the first frame refers to the depth of the object in the first acquisition coordinate system, which is unknown and needs to be predicted.
[0047] The so-called second spatial position refers to the spatial position of an object in the second acquisition coordinate system determined by visual reprojection based on the observation position of the first frame. It can be represented by (x, y, z), where z represents the depth of the object in the second acquisition coordinate system.
[0048] Specifically, the object can be visually reprojected using the observation position of the first frame and the depth to be predicted in the first frame. That is, the object is projected from the first acquisition coordinate system to the second acquisition coordinate system to obtain the second spatial position to be predicted in the second acquisition coordinate system.
[0049] S104, based on the first spatial position to be predicted and the second spatial position to be predicted, the depth of the first frame to be predicted and the depth coefficient to be predicted are processed to obtain the processing result.
[0050] In this embodiment, the processing result refers to the result obtained by predicting the first frame depth and the depth coefficient to be predicted; optionally, the processing result may include the value of the first frame depth and the value of the depth coefficient. Further, the processing result is used to initialize the visual inertial odometry (VIO) in the image acquisition unit.
[0051] Specifically, the first spatial position to be predicted and the second spatial position to be predicted can be jointly processed to obtain the value of the first frame depth to be predicted and the value of the depth coefficient to be predicted.
[0052] The technical solution provided in this disclosure transforms the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transforms the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. Then, the spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted, to obtain the first spatial position to be predicted. Then, visual reprojection is performed using the first frame observation position and the first frame depth to be predicted to determine the second spatial position to be predicted of the object in the second acquisition coordinate system. Finally, based on the first and second spatial positions to be predicted, the first frame depth and the depth coefficient to be predicted are processed to obtain the processing result. The above technical solution introduces a depth scale coefficient to determine the second spatial position of the object in the second acquisition coordinate system, thereby predicting the first frame depth and the depth coefficient, improving the accuracy of the first frame depth prediction, and thus improving the success rate of VIO initialization.
[0053] Based on the above embodiments, as an optional approach of this disclosure, the first frame depth to be predicted and the depth coefficient to be predicted are processed based on the first spatial position to be predicted and the second spatial position to be predicted. The processing result can be that the first frame depth to be predicted and the depth coefficient to be predicted are processed based on the constraint relationship between the first spatial position to be predicted and the second spatial position to be predicted, so as to obtain the value of the first frame depth and the value of the depth coefficient.
[0054] In this embodiment, the constraint relationship between the first spatial position to be predicted and the second spatial position to be predicted refers to the equality relationship between the first spatial position to be predicted and the second spatial position to be predicted.
[0055] Specifically, based on the constraint relationship between the first spatial position to be predicted and the second spatial position to be predicted, the first spatial position to be predicted and the second spatial position to be predicted can be solved to obtain the value of the depth of the first frame and the value of the depth coefficient.
[0056] It is understandable that predicting the object's first-frame depth in the first-frame acquisition coordinate system and its depth coefficient in the second-frame acquisition coordinate system based on the equality relationship between the two different spatial positions of the object in the second acquisition coordinate system can improve the accuracy of depth prediction.
[0057] Figure 2 This is a flowchart of another data processing method provided according to an embodiment of this disclosure. Based on the above embodiments, this embodiment further optimizes the process of "predicting the spatial position of an object in the second acquisition coordinate system using auxiliary frame images, auxiliary observation positions, and depth coefficients to be predicted, to obtain the first spatial position to be predicted," providing an optional implementation scheme. For example... Figure 2 As shown, the data processing method in this embodiment may include:
[0058] S201, transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0059] S202, the depth of the object in the auxiliary frame image is predicted using the auxiliary frame image and the depth coefficient to be predicted, so as to obtain the auxiliary actual depth to be predicted.
[0060] In this embodiment, the auxiliary actual depth refers to the actual depth of the object in the second acquisition coordinate system.
[0061] One alternative approach is to determine the object's depth in the second acquisition coordinate system based on the correspondence between scene and depth, and the scene in which the auxiliary frame image is located. This depth is then multiplied by a depth coefficient to be predicted to obtain the auxiliary actual depth to be predicted. The correspondence between scene and depth can be determined by those skilled in the art based on a large amount of historical scene data.
[0062] S203, using the auxiliary observation position and the actual depth to be predicted, the spatial position of the object in the second acquisition coordinate system is predicted to obtain the first spatial position to be predicted.
[0063] Specifically, based on spatial position prediction, the spatial position of an object in the second acquisition coordinate system can be predicted using an auxiliary observation position and an auxiliary actual depth to be predicted, thus obtaining the first spatial position to be predicted. For example, based on a spatial position prediction model, the first spatial position to be predicted in the second acquisition coordinate system can be obtained according to the auxiliary observation position and the auxiliary actual depth to be predicted; whereby the spatial position prediction model can be obtained based on a deep learning algorithm. Alternatively, the auxiliary observation position and the auxiliary actual depth to be predicted can be multiplied, and the result of the multiplication can be used as the first spatial position to be predicted.
[0064] S204 uses the first frame observation position and the first frame depth to be predicted for visual reprojection to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0065] S205, based on the first spatial position to be predicted and the second spatial position to be predicted, the depth of the first frame to be predicted and the depth coefficient to be predicted are processed to obtain the processing result.
[0066] The technical solution provided in this disclosure involves transforming the first frame image of an object acquired by an image acquisition device to obtain the first frame observation position of the object, and transforming the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. Then, the depth of the object in the auxiliary frame image is predicted using the auxiliary frame image and the depth coefficient to be predicted to obtain the auxiliary actual depth to be predicted. The spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary observation position and the auxiliary actual depth to be predicted to obtain the first spatial position to be predicted. Then, visual reprojection is performed using the first frame observation position and the first frame depth to be predicted to determine the second spatial position to be predicted of the object in the second acquisition coordinate system. Finally, based on the first spatial position and the second spatial position to be predicted, the first frame depth to be predicted and the depth coefficient to be predicted are processed to obtain the processing result. Compared with related technologies that determine the depth of an object in the second acquisition coordinate system based on empirical values, the above-mentioned technical solution determines the auxiliary actual depth to be predicted by using auxiliary images and depth coefficients, and then determines the first spatial position to be predicted based on the auxiliary observation position and the auxiliary actual depth to be predicted. This makes the depth of the object in the second acquisition coordinate system more reasonable and accurate, thus laying the foundation for determining the depth and depth coefficient of the first frame.
[0067] Based on the above embodiments, as an optional method of this disclosure, the depth of an object in the auxiliary frame image is predicted using an auxiliary frame image and a depth coefficient to be predicted, so as to obtain the auxiliary actual depth to be predicted. This can be achieved by performing depth prediction on the auxiliary frame image to obtain the auxiliary depth of the object in the second acquisition coordinate system; or by using the depth coefficient to be predicted and the auxiliary depth to predict the depth of the object in the auxiliary frame image to obtain the auxiliary actual depth to be predicted.
[0068] The auxiliary depth refers to the depth of the object in the second acquisition coordinate system obtained by performing depth prediction on the auxiliary frame image.
[0069] Specifically, depth prediction can be performed on the auxiliary image based on a depth prediction model to obtain the auxiliary depth of the object in the second acquisition coordinate system. The depth prediction model can be a deep learning network, such as a monocular depth estimation network. Then, the auxiliary depth and the depth coefficient to be predicted can be multiplied to obtain the actual auxiliary depth of the object in the second acquisition coordinate system.
[0070] It is understandable that by performing depth prediction on the auxiliary frame image, the depth of the object in the second acquisition coordinate system can be obtained. Thus, based on the depth coefficient to be predicted, the actual depth of the object in the second acquisition coordinate system to be predicted can be obtained, laying the foundation for determining the first spatial position to be predicted.
[0071] Figure 3This is a flowchart of another data processing method provided according to an embodiment of this disclosure. Based on the above embodiments, this embodiment further optimizes the step of "using the first frame observation position and the first frame depth to be predicted for visual reprojection to determine the second spatial position of the object to be predicted in the second acquisition coordinate system," providing an optional implementation scheme. For example... Figure 3 As shown, the data processing method in this embodiment may include:
[0072] S301, transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0073] S302, using auxiliary frame images, auxiliary observation positions, and depth coefficients to be predicted, predicts the spatial position of the object in the second acquisition coordinate system to obtain the first spatial position to be predicted.
[0074] S303 uses the first frame observation position and the first frame depth to be predicted to predict the spatial position of the object in the first acquisition coordinate system, and obtains the first frame spatial position to be predicted.
[0075] In this embodiment, the spatial position of the first frame refers to the spatial position of the object in the first acquisition coordinate system.
[0076] Specifically, the observation position of the first frame and the depth of the first frame to be predicted can be multiplied, and the result of the multiplication can be used as the spatial position of the object in the first frame to be predicted in the coordinate system of the first acquisition.
[0077] S304. Visual reprojection is performed using the spatial position of the first frame to be predicted to obtain the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0078] An alternative approach is to use the transformation relationship between the first and second acquisition coordinate systems, and then perform visual reprojection using the spatial position of the first frame to be predicted, to obtain the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0079] S305, based on the first spatial position to be predicted and the second spatial position to be predicted, processes the first frame depth to be predicted and the depth coefficient to be predicted to obtain the processing result.
[0080] The technical solution provided in this disclosure transforms the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transforms the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. Then, the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted are used to predict the spatial position of the object in the second acquisition coordinate system to obtain the first spatial position to be predicted. Then, the first frame observation position and the first frame depth to be predicted are used to predict the spatial position of the object in the first acquisition coordinate system to obtain the first frame spatial position to be predicted. Then, the first frame spatial position to be predicted is used for visual reprojection to obtain the second spatial position to be predicted of the object in the second acquisition coordinate system. Finally, based on the first and second spatial positions to be predicted, the first frame depth and the depth coefficient to be predicted are processed to obtain the processing result. The above technical solution determines the second spatial position to be predicted based on the first frame spatial position through visual reprojection, thereby laying the foundation for determining the first frame depth and the depth coefficient to be predicted.
[0081] Based on the above embodiments, as an optional method of this disclosure, visual reprojection using the spatial position to be predicted in the first frame to obtain the second spatial position to be predicted in the second acquisition coordinate system can be performed by: using the spatial position to be predicted in the first frame and the transformation matrix between the inertial measurement unit (IMU) and the image acquisition device to transform the position of the object in the first IMU coordinate system corresponding to the first frame image to obtain the first IMU position; using the first IMU position and the transformation matrix between the first IMU coordinate system and the second IMU coordinate system to transform the position of the object in the second IMU coordinate system to obtain the second IMU position; wherein, the second IMU coordinate system is the coordinate system for the first acquisition of IMU data; using the second IMU position and the transformation matrix between the third IMU coordinate system and the second IMU coordinate system to transform the position of the object in the third IMU coordinate system to obtain the third IMU position; wherein, the third IMU coordinate system is the IMU coordinate system corresponding to the auxiliary frame image; using the third IMU position and the transformation matrix between the IMU and the image acquisition device to transform the spatial position of the object in the second acquisition coordinate system to obtain the second spatial position to be predicted.
[0082] The transformation matrix between the IMU and the image acquisition device characterizes the pose transformation relationship between the IMU coordinate system and the image acquisition coordinate system. It can include a rotation matrix and a translation vector, and can be denoted as R. CI , t CI .
[0083] The first IMU coordinate system refers to the IMU coordinate system corresponding to the first frame image. The first IMU position refers to the position of the object in the first IMU coordinate system. The second IMU coordinate system refers to the IMU coordinate system in which the first IMU data was acquired. The transformation matrix between the first and second IMU coordinate systems is used to characterize the pose transformation relationship between the first and second IMU coordinate systems, and may include a rotation matrix (…). Translation vector The second IMU position refers to the object's position in the second IMU coordinate system. The third IMU coordinate system refers to the IMU coordinate system corresponding to the auxiliary frame image. The transformation matrix between the third and second IMU coordinate systems is used to identify the pose transformation relationship between the first and second IMU coordinate systems, and may include a rotation matrix. Translation vector ).
[0084] Specifically, the spatial position of the object in the first frame to be predicted can be subtracted from the translation vector in the transformation matrix between the IMU and the image acquisition device. Then, the transpose of the rotation matrix in the transformation matrix between the IMU and the image acquisition device and the result of the subtraction are multiplied to obtain the first IMU position of the object in the first IMU coordinate system (denoted as ). ),Right now Next, the translation vector in the transformation matrix between the first and second IMU coordinate systems is subtracted from the first IMU position. Then, the transpose of the rotation matrix in the transformation matrix between the first and second IMU coordinate systems is multiplied by the subtraction result to obtain the second IMU position of the object in the second IMU coordinate system (denoted as ). ),Right now Then, subtract the translation vectors from the transformation matrices between the second and third IMU coordinate systems, and multiply the transpose and subtraction results of the rotation matrices from the transformation matrices between the second and third IMU coordinate systems to obtain the object's third IMU position in the third IMU coordinate system (denoted as ). ),Right now Finally, the third IMU position is multiplied by the rotation matrix in the transformation matrix between the IMU and the image acquisition device. The result of this multiplication is then added to the translation vector in the transformation matrix between the IMU and the image acquisition device. This sum is used as the second spatial position to be predicted. ),Right now It should be noted that j represents the j-th object in the environment. Object j can be captured by images [i, ..., k, ...], that is, image i is the first frame image; image k is the k-th auxiliary frame image.
[0085] Understandably, by using the IMU coordinate system, the spatial position of the first frame is reprojected onto the second acquisition coordinate system to obtain the second spatial position, which provides an alternative method for determining the second spatial position to be predicted.
[0086] Figure 4 This is a flowchart of another data processing method provided according to an embodiment of the present disclosure. Based on the above embodiments, this embodiment further optimizes the process of "transforming the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position, and transforming the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object", and provides an optional implementation scheme.
[0087] like Figure 4 As shown, the data processing method in this embodiment may include:
[0088] S401 uses the intrinsic parameter matrix of the image acquisition device to transform the first frame image acquired by the image acquisition device to obtain the first frame observation position of the object, and transforms the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0089] In this embodiment, the intrinsic parameter matrix of the image acquisition device refers to the projection relationship used to transform three-dimensional spatial coordinates to two-dimensional image coordinates.
[0090] Specifically, the intrinsic parameter matrix of the image acquisition device can be used to transform the pixel positions of the object's first frame image into the corresponding first acquisition coordinate system to obtain the object's first frame observation position. Then, the intrinsic parameter matrix of the image acquisition device can be used to transform the pixel positions of the object's auxiliary frame images into the corresponding second acquisition coordinate system to obtain the object's auxiliary observation position.
[0091] S402 uses auxiliary frame images, auxiliary observation positions, and depth coefficients to be predicted to predict the spatial position of the object in the second acquisition coordinate system, thus obtaining the first spatial position to be predicted.
[0092] S403 uses the first frame observation position and the first frame depth to be predicted for visual reprojection to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0093] S404, based on the first spatial position to be predicted and the second spatial position to be predicted, processes the first frame depth to be predicted and the depth coefficient to be predicted to obtain the processing result.
[0094] The technical solution provided in this disclosure transforms the first frame image acquired by the image acquisition device using its intrinsic parameter matrix to obtain the first frame observation position of the object. It then transforms the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. Next, it uses the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted to predict the spatial position of the object in the second acquisition coordinate system, obtaining the first spatial position to be predicted. Then, it performs visual reprojection using the first frame observation position and the first frame depth to be predicted to determine the second spatial position to be predicted of the object in the second acquisition coordinate system. Finally, based on the first and second spatial positions to be predicted, it processes the first frame depth and the depth coefficient to be predicted to obtain the processing result. This technical solution transforms the image into the acquisition coordinate system, which facilitates the subsequent determination of the spatial position, thus providing a basis for the first frame depth and the depth coefficient to be predicted.
[0095] Based on the above embodiments, as an optional method of this disclosure, transforming the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transforming the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object, can also be done by extracting features from the first frame image and the auxiliary frame image respectively to obtain the feature points of the object; transforming the first frame pixel position of the feature points in the first frame image to obtain the first frame observation position of the feature points, and transforming the auxiliary pixel position of the feature points in the auxiliary frame image to obtain the auxiliary observation position of the feature points.
[0096] Specifically, feature extraction can be performed on the first frame image to obtain the object's feature points, thereby determining the pixel positions of these feature points in the first frame image. Similarly, feature extraction can be performed on the auxiliary frame image to obtain the object's feature points, thereby determining the auxiliary pixel positions of these feature points in the auxiliary frame image. Furthermore, an optical flow tracing algorithm can be used to determine the auxiliary pixel positions of the object's feature points based on the pixel positions in the first frame.
[0097] Then, based on the intrinsic parameter matrix of the image acquisition device, the first frame pixel position of the feature point in the first frame image can be transformed to the first acquisition coordinate system to obtain the first frame observation position of the feature point in the first acquisition coordinate system; and based on the intrinsic parameter matrix of the image acquisition device, the auxiliary pixel position of the feature point in the auxiliary frame image can be transformed to the second acquisition coordinate system to obtain the auxiliary observation position of the feature point in the second acquisition coordinate system.
[0098] It is understandable that in this embodiment, the observation position of the object's feature points is determined, and the observation position of the object is determined at a finer granularity, which further makes the subsequent processing of the first frame depth to be predicted and the depth coefficient to be predicted more accurate.
[0099] Figure 5 This is a flowchart of another data processing method provided according to an embodiment of this disclosure. Based on the above embodiments, this embodiment further optimizes the process of "processing the first frame depth and the depth coefficient to be predicted based on the first spatial position to be predicted and the second spatial position to be predicted, to obtain the processing result," providing an optional implementation scheme. For example... Figure 5 As shown, the data processing method in this embodiment may include:
[0100] S501, transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0101] S502 uses auxiliary frame images, auxiliary observation positions, and depth coefficients to be predicted to predict the spatial position of the object in the second acquisition coordinate system, thus obtaining the first spatial position to be predicted.
[0102] S503 uses the first frame observation position and the first frame depth to be predicted for visual reprojection to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0103] S504. Using the first spatial location to be predicted and the second spatial location to be predicted, construct the first constraint relationship.
[0104] In this embodiment, the first constraint relationship refers to the constraint relationship established based on the first spatial position and the second spatial position.
[0105] Specifically, since the first spatial position to be predicted and the second spatial position to be predicted are the spatial positions of the object in the second acquisition coordinate system, the error between the first spatial position and the second spatial position should be close to 0, that is, the first spatial position to be predicted is approximately equal to the second spatial position to be predicted. Based on the approximate equality between the first spatial position to be predicted and the second spatial position to be predicted, the first constraint relationship is constructed.
[0106] S505 uses a preset depth value, auxiliary observation position, and second spatial position to be predicted to construct a second constraint relationship.
[0107] In this embodiment, the second constraint relationship refers to the constraint relationship established based on the preset depth value, the auxiliary observation position, and the second spatial position to be predicted.
[0108] Specifically, based on the depth prediction model, a preset depth value and auxiliary observation position can be used to obtain the third spatial position of the object in the second acquisition coordinate system. Then, the equality relationship between the third spatial position and the second spatial position to be predicted can be used to construct the second constraint relationship.
[0109] S506 uses IMU data acquired by the IMU and the initial variables to be predicted in the IMU coordinate system to construct a third constraint relationship.
[0110] In this embodiment, IMU data refers to the data of the accelerometer and gyroscope in the IMU at every moment; optionally, it may include, but is not limited to, the measured values of the accelerometer and gyroscope, the actual motion values of the accelerometer and gyroscope, the bias of the accelerometer and gyroscope, and the Gaussian noise of the accelerometer and gyroscope.
[0111] The initial variables to be predicted include the position and velocity of the image acquisition device in the IMU coordinate system, as well as the direction of gravity in the IMU coordinate system.
[0112] The so-called third constraint relationship refers to the constraint relationship established based on IMU-related data.
[0113] Specifically, based on the IMU measurement model and the IMU integration model, the predicted position and velocity values in the IMU can be obtained by integration prediction based on the IMU data. Then, based on the constraint relationship between position and velocity, and the equality relationship of gravity at different times, the actual position and velocity of the image acquisition device in the IMU coordinate system can be determined according to the position to be predicted, the velocity to be predicted, and the direction of gravity in the IMU coordinate system. Furthermore, a third constraint relationship is constructed by using the position prediction value and the actual position to be predicted, the velocity prediction value and the actual velocity to be predicted, and the equality relationship of gravity at different times.
[0114] It should be noted that the IMU measurement model and IMU integration model are not specifically limited in this disclosure, and can be selected by those skilled in the art according to actual needs.
[0115] S507, combining the first constraint relationship, the second constraint relationship and the third constraint relationship for processing, to obtain the processing result.
[0116] In this embodiment, the processing result may also include the value of the image acquisition device's position to be predicted in the IMU coordinate system, the value of the velocity to be predicted, and the direction of gravity in the IMU coordinate system.
[0117] Specifically, by combining the first constraint relationship, the second constraint relationship, and the third preset relationship, the depth of the first frame to be predicted, the depth coefficient to be predicted, the value of the position to be predicted, the value of the velocity to be predicted, and the direction of gravity in the IMU coordinate system are jointly solved to obtain the processing result.
[0118] The technical solution provided in this disclosure transforms the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transforms the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. Then, the spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted to obtain the first spatial position to be predicted. Then, the first frame observation position and the first frame depth to be predicted are used for visual reprojection to determine the second spatial position to be predicted of the object in the second acquisition coordinate system. Then, the first spatial position and the second spatial position to be predicted are used to construct a first constraint relationship, and the second constraint relationship is constructed using a preset depth value, the auxiliary observation position, and the second spatial position to be predicted. At the same time, the IMU data acquired by the IMU and the initialization variables to be predicted in the IMU coordinate system are used to construct a third constraint relationship. Finally, the first constraint relationship, the second constraint relationship, and the third constraint relationship are combined for processing to obtain the processing result. The above technical solution, by combining multiple constraint relationships to determine the first frame depth and depth coefficient to be predicted, can improve the accuracy of determining the first frame depth and depth coefficient, thereby improving the initialization efficiency of VIO.
[0119] Based on the above embodiments, as an optional approach of this disclosure, constructing the second constraint relationship using a preset depth value, an auxiliary observation position, and a second spatial position to be predicted can be achieved by transforming the spatial position of the object in the second acquisition coordinate system using the auxiliary observation position and the preset depth value to obtain a third spatial position; and constructing the second constraint relationship based on the third spatial position and the second spatial position to be predicted.
[0120] The third spatial position refers to the spatial position of the object in the second acquisition coordinate system, which is determined based on the auxiliary observation position and the preset depth value.
[0121] Specifically, the auxiliary observation position can be multiplied by the preset depth value, and the result of the multiplication can be used as the third spatial position of the object in the second acquisition coordinate system. Then, based on the approximate relationship between the third spatial position and the second spatial position to be predicted, a second constraint relationship can be constructed.
[0122] Understandably, a second constraint relationship has been provided to make the determination of the first frame depth and depth coefficient more accurate, thus providing a foundation for the successful initialization of VIO.
[0123] Figure 6This is a schematic diagram of a data processing device according to an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the initialization of VIO in AR, VR, or XR scenarios. This method can be executed by a data processing device, which can be implemented in software and / or hardware and can be integrated into an electronic device that carries data processing functions, such as an image acquisition device. Optionally, the image acquisition device may include a camera or a head-mounted display device (head-mounted display), etc. Figure 6 As shown, the data processing apparatus 600 of this embodiment includes:
[0124] The observation position determination module 601 is used to transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and to transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object.
[0125] The first spatial position determination module 602 is used to predict the spatial position of an object in the second acquisition coordinate system using an auxiliary frame image, an auxiliary observation position, and a depth coefficient to be predicted, so as to obtain the first spatial position to be predicted.
[0126] The second spatial position determination module 603 is used to perform visual reprojection using the first frame observation position and the first frame depth to be predicted, and to determine the second spatial position of the object to be predicted in the second acquisition coordinate system.
[0127] The processing result determination module 604 is used to process the first frame depth to be predicted and the depth coefficient to be predicted based on the first spatial position to be predicted and the second spatial position to be predicted, and obtain the processing result.
[0128] The technical solution provided in this disclosure transforms the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transforms the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. Then, the spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted, to obtain the first spatial position to be predicted. Then, visual reprojection is performed using the first frame observation position and the first frame depth to be predicted to determine the second spatial position to be predicted of the object in the second acquisition coordinate system. Finally, based on the first and second spatial positions to be predicted, the first frame depth and the depth coefficient to be predicted are processed to obtain the processing result. The above technical solution introduces a depth scale coefficient to determine the second spatial position of the object in the second acquisition coordinate system, thereby predicting the first frame depth and the depth coefficient, improving the accuracy of the first frame depth prediction, and thus improving the success rate of VIO initialization.
[0129] Furthermore, the processing results are used to initialize the visual inertial odometry (VIO) in the image acquisition unit.
[0130] Furthermore, the first spatial location determination module 602 includes:
[0131] The auxiliary actual depth determination unit is used to predict the depth of an object in the auxiliary frame image using the auxiliary frame image and the depth coefficient to be predicted, so as to obtain the auxiliary actual depth to be predicted.
[0132] The first spatial position determination unit is used to predict the spatial position of an object in the second acquisition coordinate system by using the auxiliary observation position and the auxiliary actual depth to be predicted, so as to obtain the first spatial position to be predicted.
[0133] Furthermore, the auxiliary actual depth determination unit is specifically used for:
[0134] Depth prediction is performed on the auxiliary frame image to obtain the auxiliary depth of the object in the second acquisition coordinate system;
[0135] The depth of the object in the auxiliary frame image is predicted using the depth coefficient to be predicted and the auxiliary actual depth to be predicted.
[0136] Furthermore, the second spatial location determination module 603 includes:
[0137] The first frame spatial position determination unit is used to predict the spatial position of an object in the first acquisition coordinate system using the first frame observation position and the first frame depth to be predicted, so as to obtain the first frame spatial position to be predicted.
[0138] The second spatial position determination unit is used to perform visual reprojection using the spatial position to be predicted in the first frame to obtain the second spatial position to be predicted of the object in the second acquisition coordinate system.
[0139] Furthermore, the second spatial location determination unit is specifically used for:
[0140] Using the spatial position to be predicted in the first frame, and the transformation matrix between the inertial measurement unit (IMU) and the image acquisition unit, the position of the object in the first IMU coordinate system corresponding to the first frame image is transformed to obtain the first IMU position;
[0141] The object's position in the second IMU coordinate system is transformed using the first IMU position and the transformation matrix between the first and second IMU coordinate systems to obtain the second IMU position; wherein, the second IMU coordinate system is the coordinate system in which the IMU data was first acquired;
[0142] The object's position in the third IMU coordinate system is transformed using the second IMU position and the transformation matrix between the third IMU coordinate system and the second IMU coordinate system to obtain the third IMU position; where the third IMU coordinate system is the IMU coordinate system corresponding to the auxiliary frame image;
[0143] The spatial position of the object in the second acquisition coordinate system is transformed using the third IMU position and the transformation matrix between the IMU and the image acquisition device to obtain the second spatial position to be predicted.
[0144] Furthermore, the observation location determination module 601 is used for:
[0145] The first frame observation position of the object is obtained by transforming the first frame image acquired by the image acquisition device using the intrinsic parameter matrix of the image acquisition device, and the auxiliary frame image of the object acquired by the image acquisition device is transformed to obtain the auxiliary observation position of the object.
[0146] Furthermore, the processing result determination module 604 is used for:
[0147] Based on the constraint relationship between the first spatial position to be predicted and the second spatial position to be predicted, the first frame depth to be predicted and the depth coefficient to be predicted are processed to obtain the value of the first frame depth and the value of the depth coefficient.
[0148] Furthermore, the processing result determination module 604 includes:
[0149] The first constraint relationship determination unit is used to construct the first constraint relationship using the first spatial location to be predicted and the second spatial location to be predicted.
[0150] The second constraint relationship determination unit is used to construct the second constraint relationship using a preset depth value, auxiliary observation position, and the second spatial position to be predicted;
[0151] The third constraint relationship determination unit is used to construct the third constraint relationship using IMU data acquired by the IMU and the initialization variables to be predicted in the IMU coordinate system.
[0152] The processing result determination unit is used to combine the first constraint relationship, the second constraint relationship, and the third constraint relationship to process and obtain the processing result.
[0153] Furthermore, the second constraint relationship determination unit is specifically used for:
[0154] The spatial position of the object in the second acquisition coordinate system is transformed using the auxiliary observation position and the preset depth value to obtain the third spatial position;
[0155] Based on the third spatial location and the second spatial location to be predicted, construct the second constraint relationship.
[0156] Furthermore, the observation location determination module 601 is also used for:
[0157] Feature extraction is performed on the first frame image and the auxiliary frame image respectively to obtain the feature points of the object;
[0158] The first-frame observation position of the feature point is obtained by transforming the first-frame pixel position of the feature point in the first frame image, and the auxiliary observation position of the feature point is obtained by transforming the auxiliary pixel position of the feature point in the auxiliary frame image.
[0159] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0160] Figure 7 This is a block diagram of an electronic device used to implement the data processing method of the embodiments of this disclosure; Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0161] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0162] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0163] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform data processing methods by any other suitable means (e.g., by means of firmware).
[0164] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0165] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0166] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0167] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0168] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0169] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0170] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0171] Cloud computing refers to a technology system that enables access to a shared pool of physical or virtual resources via a network. These resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for applications such as artificial intelligence and blockchain, as well as for model training.
[0172] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0173] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A data processing method, comprising: The first frame image of the object acquired by the image acquisition device is transformed to obtain the first frame observation position of the object, and the auxiliary frame image of the object acquired by the image acquisition device is transformed to obtain the auxiliary observation position of the object. The spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted, to obtain the first spatial position to be predicted; The second acquisition coordinate system is the acquisition coordinate system corresponding to the auxiliary frame image; Visual reprojection is performed using the first frame observation position and the first frame depth to be predicted to determine the second spatial position of the object to be predicted in the second acquisition coordinate system. Based on the first spatial position to be predicted and the second spatial position to be predicted, the first frame depth to be predicted and the depth coefficient to be predicted are processed to obtain a processing result; wherein, the first frame depth refers to the depth of the object in the first acquisition coordinate system; the first acquisition coordinate system refers to the acquisition coordinate system corresponding to the first frame image; the depth coefficient is used to characterize the difference between the actual depth and the predicted depth of the object in the second acquisition coordinate system; the processing result is used to initialize the visual inertial odometry (VIO) in the image acquisition device.
2. The method according to claim 1, wherein, The step of predicting the spatial position of an object in the second acquisition coordinate system using the auxiliary frame image, the auxiliary observation position, and the depth coefficient to be predicted, to obtain the first spatial position to be predicted, includes: The depth of the object in the auxiliary frame image is predicted using the auxiliary frame image and the depth coefficient to be predicted, so as to obtain the auxiliary actual depth to be predicted. The spatial position of the object in the second acquisition coordinate system is predicted using the auxiliary observation position and the actual depth to be predicted, thus obtaining the first spatial position to be predicted.
3. The method according to claim 2, wherein, The step of predicting the depth of an object in the auxiliary frame image using the auxiliary frame image and the depth coefficient to be predicted, to obtain the auxiliary actual depth to be predicted, includes: Depth prediction is performed on the auxiliary frame image to obtain the auxiliary depth of the object in the second acquisition coordinate system; The depth of the object in the auxiliary frame image is predicted using the depth coefficient to be predicted and the auxiliary depth to be predicted.
4. The method according to claim 1, wherein, The step of visual reprojection using the first frame observation position and the first frame depth to be predicted to determine the second spatial position of the object in the second acquisition coordinate system includes: The spatial position of the object in the first acquisition coordinate system is predicted using the first frame observation position and the first frame depth to be predicted, thus obtaining the first frame spatial position to be predicted. Visual reprojection is performed using the first frame spatial position to be predicted to obtain the second spatial position to be predicted of the object in the second acquisition coordinate system.
5. The method according to claim 4, wherein, The step of visually reprojecting the first frame spatial position to be predicted to obtain the second spatial position to be predicted of the object in the second acquisition coordinate system includes: Using the spatial position to be predicted in the first frame, and the transformation matrix between the inertial measurement unit (IMU) and the image acquisition unit, the position of the object in the first IMU coordinate system corresponding to the first frame image is transformed to obtain the first IMU position; The object's position in the second IMU coordinate system is transformed using the first IMU position and the transformation matrix between the first and second IMU coordinate systems to obtain the second IMU position; wherein, the second IMU coordinate system is the coordinate system in which the IMU data was first acquired; The object's position in the third IMU coordinate system is transformed using the second IMU position and the transformation matrix between the third IMU coordinate system and the second IMU coordinate system to obtain the third IMU position; wherein, the third IMU coordinate system is the IMU coordinate system corresponding to the auxiliary frame image; The spatial position of the object in the second acquisition coordinate system is transformed using the third IMU position and the transformation matrix between the IMU and the image acquisition device to obtain the second spatial position to be predicted.
6. The method according to any one of claims 1-5, wherein, The process of transforming the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position, and transforming the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object, includes: The first frame observation position of the object is obtained by transforming the first frame image acquired by the image acquisition device using the intrinsic parameter matrix of the image acquisition device, and the auxiliary frame image of the object acquired by the image acquisition device is transformed to obtain the auxiliary observation position of the object.
7. The method according to any one of claims 1-5, wherein, The process involves processing the first frame depth and the depth coefficient to be predicted based on the first spatial location and the second spatial location to be predicted, to obtain the processing result, including: Based on the constraint relationship between the first spatial position to be predicted and the second spatial position to be predicted, the first frame depth to be predicted and the depth coefficient to be predicted are processed to obtain the value of the first frame depth and the value of the depth coefficient.
8. The method according to any one of claims 1-5, wherein, The process involves processing the first frame depth and the depth coefficient to be predicted based on the first spatial location and the second spatial location to be predicted, to obtain the processing result, including: A first constraint relationship is constructed using the first spatial location to be predicted and the second spatial location to be predicted. A second constraint relationship is constructed using a preset depth value, the auxiliary observation position, and the second spatial position to be predicted; A third constraint relationship is constructed using IMU data acquired by the IMU and the initial variables to be predicted in the IMU coordinate system; The processing results are obtained by combining the first constraint relationship, the second constraint relationship, and the third constraint relationship.
9. The method according to claim 8, wherein, The second constraint relationship is constructed by using a preset depth value, the auxiliary observation position, and the second spatial position to be predicted, including: The spatial position of the object in the second acquisition coordinate system is transformed using the auxiliary observation position and the preset depth value to obtain the third spatial position; Based on the third spatial location and the second spatial location to be predicted, a second constraint relationship is constructed.
10. The method according to any one of claims 1-5, wherein, The process of transforming the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and transforming the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object, includes: Feature extraction is performed on the first frame image and the auxiliary frame image respectively to obtain the feature points of the object; The first-frame observation position of the feature point is obtained by transforming the first-frame pixel position of the feature point in the first frame image, and the auxiliary observation position of the feature point is obtained by transforming the auxiliary pixel position of the feature point in the auxiliary frame image.
11. A data processing apparatus, comprising: The observation position determination module is used to transform the first frame image of the object acquired by the image acquisition device to obtain the first frame observation position of the object, and to transform the auxiliary frame image of the object acquired by the image acquisition device to obtain the auxiliary observation position of the object. The first spatial position determination module is used to predict the spatial position of the object in the second acquisition coordinate system using the auxiliary frame image, the auxiliary observation position and the depth coefficient to be predicted, so as to obtain the first spatial position to be predicted. The second acquisition coordinate system is the acquisition coordinate system corresponding to the auxiliary frame image; The second spatial position determination module is used to perform visual reprojection using the first frame observation position and the first frame depth to be predicted, and to determine the second spatial position of the object to be predicted in the second acquisition coordinate system. The processing result determination module is used to process the first frame depth to be predicted and the depth coefficient to be predicted based on the first spatial position to be predicted and the second spatial position to be predicted, to obtain a processing result; wherein, the first frame depth refers to the depth of the object in the first acquisition coordinate system; the first acquisition coordinate system refers to the acquisition coordinate system corresponding to the first frame image; the depth coefficient is used to characterize the difference between the actual depth and the predicted depth of the object in the second acquisition coordinate system; the processing result is used to initialize the visual inertial odometry (VIO) in the image acquisition device.
12. The apparatus according to claim 11, wherein, The first spatial location determination module includes: An auxiliary actual depth determination unit is used to predict the depth of an object in the auxiliary frame image using the auxiliary frame image and the depth coefficient to be predicted, so as to obtain the auxiliary actual depth to be predicted. The first spatial position determination unit is used to predict the spatial position of an object in the second acquisition coordinate system using the auxiliary observation position and the auxiliary actual depth to be predicted, so as to obtain the first spatial position to be predicted.
13. The apparatus according to claim 12, wherein, The auxiliary actual depth determination unit is specifically used for: Depth prediction is performed on the auxiliary frame image to obtain the auxiliary depth of the object in the second acquisition coordinate system; The depth of the object in the auxiliary frame image is predicted using the depth coefficient to be predicted and the auxiliary depth to be predicted.
14. The apparatus according to claim 11, wherein, The second spatial location determination module includes: The first frame spatial position determination unit is used to predict the spatial position of the object in the first acquisition coordinate system using the first frame observation position and the first frame depth to be predicted, so as to obtain the first frame spatial position to be predicted. The second spatial position determination unit is used to perform visual reprojection using the spatial position to be predicted in the first frame to obtain the second spatial position to be predicted of the object in the second acquisition coordinate system.
15. The apparatus according to claim 14, wherein, The second spatial position determination unit is specifically used for: Using the spatial position to be predicted in the first frame, and the transformation matrix between the inertial measurement unit (IMU) and the image acquisition unit, the position of the object in the first IMU coordinate system corresponding to the first frame image is transformed to obtain the first IMU position; The object's position in the second IMU coordinate system is transformed using the first IMU position and the transformation matrix between the first and second IMU coordinate systems to obtain the second IMU position; wherein, the second IMU coordinate system is the coordinate system in which the IMU data was first acquired; The object's position in the third IMU coordinate system is transformed using the second IMU position and the transformation matrix between the third IMU coordinate system and the second IMU coordinate system to obtain the third IMU position; wherein, the third IMU coordinate system is the IMU coordinate system corresponding to the auxiliary frame image; The spatial position of the object in the second acquisition coordinate system is transformed using the third IMU position and the transformation matrix between the IMU and the image acquisition device to obtain the second spatial position to be predicted.
16. The apparatus according to any one of claims 11-15, wherein, The observation location determination module is used for: The first frame observation position of the object is obtained by transforming the first frame image acquired by the image acquisition device using the intrinsic parameter matrix of the image acquisition device, and the auxiliary frame image of the object acquired by the image acquisition device is transformed to obtain the auxiliary observation position of the object.
17. The apparatus according to any one of claims 11-15, wherein, The processing result determination module is used for: Based on the constraint relationship between the first spatial position to be predicted and the second spatial position to be predicted, the first frame depth to be predicted and the depth coefficient to be predicted are processed to obtain the value of the first frame depth and the value of the depth coefficient.
18. The apparatus according to any one of claims 11-15, wherein, The processing result determination module includes: The first constraint relationship determination unit is used to construct the first constraint relationship using the first spatial location to be predicted and the second spatial location to be predicted. The second constraint relationship determination unit is used to construct a second constraint relationship using a preset depth value, the auxiliary observation position, and the second spatial position to be predicted; The third constraint relationship determination unit is used to construct the third constraint relationship using IMU data acquired by the IMU and the initialization variables to be predicted in the IMU coordinate system. The processing result determination unit is used to process the first constraint relationship, the second constraint relationship and the third constraint relationship to obtain the processing result.
19. The apparatus according to claim 18, wherein, The second constraint relationship determination unit is specifically used for: The spatial position of the object in the second acquisition coordinate system is transformed using the auxiliary observation position and the preset depth value to obtain the third spatial position; Based on the third spatial location and the second spatial location to be predicted, a second constraint relationship is constructed.
20. The apparatus according to any one of claims 11-15, wherein, The observation location determination module is also used for: Feature extraction is performed on the first frame image and the auxiliary frame image respectively to obtain the feature points of the object; The first-frame observation position of the feature point is obtained by transforming the first-frame pixel position of the feature point in the first frame image, and the auxiliary observation position of the feature point is obtained by transforming the auxiliary pixel position of the feature point in the auxiliary frame image.
21. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the data processing method according to any one of claims 1-10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the data processing method according to any one of claims 1-10.
23. A computer program product comprising a computer program that, when executed by a processor, implements the data processing method according to any one of claims 1-10.
Citation Information
Patent Citations
Spatial joint calibration method of underwater camera-IMU-depthometer
CN115330852A
KR20220153535A