Model training methods and devices
By using image triples as training samples, a matching keypoint localization model is trained based on image feature information. This solves the problem of high-cost dense depth information requirements in existing technologies, achieves accuracy and cost-effectiveness in model training, and improves the stability of the SLAM algorithm and the localization and tracking capabilities of AR/VR applications.
Patent Information
- Application Number
- CN202311185409.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-13
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-09-13
AI Technical Summary
In existing technologies, electronic devices have a high demand for dense depth information when training models, which increases training costs and makes it difficult to effectively train feature point extraction networks with large amounts of data.
Image triples are used as training samples. By acquiring three images from different views, similarity information is determined based on the image feature information of key points. A loss function is constructed to train the matching key point localization model, reducing the requirements for acquiring dense depth images.
While reducing training costs, it ensures the accuracy and generalization performance of the model, provides more reliable image feature points, improves the stability of the SLAM algorithm, and provides more stable localization and tracking capabilities for AR/VR applications.
Smart Images

Figure CN117197616B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a model training method and apparatus. Background Technology
[0002] Currently, electronic devices can obtain matching keypoints in images through neural networks, thereby enabling localization and mapping based on these keypoints. In related technologies, electronic devices can train matching keypoints on image pairs with accompanying dense depth and camera pose. Specifically, electronic devices can project keypoints from one image onto another image using dense depth and camera pose, thereby determining the ground truth matching relationship of keypoints in the image pair, and then obtaining matching keypoints based on this ground truth matching relationship.
[0003] However, the above methods have high requirements for data quality, that is, they require accurate and dense depth information. Usually, the cost of obtaining accurate and dense depth information is extremely high, which makes the cost of training models on electronic devices high. Summary of the Invention
[0004] The purpose of this application is to provide a model training method and apparatus that can ensure the accuracy of the model without increasing training costs.
[0005] In a first aspect, embodiments of this application provide a model training method, which includes: acquiring a first training sample, the first training sample including a first image triplet, each image in the first image triplet corresponding to a shooting viewpoint, and each image in the first image triplet including the same shooting object; determining similarity information between the same matching keypoint between any two images in the first image triplet based on image feature information of key points in each image in the first image triplet; and training a matching keypoint localization model based on the similarity information; wherein the matching keypoint is the keypoint corresponding to the shooting object in each image.
[0006] Secondly, embodiments of this application provide a model training apparatus, comprising: an acquisition module, a determination module, and a training module. The acquisition module acquires a first training sample, which includes a first image triplet, where each image in the first image triplet corresponds to a shooting viewpoint, and each image in the first image triplet includes the same shooting object. The determination module determines similarity information between any two images in the first image triplet based on image feature information of key points in each image of the first image triplet, for the same matching key point. The training module trains a matching key point localization model based on the similarity information; wherein the key point is the key point corresponding to the shooting object in each image.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0011] In this embodiment, the electronic device can acquire a first training sample, which includes a first image triplet. Each image in the first image triplet corresponds to a shooting viewpoint, and each image in the first image triplet includes the same shooting object. Based on the image feature information of key points in each image in the first image triplet, the similarity information between the same matching key point between any two images in the first image triplet is determined. Based on the similarity information, the matching key point localization model is trained. The matching key point is the key point corresponding to the shooting object in each image. In this solution, since the image triplet in the training sample includes three images with different views, when the electronic device uses the image triplet to train the feature point extraction model, it can determine the matching key point based on the three images with different views. Therefore, this solution does not require the acquisition of costly dense depth images, and it detects the matching relationship between key points based on the three images with different views in the first image triplet, thereby achieving the training of the feature point extraction model. That is, this invention reduces the requirements for the acquired data when training the feature point extraction model while ensuring the accuracy of the model. Attached Figure Description
[0012] Figure 1 This is a polar diagram in related technologies;
[0013] Figure 2 This is one of the flowcharts of a model training method provided in the embodiments of this application;
[0014] Figure 3This is the second flowchart of a model training method provided in the embodiments of this application;
[0015] Figure 4 This is a schematic diagram of a camera view cone provided in an embodiment of this application;
[0016] Figure 5 This is the third flowchart of a model training method provided in the embodiments of this application;
[0017] Figure 6 This is one of the schematic diagrams of an epipolar projection provided in the embodiments of this application;
[0018] Figure 7 This is a second schematic diagram of an epipolar projection provided in an embodiment of this application;
[0019] Figure 8 This is the fourth flowchart of a model training method provided in the embodiments of this application;
[0020] Figure 9 This is one of the structural schematic diagrams of a model training device provided in the embodiments of this application;
[0021] Figure 10 This is a second schematic diagram of the structure of a model training device provided in an embodiment of this application;
[0022] Figure 11 This is the third schematic diagram of a model training device provided in the embodiments of this application;
[0023] Figure 12 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;
[0024] Figure 13 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0028] The following explains the technical terms involved in the model training method, apparatus, electronic device, and storage medium provided in the embodiments of this application.
[0029] Augmented Reality (AR): A technology that overlays virtual information onto the real world. Users observe real-world scenes through display devices such as mobile phones and tablets, and the system identifies the real-world scene and overlays virtual information onto it.
[0030] Virtual Reality (VR): A technology that uses special equipment to create a completely new virtual environment, immersing users in a virtual world. For example, by wearing VR headsets or similar devices, users can experience immersive virtual scenes.
[0031] Structure from Motion (SFM): Recovering the structural information of a scene from photographs. For example, an electronic device takes a large number of photos of the same scene from different positions and angles using a camera, and then uses the poses from these photos to recover the three-dimensional structural information of the scene.
[0032] Simultaneous localization and mapping (SLAM): Simultaneous localization and map building, or concurrent mapping and localization. For example, when a robot is placed in an unknown location in an unknown environment, SLAM allows the robot to gradually map the environment as it moves.
[0033] Image feature points: This includes two aspects: keypoints and descriptors of keypoints. Keypoints are points that remain stable in an image under varying lighting, viewing angles, and other conditions (in this invention, keypoints refer to image keypoints in a broad sense). Descriptors are feature descriptions of keypoints.
[0034] Pose: Position and attitude are used to describe the three-dimensional position and orientation of an object in space. Position refers to spatial location, which includes three quantities: x, y, and z; attitude describes orientation information, including rotation angle information in three rotational directions: pitch, roll, and yaw.
[0035] Epipolar lines: When the camera is in a relative pose, key points in one image are projected onto another image and appear as a straight line.
[0036] For example, such as Figure 1 As shown, C and C' are the centers of the two cameras, X is a point in space, and the projections of X onto the corresponding image planes of C and C' are X' and X'', respectively. The focal points e and e' of the line connecting C and C' to the image plane are called the poles, and l' is called the epipolar line.
[0037] The model training method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0038] Currently, with the development of electronic devices and communication technologies, electronic devices are increasingly equipped with more and more functions. For example, electronic devices can perform localization and mapping using image feature points. These image feature points are points that can be stably extracted and described in images from different viewpoints and under different lighting conditions, and they have wide applications in the field of computer vision. Among them, visual SLAM technology is based on stable image feature points to calculate the device's own posture, and it is the underlying algorithm for electronic devices or AR / VR head-mounted displays to achieve self-localization.
[0039] Currently, image feature point extraction algorithms based on human experience and designed manually are widely used in SLAM algorithms. However, these methods are usually relatively simple and cannot effectively extract stable feature points from images taken in complex scenes.
[0040] In recent years, feature point extraction algorithms based on neural networks have been widely researched and applied. Compared with manually designed image feature points based on human experience, neural network-based algorithms learn feature points from a large number of images, exhibiting better stability and robustness. This significantly improves the accuracy of image matching, thereby ensuring the stability of the SLAM algorithm. Consequently, it provides more stable localization and tracking capabilities for AR / VR applications.
[0041] Since neural network-based feature point extraction algorithms are data-driven, they first require acquiring a large number of image pairs with correct matching relationships. These matching relationships can then be used for supervised training of the neural network.
[0042] Currently, there are generally three ways to train feature point extraction networks.
[0043] Method 1: Feature point training using homography image pairs. In this method, an image is homography-transformed to obtain a transformed image, thus forming an image pair. The matching relationship between key points in the two images can be determined using the homography matrix generated during the transformation. However, image homography transformation is only a simple two-dimensional planar transformation and cannot simulate the three-dimensional spatial relationship between images taken from different angles.
[0044] Method 2: Feature point training using image pairs with accompanying dense depth and camera pose. This method uses dense depth and camera pose to project keypoints from one image onto another, thereby determining the ground truth matching relationship of keypoints in the image pair. However, this method has high requirements for data quality, i.e., it requires accurate and dense depth information, but obtaining accurate and dense depth information is extremely costly.
[0045] Method 3: Use image pairs with accompanying camera poses for feature point training. This method abandons image depth, which has a higher acquisition cost, and only uses the relative poses between image pairs for network training. However, with only relative poses, keypoints in one image projected onto another image only result in an epipolar line, and any point on the epipolar line could potentially be a matching keypoint. Therefore, this method has significant ambiguity and is unlikely to yield satisfactory results.
[0046] Of the above methods, Method 2 is the most ideal training scheme for the feature point extraction neural network. However, dense image depth acquisition is difficult, and currently only a small amount of training data can meet this requirement. The richness of the training data for the neural network is crucial to the training results. Although Methods 1 and 3 can obtain a large amount of data, the image pairs generated by Method 1 do not perfectly match the real situation, and the data obtained by Method 3 contains significant ambiguity. Therefore, none of the above methods can effectively train the feature point extraction network with a large amount of data.
[0047] In this embodiment, since the image triplets in the training samples contain three images with different views, when the electronic device uses the image triplets to train the feature point extraction model, it can determine the matching key points based on the three images with different views. Therefore, this solution does not require the acquisition of costly dense depth images, and it detects the matching relationship between key points based on the three images with different views in the first image triplet, thereby achieving the training of the feature point extraction model. That is, this invention reduces the requirements for the collected data when training the feature point extraction model while ensuring the accuracy of the model. Moreover, this invention reduces the requirements for the collected data when training the feature point extraction model, and can train the feature point extraction network on richer data rather than data that meets specific requirements, so that the trained feature point extraction model has better generalization performance. This provides more reliable and robust image feature points for downstream SLAM algorithms, thereby providing more stable positioning and tracking capabilities for AR / VR and other applications.
[0048] The execution entity of the model training method provided in this application embodiment can be a model training device, which can be an electronic device or a functional module in an electronic device. The following uses an electronic device as an example to illustrate the technical solution provided in this application embodiment.
[0049] This application provides a model training method. Figure 2 A flowchart of a model training method provided in an embodiment of this application is shown. Figure 2 As shown, the model training method provided in this application embodiment may include the following steps 201 to 203.
[0050] Step 201: The electronic device acquires the first training sample.
[0051] In this embodiment of the application, the first training sample includes a first image triplet, each image in the first image triplet corresponds to a shooting angle, and each image in the first image triplet includes the same shooting object.
[0052] In this embodiment of the application, the first training sample may contain one or more image triples. In other words, the first image triple may be one of the image triples in the first training sample.
[0053] Optionally, in the embodiments of this application, the first image triplet may include three images, each of which may correspond to a dimension, namely the X-axis, Y-axis and Z-axis, and each image in the triplet is ordered.
[0054] For example, if three images are labeled 1, 2, and 3, then each image in the first image triplet can be sorted as 1, 2, 3; or 3, 2, 1. However, an image triplet sorted as 1, 2, 3 is not equal to an image triplet sorted as 3, 2, 1.
[0055] Optionally, in the embodiments of this application, each image in the first image triplet can be captured by an electronic device; or, the electronic device can obtain it from a cloud server.
[0056] For example, the fact that each image in the first image triplet can be captured by an electronic device can be understood as follows: the electronic device can capture images of the same subject or the same scene from different perspectives in a normal static environment using a camera, thereby obtaining the first image triplet.
[0057] It should be noted that each image in the first image triplet mentioned above needs to cover as many viewpoints as possible in the scene being captured. Therefore, the number of images to be captured is related to the size of the scene; the larger the scene, the more images need to be captured to cover the entire scene.
[0058] Understandably, the more initial training samples a device has, the better it will perform when training the model based on those samples.
[0059] Optionally, in the embodiments of this application, the images captured by the above-mentioned electronic device from different perspectives can be captured in multiple shots or can be non-continuous.
[0060] Optionally, in the embodiments of this application, each image in the first image triplet can be randomly obtained from images captured by the electronic device from different perspectives; or, the first image triplet can be an image triplet in the first training sample that meets the following preset conditions.
[0061] For example, each image in the first image triplet corresponds to a different shooting angle.
[0062] For example, the shooting angle may include at least one of the following: any angle such as top view, bottom view, left view, and right view. The specific angle can be determined according to the actual use, and this application embodiment does not impose any restrictions.
[0063] Optionally, in this embodiment of the application, the number of the first image triples can be one or more.
[0064] Optionally, in this embodiment, the subject of the photograph can be a person, animal, or object. The specific subject can be determined based on actual usage, and this embodiment does not impose any limitations.
[0065] Step 202: Based on the image feature information of the key points in each image of the first image triplet, the electronic device determines the similarity information between the same matching key points between any two images in the first image triplet.
[0066] Optionally, in the embodiments of this application, the above-mentioned image feature information may include at least one of the following: image brightness, image color difference, image contrast, and imaging information.
[0067] In this embodiment of the application, after obtaining the key points in each image, the electronic device can obtain the image feature information corresponding to the key point, i.e., the descriptor, based on the key point matching network model. Then, the electronic device can determine the similarity information between the same matching key points between any two images in the first image triplet through the descriptor.
[0068] For example, the electronic device can determine the matching key points in other images corresponding to the key points of each image in the first image triplet, and then obtain the descriptor corresponding to the matching key point through the key point matching network model; then, the electronic device can determine the similarity information between the same matching key point between any two images in the first image triplet by matching the descriptor corresponding to the key point.
[0069] For example, after obtaining the descriptors corresponding to the key points of any two images in the first image triplet, the electronic device can calculate the similarity information of the descriptors corresponding to the matching key points of any two images.
[0070] In this embodiment of the application, the aforementioned similarity information is used to indicate the cosine similarity of the descriptors corresponding to the matching key points of two pairs of images.
[0071] For example, an electronic device can use a first algorithm to calculate the cosine similarity of descriptors corresponding to matching key points of any two images.
[0072] For example, the first algorithm described above can be an artificial intelligence (AI) algorithm or a neural network algorithm.
[0073] Step 203: The electronic device trains the matching key point localization model based on similarity information.
[0074] In this embodiment of the application, the aforementioned key points are the key points corresponding to the photographed object in each image of the first image triplet.
[0075] In this embodiment of the application, after obtaining the above-mentioned similarity information, the electronic device can construct a loss function through the similarity information, thereby training the matching key point localization model.
[0076] For example, electronic devices can construct a loss function based on cosine similarity to train a matching keypoint localization model.
[0077] For example, the loss function described above can be any of the following: cross-entropy loss function, squared loss function, absolute value loss function, or logarithmic loss function, etc. The specific loss function can be determined based on actual usage, and this application embodiment does not impose any limitations.
[0078] In the model training method provided in this application embodiment, the electronic device can acquire a first training sample, which includes a first image triplet. Each image in the first image triplet corresponds to a shooting viewpoint, and each image in the first image triplet contains the same shooting object. Based on the image feature information of key points in each image in the first image triplet, the similarity information between the same matching key points between any two images in the first image triplet is determined. Based on the similarity information, the matching key point localization model is trained. Here, the key point is the key point corresponding to the shooting object in each image. In this scheme, since the image triplet in the training sample contains three images with different views, when the electronic device uses the image triplet to train the feature point extraction model, it can determine the matching key points based on the three images with different views. Therefore, this scheme does not require the acquisition of costly dense depth images, and it detects the matching relationship between key points based on the three images with different views in the first image triplet, thereby realizing the training of the feature point extraction model. In other words, this invention reduces the requirements for the collected data when training the feature point extraction model, while also ensuring the accuracy of the model.
[0079] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, after step 201 above, the model training method provided in this application embodiment further includes step 301 below, and step 202 above can be specifically implemented by step 202a below.
[0080] Step 301: The electronic device calculates the image overlap between any two images in the first image triplet.
[0081] In this embodiment of the application, the electronic device can calculate the image overlap between any two images by using the camera intrinsic parameters and camera pose corresponding to each image in the first image triplet.
[0082] For example, the camera intrinsic parameters mentioned above may include: camera focal length and camera center point position.
[0083] For example, the aforementioned camera intrinsic parameters may also include at least one of the following: camera lens, image sensor type, and position information of pixels in the image sensor, i.e., horizontal axis coordinates.
[0084] Optionally, in the embodiments of this application, each image in the first image triplet corresponds to a camera pose and camera intrinsic parameters.
[0085] For example, step 301 above can be implemented by steps 301a to 301e below.
[0086] Step 301a: The electronic device obtains the relative pose information between the third image and the fourth image based on the camera pose corresponding to the third image and the camera pose corresponding to the fourth image.
[0087] In this embodiment of the application, the third image and the fourth image are two of the images in the first image triplet.
[0088] In this embodiment, before obtaining the relative pose information between the third and fourth images based on the camera pose corresponding to the third image and the camera pose corresponding to the fourth image, the electronic device can obtain the camera pose and camera intrinsic parameters corresponding to the third image, as well as the camera pose and camera intrinsic parameters corresponding to the fourth image, through the SFM algorithm. Thus, the electronic device can obtain the relative pose information between the third and fourth images using the camera pose and camera intrinsic parameters corresponding to the third image and the fourth image, respectively.
[0089] For example, taking the third image as image A and the fourth image as image B, let the pose of image A be T. A The pose of image A is T B Then the relative pose between image A and image B is
[0090] Step 301b: The electronic device determines the overlap volume of the imaging cones between the third image and the fourth image based on the relative pose information, the imaging cone corresponding to the third image, and the imaging cone corresponding to the fourth image.
[0091] In this embodiment of the application, the imaging cone corresponding to the above image refers to the field of view of the camera when shooting the first object.
[0092] In this embodiment of the application, when the electronic device determines that the pose of the third image and the pose of the fourth image are relative poses, and the electronic device obtains the imaging frustum corresponding to the third image and the imaging frustum corresponding to the fourth image, the electronic device can calculate the overlap volume of the imaging frustum between the third image and the fourth image by integration.
[0093] It should be noted that this application can take into account the camera frustum between the third image and the fourth image, and calculate the overlap volume of the imaging frustum between the third image and the fourth image based on the actual pose of the camera in space.
[0094] Step 301c: The electronic device determines the three-dimensional volume of the imaging cone corresponding to the third image based on the relative pose information and the imaging cone corresponding to the third image.
[0095] In this embodiment, the electronic device obtains the three-dimensional volume of the imaging cone corresponding to the third image by integration based on the actual pose in the camera space corresponding to the third image.
[0096] Step 301d: The electronic device determines the three-dimensional volume of the imaging cone corresponding to the fourth image based on the relative pose information and the imaging cone corresponding to the fourth image.
[0097] In this embodiment, the electronic device obtains the three-dimensional volume of the imaging cone corresponding to the fourth image by integration based on the actual pose in the camera space corresponding to the fourth image.
[0098] Step 301e: The electronic device obtains the overlap between the third image and the fourth image based on the overlap volume, the three-dimensional volume of the third image, and the three-dimensional volume of the fourth image.
[0099] For example, the electronic device can obtain the overlap between the third image and the fourth image using the following formula 1, which is as follows:
[0100] O AB =2*V overlap / (V frustumA +V frustumB (1)
[0101] Among them, O AB V represents the overlap between the third and fourth images. overlap V is the overlap volume of the imaging cones between the third and fourth images. frustumA V is the three-dimensional volume of the imaging cone of the third image. frustumB Let be the three-dimensional volume of the imaging cone of the fourth image.
[0102] It should be noted that the above description only applies to two images in the first image triplet. The image overlap between any two images in the first image triplet can be obtained using the above method.
[0103] In this embodiment, the order of steps 301b, 301c, and 301d is not limited. For example, the electronic device may execute step 301c first, then step 301b, and finally step 301d; or, the electronic device may execute step 301d first, then step 301b, and finally step 301c. The specific order can be determined based on actual usage.
[0104] Step 202a: When the image overlap between any two images in the first image triplet is greater than or equal to the first threshold, the electronic device determines the similarity information between the same matching key point between any two images in the first image triplet based on the image feature information of the key point in each image in the first image triplet.
[0105] Optionally, in this embodiment of the application, if any one of the image overlaps between any two images in the first image triplet is less than a preset threshold, the electronic device can randomly select a second image triplet from the first training sample again and perform the above calculation again.
[0106] Optionally, in the embodiments of this application, the above-mentioned "the image overlap between any two images in the first image triplet is greater than or equal to a preset threshold" means / characterizes that the same image area captured between any two images is greater than or equal to a preset threshold.
[0107] In this embodiment, the electronic device estimates the image overlap rate by using a camera imaging frustum, avoiding reliance on dense depth images which have high acquisition costs. Moreover, since the computational cost required to estimate the image overlap rate using a camera imaging frustum is relatively small compared to dense depth images, the efficiency of image overlap rate estimation is also improved.
[0108] Optionally, in this embodiment of the application, before step 301b above, the model training method provided in this embodiment of the application includes the following steps 401 to 404.
[0109] Step 401: The electronic device sets the first plane and the second plane based on the camera pose corresponding to the fifth image.
[0110] In this embodiment of the application, the distance between the first plane and the camera is less than the distance between the second plane and the camera. The first plane includes the camera's imaging range, and the second plane also includes the camera's imaging range.
[0111] In this embodiment of the application, the phrase "the first plane includes the camera imaging range" means that the first plane can include the complete camera imaging range.
[0112] In this embodiment of the application, the phrase "the second plane includes the camera imaging range" means that the second plane can include the complete camera imaging range.
[0113] In this embodiment of the application, the phrase "the distance between the first plane and the camera is less than the distance between the second plane and the camera" means that the first plane has a small imaging range and the second plane has a large imaging range.
[0114] For example, the electronic device can set a first plane and a second plane at a first distance and a second distance in front of the camera based on the camera pose and the camera centerline, and the electronic device can set the width and height corresponding to the first plane and the width and height corresponding to the second plane according to the camera imaging range, wherein the second distance is greater than the first distance.
[0115] Step 402: The electronic device obtains the vertex information of the first plane based on the distance between the first plane and the camera, the camera's intrinsic parameters, and the width and height of the first plane.
[0116] For example, the electronic device can obtain the vertex information of the first plane using the following formula 2.
[0117] Specifically, Formula 2 can be:
[0118]
[0119] Where corner1 is the vertex information of the first plane, d1 is the first distance, K is the camera intrinsic parameter, w1 is the width of the first plane, and h1 is the height of the first plane.
[0120] It should be noted that the vertex information of the first plane mentioned above includes the coordinates of the four vertices of the first plane.
[0121] Step 403: The electronic device obtains the vertex information of the second plane based on the distance between the second plane and the camera, the camera intrinsic parameters, and the width and height of the second plane.
[0122] For example, the electronic device can obtain the vertex information of the second plane using the following formula 3.
[0123] Specifically, Formula 3 can be:
[0124]
[0125] Where corner2 is the vertex information of the second plane, d2 is the second distance, K is the camera intrinsic parameter, w2 is the width of the first plane, and h2 is the height of the first plane.
[0126] It should be noted that the vertex information of the second plane mentioned above includes the coordinates of the four vertices of the second plane.
[0127] Step 404: The electronic device obtains the imaging frustum corresponding to the fifth image based on the vertex information of the first plane and the vertex information of the second plane.
[0128] In this embodiment of the application, the fifth image is one of the images in the first image triplet.
[0129] In this embodiment of the application, after obtaining the vertex information of the first plane and the vertex information of the second plane, the electronic device can intersect the center line of the camera based on the coordinates of the four vertices of the first plane and the four vertices of the second plane to obtain the imaging cone corresponding to the fifth image.
[0130] For example, such as Figure 4 As shown, Figure 4 It includes a first plane 10 and a second plane 11. After the electronic device obtains the coordinates of the four vertices corresponding to the first plane 10 and the four vertices corresponding to the second plane 11, the electronic device can extend the lines from the four vertices of the second plane 11 until they intersect with the center line of the camera. In this way, the imaging cone 12 corresponding to the fifth image is obtained.
[0131] It should be noted that the above explanation is based solely on one image in the first image triplet. Each image in the first image triplet can be used to obtain the corresponding imaging frustum through the above embodiments. To avoid repetition, this will not be elaborated further here.
[0132] In this embodiment, the execution order of steps 402 and 403 is not limited. For example, the electronic device may execute step 402 first, then step 403; or, the electronic device may execute step 403 first, then step 402. The specific order can be determined based on actual usage, and this embodiment is not limited thereto.
[0133] In this embodiment, the electronic device can obtain the imaging frustum corresponding to the fifth image through the camera intrinsic parameters corresponding to the fifth image and the set first and second planes, thus ensuring that the electronic device can estimate the overlap of the images through the imaging frustum.
[0134] Optionally, in this embodiment of the application, the first image triplet includes a reference image, a first image, and a second image. The reference image in the first image triplet can be any image in the first image triplet.
[0135] For example, combined Figure 2 ,like Figure 5 As shown, prior to step 202 above, the model training method provided in this application embodiment further includes steps 501 and 502 as described below.
[0136] Step 501: The electronic device determines the same matching key point from the reference image that corresponds to the first key point in the first image, based on the distance between all key points in the first image and the first epipolar line.
[0137] Optionally, in the embodiments of this application, the reference image is randomly set by the electronic device, that is, any one of the three images in the first image triplet can be the reference image.
[0138] In this embodiment of the application, all key points in the first image are determined by the electronic device through a feature point extraction network model.
[0139] In this embodiment of the application, the first epipolar line is the projection epipolar line of the first key point in the reference image onto the first image, and the first key point is one of the key points in the reference image.
[0140] In this embodiment of the application, the electronic device can determine the key point with the smallest distance between all key points in the first image and the first epipolar line as the same matching key point corresponding to the first key point in the reference image.
[0141] It is understood that the above is only an explanation using one key point. For all key points in the reference image, the electronic device can obtain the same matching key point in the first image corresponding to all key points in the reference image in the above manner.
[0142] In other words, the electronic device can traverse each key point in the reference image to obtain the matching key point corresponding to each key point in the first image.
[0143] Step 502: The electronic device determines the same matching key point from the reference image that corresponds to the first key point in the second image, based on the distance between all key points in the second image and the second epipolar line.
[0144] In this embodiment of the application, the second epipolar line is the projection epipolar line of the first key point in the reference image onto the second image, and the first key point is one of the matching key points in the reference image.
[0145] In this embodiment of the application, all key points in the first image are determined by the electronic device through a feature point extraction network.
[0146] In this embodiment of the application, the electronic device can determine the key point with the smallest distance between all key points in the second image and the second polar line as the matching key point corresponding to the first key point in the reference image.
[0147] It is understood that the above is only an explanation using one key point. For all key points in the reference image, the electronic device can obtain the corresponding matching key points in the second image using the above method.
[0148] In other words, the electronic device can traverse each key point in the reference image to obtain the matching key point of each key point in the second image.
[0149] Optionally, in this embodiment of the application, before the electronic device obtains the matching key point corresponding to the first image and the matching key point corresponding to the second image, the electronic device can obtain the first candidate matching key point in the first image and the second candidate matching key point in the second image through the first polar line and the second polar line respectively; and then obtain the same matching key point based on the first candidate matching key point and the second candidate matching key point.
[0150] For example, the number of the first candidate matching key points can be one or more, and the number of the second candidate matching key points can be one or more.
[0151] For example, such as Figure 6 As shown, assuming image A is the reference image, image B is the first image, and image C is the second image, the first key point in image A is located at... Figure 6 The first polar line is represented by stars in image B, and is indicated by a line l. AB This indicates that image B can contain 8 key points. Figure 6 The second polar line is represented by a hollow circle in image C, and the second polar line is represented by a l. AC This indicates that image C can contain 8 key points. Figure 6 Represented by hollow circles, in images B and C, the electronic device can calculate the distance from all detected keypoints to the epipolar line l. AB and l AC The distance. Taking image B as an example, let the coordinates of the first keypoint be P. BM =(x b ,y b ), then the first polar line l AB The parametric equation is: a AB x b +b AB y b +c AB =0, where a, b, c are constants, then all keypoints in image B will be connected to the first epipolar line l. AB The distance is: If the distance is less than a preset threshold (usually set to 3 pixels), the keypoint is retained. Figure 6Solid keypoints in the data are selected as candidate matching keypoints; if the distance is greater than the threshold, the keypoint is discarded. Figure 6 (The hollow key point in the middle).
[0152] It should be noted that the candidate matching key points in image C can also be obtained through the above example. To avoid repetition, it will not be repeated here.
[0153] Optionally, in this embodiment of the application, the electronic device can determine the matching key point from the first candidate matching key point and the second candidate matching key point.
[0154] For example, the electronic device can further determine matching keypoints by the projection relationship between images A, B, and C in the first image triplet. Combined with... Figure 6 ,like Figure 7 As shown, taking image B as an example, after the electronic device determines all candidate matching points in image B, the electronic device can use the above projection method to project all candidate matching points in image B ( Figure 7 China-Israel P B3 P B2 and P B3 (represented) are all projected onto image C to obtain the epipolar line from image B to image C, thus obtaining the epipolar line from image B to image C (l B1C , l B2C and l B3C ), and calculate all candidate matching points (P) in image C respectively. C1 P C2 and P C3 The distance from each candidate matching point in image B to all epipolar lines in image C can be calculated similarly. Thus, the symmetrical epipolar distance between a pair of candidate matching points in images B and C can be obtained. Then construct the midpoint P in image A A With point P in image B BM And point P in image C CM The polar distance between them:
[0155]
[0156] like If it is less than the preset threshold, then p is determined. A ,p Bm ,p cn The three points are the key matching points. Figure 7 (Represented by a pentagram); otherwise, the three points are not matching keypoints. Following this step, all keypoints in images A, B, and C are identified to obtain a set of matching keypoints for descriptor training.
[0157] It is understandable that when there is only one first candidate matching key point, the electronic device can directly use the first candidate matching key point as the matching key point; when there is only one second candidate matching key point, the electronic device can directly use the second candidate matching key point as the matching key point.
[0158] In this embodiment, the electronic device can determine the matching key point corresponding to the first key point in the reference image in the first image based on the distance between all key points in the first image and the first epipolar line, and determine the matching key point corresponding to the first key point in the reference image in the first image based on the distance between all key points in the second image and the second epipolar line. In other words, the electronic device can determine the matching key point based on the epipolar line projected from the key points between the other images in the first image triplet (excluding the reference image) and the distance from the key point to the epipolar line projected from the reference image.
[0159] Optionally, in this embodiment of the application, the first image triplet includes a reference image, a first image, and a second image.
[0160] For example, combined Figure 2 ,like Figure 8 As shown, prior to step 201 above, the model training method provided in this application embodiment further includes step 601 as described below.
[0161] Step 601: When the angle between the first polar line and the second polar line is greater than or equal to the second threshold, the electronic device uses the first image triplet as a training sample.
[0162] In this embodiment of the application, the first epipolar line is the projection epipolar line of the first key point in the reference image onto the first image, and the second epipolar line is the projection epipolar line of the first key point in the reference image onto the second image. The first key point is any matching key point in the reference image.
[0163] Optionally, in this embodiment of the application, when the angle between the first polarity and the second polarity is less than the second threshold, the electronic device can re-obtain the third image triplet from the first training sample and redetermine the key matching point in each image of the third image triplet according to the above method.
[0164] In this embodiment of the application, when the angle between the first polar line and the second polar line is greater than or equal to the second threshold, the electronic device can use the first image triplet as a training sample to ensure the accuracy of the electronic device in obtaining the key matching point based on the first image triplet.
[0165] For example, the model training method provided in this application will be explained in detail below through specific examples. Specifically, it can be implemented through steps 20 to 32 below.
[0166] Step 20: The electronic device acquires N images, where N is an integer greater than or equal to 3.
[0167] For example, the above N images are taken by an electronic device from different perspectives of the same scene in a normal static environment.
[0168] For example, there are at least three different perspectives mentioned above.
[0169] Step 21: The electronic device calculates the camera intrinsic parameters and camera pose for each of the N images.
[0170] Step 22: The electronic device samples the image triplet.
[0171] For example, the electronic device can randomly sample three images from all images of the same scene acquired in step 20, based on different viewpoints. These three images are denoted as A, B, and C, respectively. Each of these three images has a corresponding in-camera pose.
[0172] Step 24: For each image in the image triplet, the electronic device constructs the camera imaging frustum corresponding to each image.
[0173] Step 25: The electronic device filters the image triplets based on the camera imaging cone corresponding to each image.
[0174] For example, an electronic device can filter image triples by the degree of repetition between any two images in the image triplet.
[0175] Step 26: If the repetition between any two images in the above image triplet meets the first threshold, the electronic device filters between any two images in the image triplet again based on the epipolar geometry relationship.
[0176] For example, when the angle between the first polar line and the second polar line is greater than or equal to the second threshold, the electronic device uses the first image triplet as a training sample.
[0177] Step 27: Count the sampled image triplets and save the filtered image triplets.
[0178] Step 28: For each image in the image triplet, the electronic device extracts the key points and descriptors of each image through a feature point extraction network.
[0179] Step 29: The electronic device calculates the projected epipolar line through key points in the image.
[0180] Step 30: The electronic device obtains candidate matching key points by projecting epipolar lines.
[0181] Step 31: The electronic device obtains the matching key points through the projected epipolar lines corresponding to the candidate matching key points.
[0182] Step 32: The electronic device calculates the cosine similarity of the descriptors between matching points of any two images (AB, BC, AC) in the image triplet, and constructs a loss function based on the cosine similarity for network training.
[0183] It should be noted that the model training method provided in this application can be executed by a model training device, an electronic device, or a functional module or entity within an electronic device. This application uses a model training device executing the model training method as an example to illustrate the model training device provided in this application.
[0184] Figure 9 A schematic diagram of a possible structure of the model training apparatus involved in an embodiment of this application is shown. For example... Figure 9 As shown, the model training device 70 may include: an acquisition module 71, a determination module 72, and a training module 73.
[0185] The acquisition module 71 is used to acquire a first training sample, which includes a first image triplet. Each image in the first image triplet corresponds to a shooting viewpoint, and each image in the first image triplet includes the same shooting object. The determination module 72 is used to determine the similarity information between any two images in the first image triplet for the same matching keypoint based on the image feature information of keypoints in each image of the first image triplet. The training module 73 is used to train the matching keypoint localization model based on the similarity information; wherein the keypoint is the keypoint corresponding to the shooting object in each image.
[0186] In one possible implementation, the first image triplet includes a reference image, a first image, and a second image. The determining module 72 is further configured to, before determining the similarity information between the same matching keypoint between any two images in the first image triplet based on the image feature information of keypoints in each image of the first image triplet, determine the same matching keypoint corresponding to the first keypoint in the first image from the reference image based on the distance between all keypoints in the first image and the first epipolar line; and determine the same matching keypoint corresponding to the first keypoint in the second image from the reference image based on the distance between all keypoints in the second image and the second epipolar line; wherein the first epipolar line is the projection epipolar line of the first keypoint in the reference image onto the first image, the second epipolar line is the projection epipolar line of the first keypoint in the reference image onto the second image, and the first keypoint is any matching keypoint in the reference image.
[0187] In one possible implementation, combining Figure 9 ,like Figure 10 As shown, the model training apparatus 70 provided in this embodiment further includes a processing module 74. The processing module 74 is used to calculate the image overlap between any two images in the first image triplet after the acquisition module 71 acquires the first training sample. The determination module 72 is specifically used to determine the similarity information between the same matching keypoint between any two images in the first image triplet, based on the image feature information of keypoints in each image of the first image triplet, when the image overlap between any two images in the first image triplet is greater than or equal to a preset threshold.
[0188] In one possible implementation, each image in the first image triplet corresponds to a camera pose and camera intrinsic parameters. The processing module 74 is specifically configured to: obtain relative pose information between the third and fourth images based on the camera poses corresponding to the third and fourth images; determine the overlap volume of the imaging cones between the third and fourth images based on the relative pose information, the imaging cone corresponding to the third image, and the imaging cone corresponding to the fourth image; determine the three-dimensional volume of the imaging cone corresponding to the third image based on the relative pose information and the imaging cone corresponding to the third image; determine the three-dimensional volume of the imaging cone corresponding to the fourth image based on the relative pose information and the imaging cone corresponding to the fourth image; and obtain the overlap degree between the third and fourth images based on the overlap volume, the three-dimensional volume of the third image, and the three-dimensional volume of the fourth image; wherein the third and fourth images are two images in the first image triplet.
[0189] In one possible implementation, combining Figure 10 ,like Figure 11As shown, the model training apparatus 70 provided in this embodiment further includes: a setting module 75; the setting module 75 is used to, before the processing module 74 obtains the relative pose information between the third image and the fourth image based on the camera pose corresponding to the third image and the fourth image, set a first plane and a second plane based on the camera pose corresponding to the fifth image, wherein the distance between the first plane and the camera is less than the distance between the second plane and the camera, the first plane includes the camera imaging range, and the second plane includes the camera imaging range. The processing module 74 is also used to obtain vertex information of the first plane based on the distance between the first plane and the camera, camera intrinsic parameters, and the width and height of the first plane; and to obtain vertex information of the second plane based on the distance between the second plane and the camera, camera intrinsic parameters, and the width and height of the second plane; and to obtain the imaging frustum corresponding to the fifth image based on the vertex information of the first plane and the second plane; wherein the fifth image is any one of the images in the first image triplet.
[0190] This application provides a model training apparatus. Since the image triplets in the training samples contain three images with different views, when the model training apparatus uses image triplets to train a feature point extraction model, it can determine matching key points based on the three images with different views. Therefore, this solution does not require the acquisition of costly dense depth images, and it detects the matching relationship between key points based on the three images with different views in the first image triplet, thereby realizing the training of the feature point extraction model. That is, this invention reduces the requirements for the acquired data when training the feature point extraction model while ensuring the accuracy of the model.
[0191] The model training device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0192] The model training device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0193] The model training apparatus provided in this application embodiment can implement all the processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0194] Optionally, such as Figure 12 As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described model training method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0195] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0196] Figure 13 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0197] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0198] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0199] The input unit 104 is used to acquire a first training sample, which includes a first image triplet. Each image in the first image triplet corresponds to a shooting viewpoint, and each image in the first image triplet includes the same shooting object. The processor 110 is used to determine the similarity information between the same matching keypoint between any two images in the first image triplet based on the image feature information of keypoints in each image in the first image triplet; and to train the matching keypoint localization model based on the similarity information; wherein the matching keypoint is the keypoint corresponding to the shooting object in each image.
[0200] This application provides an electronic device where, since the image triplets in the training samples contain three images with different views, when the electronic device trains a feature point extraction model using these image triplets, it can determine matching key points based on the three different views. Therefore, this solution does not require the acquisition of costly dense depth images, and it detects the matching relationship between key points based on the three different views in the first image triplet, thereby enabling the training of the feature point extraction model. In other words, this invention reduces the requirements for the acquired data when training the feature point extraction model while ensuring the accuracy of the model.
[0201] Optionally, in this embodiment, the first image triplet includes a reference image, a first image, and a second image. The processor 110 is further configured to, before determining the similarity information between the same matching keypoint between any two images in the first image triplet based on the image feature information of keypoints in each image of the first image triplet, determine the same matching keypoint corresponding to the first keypoint in the first image from the reference image based on the distance between all keypoints in the first image and the first epipolar line; and determine the same matching keypoint corresponding to the first keypoint in the second image from the reference image based on the distance between all keypoints in the second image and the second epipolar line; wherein the first epipolar line is the projection epipolar line of the first keypoint in the reference image onto the first image, the second epipolar line is the projection epipolar line of the first keypoint in the reference image onto the second image, and the first keypoint is any matching keypoint in the reference image.
[0202] Optionally, in this embodiment of the application, the processor 110 is further configured to calculate the image overlap between any two images in the first image triplet after acquiring the first training sample. Specifically, the processor 110 is configured to determine the similarity information between the same matching key point between any two images in the first image triplet, based on the image feature information of the key points in each image of the first image triplet, when the image overlap between any two images in the first image triplet is greater than or equal to a first threshold.
[0203] Optionally, in this embodiment, each image in the first image triplet corresponds to a camera pose and camera intrinsic parameters; the processor 110 is specifically configured to obtain relative pose information between the third image and the fourth image based on the camera pose corresponding to the third image and the camera pose corresponding to the fourth image; determine the overlap volume of the imaging cones between the third image and the fourth image based on the relative pose information, the imaging cone corresponding to the third image, and the imaging cone corresponding to the fourth image; determine the three-dimensional volume of the imaging cone corresponding to the third image based on the relative pose information and the imaging cone corresponding to the third image; determine the three-dimensional volume of the imaging cone corresponding to the fourth image based on the relative pose information and the imaging cone corresponding to the fourth image; and obtain the overlap degree between the third image and the fourth image based on the overlap volume, the three-dimensional volume of the imaging cone corresponding to the third image, and the three-dimensional volume of the imaging cone corresponding to the fourth image; wherein the third image and the fourth image are two images in the first image triplet.
[0204] Optionally, in this embodiment, the processor 110 is further configured to, before determining the overlap volume of the imaging frustums between the third image and the fourth image based on the relative pose information, the imaging frustum corresponding to the third image, and the imaging frustum corresponding to the fourth image, set a first plane and a second plane based on the camera pose corresponding to the fifth image, wherein the distance between the first plane and the camera is less than the distance between the second plane and the camera, the first plane includes the camera imaging range, and the second plane includes the camera imaging range; obtain vertex information of the first plane based on the distance between the first plane and the camera, camera intrinsic parameters, and the width and height of the first plane; obtain vertex information of the second plane based on the distance between the second plane and the camera, camera intrinsic parameters, and the width and height of the second plane; and obtain the imaging frustum corresponding to the fifth image based on the vertex information of the first plane and the vertex information of the second plane; wherein the fifth image is any one of the images in the first image triplet.
[0205] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0206] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.
[0207] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.
[0208] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0209] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0210] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0211] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0212] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0213] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0214] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the model training method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0215] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0216] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0217] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A model training method, characterized in that, The method comprises: obtaining a first training sample, the first training sample comprising a first image triplet, each image in the first image triplet corresponding to a shooting angle, and each image in the first image triplet comprising a same shooting object; determining similarity information between a same matching key point in any two images in the first image triplet based on image feature information of the key point in each image in the first image triplet; training a key point positioning model based on the similarity information; wherein the key point is a key point corresponding to the shooting object in each image; the first image triplet comprises a reference image, a first image and a second image; before the determining similarity information between a same matching key point in any two images in the first image triplet based on image feature information of the key point in each image in the first image triplet, the method further comprises: determining a same matching key point corresponding to a first key point in the first image from the reference image based on a distance between all key points in the first image and a first epipolar line; determining a same matching key point corresponding to a first key point in the second image from the reference image based on a distance between all key points in the second image and a second epipolar line; wherein the first epipolar line is a projection epipolar line of the first key point in the reference image projected into the first image, the second epipolar line is a projection epipolar line of the first key point in the reference image projected into the second image, and the first key point is any key point in the reference image.
2. The method of claim 1, wherein, after the obtaining the first training sample, the method further comprises: calculating image overlap degrees between any two images in the first image triplet; the determining similarity information between a same matching key point in any two images in the first image triplet based on image feature information of the key point in each image in the first image triplet comprises: in a case where the image overlap degrees between any two images in the first image triplet are all greater than or equal to a first threshold, determining similarity information between a same matching key point in any two images in the first image triplet based on image feature information of the key point in each image in the first image triplet.
3. The method of claim 2, wherein, each image in the first image triplet corresponds to a camera pose and camera intrinsic parameter; the calculating image overlap degrees between any two images in the first image triplet comprises: obtaining relative pose information between the third image and the fourth image according to the camera pose corresponding to the third image and the camera pose corresponding to the fourth image; determining an overlapping volume of imaging frustums between the third image and the fourth image according to the relative pose information, an imaging frustum corresponding to the third image and an imaging frustum corresponding to the fourth image; determining a three-dimensional volume of the imaging frustum corresponding to the third image according to the relative pose information and the imaging frustum corresponding to the third image. determine a three-dimensional volume of the imaging frustum corresponding to the fourth image according to the relative pose information and the imaging frustum corresponding to the fourth image; obtain an overlap degree between the third image and the fourth image according to the overlap volume, the three-dimensional volume of the imaging frustum corresponding to the third image, and the three-dimensional volume of the imaging frustum corresponding to the fourth image; wherein the third image and the fourth image are any two images in the first image triplet.
4. The method of claim 3, wherein, Before the determining the overlap volume of the imaging frustum between the third image and the fourth image according to the relative pose information, the imaging frustum corresponding to the third image, and the imaging frustum corresponding to the fourth image, the method further comprises: set a first plane and a second plane based on a camera pose corresponding to a fifth image, the distance between the first plane and the camera is less than the distance between the second plane and the camera, the first plane contains the imaging range of the camera, and the second plane contains the imaging range of the camera; obtain vertex information of the first plane based on the distance between the first plane and the camera, camera intrinsic parameters, and the width and height of the first plane; obtain vertex information of the second plane based on the distance between the second plane and the camera, camera intrinsic parameters, and the width and height of the second plane; obtain an imaging frustum corresponding to the fifth image based on the vertex information of the first plane and the vertex information of the second plane; wherein the fifth image is any one image in the first image triplet.
5. A model training apparatus characterized by comprising: The device comprises an acquisition module, a determination module, and a training module. The acquisition module is configured to acquire a first training sample, the first training sample comprising a first image triplet, each image in the first image triplet corresponding to a shooting angle, and each image in the first image triplet comprising a same shooting object. The determination module is configured to determine similarity information between a same matching key point between any two images in the first image triplet based on image feature information of the key point in each image in the first image triplet acquired by the acquisition module. The training module is configured to train a matching key point positioning model based on the similarity information determined by the determination module. wherein the key point is a key point corresponding to the shooting object in each image. The first image triplet includes a reference image, a first image and a second image; the determination module is further configured to, before determining the similarity information between the same matching key points between any two images in the first image triplet based on the image feature information of the key points in each image in the first image triplet, determine, based on distances between all key points in the first image and a first epipolar line, a same matching key point in the reference image corresponding to a first key point in the first image; and determine, based on distances between all key points in the second image and a second epipolar line, a same matching key point in the reference image corresponding to a first key point in the second image; wherein the first epipolar line is a projection epipolar line of the first key point in the reference image projected into the first image, the second epipolar line is a projection epipolar line of the first key point in the reference image projected into the second image, and the first key point is any one of the matching key points in the reference image.
6. The apparatus of claim 5, wherein, The model training apparatus further comprises a processing module; The processing module is configured to, after the acquisition module acquires the first training sample, calculate the image overlap degree between any two images in the first image triplet; The determination module is specifically configured to, in a case where the image overlap degree between any two images in the first image triplet is greater than or equal to a first threshold, determine the similarity information between the same matching key points between any two images in the first image triplet based on the image feature information of the key points in each image in the first image triplet.
7. The apparatus of claim 6, wherein, Each image in the first image triplet corresponds to a camera pose and camera intrinsic parameter; The processing module is specifically configured to obtain relative pose information between the third image and the fourth image according to the camera pose corresponding to the third image and the camera pose corresponding to the fourth image; determine an overlap volume of imaging frustums between the third image and the fourth image according to the relative pose information, the imaging frustum corresponding to the third image and the imaging frustum corresponding to the fourth image; determine a three-dimensional volume of the imaging frustum corresponding to the third image according to the relative pose information and the imaging frustum corresponding to the third image; determine a three-dimensional volume of the imaging frustum corresponding to the fourth image according to the relative pose information and the imaging frustum corresponding to the fourth image; and obtain an overlap degree between the third image and the fourth image according to the overlap volume, the three-dimensional volume of the third image and the three-dimensional volume of the fourth image; wherein the third image and the fourth image are any two images in the first image triplet.
8. The apparatus of claim 7, wherein, The model training apparatus further comprises a setting module; The setting module is configured to set a first plane and a second plane based on a camera pose corresponding to a fifth image before the processing module obtains relative pose information between the third image and the fourth image based on a camera pose corresponding to the third image and a camera pose corresponding to the fourth image, a distance between the first plane and the camera is less than a distance between the second plane and the camera, the first plane contains an imaging range of the camera, and the second plane contains the imaging range of the camera. The processing module is further configured to obtain vertex information of the first plane based on the distance between the first plane and the camera, camera intrinsic parameters, and a width and a height of the first plane, obtain vertex information of the second plane based on the distance between the second plane and the camera, the camera intrinsic parameters, and a width and a height of the second plane, and obtain an imaging frustum corresponding to the fifth image based on the vertex information of the first plane and the vertex information of the second plane, wherein the fifth image is any one of the first image triplet.
Citation Information
Patent Citations
Feature extraction and descriptor generation method and system based on convolutional neural network
CN114119987A
Method and device for correcting shooting center point during annular shooting and electronic equipment
CN115297315A