Attitude Determination Method and Device, Electronic Device, and Storage Medium
Through the method of feature point matching and three-dimensional point cloud construction, the problem of easy deviation of pose estimation results in the prior art is solved, and more accurate target pose determination is achieved.
Patent Information
- Application Number
- CN202210239443.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-03-11
AI Technical Summary
The prior art pose estimation results in the case of changing the background of the item, etc., are prone to deviations, and it is difficult to distinguish the pose difference of the item in the item image collected at two approximate position angles.
By determining the feature points on the target item in the image to be identified and matching these feature points in the pre-stored image, a three-dimensional point cloud is constructed, thereby accurately determining the target posture of the image to be identified.
Improve the accuracy of the target pose and reduce the pose estimation deviation due to background changes or angle changes.
Smart Images

Figure CN114581525B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for determining a pose, an electronic device, and a storage medium. Background Art
[0002] One of the important applications of augmented reality technology is to interact with real objects in the real world and render virtual effects on this basis. Accurately estimating and tracking the 6D pose of an object is a prerequisite for interactive rendering and is also a very important research issue in the field of computer vision. Among them, the specific definition of the 6D pose of an object is three degrees of freedom of displacement plus three degrees of freedom of rotation. In related technologies, there will be certain deviations in the pose estimation results in cases where the object background changes. At the same time, it is also difficult to distinguish the pose differences of the object in the images acquired at two approximate positions and angles. Summary of the Invention
[0003] The present disclosure provides a method and apparatus for determining a pose, an electronic device, and a storage medium, aiming to improve the accuracy of the determined target pose.
[0004] According to a first aspect of the present disclosure, a method for determining a pose is provided, including:
[0005] Determine at least one first feature point on a target object in a to-be-recognized image;
[0006] Determine at least one target image from pre-stored images according to the target object in the to-be-recognized image, where the target image has second feature points and corresponding three-dimensional point clouds, and the second feature points of the target image have corresponding points in the three-dimensional point clouds;
[0007] Determine target points of the at least one first feature point in the three-dimensional point clouds according to the at least one first feature point and the second feature points of the target image, and the points corresponding to the second feature points in the three-dimensional point clouds;
[0008] Determine a target pose corresponding to the to-be-recognized image according to the target points of the at least one first feature point in the three-dimensional point clouds.
[0009] In a possible implementation manner, the method further includes:
[0010] Determine at least one pre-stored image, and obtain at least one second feature point of the pre-stored image;
[0011] Match the second feature points of each pre-stored image to obtain a plurality of second feature point groups, and at least one of the second feature points in each second feature point group is used to represent the same position on the target object;
[0012] Determine the three-dimensional point cloud corresponding to the pre-stored image according to the multiple second feature point groups corresponding to the target item, where each point in the three-dimensional point cloud corresponds to at least one of the second feature points in one of the second feature point groups.
[0013] In a possible implementation, the determining the three-dimensional point cloud corresponding to the pre-stored image according to the multiple second feature point groups corresponding to the target item includes:
[0014] Solve the three-dimensional point cloud of the target item based on the multiple second feature point groups through the structure from motion algorithm.
[0015] In a possible implementation, the determining at least one target image from the pre-stored images according to the target item in the image to be recognized includes:
[0016] Determine the pose information corresponding to each pre-stored image;
[0017] Determine the initial pose corresponding to the image to be recognized;
[0018] Determine the at least one target image from the pre-stored images according to the initial pose of the image to be recognized and the pose information corresponding to the pre-stored images.
[0019] In a possible implementation, the determining the pose information corresponding to each pre-stored image includes:
[0020] For each pre-stored image, perform the N-point perspective algorithm on the points corresponding to each of the second feature points included therein in the three-dimensional point cloud to obtain the pose information corresponding to the pre-stored image.
[0021] In a possible implementation, the image to be recognized is a frame in a continuously acquired image sequence, and the determining the initial pose corresponding to the image to be recognized includes:
[0022] Determine the prior poses corresponding to multiple frames of images before the image to be recognized in the image sequence;
[0023] Perform extrapolation based on the multiple prior poses to obtain the initial pose corresponding to the image to be recognized.
[0024] In a possible implementation, the determining the target points of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point, the second feature points of the target image, and the points corresponding to the second feature points in the three-dimensional point cloud includes:
[0025] Perform feature point matching between the image to be recognized and the target image to obtain the second feature points matched with each of the first feature points;
[0026] Determine the points corresponding to the second feature points matched by each of the first feature points in the 3D point cloud as the target points.
[0027] In a possible implementation manner, the determining the target pose corresponding to the image to be recognized according to the target points of the at least one first feature point in the 3D point cloud includes:
[0028] Execute the N-point perspective algorithm based on the target points of the at least one first feature point in the 3D point cloud to obtain the target pose corresponding to the image to be recognized.
[0029] In a possible implementation manner, the image to be recognized is a frame in a continuously acquired image sequence, and the method further includes:
[0030] Determine the next frame image of the image to be recognized in the image sequence as the reference image, where the reference image includes at least one third feature point on the target item;
[0031] According to the target points corresponding to each of the first feature points on the image to be recognized, determine the target points corresponding to each of the third feature points on the reference image;
[0032] Determine the reference pose corresponding to the reference image according to the corresponding relationship between each of the third feature points and the target points.
[0033] In a possible implementation manner, the determining the target points corresponding to each of the third feature points on the reference image according to the target points corresponding to each of the first feature points on the image to be recognized includes:
[0034] Track each of the first feature points on the image to be recognized according to the sparse optical flow algorithm to obtain the third feature points matched by each of the first feature points on the reference image;
[0035] Determine that the third feature points matched by each of the first feature points correspond to the target points corresponding to the first feature points in the 3D point cloud.
[0036] According to a second aspect of the present disclosure, there is provided a pose determination device, including:
[0037] A first information determination module, configured to determine at least one first feature point on a target item in an image to be recognized;
[0038] A second information determination module, configured to determine at least one target image from pre-stored images according to the target item in the image to be recognized, where the target image has second feature points and corresponding 3D point clouds, and the second feature points of the target image have corresponding points in the 3D point cloud;
[0039] A target point matching module, configured to determine a target point of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point, a second feature point of the target image, and a point corresponding to the second feature point in the three-dimensional point cloud;
[0040] An attitude determination module, configured to determine a target attitude corresponding to the image to be recognized according to the target point of the at least one first feature point in the three-dimensional point cloud.
[0041] In a possible implementation manner, the apparatus further includes:
[0042] A pre-stored image determination module, configured to determine at least one pre-stored image and obtain at least one second feature point of the pre-stored image;
[0043] A first feature point matching module, configured to match the second feature points of each pre-stored image to obtain a plurality of second feature point groups, and at least one of the second feature points in each second feature point group is used to represent the same position on the target item;
[0044] A three-dimensional point matching module, configured to determine a three-dimensional point cloud corresponding to the pre-stored image according to the plurality of second feature point groups corresponding to the target item, and each point in the three-dimensional point cloud corresponds to at least one of the second feature points in a second feature point group.
[0045] In a possible implementation manner, the three-dimensional point matching module includes:
[0046] A point cloud generation sub-module, configured to solve the three-dimensional point cloud of the target item based on the plurality of second feature point groups through a structure from motion algorithm.
[0047] In a possible implementation manner, the second information determination module includes:
[0048] A first attitude determination sub-module, configured to determine attitude information corresponding to each pre-stored image;
[0049] A second attitude determination sub-module, configured to determine an initial attitude corresponding to the image to be recognized;
[0050] An attitude screening sub-module, configured to determine the at least one target image from the pre-stored images according to the initial attitude of the image to be recognized and the attitude information corresponding to the pre-stored images.
[0051] In a possible implementation manner, the first attitude determination sub-module includes:
[0052] A first pose determination unit, configured to perform an N-point perspective algorithm on the points corresponding to each of the second feature points included in each of the pre-stored images in the three-dimensional point cloud, so as to obtain the pose information corresponding to the pre-stored image.
[0053] In a possible implementation manner, the image to be recognized is a frame in a continuously acquired image sequence, and the second pose determination sub-module includes:
[0054] A second pose determination unit, configured to determine the prior poses corresponding to multiple frames of images before the image to be recognized in the image sequence;
[0055] A third pose determination unit, configured to perform an extrapolation method based on the multiple prior poses to obtain the initial pose corresponding to the image to be recognized.
[0056] In a possible implementation manner, the target point matching module includes:
[0057] A first feature point matching sub-module, configured to perform feature point matching between the image to be recognized and the target image to obtain the second feature points matched with each of the first feature points;
[0058] A first target point matching sub-module, configured to determine the points corresponding to the second feature points matched with each of the first feature points in the three-dimensional point cloud as the target points.
[0059] In a possible implementation manner, the pose determination module includes:
[0060] A third pose determination sub-module, configured to perform an N-point perspective algorithm on the target points of the at least one first feature point in the three-dimensional point cloud to obtain the target pose corresponding to the image to be recognized.
[0061] In a possible implementation manner, the image to be recognized is a frame in a continuously acquired image sequence, and the device further includes:
[0062] A reference image determination module, configured to determine the next frame image of the image to be recognized in the image sequence as a reference image, where the reference image includes at least one third feature point on the target item;
[0063] A second feature point matching module, configured to determine the target points corresponding to each of the third feature points on the reference image according to the target points corresponding to each of the first feature points on the image to be recognized;
[0064] A reference pose determination module, configured to determine the reference pose corresponding to the reference image according to the correspondence between each of the third feature points and the target points.
[0065] In a possible implementation, the second feature point matching module includes:
[0066] A second feature point matching sub-module, configured to track each of the first feature points on the image to be recognized according to the sparse optical flow algorithm, and obtain a third feature point matched by each of the first feature points on the reference image;
[0067] A second target point matching sub-module, configured to determine that the third feature point matched by each of the first feature points corresponds to the target point corresponding to the first feature point in the three-dimensional point cloud.
[0068] According to a third aspect of the present disclosure, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.
[0069] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0070] In the embodiments of the present disclosure, by determining the two-dimensional second feature point corresponding to each two-dimensional first feature point in the image to be recognized through local feature matching, and then determining the pose of the image acquisition device when the image to be recognized is acquired according to the three-dimensional feature point corresponding to the second feature point matched with each first feature point, the accuracy of the determined target pose is improved.
[0071] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, which illustrate embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure.
[0073] Figure 1 A flowchart showing a method for determining a pose according to an embodiment of the present disclosure;
[0074] Figure 2 A schematic diagram showing a three-dimensional point cloud according to an embodiment of the present disclosure;
[0075] Figure 3 A schematic diagram showing a second feature point matching process according to an embodiment of the present disclosure;
[0076] Figure 4 A schematic diagram showing a determination of a target pose according to an embodiment of the present disclosure;
[0077] Figure 5 A schematic diagram showing a method for determining a reference pose according to an embodiment of the present disclosure;
[0078] Figure 6 A schematic diagram showing a pose determination device according to an embodiment of the present disclosure;
[0079] Figure 7 A schematic diagram showing an electronic device according to an embodiment of the present disclosure;
[0080] Figure 8 A schematic diagram showing another electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0081] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0082] The special term "exemplary" herein means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior to or better than other embodiments.
[0083] The term "and / or" in this document merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this document means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0084] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0085] The posture determination method of the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be a fixed or mobile device such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can implement the posture determination method of the embodiments of the present disclosure by the processor calling computer-readable instructions stored in the memory.
[0086] In a possible implementation manner, the embodiments of the present disclosure are used to accurately determine the posture of the image acquisition device when acquiring the to-be-recognized image according to the three-dimensional point cloud constructed from multiple pre-stored images including the target item after acquiring the to-be-recognized image including the target item.
[0087] Figure 1 The flowchart showing a posture determination method according to an embodiment of the present disclosure is as follows Figure 1 As shown, the item posture method of the embodiments of the present disclosure may include the following steps S10 - S40.
[0088] Step S10, determine at least one first feature point on the target item in the to-be-recognized image.
[0089] In a possible implementation manner, the posture determination method of the embodiments of the present disclosure is used to determine the target posture corresponding to the two-dimensional to-be-recognized image, that is, the posture of the image acquisition device when acquiring the to-be-recognized image. The to-be-recognized image can be determined by the image acquisition device acquiring the target item, and can be a single image obtained by a single acquisition method, or a frame image in an image sequence continuously acquired by a moving image acquisition device during the movement process. After the electronic device determines the to-be-recognized image, it determines at least one first feature point on the target item in the to-be-recognized image. Among them, the target item is a static item whose posture does not change during the image acquisition process, and can be an inanimate still life or a living thing that remains stationary during the image acquisition process.
[0090] Optionally, the determination process of the first feature point on the to-be-recognized image can be to first obtain the target item in the image through object recognition, and then further determine at least one first feature point on the target item for characterizing the local area of the target item. Among them, each first feature point can also have a first descriptor for describing the local feature of the first feature point, that is, the first descriptor is used to characterize the local area feature of the target item in the to-be-recognized image.
[0091] Further, the determination method of the first descriptor may be to extract a preset feature in the position where the first feature point is located, and determine a feature vector as the first descriptor according to the distribution of the preset feature. The preset feature can be set according to actual needs and may include at least one of color features and texture features. The first descriptor can be any feature descriptor. For example, it can be a Histogram of Oriented Gradients (HOG) feature descriptor. The determination process of the first descriptor may be to determine the gradient changes of each pixel in the horizontal and vertical directions within the position where the first feature point is located, and determine the amplitude and direction of each pixel according to the gradient changes to generate a histogram based on the gradient amplitude and direction of each pixel. Further, the feature vector is obtained as the first descriptor through histogram normalization.
[0092] Step S20: Determine at least one target image from the pre-stored images according to the target object in the image to be recognized.
[0093] In a possible implementation manner, after determining the target object in the image to be recognized, further determine the three-dimensional point cloud of the target object and at least one target image. Each of the pre-stored images is an image obtained by an image acquisition device pre-collecting the target object and can be stored locally in the electronic device or in a database connected remotely. The at least one target image determined by the electronic device may be all the pre-stored images or some of the pre-stored images. Optionally, the electronic device may obtain at least one target image from the pre-stored images according to any preset rule. For example, the electronic device may determine the pose information corresponding to each pre-stored image and the initial pose corresponding to the image to be recognized, and determine at least one target image from the pre-stored images according to the initial pose of the image to be recognized and the pose information corresponding to the pre-stored images.
[0094] Further, the preliminary estimation of the pose of the target object in the image to be recognized can be determined according to any existing method. For example, when the image to be recognized is a frame in a continuously acquired image sequence, the prior poses corresponding to multiple frames of images before the image to be recognized in the image sequence can be determined, that is, the poses of the image acquisition device when acquiring the multiple frames of images before. The initial pose corresponding to the image to be recognized is obtained by performing extrapolation based on the multiple prior poses. Alternatively, the image to be recognized can also be directly input into a pre-trained pose estimation model, and the corresponding initial pose is output after pose estimation, etc. The initial pose can be represented by a vector, including three displacement parameters and three rotation parameters of the image acquisition device in the target three-dimensional coordinate system when acquiring the image to be recognized.
[0095] Optionally, the initial pose of the image to be recognized and the pose information of each pre-stored image can both be represented by vectors. Therefore, the pose matching between the image to be recognized and the pre-stored images can be directly determined by calculating the vector distance, and the pre-stored images with a vector distance less than a preset distance threshold are filtered out as target images.
[0096] Optionally, the target image has second feature points and corresponding 3D point clouds, and the second feature points of the target image have corresponding points in the 3D point cloud. The 3D point cloud of the target item can be determined in advance according to at least one pre-stored image. Alternatively, after the target item in the image to be recognized is determined, multiple pre-stored images including the target item are filtered out in the database, and then the 3D point cloud of the target item is determined according to the multiple pre-stored images. Optionally, the 3D point cloud of the target item can be generated by an electronic device according to at least one pre-stored image, or generated by other devices such as a server according to at least one pre-stored image. In the case where the electronic device generates the 3D point cloud, the electronic device can generate the 3D point cloud before executing the pose determination method of the embodiments of the present disclosure, or generate the 3D point cloud of the target item during the process of executing the pose determination method of the embodiments of the present disclosure.
[0097] Optionally, when the 3D point cloud of the target item in the embodiments of the present disclosure is generated by an electronic device, the process of determining the 3D point cloud of the target item according to multiple pre-stored images can first determine at least one pre-stored image and obtain at least one second feature point of the predicted image. The second feature points of each pre-stored image are matched to obtain multiple groups of second feature points, and at least one second feature point in each group of second feature points is used to represent the same position on the target item. The 3D point cloud corresponding to the pre-stored image is determined according to the multiple groups of second feature points corresponding to the target item, and each point in the 3D point cloud corresponds to at least one second feature point in a group of second feature points. That is to say, the images with a pose similar to that of the item in the image to be recognized can be filtered out from multiple pre-stored images by means of pose screening as target images.
[0098] Among them, each pre-stored image may include a target object, and at least one second feature point is located on the target object in the corresponding pre-stored image. The second feature point is used to characterize a local area of the target object. Optionally, each second feature point may also have a corresponding second descriptor for describing local features, that is, the second descriptor is used to characterize the local area features of the target object in the pre-stored image. Further, the process of matching the second feature points may be implemented based on the second descriptor. That is, the electronic device matches the second feature points of each pre-stored image according to the second descriptors of at least one second feature point of each pre-stored image to obtain a plurality of second feature point groups. At least one second feature point in each second feature point group is used to characterize the same position on the target object. Then, a three-dimensional point cloud is determined according to the plurality of second feature point groups corresponding to the target object. Each point in the three-dimensional point cloud corresponds to at least one of the second feature points in a second feature point group.
[0099] In a possible implementation manner, the process of determining a plurality of second feature point groups may be as follows: randomly select one of the plurality of pre-stored images as a target pre-stored image, perform local feature matching on each target second feature point in the target pre-stored image with other pre-stored images based on the second descriptor, and form a second feature point group by combining the matched plurality of second feature points with the target feature point. Further, re-determine a pre-stored image where an unmatched second feature point is located as the target pre-stored image, and then perform second descriptor matching on the unmatched target second feature points in it with other unmatched second feature points to determine the second feature point group until all second feature points are matched, or the matching process of each second feature point has passed through all second feature points.
[0100] Optionally, when matching every two second feature points based on the second descriptor, the distance between the second descriptors of the two second feature points may be calculated to obtain the matching degree of the local features. The closer the distance, the higher the matching degree. Among them, the second descriptor is a vector, and the distance between the two second descriptors can be obtained by directly taking the inner product of the vectors. Further, the second feature points may also be matched by two parameters, namely local features and spatial positions. Further, after obtaining a plurality of second feature point groups, a three-dimensional point cloud of the target object may be calculated based on the multiple second feature point groups through a structure from motion algorithm. That is, the feature point tracking trajectory is determined according to the acquisition time of the image where the second feature point in each second feature point group is located, and a three-dimensional point of the target object is determined according to the multiple feature point tracking trajectories to form a three-dimensional point cloud. Therefore, each three-dimensional point in the three-dimensional point cloud corresponds to all the second feature points in a second feature point group, that is, the corresponding relationship between each three-dimensional point and at least one two-dimensional second feature point can be obtained during the construction process of the three-dimensional point cloud.
[0101] Figure 2A schematic diagram showing a three-dimensional point cloud according to an embodiment of the present disclosure. As Figure 2 shown, the three-dimensional point cloud 20 includes a plurality of three-dimensional points in a three-dimensional coordinate system, and each three-dimensional point has at least one corresponding two-dimensional second feature point.
[0102] Figure 3 A schematic diagram showing a second feature point matching process according to an embodiment of the present disclosure. As Figure 3 shown, for the first pre-stored image 30 and the second pre-stored image 31, both of which include the target object 32. There are a plurality of second feature points on the target object 32 in the first pre-stored image 30 and the second pre-stored image 31. When matching based on the corresponding second descriptors, the second feature points representing the same position of the target object 32 can be matched together according to local features.
[0103] Furthermore, each pre-stored image also has corresponding pose information, which is used to characterize the pose of the image acquisition device when the pre-stored image is acquired. This pose information can be represented by a vector, including three displacement parameters and three rotation parameters of the image acquisition device in the target three-dimensional coordinate system when the pre-stored image is acquired. The target three-dimensional coordinate system can be preset or selected. Optionally, the pose information can be determined when the pre-stored image is acquired and stored together with the pre-stored image. Alternatively, the pose information can also be determined after the three-dimensional point cloud is determined, according to the three-dimensional points corresponding to the plurality of second feature points included therein. That is to say, for each pre-stored image, the N-point perspective algorithm can be executed based on the correspondence between each second feature point included therein and the three-dimensional points in the three-dimensional point cloud, to obtain the pose information corresponding to the pre-stored image, that is, the pose of the image acquisition device when each pre-stored image is acquired.
[0104] Step S30: Determine the target points of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point, the second feature points of the target image, and the points corresponding to the second feature points in the three-dimensional point cloud.
[0105] In a possible implementation manner, after the electronic device determines at least one feature point of the target object in the image to be recognized and the second feature points of the at least one target image corresponding to the image to be recognized, it can determine the target points of the at least one first feature point in the three-dimensional point cloud according to the plurality of first feature points, the second feature points, and the points corresponding to the second feature points in the three-dimensional point cloud. That is to say, the matching of the two-dimensional points and three-dimensional points of the first feature point and the target point is indirectly realized through the two-dimensional point matching of the first feature point and the second feature point. Optionally, the electronic device can perform feature point matching between the image to be recognized and the target image to obtain the second feature points matched with each first feature point. Then, the points corresponding to the second feature points matched with each first feature point in the three-dimensional point cloud are determined as the target points.
[0106] Exemplarily, embodiments of the present disclosure can perform matching of the first feature points and the second feature points through descriptors. For example, the electronic device can perform feature point matching according to the first descriptors of each first feature point in the image to be recognized and the second descriptors of each second feature point in at least one target image, so as to determine the second feature points that match multiple first feature points, and indirectly determine the points corresponding to the first feature points in the three-dimensional point cloud as target points according to the correspondence between the second feature points and the three-dimensional points in the target point cloud. Optionally, the process of matching the second feature points corresponding to the first feature points according to the first descriptors and the second descriptors is the same as the process of matching two second feature points in step S20, and will not be elaborated here.
[0107] Further, for each first feature point in the image to be recognized, the second feature point with the closest distance can be determined by calculating the distances between the first descriptor and the second descriptors of the second feature points on each candidate image, and the points corresponding to the second feature points that match each first feature point in the three-dimensional point cloud are determined as target points.
[0108] Step S40: Determine the target pose corresponding to the image to be recognized according to the target points of the at least one first feature point in the three-dimensional point cloud.
[0109] In a possible implementation manner, after determining the target points in the three-dimensional point cloud that match each first feature point in the image to be recognized, that is, determining the positions of the target items represented by each first feature point in the image to be recognized in three-dimensional space. Optionally, the target pose corresponding to the image to be recognized can be determined according to the correspondence between at least one first feature point and each target point. For example, the N-point perspective algorithm can be executed according to the target points of at least one first feature point in the three-dimensional point cloud to obtain the target pose corresponding to the image to be recognized. That is, the target pose can be obtained by constructing and solving a prior equation according to the correspondence between multiple target points and the first feature points.
[0110] Figure 4 A schematic diagram showing a method for determining a target pose according to an embodiment of the present disclosure is shown. As Figure 4 shown, when it is necessary to determine the target pose corresponding to the image to be recognized 41, first perform pose matching between the image to be recognized 41 and the pre-stored images 40, screen at least one target image 42 from the multiple pre-stored images 40, and then perform local feature matching between the first feature points in the image to be recognized 41 and the second feature points in each target image 42 to obtain at least one second feature point that matches the first feature point as the matching feature point 43. Further, determine the points in the three-dimensional point cloud 44 that match each matching feature point 43 as target points 45, and then execute the N-point perspective algorithm according to the target points 45 that match each first feature point in the image to be recognized 41 to obtain the target pose 46 corresponding to the image to be recognized.
[0111] Further, when the image to be recognized is a frame in a continuously acquired image sequence and it is necessary to determine the pose of the image acquisition device when each frame of the acquired image is collected, the pose of the image acquisition device when the current frame of the image is collected can be determined based on the pose of the image acquisition device when the previous frame of the image is collected. For example, when acquiring a continuously acquired image sequence, image frame 1 can be extracted from the image sequence as the image to be recognized, and image frames 2 - N in the image sequence can be determined as reference images. After determining the correspondence between each first feature point in the image to be recognized and the points in the three-dimensional point cloud, according to the target points corresponding to each first feature point in the three-dimensional point cloud, determine the reference points corresponding to multiple second feature points in the reference image in the three-dimensional point cloud, and determine the reference pose of each reference image frame according to the correspondence between the third feature points and the reference points. Further, since the reference poses determined multiple times are all obtained based on the accurate target pose, as the distance between the reference image and the image to be recognized in the image sequence becomes farther and farther, the accuracy will become lower and lower. To avoid the above problems, after determining that image frame N is the reference image, extract image frame N + 1 as the next image to be recognized. Determine the target pose of the current image to be recognized through the pose determination method of the embodiments of the present disclosure to continue to determine the reference poses of frames N + 2 - N + N through the transfer method.
[0112] In a possible implementation manner, the process of determining the reference pose of the adjacent next frame of the reference image according to the target pose of the image to be recognized may include the following steps: First, determine the next frame of the image in the image sequence as the reference image, and the reference image includes at least one third feature point on the target item. According to the target points corresponding to each first feature point on the image to be recognized, determine the target points corresponding to each third feature point on the reference image. Then, determine the reference pose corresponding to the reference image according to the correspondence between each third feature point and the target point. Among them, the correspondence between each third feature point and the target point on the reference image can be determined based on the sparse optical flow algorithm, that is, track each first feature point on the image to be recognized according to the sparse optical flow algorithm to obtain the third feature point matched by each first feature point on the reference image. Determine that the third feature point matched by each first feature point corresponds to the target point corresponding to the first feature point in the three-dimensional point cloud. Further, perform the N-view algorithm according to each third feature point and the corresponding target point to obtain the reference pose corresponding to the reference image. Optionally, after determining the reference pose of the current reference image, the reference pose of the next frame of the image in the image sequence can be determined in sequence according to the correspondence between the third feature points and the target points in the current reference image.
[0113] Figure 5 FIG. shows a schematic diagram of determining a reference pose according to an embodiment of the present disclosure. As Figure 5As shown, for the image 51 to be recognized determined in the image sequence 50, the second feature point 53 matched by each first feature point in the image 51 to be recognized in the target image 52 screened from the pre-stored images can be determined according to the local feature matching method. Further, the three-dimensional point corresponding to the second feature point 53 matched by each first feature point in the three-dimensional point cloud 54 is determined as the target point 55, and the corresponding relationship between the first feature point and the target point 55 is determined to determine the target pose 56 of the image to be recognized.
[0114] Further, for the reference image 57 at the next frame position of the image 51 to be recognized in the image sequence 50, the reference pose of the reference image 57 can be determined based on the tracking algorithm according to the image 51 to be recognized, the target point 55, and the reference image 57. Optionally, the corresponding relationship between each first feature point in the image 51 to be recognized and each third feature point in the reference image 57 can be determined according to the sparse optical flow algorithm, and the corresponding relationship between each third feature point and the target point 55 can be determined according to the corresponding relationship between each first feature point and the target point 55. The reference pose 58 of the reference image 57 is determined according to the corresponding relationship between each third feature point and the target point 55.
[0115] The embodiments of the present disclosure can determine the corresponding relationship between each two-dimensional feature point in the image to be recognized and each two-dimensional feature point in the target image through local feature matching, generate a three-dimensional point cloud according to multiple target images, and further obtain the corresponding relationship between each two-dimensional feature point in the image to be recognized and the three-dimensional feature points in the three-dimensional point cloud, so as to accurately determine the pose of the image acquisition device when the image to be recognized is acquired, and improve the accuracy of the determined target pose. At the same time, the embodiments of the present disclosure can also preliminarily screen the target images by estimating the preliminary pose of the image to be recognized, reduce the calculation amount in the two-dimensional feature point matching process, and improve the efficiency of the pose determination process.
[0116] Further, when determining the poses of multiple continuously acquired images, the three-dimensional feature points corresponding to each two-dimensional feature point on the image to be recognized can be indirectly matched through two-dimensional feature point matching to determine the target pose. Further, based on the sparse optical flow algorithm, the corresponding relationship between the two-dimensional feature points of the subsequent acquired multiple frames of images and the three-dimensional feature points on the image to be recognized is tracked to quickly and accurately determine the reference poses of the multiple frames of images acquired after the image to be recognized. At the same time, this method can also accurately sense the pose change when the image acquisition position changes little.
[0117] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0118] In addition, the present disclosure also provides an attitude determination device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any of the attitude determination methods provided by the present disclosure. For the corresponding technical solutions and descriptions, refer to the corresponding records in the method section, which will not be elaborated here.
[0119] Figure 6 A schematic diagram showing an attitude determination device according to an embodiment of the present disclosure. As Figure 6 shown, the attitude determination device according to the embodiment of the present disclosure includes:
[0120] A first information determination module 60, configured to determine at least one first feature point on a target object in a to-be-recognized image;
[0121] A second information determination module 61, configured to determine at least one target image from pre-stored images according to the target object in the to-be-recognized image, where the target image has second feature points and corresponding three-dimensional point clouds, and the second feature points of the target image have corresponding points in the three-dimensional point clouds;
[0122] A target point matching module 62, configured to determine target points of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point and the second feature points of the target image, and the points corresponding to the second feature points in the three-dimensional point cloud;
[0123] An attitude determination module 63, configured to determine a target attitude corresponding to the to-be-recognized image according to the target points of the at least one first feature point in the three-dimensional point cloud.
[0124] In a possible implementation manner, the device further includes:
[0125] A pre-stored image determination module, configured to determine at least one pre-stored image and obtain at least one second feature point of the pre-stored image;
[0126] A first feature point matching module, configured to match the second feature points of each pre-stored image to obtain a plurality of second feature point groups, and at least one of the second feature points in each second feature point group is used to represent the same position on the target object;
[0127] A three-dimensional point matching module, configured to determine a three-dimensional point cloud corresponding to the pre-stored image according to the plurality of second feature point groups corresponding to the target object, and each point in the three-dimensional point cloud corresponds to at least one of the second feature points in one of the second feature point groups.
[0128] In a possible implementation manner, the three-dimensional point matching module includes:
[0129] A point cloud generation sub-module, configured to solve the three-dimensional point cloud of the target item based on multiple groups of the second feature points through a structure from motion algorithm.
[0130] In a possible implementation manner, the second information determination module 61 includes:
[0131] A first pose determination sub-module, configured to determine the pose information corresponding to each of the pre-stored images;
[0132] A second pose determination sub-module, configured to determine the initial pose corresponding to the image to be recognized;
[0133] A pose screening sub-module, configured to determine the at least one target image from the pre-stored images according to the initial pose of the image to be recognized and the pose information corresponding to the pre-stored images.
[0134] In a possible implementation manner, the first pose determination sub-module includes:
[0135] A first pose determination unit, configured to perform an N-point perspective algorithm on the points corresponding to each of the second feature points included in each of the pre-stored images in the three-dimensional point cloud, to obtain the pose information corresponding to the pre-stored images.
[0136] In a possible implementation manner, the image to be recognized is a frame in a continuously acquired image sequence, and the second pose determination sub-module includes:
[0137] A second pose determination unit, configured to determine the prior poses corresponding to multiple frames of images before the image to be recognized in the image sequence;
[0138] A third pose determination unit, configured to perform an extrapolation method based on the multiple prior poses, to obtain the initial pose corresponding to the image to be recognized.
[0139] In a possible implementation manner, the target point matching module 62 includes:
[0140] A first feature point matching sub-module, configured to perform feature point matching between the image to be recognized and the target image, to obtain the second feature points matched with each of the first feature points;
[0141] A first target point matching sub-module, configured to determine the points corresponding to the second feature points matched with each of the first feature points in the three-dimensional point cloud as the target points.
[0142] In a possible implementation manner, the pose determination module 63 includes:
[0143] A third posture determination sub-module, configured to perform an N-point perspective algorithm on target points of the at least one first feature point in the three-dimensional point cloud to obtain a target posture corresponding to the image to be recognized.
[0144] In a possible implementation, the image to be recognized is a frame in a sequence of continuously acquired images, and the apparatus further includes:
[0145] A reference image determination module, configured to determine a next-frame image of the image to be recognized in the image sequence as a reference image, where the reference image includes at least one third feature point on the target item;
[0146] A second feature point matching module, configured to determine target points corresponding to each of the third feature points on the reference image according to the target points corresponding to each of the first feature points on the image to be recognized;
[0147] A reference posture determination module, configured to determine a reference posture corresponding to the reference image according to the correspondence between each of the third feature points and the target points.
[0148] In a possible implementation, the second feature point matching module includes:
[0149] A second feature point matching sub-module, configured to track each of the first feature points on the image to be recognized according to an optical flow algorithm to obtain a third feature point matched by each of the first feature points on the reference image;
[0150] A second target point matching sub-module, configured to determine that the third feature point matched by each of the first feature points corresponds to the target point corresponding to the first feature point in the three-dimensional point cloud.
[0151] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present disclosure may be used to execute the methods described in the foregoing method embodiments. The specific implementation may refer to the description of the foregoing method embodiments. For the sake of brevity, it will not be described herein again.
[0152] The embodiments of the present disclosure further propose a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing methods are implemented. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0153] The embodiments of the present disclosure further propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the foregoing methods.
[0154] Embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0155] The electronic device may be provided as a terminal, a server, or other forms of devices.
[0156] Figure 7 The schematic diagram of an electronic device 800 according to an embodiment of the present disclosure is shown. For example, the electronic device 800 may be a terminal such as a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0157] Refer to Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0158] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0159] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0160] The power component 806 provides power for various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0161] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0162] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0163] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0164] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects for the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby items without any physical contact. The sensor assembly 814 can also include a light sensor, such as a complementary metal oxide semiconductor (CMOS) or charge coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0165] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as a wireless local area network (WiFi), a second-generation mobile communication technology (2G), or a third-generation mobile communication technology (3G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0166] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.
[0167] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions, and the above computer program instructions can be executed by a processor 820 of the electronic device 800 to complete the above-described method.
[0168] The present disclosure relates to the field of augmented reality. By acquiring image information of a target object in a real environment, relevant features, states, and attributes of the target object are detected or recognized through various vision-related algorithms, thereby obtaining an AR effect that combines virtual and real and matches a specific application. Exemplarily, the target object may relate to the face, limbs, gestures, actions, etc. related to the human body, or identification markers, landmarks related to objects, or sand tables, display areas, or display items related to venues or places. The vision-related algorithms may involve visual positioning, SLAM, 3D reconstruction, image registration, background segmentation, key point extraction and tracking of objects, pose or depth detection of objects, etc. The specific application may not only involve interaction scenarios such as guided tours, navigation, explanations, reconstructions, virtual effect overlay displays related to real scenes or objects, but also involve special effect processing related to people, such as makeup beautification, limb beautification, special effect display, virtual model display, etc. The relevant features, states, and attributes of the target object can be detected or recognized through a convolutional neural network. The convolutional neural network is a network model obtained by training a model based on a deep learning framework.
[0169] Figure 8 FIG. shows a schematic diagram of another electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server. Referring to Figure 8 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0170] The electronic device 1900 may further include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Microsoft Server Operating System (Windows Server TM ), the graphical user interface-based operating system (Mac OS X TM ) launched by Apple Inc., the multi-user multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ) or the like.
[0171] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.
[0172] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0173] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not construed to be an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0174] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0175] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0176] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0177] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0178] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0179] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0180] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0181] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0182] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of the steps does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0183] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0184] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A posture determination method, characterized in that, The method includes: Determining at least one first feature point on a target object in the image to be recognized; Determining at least one target image from pre-stored images according to the target object in the image to be recognized, where the target image has second feature points and corresponding three-dimensional point clouds, and the second feature points of the target image have corresponding points in the three-dimensional point cloud; Determining target points of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point and the second feature points of the target image, and the points corresponding to the second feature points in the three-dimensional point cloud; Determining the target pose corresponding to the image to be recognized according to the target points of the at least one first feature point in the three-dimensional point cloud; Wherein, the determining at least one target image from pre-stored images according to the target object in the image to be recognized includes: Determining the pose information corresponding to each pre-stored image, where the pose information corresponding to any one of the pre-stored images is used to characterize the pose of the image acquisition device when collecting the pre-stored image, and includes three displacement parameters and three rotation parameters of the image acquisition device in the target three-dimensional coordinate system when collecting the pre-stored image; Determining the initial pose corresponding to the image to be recognized, where the initial pose corresponding to the image to be recognized is used to preliminarily estimate the pose of the image acquisition device when collecting the image to be recognized, and includes three displacement parameters and three rotation parameters of the image acquisition device in the target three-dimensional coordinate system when collecting the image to be recognized; Determining the at least one target image from the pre-stored images according to the initial pose of the image to be recognized and the pose information corresponding to the pre-stored images.
2. The method according to claim 1, wherein The method further includes: Determining at least one pre-stored image, and obtaining at least one second feature point of the pre-stored image; Matching the second feature points of each pre-stored image to obtain a plurality of second feature point groups, and at least one of the second feature points in each second feature point group is used to characterize the same position on the target object; Determining the three-dimensional point cloud corresponding to the pre-stored image according to the plurality of second feature point groups corresponding to the target object, and each point in the three-dimensional point cloud corresponds to at least one of the second feature points in one of the second feature point groups.
3. The method according to claim 2, wherein The determining the three-dimensional point cloud corresponding to the pre-stored image according to the plurality of second feature point groups corresponding to the target object includes: Solving the three-dimensional point cloud of the target object based on the plurality of second feature point groups through a structure from motion algorithm.
4. The method according to claim 1, wherein The determining the pose information corresponding to each pre-stored image includes: For each pre-stored image, performing an N-point perspective algorithm based on the points corresponding to each second feature point included therein in the three-dimensional point cloud to obtain the pose information corresponding to the pre-stored image.
5. The method according to any one of claims 1-4, characterized in that, The image to be recognized is a frame in a continuously acquired image sequence, and the determining the initial pose corresponding to the image to be recognized includes: Determining the prior poses corresponding to multiple frames of images before the image to be recognized in the image sequence; Performing an extrapolation method based on the multiple prior poses to obtain the initial pose corresponding to the image to be recognized.
6. The method according to any one of claims 1 to 4, characterized in that Determining the target point of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point and the second feature point of the target image, and the point corresponding to the second feature point in the three-dimensional point cloud, includes: Performing feature point matching between the image to be recognized and the target image to obtain the second feature point matched with each first feature point; Determining the point corresponding to the second feature point matched with each first feature point in the three-dimensional point cloud as the target point.
7. The method according to any one of claims 1-4, characterized in that Determining the target pose corresponding to the image to be recognized according to the target point of the at least one first feature point in the three-dimensional point cloud includes: Performing an N-point perspective algorithm according to the target point of the at least one first feature point in the three-dimensional point cloud to obtain the target pose corresponding to the image to be recognized.
8. The method according to any one of claims 1-4, characterized in that The image to be recognized is a frame in a continuously acquired image sequence, and the method further includes: Determining the next frame image in the image sequence as the reference image, where the reference image includes at least one third feature point on the target item; Determining the target point corresponding to each third feature point on the reference image according to the target point corresponding to each first feature point on the image to be recognized; Determining the reference pose corresponding to the reference image according to the corresponding relationship between each third feature point and the target point.
9. The method according to claim 8, wherein Determining the target point corresponding to each third feature point on the reference image according to the target point corresponding to each first feature point on the image to be recognized includes: Tracking each first feature point on the image to be recognized according to the sparse optical flow algorithm to obtain the third feature point matched with each first feature point on the reference image; Determining that the third feature point matched with each first feature point corresponds to the target point corresponding to the first feature point in the three-dimensional point cloud.
10. A posture determination device, characterized in that, The apparatus includes: A first information determination module, configured to determine at least one first feature point on a target item in the image to be recognized; A second information determination module, configured to determine at least one target image from pre-stored images according to the target item in the image to be recognized, where the target image has second feature points and a corresponding three-dimensional point cloud, and the second feature points of the target image have corresponding points in the three-dimensional point cloud; A target point matching module, configured to determine the target point of the at least one first feature point in the three-dimensional point cloud according to the at least one first feature point and the second feature points of the target image, and the points corresponding to the second feature points in the three-dimensional point cloud; A pose determination module, configured to determine the target pose corresponding to the image to be recognized according to the target point of the at least one first feature point in the three-dimensional point cloud; Wherein, the second information determination module includes: A first pose determination sub-module, configured to determine the pose information corresponding to each pre-stored image, where the pose information corresponding to any one of the pre-stored images is used to characterize the pose of the image acquisition device when acquiring the pre-stored image, and includes three displacement parameters and three rotation parameters of the image acquisition device in the target three-dimensional coordinate system; The second posture determination sub-module is configured to determine the initial posture corresponding to the image to be recognized, wherein the initial posture corresponding to the image to be recognized is used to preliminarily estimate the posture of the image acquisition device when the image to be recognized is acquired, and includes three displacement parameters and three rotation parameters of the image acquisition device in the target three-dimensional coordinate system when the image to be recognized is acquired; The posture screening sub-module is configured to determine the at least one target image from the pre-stored images according to the initial posture of the image to be recognized and the posture information corresponding to the pre-stored images.
11. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Image processing method and device
CN109753940A
Image processing method and device, appartaus and computer storage medium
CN110246163A
Vehicle identification method and device, electronic equipment and storage medium
CN113569911A