Method, apparatus, device and medium for determining relative positions of multiple entities
By using vision sensors to acquire image data, extract feature points, build sparse maps and align them in multi-entity collaboration scenarios, the flexibility problem of relying on external calibration objects in the prior art is solved, and the method of automatically determining the relative position of multiple entities is realized, which improves the robustness and accuracy of positioning.
Patent Information
- Application Number
- CN202510121923.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-26
AI Technical Summary
In multi-entity collaboration scenarios, the dependence of external calibration objects when determining relative positions between entities limits the flexibility of the usage scenarios, and requires pre-defined patterns or features of the calibration objects, which increases the complexity of algorithm development and adaptation.
By equipping each entity with vision sensors, acquiring environmental image data, extracting feature points, building sparse maps, and aligning sparse maps of multiple entities to automatically determine the relative position between multiple entities.
It realizes that the relative position of multiple entities can be automatically determined without external calibration objects. It is suitable for complex and dynamic environments, reduces the complexity and cost of algorithm development and adaptation, and improves the robustness, accuracy and reliability of positioning.
Smart Images

Figure CN119594985B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of positioning technology, and particularly to a method, device, equipment and medium for determining the relative positions of multiple entities. Background Art
[0002] In scenarios of multi-entity collaboration, whether it is robot collaborative operation, formation driving of driverless vehicles, or interaction of multiple terminal devices in a complex environment, determining the relative positions between these entities is crucial for the efficiency, safety, etc. of multi-entity collaboration. Summary of the Invention
[0003] The present application aims to at least solve one of the technical problems in the related art to some extent.
[0004] To this end, the first object of the present application is to propose a method for determining the relative positions of multiple entities.
[0005] The second object of the present application is to propose a device for determining the relative positions of multiple entities.
[0006] The third object of the present application is to propose an electronic device.
[0007] The fourth object of the present application is to propose a computer-readable storage medium.
[0008] The fifth object of the present application is to propose a computer program product.
[0009] To achieve the above object, an embodiment of the first aspect of the present application proposes a method for determining the relative positions of multiple entities, including: obtaining image data of the environment where each entity is located through a visual sensor equipped on each entity among multiple entities; for any entity, extracting feature points from the image data corresponding to the entity to obtain multiple feature points corresponding to the entity; constructing a sparse map of the entity in the corresponding entity coordinate system based on the multiple feature points corresponding to the entity; aligning the sparse maps corresponding to any two entities among the multiple entities to determine the relative positions between any two entities.
[0010] To achieve the above object, an embodiment of the second aspect of the present application proposes a device for determining the relative positions of multiple entities, including: an acquisition module, configured to obtain image data of the environment where each entity is located through a visual sensor equipped on each entity among multiple entities; an extraction module, configured to extract feature points from the image data corresponding to any one of the entities to obtain multiple feature points corresponding to the entity; a construction module, configured to construct a sparse map of the entity based on the multiple feature points corresponding to the entity; an alignment module, configured to align the sparse maps corresponding to any two entities among the multiple entities to determine the relative positions between any two entities.
[0011] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method described in the embodiment of the first aspect above.
[0012] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the embodiment of the first aspect above.
[0013] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in the embodiment of the first aspect above.
[0014] The method, device, equipment and medium for determining the relative positions of multiple entities provided by the present application can automatically determine the relative positions between multiple entities by constructing and aligning the sparse maps of each entity among multiple entities, without relying on external calibration objects, and is applicable to various complex and dynamic environments, with strong adaptability and high flexibility. Moreover, since there is no need to pre-define the patterns or features of the calibration objects, the complexity of algorithm development and adaptation is reduced, thereby reducing the adaptation cost and time; even when some feature points are occluded or the environment changes, the sparse map alignment and positioning can still be performed through the remaining feature points, improving the robustness, accuracy and reliability of relative position determination.
[0015] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings
[0016] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0017] Figure 1 is a schematic flowchart of the method for determining the relative positions of multiple entities provided by an embodiment of the present application;
[0018] Figure 2 is a schematic flowchart of the method for determining the relative positions of multiple entities provided by another embodiment of the present application;
[0019] Figure 3 is a schematic flowchart of the method for determining the relative positions of multiple entities provided by another embodiment of the present application;
[0020] Figure 4Schematic flowchart of a method for determining the relative positions of multiple entities provided by another embodiment of the present application;
[0021] Figure 5 Schematic flowchart of a method for determining the relative positions of multiple entities provided by another embodiment of the present application;
[0022] Figure 6 Schematic structural diagram of a device for determining the relative positions of multiple entities provided by another embodiment of the present application. Detailed implementation manners
[0023] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.
[0024] In the related art, in the process of determining the relative positions between multiple entities, it generally relies on external markers (such as two-dimensional codes with prior information) or calibration plates, that is, multiple entities with visual sensors observe the same characteristic object to determine the relative positions between the multiple entities. Specifically, the following steps may be included:
[0025] 1. External calibration object: Place an object with known characteristics in the working environment. This object can be a two-dimensional code or a calibration plate, etc. Among them, these objects usually have characteristic points that are easy to identify and calculate, such as the positioning markers of the two-dimensional code or the specific patterns on the calibration plate.
[0026] 2. Relative position calculation: The relative positions between each entity and the calibration object can be calculated, and a calibration algorithm can be used to determine the positional relationship between the calibration object and the corresponding entity; furthermore, after obtaining the positional relationships between the entities and the calibration object, the position of the calibration object can be used as an intermediary to indirectly deduce the relative positions between the multiple entities.
[0027] However, this method has the following disadvantages:
[0028] 1. Dependence on external markers: It is necessary to place calibration objects in the environment, which limits the flexibility of the usage scenarios.
[0029] 2. Marker types need to be predefined in advance: Taking the calibration plate as an example, the patterns drawn on the calibration plate need to be precisely predefined in advance, and the algorithm needs to be adapted to this definition, resulting in a high algorithm adaptation cost.
[0030] 3. Poor reliability: When the environment changes or the calibration object is blocked, the reliability of the system will be affected.
[0031] In view of at least one of the above problems, the present application proposes a method, apparatus, device and medium for determining the relative positions of multiple entities.
[0032] The following describes the method, apparatus, device and medium for determining the relative positions of multiple entities according to the embodiments of the present application with reference to the accompanying drawings.
[0033] Figure 1 It is a schematic flowchart of the method for determining the relative positions of multiple entities provided by an embodiment of the present application.
[0034] In the embodiment of the present application, the method for determining the relative positions of multiple entities is configured in a device for determining the relative positions of multiple entities as an example. The device for determining the relative positions of multiple entities can be applied to any electronic device so that the electronic device can perform the function of determining the relative positions of multiple entities.
[0035] Among them, the electronic device can be any device with computing capabilities. For example, it can be a personal computer (PC for short), an industrial computer, a host computer, a mobile terminal, a server, etc. The mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a personal digital assistant, a wearable device, etc. with various operating systems, touch screens and / or display screens.
[0036] As Figure 1 shown, the method includes the following steps:
[0037] Step S101, obtain image data of the environment where each entity is located through a vision sensor equipped with each entity among the multiple entities.
[0038] It should be noted that the present application does not limit the entity too much. For example, in a robot collaborative operation scenario, the entity can be a robot; in a formation driving scenario of an autonomous vehicle, the entity can be a vehicle; in an interaction scenario of multiple terminal devices, the entity can be a terminal device, and so on.
[0039] Among them, the vision sensor can be a monocular camera, a binocular camera, an RGB-D (Red-Green-Blue-Depth) camera, etc. The present application does not limit this.
[0040] In the embodiment of the present application, each entity among the multiple entities is equipped with a corresponding vision sensor. Thus, image data of the environment where each entity is located can be collected and obtained through the vision sensor equipped with each entity among the multiple entities.
[0041] As a possible implementation, the image data may include multiple consecutive image frames.
[0042] Step S102, for any entity, extract feature points from the image data corresponding to the entity to obtain multiple feature points corresponding to the entity.
[0043] Among them, the feature points can include corner points, edge points, spots, etc., and the present application does not limit this.
[0044] In the embodiments of the present application, for any entity, feature points can be extracted from the image data corresponding to the entity to obtain multiple feature points corresponding to the entity.
[0045] Optionally, a feature point detection algorithm can be used to detect or identify the feature points in the image data, and the detected or identified feature points can be described to obtain the corresponding feature point descriptors. For example, the SIFT (Scale-Invariant Feature Transform) algorithm, SURF (Speeded-Up Robust Features), ORB (Oriented FAST and Rotated BRIEF), etc. can be used to detect the feature points in the image data and describe the detected feature points to obtain their corresponding feature point descriptors.
[0046] Optionally, before extracting the feature points from the image data corresponding to the entity, the image data can be preprocessed. Among them, the preprocessing can include denoising, grayscale conversion, etc., and the present application does not limit this. Thereby, it can facilitate the subsequent processing of the image data and improve the accuracy and effectiveness of feature point extraction.
[0047] Step S103: Based on the multiple feature points corresponding to the entity, construct a sparse map of the entity.
[0048] In the embodiments of the present application, the multiple feature points corresponding to the entity can be used to construct a sparse map of the entity.
[0049] It should be noted that each entity has a corresponding entity coordinate system, where the entity coordinate system can be, for example, the world coordinate system of the corresponding entity.
[0050] As an example, the spatial coordinates of the respective feature points corresponding to the entity can be obtained, so that based on the spatial coordinates of the respective feature points corresponding to the entity in the corresponding entity coordinate system, the respective feature points corresponding to the entity can be fused to obtain a sparse map of the entity.
[0051] Step S104: Align the sparse maps corresponding to any two entities among the multiple entities to determine the relative positions between any two entities.
[0052] Among them, the relative position may include, for example, six - degree - of - freedom data, that is, it may include the position in space, the rotation angle, and may also include the relative distance, etc. For example, in a three - dimensional space, the coordinate system of the three - dimensional space is O - XYZ. The relative position may include the position x in the X - axis direction, the position y in the Y - axis direction, the position z in the Z - axis direction, the roll angle Roll of the rotation angle around the Z - axis, the pitch angle Pitch of the rotation angle around the Y - axis, the yaw angle Yaw of the rotation angle around the X - axis, and may also include the relative distance d.
[0053] In the multi - entity relative - position determination method according to the embodiments of the present application, image data of the environment where each entity is located is obtained through a vision sensor equipped on each entity among multiple entities; for any entity, feature points are extracted from the image data corresponding to the entity to obtain multiple feature points corresponding to the entity; based on the multiple feature points corresponding to the entity, a sparse map of the entity is constructed; the sparse maps corresponding to any two entities among the multiple entities are aligned to determine the relative position between any two entities. Thus, by constructing the sparse maps of each entity among multiple entities and aligning the sparse maps, the determination of the relative position between multiple entities can be automatically realized without relying on external calibration objects, which is applicable to various complex and dynamic environments, has strong adaptability and high flexibility, and since there is no need to pre - define the pattern or features of the calibration object, the complexity of algorithm development and adaptation is reduced, thereby reducing the adaptation cost and time; even in the case where some feature points are occluded or the environment changes, the sparse - map alignment and positioning can still be performed through the remaining feature points, improving the robustness, accuracy, and reliability of relative - position determination.
[0054] In the case where any two entities include a first entity and a second entity, in order to clearly illustrate how the sparse maps corresponding to any two entities among multiple entities are aligned to determine the relative position between any two entities in the above - mentioned embodiments of the present application, the present application also proposes a multi - entity relative - position determination method.
[0055] Figure 2 It is a schematic flowchart of the multi - entity relative - position determination method provided by another embodiment of the present application.
[0056] As Figure 2 shown, the method includes the following steps:
[0057] Step S201, obtaining image data of the environment where each entity is located through a vision sensor equipped on each entity among multiple entities.
[0058] Step S202, for any entity, extracting feature points from the image data corresponding to the entity to obtain multiple feature points corresponding to the entity.
[0059] Step S203, constructing a sparse map of the entity based on the multiple feature points corresponding to the entity.
[0060] It should be noted that the execution processes of steps S201 to S203 can refer to the relevant descriptions in any embodiment of this application, and will not be elaborated here.
[0061] Step S204, obtain the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities.
[0062] As a possible implementation, as Figure 3 shown, the following steps can be adopted to determine the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities:
[0063] Step S2041, determine the first quantity of the feature points in the sparse map of the first entity that match the feature points in the sparse map of the second entity.
[0064] In the embodiment of this application, for any feature point in the sparse map of the first entity, the feature point can be matched with each feature point in the sparse map of the second entity. When there is a first feature point in each feature point in the sparse map of the second entity that matches this feature point, it is determined that this feature point matches the feature points in the sparse map of the second entity. Thus, by accumulating the feature points in the sparse map of the first entity that match the feature points in the sparse map of the second entity, the first quantity can be obtained.
[0065] Among them, when matching any feature point in the sparse map of the first entity with each feature point in the sparse map of the second entity, for example, matching algorithms such as FLANN (Fast Library for Approximate Nearest Neighbors) and BFMatcher (Brute-Force Matcher) can be used.
[0066] For any feature point in the sparse map of the first entity, in order to accurately match this feature point with each feature point in the sparse map of the second entity, in a possible implementation of the embodiment of this application, for any feature point in the sparse map of the first entity, the distance between the feature point descriptor of this feature point and the feature point descriptor of any feature point in the sparse map of the second entity can be determined; when there is a second feature point in the sparse map of the second entity, and the distance between the feature point descriptor of this second feature point and the feature point descriptor of this feature point is less than the second set threshold, it is determined that the second feature point is the first feature point that matches this feature point.
[0067] Among them, it should be noted that when determining the distance between the feature point descriptor of any feature point in the sparse map of the first entity and the feature point descriptor of any feature point in the sparse map of the second entity, for example, a distance metric method (such as Euclidean distance, Manhattan distance, cosine similarity, etc.) can be used for determination.
[0068] Among them, the second set threshold can be preset, and the present application does not limit the value of the second set threshold.
[0069] Step S2042: Determine the first coefficient according to the first quantity and the total number of feature points in the sparse map of the first entity.
[0070] As an example, the ratio of the first quantity to the total number of feature points in the sparse map of the first entity can be determined as the first coefficient.
[0071] For example, assume that the total number of feature points in the sparse map of the first entity is a1, and the first quantity of the feature points in the sparse map of the first entity that match the feature points in the sparse map of the second entity is b1, then the first coefficient is (a1 / b1).
[0072] Step S2043: Determine the second quantity of the feature points in the sparse map of the second entity that match the feature points in the sparse map of the first entity.
[0073] It should be noted that the method for determining the second quantity is similar to the method for determining the above-mentioned first quantity, and will not be elaborated here.
[0074] Among them, it should also be noted that the first quantity can be the same as or different from the second quantity, and the present application does not limit this.
[0075] Step S2044: Determine the second coefficient according to the second quantity and the total number of feature points in the sparse map of the second entity.
[0076] As an example, the ratio of the second quantity to the total number of feature points in the sparse map of the second entity can be determined as the second coefficient.
[0077] For example, assume that the total number of feature points in the sparse map of the second entity is a2, and the second quantity of the feature points in the sparse map of the second entity that match the feature points in the sparse map of the first entity is b2, then the second coefficient is (a2 / b2).
[0078] It should be noted that this application does not limit the execution timing of steps S2041 to S2042 and steps S2043 to S2044. This application only takes the example that steps S2041 to S2042 are executed before steps S2043 to S2044. Steps S2041 to S2042 can also be executed in parallel with steps S2043 to S2044, or steps S2041 to S2042 can also be executed after steps S2043 to S2044.
[0079] Step S2045, based on the first coefficient and / or the second coefficient, determine the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities.
[0080] As an example, the first coefficient can be determined as the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities.
[0081] As another example, the second coefficient can be determined as the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities.
[0082] As still another example, a weighted average can be performed on the first coefficient and the second coefficient to obtain the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities.
[0083] Thus, the similarity between the sparse maps of any two entities can be effectively and accurately determined.
[0084] Step S205, in response to the similarity being greater than the first set threshold, align the sparse map of the first entity and the sparse map of the second entity to obtain the first relative pose between the first entity and the second entity.
[0085] Among them, the first set threshold can be preset, and this application does not limit the value of the first set threshold.
[0086] Among them, the first relative pose can include a translation vector and a rotation matrix.
[0087] In the embodiment of this application, when the similarity is greater than the first set threshold, it indicates that the first entity and the second entity among any two entities have observed similar scenes and a loop can be formed. At this time, the sparse map of the first entity and the sparse map of the second entity can be aligned to obtain the first relative pose between the first entity and the second entity.
[0088] As a possible implementation, each feature point in the sparse map of the first entity and each feature point in the sparse map of the second entity can be matched to obtain a matched first feature point pair; based on the matched first feature point pair, the transformation relationship between the entity coordinate system of the first entity and the entity coordinate system of the second entity can be determined; based on the transformation relationship, the first relative pose between the first entity and the second entity can be determined.
[0089] It should be noted that the number of the first feature point pairs is not limited in this application.
[0090] In the embodiments of this application, each feature point in the sparse map of the first entity and each feature point in the sparse map of the second entity can be matched to obtain a matched first feature point pair. For example, a feature matching algorithm (such as FLANN, BFMatcher, etc.) can be used to match each feature point in the sparse map of the first entity and each feature point in the sparse map of the second entity, so as to obtain a matched first feature point pair.
[0091] In the embodiments of this application, the transformation relationship between the entity coordinate system of the first entity and the entity coordinate system of the second entity can be determined based on the matched first feature point pair. As an example, the spatial coordinates of each feature point in the matched first feature point pair in the corresponding entity coordinate system can be obtained, and then, one of the least squares method, the Iterative Closest Point (ICP) algorithm, etc. can be used to determine the transformation relationship between the entity coordinate system of the first entity and the entity coordinate system of the second entity.
[0092] In the embodiments of this application, the first relative pose between the first entity and the second entity can be determined based on the transformation relationship.
[0093] As an example, it is assumed that the transformation relationship between the entity coordinate system of the first entity and the entity coordinate system of the second entity is shown in the following formula:
[0094] ; (1)
[0095] where, P 1 represents the spatial coordinate of the feature point of the first entity in the entity coordinate system of the first entity, which is a vector with 3 rows and 1 column; P 2 is the spatial coordinate of the feature point of the second entity in the entity coordinate system of the second entity, which is a vector with 3 rows and 1 column; R is a rotation matrix with 3 rows and 3 columns, representing the rotation operation from the entity coordinate system of the second entity to the entity coordinate system of the first entity; T is a translation vector with 3 rows and 1 column, representing the translation operation from the entity coordinate system of the second entity to the entity coordinate system of the first entity;
[0096] Furthermore, according to formula (1), the first relative pose between the first entity and the second entity can be determined as follows: the rotation matrix of the first entity relative to the second entity is R, and the translation vector is T.
[0097] As another example, assume that the conversion relationship between the entity coordinate system of the first entity and the entity coordinate system of the second entity is shown in the following formula:
[0098] ; (2)
[0099] where P 1 represents the spatial coordinates of the feature point of the first entity in the entity coordinate system of the first entity, which is a vector with 3 rows and 1 column; P 2 is the spatial coordinates of the feature point of the second entity in the entity coordinate system of the second entity, which is a vector with 3 rows and 1 column; R' is a 3×3 rotation matrix representing the rotation operation from the entity coordinate system of the second entity to the entity coordinate system of the first entity; T' is a 3×1 translation vector representing the translation operation from the entity coordinate system of the second entity to the entity coordinate system of the first entity.
[0100] Furthermore, according to formula (2), the first relative pose between the first entity and the second entity can be determined as follows: the rotation matrix of the second entity relative to the first entity is R', and the translation vector is T'.
[0101] Thus, by finding the matching feature point pairs in the sparse maps of any two entities, the conversion relationship between the entity coordinate systems of any two entities can be effectively determined. Furthermore, based on the conversion relationship, the first relative pose between the first entity and the second entity can be effectively and accurately determined.
[0102] In order to improve the accuracy of the feature point matching process, in a possible implementation manner of the embodiment of the present application, after obtaining the matching first feature point pair, the RANSAC (Random Sample Consensus) algorithm can be used to perform geometric consistency verification on the obtained matching first feature point pair. Thus, the incorrect matches in the feature point matching process can be effectively removed, and the accuracy and robustness of subsequent tasks can be improved.
[0103] It can be understood that there may be a situation where the similarity is not greater than the first set threshold. At this time, it indicates that no similar scenarios have been observed between the first entity and the second entity among any two entities. Therefore, in a possible implementation manner of the embodiments of the present disclosure, when the similarity is not greater than the first set threshold, it is possible to determine whether there is a target entity among multiple entities; wherein, the similarity between the sparse map of the target entity and the sparse map of the first entity is greater than the first set threshold, and the similarity between the sparse map of the target entity and the sparse map of the second entity is greater than the first set threshold; when there is a target entity among multiple entities, based on the conversion relationship between the entity coordinate system of the first entity and the entity coordinate system of the target entity, and the conversion relationship between the entity coordinate system of the second entity and the entity coordinate system of the target entity, the first relative pose between the first entity and the second entity can be determined.
[0104] It should be noted that the method for obtaining the similarity between the sparse map of the target entity and the sparse map of the first entity, and the method for obtaining the similarity between the sparse map of the target entity and the sparse map of the second entity are similar to the method for obtaining the similarity between the sparse map of the first entity and the sparse map of the second entity in step S204, and will not be elaborated here.
[0105] In the embodiments of the present application, when there is a target entity among multiple entities, based on the conversion relationship between the entity coordinate system of the first entity and the entity coordinate system of the target entity, and the conversion relationship between the entity coordinate system of the second entity and the entity coordinate system of the target entity, the first relative pose between the first entity and the second entity can be determined.
[0106] As an example, when there is a target entity among multiple entities, assume that the conversion relationship between the entity coordinate system of the first entity and the entity coordinate system of the target entity is as shown in the following formula:
[0107] ; (3)
[0108] where, P 1 represents the spatial coordinates of the feature points of the first entity in the entity coordinate system of the first entity, which is a vector of 3 rows and 1 column; P is the spatial coordinates of the feature points of the target entity in the entity coordinate system of the target entity, which is a vector of 3 rows and 1 column; R is a 3-row and 3-column rotation matrix, representing the rotation operation from the entity coordinate system of the target entity to the entity coordinate system of the first entity; T is a 3-row and 1-column translation vector, representing the translation operation from the entity coordinate system of the target entity to the entity coordinate system of the first entity;
[0109] The conversion relationship between the entity coordinate system of the second entity and the entity coordinate system of the target entity is as shown in the following formula:
[0110] ; (4)
[0111] wherein, P 2 represents the spatial coordinates of the feature points of the second entity in the entity coordinate system of the second entity, which is a vector with 3 rows and 1 column; P is the spatial coordinates of the feature points of the target entity in the entity coordinate system of the target entity, which is a vector with 3 rows and 1 column; R' is a rotation matrix with 3 rows and 3 columns, representing the rotation operation from the entity coordinate system of the target entity to the entity coordinate system of the first entity; T' is a translation vector with 3 rows and 1 column, representing the translation operation from the entity coordinate system of the target entity to the entity coordinate system of the first entity;
[0112] Based on the conversion relationship between the entity coordinate system of the first entity and the entity coordinate system of the target entity, and the conversion relationship between the entity coordinate system of the second entity and the entity coordinate system of the target entity, the first relative pose between the first entity and the second entity can be determined according to the following formula:
[0113] ; (5)
[0114] wherein, the first relative pose between the first entity and the second entity is: the rotation matrix of the second entity relative to the first entity is R'R, and the translation vector is (R'T + T').
[0115] Thus, when the similarity between the sparse map of the first entity and the sparse map of the second entity is not greater than the first set threshold, the first relative pose between the first entity and the second entity can be effectively and accurately determined by means of the target entity.
[0116] Step S206, analyze the first relative pose between the first entity and the second entity to obtain the relative position between the first entity and the second entity.
[0117] It should be noted that the explanation of the relative position in step S104 also applies to this embodiment and will not be elaborated here.
[0118] Specifically, when the first relative pose includes a translation vector and a rotation matrix, the spatial position of the first entity in the entity coordinate system of the second entity and / or the spatial position of the first entity in the entity coordinate system of the second entity, and the relative distance between the first entity and the second entity can be obtained according to the translation vector in the first relative pose between the first entity and the second entity; the rotation angle of the first entity around the corresponding coordinate axis of the entity coordinate system of the second entity can be obtained according to the rotation matrix in the first relative pose between the first entity and the second entity.
[0119] The method for determining the relative positions of multiple entities according to the embodiments of the present application obtains the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities; in response to the similarity being greater than the first set threshold, aligning the sparse map of the first entity and the sparse map of the second entity to obtain the first relative pose between the first entity and the second entity; analyzing the first relative pose between the first entity and the second entity to obtain the relative position between the first entity and the second entity. Thus, by calculating the similarity between the sparse maps of any two entities and aligning the sparse maps when the similarity meets certain conditions, the relative position between the entities is determined, improving the accuracy and efficiency of determining the relative position.
[0120] In the case where the image data corresponding to each entity includes multiple consecutive image frames, and correspondingly, each image frame has a corresponding plurality of feature points, in order to clearly illustrate how, for any entity in the above embodiments of the present application, a sparse map of the entity is constructed based on the plurality of feature points corresponding to the entity, the present application also proposes a method for determining the relative positions of multiple entities.
[0121] Figure 4 It is a schematic flowchart of the method for determining the relative positions of multiple entities provided by another embodiment of the present application.
[0122] As Figure 4 shown, the method includes the following steps:
[0123] Step S401, obtaining, by a vision sensor equipped with each entity among multiple entities, the image data of the environment where the corresponding entity is located.
[0124] Step S402, for any entity, extracting feature points from the image data corresponding to the entity to obtain a plurality of feature points corresponding to the entity.
[0125] It should be noted that the execution processes of steps S401 to S402 can refer to the relevant descriptions in any embodiment of the present application, and will not be elaborated here.
[0126] Step S403, sorting in the order of the acquisition times of the multiple consecutive image frames corresponding to the entity to obtain the image sorting sequence of the entity.
[0127] Step S404, determining the camera coordinate system when the vision sensor equipped with the entity acquires the image frame with the serial number 1 in the image sorting sequence of the entity as the entity coordinate system of the entity.
[0128] That is to say, the camera coordinate system when the vision sensor equipped with the entity acquires the image frame with the serial number 1 in the image sorting sequence of the entity coincides with the entity coordinate system of the entity.
[0129] Step S405: According to the image sorting sequence of the entity, sequentially determine the spatial coordinates of each feature point corresponding to each image frame in the entity coordinate system of the entity.
[0130] Optionally, as Figure 5 shown, the following steps can be adopted to determine the spatial coordinates of each feature point corresponding to each image frame in the entity coordinate system of the entity:
[0131] Step S4051: For any feature point of the image frame with the serial number 1 in the image sorting sequence of the entity, based on the image coordinates of the feature point, determine the spatial coordinates of the feature point in the corresponding entity coordinate system.
[0132] Among them, the image coordinates can be used to indicate the coordinate position of the corresponding feature point in the corresponding image frame.
[0133] In the embodiments of the present application, the spatial coordinates of the feature point in the corresponding entity coordinate system can be determined based on the image coordinates of the feature point. Specifically, based on the image coordinates of the feature point and the internal parameters of the vision sensor equipped on the entity (such as focal length, principal point coordinates, etc.), triangulation of the feature point can be performed to obtain the spatial coordinates of the feature point in the corresponding entity coordinate system.
[0134] Step S4052: For the (i + 1)-th image frame in the image sorting sequence corresponding to the entity, match the multiple feature points of the i-th image frame and the multiple feature points of the (i + 1)-th image frame to obtain a matched second feature point pair.
[0135] Among them, i can be a positive integer other than 0.
[0136] It should be noted that the present application does not limit the number of the second feature point pairs.
[0137] As an example, for the (i + 1)-th image frame in the image sorting sequence corresponding to the entity, matching algorithms such as FLANN and BFMatcher can be used to match the multiple feature points of the i-th image frame and the multiple feature points of the (i + 1)-th image frame, so as to obtain a matched second feature point pair.
[0138] During the feature point matching process, due to factors such as image noise, illumination changes, and perspective differences, some incorrect matching pairs often occur, which will have a negative impact on subsequent data applications. To improve the accuracy of the feature point matching process, in a possible implementation manner of the embodiment of the present application, after obtaining the matched second feature point pairs, the RANSAC (Random Sample Consensus) algorithm can be used to perform geometric consistency verification on the obtained matched second feature point pairs. Thus, incorrect matches in the feature point matching process can be effectively removed, and the accuracy and robustness of subsequent tasks can be improved.
[0139] Step S4053: Based on the matched second feature point pairs, determine the second relative pose of the visual sensor equipped on the entity when collecting the (i + 1)-th image frame relative to when collecting the i-th image frame.
[0140] Among them, the second relative pose may include a translation vector and a rotation matrix.
[0141] In the embodiment of the present application, the second relative pose of the visual sensor equipped on the entity when collecting the (i + 1)-th image frame relative to when collecting the i-th image frame can be determined based on the matched second feature point pairs between the i-th image frame and the (i + 1)-th image frame. Specifically, the second relative pose of the visual sensor equipped on the entity when collecting the (i + 1)-th image frame relative to when collecting the i-th image frame can be determined based on the spatial coordinates of the feature points belonging to the i-th image frame and the image coordinates of the feature points belonging to the (i + 1)-th image frame among the matched second feature point pairs between the i-th image frame and the (i + 1)-th image frame. For example, the second relative pose of the visual sensor equipped on the entity when collecting the (i + 1)-th image frame relative to when collecting the i-th image frame can be determined by using a pose estimation algorithm (such as the PnP (Perspective-n-Points) algorithm) based on the spatial coordinates of the feature points belonging to the i-th image frame and the image coordinates of the feature points belonging to the (i + 1)-th image frame among the matched second feature point pairs between the i-th image frame and the (i + 1)-th image frame.
[0142] Step S4054: For any feature point of the (i + 1)-th image frame, based on the second relative pose and the image coordinates of the feature point, determine the spatial coordinates of the feature point in the corresponding entity coordinate system.
[0143] As an example, first, for any feature point in the (i + 1)-th image, based on the image coordinates of the feature point, determine the spatial coordinates of the feature point in the camera coordinate system when the vision sensor equipped on the entity captures the (i + 1)-th image frame. For example, based on the image coordinates of the feature point and the internal parameters of the vision sensor equipped on the entity (such as focal length, principal point coordinates, etc.), perform feature point triangulation on the feature point to obtain the spatial coordinates of the feature point in the camera coordinate system when the vision sensor equipped on the entity captures the (i + 1)-th image frame. Secondly, based on the second relative pose of the vision sensor equipped on the entity when capturing the (i + 1)-th image frame relative to when capturing the i-th image frame, the estimated pose of the camera coordinate system of the vision sensor equipped on the entity relative to the entity coordinate system of the entity when capturing the (i + 1)-th image frame can be determined. Finally, based on the estimated pose and the internal parameters of the vision sensor equipped on the entity (such as focal length, principal point coordinates, etc.), convert the spatial coordinates of the feature point in the camera coordinate system when the vision sensor equipped on the entity captures the (i + 1)-th image frame into the spatial coordinates in the entity coordinate system.
[0144] Among them, when determining the estimated pose of the camera coordinate system of the vision sensor equipped on the entity relative to the entity coordinate system of the entity based on the second relative pose of the vision sensor equipped on the entity when capturing the (i + 1)-th image frame relative to when capturing the i-th image frame, for example, based on the second relative pose of the vision sensor equipped on the entity when capturing the (i + 1)-th image frame relative to when capturing the i-th image frame, and the second relative poses corresponding to any adjacent image frames before the vision sensor captures the (i + 1)-th image frame, determine the estimated pose of the camera coordinate system of the vision sensor equipped on the entity relative to the entity coordinate system of the entity when capturing the (i + 1)-th image frame.
[0145] For example, assume i is 3, and the second relative pose of the vision sensor equipped on the entity when capturing the 4th image frame relative to when capturing the 3rd image frame is: translation vector T 34 and rotation matrix R 34 ; the second relative poses corresponding to any adjacent image frames before the vision sensor captures the 4th image frame include: the second relative pose of the vision sensor when capturing the 3rd image frame relative to when capturing the 2nd image frame is: translation vector T 23 and rotation matrix R 23 , and the second relative pose of the vision sensor when capturing the 2nd image frame relative to when capturing the 1st image frame is: translation vector T 12 and rotation matrix R 12 ; then the estimated pose of the camera coordinate system of the vision sensor equipped on the entity relative to the entity coordinate system of the entity when capturing the (i + 1)-th image frame is: rotation matrix (R 34 R 23 R12 ), translation vector (R 34 R 23 T 12 +R 34 T 23 +T 34 ). Thus, the estimated pose of the camera coordinate system of the vision sensor equipped on the entity relative to the entity coordinate system of the entity can be effectively determined when the (i + 1)-th image frame is collected.
[0146] In summary, according to the acquisition times of multiple consecutive image frames corresponding to the entity, the spatial coordinates of each feature point of the corresponding image frame can be determined frame by frame.
[0147] Step S406: Based on the spatial coordinates of each feature point corresponding to the entity in the corresponding entity coordinate system, construct a sparse map of the entity.
[0148] In the embodiment of the present application, based on the spatial coordinates of each feature point corresponding to the entity in the corresponding entity coordinate system, each feature point corresponding to the entity can be fused to obtain a sparse map of the entity.
[0149] Step S407: Align the sparse maps corresponding to any two entities among multiple entities to determine the relative positions between any two entities.
[0150] It should be noted that the execution process of step S407 can refer to the relevant descriptions in any embodiment of the present application, and will not be elaborated here.
[0151] In the method for determining the relative positions of multiple entities in the embodiment of the present application, by sorting in the order of the acquisition times of multiple consecutive image frames corresponding to the entity, an image sorting sequence of the entity is obtained; the camera coordinate system when the vision sensor equipped on the entity collects the image frame with the serial number 1 in the image sorting sequence of the entity is determined as the entity coordinate system of the entity; according to the image sorting sequence of the entity, the spatial coordinates of each feature point corresponding to each image frame in the entity coordinate system of the entity are determined in sequence; based on the spatial coordinates of each feature point corresponding to the entity in the corresponding entity coordinate system, a sparse map of the entity is constructed. Thus, the construction of the sparse maps of each entity is effectively realized, providing key data support for subsequent determination of the relative positions between multiple entities.
[0152] To clearly illustrate the method for determining the relative positions of multiple entities in the present application, the following will be described in detail with examples.
[0153] As an example, the method for determining the relative positions of multiple entities may include the following steps:
[0154] 1. For any one of multiple entities, perform real-time SLAM (Simultaneous Localization and Mapping) during the operation of the entity. Specifically:
[0155] 1.1 Data acquisition
[0156] Collect image data of the environment where the entity is located through at least one visual sensor equipped on the entity;
[0157] Among them, the visual sensor can be a monocular camera, a binocular camera, an RGB-D camera, etc., and this application does not limit this.
[0158] 1.2 Feature point extraction
[0159] Extract feature points from the image data to obtain multiple feature points corresponding to the entity. For example, algorithms such as SIFT, SURF, and ORB can be used to detect or identify the feature points of the image data, and describe the detected or identified feature points to obtain the feature point descriptors of the feature points;
[0160] Among them, the feature points can be but are not limited to corner points, edges, etc.
[0161] 1.3 Feature point matching
[0162] In the case where the image data includes multiple consecutive image frames, feature point matching can be performed between adjacent image frames among the multiple consecutive image frames to obtain a second pair of matching feature points. For example, matching algorithms such as FLANN and BFMatcher can be used to obtain a second pair of matching feature points between adjacent image frames.
[0163] It should be noted that incorrect matches may occur during the feature point matching process. To improve the accuracy of the matching, the RANSAC (Random Sample Consensus) algorithm can be used to perform geometric consistency verification on the obtained second pair of feature points to remove incorrect matches.
[0164] 1.4 Pose estimation
[0165] Based on the second pair of matching feature points between adjacent image frames, use a pose estimation algorithm (such as the PnP algorithm) to estimate the camera pose to obtain the second relative pose of the visual sensor when collecting the corresponding adjacent image frames. And through the multi-view set relationship, calculate the estimated poses of the visual sensors equipped on the entity when collecting different image frames.
[0166] 1.5 Sparse map construction
[0167] Based on the estimated pose, back-project the feature points to obtain the spatial coordinates of the feature points in the three-dimensional space, and based on the spatial coordinates of the feature points corresponding to the entity, fuse the feature points corresponding to the entity to obtain the sparse map of the entity.
[0168] Optionally, the Bundle Adjustment (BA) algorithm or filtering methods (such as EKF (Extended Kalman Filter), UKF (Unscented Kalman Filter), etc.) can be used to globally optimize the estimated pose and the spatial coordinates of the feature points to reduce errors and improve the accuracy of the map.
[0169] 2. Align the sparse maps of any two entities to determine the relative position between any two entities. Specifically:
[0170] 2.1 Feature point matching
[0171] Perform feature point matching on the feature points of the sparse maps of any two entities to obtain the first pair of matching feature points. For example, matching algorithms such as FLANN and BFMatcher can be used to obtain the first pair of matching feature points.
[0172] It should be noted that in step 1.2, after extracting the feature points of the entity, the feature descriptors of the feature points of the entity can be saved for subsequent data applications, such as feature point matching.
[0173] 2.2 Determination of the relative position between any two entities
[0174] Determine the similarity between the sparse maps of any two entities; when the similarity is greater than the first set threshold, it is determined that the visual sensors of the corresponding two entities have observed similar scenes and can form a loop. Thus, based on the first pair of matching feature points between the two entities, the transformation relationship between the world coordinate systems (denoted as entity coordinate systems in this application) of the two entities can be determined, and based on this transformation relationship, the relative position between the two entities can be determined.
[0175] Among them, based on the transformation relationship, the coordinates of one entity among any two entities can be transformed into the world coordinate system of the other entity, realizing the alignment of the sparse maps of the two entities.
[0176] Optionally, the aligned sparse maps can be fused to obtain the merged maps of the corresponding two entities, and so on, to achieve map fusion between multiple entities, which will not be elaborated here.
[0177] It is understandable that as the SLAM algorithms of each entity continue to run, the relative positions between entities can be updated and dynamically adjusted in real time, and the continuous operation of the SLAM process and map fusion algorithm can ensure that the complete merged map can be synchronously updated to each continuously running entity, and the new map areas explored by each continuously running entity can also be continuously passed to other entities by merging maps. In this way, the three-dimensional map reconstruction of a single entity can be expanded to the three-dimensional map reconstruction and fusion of multiple entities. As long as there is a common view or a common view area between any two entities in the multiple entities, the maps between the multiple entities can be fused by fusing map feature points, and the relative positions between the multiple entities can be determined. This method does not require the introduction of external markers or base stations, reduces deployment costs and usage costs, is easy to expand to multiple devices, and has strong environmental adaptability.
[0178] Optionally, for any entity, the entity can also be equipped with an IMU (Inertial Measurement Unit) or other auxiliary sensors to provide additional position information and motion data.
[0179] The multi-entity relative position determination method can be applied to a multi-entity relative position determination system, wherein the multi-entity relative position determination system comprises:
[0180] 1. Visual sensor, used to capture images in the environment and generate image data.
[0181] 2. IMU (optional), used to provide auxiliary positioning information, including acceleration and angular velocity data; it should be noted that IMU can help improve the positioning accuracy of the system in the case of rapid movement or visual sensor failure.
[0182] 3. SLAM algorithm unit, used to process visual sensor data and IMU data for real-time 3D map reconstruction and positioning. The SLAM algorithm unit includes front-end and back-end processing modules; the front-end is responsible for feature point extraction, matching and preliminary pose estimation, and the back-end is responsible for global optimization and map construction.
[0183] 4. Feature point extraction and matching module, used to extract feature points from image data and perform feature point matching on the feature points.
[0184] 5. Position optimization module, used to optimize and correct the estimated pose and constructed map to ensure the accuracy of the relative position.
[0185] 6. The map construction module is used to construct a sparse three-dimensional point cloud map (denoted as the sparse map in this application) based on the poses and feature points obtained from the SLAM algorithm unit, and update and maintain the map in real time, and perform global consistency optimization by detecting loop closure.
[0186] 7. The map fusion module is used to fuse the sparse three-dimensional point cloud maps corresponding to multiple entities. Specifically, by matching the feature points in the sparse three-dimensional point cloud maps of different entities, calculating the coordinate transformation matrix, and based on the coordinate transformation matrix, converting the map coordinate system of one entity to the map coordinate system of another entity to achieve the unification of multi-entity maps.
[0187] 8. The transmission and synchronization module is used to exchange map and location information between multiple entities through wired or wireless communication methods to ensure data synchronization among entities. It should be noted that the transmission and synchronization module supports real-time map fusion and location update to ensure the global consistency and real-time performance of the system.
[0188] The method for determining the relative positions of multiple entities in this application has at least the following advantages:
[0189] 1. No external markers are required: It gets rid of the dependence on external calibration objects and has a wider scope of application.
[0190] 2. Strong environmental adaptability: It has strong adaptability and is suitable for complex and dynamic environments. It is not restricted by the installation and maintenance of calibration objects, nor affected by occlusion.
[0191] 3. Quick addition of new devices: When a new device joins the corresponding system, no complex configuration and calibration process are required, which helps to ensure seamless integration and interoperability between the new device and the existing system.
[0192] For the method for determining the relative positions of multiple entities in this application, the inventors of this application have verified the feasibility of its technology through experiments. They have performed three-dimensional reconstruction and map fusion on multiple entities equipped with visual sensors in different environments, accurately determined the relative positions between multiple entities, and the experimental results show that this method can quickly and accurately determine the relative positions between multiple entities and maintain high real-time performance and stability in a dynamic environment.
[0193] To implement the above embodiments, this application also proposes a device for determining the relative positions of multiple entities.
[0194] Figure 6 It is a schematic structural diagram of the device for determining the relative positions of multiple entities provided by another embodiment of this application.
[0195] As Figure 6As shown in the figure, the multi-entity relative position determination device 600 includes: an acquisition module 610, an extraction module 620, a construction module 630, and an alignment module 640.
[0196] The acquisition module 610 is configured to obtain image data of the environment where the corresponding entity is located through a visual sensor equipped with each entity among the multiple entities.
[0197] The extraction module 620 is configured to, for any entity, extract feature points from the image data corresponding to the entity to obtain multiple feature points corresponding to the entity.
[0198] The construction module 630 is configured to construct a sparse map of the entity based on the multiple feature points corresponding to the entity.
[0199] The alignment module 640 is configured to align the sparse maps corresponding to any two entities among the multiple entities to determine the relative position between any two entities.
[0200] Further, in a possible implementation manner of the embodiment of the present application, any two entities include a first entity and a second entity; the alignment module 640 is configured to: obtain the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities; in response to the similarity being greater than a first set threshold, align the sparse map of the first entity and the sparse map of the second entity to obtain a first relative pose between the first entity and the second entity; analyze the first relative pose between the first entity and the second entity to obtain the relative position between the first entity and the second entity among any two entities.
[0201] Further, in a possible implementation manner of the embodiment of the present application, the alignment module 640 is configured to: determine a first quantity of feature points in the sparse map of the first entity that match the feature points in the sparse map of the second entity; determine a first coefficient according to the first quantity and the total quantity of feature points in the sparse map of the first entity; determine a second quantity of feature points in the sparse map of the second entity that match the feature points in the sparse map of the first entity; determine a second coefficient according to the second quantity and the total quantity of feature points in the sparse map of the second entity; determine the similarity between the sparse map of the first entity and the sparse map of the second entity among any two entities based on the first coefficient and / or the second coefficient.
[0202] Further, in a possible implementation manner of the embodiment of the present application, the alignment module 640 is configured to: match each feature point in the sparse map of the first entity with each feature point in the sparse map of the second entity to obtain a first pair of matching feature points; determine a transformation relationship between the entity coordinate system of the first entity and the entity coordinate system of the second entity based on the first pair of matching feature points; determine a first relative pose between the first entity and the second entity based on the transformation relationship.
[0203] Further, in a possible implementation manner of the embodiment of the present application, the device further includes a first determination module and a second determination module. The first determination module is configured to: in response to the similarity not being greater than a first set threshold, determine whether there is a target entity among multiple entities; wherein, the similarity between the sparse map of the target entity and the sparse map of the first entity is greater than the first set threshold, and the similarity between the sparse map of the target entity and the sparse map of the second entity is greater than the first set threshold. The second determination module is configured to: in response to the existence of the target entity among the multiple entities, determine a first relative pose between the first entity and the second entity based on the conversion relationship between the entity coordinate system of the first entity and the entity coordinate system of the target entity, and the conversion relationship between the entity coordinate system of the second entity and the entity coordinate system of the target entity.
[0204] Further, in a possible implementation manner of the embodiment of the present application, the image data includes multiple consecutive image frames. Correspondingly, each image frame has a corresponding multiple feature points. The construction module 630 is configured to: sort the multiple consecutive image frames collected by the entity in the order of acquisition time to obtain an image sorting sequence of the entity; determine the camera coordinate system when the visual sensor equipped with the entity collects the image frame with the serial number 1 in the image sorting sequence of the entity as the entity coordinate system of the entity; according to the image sorting sequence of the entity, sequentially determine the spatial coordinates of each feature point corresponding to each image frame in the entity coordinate system of the entity; and construct a sparse map of the entity based on the spatial coordinates of each feature point corresponding to the entity in the corresponding entity coordinate system.
[0205] Further, in a possible implementation manner of the embodiment of the present application, the construction module 630 is configured to: for any feature point of the image frame with the serial number 1 in the image sorting sequence of the entity, determine the spatial coordinate of the feature point in the corresponding entity coordinate system based on the image coordinate of the feature point; for the (i + 1)-th image frame in the image sorting sequence corresponding to the entity, match the multiple feature points of the i-th image frame and the multiple feature points of the (i + 1)-th image frame to obtain a matched second feature point pair; wherein, i is a positive integer other than 0; determine a second relative pose of the visual sensor equipped with the entity when collecting the (i + 1)-th image frame relative to when collecting the i-th image frame based on the matched second feature point pair; and for any feature point of the (i + 1)-th image frame, determine the spatial coordinate of the feature point in the corresponding entity coordinate system based on the second relative pose and the image coordinate of the feature point.
[0206] It should be noted that the foregoing explanation of the embodiment of the multi-entity relative position determination method is also applicable to the multi-entity relative position determination device of this embodiment, and will not be elaborated here.
[0207] In summary, the multi-entity relative position determination device according to the embodiments of the present application can automatically determine the relative positions between multiple entities by constructing a sparse map for each entity among the multiple entities and aligning the sparse maps, without relying on external calibration objects. It is applicable to various complex and dynamic environments, with strong adaptability and high flexibility. Moreover, since there is no need to pre-define the patterns or features of the calibration objects, the complexity of algorithm development and adaptation is reduced, thereby reducing the adaptation cost and time. Even when some feature points are occluded or the environment changes, it can still perform sparse map alignment and positioning through the remaining feature points, improving the robustness, accuracy, and reliability of relative position determination.
[0208] To implement the above embodiments, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the multi-entity relative position determination method provided in the foregoing embodiments.
[0209] To implement the above embodiments, the present application also proposes a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are used to implement the multi-entity relative position determination method provided in the foregoing embodiments when executed by a processor.
[0210] To implement the above embodiments, the present application also proposes a computer program product including a computer program, and the computer program implements the multi-entity relative position determination method provided in the foregoing embodiments when executed by a processor.
[0211] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0212] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.
[0213] This application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of users.
[0214] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0215] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as controlling or implying relative importance or implicitly indicating the quantity of the controlled technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0216] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred implementation of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0217] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0218] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0219] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above-described embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0220] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0221] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A method for automatically determining the relative positions of multiple entities, characterized in that: The method comprises: Acquire image data of the environment in which the corresponding entity is located through a visual sensor equipped by each of the multiple entities, wherein the image data includes multiple continuous image frames, and each of the image frames has corresponding multiple feature points; Extracting feature points from a plurality of continuous image frames corresponding to each entity to obtain a plurality of feature points corresponding to each entity; Sorting the multiple continuous image frames in the order of acquisition time to obtain an image sorting sequence for each entity, determining the camera coordinate system of the image frame with a sequence number of 1 in the image sorting sequence of each entity as the entity coordinate system, and sequentially determining the spatial coordinates of each feature point corresponding to each image frame in the image sorting sequence of each entity in the entity coordinate system according to the image sorting sequence of each entity, and constructing a sparse map of each entity based on the spatial coordinates of each feature point corresponding to each image frame in the image sorting sequence of each entity in the entity coordinate system; Obtain a similarity between sparse maps of a first entity and a second entity among the multiple entities, and when the similarity is greater than a first set threshold, align the sparse map of the first entity with the sparse map of the second entity to obtain a first relative pose between the first entity and the second entity; when the similarity is not greater than the first set threshold and there is a target entity among the multiple entities, determine a first relative pose between the first entity and the second entity based on a conversion relationship between a physical coordinate system of the first entity and a physical coordinate system of the target entity, and a conversion relationship between a physical coordinate system of the second entity and a physical coordinate system of the target entity; and parse the first relative pose to determine the relative position between any two entities.
2. The method according to claim 1, characterized in that Obtaining a similarity between a sparse map of a first entity and a second entity in the plurality of entities, comprising: determining a first number of feature points in the sparse map of the first entity that match feature points in the sparse map of the second entity; determining a first coefficient according to the first number and a total number of feature points in the sparse map of the first entity; determining a second number of feature points in the sparse map of the second entity that match feature points in the sparse map of the first entity; determining a second coefficient according to the second number and a total number of feature points in the sparse map of the second entity; Based on the first coefficient and / or the second coefficient, a similarity between the sparse map of the first entity and the sparse map of the second entity in any two entities is determined.
3. The method according to claim 1, characterized in that When the similarity is greater than a first set threshold, aligning the sparse map of the first entity and the sparse map of the second entity to obtain a first relative position between the first entity and the second entity, including: Matching each feature point in the sparse map of the first entity with each feature point in the sparse map of the second entity to obtain a matched first feature point pair; Determining a conversion relationship between a physical coordinate system of the first entity and a physical coordinate system of the second entity based on the matched first feature point pair; Based on the transformation relationship, a first relative position and posture between the first entity and the second entity is determined.
4. The method according to claim 1, characterized in that: According to the image sorting sequence of each entity, sequentially determining the spatial coordinates of each feature point corresponding to each image frame in the image sorting sequence of each entity in the entity coordinate system, including: For any of the feature points of the image frame with a sequence number of 1 in the image sorting sequence of each entity, determine the spatial coordinates of the feature point in the corresponding entity coordinate system based on the image coordinates of the feature point; For the i+1th image frame in the image sorting sequence corresponding to each entity, multiple feature points of the i+1th image frame are matched with multiple feature points of the i+1th image frame to obtain a matched second feature point pair; wherein i is a positive integer not equal to 0; Based on the matched second feature point pair, determining a second relative posture of the visual sensor equipped by the entity when capturing the (i+1)th image frame relative to when capturing the (i)th image frame; For any of the feature points in the (i+1)th image frame, based on the second relative posture and the image coordinates of the feature point, determine the spatial coordinates of the feature point in the corresponding physical coordinate system.
5. A device for automatically determining the relative positions of multiple entities, characterized in that: The device comprises: An acquisition module, configured to acquire image data of an environment in which a corresponding entity is located through a visual sensor equipped by each of the multiple entities, wherein the image data includes multiple continuous image frames, and each of the image frames has a corresponding plurality of feature points; An extraction module, used for extracting feature points from a plurality of continuous image frames corresponding to each entity, so as to obtain a plurality of feature points corresponding to each entity; A construction module, used to sort the multiple continuous image frames according to the order of acquisition time to obtain an image sorting sequence of each entity, determine the camera coordinate system of the image frame with sequence number 1 in the image sorting sequence of each entity as the entity coordinate system, and sequentially determine the spatial coordinates of each feature point corresponding to each image frame in the image sorting sequence of each entity in the entity coordinate system according to the image sorting sequence of each entity, and construct a sparse map of each entity based on the spatial coordinates of each feature point corresponding to each image frame in the image sorting sequence of each entity in the entity coordinate system; An alignment module is used to obtain the similarity between the sparse maps of a first entity and a second entity among the multiple entities, and when the similarity is greater than a first set threshold, align the sparse map of the first entity and the sparse map of the second entity to obtain a first relative posture between the first entity and the second entity; when the similarity is not greater than the first set threshold and there is a target entity among the multiple entities, determine the first relative posture between the first entity and the second entity based on the conversion relationship between the physical coordinate system of the first entity and the physical coordinate system of the target entity, and the conversion relationship between the physical coordinate system of the second entity and the physical coordinate system of the target entity; and parse the first relative posture to determine the relative position between any two entities.
6. The device according to claim 5, characterized in that Obtaining a similarity between a sparse map of a first entity and a second entity in the plurality of entities, comprising: determining a first number of feature points in the sparse map of the first entity that match feature points in the sparse map of the second entity; determining a first coefficient according to the first number and a total number of feature points in the sparse map of the first entity; determining a second number of feature points in the sparse map of the second entity that match feature points in the sparse map of the first entity; determining a second coefficient according to the second number and a total number of feature points in the sparse map of the second entity; Based on the first coefficient and / or the second coefficient, a similarity between the sparse map of the first entity and the sparse map of the second entity in any two entities is determined.
7. The device according to claim 6, characterized in that When the similarity is greater than a first set threshold, aligning the sparse map of the first entity and the sparse map of the second entity to obtain a first relative position between the first entity and the second entity, including: Matching each feature point in the sparse map of the first entity with each feature point in the sparse map of the second entity to obtain a matched first feature point pair; Determining a conversion relationship between a physical coordinate system of the first entity and a physical coordinate system of the second entity based on the matched first feature point pair; Based on the transformation relationship, a first relative position and posture between the first entity and the second entity is determined.
8. The device according to claim 5, characterized in that According to the image sorting sequence of each entity, sequentially determining the spatial coordinates of each feature point corresponding to each image frame in the image sorting sequence of each entity in the entity coordinate system, including: For any of the feature points of the image frame with a sequence number of 1 in the image sorting sequence of each entity, determine the spatial coordinates of the feature point in the corresponding entity coordinate system based on the image coordinates of the feature point; For the i+1th image frame in the image sorting sequence corresponding to each entity, multiple feature points of the i+1th image frame are matched with multiple feature points of the i+1th image frame to obtain a matched second feature point pair; wherein i is a positive integer not equal to 0; Based on the matched second feature point pair, determining a second relative posture of the visual sensor equipped by the entity when capturing the (i+1)th image frame relative to when capturing the (i)th image frame; For any of the feature points in the (i+1)th image frame, based on the second relative posture and the image coordinates of the feature point, determine the spatial coordinates of the feature point in the corresponding physical coordinate system.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 4 when executed by a processor.
11. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 4 when being executed by a processor.
Citation Information
Patent Citations
Crowdsourcing and distributing a sparse map, and lane measurements for autonomous vehicle navigation
CN109643367A