Robot multi-camera positioning key frame selection method and device and storage medium

CN121999033APending Publication Date: 2026-05-08ZHEJIANG SUNSEEKER IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG SUNSEEKER IND CO LTD
Filing Date
2024-11-03
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing robot vision positioning systems cannot obtain accurate robot poses when the field of view of a single vision sensor is obstructed by obstacles or there are no obvious features within the field of view, resulting in inaccurate positioning.

Method used

A multi-camera positioning method is adopted, which generates a keyframe sequence by selecting keyframes from real-time frames of multiple cameras, and optimizes pose calculation using bundle adjustment method to improve positioning accuracy.

Benefits of technology

By using multiple cameras in collaboration, the robot's pose calculation is optimized, improving the accuracy and stability of robot localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999033A_ABST
    Figure CN121999033A_ABST
Patent Text Reader

Abstract

The invention provides a robot multi-camera positioning key frame selection method and device and a storage medium, and is applied to the field of automatic control, and the method comprises the steps: carrying out the key frame selection of a real-time frame of a first camera, obtaining a plurality of first key frames, and generating a first key frame sequence; performing key frame selection on the real-time frame of the second camera to obtain a plurality of second key frames, and generating a second key frame sequence; and performing optimization pose calculation based on the first key frame sequence and the second key frame sequence to obtain a plurality of first poses corresponding to the first camera and a plurality of second poses corresponding to the second camera. Optimization calculation is carried out on the pose of the robot by using the key frame, and the positioning accuracy of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot localization, and in particular to a method, apparatus and storage medium for selecting keyframes for multi-camera robot localization. Background Technology

[0002] With the development of technology, robots are increasingly used in human production and daily life, often replacing human labor in tasks such as handling, inspection, and cleaning. Robots are typically equipped with sensors, navigation and positioning systems, and control algorithms, enabling them to perceive their surroundings and take appropriate actions. Currently, there are various types of robots on the market that assist people in completing tasks, such as robot vacuums, lawnmowers, and vacuum cleaners, providing convenience for people's lives and work.

[0003] One of the keys to achieving robot intelligence and automation is robot localization technology. Visual localization systems are widely used in robot localization due to their wide applicability and low cost. Existing visual localization systems mainly consist of a single visual sensor, such as a single monocular camera, a single binocular camera, or a single depth camera. However, since visual localization systems primarily rely on visual sensors to identify environmental features and calculate their pose, when the robot's single visual sensor's field of view is obstructed by obstacles or lacks obvious features, it may not be able to obtain a relatively accurate robot pose. This can lead to the visual localization system failing to function effectively and affecting the robot's overall operation. Therefore, when a robot has multiple visual sensors, how to rationally select real-time frame images acquired by the visual sensors as keyframes to optimize robot pose calculations is a problem that urgently needs to be solved. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, and storage medium for selecting keyframes in multi-camera robot localization. This method enables the reasonable selection of keyframes acquired by multiple cameras during multi-camera localization, allowing for optimized calculation of the robot's pose using these keyframes and improving the accuracy of robot localization. The technical solution is as follows:

[0005] In a first aspect, embodiments of this application provide a method for selecting keyframes for robot multi-camera localization, the method comprising: selecting keyframes from real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence;

[0006] Keyframes are selected from the real-time frames of the second camera to obtain multiple second keyframes, and a second keyframe sequence is generated.

[0007] Based on the first keyframe sequence and the second keyframe sequence, optimized pose calculation is performed to obtain multiple first poses corresponding to the first camera and multiple second poses corresponding to the second camera. The multiple first poses and the multiple second poses are used to optimize the robot pose within a specific time period.

[0008] Furthermore, the second keyframe sequence includes second camera real-time frames at the same time as the plurality of first keyframes.

[0009] Furthermore, the step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence includes:

[0010] Feature matching and filtering are performed on the real-time frames of the first camera to obtain the plurality of first key frames and generate the first key frame sequence;

[0011] The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence includes:

[0012] Select a real-time frame from the second camera that occurs at the same time as the plurality of first keyframes, and use it as the second keyframe.

[0013] Furthermore, the step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence includes:

[0014] Feature matching and filtering are performed on the real-time frames of the first camera to obtain the plurality of first key frames and generate the first key frame sequence;

[0015] The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence includes:

[0016] Feature matching and filtering are performed on the real-time frames of the second camera to obtain the plurality of second key frames, and the second key frame sequence is generated.

[0017] Furthermore, the step of selecting keyframes from the real-time frames of the second camera... Multiple second keyframes are obtained, and a second keyframe sequence is generated, including:

[0018] Select a real-time frame from the second camera that is at the same time as the plurality of first keyframes, and add it to the second keyframe sequence as the second keyframe.

[0019] Furthermore, the step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes, and the step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes, includes:

[0020] Feature matching and filtering are performed on the real-time frames of the first camera to obtain multiple first candidate keyframes;

[0021] Feature matching and filtering are performed on the real-time frames of the second camera to obtain multiple second candidate keyframes;

[0022] Filter the first candidate keyframe and the second candidate keyframe at the same time. These are respectively used as the first keyframe and the second keyframe.

[0023] Furthermore, before selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence, the method further includes:

[0024] The real-time frame of the first camera is selected as the first candidate real-time frame based on a specific inter-frame interval.

[0025] The step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence includes:

[0026] From the first candidate real-time frame, key frames are selected from the real-time frames of the first camera to obtain multiple first key frames.

[0027] Furthermore, before selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence, the method further includes:

[0028] The real-time frame of the second camera is selected as the second candidate real-time frame based on a specific inter-frame interval.

[0029] The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a first keyframe sequence includes:

[0030] From the second candidate real-time frames, key frames are selected from the real-time frames of the second camera to obtain multiple second key frames.

[0031] Furthermore, the feature matching and filtering includes the following methods:

[0032] The image features of a real-time frame from the first camera are matched with at least one other first keyframe to obtain the number of matched feature points; or the real-time frame from the second camera is matched with at least one other second keyframe to obtain the number of matched feature points.

[0033] If the number of matched feature points is greater than or equal to the matching threshold, the corresponding real-time frame of the first camera is used as the first key frame, or the corresponding real-time frame of the second camera is used as the second key frame.

[0034] Furthermore, the optimized pose calculation based on the first keyframe sequence and the second keyframe sequence includes:

[0035] Based on the first keyframe and the second keyframe in the first keyframe sequence and the second keyframe sequence, bundle adjustment optimization is performed to obtain the first pose corresponding to the first camera and the second pose corresponding to the second camera.

[0036] Secondly, embodiments of this application provide a robot multi-camera keyframe selection device, the device comprising:

[0037] The first selection module is used to select keyframes from the real-time frames of the first camera. Multiple first keyframes are obtained, and a first keyframe sequence is generated. The first keyframes reflect the key poses of the first camera, and the first keyframe sequence reflects the key pose changes of the first camera.

[0038] The second selection module is used to select keyframes from the real-time frames of the second camera. Multiple second keyframes are obtained, and a second keyframe sequence is generated. The second keyframes reflect the key poses of the second camera, and the second keyframe sequence reflects the key pose changes of the second camera.

[0039] The pose calculation module is used to perform optimized pose calculation based on the first keyframe sequence and the second keyframe sequence to obtain multiple first poses corresponding to the first camera and multiple second poses corresponding to the second camera. The multiple first poses and the multiple second poses are used to optimize the robot's historical robot poses within a specific time period.

[0040] Thirdly, embodiments of this application provide a robot multi-camera keyframe selection device. The device includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to achieve robot multi-camera keyframe selection as described above.

[0041] Fourthly, embodiments of this application provide a computer-readable storage medium storing at least one program, which is loaded and executed by a processor to implement the robot multi-camera localization keyframe selection method as described above. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 A flowchart illustrating a robot multi-camera localization keyframe selection method provided in an exemplary embodiment of this application is shown.

[0044] Figure 2 This illustration shows a schematic diagram of a keyframe selection process provided in an exemplary embodiment of this application;

[0045] Figure 3 This illustration shows a schematic diagram of a keyframe selection process provided by another exemplary embodiment of this application;

[0046] Figure 4 This illustration shows a schematic diagram of a keyframe selection process provided by another exemplary embodiment of this application;

[0047] Figure 5 This illustration shows a schematic diagram of a keyframe selection process provided by another exemplary embodiment of this application;

[0048] Figure 6 A flowchart illustrating a robot multi-camera localization optimization method provided in an exemplary embodiment of this application is shown.

[0049] Figure 7 This invention provides a schematic diagram of a robot localization optimization process according to an exemplary embodiment of the present application.

[0050] Figure 8 This application illustrates a schematic diagram of robot movement provided in an exemplary embodiment.

[0051] Figure 9 This invention provides a structural block diagram of a robot multi-camera localization keyframe selection device according to an exemplary embodiment of the present application.

[0052] Figure 10 A structural block diagram of a robot multi-camera positioning keyframe selection device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0053] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0054] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] It should be noted that the terms "first," "second," etc., used in the specification, claims, and drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0056] The following describes the robot multi-camera localization keyframe selection method of the present invention. This specification provides the operation steps of the method as described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operation steps may be included. The order of steps listed in the embodiments is merely one of many possible execution orders and does not represent the only execution order. In actual execution, the method can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment) as shown in the embodiments or drawings.

[0057] The robot is a device with autonomous walking capabilities, including but not limited to lawnmowers, sweeping robots, and transport robots. Furthermore, to achieve multi-camera positioning, the robot is equipped with a first camera and at least one second camera. The first camera may be a single monocular camera, a single binocular camera, or a single depth camera, and the second camera may be a single monocular camera, a single binocular camera, or a single depth camera.

[0058] When a robot performs multi-camera localization, multiple cameras can acquire environmental images in real time. By processing these images, the camera pose at a given moment can be obtained, and the robot pose can then be derived from these poses to complete the localization process. However, because each camera is positioned and oriented differently, the acquired images differ. Although camera and robot poses can be calculated instantly from the real-time frames acquired by each camera, these frames and the keyframes selected from each frame are independent of each other. This makes it difficult to effectively utilize information from each other to constrain localization calculations and optimize the localization effect. Therefore, it is crucial to accurately select keyframes from the real-time frames acquired by the cameras that reflect the constraints between the poses of multiple cameras. Robot poses derived from keyframes with constraints are more accurate, allowing for updates to historical robot poses and increasing the accuracy of robot localization.

[0059] Specifically, such as Figure 1 As shown, Figure 1 This is a schematic diagram of a process for selecting keyframes for multi-camera localization of a robot, provided by an embodiment of the present invention. The above method may include:

[0060] S101. Select keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generate a first keyframe sequence.

[0061] The aforementioned first keyframe reflects the key pose of the first camera, and the sequence of first keyframes reflects the changes in the key pose of the first camera. For example, in one possible implementation, the first and second cameras can acquire real-time frames during the robot's movement, thus allowing direct extraction of multiple first and second keyframes from the real-time frames of the first and second cameras within a specific time period.

[0062] In some embodiments, keyframes may be selected in the following manner.

[0063] 1. Determine whether a specific frame / time has elapsed since the last keyframe insertion; specifically, a specific frame interval can be set to uniformly extract real-time frames as keyframes within a specific time period; or, a specific frame interval can be set to select the next real-time frame as a keyframe after a specific frame interval, i.e., a specific time or a specific number of real-time frames, following the last keyframe selection.

[0064] 2. Determine the number of image feature points that match the real-time frame with other selected keyframes, or determine whether the number of interior points among the matched feature points that meet the reprojection error threshold is greater than a specific threshold.

[0065] The two methods described above can be used individually or in combination. That is, after a specific time / frame count, it is determined whether the real-time frame meets the keyframe condition. If it does, another keyframe is added. After adding a keyframe, the above keyframe selection steps are repeated until keyframes have been selected for all real-time frames within the specific time period.

[0066] S102. Select keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generate a second keyframe sequence.

[0067] The aforementioned second keyframe reflects the key pose of the second camera, and the sequence of second keyframes reflects the key pose changes of the second camera.

[0068] Similar to the first keyframe selection method, the second keyframe can also be selected by determining whether a specific frame / time has passed since the last keyframe insertion, and by determining the number of image feature points that match the real-time frame of the second camera with other selected keyframes.

[0069] In existing technologies, if a robot is equipped with multiple cameras (vision sensors), and keyframes from all cameras are selected uniformly, when the fields of view of the multiple cameras do not overlap or have a small overlap range, the keyframes of one camera may have fewer matching feature points, or even no matching feature points for a certain period of time. This can cause previous keyframes from one camera to exclude real-time frames from another camera, preventing the real-time frames from being added to the keyframes when the angle between the multiple cameras is large, thus making it impossible to utilize the information from that camera for optimization. In the embodiments of this application, the first camera and the second camera each maintain a keyframe sequence, and the first camera and the second camera each select real-time frame images to add to their respective keyframe sequences, which can avoid the problem of keyframe exclusion between different cameras in existing technologies.

[0070] S103. Optimize pose calculation based on the first keyframe sequence and the second keyframe sequence to obtain multiple first poses corresponding to the first camera and multiple second poses corresponding to the second camera.

[0071] The aforementioned multiple first poses and multiple second poses are used to optimize the robot's historical pose within a specific time period. Optimizing the pose using the first and second keyframes yields relatively accurate first and second camera poses. The accurate robot pose can be obtained by utilizing the positional transformation relationship between the first and second cameras and the robot, thus optimizing the robot's historical pose within a specific time period. The optimized pose calculation can employ localization optimization algorithms such as bundle adjustment.

[0072] In some embodiments, the first camera and the second camera can acquire real-time frames during the robot's movement and calculate the real-time pose of the first camera and the second camera in real time.

[0073] Specifically, based on the real-time frames of the first and second cameras, feature points are extracted from the real-time frame images of the first and second cameras using feature detection algorithms such as ORB-FAST corner detection. The feature points of the real-time frame images of the first and second cameras are then matched with the feature points of their respective previous frames. Feature point matching can employ matching algorithms such as ORB-BRIEF descriptors and optical flow methods to obtain feature matching relationships and acquire multiple matching feature point pairs. Then, the pixel coordinates of the matching feature points in the two frames are calculated. The essential matrix E or homography matrix H is obtained using epipolar geometry constraints. Based on the essential matrix E or homography matrix H, the camera transformation matrix T or the [R, t] rotation / translation relationship is calculated. Then, based on the camera transformation and the camera's initial pose or the pose of the previous frame, the current position and pose of the first and second cameras are obtained. Alternatively, after obtaining multiple pairs of matching feature points, the pixel coordinates of the matching feature points in the current frame and the 3D coordinates of the feature points in the camera coordinate system or world coordinate system of the previous frame are obtained. The camera pose transformation matrix T or the [R, t] rotation / translation relationship is solved using the PNP method or the BundleAdjustment method. Then, based on the camera transformation and the initial pose or the pose of the previous frame, the current position and pose of the camera are obtained. The 3D coordinates of the feature points in the camera coordinate system or world coordinate system of the previous frame can be calculated by a stereo camera or a depth camera, or by a monocular camera after obtaining the camera poses of the two frames, calculating the depth of the matching feature points using a triangulation method, and then obtaining the 3D coordinates of the feature points.

[0074] The robot's movement is controlled based on the current pose information of the first or second camera. Specifically, when controlling the robot's movement based on the current pose information of the first or second camera, the robot's movement can be controlled directly using the pose information of the first or second camera, or it can be controlled using other pose information calculated from the pose information of the first or second camera and the robot's own attributes. Examples include the body center pose, the four corner poses, the drive wheel center pose, and the left and right drive wheel poses calculated from the first / second camera, all within the scope defined in this solution. For instance, the body pose can be obtained from the pose of the first camera through a pre-determined transformation relationship between the first camera and the body center, or it can be obtained from the pose of the second camera through a pre-determined transformation relationship between the second camera and the body center.

[0075] The real-time camera pose obtained above contains some error due to the use of real-time frames. The robot pose obtained based on the camera-robot position transformation relationship also contains errors, leading to inaccurate robot localization. However, after selecting keyframes, the first and second poses calculated based on the first and second keyframes obtained from the two cameras are more accurate. This allows for updating the poses obtained by the cameras in real-time within a specific time period, resulting in a more accurate robot pose and thus improving the accuracy of robot localization.

[0076] When multiple cameras maintain their own keyframe sequences and a real-time frame is selected as a keyframe, a time asynchrony may occur between the keyframes of different cameras. This prevents the utilization of the fixed pose transformation relationships of synchronized image frames from different cameras, hindering subsequent optimization of the robot's pose based on the keyframes. Therefore, the following method can be used to select the first and second keyframes:

[0077] First, feature matching and filtering are performed on the real-time frames of the first camera to obtain multiple first key frames and generate a first key frame sequence.

[0078] Then, real-time frames from the second camera at the same time as the multiple first keyframes are selected as the second keyframes to generate the second keyframe sequence.

[0079] like Figure 2 As shown, Figure 2 A schematic diagram of a keyframe selection process is shown. In the diagram, a1-a10 are real-time frames of the first camera, b1-b10 are real-time frames of the second camera synchronized with a1-a10, a2 and a6 are the first keyframes obtained after feature matching and filtering, and b2 and b6 are the real-time frames of the second camera at the same time as a2 and a6, which are selected together as the second keyframes.

[0080] In some embodiments, to improve the efficiency and accuracy of keyframe selection, real-time frames can be screened before selecting the first and second keyframes.

[0081] Specifically, a real-time frame from the first camera is selected as a first candidate real-time frame based on a specific inter-frame interval. Correspondingly, keyframe selection is performed on the real-time frame of the first camera to obtain multiple first keyframes, and a first keyframe sequence is generated, including: selecting keyframes from the first candidate real-time frame to obtain multiple first keyframes, and generating a first keyframe sequence.

[0082] Similarly, real-time frames from the second camera are selected as second candidate real-time frames based on a specific inter-frame interval. Correspondingly, keyframe selection is performed on the real-time frames of the second camera to obtain multiple second keyframes, generating a first keyframe sequence. This includes: selecting keyframes from the second candidate real-time frames to obtain multiple second keyframes, and generating a second keyframe sequence.

[0083] For example, real-time frames from the first camera or the second camera are selected as the first or second candidate real-time frames based on a specific inter-frame interval. Specifically, after a specific inter-frame interval, i.e. a specific number of real-time frames, since the last inserted keyframe, multiple consecutive real-time frames are selected as candidate real-time frames.

[0084] In some embodiments, the above feature matching and filtering can be implemented in the following way: Image feature matching is performed on a real-time frame or a first candidate real-time frame from a first camera with at least one other first keyframe to obtain the number of matching feature points. If the number of matching feature points is greater than or equal to a matching threshold, the corresponding real-time frame or first candidate frame from the first camera is used as the first keyframe. Similarly, image feature matching is performed on a real-time frame or a second candidate real-time frame from a second camera with at least one other second keyframe to obtain the number of matching feature points. If the number of matching feature points is greater than or equal to a matching threshold, the corresponding real-time frame or second candidate frame from the second camera is used as the second keyframe. For example, the specific process of obtaining matching feature points through feature matching is as follows: First, feature points of the first candidate frame and other first keyframes are obtained through feature detection algorithms such as ORB-FAST corner detection. The feature points of the first candidate frame and other first candidate frame images are then matched. Feature point matching can employ matching algorithms such as ORB-BRIEF descriptors and optical flow methods.

[0085] In the above embodiments, the selection of the second keyframe is based on the real-time frame synchronized with the first keyframe. The focus is on the selection of the first keyframe, which only judges whether the real-time frame of the first camera meets the keyframe conditions (such as the number of matching feature points reaching a threshold). This method may result in the selected second camera keyframe not meeting the keyframe conditions, and the real-time frame of the second camera that meets the keyframe conditions cannot be selected as a keyframe, resulting in insufficient quality of the second keyframe, affecting the pose optimization effect of the second camera keyframe, and causing the second pose to be inaccurate.

[0086] Therefore, the selection of the first and second keyframes can be further improved. In some embodiments, when the two cameras select keyframes to generate a keyframe sequence: it is determined whether the real-time frame of the first camera meets the keyframe condition. If so, the real-time frame of the first camera is used as the first keyframe. At the same time, it is determined whether the real-time frame of the second camera meets the keyframe condition. If so, the real-time frame of the second camera is used as the second keyframe. The first and second cameras maintain their respective keyframe sequences.

[0087] For example, in some embodiments, when the two cameras select keyframes to generate a keyframe sequence: it is determined whether the real-time frame of the first camera meets the keyframe condition. If so, the real-time frame of the first camera is used as the first keyframe, and the frame of the second camera at the same time is used as the second keyframe. Furthermore, it is determined whether the real-time frame of the second camera meets the keyframe condition. If so, the real-time frame of the second camera is used as the second keyframe, and the frame of the first camera that is time-synchronized with it is also selected as the first keyframe. Thus, a keyframe sequence of two cameras is obtained.

[0088] In one alternative implementation, if a first keyframe sequence and a second keyframe sequence are generated based on keyframe selection conditions (inter-frame interval, number of matching feature points), real-time frames can be supplemented as keyframes from the real-time frames of the first and second cameras based on conditions at the same time.

[0089] Specifically, a first real-time frame from the first camera that is at the same time as multiple second keyframes can be selected and added to the first keyframe sequence as a first keyframe; correspondingly, a real-time frame from the second camera that is at the same time as multiple first keyframes can be selected and added to the second keyframe sequence as a second keyframe.

[0090] like Figure 3 As shown, Figure 3 This diagram illustrates another keyframe selection process. In the diagram, a1-a10 represent real-time frames from the first camera, and b1-b10 represent real-time frames from the second camera synchronized with a1-a10. The colored a2 and a6 represent the first keyframes selected based on the keyframe selection criteria. Similarly, the colored b3 and b7 represent the second keyframes selected based on the same criteria. Correspondingly, b2 and b6 at the same time as a2 and a6, and a3 and a7 at the same time as b3 and b7, are supplementary real-time frames selected from the first and second real-time frame sequences as keyframes, based on the condition of being at the same time.

[0091] In the above embodiments, to avoid the problem that the current camera cannot add a valid keyframe that meets the conditions due to the selection of another camera's keyframe, the keyframes of the first camera and the second camera can retain the keyframes at the same moment, without reducing the quality of the keyframes of the two cameras.

[0092] While the previous embodiment ensured the quality of the keyframes selected by the two cameras, the large number of keyframes increased the storage and computational burden on the processing module. Therefore, the keyframe selection method was further optimized. In some embodiments, multiple candidate keyframes that meet the keyframe conditions can be obtained within a specific time range based on the real-time frames of the first and second cameras. It is then determined whether there are candidate keyframes at the same time among the candidate keyframes of the two cameras. If so, the two candidate keyframes at the same time are taken as keyframes.

[0093] First, determine whether the real-time frames from the first and second cameras meet the keyframe conditions, that is, First, the real-time frames of the first camera are filtered according to the keyframe selection criteria (inter-frame interval, number of matching feature points) to obtain multiple first candidate keyframes. Then, the real-time frames of the second camera are filtered to obtain multiple second candidate keyframes. Then, based on the condition of the same time, the multiple first candidate keyframes and multiple second candidate keyframes are further filtered to obtain time-synchronized first keyframes and second keyframes, which are used for subsequent camera / robot pose calculation.

[0094] like Figure 4 As shown, Figure 4 This diagram illustrates another keyframe selection process. In the diagram, a1-a10 represent real-time frames from the first camera, and b1-b10 represent real-time frames from the second camera synchronized with a1-a10. The colored frames a2, a3, a4, and a6 are the first candidate keyframes selected based on the keyframe selection criteria. Similarly, the colored frames b1, b3, and b7 are the second candidate keyframes selected based on the same keyframe selection criteria. In the real-time frame sequences of these two cameras, it can be seen that the boxed frame a3 and the frame b3 at the same time are the first and second keyframes, satisfying both the keyframe selection criteria and the requirement of time synchronization.

[0095] By using the above settings, the keyframe selection of the other camera can be considered when the two cameras select keyframe sequences, which can maintain the synchronization of the two cameras when selecting keyframes, while ensuring the quality of keyframes as much as possible and reducing the number of keyframes added.

[0096] In one possible implementation, when there are no synchronized keyframes, or the number of synchronized keyframes is too small for subsequent calculations, any of the keyframe selection embodiments described above can be adopted. For example, the first and second keyframes can be selected separately, and real-time frames from another camera at the same time can be added to the keyframe selection to ensure the subsequent calculation of the first and second poses of the camera. Figure 5 As shown, Figure 5 This diagram illustrates another keyframe selection process.

[0097] The following describes a method for optimizing robot multi-camera localization using the first and second keyframes after keyframe selection. This specification provides the operational steps of the method as described in the embodiments or flowcharts, but may include more or fewer operational steps based on conventional or non-creative labor.

[0098] When a robot performs multi-camera localization, multiple cameras can acquire environmental images in real time. By processing these images, the camera pose at a given moment can be obtained, and the robot pose can then be derived from these poses to complete the localization process. However, because each camera is positioned and oriented differently, the acquired images differ. Although camera and robot poses can be calculated instantly from the real-time frames acquired by each camera, these frames and the keyframes selected from each frame are independent of each other. This makes it difficult to effectively utilize information from each other to constrain localization calculations and optimize the localization effect. Therefore, it is crucial to accurately select keyframes from the real-time frames acquired by the cameras that reflect the constraints between the poses of multiple cameras. Robot poses derived from keyframes with constraints are more accurate, allowing for updates to historical robot poses and increasing the accuracy of robot localization.

[0099] Specifically, such as Figure 6 As shown, Figure 6 This is a schematic diagram of a process for selecting keyframes for multi-camera localization of a robot, provided by an embodiment of the present invention. The above method may include:

[0100] S601, Obtain at least one first keyframe from the first camera within a specific time period.

[0101] The first keyframe described above reflects the key pose of the first camera. The method for obtaining the first keyframe can be found in the above embodiments.

[0102] S602, Obtain at least one second keyframe from the second camera within a specific time period.

[0103] The second keyframe described above reflects the key pose of the second camera. The method for obtaining the second keyframe can be found in the above embodiments.

[0104] S603. Perform bundle adjustment optimization on the first keyframe and the second keyframe to obtain the first pose corresponding to the first camera and the second pose corresponding to the second camera.

[0105] Specifically, the system maintains the first and second keyframes within a specific time or range, and acquires information such as the camera pose, pixel coordinates, or 3D coordinates of image feature points corresponding to the keyframes. When conditions such as a specific time and number of keyframes are met, the keyframes within a specific interval are optimized to improve the camera poses corresponding to each keyframe of the first and second cameras. The optimized pose information of the first and second cameras is then used to update the historical and current poses of the first and second cameras.

[0106] For example, a first keyframe and a second keyframe are received. Feature points from the newly added first and second keyframes are matched with feature points from existing keyframes. Based on the feature matching relationship between keyframes, the 3D coordinates of the feature points are calculated using geometric (e.g., triangulation) or optimization methods. Bundle adjustment is then applied to the historical camera pose and the already obtained 3D coordinates of the feature points. Specifically, the observed pixel coordinates of the feature points across multiple keyframes are extracted. Based on the 3D projection model of the feature points, estimated pixel coordinates are established across multiple keyframes. The difference between the observed pixel coordinates and the estimated pixel coordinates of each feature point across multiple frames is calculated, and the errors are accumulated to obtain the reprojection error. Here, the observed pixel coordinates are the known pixel coordinates of the feature points, and the estimated pixel coordinates consist of the keyframe camera pose and the parameters to be optimized for the 3D coordinates of the feature points. After obtaining the reprojection error function, the known historical camera pose and 3D coordinates of the feature points are used as initial values. An iterative solution is performed using a nonlinear least squares optimization method, such as the Gauss-Newton method, to obtain the optimized keyframe camera pose and 3D coordinates of the feature points.

[0107] S604. Update the robot's historical robot pose within a specific time period based on the first pose and the second pose.

[0108] The optimized robot pose can be obtained by transforming the first pose of the first camera through a pre-determined transformation relationship between the first camera and the robot's center, or by transforming the second pose of the second camera through a pre-determined transformation relationship between the second camera and the robot's center. The robot pose can also be any pose other than the robot's center, including but not limited to the robot's center pose, the four corner poses, the drive wheel center pose, and the left and right drive wheel poses. Other pose information can be obtained using the same method. Furthermore, the optimized robot pose replaces the robot's historical pose within a specific time period.

[0109] In one possible implementation, further optimization of the first and second keyframes using bundle adjustment includes the following steps:

[0110] First, the pixel coordinate information of the first keyframe and the second keyframe is obtained. The pixel coordinate information includes the pixel coordinates of the feature points of the first keyframe and the second keyframe. Then, the difference between the pixel coordinate information and the pixel coordinate estimate is calculated to establish a reprojection error function. The pixel coordinate estimate includes a first predictor variable and a second predictor variable. The first predictor variable includes the pose variable of the first keyframe and the 3D coordinates of the feature points of the first keyframe. The second predictor variable includes the pose variable of the second keyframe and the 3D coordinates of the feature points of the second keyframe. The first and second predictor variables consist of two parameters to be optimized: the camera key pose reflected by the keyframe and the 3D coordinates of the map points or feature points.

[0111] In one possible implementation, the reprojection error function is exemplary as follows:

[0112]

[0113] In the formula, g(X,R,T) represents the reprojection error function with camera pose (first pose or second pose) and 3D coordinates of feature points as variables, X represents the 3D coordinates of feature points, R and T are camera pose parameters, m 3D points, n frames (n images), the i-th feature point, the j-th frame, and w ij As an indicator variable, P(x) indicates whether the i-th feature point appears in the j-th frame. i ,R j ,t j () represents the pixel coordinate estimate composed of the parameters to be optimized. The pixel coordinates of the known feature points are given.

[0114] Finally, based on the reprojection error function, the historical pose and the three-dimensional coordinates of historical feature points are used as initial values. The nonlinear least squares method is used to iteratively solve the problem to obtain the first pose corresponding to the first camera and the second pose corresponding to the second camera.

[0115] like Figure 7 As shown, Figure 7 A schematic diagram of a robot localization optimization process is shown, for reference. Figure 7 a1-a11 represent image frames obtained by the first camera in different poses, b1-b11 represent image frames obtained by the second camera in different poses, and c1-c13 represent landmarks (map points) on the robot's travel map. In some embodiments, since the first camera and the second camera have different field of view, keyframes within a specific range of the two cameras may not be able to establish a sufficient or effective co-view relationship.

[0116] Since BA optimization mainly establishes constraints based on the projections and co-view relationships of multiple map points across multiple image frames, iteratively solves for minimizing the poses of each frame corresponding to the reprojection error. As shown in the figure above, image frame a1 observes multiple map points c1, c2, and c3; map point c2 is observed by image frames a1 and a2; and image frame a2 observes c2, c3, and c4. When the coordinates of map point c2 are adjusted to reduce the projection error from c2 to a1, the projection error from c2 to a2 also changes; further adjustments to the pose of a2 to reduce the projection from c2 to a2 also change the projection errors from c3 and c4 to a2. Therefore, the impact of adjusting the first and second predictor variables (two parameters to be optimized) on the error propagates through the connected graph of the projections and co-view relationships between map points and image frames. BundleAdjustment, on the other hand, needs to find a set of values ​​for all parameters to be optimized that minimizes the sum of all projection errors.

[0117] In the above bundle adjustment optimization (BA optimization), only the projection relationship between all map points and all image frames is considered, and no fixed pose constraint between the two cameras is added.

[0118] Alternatively, in another embodiment of this solution, before performing BA optimization on the keyframes of the first and second cameras, the pose variable to be optimized for one camera is converted into the form of a transformation matrix and the pose variable of the other camera using the transformation relationship between the first and second cameras, including:

[0119] A transformation matrix is ​​generated based on the position transformations of the first and second cameras; the pose variables of the second keyframe are converted into expressions of the transformation matrix and the pose variables of the first keyframe. The second keyframe and the first keyframe are keyframes at the same time, and there is a fixed pose transformation relationship T12 between the first and second keyframes at the same time.

[0120] Specifically, the original pose T2 of the second camera is represented by a predetermined transformation matrix T12 and the pose representation T1 of the first camera, such that T2 = T12 * T1. Then, the keyframes of the first and second cameras are combined for BA optimization to obtain the optimized pose information of the first or second camera.

[0121] In the above embodiments, on the one hand, the fixed pose transformation relationship constraint T2 = T12 * T1 of the first camera and the second camera is used during BA optimization, which avoids the problem that the pose estimation of the two camera keyframes is independent and cannot make full use of the information of the two cameras when the constraints between the cameras are not considered, thus improving the accuracy of pose estimation.

[0122] On the other hand, converting the parameters to be optimized from two cameras into the parameters to be optimized from a single camera significantly reduces the dimensionality of the parameters to be optimized and lowers computational complexity. Specifically, if the keyframes and map points of the first and second cameras are directly optimized, the parameters to be optimized would include the 6-dimensional pose parameters of all image frames from the first and second cameras. However, through the above conversion, the 6-dimensional pose parameters of all image frames from a single camera can be eliminated, reducing the dimensionality of the error function and thus lowering computational complexity.

[0123] To further improve the accuracy of robot localization, in some embodiments, before performing BA optimization on the keyframes of the first and second cameras, the robot's real-time pose is first optimized based on the real-time frames of the first and second cameras within a specific time period to improve the accuracy of the robot's real-time pose.

[0124] Specifically, feature points are extracted from the first real-time frame of the first camera and the second real-time frame of the second camera. Then, feature points of the first and second real-time frames are matched with the preceding real-time frames to obtain feature matching relationships. Based on these feature matching relationships, a pose algorithm is used to calculate the first real-time pose corresponding to the first real-time frame and the second real-time pose corresponding to the second real-time frame. Feature point matching can employ matching algorithms such as ORB-BRIEF descriptors and optical flow methods, while pose algorithms include epipolar geometry methods.

[0125] Furthermore, the first and second real-time poses corresponding to the first and second real-time frames at the same time can be fused and optimized to obtain the pose information of the first camera, the second camera, or the robot after fusion and optimization.

[0126] Specifically, this can be achieved through the following steps: First, obtain the first real-time pose and the second real-time pose corresponding to the first and second real-time frames synchronized with time, respectively, as the poses to be fused; then, obtain the position transformation relationship between the first camera, the second camera, and the robot; based on the position transformation relationship, optimize the poses to be fused using a fusion algorithm to obtain the optimized pose of the robot, the fusion algorithm including the extended Kalman filter algorithm; and then update the robot's real-time pose based on the optimized pose.

[0127] For example, when performing pose fusion, data fusion algorithms such as Extended Kalman Filter (EKF) can be used. Specifically, feature points of the first camera in two consecutive frames are identified and matched. Based on geometric constraints or optimization methods (such as epipolar geometry, PNP, ICP algorithms), the pose transformation between the two frames of the first camera is calculated to obtain the pose of the first camera. The pose of the second camera is inferred based on the pose of the first camera and the pose transformation between the first and second cameras, serving as the observation equation for the second camera. Similarly, the pose transformation between the two frames of the second camera is calculated. Based on the pose of the previous frame and the pose transformation between the two frames, the pose of the second camera is calculated as the state equation for the pose of the second camera. Then, the inferred pose of the second camera is fused with the calculated pose of the second camera using the EKF method to obtain the fused pose of the second camera. Based on the pose transformation relationship between the second camera, the first camera, and the center of the robot, the poses of the first camera, the center of the robot, the center of the drive wheels, etc., can be calculated. The pose is selected as the optimized pose according to actual needs, thereby updating the robot's real-time pose.

[0128] In the above-described BA (Balanced Analytical Base) positioning optimization embodiment, directly performing BA optimization on the first and second keyframes may not effectively fuse and utilize the observation information from the two cameras. This is because BA optimization is based on establishing constraints through the projection and co-view relationships of multiple map points and multiple image frames. However, when multiple cameras are facing different directions, keyframes from different cameras at the same time may not establish sufficient co-view relationships, thus failing to effectively utilize the observation information from different cameras. (Reference) Figure 7 When performing localization optimization, the keyframes a1-a5 and b1-b5 may be optimized. However, at this time, the map points observed by a1-a5 are c1-c7, and the map points observed by b1-b5 are c10-c17. The two cameras do not form a co-view relationship.

[0129] However, as the robot moves, the first camera a moves to position a10, where a10 can observe three map points c10, c11, and c12. The second camera b, at position b1, also observed c10, c11, and c12, establishing a relatively strong co-view relationship between a10 and b1. Understandably, during BA optimization, adjusting the coordinate parameters of map points c10, c11, and c12 to reduce the projection error from c10, c11, and c12 to a10 will also change the projection error from c10, c11, and c12 to b1. That is, the error caused by adjusting the parameters of one camera will propagate to the other camera through the co-view frame. From an error perspective, in image frame b1 without the second camera b, the BA optimization still has errors in estimating the poses of a1-a10 and the coordinates of c1-c12. However, without more information, this error cannot be further corrected. After introducing image frame b1 from the second camera b, when the estimated coordinates of map points c10-c12 are projected onto image frame b1, there will be a large error compared to the observed values ​​of b1, causing the BA optimizer to search for a better estimate. Therefore, introducing the observations from the second camera b can further detect the estimation error of the first camera a, and then, by minimizing the accumulated error, the optimizer can find a better estimate, resulting in a more accurate first pose and second pose.

[0130] Specifically, Business Architecture (BA) optimization can be performed using common-vision constraints in the following ways:

[0131] Determine whether there is a shared keyframe between the first keyframe and the second keyframe. If there is a shared keyframe, acquire the shared keyframe, as well as multiple first keyframes and second keyframes preceding the shared keyframe. The shared keyframe is a keyframe acquired by the first camera or the second camera, and the image of the shared keyframe has a similar image to the keyframe acquired by the other camera; that is, the shared keyframe is a first keyframe and a second keyframe with image similarity.

[0132] Bundle adjustment optimization is performed on the common-view keyframe, as well as multiple first and second keyframes preceding the common-view keyframe, to obtain the first pose corresponding to the first camera and the second pose corresponding to the second camera.

[0133] During the robot's movement, two cameras may observe the same environment at different times. In this case, image frames from different cameras at different times establish a co-view relationship, such as... Figure 8 As shown, Figure 8A schematic diagram of robot movement is shown, including map point 810, robot 820, first camera 830, and second camera 840. Hollow arrows indicate the orientation of the first and second cameras. The diagram shows the movement changes of robot 820 at four times from t1 to t4. It can be seen that from t1 to t2, the first camera 830 cannot observe map point 810, while the second camera 840 can observe map point 810. Therefore, there is no co-view relationship between the first and second keyframes acquired during t1-t2. However, from t3 (or the aforementioned t1-t2) to t4, the first camera 830 cannot observe map point 810 at t3, but at t4, robot 820 turns and moves forward. The first camera 830 can observe map point 810 at t4, and the second camera 840 can observe map point 810 continuously from t3 to t4. Therefore, there is a co-view relationship between the first and second keyframes acquired during t3-t4, and the co-view constraint can be used for BA optimization.

[0134] In some embodiments, shared keyframes can be filtered in the following manner:

[0135] In the presence of shared keyframes, acquire the shared keyframes, as well as multiple first and second keyframes preceding the shared keyframes, including:

[0136] Calculate the bag-of-words model vectors of multiple first keyframes, and use them as multiple first vectors;

[0137] Calculate the bag-of-words model vectors of multiple second keyframes, and use them as multiple second vectors;

[0138] The similarity between the first vector and the second vector is calculated to obtain the similarity result;

[0139] If the similarity results indicate that the first vector and the second vector are similar, the corresponding first keyframe and second keyframe are obtained as co-view keyframes.

[0140] Acquire multiple first and second keyframes prior to the shared keyframe.

[0141] In the above embodiments, by utilizing the same image observed by the camera at different times, a co-view relationship between different cameras can be established. The observations of different cameras can be used to optimize local pose estimation, reduce cumulative errors, and improve positioning accuracy.

[0142] Figure 9 This is a structural block diagram of a robot multi-camera localization keyframe selection device provided in an exemplary embodiment of this application. The device includes:

[0143] The first selection module 901 is used to select key frames from the real-time frames of the first camera, obtain multiple first key frames, and generate a first key frame sequence.

[0144] The second selection module 902 is used to select key frames from the real-time frames of the second camera, obtain multiple second key frames, and generate a second key frame sequence.

[0145] The pose calculation module 903 is used to perform optimized pose calculation based on the first keyframe sequence and the second keyframe sequence to obtain multiple first poses corresponding to the first camera and multiple second poses corresponding to the second camera. The multiple first poses and multiple second poses are used to optimize the robot's historical robot pose within a specific time period.

[0146] Optionally, the second selection module 902 selects a second keyframe sequence that includes second camera real-time frames at the same time as the plurality of first keyframes;

[0147] Optionally, the first selection module 901 is also used for:

[0148] Feature matching and filtering are performed on the real-time frames of the first camera to obtain multiple first keyframes. Generate the first keyframe sequence;

[0149] Optionally, the second selection module 902 is also used for:

[0150] Select a real-time frame from a second camera that occurs at the same time as multiple first keyframes, and add it to the second keyframe sequence as a second keyframe.

[0151] Optionally, the first selection module 901 is also used for:

[0152] Feature matching and filtering are performed on the real-time frames of the first camera to obtain multiple first keyframes. Generate the first keyframe sequence;

[0153] Optionally, the second selection module 902 is also used for:

[0154] Feature matching and filtering are performed on real-time frames from the second camera to obtain multiple second keyframes. Generate the second keyframe sequence.

[0155] Optionally, the above apparatus further includes a feature matching module, used for:

[0156] Perform image feature matching between a real-time frame from a first camera and other first keyframes to obtain the number of matching feature points; or perform feature matching between a real-time frame from a second camera and other second keyframes to obtain the number of matching feature points.

[0157] If the number of matched feature points is greater than or equal to the matching threshold, the corresponding first candidate frame is used as the first keyframe, or the corresponding second candidate frame is used as the second keyframe.

[0158] Optionally, the above apparatus further includes a first real-time frame module, used for:

[0159] The real-time frame of the first camera is selected as the first candidate real-time frame based on a specific inter-frame interval.

[0160] Correspondingly, the first selection module 901 is also used for:

[0161] From the first candidate real-time frame, key frames are selected from the real-time frames of the first camera to obtain multiple first key frames, and a first key frame sequence is generated.

[0162] Optionally, the above apparatus further includes a second real-time frame module, used for:

[0163] The real-time frame of the second camera is selected as the second candidate real-time frame based on a specific inter-frame interval.

[0164] Correspondingly, the second selection module 902 is also used for:

[0165] From the second candidate real-time frames, keyframes are selected from the real-time frames of the second camera to obtain multiple second keyframes, and a second keyframe sequence is generated.

[0166] Optionally, the first selection module 901 is also used for:

[0167] Select a real-time frame from the first camera that is time-synchronized with the real-time frame of the second camera, and use it as the first keyframe, adding it to the first keyframe sequence.

[0168] Correspondingly, the second selection module 902 is also used for:

[0169] Select a real-time frame from a second camera that is time-synchronized with the real-time frame of the first camera, and use it as the second keyframe, adding it to the second keyframe sequence.

[0170] Optionally, the first selection module 901 and the second selection module 902 are further used for:

[0171] Feature matching and filtering are performed on the real-time frames of the first camera to obtain the first candidate keyframes;

[0172] Feature matching and filtering are performed on the real-time frames of the second camera to obtain the second candidate keyframes;

[0173] Filter the first and second candidate keyframes that are time-synchronized. These are respectively used as the first keyframe and the second keyframe.

[0174] Optionally, the above-mentioned device further includes a selection module, which, before performing optimized pose calculation based on the first keyframe sequence and the second keyframe sequence, is used to:

[0175] A real-time frame from the first camera that is at the same time as multiple second keyframes is selected and added to the first keyframe sequence as the first keyframe.

[0176] A real-time frame from a second camera at the same time as multiple first keyframes is selected and added to the second keyframe sequence as a second keyframe.

[0177] Optionally, the pose calculation module 904 is also used for:

[0178] Based on the first keyframe and the second keyframe in the first keyframe sequence and the second keyframe sequence, the bundle adjustment method is used to optimize and obtain the first pose corresponding to the first camera and the second pose corresponding to the second camera.

[0179] In an exemplary embodiment, an electronic device for selecting keyframes for multi-camera positioning of a robot is also provided. Figure 10 This is a block diagram of an electronic device according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for selecting keyframes for multi-camera positioning in a robot.

[0180] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the robot multi-camera positioning keyframe selection method described in the above embodiments.

[0181] Optionally, the computer-readable storage medium may include ROM, RAM, solid-state drives (SSDs), or optical discs, etc. The RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0182] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the robot multi-camera localization keyframe selection method described in the above embodiments.

[0183] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0184] In this specification, the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the descriptions of the embodiments described later are relatively simple, and relevant parts can be referred to the descriptions of the foregoing embodiments.

[0185] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for selecting keyframes for multi-camera positioning in a robot, the robot comprising a first camera and at least one second camera, characterized in that, The method includes: Keyframes are selected from the real-time frames of the first camera to obtain multiple first keyframes and generate a first keyframe sequence. Keyframes are selected from the real-time frames of the second camera to obtain multiple second keyframes, and a second keyframe sequence is generated. Based on the first keyframe sequence and the second keyframe sequence, optimized pose calculation is performed to obtain multiple first poses corresponding to the first camera and multiple second poses corresponding to the second camera. The multiple first poses and the multiple second poses are used to optimize the robot pose within a specific time period.

2. The method according to claim 1, characterized in that, The second keyframe sequence contains second camera real-time frames at the same time as the plurality of first keyframes.

3. The method according to claim 1, characterized in that, The step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence includes: Feature matching and filtering are performed on the real-time frames of the first camera to obtain the plurality of first key frames and generate the first key frame sequence; The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence includes: Select a real-time frame from the second camera that occurs at the same time as the plurality of first keyframes, and use it as the second keyframe.

4. The method according to claim 1, characterized in that, The step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence includes: Feature matching and filtering are performed on the real-time frames of the first camera to obtain the plurality of first key frames and generate the first key frame sequence; The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence includes: Feature matching and filtering are performed on the real-time frames of the second camera to obtain the plurality of second key frames, and the second key frame sequence is generated.

5. The method according to claim 4, characterized in that, The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence includes: Select a real-time frame from the second camera that is at the same time as the plurality of first keyframes, and add it to the second keyframe sequence as the second keyframe.

6. The method according to claim 1, characterized in that, The step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes, and the step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes, includes: Feature matching and filtering are performed on the real-time frames of the first camera to obtain multiple first candidate keyframes; Feature matching and filtering are performed on the real-time frames of the second camera to obtain multiple second candidate keyframes; The first candidate keyframe and the second candidate keyframe at the same time are selected and used as the first keyframe and the second keyframe, respectively.

7. The method according to claims 3-6, characterized in that, Before selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence, the method further includes: The real-time frame of the first camera is selected as the first candidate real-time frame based on a specific inter-frame interval. The step of selecting keyframes from the real-time frames of the first camera to obtain multiple first keyframes and generating a first keyframe sequence includes: From the first candidate real-time frame, key frames are selected from the real-time frames of the first camera to obtain multiple first key frames.

8. The method according to claims 3-6, characterized in that, Before selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a second keyframe sequence, the method further includes: The real-time frame of the second camera is selected as the second candidate real-time frame based on a specific inter-frame interval. The step of selecting keyframes from the real-time frames of the second camera to obtain multiple second keyframes and generating a first keyframe sequence includes: From the second candidate real-time frames, key frames are selected from the real-time frames of the second camera to obtain multiple second key frames.

9. The method according to any one of claims 3-6, characterized in that, The feature matching and filtering includes the following methods: The image features of a real-time frame from the first camera are matched with at least one other first keyframe to obtain the number of matched feature points; or the real-time frame from the second camera is matched with at least one other second keyframe to obtain the number of matched feature points. If the number of matched feature points is greater than or equal to the matching threshold, the corresponding real-time frame of the first camera is used as the first key frame, or the corresponding real-time frame of the second camera is used as the second key frame.

10. The method according to claim 1, characterized in that, The optimized pose calculation based on the first keyframe sequence and the second keyframe sequence includes: Based on the first keyframe and the second keyframe in the first keyframe sequence and the second keyframe sequence, bundle adjustment optimization is performed to obtain the first pose corresponding to the first camera and the second pose corresponding to the second camera.

11. A keyframe selection device for multi-camera positioning in a robot, characterized in that, The control device executes the robot multi-camera localization keyframe selection method as described in any one of claims 1 to 10, and the device includes: The first selection module is used to select key frames from the real-time frames of the first camera to obtain multiple first key frames and generate a first key frame sequence. The second selection module is used to select key frames from the real-time frames of the second camera to obtain multiple second key frames and generate a second key frame sequence. The pose calculation module is used to perform optimized pose calculation based on the first keyframe sequence and the second keyframe sequence to obtain multiple first poses corresponding to the first camera and multiple second poses corresponding to the second camera. The multiple first poses and the multiple second poses are used to optimize the robot pose within a specific time period.

12. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement the robot multi-camera localization keyframe selection method as described in any one of claims 1 to 10.

13. A keyframe selection device for multi-camera positioning in robots, characterized in that, include: The readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement the robot multi-camera localization keyframe selection method as described in any one of claims 1 to 10.