Fast selection and tracking method and system of dynamic visual servo target in complex background
By combining image texture analysis and deep learning technology, the system achieves rapid target selection and tracking in dynamic visual servoing systems under complex backgrounds. This solves the problems of target loss and positioning errors in traditional visual servoing systems under complex backgrounds, and improves the system's response speed and anti-interference ability.
Patent Information
- Application Number
- CN202411862043.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2044-12-17
AI Technical Summary
Traditional visual servoing systems struggle to quickly identify and accurately locate dynamic targets in complex backgrounds, especially when there is significant background interference or the target moves at high speeds. This can easily lead to target loss and positioning errors, as well as insufficient system response speed and anti-interference capabilities.
By combining image texture analysis and deep learning technology, and by acquiring the pose of the master device in real time and detecting interactive tools, the image sampling is expanded outward step by step to perform layer texture analysis and similarity matching, thereby realizing the rapid selection and tracking of dynamic visual servo targets.
It effectively reduces the risk of target loss, simplifies the operation process, and improves the accuracy and reliability of target selection, making it particularly suitable for the identification and tracking of fast-moving objects.
Smart Images

Figure CN119704186B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of machine vision technology, specifically to a method and system for rapidly selecting dynamic visual servo targets in complex backgrounds. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] In the field of robotics and automation control, visual servoing technology is one of the important means to achieve high-precision target positioning and motion control. Visual servoing systems guide robots to perform precise operational tasks by acquiring visual images of the target object, and are widely used in various fields such as industrial manufacturing, autonomous driving, and medical surgery. However, traditional visual servoing systems face some technical challenges and limitations in practical applications.
[0004] First, due to the complexity of the background and the dynamic movement of the target object, visual servoing systems face significant challenges in rapid target recognition and high-precision positioning. Traditional target selection methods mainly rely on static or semi-dynamic scenes, suitable for applications with relatively simple backgrounds. When there is significant background interference or the target object moves at high speed, these methods often fail to adapt, easily leading to target recognition failure or increased positioning errors. During the target selection stage, the phenomenon of the target leaving the frame (target loss) is quite common, severely impacting the robustness and practicality of the system.
[0005] Secondly, due to the non-uniform motion of the target object in space, visual servoing systems need to possess dynamic tracking and adaptive processing capabilities. However, traditional methods often lack effective modeling of the target's dynamic motion characteristics, resulting in insufficient system response speed and anti-interference capabilities. This not only limits the application of visual servoing technology in complex dynamic scenes but also places higher demands on improving the system's real-time performance. Summary of the Invention
[0006] To address the aforementioned issues, this disclosure proposes a method and system for the rapid selection and tracking of dynamic visual servoing targets in complex backgrounds. By combining image texture analysis with deep learning techniques, it achieves efficient target selection and real-time tracking with limited time and computing resources, providing an innovative solution for the application of dynamic visual servoing systems.
[0007] To achieve the above objectives, the present disclosure adopts the following technical solution:
[0008] One or more embodiments provide a method for rapid selection and tracking of dynamically visual servoing targets in complex backgrounds, applied to a dynamic visual servoing system, including a master end and a slave end, wherein the master end controls the slave end to perform target tracking, including the following steps:
[0009] The pose of the master terminal is acquired in real time, and a follow command is sent to the slave terminal so that the slave terminal follows the master terminal's movements and maintains the same pose.
[0010] The interactive tools on the master end are detected and mapped to the coordinate system of the screen captured by the slave robot;
[0011] The sampling time interval is dynamically adjusted according to the movement speed of the interactive tool to sample the images collected by the robot at the end.
[0012] For the sampled image, continuous screenshots are taken outward from the center of the interactive tool coordinates, gradually expanding outward to obtain multi-scale regional images. Layer texture analysis and similarity matching are then performed to obtain the final set of matching results for the selected target.
[0013] Target tracking is performed on the set of matching results for the selected target.
[0014] One or more embodiments provide a rapid selection and tracking system for dynamic visual servoing targets in complex backgrounds, including master and slave robots with communication connections;
[0015] The master end includes a master end microprocessor, which is configured to perform the steps of the above-described method for the rapid selection and tracking of dynamic visual servoing targets in complex backgrounds.
[0016] Compared with the prior art, the beneficial effects of this disclosure are as follows:
[0017] This disclosure ensures the target remains in the frame by continuously taking screenshots and allowing the operator to control the robot's (or camera's) movement during target selection, effectively reducing the risk of target loss. Simultaneously, the proposed step-by-step outward expansion automatic selection method simplifies the selection process and mitigates human error to some extent. Furthermore, the automatic adjustment of the target selection tolerance based on the movement speed of the interactive tool (such as a finger) adapts to dynamic changes in the target, further reducing the operator's workload and achieving more accurate and reliable target selection, particularly suitable for the identification and tracking of fast-moving objects.
[0018] The advantages of this disclosure, as well as its additional advantages, will be described in detail in the following specific embodiments. Attached Figure Description
[0019] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute a limitation thereof.
[0020] Figure 1 Here is a block diagram of the control and interaction system structure for the main terminal in Embodiment 1 of this disclosure;
[0021] Figure 2 This is a block diagram of the slave robot system structure of Embodiment 1 of this disclosure;
[0022] Figure 3 This is a schematic diagram of the rapid target selection method in the dynamic visual servoing target rapid selection and tracking method under complex backgrounds according to Embodiment 2 of this disclosure;
[0023] Figure 4 This is an overall flowchart of the method for rapid selection and tracking of dynamic visual servoing targets in complex backgrounds according to Embodiment 2 of this disclosure;
[0024] Figure 5 This is a layer acquisition region division map for Embodiment 2 of this disclosure;
[0025] Figure 6 This is a flowchart of layer texture analysis in Embodiment 2 of this disclosure;
[0026] Figure 7 This is a flowchart of image feature similarity matching for Embodiment 2 of this disclosure. Detailed Implementation
[0027] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0028] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0029] It should be noted that the terminology used herein is for descriptive purposes only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features within those embodiments can be combined with each other. The embodiments will now be described in detail with reference to the accompanying drawings.
[0030] Example 1
[0031] In one or more of the technical solutions disclosed in the embodiments, such as Figures 1 to 2 As shown, a rapid selection and tracking system for dynamic visual servoing targets in complex backgrounds includes: master and slave robots;
[0032] The main unit includes a main unit microprocessor, a wireless communication unit, an image display unit, an image processing unit, a front-facing camera unit, and an inertial measurement unit; the main unit microprocessor is connected to the image processing unit, the image display unit, the wireless communication unit, the front-facing camera unit, and the inertial measurement unit, respectively.
[0033] One feasible technical solution is that the master end can be set on a wearable device, which can be easily worn on the operator's body, such as on the head, and is used to send control commands and receive information from the slave robot system.
[0034] The master wireless communication unit is used to realize data and image transmission between the master and slave ends;
[0035] The image display unit is used to display real-time images collected from the end-robot system;
[0036] The image processing unit is used for texture and similarity analysis of images captured by the slave robot, etc.
[0037] The front-facing camera unit is used to capture images, identify and detect human hands, and determine the relative positions of the fingers;
[0038] An inertial measurement unit is used to analyze the operator's motion state and calculate the operator's pose.
[0039] The master microprocessor unit is used to generate instructions and information to control each unit, and to send the processed results from the master end to the slave robot system via wireless communication.
[0040] The slave robot includes an image acquisition unit, a robot motion control unit, and a slave wireless communication unit, wherein the robot motion control unit is connected to the image acquisition unit and the slave wireless communication unit respectively.
[0041] The image acquisition unit is used to collect information about the robot's surrounding environment, and the robot motion control unit is used to control the robot's movement.
[0042] Example 2
[0043] Based on Embodiment 1, this embodiment provides a method for fast selection and tracking of dynamic visual servoing targets in complex backgrounds, which can be configured to be implemented in a host microprocessor, such as... Figures 3 to 7 As shown, it includes the following steps:
[0044] Step 1: Acquire the pose of the master end in real time and send a follow command to the slave end so that the slave end can follow the master end's movements and maintain the same pose.
[0045] Step 2: Detect and map the interactive tools on the master end to the coordinate system of the screen captured by the slave robot;
[0046] Step 3: Dynamically adjust the sampling time interval according to the movement speed of the interactive tool to sample the images collected by the robot at the end;
[0047] For the sampled image, continuous screenshots are taken outwards from the center, gradually increasing in size, using the coordinates of the interactive tool as the center, to obtain multi-scale region images. Layer texture analysis and similarity matching are then performed to obtain the final matching result set O for the selected target. f2 ;
[0048] Step 4: Matching result set O for the selected target f2 Perform target tracking;
[0049] In this embodiment, by continuously taking screenshots and allowing the operator to control the robot (or camera) movement during target selection, the target is ensured to remain in the frame, effectively reducing the risk of target loss. Simultaneously, the proposed step-by-step outward expansion automatic selection method simplifies the selection process and, to some extent, avoids human error. Furthermore, the target selection tolerance is automatically adjusted based on the movement speed of the interactive tool (such as a finger) to adapt to dynamic changes in the target, further reducing the operator's burden and achieving more accurate and reliable target selection, particularly suitable for the identification and tracking of fast-moving objects.
[0050] Step 1: Acquire the pose of the master device in real time and send a follow command to the slave device, so that the slave device follows the master device's movements and maintains the same pose; specifically:
[0051] Step 11: Detect the operator's head posture T using the inertial measurement unit on the wearable device at the main end;
[0052] The inertial measurement unit on the wearable device detects the operator's head posture and sends the head posture T as the target posture to the slave robot system;
[0053] Step 12: Control the end-effector posture of the slave robot to match the operator's head posture, thereby achieving synchronization between the slave robot and the operator;
[0054] The master end sends the target pose T to the slave robot system, controlling it to follow the human head movement, so that the robot's end pose is consistent with the operator's head pose T.
[0055] In this step, by controlling the master robot to follow the slave robot, the operator can control the robot (camera) to move during target selection, thereby ensuring that the target is always in the frame and effectively reducing the risk of target loss.
[0056] Step 2: Detect and map the interactive tools on the master end to the screen captured by the slave robot to obtain the coordinates of the interactive tools;
[0057] Among them, the interactive tool is used to point to the target object, which can be a person's finger or a handheld object. The master end is generally the device worn by the operator, and the slave end is an intelligent agent such as a robot that performs certain operations based on the control of the master end.
[0058] Specifically, using the index finger as the interaction tool, finger coordinate detection and mapping includes the following steps:
[0059] Step 21: Use the front-facing camera unit of the wearable device on the main end to identify the index finger of the hand and determine its coordinates p1(x1,y1) in the first frame of the front-facing camera unit.
[0060] Step 22, Image Acquisition and Display: Acquire image two captured by the image acquisition unit installed at the end of the slave robot;
[0061] Specifically, a real-time frame is captured using an image acquisition unit installed at the end of the slave robot, and the image is transmitted back to the image display unit worn by the operator via a wireless communication unit.
[0062] Step 23, Coordinate Mapping: Map the coordinates p1(x1,y1) in Image 1 to Image 2 in a proportional manner to obtain the coordinates p2(x2,y2) of the index finger of the hand;
[0063] Specifically, the coordinate position p1(x1,y1) is mapped proportionally to the real-time display screen of the main image display unit, and its coordinates p2(x2,y2) in the image display system are obtained. The calculation formula for the mapping is as follows:
[0064]
[0065] in, and This indicates the range of finger movement within the front-facing camera unit's view; and This indicates the display area of the image display unit; it defines that the origin of the main image display unit is located at the upper left corner.
[0066] In this embodiment, the finger is used as an auxiliary tool. The finger plays the role of an interaction tool between the operator and the robot system. By recognizing the coordinates of the finger, the master system can convert the operator's intention (the target area being pointed to) into coordinates that the system can process. The coordinates of the finger are used to guide the visual servo system to lock onto the target object.
[0067] Step 3: Dynamically adjust the sampling time interval to sample the images collected by the robot; for the sampled images, continuously capture images outward from the center of the interactive tool coordinates, gradually expanding to obtain multi-scale regional images, and then identify the initially selected targets.
[0068] Step 31: The master end dynamically adjusts the sampling time interval to sample the images collected by the slave robot;
[0069] Step 32: Using the coordinates of the interactive tool as the center, continuously capture the sampled image outward from small to large scale, selecting areas of multiple scales to form a set of layers V of the selected areas;
[0070] Step 33: Select target based on texture analysis: Perform layer texture analysis on each layer in the set and calculate the wrap rate; if the wrap rate requirement is met, determine the smallest layer of the target object as the selected target; if the requirement is not met, return to step 1 to select again.
[0071] Specifically, according to the set time interval ΔT seconds, a screenshot of the real-time image displayed by the image display unit is taken, and the currently acquired screenshot is defined as img, where ΔT is the sampling time interval; multiple scale regions are selected in the screenshot img with coordinates p2(x2,y2) as the center, which together constitute the layer set V of the region to be selected; the system performs texture analysis on the layers of different sizes in the set V and calculates the wrapping rate. If the analysis result meets the requirements and the minimum layer w that wraps the target object is determined, then the next step (step 34) is continued; if the analysis result does not meet the requirements, the system returns to step 1 and re-selects the target.
[0072] In the above steps, the selection process is simplified by automatically expanding outwards step by step, which avoids human error to a certain extent. Furthermore, the target selection tolerance is automatically adjusted according to the finger's movement speed, and the sampling interval is dynamically adjusted to adapt to dynamic changes in the target, further reducing the operator's burden and achieving more accurate and reliable target selection. The operator only needs to keep their finger in a suitable position on the screen, and the system will begin texture analysis and wrap-around calculation from the fingertip position outwards step by step, automatically selecting the area containing the target object.
[0073] A further technical solution, in step 31, dynamically adjusting the sampling time interval based on the movement speed of the interactive tool, can dynamically determine the sampling time interval ΔT based on the frame reduction parameter d of the previous moment. The specific process is as follows:
[0074] Step 311: Calculate the change in position Δp of the interactive tool (finger tip) in the display screen uploaded by the slave robot before and after the previous time.
[0075] The formula for calculating the change in position Δp is:
[0076] Δp=||p1-p′1||;
[0077] Where p′1 represents the finger coordinate position at the previous moment, and ||·|| represents the Euclidean norm calculation symbol;
[0078] Step 312: Based on the position change Δp and the previous sampling time interval ΔT last The average velocity v of the interactive tool (finger tip) moving on the screen is calculated using the following formula:
[0079]
[0080] Step 313: Calculate the frame rate reduction parameter d based on the movement speed v of the interactive tool (finger tip). The specific calculation formula is as follows:
[0081]
[0082] in, This indicates rounding down to the nearest integer, ensuring that d is an integer; parameters a and b are chosen appropriately to control the growth rate of d; the frame rate reduction parameter f is a non-negative integer, ranging from 0 to f-1; f represents the set system standard frame rate.
[0083] In this embodiment, the frame reduction parameter f is calculated using an exponential growth relationship, so that the growth rate of d increases as the average speed v of the fingertip moving in the screen increases;
[0084] Step 314: Calculate the current sampling time interval ΔT based on the frame reduction parameter d and the set system standard frame rate. The specific calculation formula is as follows:
[0085]
[0086] In step 32, the method for constructing the set of layers V for the selected area is as follows:
[0087] Step 321: Set the shape of the selected area to a rectangle, and determine the aspect ratio of the rectangle (div).
[0088] The resolution of the main image display unit is m×n. The method for determining the aspect ratio of the rectangle div is as follows:
[0089]
[0090] Where max(·) and min(·) represent the operations of finding the maximum and minimum values, respectively; for example Figure 5 As shown, D1, D2, D3, and D4 represent the quadrant regions divided within the image display unit; p2∈Dn This indicates that coordinate p2 is located in region D. n Inside;
[0091] Region D1 can be represented by the following constraints:
[0092]
[0093] Region D2 can be represented by the following constraints:
[0094]
[0095] Region D3 can be represented by the following constraints:
[0096]
[0097] Region D4 can be represented by the following constraints:
[0098]
[0099] Step 322: Set the number of rectangular regions q for the outward expansion screenshot, and calculate the side length increment Δl(Δx,Δy) of the rectangular regions. The calculation formula is as follows:
[0100]
[0101] Step 323: Using the coordinates p2(x2,y2) of the interactive tool (operator's finger position) as the center, determine a rectangle that expands outwards in a progressively larger manner according to the aspect ratio of div, where the size of each rectangular area differs by Δl, until it reaches the edge of the screen;
[0102] Step 324: Take screenshots of the rectangles from smallest to largest, forming a total of q layers. These q layers together constitute the layer set V of the area to be selected.
[0103] Further, in step 33, the target is selected based on layer texture analysis, and layer texture analysis is performed, including the following steps:
[0104] (1) Traverse the set of layers V of the region to be selected in ascending order of image area, and select the current target analysis layer;
[0105] (2) Use the Gray Co-occurrence Matrix (GLCM) algorithm to analyze the texture features (such as contrast, homogeneity, energy, etc.) of the current target analysis layer to separate the target object from the background of the image;
[0106] (3) Calculate the target wrap-around rate (cov) by using the area ratio of the target to the background in the current target analysis layer image; the calculation formula is as follows:
[0107]
[0108] Among them, S t S represents the area of the image that is similar to the texture features of the target. b The area represented by the background;
[0109] Determine if the target's encapsulation rate is less than a set threshold. If the condition is met, the target is considered successfully separated, the texture analysis ends, and the selected layer w obtained from the texture analysis is added to the initial target selection set O. f Otherwise, continue with the texture analysis of the next layer; if target separation fails multiple times consecutively, and the number of attempts exceeds the texture analysis frame skipping parameter, then the texture analysis process ends; the specific implementation process is as follows:
[0110] (4) If the wrapping rate is greater than the threshold ε1, then proceed to step (5); if the wrapping rate is less than the threshold ε1, then it is determined that the texture analysis results based on the current set V meet the requirements, the target separation is completed, and the layer texture analysis process ends.
[0111] (5) Determine whether the traversal in step (1) has ended. If the traversal has not ended, then execute step (6); if all layers have been traversed, then determine that the target separation has not been completed based on the current set V, and execute step (7).
[0112] (6) Perform the traversal process of step (1) and update the analysis target layer, and then perform step (2);
[0113] (7) Determine whether the number of consecutive occurrences of this situation is greater than the texture analysis frame skipping parameter n1. If it is less than n1, proceed to step (8); if it is greater than n1, proceed to step (9).
[0114] Among them, the texture analysis frame skipping parameter n1 is related to the average movement speed v of the interactive tool (finger) in the picture, and the logarithmic function normalization method is used. The specific calculation formula is as follows:
[0115]
[0116] in, This represents the maximum value set for n1. The range of the texture analysis frame skipping parameter n1 is... v max This represents the maximum value of the average finger movement speed v in the image, where v ranges from [0, v]. max c1 is a positive constant, designed to prevent the logarithmic function from becoming undefined when v = 0;
[0117] (8) If the texture analysis results based on the current set V fail to meet the requirements and fail to separate the target object from the background, the layer texture analysis process ends; and the initial target selection set O is cleared. fThe initial target set O is selected. f Used to store the selected target image;
[0118] (9) Based on step (8), set O also needs to be cleared. f ;
[0119] Furthermore, in step 3, the selected target is matched to obtain the final selected target. The similarity matching method is used, including the following steps:
[0120] Step 34: Add the selected layer w obtained based on texture analysis to the initial target selection set O. f ;
[0121] Step 35: Initially select set O for the target. f Image feature similarity matching is performed on the layers in the image, and images with a similarity greater than the similarity matching threshold ε2 are extracted as the final target image matching result set O. f2 ;
[0122] Initially select set O for the target f Image feature similarity matching is performed on the layers in the image. If the similarity meets the requirements, the final matching result set O is obtained. f2 If the matching is not yet complete, proceed to step 4; if the analysis results do not meet the requirements, end the matching process.
[0123] In step 35, a preliminary set O is selected for the target. f In the layers, when the target is initially selected from set O f If the number of layers meets the requirements, the steps for image feature similarity matching are as follows:
[0124] Step 351: Initially select set O for the target. f In the images, similarity matching is performed pairwise, and the similarity is calculated for each pair.
[0125] Step 352: If the number of similarity scores less than the similarity matching threshold ε2 is less than the similarity analysis frame skipping parameter n2, extract the initial target selection set O. f Layers with a similarity greater than the threshold ε2 constitute a set O. f2 The final set of matching results O for selecting the target f2 ;
[0126] Specifically, the example steps for implementing step 35 are as follows:
[0127] (35.1) Set the number of layers n3 required for similarity analysis matching, and the judgment set O fIf the number of existing layers is greater than n3, then continue to step (352); otherwise, it is determined that the number of collected layers has not yet met the requirements, that is, the similarity matching process has not been completed, and the current similarity analysis process ends.
[0128] (35.2) Using the Structural Similarity (SSIM) algorithm, the set O f The similarity scores of the images in the diagram are calculated by performing pairwise similarity matching. The specific calculation formula is as follows:
[0129]
[0130] Among them, s ij Represents set O f Similarly, the similarity between the i-th layer and the j-th layer is calculated. This represents the similarity between the (n3-1)th layer and the n3rd layer; s ij The range is [0,1], s ij The larger the value, the higher the similarity between the i-th layer and the j-th layer; μ i With μ j These represent the average brightness of the i-th layer and the j-th layer, respectively, reflecting their brightness levels; and σ represents the luminance variance of the i-th layer and the j-th layer, respectively, reflecting their contrast levels; ij c3 and c4 represent the covariance between the i-th layer and the j-th layer, reflecting their structural information; c3 and c4 are positive constants, designed to avoid the denominator being zero during the calculation process.
[0131] Where i and j are both integers and satisfy the following relationship:
[0132]
[0133] (35.3) Guarantee set O f The continuity and stability of similarity between mid-layers, i.e., the judgment If the number of values less than the similarity matching threshold ε2 is less than the similarity analysis frame skipping parameter n2, then proceed to step (35.4); otherwise, proceed to step (35.6).
[0134] The calculation method for the similarity analysis frame skipping parameter n2 is consistent with the calculation method for the texture analysis frame skipping parameter n1. The specific calculation formula is as follows:
[0135]
[0136] in, This represents the maximum value set for n2. The range of the texture analysis frame skipping parameter n2 is... v max This represents the maximum value of the average finger movement speed v in the image, where v ranges from [0, v]. max c2 is a positive constant, designed to prevent the logarithmic function from becoming undefined when v = 0;
[0137] (35.4) Extract set O f Layers with a similarity greater than the threshold ε2 constitute a set O. y2 The final set of matching results for the selected target, O f2 The number of layers in the array is n4;
[0138] (35.5) Guarantee set O f2 The overall similarity of the middle layers is calculated. The average value E is used to determine whether the average value E is greater than the similarity average threshold ε3. If it is greater than ε3, then set O is determined. f Once the image feature similarity matching process of the layer is completed, meaning the layer similarity analysis meets the requirements, the current similarity analysis process ends; otherwise, continue to step (35.6).
[0139] in, This represents the number of all distinct combinations of selecting 2 layers from n4 layers;
[0140] (35.6) Determine based on the current set O f Unable to complete layer similarity analysis matching, meaning the analysis results do not meet the requirements, the current similarity analysis process ends, and step 1 is executed to re-acquire images;
[0141] Step 4: Matching result set O for the selected target f2 Target tracking includes the following steps:
[0142] Step 41: For the matching result set O of the selected target f2 Feature extraction is performed on the layers in the image to obtain depth features, which are then used as the target feature library.
[0143] Specifically, the final matching result set O f2 The layers in the network are input into a lightweight feature extraction network, which extracts and stores the depth features of each layer to form a target feature library.
[0144] Step 42: Acquire each real-time frame uploaded by the robot, extract features, calculate the similarity between the extracted features and the features in the target feature library, and locate and track targets with similarity higher than a set threshold.
[0145] In this step, we enter the target detection and tracking stage, which involves displaying each real-time frame A on the image display unit. 11 Cosine similarity can be used to calculate the real-time image A. 11 The angle between the included feature vector and the feature vector in the target feature library is used to determine their similarity and to achieve target detection, that is, to detect the position of the target in the picture in real time, thus preparing for subsequent visual servo control.
[0146] This embodiment employs a real-time, rapid, dynamic target selection method, which reduces the risk of target loss after selection. Some traditional target selection methods require taking screenshots of the displayed screen and then performing a relatively slow target selection process on the static screenshot. When the selection is complete and the detection phase begins, the target may have already left the camera's view, resulting in target loss, especially with fast-moving targets. The method in this embodiment, by continuously taking screenshots of the real-time camera view in step 3 throughout the target selection process, and allowing the operator to control the robot's (camera's) movement during this time, ensures the target remains within the camera's view, effectively solving the aforementioned target loss problem.
[0147] The outward-expanding automatic selection method of this embodiment reduces the operational difficulty of the selection process. In some traditional target selection processes, a large amount of tedious manual operation is required, such as gesture selection. This process is highly susceptible to subjective human error and prone to errors. Furthermore, even with continuous screenshotting, when facing a rapidly moving target, the speed and uncertainty of the target's movement can cause previous selection operations to fail, such as the target leaving the selection area that was about to be closed by the gesture selection. In the method proposed in this embodiment, the operator only needs to keep their finger in a suitable position on the screen, and the system will begin to perform texture analysis and wrap-around calculation from the fingertip position outwards in a progressively larger manner, automatically selecting the area containing the target object.
[0148] The method proposed in this embodiment can adjust the system's dynamic selection tolerance rate based on the speed of finger movement, thereby reducing the operational burden of selection. Since the finger's position on the screen needs to change with the movement of the target object, the finger's movement speed can indirectly reflect the target object's movement speed. For example, when the finger moves quickly on the screen, it means the target object is likely also moving relatively quickly. In this case, the probability of target selection error increases, requiring dynamic adjustment of relevant parameters to affect the system's tolerance rate during dynamic selection, thus reducing the finger's movement burden and avoiding or minimizing the adverse effects of operational errors on target selection.
[0149] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
[0150] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.
Claims
1. A method for fast selection and tracking of dynamic visual servoing target in complex background, applied to a dynamic visual servoing system, including a master end and a slave end, the master end controls the slave end to track the target, characterized in that, The method comprises the following steps: Real-time acquisition of the pose of the master end, sending a following instruction to the slave end, so that the slave end follows the master end and keeps the pose consistent; Detection and mapping of the interactive tool of the master end to the picture coordinate system collected by the slave end robot; Dynamic adjustment of the sampling time interval according to the moving speed of the interactive tool, and sampling of the image collected by the slave end robot; For the sampled image, continuously capture the image outward from small to large step by step with the interactive tool coordinates as the center, obtain a multi-scale region image, and perform layer texture analysis and similarity matching to obtain a matching result set of the selected target; Target tracking for the matching result set of the selected target; The obtained sampling image is centered on the interactive tool coordinates, and the continuous screenshots are expanded outward in stages from small to large, a plurality of scale regions are selected, and a to-be-selected region layer set is formed , the construction method of the to-be-selected region layer set is as follows: The shape of the selection region is set to a rectangle, and the length-width ratio of the rectangle is determined ; Setting the number of rectangular regions of the screenshot to be expanded outward , calculating the side length increment of the rectangular region; Determine the outward proportion of length and width with the interactive tool coordinates as the center Rectangles of gradually increasing size, where the size of each rectangular region differs from the next by a factor of 2, until the picture border is reached. For the above from small to large rectangle respectively screenshot, a total of layers, layers together constitute the selected area layer set ; Target selection based on layer texture analysis, comprising the following steps: traversing the to-be-selected region layer set in order of image area from small to large selecting a current target analysis layer; Separation of the target object and the picture background by performing texture feature analysis on the current target analysis layer by using the gray level co-occurrence matrix algorithm; Calculation of the target wrapping rate by using the area proportion of the target and the background in the current target analysis layer picture; determining whether the target package rate is greater than a set threshold value, if the condition is met, it is considered that the target separation is successful, the texture analysis is ended, and the selected layer adding the target preliminary selection set ; otherwise, continue the texture analysis of the next layer; if the target separation is not completed for a plurality of times, and the number of times exceeds the texture analysis frame skipping parameter, the texture analysis process is ended. Matching to obtain the final selected target by using the similarity matching method, comprising the following steps: selecting layers based on texture analysis adding target preliminary selection set ; Preliminary selection set of target The similarity is calculated by matching the figures in the set two by two. If the similarity is less than a similarity match threshold The number of layers is less than a similarity analysis skip frame parameter Extract a target preliminary selection set The similarity of which is greater than a threshold The layer set consisting of the set The matching result set as the final selected target .
2. The method of claim 1, wherein the method comprises: For the interactive tool being a finger, detection and mapping of the interactive tool of the master end, comprising the following steps: The front camera unit of the wearable device of the host end recognizes the finger and determines the coordinate of the finger in a picture one of the front camera unit ; Acquisition of the image captured by the image acquisition unit installed at the end of the slave end robot; The coordinates of the hand in the frame are mapped to the frame of the image two in a scale manner to get the coordinates of the index finger of the hand .
3. The method of claim 1, wherein the method comprises: Dynamic adjustment of the sampling time interval according to the moving speed of the interactive tool, comprising the following steps: The position change amount of the interaction tool in the display screen uploaded from the terminal robot before and after the calculation ; According to the position change amount and the previous sampling time interval , the average speed of the interactive tool moving in the picture is calculated ; Calculating a frame reduction parameter based on average speed of interaction tool ; According to the frame reduction parameter and the set system standard frame number, the current sampling time interval is calculated .
4. The method of claim 1, wherein: Similarity analysis skip frame parameters The calculation formula is as follows: wherein represents the set maximum value, texture analysis frame skipping parameter ranging from ; represents the maximum value of the average movement speed of the finger in the picture , ranging from ; is a positive constant.
5. The method of claim 1, wherein the dynamic visual servoing target is selected from a plurality of targets in a complex background. Target tracking for the matching result set of the selected target, comprising the following steps: Feature extraction is performed on the layers in the matching result set of the selected target to obtain depth features as a target feature library; Real-time picture uploaded by the slave end robot is acquired, feature extraction is performed, similarity between the extracted features and the features in the target feature library is calculated, and the target with a similarity higher than a set threshold is positioned and tracked.
6. A fast selection and tracking system of dynamic visual servo target in complex background, characterized in that: The master end and the slave end robot are communicatively connected. The master end comprises a master end microprocessor configured to perform the steps of the method of claim 1-5.
7. The fast selection and tracking system of dynamic visual servoing target in complex background according to claim 6, characterized in that: The master end further comprises a wireless communication unit, an image display unit, an image processing unit, a front camera unit, and an inertial measurement unit; the master end microprocessor is connected with the image processing unit, the image display unit, the wireless communication unit, the front camera unit, and the inertial measurement unit, respectively; The slave end robot comprises an image acquisition unit, a robot motion control unit, and a slave end wireless communication unit; the robot motion control unit is connected with the image acquisition unit and the slave end wireless communication unit, respectively; The image acquisition unit is used to collect the environmental information around the robot, and the robot motion control unit is used to control the motion of the robot.
Citation Information
Patent Citations
Robot remote control system and method based on multi-modal interaction technology
CN113821108A
Multi-robot collaborative visual monitoring method and system
CN115331160A