Multi-mobile intelligent body anti-shake automatic aiming method, system, device and medium
By using the YOLOv8-pose detector and multi-object tracking algorithm in the automated aiming system, combined with adaptive PID control, the problems of sight jitter and tracking delay under multiple high-speed displacement targets are solved, and the target tracking and anti-interference ability with high smoothness is achieved.
Patent Information
- Application Number
- CN202411720857.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the case of multiple high-speed displacement targets, existing automated aiming technologies are difficult to make decisions quickly and accurately, resulting in a large swing of the sight between multiple targets, resulting in the tracking delay and sight jitter problems.
The YOLOv8-pose detector is used to receive the camera airport scene image, predict, match and filter target points through a multi-object tracking algorithm, calculate the offset in the two-dimensional picture and perform three-dimensional viewing angle transformation to obtain the camera's rotation vector. Then, using adaptive position PID control, the rotation vector is fine-tuned to realize adaptive adjustment of the actual execution amount and ensuring smooth rotation of the camera's viewing angle.
It effectively offsets the tracking delay, realizes high-smooth target tracking that resists multi-target interference, reduces sight jitter, and predicts emergencies such as large changes in tracking direction and loss of tracking points, and makes decisions in advance.
Smart Images

Figure CN119200387B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection and tracking, and in particular to a multi-mobile intelligent body anti-shake automatic aiming method, system, equipment and medium. Background Art
[0002] Automatic aiming is an artificial intelligence application technology based on machine vision. It uses a high-precision camera to capture the target scene in real time, and then sends the acquired video stream to the target detection model for reasoning. It then automatically aims and tracks the target results based on the detection results output by the model, thereby realizing the full process operation from recognition to aiming.
[0003] The automated aiming scheme based on machine vision has the advantages of high efficiency, low cost and high autonomy. Some existing studies divide the automated aiming scheme into three parts: target detection, aiming position selection and automatic aiming. The YOLOv5 target detector is used to perform target recognition operations on the two-dimensional plane. However, when faced with extremely complex real-life scenarios, most of the current automated aiming technologies are only aimed at single target tracking. When multiple targets appear in the picture at the same time, it is difficult for the algorithm to make decisions quickly and accurately, which will cause the crosshairs to swing greatly between multiple targets. For this reason, we propose a multi-mobile intelligent body anti-shake automatic aiming method, system, equipment and medium. Summary of the invention
[0004] The purpose of the present invention is to provide a multi-mobile intelligent body anti-shake automatic aiming method, system, equipment and medium, which can solve the problems of tracking delay and sight shaking when multiple high-speed displacement targets are encountered.
[0005] According to a first aspect of the present invention, in order to achieve the above-mentioned purpose, the present invention provides the following technical solution: a multi-mobile intelligent body anti-shake automatic aiming method, comprising the following steps:
[0006] Receive the camera scene image frame and input it into the Yolov8-pose detector to obtain the target point position information recorded as ,in The Elements Indicates time The corresponding screen The coordinate position of the tracking target point in the two-dimensional plane of the camera;
[0007] The multi-target tracking algorithm is used to predict, match and filter the location information of multiple target points to obtain new tracking points. ;
[0008] Calculate the two-dimensional image Offset from the crosshair , according to the horizontal and vertical field of view of the camera, the offset Perform three-dimensional perspective transformation on each dimension to obtain the rotation vector of the camera ;
[0009] Define the original execution volume as , the two states of "crosshair chasing" and "crosshair overspeed" that appear in the process of tracking the target affect the rotation vector Perform adaptive position PID control to achieve the original execution amount Adaptive adjustment to get the actual execution amount ;
[0010] The actual execution Input external control device to rotate the camera's viewing angle vector The actual execution volume equal.
[0011] Furthermore, the camera scene image frame is obtained by capturing the real-time image of the camera, and the aspect ratio of the captured image is , the side length is .
[0012] Furthermore, the field of view of the camera is denoted as ,in and Represents the camera's field of view in the horizontal and vertical directions.
[0013] Furthermore, a multi-target tracking algorithm is used to predict, match and filter multiple target points to obtain new tracking points. , specifically including the following:
[0014] (41) Prediction phase: Defining the mean of the Kalman filter for ,in Representative Points The estimated coordinates, Representative Points Normalized results of estimated moving speed in horizontal and vertical dimensions;
[0015] At the moment , the Kalman filter is based on the mean of a point at the previous moment Predict the posterior mean of the point at the current moment , and obtain the posterior distribution ;take out The first two components of all elements in get the predicted point set ,in , yes The corresponding prediction point is denoted as ;
[0016] (42) Matching stage, used to match point sets and Points in:
[0017] (421) First, calculate the weighted Mahalanobis distance between different points in the two point sets ,in , , is the covariance matrix between two point sets;
[0018] (422) To ensure that the tracking point at the previous moment is matched first, define the weight coefficient vector , Indicate point The corresponding weight coefficient is calculated as follows:
[0019] (1)
[0020] (2)
[0021] (3)
[0022] In the formula, Responsible for improving forecast points The matching weight of Indicate point The offset from the center of the picture, point The coordinates are expressed as ; is the offset component coefficient, which is used to control the effect of the offset in the weight coefficient, usually ,in is the side length of the square camera image; during the matching process, even if If no suitable detection point is matched, the program can also give priority to selecting detection points closer to the center point as new tracking points to achieve small-scale tracking transition;
[0023] (423) Define the cost matrix , whose elements are the normalized weighted Mahalanobis distances , after Hungarian matching, we get an ordered pair set ,in express and Match;
[0024] (43) Screening phase: After matching is completed Tracking points is defined as ,definition Offset from the crosshair .
[0025] Furthermore, since the 3D scene is reduced to 2D after being captured by the camera, the scene depth information is discarded, so the offset needs to be Optimize the viewing field and get , the calculation formula is as follows:
[0026] (4)
[0027] (5)
[0028] In the formula, Indicates the side length of the screen, .
[0029] Furthermore, before performing PID control, the proportional coefficient To set the parameters, the calculation formula is as follows:
[0030] (6)
[0031] In the formula, Indicates the movable range of the hardware device. Indicates the field of view of the camera, confirm After that, according to the tracking situation Start fine tuning the integral coefficient and the differential coefficient , the former can make up for Too small, the latter is responsible for predicting and suppressing Oversized situations;
[0032] Will Considered as deviation, inspired by position PID, the execution amount is defined The calculation formula is as follows:
[0033] (7)
[0034] (8) (9)
[0035] In the formula, represents the rotation vector, Represents the rotation vector relative to time The derivative of is the integral value of the limited deviation, and Indicates the upper and lower limits of the integral value of the deviation; is the conditional coefficient, Indicates when the integral fine-tuning takes effect In order to prevent the long-term lag of the crosshairs during high real-time tracking, which leads to a large accumulation of integrals, the integral limit and conditional integral are used to limit the integral term.
[0036] Furthermore, the two states of "crosshair chasing" and "crosshair overspeed" that occur in the process of the crosshair tracking the target affect the rotation vector Perform adaptive position PID control, as follows:
[0037] (71) First, determine whether the current tracking point is a new tracking point. If so, immediately perform full control to accelerate aiming, i.e. ; If not, enter the state subdivision control:
[0038] (711) When the target moves at high speed and in the same direction as the crosshair, due to the delay in model reasoning, the crosshair will always lag behind the target, that is, the "crosshair chasing" state. At this time, full control is performed, that is, ;
[0039] (712) When the target suddenly slows down or changes direction drastically, the crosshair will exceed the target, which is the "crosshair overspeed" state. At this time, define the force scaling factor vector To reduce the impact of execution volume and use the freeze frame number To control the number of frames of force reduction, ,in It is related to the motion state of the tracking point and the calculation formula is as follows:
[0040] (10)
[0041] Where e represents the base of the natural logarithm and defines the force scaling factor in the horizontal and vertical directions, where and Used to control the strength of force scaling in each direction, and the offset component accounts for The larger the offset, the more significant the force scaling. represents the posterior mean at time k, , They respectively represent the normalized results of the estimated values of the moving speed of point i in the horizontal and vertical directions at time k.
[0042] According to a second aspect of the present invention, the present invention provides a multi-mobile agent anti-shake automatic aiming system, comprising:
[0043] The first input module is used to receive the camera scene image frame and input it into the Yolov8-pose detector to obtain the target point position information recorded as ,in The Elements Indicates time The corresponding screen The coordinate position of the tracking target point in the two-dimensional plane of the camera;
[0044] The target point processing module is used to predict, match and filter the location information of multiple target points using a multi-target tracking algorithm to obtain new tracking points. ;
[0045] The calculation module is used to calculate the Offset from the crosshair , according to the camera's horizontal and vertical field of view, Perform three-dimensional perspective transformation on each dimension to obtain the rotation vector of the camera ;
[0046] PID control module, used to define the original execution quantity as , the two states of "crosshair chasing" and "crosshair overspeed" that appear in the process of tracking the target affect the rotation vector Perform adaptive position PID control to achieve the original execution amount Adaptive adjustment to get the actual execution amount ;
[0047] The second input module is used to convert the actual execution amount Input external control equipment to make the camera's viewing angle rotation amount and actual execution amount equal.
[0048] According to a third aspect of the present invention, the present invention provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, the above-mentioned multi-mobile intelligent body anti-shake automatic aiming method is adopted.
[0049] According to a third aspect of the present invention, the present invention provides a storage medium containing computer executable instructions, which are used to execute the above-mentioned multi-mobile agent anti-shake automatic aiming method when executed by a computer processor.
[0050] The present invention has at least the following beneficial effects:
[0051] This paper takes the YOLOv8-pose key point detector as the basis, integrates the multi-target screening strategy based on MOT, and for the first time applies the PID (proportional integral differential control) control method in mechanical control to the post-processing stage of tracking. It offsets the tracking delay through adaptive fine-tuning, realizes high-smoothness target tracking with resistance to redundant target interference, and compensates for the tracking lag and crosshair jitter caused by inference delay to the greatest extent. At the same time, it can predict emergencies such as large changes in tracking direction and loss of tracking points, and make decisions in advance.
[0052] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the process of the method of the present invention;
[0054] Figure 2 A schematic diagram of the structure of an embodiment of the present invention
[0055] Figure 3 This is a schematic diagram of the multi-target screening strategy based on MOT of the present invention;
[0056] Figure 4 Schematic diagram of nonlinear transformation of distance vector in the viewing field of the present invention;
[0057] Figure 5 This is a subdivided PID control logic process diagram of the present invention. DETAILED DESCRIPTION
[0058] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0059] Embodiment 1:
[0060] like Figure 2 As shown in the figure, the multi-target tracking algorithm architecture proposed in the present invention mainly includes six parts: picture capture, target point detection, multi-target screening, nonlinear transformation of distance vector in the field of view, angle vector fine-tuning, and hardware control aiming. The image input to the detector is .
[0061] See also Figure 1-Figure 5 The present invention provides a technical solution: a multi-mobile intelligent body anti-shake automatic aiming method, comprising the following steps:
[0062] S1. Receive the camera scene image frame and input it into the Yolov8-pose detector to obtain the target point position information recorded as ,in The Elements Indicates time The corresponding screen The coordinate position of the tracking target point in the two-dimensional plane of the camera;
[0063] According to the technical solution of this embodiment, the camera scene image frame is obtained by capturing the real-time image of the camera, and the aspect ratio of the captured image is , the side length is ;
[0064] Furthermore, the field of view of the camera is denoted as ,in and Represents the camera's field of view in the horizontal and vertical directions;
[0065] S2. Use the multi-target tracking algorithm (MOT) to predict, match and filter the location information of multiple target points to obtain new tracking points ,like Figure 3 As shown, the details are as follows:
[0066] (S21) Prediction phase: Defining the mean of the Kalman filter for ,in Representative Points The estimated coordinates, Representative Points Normalized results of estimated moving speed in horizontal and vertical dimensions;
[0067] At the moment , the Kalman filter is based on the mean of a point at the previous moment Predict the posterior mean of the point at the current moment , and obtain the posterior distribution Since the distribution dimension of the detection point set (two-dimensional) is inconsistent with the dimension of the posterior distribution (four-dimensional), it is necessary to remove The first two components of all elements in get the predicted point set ,in , yes The corresponding prediction point is denoted as ;
[0068] (S22) Matching stage, used to match point sets and Points in:
[0069] (S221) First, calculate the weighted Mahalanobis distance between different points in the two point sets ,in , , is the covariance matrix between two point sets;
[0070] (S222) To ensure that the tracking point at the previous moment is matched first, define a weight coefficient vector , Indicate point The corresponding weight coefficient is calculated as follows:
[0071] (1)
[0072] (2)
[0073] (3)
[0074] In the formula, Responsible for improving forecast points The matching weight of Indicate point The offset from the center of the picture, point The coordinates are expressed as ; is the offset component coefficient, which is used to control the effect of the offset in the weight coefficient, usually ,in is the side length of the square camera image; during the matching process, even if If no suitable detection point is matched, the program can also give priority to selecting detection points closer to the center point as new tracking points to achieve small-scale tracking transition;
[0075] (S223) Define the cost matrix , whose elements are the normalized weighted Mahalanobis distances , after Hungarian matching, we get an ordered pair set ,in express and Match;
[0076] (S23) Screening stage: After matching is completed Tracking points is defined as ,definition Offset relative to the crosshair (center of the screen) ;
[0077] It should be noted that since the 3D scene is reduced to 2D after being captured by the camera, the scene depth information is discarded. Optimize the viewing field and get , the calculation formula is as follows:
[0078] (4) (5)
[0079] In the formula, Indicates the side length of the screen, ;
[0080] S3. Figure 4 As shown, calculate the new tracking point in the two-dimensional picture Offset relative to the crosshair (center of the screen) , according to the horizontal and vertical field of view of the camera, the offset Perform three-dimensional perspective transformation on each dimension to obtain the rotation vector of the camera ;
[0081] S4. Define the original execution volume as , the two states of "crosshair chasing" and "crosshair overspeed" that appear in the process of tracking the target affect the rotation vector Perform adaptive position PID control to achieve the original execution amount Adaptive adjustment to get the actual execution amount ;
[0082] It should be noted that before performing PID control, the proportional coefficient To set the parameters, the calculation formula is as follows:
[0083] (6)
[0084] In the formula, Indicates the movable range of the hardware device. Indicates the field of view of the camera, confirm After that, according to the tracking situation Start fine tuning the integral coefficient and the differential coefficient , which can make up for Too small, the latter is responsible for predicting and suppressing Oversized situations;
[0085] Will Considered as deviation, inspired by position PID, the execution amount is defined The calculation formula is as follows ( Contains two execution volume components: horizontal and vertical):
[0086] (7) (8) (9)
[0087] In the formula, represents the rotation vector, Represents the rotation vector relative to time The derivative of is the integral value of the limited deviation, and Indicates the upper and lower limits of the integral value of the deviation; is the conditional coefficient, Indicates when the integral fine-tuning takes effect In order to prevent the long-term lag of the crosshairs during high real-time tracking, which leads to a large accumulation of integrals, the integral limit (Formula 8) and conditional integral (Formula 9) are used to limit the integral term; the integral limit is set by setting the integral threshold. Restricted to Within the range; conditional integration is achieved by setting the limit coefficient Some of the large deviations that may be caused by unexpected circumstances Filtering further improves the stability of the system;
[0088] The rotation vector is affected by the two states of "crosshair chasing" and "crosshair overspeed" that occur when the crosshair is tracking the target. Perform adaptive position PID control, such as Figure 5 As shown, the details are as follows:
[0089] (S41) First, determine whether the current tracking point is a new tracking point. If so, immediately perform full control to accelerate aiming, i.e. ; If not, enter the state subdivision control:
[0090] (S411) When the target moves at high speed and in the same direction as the crosshair, due to the delay in model reasoning, the crosshair will always lag behind the target, that is, the "crosshair chasing" state. At this time, full control is performed, that is, ;
[0091] (S412) When the target suddenly slows down or changes its moving direction significantly, the crosshair will exceed the target, which is the "crosshair overspeed" state. At this time, the force scaling factor vector is defined To reduce the impact of execution volume and use the freeze frame number To control the number of frames of force reduction, ,in It is related to the motion state of the tracking point and the calculation formula is as follows:
[0092] (10)
[0093] Where e represents the base of the natural logarithm, defining the force scaling factor in the horizontal and vertical directions, where and Used to control the strength of force scaling in each direction, and the offset component accounts for The larger the offset, the more significant the force scaling. represents the posterior mean at time k, , They represent the normalized results of the estimated moving speed of point i in the horizontal and vertical directions at time k respectively;
[0094] S5. The actual execution amount Input external control device to rotate the camera's viewing angle vector and Equal, the camera's line of sight moves to the tracking point, and now the tracking operation within one frame is completed.
[0095] In summary, this embodiment is based on the YOLOv8-pose key point detector, integrates the MOT-based multi-target screening strategy, and for the first time applies the PID (proportional integral differential control) control method in mechanical control to the post-processing stage of tracking. It offsets the tracking delay through adaptive fine-tuning, realizes high-smoothness target tracking that is resistant to redundant target interference, and compensates for the tracking lag caused by the inference delay to the greatest extent. At the same time, it can predict sudden situations such as large changes in tracking direction and loss of tracking points and make decisions in advance.
[0096] Embodiment 2:
[0097] The present invention provides a multi-mobile intelligent body anti-shake automatic aiming system, comprising:
[0098] The first input module is used to receive the camera scene image frame and input it into the Yolov8-pose detector to obtain the target point position information recorded as The Elements Indicates time The corresponding screen The coordinate position of the tracking target point in the two-dimensional plane of the camera;
[0099] The target point processing module is used to predict, match and filter the position information of multiple target points using a multi-target tracking algorithm to obtain new tracking points;
[0100] The calculation module is used to calculate the offset relative to the crosshair in the two-dimensional image, and perform three-dimensional perspective transformation on each dimension according to the horizontal and vertical field of view of the camera, so as to obtain the rotation vector of the camera;
[0101] The PID control module is used to define the original execution amount as follows: for the two states of "crosshair chasing" and "crosshair overspeed" that appear in the process of the crosshair tracking the target, the rotation vector is adaptively controlled by position-based PID control, so as to realize the adaptive adjustment of the original execution amount to obtain the actual execution amount;
[0102] The second input module is used to input the actual execution amount into the external control device to rotate the camera's viewing angle vector The actual execution volume equal.
[0103] Specifically, the above-mentioned first input module, target point processing module, calculation module, PID control module and second input module can be embedded in a computer processing system, and the computer calls the above-mentioned modules to complete the task of tracking and aiming at multiple targets based on the above-mentioned multi-mobile intelligent body anti-shake automatic aiming method; the above-mentioned first input module, target point processing module, calculation module, PID control module and second input module can perform operations according to the specific steps given in the multi-mobile intelligent body anti-shake automatic aiming method.
[0104] It should be noted that it should be understood that the division of the various modules of the above system is only the division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated, and these modules can all be implemented in the form of software calling through processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. For example, the calculation module can be a separately established processing element, or it can be integrated in a chip of the above-mentioned device. In addition, it can also be stored in the memory of the above-mentioned device in the form of program code, and called and executed by a processing element of the above-mentioned device. The functions of the above signal processing module, and the implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the hardware integrated logic circuit in the processor element or the instructions in the form of software.
[0105] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more digital singnal processors (DSP), or one or more field programmable gate arrays (FPGA). For another example, when a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0106] Embodiment three:
[0107] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, the above-mentioned multi-mobile intelligent body anti-shake automatic aiming method is adopted.
[0108] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer or a cloud server, and the terminal device includes but is not limited to a processor and a memory. For example, the terminal device can also include input and output devices, network access devices and buses.
[0109] Furthermore, the processor may adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. may also be adopted. The general-purpose processor may adopt a microprocessor or any conventional processor, etc., and this application does not impose any restrictions on this.
[0110] Embodiment 4:
[0111] The present invention provides a storage medium containing computer executable instructions, characterized in that the computer executable instructions are used to execute the above-mentioned multi-mobile intelligent body anti-shake automatic aiming method when executed by a computer processor.
[0112] Among them, the computer program can be stored in a computer-readable medium, the computer program includes computer program code, the computer program code can be in the form of source code, object code, executable file or certain middleware, etc. The computer-readable medium includes any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the computer-readable medium includes but is not limited to the above-mentioned components.
[0113] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0114] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. When an element is referred to as being "assembled on", "installed on", "fixed on" or "set on" another element, it can be directly on the other element or there can also be a centered element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be a centered element at the same time. The terms "vertical", "horizontal", "up", "down", "left", "right" and similar expressions used herein are only for illustrative purposes and are not intended to be the only implementation method.
[0115] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
[0116] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
Claims
1. A multi-mobile agent anti-shake automatic aiming method, characterized in that: The following steps are involved: Receive the camera scene image frame and input it into the Yolov8-pose detector to obtain the target point position information recorded as ,in The Elements Indicates time The corresponding screen The coordinate position of the tracking target point in the two-dimensional plane of the camera; The improved multi-target tracking algorithm is used to predict, match and filter the location information of multiple target points to obtain new tracking points. ; Calculate the two-dimensional image Offset from the crosshair , according to the horizontal and vertical field of view of the camera, the offset Perform three-dimensional perspective transformation on each dimension to obtain the rotation vector of the camera ; Define the original execution volume as , the two states of "crosshair chasing" and "crosshair overspeed" that appear in the process of tracking the target affect the rotation vector Perform adaptive position PID control to achieve the original execution amount Adaptive adjustment to get the actual execution amount ; The actual execution Input external control device to rotate the camera's viewing angle vector The actual execution volume equal.
2. The multi-mobile agent anti-shake automatic aiming method according to claim 1, characterized in that: The camera scene image frame is obtained by capturing the real-time image of the camera, and the aspect ratio of the captured image is , the side length is .
3. The multi-mobile agent anti-shake automatic aiming method according to claim 2, characterized in that: The field of view of the camera is denoted by ,in and Represents the camera's field of view in the horizontal and vertical directions.
4. The multi-mobile agent anti-shake automatic aiming method according to claim 3, characterized in that: Use multi-target tracking algorithm to predict, match and filter multiple target points to obtain new tracking points , specifically including the following: (41) Prediction phase: Defining the mean of the Kalman filter for ,in Representative Points The estimated coordinates, Representative Points Normalized results of estimated moving speed in horizontal and vertical dimensions; At the moment , the Kalman filter is based on the mean of a point at the previous moment Predict the posterior mean of the point at the current moment , and obtain the posterior distribution ;take out The first two components of all elements in get the predicted point set ,in , yes The corresponding prediction point is denoted as ; (42) Matching stage, used to match point sets and Points in: (421) First, calculate the weighted Mahalanobis distance between different points in the two point sets ,in , , is the covariance matrix between two point sets; (422) To ensure that the tracking point at the previous moment is matched first, define the weight coefficient vector , Indicate point The corresponding weight coefficient is calculated as follows: (1) (2) (3) In the formula, Responsible for improving forecast points The matching weight of Indicate point The offset from the center of the picture, point The coordinates are expressed as ; is the offset component coefficient, which is used to control the effect of the offset in the weight coefficient, usually ,in is the side length of the camera's square image; during the matching process, even if If no suitable detection point is matched, the program can also give priority to selecting detection points closer to the center point as new tracking points to achieve small-scale tracking transition; (423) Define the cost matrix , whose elements are the normalized weighted Mahalanobis distances , after Hungarian matching, we get an ordered pair set ,in express and Match; (43) Screening phase: After matching is completed Tracking points is defined as ,definition Offset from the crosshair .
5. The multi-mobile agent anti-shake automatic aiming method according to claim 4, characterized in that: Offset Optimize the viewing angle and get , the calculation formula is as follows: (4) (5) In the formula, Indicates the side length of the screen, .
6. The multi-mobile agent anti-shake automatic aiming method according to claim 1, characterized in that: Before performing PID control, you first need to adjust the proportional coefficient To set the parameters, the calculation formula is as follows: (6) In the formula, Indicates the movable range of the hardware device. Indicates the field of view of the camera, confirm After that, according to the tracking situation Start fine-tuning the integral coefficient and the differential coefficient , the former can make up for Too small, the latter is responsible for predicting and suppressing Oversized situations; Will Considered as deviation, inspired by position PID, the execution amount is defined The calculation formula is as follows: (7) (8) (9) In the formula, represents the rotation vector, Represents the rotation vector relative to time The derivative of is the integral value of the limited deviation, and Indicates the upper and lower limits of the integral value of the deviation; is the conditional coefficient, Indicates when the integral fine-tuning takes effect In order to prevent the long-term lag of the crosshairs during high real-time tracking, which leads to a large accumulation of integrals, the integral limit and conditional integral are used to limit the integral term.
7. The multi-mobile agent anti-shake automatic aiming method according to claim 6, characterized in that: The rotation vector is affected by the two states of "crosshair chasing" and "crosshair overspeed" that occur when the crosshair is tracking the target. Perform adaptive position PID control, as follows: (71) First, determine whether the current tracking point is a new tracking point. If so, immediately perform full control to accelerate aiming, i.e. ; If not, enter the state subdivision control: (711) When the target moves at high speed and in the same direction as the crosshair, due to the delay in model reasoning, the crosshair will always lag behind the target, that is, the "crosshair chasing" state. At this time, full control is performed, that is, ; (712) When the target suddenly slows down or changes direction drastically, the crosshair will exceed the target, which is the "crosshair overspeed" state. At this time, define the force scaling factor vector To reduce the impact of execution volume and use the freeze frame number To control the number of frames of force reduction, ,in It is related to the motion state of the tracking point and the calculation formula is as follows: (10) Where e represents the base of the natural logarithm and defines the force scaling factors in the horizontal and vertical directions, where and Used to control the strength of force scaling in each direction, and the offset component accounts for The larger the offset, the more significant the force scaling. represents the posterior mean at time k, , They respectively represent the normalized results of the estimated values of the moving speed of point i in the horizontal and vertical directions at time k.
8. A multi-mobile agent anti-shake automatic aiming system, used to implement the multi-mobile agent anti-shake automatic aiming method according to any one of claims 1 to 7, characterized in that: include: The first input module is used to receive the camera scene image frame and input it into the Yolov8-pose detector to obtain the target point position information recorded as ,in The Elements Indicates time The corresponding screen The coordinate position of the tracking target point in the two-dimensional plane of the camera; The target point processing module is used to predict, match and filter the location information of multiple target points using an improved multi-target tracking algorithm to obtain new tracking points. ; The calculation module is used to calculate the Offset from the crosshair , according to the horizontal and vertical field of view of the camera, Perform three-dimensional perspective transformation on each dimension to obtain the rotation vector of the camera ; PID control module, used to define the original execution quantity as , the two states of "crosshair chasing" and "crosshair overspeed" that appear in the process of tracking the target affect the rotation vector Perform adaptive position PID control to achieve the original execution amount Adaptive adjustment to get the actual execution amount ; The second input module is used to convert the actual execution amount Input external control device to rotate the camera's viewing angle vector The actual execution volume equal.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the multi-mobile intelligent body anti-shake automatic aiming method described in any one of claims 1 to 7 is adopted.
10. A storage medium containing computer executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to perform the multi-mobile agent anti-shake automatic aiming method according to any one of claims 1 to 7.
Citation Information
Patent Citations
System and method for visual tracking
CN103268480A
Mechanical anti-shake processing method and device
CN105578146A