Detection method and system for automatically tracking human shape

By performing moving target and human figure detection on image frames and combining near-field or far-field position information, differentiated motor drive commands are generated, solving the problems of gimbal jitter and abnormal speed, and realizing a method and system for stable tracking of human targets.

CN120894397AInactive Publication Date: 2025-11-04HANGZHOU CLOSELI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511064183.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing automatic human tracking technologies are prone to problems such as target loss, gimbal movement irregularities, jitter, and abnormal speed in complex scenarios, especially in high-precision tracking applications.

Method used

By performing moving target detection and human detection on image frames, the coordinate information of human targets is saved. Combined with the near or far position, it is mapped to the 5×5 matrix coordinate system of the gimbal to generate differentiated motor drive commands to control the operation of the gimbal motor.

Benefits of technology

It achieves stable tracking of humanoid targets, avoids gimbal jitter and abnormal speed, and improves tracking accuracy and operational comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894397A_ABST
    Figure CN120894397A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic tracking, and particularly discloses a detection method and system for automatically tracking a human shape, and the method comprises the steps: firstly carrying out the moving target detection of a current image frame, further executing the human shape detection if a moving target exists, storing the coordinate information of a human shape target after the human shape target is detected, and determining that the human shape target is located at a close-shot or long-shot position; and then, mapping the coordinate information of the human-shaped target to a 5 * 5 matrix coordinate system of the holder, determining the offset level of the human-shaped target according to the positioning information of the grid where the target is located, generating a motor driving instruction containing the rotating speed and the step length from a preset adjustment strategy query table in combination with the close-shot or long-shot position information of the human-shaped target, and controlling the operation of a holder motor. According to the scheme, by positioning the target position, dynamically matching the coordinate system and differentially adjusting the motor parameters, stable tracking of the human-shaped target can be achieved, cradle head shaking and speed abnormity are avoided, and therefore tracking accuracy and operation comfort are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic tracking, and more particularly, to a detection method and system for automatically tracking a human form. BACKGROUND

[0002] With the rapid development of artificial intelligence and computer vision technology, the demand for automatically tracking human form targets is increasingly prominent in the fields of security, intelligent monitoring, robot navigation, etc. In complex scenes, real-time and stable tracking of moving human form targets can not only improve the response efficiency of security systems, but also optimize user experience. However, existing technologies often face problems such as target loss and unsmooth gimbal movement in practical applications, for example, the gimbal may appear jitter or lag due to the mismatch between response speed and target movement, resulting in tracking interruption or unstable picture. These defects not only reduce the reliability of the system, but also may cause misjudgment or operation delay, especially in situations requiring high-precision tracking, the limitations of existing solutions become even more significant.

[0003] Current mainstream automatic tracking solutions mostly rely on the simple combination of movement detection and human form recognition, by directly transmitting the detected target coordinates to the gimbal control system to achieve tracking control. However, this type of method often lacks dynamic perception ability for target positions (such as long-range or close-range), leading to the singleization of gimbal adjustment strategies. For example, when the target is in the long-range, if the fast rotation mode corresponding to the close-range is still used, it may cause the gimbal to respond excessively; conversely, if the target moves quickly to the close-range area, the gimbal may not be able to keep up due to the small step size, ultimately causing the target to be lost. In addition, traditional coordinate mapping methods usually use fixed partitioning or linear control, which is difficult to adapt to dynamic changes in complex scenes, further exacerbating the instability of gimbal movement.

[0004] Therefore, an optimized detection method and system for automatically tracking a human form are expected. SUMMARY

[0005] To solve the above technical problems, the present application is proposed. The embodiments of the present application provide a detection method and system for automatically tracking a human form, which first performs mobile target detection on the current image frame, and if there is a mobile target, further performs human form detection, saves the coordinate information of the human form target after detection, and confirms its position in the long-range or close-range. Then, the coordinate information of the human form target is mapped to the 5x5 matrix coordinate system of the gimbal, the offset level of the human form target is determined according to the positioning information of the target grid, and the motor driving instructions containing the rotation speed and step size are generated from the pre-set adjustment strategy query table in combination with the long-range or close-range position information of the human form target, to control the gimbal motor to run. This solution can realize stable tracking of the human form target by positioning the target position, dynamically matching the coordinate system, and differentiating the motor parameters, avoiding gimbal jitter and speed abnormalities, thereby improving tracking accuracy and running comfort.

[0006] Accordingly, according to an aspect of the present application, there is provided a method for automatically tracking a human form, comprising:

[0007] S1: performing mobile target detection on a first frame of image to obtain a mobile target detection result;

[0008] S2: in response to the mobile target detection result being that there is a mobile target, performing human form detection on the first frame of image to obtain a human form detection result;

[0009] S3: in response to the human form detection result being that there is a human form target, saving coordinate information of the human form target and confirming position information of the human form target;

[0010] S4: mapping the coordinate information of the human form target to a coordinate system of a pan-tilt head to obtain a coordinate mapping result, and generating a motor driving instruction in combination with the position information of the human form target, the motor driving instruction being used to adjust a rotation speed and a step length of a motor.

[0011] According to another aspect of the present application, there is provided a system for automatically tracking a human form, comprising:

[0012] a mobile target detection module, configured to perform mobile target detection on a first frame of image to obtain a mobile target detection result;

[0013] a human form detection module, configured to, in response to the mobile target detection result being that there is a mobile target, perform human form detection on the first frame of image to obtain a human form detection result;

[0014] a human form target position confirmation module, configured to, in response to the human form detection result being that there is a human form target, save coordinate information of the human form target and confirm position information of the human form target;

[0015] a motor driving instruction generation module, configured to map the coordinate information of the human form target to a coordinate system of a pan-tilt head to obtain a coordinate mapping result, and generate a motor driving instruction in combination with the position information of the human form target, the motor driving instruction being used to adjust a rotation speed and a step length of a motor.

[0016] Compared with the prior art, the detection method and system for automatically tracking a human figure provided by the application first perform mobile target detection on a current image frame, further perform human figure detection if there is a mobile target, save the coordinate information of the detected human figure target and confirm whether the human figure target is in a close-up or long shot position. Then, the coordinate information of the human figure target is mapped to a 5x5 matrix coordinate system of a pan-tilt head, the offset level of the human figure target is determined according to the positioning information of the grid where the target is located, and the motor driving instructions containing the rotation speed and the step length are generated from the preset adjustment strategy query table in combination with the close-up or long shot position information of the human figure target, so as to control the operation of the motor of the pan-tilt head. The scheme can realize stable tracking of the human figure target by positioning the target position, dynamically matching the coordinate system and differentiating the motor parameters, avoid the shaking and speed abnormality of the pan-tilt head, and thus improve the tracking accuracy and operation comfort. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application when taken in conjunction with the accompanying drawings. The drawings provided in the present application are used to provide further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally indicate the same components or steps.

[0018] Figure 1 A flowchart of the detection method for automatically tracking a human figure according to the embodiments of the present application.

[0019] Figure 2 A data flow diagram of the detection method for automatically tracking a human figure according to the embodiments of the present application.

[0020] Figure 3 A flowchart of step S3 in the detection method for automatically tracking a human figure according to the embodiments of the present application.

[0021] Figure 4 A flowchart of step S4 in the detection method for automatically tracking a human figure according to the embodiments of the present application.

[0022] Figure 5 A flowchart of step S43 in the detection method for automatically tracking a human figure according to the embodiments of the present application.

[0023] Figure 6 A flowchart of step S433 in the detection method for automatically tracking a human figure according to the embodiments of the present application.

[0024] Figure 7 A block diagram of the detection system for automatically tracking a human figure according to the embodiments of the present application. DETAILED DESCRIPTION

[0025] Hereinafter, the example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and the present application is not limited to the described example embodiments. It is worth noting that in the present application, all the actions of obtaining data are carried out in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the corresponding device owner.

[0026] Embodiment 1

[0027] Figure 1 Flowchart of the detection method for automatically tracking human figures according to the embodiments of the present application. Figure 2 Data flow diagram of the detection method for automatically tracking human figures according to the embodiments of the present application. As shown in Figure 1 and Figure 2 As shown in the detection method for automatically tracking human figures according to the embodiments of the present application, the method comprises the following steps: S1, performing moving target detection on the obtained first frame image to obtain a moving target detection result; S2, in response to the moving target detection result being that there is a moving target, performing human figure detection on the first frame image to obtain a human figure detection result; S3, in response to the human figure detection result being that there is a human figure target, saving the coordinate information of the human figure target and confirming the position information of the human figure target; S4, mapping the coordinate information of the human figure target to the coordinate system of the pan-tilt to obtain a coordinate mapping result, and generating a motor driving instruction in combination with the position information of the human figure target, the motor driving instruction being used to adjust the rotation speed and step length of the motor.

[0028] In the above detection method for automatically tracking human figures, the step S1, performing moving target detection on the obtained first frame image to obtain a moving target detection result. It can be understood that, in actual monitoring scenarios, there may be a large amount of static background information in the image, and directly using a human figure detection algorithm will consume a large amount of computing resources and be low in efficiency. Therefore, in order to quickly filter out invalid information and narrow the detection range, the present application identifies dynamic changes in the image based on a moving detection algorithm, and determines whether there is a moving target by analyzing pixel value changes, optical flow field and other features. In the embodiments of the present application, by comparing the pixel changes between consecutive image frames, when the pixel value difference between the first frame image and the next frame image exceeds a preset threshold, it is determined that there is a moving target. In this way, the image frame that needs to be further subjected to human figure detection can be quickly located, and the still image frame can be filtered out, so as to improve the detection efficiency and avoid invalid processing of static pictures.

[0029] In the above-described automatic human detection method, step S2, in response to the detection result indicating the presence of a moving target, performs human detection on the first frame image to obtain a human detection result. It should be understood that, considering the possibility that moving targets may include non-human objects (such as vehicles, animals, or floating objects), directly controlling the gimbal to track all moving targets would lead to mistracking and wasted resources. Therefore, to ensure the uniqueness and accuracy of the tracked target, after detecting a moving target in the first frame image, this application further uses a deep learning-based YOLO network to perform secondary detection on the first frame image, extracting moving target features to determine whether the target is human. Specifically, the YOLO network, as an advanced target detection algorithm, has been trained on a large dataset and can efficiently extract features from images and accurately determine the target category. In the embodiments of this application, the YOLO network performs convolution operations on the first frame image to extract feature information from the image, and performs target localization and human recognition based on the extracted image feature information, outputting the bounding box coordinates and category confidence of the moving target. In this system, bounding box coordinates are used to locate the specific position of the moving target in the image, while the class confidence score indicates the probability that a human-shaped target exists at that position. When the class confidence score of a human-shaped target exceeds a preset threshold, it is identified as a human-shaped target and proceeds to the next step; otherwise, the current frame is skipped and re-detection is performed. In this way, non-human interference can be effectively filtered out, ensuring the accuracy of target tracking.

[0030] In the above-described automatic human detection method, step S3, in response to the detection result indicating the presence of a human target, involves saving the coordinate information of the human target and confirming its location. Figure 3 This is a flowchart of step S3 in the automatic human figure tracking detection method according to an embodiment of this application. Figure 3 As shown, step S3 includes: S31, extracting the human-shaped target bounding box from the first frame image; S32, extracting the center point of the human-shaped target bounding box as the coordinate information of the human-shaped target; S33, determining the position information of the human-shaped target based on the size of the human-shaped target bounding box, wherein the position information is either a distant view or a close-up view. More specifically, step S33 includes: determining the position information of the human-shaped target based on a comparison between the height of the human-shaped target bounding box and a preset threshold; or, determining the position information of the human-shaped target based on a comparison between the area of ​​the human-shaped target bounding box and a preset threshold.

[0031] Specifically, the present application takes into account that the imaging size of a human-shaped target in an image is different at different distances, and the distance will affect the speed and accuracy requirements of the gimbal tracking. For example, a close-range target is imaged larger, and the gimbal needs to be finely adjusted to avoid the target out of the frame; a long-range target is imaged smaller, and the gimbal can be allowed to rotate at a faster speed to cover a larger range. Therefore, in order to improve the adaptability of the gimbal control, after detecting the human-shaped target, the present application further extracts the human-shaped target box in the first frame image, that is, the bounding box coordinates output by the above-mentioned YOLO network, and calculates the center point coordinates (x, y) of the human-shaped target based on the human-shaped target box as its coordinate position in the image coordinate system. At the same time, since the size of the human-shaped target in the picture reflects its actual distance relative to the camera, the present application compares the height size or area of the human-shaped target box with the preset height threshold and area threshold to determine whether the human-shaped target is in the close range or the long range. When the height size or area of the human-shaped target box is greater than or equal to the preset close-range threshold, it is determined that the human-shaped target is in the close range; otherwise, it is determined that the human-shaped target is in the long range (for example, when the target box height is greater than or equal to 1 / 3 of the image height, it is determined to be in the close range, and less than 1 / 5, it is determined to be in the long range). In this way, the distance of the human-shaped target can be preliminarily judged, not only the planar coordinates of the target are recorded, but also the spatial depth information is quantified, laying a data foundation for the dynamic adjustment of the motor driving parameters, which helps to solve the problem of single strategy caused by lack of distance perception in the traditional scheme.

[0032] In the above-mentioned detection method for automatically tracking a human-shaped target, the step S4 maps the coordinate information of the human-shaped target to the coordinate system of the gimbal to obtain a coordinate mapping result, and generates a motor driving instruction in combination with the position information of the human-shaped target, the motor driving instruction being used to adjust the rotation speed and step length of the motor. Wherein, Figure 4 The flow chart of step S4 in the detection method for automatically tracking a human-shaped target according to the embodiment of the present application is shown in FIG. 4. As shown in FIG. 4, the step S4 includes: S41, uniformly dividing the current camera field of view of the gimbal to obtain a 5x5 matrix as the coordinate system of the gimbal; S42, judging which grid the coordinate information of the human-shaped target is located in the 5x5 matrix to obtain grid positioning information; S43, generating the motor driving instruction based on the grid positioning information and the position information of the human-shaped target. Figure 4

[0033] ​In a specific example of the present application, the step S42 comprises: if the lattice positioning information is a center area, the offset level is level 0, wherein the center area is the center cell of the 5x5 matrix; if the lattice positioning information is a near-center area, the offset level is level 1, wherein the near-center area is the 8 cells adjacent to the center cell in the 5x5 matrix; and if the lattice positioning information is an edge area, the offset level is level 2, wherein the edge area is the 16 cells in the outermost circle of the 5x5 matrix. Specifically, the present application takes into account that the traditional fixed partition or linear coordinate mapping method cannot accurately reflect the degree of position deviation of the target in the gimbal field of view, and it is difficult to deal with the gimbal jitter problem caused by the rapid movement or position mutation of the target. Therefore, in order to realize the accurate matching of the position of the human-shaped target and the control of the gimbal, the present application divides the current camera field of view of the gimbal into a 5x5 matrix as a coordinate system based on the hierarchical control logic, and each cell corresponds to a different offset level: the center cell is level 0 (the smallest deviation, defined as the center area), the 8 cells adjacent to the center cell are level 1 (moderate deviation, defined as the near-center area), and the 16 cells in the outermost circle are level 2 (the largest deviation, defined as the edge area), so as to dynamically adjust the control strategy of the gimbal according to the different offset degrees of the human-shaped target in the 5x5 matrix.

[0034] In a specific example of the present application, the step S43 comprises: determining the offset level based on the lattice positioning information; and inputting the offset level and the position information into an adjustment strategy query table to obtain the motor driving instruction. More specifically, the motor driving instruction includes a target rotating speed and a target step length, the target rotating speed includes a low speed, a medium speed and a high speed, and the target step length includes a small step length, a medium step length and a large step length. That is, in actual application, by judging which cell in the 5x5 matrix the coordinate information of the human-shaped target is located, it is determined whether it is located in the center area, the near-center area or the edge area, so as to determine its offset level relative to the gimbal camera, and at the same time, in combination with the close / far position information of the human-shaped target, the corresponding motor driving instruction is called from the preset adjustment strategy query table, so as to realize the adaptive adjustment of the motor rotating speed and the step length. For example, close view + edge area corresponds to low speed + large step length, in order to avoid target loss caused by high-speed rotation; far view + center area corresponds to high speed + small step length, in order to quickly cover possible slight deviation while maintaining stability. In this way, the target position deviation is quantized as a level parameter, and then the target rotating speed (low speed / medium speed / high speed) and the step length (small step length / medium step length / large step length) are dynamically matched in combination with the distance information, so that the gimbal motor can smoothly adjust the motor operation according to the real-time position and distance of the target, effectively solving the problems of target loss, jitter and abnormal speed in the traditional scheme, and improving the stability and comfort of tracking.

[0035] Embodiment 2

[0036] In particular, considering that generating motor drive commands directly by querying a preset adjustment strategy table is limited by the fixed rules of the strategy table, it is prone to insufficient scene adaptability. Specifically, since the adjustment strategy relies entirely on manually designed mapping relationships (such as the combination of offset level and position information corresponding to fixed speed and step size), when facing complex dynamic scenes (such as the target rapidly switching between near and far views, non-uniform motion, or occlusion interference), the preset rules are difficult to cover all possible parameter combinations, resulting in gimbal response lag or jitter. In addition, the fixed strategy table lacks in-depth mining of multimodal correlation features (such as the correlation between position and offset), and cannot adaptively adjust the continuous changes of control parameters, limiting the generalization ability and tracking accuracy of human figure tracking. To address this, this application proposes an adaptive control method based on deep learning technology, which dynamically generates gimbal motor control parameters based on a data-driven approach, so as to make the gimbal rotation more closely match the target motion trajectory through continuous parameter control, while reducing the dependence on manual rules.

[0037] Figure 5 This is a flowchart of step S43 in the automatic human figure tracking detection method according to Embodiment 2 of this application. Figure 5 As shown, step S43 includes: S431, performing structured encoding on the grid positioning information and the position information of the humanoid target to obtain a structured encoding vector for the grid positioning information and a structured encoding vector for the target position information; S432, fusing the structured encoding vector for the grid positioning information and the structured encoding vector for the target position information to obtain a multimodal representation encoding vector for the humanoid target position information; S433, performing position-aware enhancement encoding on the multimodal representation encoding vector for the humanoid target position information to obtain an enhanced multimodal representation encoding vector for the humanoid target position information; S434, performing feature decoding on the enhanced multimodal representation encoding vector for the humanoid target position information to obtain a target rotational speed value and a target step size value; S435, generating the motor drive command based on the target rotational speed value and the target step size value.

[0038] Specifically, the step S431, the lattice positioning information and the position information of the human-shaped target are structured and coded to obtain a lattice positioning information structured coding vector and a target position information structured coding vector. It should be understood that, since the original lattice positioning information (such as the matrix coordinates of the lattice, the center area, the near center area or the edge area) and the target position information (such as the close shot and the long shot) are discrete category data, they cannot be directly used for training of the deep learning model. Therefore, in order to convert the two into learnable features in a high-dimensional continuous space, the present application is based on the Word Embedding idea in natural language processing, and the lattice positioning information and the position information of the human-shaped target are respectively structured and coded, the discrete classification labels are mapped to continuous vector representations with high-dimensional semantic information, the lattice positioning information structured coding vector and the target position information structured coding vector are formed, so as to retain the original information while enhancing the understanding and generalization ability of the model to the data.

[0039] Specifically, the step S432, the lattice positioning information structured coding vector and the target position information structured coding vector are fused to obtain a human-shaped target position information multi-modal representation coding vector. It should be understood that, since the lattice positioning information and the target position information respectively reflect the offset degree and distance attribute of the target in space, a single modal feature cannot comprehensively describe the influence of the target state on the pan-tilt control. Therefore, in order to capture the complementarity and interaction of the two types of information, the present application fuses the lattice positioning information structured coding vector and the target position information structured coding vector into a unified multi-modal representation through feature concatenation, and generates a human-shaped target position information multi-modal representation coding vector. In this way, it is helpful to simultaneously perceive the spatial offset trend and depth distance characteristics of the target, and to enhance the representation ability of complex motion patterns.

[0040] Specifically, the step S433, the human-shaped target position information multi-modal representation coding vector is subjected to position state perception enhancement coding to obtain a human-shaped target position information reinforced multi-modal representation coding vector. Specifically, considering that the human-shaped target position information multi-modal representation coding vector has fused the lattice positioning information and the target position information, but simple feature concatenation may not be able to fully explore the deep-level correlation and potential law between the two types of information, resulting in limited perception and representation ability of the target motion state. Therefore, in order to further improve the accurate perception and deep understanding of the target position state, the present application proposes a position state perception enhancement coding method, which aims to extract more discriminative multi-modal features through transformation and refinement of the feature space, remove redundant information, and form a human-shaped target position information reinforced multi-modal representation coding vector.

[0041] Figure 6The flow chart of step S433 in the automatic tracking human shape detection method according to embodiment 2 of the present application is shown. As shown in Figure 6 S4331, one-dimensional convolution kernel-based feature deconstruction is performed on the human target position information multi-modal representation encoding vector to obtain a set of human target position information local feature multi-modal representation encoding vectors; S4332, based on the context feature association topology of the set of human target position information local feature multi-modal representation encoding vectors, feature enhancement mapping is performed on each human target position information local feature multi-modal representation encoding vector in the set of human target position information local feature multi-modal representation encoding vectors to obtain a set of human target position information local feature multi-modal representation enhanced encoding vectors; S4333, self-attention-based feature enhancement fusion is performed on the set of human target position information local feature multi-modal representation enhanced encoding vectors to obtain the human target position information enhanced multi-modal representation encoding vector.

[0042] More specifically, the step S4331 can be expressed by the formula:

[0043]

[0044]

[0045] wherein, represents the human target position information multi-modal representation encoding vector, represents one-dimensional convolution operation based on convolution kernel, is the scale of one-dimensional convolution kernel, represents the set of human target position information local feature multi-modal representation encoding vectors, , , and respectively represent the first, second, third and fourth human target position information local feature multi-modal representation encoding vectors in the set of human target position information local feature multi-modal representation encoding vectors, is the number of human target position information local feature multi-modal representation encoding vectors.

[0046] ​​Specifically, since the original human target position information multi-modal representation encoding vector can contain redundant global information or noise, direct use for control instruction generation is easy to be disturbed by irrelevant features. Therefore, in order to extract more discriminative local features and reveal the potential correlation patterns inside, the present application first performs feature decomposition on the human target position information multi-modal representation encoding vector based on one-dimensional convolution coding, captures the local structure information in the human target position information multi-modal representation encoding vector through a sliding window of one-dimensional convolution kernel, realizes efficient pattern decoupling by using weight sharing and local receptive field characteristics, and generates a set of human target position information local feature multi-modal representation encoding vectors. In this way, the original high-dimensional features are decomposed into multiple low-order, interpretable local basis vectors, which helps to explicit the key feature components and the inherent implicit correlation patterns in the target position state, and lays a foundation for subsequent feature enhancement modeling.

[0047] More specifically, the step S4332 comprises: first, calculating the feature correlation factor between any two human target position information local feature multi-modal representation encoding vectors in the set of human target position information local feature multi-modal representation encoding vectors to obtain a multi-modal human target position information local feature correlation topology matrix composed of multiple feature correlation factors, which is expressed by the formula:

[0048]

[0049] wherein, represents the i-th human target position information local feature multi-modal representation encoding vector in the set of human target position information local feature multi-modal representation encoding vectors, the i-th human target position information local feature multi-modal representation encoding vector, represents the transpose of the vector, represents the 2-norm of the vector, represents the bandwidth parameter, represents the exponential function with base e, represents and the feature correlation factor between the two vectors, i.e. the element value in the i-th position of the multi-modal human target position information local feature correlation topology matrix.

[0050] ​Here, since there is a potential spatial or semantic correlation between the local feature multi-modal representation encoding vectors of the decomposed each human target position information, the influence of the synergistic effect on the PTZ control may be ignored if it is processed independently. Therefore, in order to quantify the interaction relationship between the local features and construct a global context awareness model, the scheme calculates the feature correlation factor of any two local feature multi-modal representation encoding vectors of the human target position information, and arranges all feature correlation factors according to the row and column positions to form a multi-modal human target position information local feature correlation topology matrix, which explicitly describes the dependence strength between local features, and provides structured guidance for subsequent noise suppression and information fusion.

[0051] Then, the multi-modal human target position information local feature correlation topology matrix is input into a gating mask function to obtain a multi-modal human target position information local feature fine-grained correlation mask topology matrix, which is expressed by a formula as follows:

[0052]

[0053] wherein, the gating mask weight matrix is represented by, the multi-modal human target position information local feature correlation topology matrix is represented by, the gating mask bias matrix is represented by, the sigmoid activation function is represented by, the multi-modal human target position information local feature fine-grained correlation mask topology matrix is represented by.

[0054] It can be understood that, considering that the original multi-modal human target position information local feature correlation topology matrix may contain invalid correlations introduced by noise or redundant connections, direct use will affect the effective interaction and information transmission between features. Therefore, in order to dynamically focus on key correlations and suppress interference, the scheme introduces a learnable gating mask network to modulate the multi-modal human target position information local feature correlation topology matrix, and realizes soft screening of feature correlation through adaptive weight adjustment, for example, strong correlations (such as the correlation between the center area and the near-range features) are given high weights, while weak correlations are suppressed, thereby constructing a multi-modal human target position information local feature fine-grained correlation mask topology matrix, retaining feature interaction paths strongly related to PTZ control, and improving the robustness of the subsequent feature enhancement process.

[0055] Then, in one preferred example of the present application, the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix is locally graphically correlated and balanced to obtain an optimized multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix. In particular, considering that the multi-modal human-shaped target position information local feature correlation topology matrix may have a response undersaturation in the feature coupling configuration layout, so that the global correlation form framework of the multi-modal human-shaped target position information local feature correlation topology matrix is contracted due to nonlinear coupling correlation, and the interaction polarization gain effect of the gating mechanism aggravates the representation distortion of the underlying structure micro-correlation of the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix. Based on this, the present application further balances the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix by locally graphically correlating and balancing to alleviate the compression problem of its geometric configuration balance solution and enhance its feature expression effect.

[0056] Specifically, first, for each feature value of the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix , an adaptive gradient field vector is introduced to correct the local non-uniform structure interaction, so as to realize the micro-scale structure balancing of the morphological coupling field:

[0057]

[0058] wherein represents the element in the th row and the th column, represents the corresponding dynamic gradient field quantity in the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix , represents the element in the th row and the th column in , represents the partial derivative.

[0059] Then, the dynamic gradient field quantity is used as an external excitation term to dynamically adjust the average field of each feature value in the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix :

[0060]

[0061] wherein, is the multi-modal human-shaped target position information local feature fine-grained correlation mask topology matrix​​ a feature mean value of all feature values of, denotes an adjustment gain factor, denotes an element in the i-th row and the j-th column of the optimized received signal strength local timing feature structure fine-grained correlation mask topology matrix. denotes an element in the i-th row and the j-th column of the optimized received signal strength local timing feature structure fine-grained correlation mask topology matrix. denotes an element in the i-th row and the j-th column of the optimized received signal strength local timing feature structure fine-grained correlation mask topology matrix.

[0062] In this way, by the action of the high-order differential excitation component, the negative feedback adjusts the structure interaction configuration of the multi-modal humanoid target position information local feature fine-grained correlation mask topology matrix under the nonlinear response saturation of the macroscopic mean field, and the average field harmonic correction is used to offset the substructure interaction decoupling caused by the interaction polarization gain effect, thereby improving the expression effect of the underlying micro-interaction structure of the multi-modal humanoid target position information local feature fine-grained correlation mask topology matrix.

[0063] Further, each humanoid target position information local feature multi-modal representation encoding vector in the set of humanoid target position information local feature multi-modal representation encoding vectors and the optimized multi-modal humanoid target position information local feature fine-grained correlation mask topology matrix are input into an attribute dense feedback enhancement unit to obtain a set of humanoid target position information local feature multi-modal representation reinforced encoding vectors, which is expressed by the formula:

[0064]

[0065] wherein, denotes an optimized multi-modal humanoid target position information local feature fine-grained correlation mask topology matrix, denotes an enhancement weight matrix, denotes a dot product, denotes a matrix multiplication operation, is a nonlinear activation function, denotes a feature scale value of, denotes the i-th humanoid target position information local feature multi-modal representation reinforced encoding vector in the set of humanoid target position information local feature multi-modal representation reinforced encoding vectors. denotes the i-th humanoid target position information local feature multi-modal representation reinforced encoding vector in the set of humanoid target position information local feature multi-modal representation reinforced encoding vectors.

[0066] ​Here, considering that the information of the local feature fragments may exist ambiguity or missing (such as the single-view displacement cannot reflect the global motion trend), isolated optimization can lead to local optimization of the control command. Therefore, in order to integrate multi-source associated information and extract core features, the application optimizes each local feature based on the multi-modal human target position information local feature fine-grained association mask topology matrix to fuse the associated information between local features into the current feature through neighborhood message passing, realizes context-aware feature enhancement, and finally generates a set of human target position information local feature multi-modal representation reinforcement encoding vectors. In this way, each local feature not only contains its own information, but also fuses the context clues of the associated features, eliminates the ambiguity of a single fragment, and forms a more consistent and more discriminative feature representation.

[0067] More specifically, the step S4333 can be represented by a formula as follows:

[0068]

[0069]

[0070] wherein, represents a set of human target position information local feature multi-modal representation reinforcement encoding vectors, , and represent the first, second and third of the set of human target position information local feature multi-modal representation reinforcement encoding vectors, personal human target position information local feature multi-modal representation reinforcement encoding vector, represents a feature reinforcement reconstruction operation, , and respectively represent a query matrix, a key matrix and a value matrix, , and respectively represent a query embedding matrix, a key embedding matrix and a value embedding matrix, represents a normalized exponential function, represents the human target position information reinforcement multi-modal representation encoding vector.

[0071] It should be appreciated that the enhanced local features still need to be integrated into a global representation to generate coherent control instructions, and simple concatenation or pooling may lose long-range dependencies. Therefore, in order to dynamically capture the global synergistic effect between local features, the scheme is based on the self-attention mechanism for feature reconstruction, and the importance weight between features is calculated to realize context-adaptive information aggregation. In this way, the model can adaptively focus on the key state information of the human target, intelligently integrate the dispersed local features into a unified enhanced representation, and obtain a human target position information enhanced multi-modal representation encoding vector.

[0072] Specifically, the step S434, the human target position information enhanced multi-modal representation encoding vector is feature decoded to obtain a target rotation speed value and a target step value. It should be appreciated that, considering that compared to the predetermined rotation speed / step level division, continuous parameter control can more subtly match the dynamic changes of the target, and realize more smooth and accurate pan-tilt rotation. Therefore, in the present application, the decoder structure in the deep learning network is used to perform feature decoding on the human target position information enhanced multi-modal representation encoding vector, and map it to a specific target rotation speed value and a target step value. In a specific embodiment of the present application, the human target position information enhanced multi-modal representation encoding vector is input into a fully connected network containing two output branches, one of which is used to predict the target rotation speed value, and the other of which is used to predict the target step value, the fully connected network realizes synchronous optimization of motor rotation speed and step through double-branch parallel processing, and outputs the target rotation speed value and the target step value. In this way, smooth motor instructions can be generated according to the dynamic state of the human target, avoiding the pan-tilt jitter caused by parameter jumps in traditional methods.

[0073] Specifically, the step S435, based on the target rotation speed value and the target step value, the motor driving instruction is generated. It should be appreciated that the motor driving system needs explicit control instructions (such as PWM signal duty cycle, pulse number) to adjust the rotation speed and step. Therefore, in order to convert the continuous rotation speed value and step value into digital signals that can be recognized by the motor controller, and realize precise driving, the present application establishes a rotation speed-duty cycle mapping table (such as rotation speed 50rpm corresponding to PWM duty cycle 30%, 100rpm corresponding to 60%) and a step value-pulse number mapping table (such as step 100steps corresponding to sending 100 pulse signals) based on the underlying protocol of motor control, and converts the decoded rotation speed value and step value into corresponding PWM signal parameters and pulse count through a microcontroller (such as STM32), and sends it to the motor driver through the bus. In this way, the pan-tilt motor can be ensured to rotate smoothly at the optimal rotation speed and step according to the real-time position and dynamic state of the human target, improving the stability and comfort of tracking, and enhancing the scene adaptability of the pan-tilt system.

[0074] In summary, the automatic human detection tracking method according to the embodiments of this application is explained. First, moving target detection is performed on the current image frame. If a moving target is found, human detection is further performed. After detecting a human target, its coordinate information is saved, and its position (near or far) is confirmed. Next, the coordinate information of the human target is mapped onto the 5×5 matrix coordinate system of the gimbal. The offset level of the human target is determined based on the positioning information of the grid where the target is located. Then, combined with its near or far position information, a motor drive command containing rotational speed and step size is generated from a preset adjustment strategy lookup table to control the operation of the gimbal motor. This scheme, by locating the target position, dynamically matching the coordinate system, and differentially adjusting motor parameters, can achieve stable tracking of human targets, avoiding gimbal jitter and abnormal speed, thereby improving tracking accuracy and operational comfort.

[0075] Furthermore, this application also provides an automatic human figure tracking detection system.

[0076] Figure 7 This is a block diagram of an automatic human figure tracking detection system according to an embodiment of this application. Figure 7 As shown, the automatic human detection system 100 according to an embodiment of this application includes: a moving target detection module 110, used to perform moving target detection on a first frame image to obtain a moving target detection result; a human detection module 120, used to perform human detection on the first frame image in response to the moving target detection result indicating the presence of a moving target to obtain a human detection result; a human target position confirmation module 130, used to save the coordinate information of the human target and confirm the position information of the human target in response to the human target detection result indicating the presence of a human target; and a motor drive command generation module 140, used to map the coordinate information of the human target to the coordinate system of the gimbal to obtain a coordinate mapping result, and generate a motor drive command in combination with the position information of the human target, wherein the motor drive command is used to adjust the speed and step size of the motor.

[0077] Here, those skilled in the art will understand that the specific operation of each module in the aforementioned automatic human figure tracking detection system has been described above according to... Figures 1 to 6 The method for automatically tracking human figures is described in detail in the description of the method, and therefore, its repeated description will be omitted.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for automatically tracking and detecting human figures, characterized in that, include: S1: Perform moving target detection on the acquired first frame image to obtain the moving target detection result; S2: In response to the moving target detection result indicating the presence of a moving target, perform human detection on the first frame image to obtain a human detection result; S3: In response to the human detection result indicating the presence of a human target, save the coordinate information of the human target and confirm the location information of the human target; S4: Map the coordinate information of the humanoid target to the coordinate system of the gimbal to obtain the coordinate mapping result, and generate motor drive commands in combination with the position information of the humanoid target. The motor drive commands are used to adjust the speed and step size of the motor.

2. The automatic human figure tracking detection method according to claim 1, characterized in that, Step S3 includes: Extract the human-shaped bounding box from the first frame image; Extract the center point of the humanoid target bounding box as the coordinate information of the humanoid target; Based on the size of the human-shaped target frame, the position information of the human-shaped target is determined, wherein the position information is either a distant view or a close-up view.

3. The automatic human figure tracking detection method according to claim 2, characterized in that, Based on the dimensions of the human-shaped target bounding box, the position information of the human-shaped target is determined, wherein the position information is either a distant view or a close-up view, including: The location information of the human-shaped target is determined by comparing the height of the human-shaped target frame with a preset threshold; or, the location information of the human-shaped target is determined by comparing the area of ​​the human-shaped target frame with a preset threshold.

4. The automatic human figure tracking detection method according to claim 1, characterized in that, Step S4 includes: The current camera field of view of the gimbal is evenly divided to obtain a 5×5 matrix as the coordinate system of the gimbal; Determine which cell in the 5×5 matrix the coordinates of the humanoid target are located to obtain cell positioning information; Based on the grid positioning information and the position information of the humanoid target, the motor drive command is generated.

5. The automatic human figure tracking detection method according to claim 4, characterized in that, Based on the grid positioning information and the position information of the humanoid target, the motor drive command is generated, including: Based on the grid positioning information, the offset level is determined; The offset level and the position information are input into the adjustment strategy lookup table to obtain the motor drive command.

6. The automatic human figure tracking detection method according to claim 5, characterized in that, Based on the grid positioning information, the offset level is determined, including: If the grid positioning information is the center area, then the offset level is level 0, where the center area is the center grid of a 5×5 matrix; If the grid positioning information is near the center area, then the offset level is level 1, where the near center area is the 8 grids in a 5×5 matrix that are immediately adjacent to the center grid. If the grid positioning information is an edge area, then the offset level is level 2, where the edge area is the outermost 16 grids of the 5×5 matrix.

7. The automatic human figure tracking detection method according to claim 5, characterized in that, The motor drive command includes a target speed and a target step size. The target speed includes low speed, medium speed and high speed, and the target step size includes small step size, medium step size and large step size.

8. The automatic human figure tracking detection method according to claim 4, characterized in that, Based on the grid positioning information and the position information of the humanoid target, the motor drive command is generated, including: The grid positioning information and the humanoid target's position information are structured and encoded to obtain a grid positioning information structured encoding vector and a target position information structured encoding vector; The structured encoding vectors of the grid positioning information and the target position information are fused to obtain the multimodal representation encoding vector of the human-shaped target position information; The humanoid target position information multimodal representation encoding vector is subjected to position-aware enhanced encoding to obtain the humanoid target position information enhanced multimodal representation encoding vector; The enhanced multimodal representation encoding vector of the humanoid target position information is used for feature decoding to obtain the target rotation speed value and the target step size value; The motor drive command is generated based on the target speed value and the target step size value.

9. The automatic human figure tracking detection method according to claim 8, characterized in that, The humanoid target location information multimodal representation encoding vector is subjected to position-aware enhanced encoding to obtain the humanoid target location information enhanced multimodal representation encoding vector, including: The humanoid target location information multimodal representation encoding vector is subjected to feature deconstruction based on a one-dimensional convolution kernel to obtain a set of local feature multimodal representation encoding vectors of humanoid target location information; Based on the context feature association topology of the set of local feature multimodal representation encoding vectors of human target location information, feature enhancement mapping is performed on each local feature multimodal representation encoding vector of human target location information in the set of local feature multimodal representation encoding vectors of human target location information to obtain a set of local feature multimodal representation enhancement encoding vectors of human target location information; The set of enhanced encoding vectors for local feature multimodal representation of the humanoid target's location information is subjected to self-attention-based feature enhancement fusion to obtain the enhanced multimodal representation encoding vector for the humanoid target's location information.

10. A detection system for automatically tracking human figures, characterized in that, include: The moving target detection module is used to perform moving target detection on the acquired first frame image to obtain the moving target detection result; The human detection module is used to perform human detection on the first frame image to obtain a human detection result in response to the moving target detection result indicating the presence of a moving target. The human target location confirmation module is used to, in response to the human detection result indicating the presence of a human target, save the coordinate information of the human target and confirm the location information of the human target; The motor drive command generation module is used to map the coordinate information of the humanoid target to the coordinate system of the gimbal to obtain the coordinate mapping result, and to generate motor drive commands in combination with the position information of the humanoid target. The motor drive commands are used to adjust the speed and step size of the motor.