Target Tracking Method, Device, Electronic Device and Storage Medium

By using multiple scale adjustment coefficient groups in target tracking, the problem of target tracking stability and accuracy in complex scenarios is solved, and more efficient target tracking effect is achieved.

CN115471527BActive Publication Date: 2025-06-13CHONGQING UNISINSIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211174875.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2025-06-13
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

In complex scenarios, target tracking is prone to loss or missed targets due to target scale deformation, large scale changes and target occlusion.

Method used

By obtaining the current video frame and its template parameters, selecting multiple scale adjustment coefficient groups, obtaining candidate areas in the current video frame based on these coefficient groups, performing feature extraction and calculating response values, and selecting the scale adjustment coefficient group corresponding to the maximum response value as the optimal scale adjustment coefficient group to achieve adaptive tracking of the target.

Benefits of technology

It improves the stability and accuracy of target tracking and can effectively deal with target scale changes and occlusion situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115471527B_ABST
    Figure CN115471527B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of target tracking, and provides a target tracking method, device, electronic device and storage medium. By obtaining the current video frame and the template parameters of the current video frame calculated from the position and scale of the target in the previous video frame; selecting multiple scale adjustment coefficient groups from a preset set of scale adjustment coefficients; according to each scale adjustment coefficient group, and the position and scale of the target in the previous video frame, obtaining candidate regions corresponding to each scale adjustment coefficient group in the current video frame, and performing feature extraction to obtain each target feature matrix, calculating a response value according to each target feature matrix and the template parameters of the current video frame, and obtaining the response value corresponding to each scale adjustment coefficient group; taking the scale adjustment coefficient group corresponding to the maximum response value as the optimal scale adjustment coefficient group of the current video frame and obtaining the position and scale of the target in the current video frame, so as to track the target. The stability and accuracy of target tracking are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and in particular, to a target tracking method, device, electronic device, and storage medium. Background Art

[0002] Target tracking is widely used in security monitoring systems. Commonly used target tracking methods are mainly divided into two categories: generative and discriminative. The generative method is to establish a model in advance and then use the model to search and compare similar features during subsequent tracking. The discriminative method is to extract the foreground target by comparing the differences between the target and the background information. However, in the face of actual scenarios such as urban streets, complex roads, parks, etc., there are often situations where the target scale deforms, the scale changes greatly, and the target is occluded, resulting in problems such as target loss and missed tracking. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a target tracking method, device, electronic device, and storage medium.

[0004] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:

[0005] In a first aspect, the present invention provides a target tracking method, and the method includes:

[0006] Obtain the current video frame and the template parameters of the current video frame, where the template parameters of the current video frame are calculated based on the position and scale of the target in the previous video frame;

[0007] Select multiple scale adjustment coefficient groups from a preset set of scale adjustment coefficients;

[0008] According to each scale adjustment coefficient group, and the position and scale of the target in the previous video frame, obtain the candidate region corresponding to each scale adjustment coefficient group in the current video frame;

[0009] Extract features from each candidate region to obtain each target feature matrix, and calculate a response value based on each target feature matrix and the template parameters of the current video frame, to obtain the response value corresponding to each scale adjustment coefficient group; the response value represents the similarity between the candidate region and the region of the target in the previous video frame;

[0010] Take the scale adjustment coefficient group corresponding to the maximum response value as the best scale adjustment coefficient group of the current video frame, and obtain the position and scale of the target in the current video frame according to the best scale adjustment coefficient group, so as to track the target.

[0011] In an alternative embodiment, the set of scale adjustment coefficients includes a width coefficient sequence and a height coefficient sequence. The width coefficient sequence includes a set number of width adjustment coefficients, and the height coefficient sequence includes a set number of height adjustment coefficients;

[0012] The step of selecting a plurality of scale adjustment coefficient groups from a preset set of scale adjustment coefficients includes:

[0013] Obtain the optimal scale adjustment coefficient group of the previous video frame to obtain a target width adjustment coefficient and a target height adjustment coefficient;

[0014] According to a first preset rule and the target width adjustment coefficient, select a plurality of candidate width adjustment coefficients from the width coefficient sequence;

[0015] According to a second preset rule and the target height adjustment coefficient, select a plurality of candidate height adjustment coefficients from the height coefficient sequence;

[0016] Combine each of the candidate width adjustment coefficients with each of the candidate height adjustment coefficients to obtain each of the scale adjustment coefficient groups. One scale adjustment coefficient group includes one of the candidate width adjustment coefficients and one of the candidate height adjustment coefficients.

[0017] In an alternative embodiment, the set number is odd, all the width adjustment coefficients in the width coefficient sequence are arranged in ascending order, and the width coefficient sequence has a corresponding first window;

[0018] The step of selecting a plurality of candidate width adjustment coefficients from the width coefficient sequence according to the target width adjustment coefficient and the first preset rule includes:

[0019] If the target width adjustment coefficient is less than the median of all the width adjustment coefficients, control the first window to slide forward one step from the current position, and obtain each width adjustment coefficient in the slid first window to obtain each of the candidate width adjustment coefficients;

[0020] If the target width adjustment coefficient is equal to the median of all the width adjustment coefficients, obtain each width adjustment coefficient in the first window to obtain each of the candidate width adjustment coefficients;

[0021] If the target width adjustment coefficient is greater than the median of all the width adjustment coefficients, control the first window to slide backward one step from the current position, and obtain each width adjustment coefficient in the slid first window to obtain each of the candidate width adjustment coefficients.

[0022] In an alternative embodiment, the template parameters of the current video frame are calculated in the following manner:

[0023] Calculate the candidate template parameters of the previous video frame according to the position and scale of the target in the previous video frame;

[0024] Calculate the template parameters of the current video frame according to the candidate template parameters and the template parameters of the previous video frame.

[0025] In an alternative embodiment, the step of calculating the template parameters of the current video frame according to the candidate template parameters and the template parameters of the previous video frame includes:

[0026] Calculate the template parameters of the current video frame according to a preset formula, a preset learning rate, the candidate template parameters and the template parameters of the previous video frame;

[0027] The preset formula is:

[0028] α i = θα′ i-1 +(1 - θ)α i-1 ;

[0029] Wherein, α i represents the template parameters of the current video frame; α′ i-1 represents the candidate template parameters of the previous video frame; α i-1 represents the template parameters of the previous video frame; θ represents the preset learning rate.

[0030] In an alternative embodiment, the step of calculating the template parameters of the current video frame according to the candidate template parameters and the template parameters of the previous video frame includes:

[0031] Obtain the best scale adjustment coefficient group of the previous video frame to obtain the target width adjustment coefficient and the target height adjustment coefficient;

[0032] Select a target learning rate from a preset learning rate set according to a third preset rule, the target width adjustment coefficient and the target height adjustment coefficient;

[0033] Calculate the template parameters of the current video frame according to a preset formula, the target learning rate, the candidate template parameters and the template parameters of the previous video frame; the preset formula is:

[0034] α i = θα′ i-1 +(1 - θ)α i-1 ;

[0035] Wherein, α i represents the template parameters of the current video frame; α′ i-1 represents the candidate template parameters of the previous video frame; α i-1The template parameters representing the previous video frame; θ represents the target learning rate.

[0036] In an optional embodiment, the set of scale adjustment coefficients includes a width coefficient sequence and a height coefficient sequence. The width coefficient sequence includes a set number of width adjustment coefficients, and the height coefficient sequence includes a set number of height adjustment coefficients. Each width adjustment coefficient in the width coefficient sequence is equal to the height adjustment coefficient in the corresponding order in the height coefficient sequence, and the set number is odd. The set of learning rates includes a first learning rate, a second learning rate, and a third learning rate. The second learning rate is greater than the first learning rate, and the first learning rate is greater than the third learning rate.

[0037] The step of selecting a target learning rate from a preset set of learning rates according to the third preset rule, the target width adjustment coefficient, and the target height adjustment coefficient includes:

[0038] If the target width adjustment coefficient is equal to the median of all width adjustment coefficients and the target height adjustment coefficient is equal to the median of all height adjustment coefficients, then select the third learning rate as the target learning rate.

[0039] If the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is equal to the target height adjustment coefficient, then select the first learning rate as the target learning rate.

[0040] If the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is not equal to the target height adjustment coefficient, then select the second learning rate as the target learning rate.

[0041] In a second aspect, the present invention provides a target tracking device, and the device includes:

[0042] An acquisition module, configured to acquire a current video frame and the template parameters of the current video frame, where the template parameters of the current video frame are calculated according to the position and scale of the target in the previous video frame.

[0043] A processing module, configured to select a plurality of scale adjustment coefficient groups from a preset set of scale adjustment coefficients.

[0044] According to each scale adjustment coefficient group, and the position and scale of the target in the previous video frame, obtain a candidate region corresponding to each scale adjustment coefficient group in the current video frame.

[0045] Feature extraction is performed on each of the candidate regions to obtain each target feature matrix, and response values are calculated based on each target feature matrix and the template parameters of the current video frame, so as to obtain the response values corresponding to each group of scale adjustment coefficients; the response value represents the similarity between the candidate region and the region of the target in the previous video frame.

[0046] A tracking module, configured to use the group of scale adjustment coefficients corresponding to the maximum response value as the optimal group of scale adjustment coefficients for the current video frame, and obtain the position and scale of the target in the current video frame according to the optimal group of scale adjustment coefficients, so as to track the target.

[0047] In a third aspect, the present invention provides an electronic device, including a processor and a memory, where the memory stores a computer program, and when the processor executes the computer program, the method described in any one of the foregoing embodiments is implemented.

[0048] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the foregoing embodiments is implemented.

[0049] The object tracking method, device, electronic device and storage medium provided by the embodiments of the present invention obtain the current video frame and the template parameters of the current video frame, where the template parameters of the current video frame are calculated according to the position and scale of the target in the previous video frame; then multiple groups of scale adjustment coefficients are selected from a preset set of scale adjustment coefficients; and according to each group of scale adjustment coefficients, as well as the position and scale of the target in the previous video frame, candidate regions corresponding to each group of scale adjustment coefficients are obtained in the current video frame; then feature extraction is performed on each candidate region to obtain each target feature matrix, and response values are calculated based on each target feature matrix and the template parameters of the current video frame, so as to obtain the response values corresponding to each group of scale adjustment coefficients, and the response value represents the similarity between the candidate region and the region of the target in the previous video frame; finally, the group of scale adjustment coefficients corresponding to the maximum response value is used as the optimal group of scale adjustment coefficients for the current video frame, and the position and scale of the target in the current video frame are obtained according to the optimal group of scale adjustment coefficients, so as to track the target. Adaptive tracking of a target with scale changes is achieved through multiple groups of scale adjustment coefficients, and the association of the target in adjacent video frames is established based on the template parameters, thereby improving the stability and accuracy of target tracking.

[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings

[0051] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0052] Figure 1 It shows a schematic block diagram of an electronic device provided by an embodiment of the present invention;

[0053] Figure 2 It shows a schematic flowchart of a target tracking method provided by an embodiment of the present invention;

[0054] Figure 3 It shows another schematic flowchart of a target tracking method provided by an embodiment of the present invention;

[0055] Figure 4 It shows an example diagram of a target tracking method provided by an embodiment of the present invention;

[0056] Figure 5 It shows another schematic flowchart of a target tracking method provided by an embodiment of the present invention;

[0057] Figure 6 It shows a functional module diagram of a target tracking device provided by an embodiment of the present invention.

[0058] Icons: 110 - Bus; 120 - Processor; 130 - Memory; 150 - I / O Module; 170 - Communication Interface; 300 - Target Tracking Device; 310 - Acquisition Module; 330 - Processing Module; 350 - Tracking Module; 370 - Calculation Module. Detailed Embodiments

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0060] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0061] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0062] Object tracking is widely used in security monitoring systems. There are mainly two categories of common object tracking methods, namely generative and discriminative. The generative method is to establish a model in advance, and then use the model to search and compare similar features during subsequent tracking. The discriminative method is to extract the foreground object by comparing the differences between the object and the background information. However, in the face of actual scenarios such as urban streets, complex roads, parks, etc., there are often situations where the object scale deforms, the scale changes greatly, and the object is occluded, resulting in problems such as object loss and missed tracking. Therefore, the embodiments of the present invention provide an object tracking method to solve the above problems.

[0063] Please refer to Figure 1 , which is a block diagram of an electronic device provided by an embodiment of the present invention. The electronic device includes a bus 110, a processor 120, a memory 130, an I / O module 150, and a communication interface 170.

[0064] The bus 110 can be a circuit that interconnects the above elements and transmits communication (such as control messages) between the above elements.

[0065] The processor 120 can receive commands from the above other elements (such as the memory 130, the I / O module 150, the communication interface 170, etc.) through the bus 110, can interpret the received commands, and can perform calculations or data processing according to the interpreted commands.

[0066] The processor 120 can be an integrated circuit chip with signal processing capabilities. The processor 120 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0067] The memory 130 may store commands or data received from the processor 120 or other components (such as the I / O module 150, the communication interface 170, etc.) or commands or data generated by the processor 120 or other components.

[0068] The memory 130 may be, but is not limited to, a Random Access Memory (RAM), a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM).

[0069] The I / O module 150 may receive commands or data input by a user via input-output means (such as sensors, keyboards, touchscreens, etc.), and may transmit the received commands or data to the processor 120 or the memory 130 via the bus 110. And it is used to display various information received, stored, and processed from the above components (such as multimedia data, text data), and can display videos, images, data, etc. to the user.

[0070] The communication interface 170 can be used for signaling or data communication with other node devices.

[0071] It can be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may also include more or fewer components than those shown Figure 1 in the figure, or have a different configuration from that shown Figure 1 in the figure. Figure 1 Each component shown in the figure can be implemented by hardware, software, or a combination thereof.

[0072] The electronic device provided by the embodiments of the present invention may be a smart phone, a personal computer, a tablet computer, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of the present invention do not make any restrictions on this.

[0073] Next, the above-mentioned electronic device will be used as the execution subject to execute each step in each method provided by the embodiments of the present invention and achieve the corresponding technical effects.

[0074] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a target tracking method provided by an embodiment of the present invention.

[0075] Step S202: Obtain the current video frame and the template parameters of the current video frame, where the template parameters of the current video frame are calculated based on the position and scale of the target in the previous video frame;

[0076] In this embodiment, the target can be a pedestrian, a vehicle or other objects, and the target can be tracked by collecting multiple video frames. When the first video frame is obtained, a yolox detection model trained based on pytorch can be used to detect the target in the first video frame, detect the target and obtain the position and scale of the target in the first video frame, and then define an identifier, i.e., ID, for the target, so as to track the target in each subsequent video frame.

[0077] For each video frame except the first video frame, a correlation filtering tracking algorithm can be used to track the target. The correlation filtering tracking algorithm is to design a filtering template and perform a correlation operation between the filtering template and the area where the target may be in the current video frame to obtain the area of the target in the current video frame, so as to realize the tracking of the target. The template parameters in the embodiment of the present invention can be understood as the filtering template, and each video frame has template parameters.

[0078] It can be understood that for each video frame except the first video frame, the method of tracking the target is similar. Therefore, the embodiment of the present invention takes the current video frame being processed, i.e., the current video frame, as an example for description. When obtaining the current video frame, it is also necessary to obtain the template parameters of the current video frame, and the template parameters of the current video frame are calculated based on the position and scale of the target in the previous video frame of the current video frame, i.e., the previous video frame.

[0079] Step S204: Select multiple scale adjustment coefficient groups from a preset set of scale adjustment coefficients;

[0080] Step S206: According to each scale adjustment coefficient group, and the position and scale of the target in the previous video frame, obtain the candidate region corresponding to each scale adjustment coefficient group in the current video frame;

[0081] In this embodiment, since the scale of the target may change, multiple scale adjustment coefficient groups can be selected from a pre-set set of scale adjustment coefficients; then, according to the position and scale of the target in the previous video frame, an area where the target may be located in the current video frame, i.e., a pending area, is marked, and then each scale adjustment coefficient group is used to adjust this pending area, such as performing a scaling process, and each adjusted pending area obtained is a candidate area of the target in the current video frame. One scale adjustment coefficient group corresponds to one candidate area.

[0082] Step S208: Extract features from each candidate area to obtain each target feature matrix, and calculate a response value based on each target feature matrix and the template parameters of the current video frame, to obtain the response value corresponding to each scale adjustment coefficient group; the response value represents the similarity between the candidate area and the area of the target in the previous video frame.

[0083] In this implementation, the method of calculating the response value for each candidate area is similar. For the sake of easy understanding, an embodiment of the present invention takes any one candidate area as an example for illustration.

[0084] Extract features from the candidate area, such as extracting Histogram of Oriented Gradients (HOG) features, Color Names (CN) features, grayscale features, etc., then perform dimensionality reduction on these features using Principal Component Analysis (PCA), and then use a Gaussian kernel function for multi-channel fusion to obtain the target feature map, i.e., the target feature matrix, of this candidate area.

[0085] Based on the obtained target feature matrix, it is converted to the Fourier domain and multiplied point by point with the template parameters of the current video frame to obtain the target response matrix, that is, the target feature matrix and the template parameters of the current video frame are respectively Fourier-transformed and then multiplied point by point. It can be expressed as Where represents the matrix after Fourier-transforming the target feature matrix; represents the parameters after Fourier-transforming the template parameters of the current video frame; represents the target response matrix.

[0086] Performing an inverse Fourier transform on the target response matrix can obtain the response matrix in the time domain, and taking the maximum value in this response matrix as the response value. It can be expressed as Where H represents the response value; represents the inverse Fourier transform; represents the target response matrix.

[0087] Based on the candidate area corresponding to each scale adjustment coefficient group, the response value corresponding to each scale adjustment coefficient group can be obtained, and this response value can be understood as the similarity between the candidate area and the area of the target in the previous video frame.

[0088] Step S210: Take the scale adjustment coefficient group corresponding to the maximum response value as the optimal scale adjustment coefficient group of the current video frame, and obtain the position and scale of the target in the current video frame according to the optimal scale adjustment coefficient group, so as to track the target;

[0089] In this embodiment, select the maximum response value from all the response values, and take the scale adjustment coefficient group corresponding to the maximum response value as the optimal scale adjustment coefficient group of the current video frame, which means that the target features in the candidate region corresponding to this optimal scale adjustment coefficient group are the most similar to the target features in the previous video frame. Then, take the candidate region corresponding to this optimal scale adjustment coefficient group as the region where the target is located in the current video frame, that is, obtain the position and scale of the target in the current video frame, so as to realize the tracking of the target.

[0090] It can be seen that based on the above steps, by obtaining the current video frame and the template parameters of the current video frame, the template parameters of the current video frame are calculated according to the position and scale of the target in the previous video frame; then select multiple scale adjustment coefficient groups from the preset scale adjustment coefficient set; and according to each scale adjustment coefficient group, as well as the position and scale of the target in the previous video frame, obtain the candidate region corresponding to each scale adjustment coefficient group in the current video frame; then extract features from each candidate region to obtain each target feature matrix, and calculate the response value according to each target feature matrix and the template parameters of the current video frame, to obtain the response value corresponding to each scale adjustment coefficient group, and the response value represents the similarity between the candidate region and the region of the target in the previous video frame; finally, take the scale adjustment coefficient group corresponding to the maximum response value as the optimal scale adjustment coefficient group of the current video frame, and obtain the position and scale of the target in the current video frame according to the optimal scale adjustment coefficient group, so as to track the target. The target with adaptive tracking scale change is realized through multiple scale adjustment coefficient groups, and the correlation of the target in adjacent video frames is established based on the template parameters to perform target tracking, thereby improving the stability and accuracy of target tracking.

[0091] Optionally, for the above step S204, an embodiment of the present invention provides a possible implementation manner. Please refer to Figure 3 .

[0092] Step S204-1: Obtain the optimal scale adjustment coefficient group of the previous video frame to get the target width adjustment coefficient and the target height adjustment coefficient;

[0093] Step S204-3: Select multiple candidate width adjustment coefficients from the width coefficient sequence according to the first preset rule and the target width adjustment coefficient;

[0094] Step S204-5: Select multiple candidate height adjustment coefficients from the height coefficient sequence according to the second preset rule and the target height adjustment coefficient;

[0095] Step S204-7: Combine each candidate width adjustment coefficient with each candidate height adjustment coefficient to obtain each scale adjustment coefficient group. A scale adjustment coefficient group includes a candidate width adjustment coefficient and a candidate height adjustment coefficient.

[0096] In this embodiment, the scale adjustment coefficient set includes a width coefficient sequence and a height coefficient sequence. The width coefficient sequence includes a set number of width adjustment coefficients, and the height coefficient sequence includes a set number of height adjustment coefficients, that is, the total number of width adjustment coefficients is the same as the total number of height adjustment coefficients.

[0097] It can be understood that when the target is in multiple consecutive video frames, the change in its scale often has a certain pattern. Therefore, the best scale adjustment coefficient group of the previous video frame can be obtained as a reference to select the scale adjustment coefficient group for processing the current video frame.

[0098] Multiple candidate width adjustment coefficients can be selected from all width adjustment coefficients according to a pre-set first rule and the best width adjustment coefficient of the previous video frame, i.e., the target width adjustment coefficient; and multiple candidate height adjustment coefficients can be selected from all height adjustment coefficients according to a pre-set second rule and the best height adjustment coefficient of the previous video frame, i.e., the target height adjustment coefficient; then each candidate width adjustment coefficient is combined with each candidate height adjustment coefficient to obtain each scale adjustment coefficient group.

[0099] For example, the width coefficient sequence includes 5 width adjustment coefficients, and the height coefficient sequence includes 5 height adjustment coefficients. According to the first preset rule and the target width adjustment coefficient, 3 candidate width adjustment coefficients are selected from the width coefficient sequence; according to the second preset rule and the target height adjustment coefficient, 3 candidate height adjustment coefficients are selected from the height coefficient sequence; then each candidate width adjustment coefficient is combined with each candidate height adjustment coefficient to obtain 9 scale adjustment coefficient groups.

[0100] Optionally, for the above step S204-3, an embodiment of the present invention provides a possible implementation manner.

[0101] Step S204-3A: If the target width adjustment coefficient is less than the median of all width adjustment coefficients, control the first window to slide forward by one step from the current position, and obtain each width adjustment coefficient in the slid first window to obtain each candidate width adjustment coefficient.

[0102] Step S204-3B: If the target width adjustment coefficient is equal to the median of all width adjustment coefficients, obtain each width adjustment coefficient in the first window to obtain each candidate width adjustment coefficient.

[0103] Step S204-3C: if the target width adjustment coefficient is greater than the median of all width adjustment coefficients, the first window is controlled to slide backward from the current position by one step, and each width adjustment coefficient in the first window after sliding is obtained to obtain each candidate width adjustment coefficient.

[0104] In this embodiment, the width coefficient sequence has a corresponding first window, and the first window is used to slide in the width coefficient sequence according to a first preset rule to select a candidate width adjustment coefficient. The total number of all width adjustment coefficients in the width coefficient sequence is an odd number, and all width adjustment coefficients are arranged in order from small to large.

[0105] It can be understood that the median of all width adjustment coefficients can be set to 1. If a width adjustment coefficient with a value equal to the median, i.e. 1, is used to adjust the pending area, it means that the width of the pending area will not be changed; if a width adjustment coefficient with a value less than the median, i.e. less than 1, is used to adjust the pending area, it means that the width of the pending area will be reduced; if a width adjustment coefficient with a value greater than the median, i.e. greater than 1, is used to adjust the pending area, it means that the width of the pending area will be expanded.

[0106] The change trend of the undetermined area is the change trend of the scale of the target. Therefore, the embodiment of the present invention selects the candidate width adjustment coefficient of the current video frame by taking the best width adjustment coefficient of the previous frame, that is, the target width adjustment coefficient, as a reference.

[0107] For ease of understanding, an example diagram is provided in the embodiment of the present invention. Figure 4 (a), wherein the width coefficient sequence includes 5 horizontally arranged width adjustment coefficients, namely {x1, x2, x3, x4, x5}, and the values ​​of x1 to x5 gradually increase, such as {0.99, 0.995, 1, 1.005, 1.01}. The width coefficient sequence has a corresponding first window and the window length is 3.

[0108] It should be understood that the width adjustment coefficient in the width coefficient sequence and the window length of the first window can be set according to actual applications, and the embodiment of the present invention does not limit this.

[0109] If the target width adjustment coefficient is less than the median of all width adjustment coefficients, such as the target width adjustment coefficient is x1 or x2, it means that the width of the target is decreasing in the previous video frame, and the width of the target is likely to continue to decrease in the current video frame. Then, the first window is controlled to slide forward one step from the current position, and each width adjustment coefficient in the first window after sliding is obtained to obtain each candidate width adjustment coefficient, namely x1, x2 and x3.

[0110] If the target width adjustment coefficient is greater than the median of all width adjustment coefficients, such as the target width adjustment coefficient is x4 or x5, it means that the width of the target is increasing in the previous video frame, and the width of the target is likely to continue to increase in the current video frame. The first window is controlled to slide backward from the current position by one step, and each width adjustment coefficient in the first window after sliding is obtained to obtain each candidate width adjustment coefficient, namely x3, x4 and x5.

[0111] If the target width adjustment coefficient is equal to the median of all width adjustment coefficients, such as the target width adjustment coefficient is x3, it means that the width of the target has not changed in the previous video frame, and the scale of the target in the current video frame may become larger or smaller, or may remain unchanged. Then, each width adjustment coefficient of the first window when the first window is at the current position is obtained to obtain each candidate width adjustment coefficient, namely x2, x3 and x4.

[0112] Optionally, when the target width adjustment coefficient is x3, the first window may also be located at the leftmost or rightmost end of the width coefficient sequence. In this case, the first window is controlled to be located in the middle of the width coefficient sequence, and each width adjustment coefficient in the first window is obtained to obtain each candidate width adjustment coefficient, namely x2, x3 and x4. It can be understood that the width adjustment coefficient adjacent to the median of all width adjustment coefficients is used as the candidate width adjustment coefficient.

[0113] It is understandable that the method of selecting the candidate height adjustment coefficient in step S204-5 is similar to the method of selecting the candidate width adjustment coefficient in step S204-3. The height coefficient sequence has a corresponding second window, and the second window is used to slide the height coefficient sequence according to the second preset rule to select the candidate height adjustment coefficient. The total number of all height adjustment coefficients in the height coefficient sequence is an odd number, and all height adjustment coefficients are arranged in order from small to large.

[0114] It can be understood that the median of all height adjustment coefficients can be set to 1. If a height adjustment coefficient with a value equal to the median, i.e. 1, is used to adjust the pending area, it means that the height of the pending area will not be changed; if a height adjustment coefficient with a value less than the median, i.e. less than 1, is used to adjust the pending area, it means that the height of the pending area will be reduced; if a height adjustment coefficient with a value greater than the median, i.e. greater than 1, is used to adjust the pending area, it means that the height of the pending area will be expanded.

[0115] The change trend of the undetermined area is the change trend of the scale of the target. Therefore, the embodiment of the present invention selects the candidate height adjustment coefficient of the current video frame by taking the best height adjustment coefficient of the previous frame, that is, the target height adjustment coefficient, as a reference.

[0116] For ease of understanding, an example diagram is provided in the embodiment of the present invention.Figure 4 (b), where the height coefficient sequence includes 5 vertically arranged height adjustment coefficients, namely {y1, y2, y3, y4, y5}, and the values from y1 to y5 gradually increase, such as {0.99, 0.995, 1, 1.005, 1.01}. The height coefficient sequence has a corresponding second window with a window length of 3.

[0117] It should be understood that the height adjustment coefficients in the height coefficient sequence and the window length of the second window can be set according to actual applications, and the embodiments of the present invention do not make limitations.

[0118] If the target height adjustment coefficient is less than the median of all height adjustment coefficients, such as the target height adjustment coefficient is y1 or y2, it means that the height of the target in the previous video frame is decreasing. Then, in the current video frame, the height of the target may continue to decrease. Then, control the second window to slide upward by one step from the current position, and obtain each height adjustment coefficient in the second window after sliding, to obtain each candidate height adjustment coefficient, namely y1, y2, and y3.

[0119] If the target height adjustment coefficient is greater than the median of all height adjustment coefficients, such as the target height adjustment coefficient is y4 or y5, it means that the height of the target in the previous video frame is increasing. Then, in the current video frame, the height of the target may continue to increase. Then, control the second window to slide downward by one step from the current position, and obtain each height adjustment coefficient in the second window after sliding, to obtain each candidate height adjustment coefficient, namely y3, y4, and y5.

[0120] If the target height adjustment coefficient is equal to the median of all height adjustment coefficients, such as the target height adjustment coefficient is y3, it means that the width of the target in the previous video frame has not changed. Then, in the current video frame, the scale of the target may increase, decrease, or remain unchanged. Then, obtain each height adjustment coefficient in the second window when the second window is in the current position, to obtain each candidate height adjustment coefficient, namely y2, y3, and y4.

[0121] Optionally, when the target height adjustment coefficient is y3, the second window may also be located at the uppermost or lowermost end of the height coefficient sequence. In this case, control the second window to be located in the middle position of the height coefficient sequence, and obtain each height adjustment coefficient in the second window, to obtain each candidate height adjustment coefficient, namely y2, y3, and y4. It can be understood that the height adjustment coefficients adjacent to the median of all height adjustment coefficients are used as candidate height adjustment coefficients.

[0122] It can be seen that in the embodiments of the present invention, by setting the scale adjustment coefficients, namely the width coefficient sequence and the height coefficient sequence, and through different proportional changes and two sliding windows, the target with changing scale is adaptively tracked, thereby avoiding the loss of the target caused by introducing too much background information and improving the accuracy and stability of target tracking.

[0123] Optionally, for the template parameters of the current video frame in the above embodiments, the embodiments of the present invention provide a possible implementation manner. Please refer to Figure 5 .

[0124] Step S222: Calculate the candidate template parameters of the previous video frame according to the position and scale of the target in the previous video frame;

[0125] Step S224: Calculate the template parameters of the current video frame according to the candidate template parameters and the template parameters of the previous video frame.

[0126] In this embodiment, the candidate template parameters of the previous video frame can be calculated by a correlation filtering tracking algorithm according to the position and scale of the target in the previous video frame. For example, determine the image of the region where the target is located based on the position and scale of the target in the previous video frame, that is, obtain the target image, and use the target response matrix corresponding to the region where the target is located as the target response matrix of the target image; then calculate the function value according to a preset function, the target image, and the target response matrix of the target image. For example, use x to represent the image matrix of the target image, use y to represent the target response matrix of the target image, and the preset function is: Wherein, represents the function value; represents the response matrix after performing a Fourier transform on the target response matrix of the target image; represents the autocorrelation of x itself in the Fourier domain; λ represents the regularization parameter to prevent overfitting.

[0127] Then, according to the obtained function value, perform an inverse Fourier transform on the function value to obtain the candidate template parameters of the previous video frame; then obtain the template parameters of the previous video frame, and calculate the template parameters of the current video frame based on the candidate parameters and the template parameters of the previous video frame.

[0128] The template parameters of the previous video frame can be understood as the template parameters actually used to process the previous video frame; the candidate template parameters of the previous video frame can be understood as the theoretical template parameters calculated according to the position and scale of the target in the previous video frame for processing the current video frame.

[0129] Optionally, for the above step S224, an embodiment of the present invention provides a possible implementation manner, that is: step S224A, calculate the template parameters of the current video frame according to a preset formula, a preset learning rate, the candidate template parameters of the previous video frame, and the template parameters; the preset formula is: α i = θα' i-1 + (1 - θ)α i-1 ; where α i represents the template parameters of the current video frame; α' i-1 represents the candidate template parameters of the previous video frame; α i-1 represents the template parameters of the previous video frame; θ represents the preset learning rate.

[0130] In this embodiment, based on the preset learning rate θ, the candidate template parameters α' i-1 of the previous video frame, and the template parameters α i-1 of the previous video frame, the template parameters of the current video frame can be calculated according to the preset formula.

[0131] It can be understood that since the target has correlation between adjacent video frames, when calculating the template parameters of the current video frame, the embodiment of the present invention not only uses the template parameters actually used to process the previous video frame, but also uses the theoretical template parameters for processing the current video frame, so as to make the obtained scale of the target more accurate.

[0132] Optionally, for the above step S224, the embodiment of the present invention provides another possible implementation manner.

[0133] Step S224B-1, obtain the best scale adjustment coefficient group of the previous video frame to obtain the target width adjustment coefficient and the target height adjustment coefficient;

[0134] Step S224B-3, select the target learning rate from the preset learning rate set according to the third preset rule, the target width adjustment coefficient, and the target height adjustment coefficient;

[0135] Step S224B-5, calculate the template parameters of the current video frame according to the preset formula, the target learning rate, the candidate template parameters of the previous video frame, and the template parameters;

[0136] The preset formula is: α i = θα' i-1 + (1 - θ)α i-1 ; where α i represents the template parameters of the current video frame; α' i-1 represents the candidate template parameters of the previous video frame; α i-1 represents the template parameters of the previous video frame; θ represents the target learning rate.

[0137] It can be understood that since the optimal scale adjustment coefficient group of each video frame can reflect the scale change trend of the target, the embodiments of the present invention can use the optimal scale adjustment coefficient group of the previous video frame as a reference to select the learning rate for calculating the template parameters of the current video frame.

[0138] The target learning rate, i.e., θ, can be selected from a preset learning rate set according to a preset third rule and the optimal scale adjustment coefficient group of the previous video frame, i.e., the target width adjustment coefficient and the target height adjustment coefficient, and based on the target learning rate, i.e., θ, the candidate template parameters of the previous video frame, i.e., α′ i-1 and the template parameters of the previous video frame, i.e., α i-1 , calculate the template parameters of the current video frame according to a preset formula.

[0139] Optionally, for the above step S224B-3, the embodiments of the present invention provide a possible implementation.

[0140] Step S224B-3a, if the target width adjustment coefficient is equal to the median of all width adjustment coefficients and the target height adjustment coefficient is equal to the median of all height adjustment coefficients, then select the third learning rate as the target learning rate;

[0141] Step S224B-3b, if the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is equal to the target height adjustment coefficient, then select the first learning rate as the target learning rate;

[0142] Step S224B-3c, if the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is not equal to the target height adjustment coefficient, then select the second learning rate as the target learning rate.

[0143] For ease of understanding, the embodiments of the present invention use the width coefficient sequence and height coefficient sequence above Figure 4 as examples for illustration. The preset learning rate set includes the first learning rate θ 1 , the second learning rate θ 2 and the third learning rate θ 3 , and θ 2 is greater than θ 1 , θ 1 is greater than θ 3 . For example, θ 1 is 0.01, θ 2 is 0.15, θ 3is 0. It should be understood that the first learning rate, the second learning rate, and the third learning rate can be set according to actual applications, and the embodiments of the present invention do not make limitations.

[0144] If the target width adjustment coefficient is equal to the median of all width adjustment coefficients and the target height adjustment coefficient is equal to the median of all height adjustment coefficients. For example, if the target width adjustment coefficient is x3 and the target height adjustment coefficient is y3, it means that the scale of the target did not change in the previous video frame, indicating that the scale of the target may remain unchanged in the current video frame. Then, the weight of the theoretical template parameters used to process the current video frame is adjusted to the minimum value, so that the weight of the template parameters actually used to process the previous video frame is the maximum. That is, the third learning rate θ 3 is used as the target learning rate.

[0145] If the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is equal to the target height adjustment coefficient. For example, if the target width adjustment coefficient and the target height adjustment coefficient are x1 and y1, or x2 and y2, or x4 and y4, or x5 and y5 respectively, it means that the width and height of the target increased or decreased in the same proportion in the previous video frame, indicating that the scale change of the target in the previous video frame has a certain pattern. Then, the weight of the theoretical template parameters used to process the current video frame is adjusted to the middle value, that is, the first learning rate θ 1 is used as the target learning rate.

[0146] If the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is not equal to the target height adjustment coefficient. For example, if the target width adjustment coefficient and the target height adjustment coefficient are x1 and y5, or x3 and y2, etc., it means that the width of the target may increase or decrease, or may remain unchanged in the previous video frame, and the height of the target may increase or decrease, or may remain unchanged. It indicates that the scale change of the target in the previous video frame has no pattern. Then, the weight of the theoretical template parameters used to process the current video frame is adjusted to the maximum value, that is, the second learning rate θ 2 is used as the target learning rate.

[0147] It can be seen that in the embodiments of the present invention, by setting the scale adjustment coefficients, namely the width coefficient sequence and the height coefficient sequence, and the best scale adjustment coefficient of the previous frame to judge the scale change trend and change amplitude of the target, so as to update the template parameters, the robustness of the target under partial occlusion can be improved, so that the obtained scale of the target is more accurate, thereby improving the accuracy and stability of target tracking.

[0148] To perform the corresponding steps in the above embodiments and each possible manner, an implementation of an object tracking device is given below. Please refer to Figure 6 , Figure 6 FIG. Figure 6 is a functional module diagram of an object tracking device 300 provided by an embodiment of the present invention. It should be noted that the basic principle and the technical effects generated by the object tracking device 300 provided in this embodiment are the same as those in the above embodiments. For a brief description, for the parts not mentioned in this embodiment, reference may be made to the corresponding content in the above embodiments. The object tracking device 300 includes:

[0149] An acquisition module 310, configured to acquire a current video frame and template parameters of the current video frame, where the template parameters of the current video frame are calculated according to the position and scale of the object in the previous video frame;

[0150] A processing module 330, configured to select multiple scale adjustment coefficient groups from a preset set of scale adjustment coefficients;

[0151] According to each scale adjustment coefficient group, and the position and scale of the object in the previous video frame, obtain a candidate region corresponding to each scale adjustment coefficient group in the current video frame;

[0152] Extract features from each candidate region to obtain each object feature matrix, and calculate a response value according to each object feature matrix and the template parameters of the current video frame, to obtain a response value corresponding to each scale adjustment coefficient group; the response value represents the similarity between the candidate region and the region of the object in the previous video frame;

[0153] A tracking module 350, configured to use the scale adjustment coefficient group corresponding to the maximum response value as the best scale adjustment coefficient group of the current video frame, and obtain the position and scale of the object in the current video frame according to the best scale adjustment coefficient group, so as to track the object.

[0154] Optionally, the processing module 330 is further configured to: obtain the best scale adjustment coefficient group of the previous video frame, to obtain an object width adjustment coefficient and an object height adjustment coefficient; select multiple candidate width adjustment coefficients from a width coefficient sequence according to a first preset rule and the object width adjustment coefficient; select multiple candidate height adjustment coefficients from a height coefficient sequence according to a second preset rule and the object height adjustment coefficient; combine each candidate width adjustment coefficient with each candidate height adjustment coefficient, to obtain each scale adjustment coefficient group, and a scale adjustment coefficient group includes a candidate width adjustment coefficient and a candidate height adjustment coefficient.

[0155] Optionally, the processing module 330 is further configured to: if the target width adjustment coefficient is less than the median of all width adjustment coefficients, control the first window to slide forward by one step from the current position, and obtain each width adjustment coefficient in the first window after sliding, so as to obtain each candidate width adjustment coefficient;

[0156] if the target width adjustment coefficient is equal to the median of all width adjustment coefficients, obtain each width adjustment coefficient in the first window, so as to obtain each candidate width adjustment coefficient;

[0157] if the target width adjustment coefficient is greater than the median of all width adjustment coefficients, control the first window to slide backward by one step from the current position, and obtain each width adjustment coefficient in the first window after sliding, so as to obtain each candidate width adjustment coefficient.

[0158] Optionally, the target tracking device 300 further includes a calculation module 370, configured to: calculate candidate template parameters of the previous video frame according to the position and scale of the target in the previous video frame; calculate template parameters of the current video frame according to the candidate template parameters and template parameters of the previous video frame.

[0159] Optionally, the calculation module 370 is further configured to: calculate template parameters of the current video frame according to a preset formula, a preset learning rate, the candidate template parameters and template parameters of the previous video frame;

[0160] The preset formula is:

[0161] α i = θα′ i-1 +(1 - θ)α i-1 ;

[0162] where, α i represents the template parameters of the current video frame; α′ i-1 represents the candidate template parameters of the previous video frame; α i-1 represents the template parameters of the previous video frame; θ represents the preset learning rate.

[0163] Optionally, the calculation module 370 is further configured to: obtain the best scale adjustment coefficient group of the previous video frame, so as to obtain the target width adjustment coefficient and the target height adjustment coefficient;

[0164] select a target learning rate from a preset learning rate set according to a third preset rule, the target width adjustment coefficient and the target height adjustment coefficient;

[0165] calculate template parameters of the current video frame according to the preset formula, the target learning rate, the candidate template parameters and template parameters of the previous video frame; the preset formula is:

[0166] α i = θα′ i-1+(1 - θ)α i-1 ;

[0167] where α i represents the template parameter of the current video frame; α' i-1 represents the candidate template parameter of the previous video frame; α i-1 represents the template parameter of the previous video frame; θ represents the target learning rate.

[0168] Optionally, the calculation module 370 is further configured to: if the target width adjustment coefficient is equal to the median of all width adjustment coefficients and the target height adjustment coefficient is equal to the median of all height adjustment coefficients, then select the third learning rate as the target learning rate;

[0169] if the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is equal to the target height adjustment coefficient, then select the first learning rate as the target learning rate;

[0170] if the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is not equal to the target height adjustment coefficient, then select the second learning rate as the target learning rate.

[0171] An embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores a computer program. When the processor executes the computer program, the target tracking method disclosed in the embodiment of the present invention is implemented.

[0172] An embodiment of the present invention further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the target tracking method disclosed in the embodiment of the present invention is implemented.

[0173] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0174] In addition, each functional module in various embodiments of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0175] If the above-mentioned functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0176] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A target tracking method, characterized in that, the method includes: obtaining a current video frame and template parameters of the current video frame, where the template parameters of the current video frame are calculated according to the position and scale of the target in the previous video frame; selecting a plurality of scale adjustment coefficient groups from a preset set of scale adjustment coefficients; obtaining, in the current video frame, a candidate region corresponding to each scale adjustment coefficient group according to each scale adjustment coefficient group, and the position and scale of the target in the previous video frame; performing feature extraction on each candidate region to obtain each target feature matrix, and calculating a response value according to each target feature matrix and the template parameters of the current video frame, to obtain a response value corresponding to each scale adjustment coefficient group; the response value represents the similarity between the candidate region and the region of the target in the previous video frame; taking the scale adjustment coefficient group corresponding to the maximum response value as the best scale adjustment coefficient group of the current video frame, and obtaining the position and scale of the target in the current video frame according to the best scale adjustment coefficient group, so as to track the target; the set of scale adjustment coefficients includes a width coefficient sequence and a height coefficient sequence, the width coefficient sequence includes a set number of width adjustment coefficients, and the height coefficient sequence includes a set number of height adjustment coefficients; the step of selecting a plurality of scale adjustment coefficient groups from a preset set of scale adjustment coefficients includes: obtaining the best scale adjustment coefficient group of the previous video frame to obtain a target width adjustment coefficient and a target height adjustment coefficient; selecting a plurality of candidate width adjustment coefficients from the width coefficient sequence according to a first preset rule and the target width adjustment coefficient; wherein, the first preset rule represents a rule for selecting candidate width adjustment coefficients according to the size relationship between the target width adjustment coefficient and the median of all width adjustment coefficients; selecting a plurality of candidate height adjustment coefficients from the height coefficient sequence according to a second preset rule and the target height adjustment coefficient; wherein, the second preset rule represents a rule for selecting candidate height adjustment coefficients according to the size relationship between the target height adjustment coefficient and the median of all height adjustment coefficients; combining each candidate width adjustment coefficient with each candidate height adjustment coefficient to obtain each scale adjustment coefficient group, and one scale adjustment coefficient group includes one candidate width adjustment coefficient and one candidate height adjustment coefficient.

2. The method according to claim 1, characterized in that, the set number is odd, all width adjustment coefficients in the width coefficient sequence are arranged in ascending order, and the width coefficient sequence has a corresponding first window; the step of selecting a plurality of candidate width adjustment coefficients from the width coefficient sequence according to a first preset rule and the target width adjustment coefficient includes: If the target width adjustment coefficient is less than the median of all width adjustment coefficients, control the first window to slide forward by one step from the current position, and obtain each width adjustment coefficient in the first window after sliding to obtain each of the candidate width adjustment coefficients; If the target width adjustment coefficient is equal to the median of all width adjustment coefficients, obtain each width adjustment coefficient in the first window to obtain each of the candidate width adjustment coefficients; If the target width adjustment coefficient is greater than the median of all width adjustment coefficients, control the first window to slide backward by one step from the current position, and obtain each width adjustment coefficient in the first window after sliding to obtain each of the candidate width adjustment coefficients.

3. The method according to claim 1, wherein, the template parameter of the current video frame is calculated in the following manner: Calculate the candidate template parameter of the previous video frame according to the position and scale of the target in the previous video frame; Calculate the template parameter of the current video frame according to the candidate template parameter and the template parameter of the previous video frame.

4. The method according to claim 3, wherein, the step of calculating the template parameter of the current video frame according to the candidate template parameter and the template parameter of the previous video frame includes: Calculate the template parameter of the current video frame according to a preset formula, a preset learning rate, the candidate template parameter and the template parameter of the previous video frame; The preset formula is: α i = θα' i-1 + (1 - θ)α i-1 ; Among them, α i represents the template parameter of the current video frame; α′ i-1 represents the candidate template parameter of the previous video frame; α i-1 represents the template parameter of the previous video frame; θ represents the preset learning rate.

5. The method according to claim 3, wherein, the step of calculating the template parameter of the current video frame according to the candidate template parameter and the template parameter of the previous video frame includes: Obtain the optimal scale adjustment coefficient group of the previous video frame to obtain the target width adjustment coefficient and the target height adjustment coefficient; Select a target learning rate from a preset set of learning rates according to a third preset rule, the target width adjustment coefficient and the target height adjustment coefficient; Calculate the template parameter of the current video frame according to a preset formula, the target learning rate, the candidate template parameter and the template parameter of the previous video frame; the preset formula is: α i = θα′ i-1 + (1 - θ)α i-1 ; Among them, α i represents the template parameter of the current video frame; α' i-1 represents the candidate template parameter of the previous video frame; α i-1 represents the template parameter of the previous video frame; θ represents the target learning rate.

6. The method according to claim 5, wherein, the set of scale adjustment coefficients includes a width coefficient sequence and a height coefficient sequence, the width coefficient sequence includes a set number of width adjustment coefficients, and the height coefficient sequence includes a set number of height adjustment coefficients; each width adjustment coefficient in the width coefficient sequence is equal to the height adjustment coefficient in the corresponding order in the height coefficient sequence, and the set number is odd; the set of learning rates includes a first learning rate, a second learning rate and a third learning rate, the second learning rate is greater than the first learning rate, and the first learning rate is greater than the third learning rate; the step of selecting a target learning rate from a preset set of learning rates according to a third preset rule, the target width adjustment coefficient and the target height adjustment coefficient includes: If the target width adjustment coefficient is equal to the median of all width adjustment coefficients and the target height adjustment coefficient is equal to the median of all height adjustment coefficients, then select the third learning rate as the target learning rate; If the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is equal to the target height adjustment coefficient, then select the first learning rate as the target learning rate; If the target width adjustment coefficient is not equal to the median of all width adjustment coefficients, the target height adjustment coefficient is not equal to the median of all height adjustment coefficients, and the target width adjustment coefficient is not equal to the target height adjustment coefficient, then select the second learning rate as the target learning rate.

7. A target tracking device characterized in that the device includes: an acquisition module, configured to acquire a current video frame and template parameters of the current video frame, where the template parameters of the current video frame are calculated according to the position and scale of the target in the previous video frame; a processing module, configured to select multiple scale adjustment coefficient groups from a preset set of scale adjustment coefficients; acquire a candidate region corresponding to each scale adjustment coefficient group in the current video frame according to each scale adjustment coefficient group, and the position and scale of the target in the previous video frame; extract features from each candidate region to obtain each target feature matrix, and calculate a response value according to each target feature matrix and the template parameters of the current video frame, to obtain a response value corresponding to each scale adjustment coefficient group; the response value represents the similarity between the candidate region and the region of the target in the previous video frame; a tracking module, configured to use the scale adjustment coefficient group corresponding to the maximum response value as the best scale adjustment coefficient group of the current video frame, and obtain the position and scale of the target in the current video frame according to the best scale adjustment coefficient group, so as to track the target; The scale adjustment coefficient set includes a width coefficient sequence and a height coefficient sequence. The width coefficient sequence includes a set number of width adjustment coefficients, and the height coefficient sequence includes a set number of height adjustment coefficients. The processing module is further configured to: obtain the optimal scale adjustment coefficient group of the previous video frame to obtain a target width adjustment coefficient and a target height adjustment coefficient; select a plurality of candidate width adjustment coefficients from the width coefficient sequence according to a first preset rule and the target width adjustment coefficient, where the first preset rule represents a rule for selecting candidate width adjustment coefficients according to the magnitude relationship between the target width adjustment coefficient and the median of all width adjustment coefficients; select a plurality of candidate height adjustment coefficients from the height coefficient sequence according to a second preset rule and the target height adjustment coefficient, where the second preset rule represents a rule for selecting candidate height adjustment coefficients according to the magnitude relationship between the target height adjustment coefficient and the median of all height adjustment coefficients; combine each of the candidate width adjustment coefficients with each of the candidate height adjustment coefficients to obtain each scale adjustment coefficient group, and one scale adjustment coefficient group includes one of the candidate width adjustment coefficients and one of the candidate height adjustment coefficients.

8. An electronic device, characterized in that it includes a processor and a memory. The memory stores a computer program. When the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.

9. A storage medium, characterized in that a computer program is stored on the storage medium. When the computer program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Adaptive scale video target tracking method

    CN107481264A

  • Unmanned aerial vehicle adaptive target tracking method based on pseudo twin network

    CN113516713A