Object Tracking Method Based on Matching Strategy of Task-Oriented Object Tracking Algorithm
Through the task-oriented target tracking algorithm matching strategy, using image attribute information to select appropriate algorithms, the problem of poor performance of existing multimodal target tracking in complex environments is solved, and the target tracking effect with efficient and low resource consumption is achieved.
Patent Information
- Application Number
- CN202510237611.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing multimodal target tracking algorithms perform poorly in complex environments, have high demand for computing resources, and lack algorithm matching strategies for different tracking scenarios.
A task-oriented target tracking algorithm matching strategy is proposed. By acquiring image sequences, detecting and generating image attribute information, selecting an appropriate target tracking algorithm based on the differences in attribute information, reducing the calculation amount and improving the calculation efficiency.
It realizes efficient target tracking in complex environments, reduces computing resource consumption, improves tracking performance and efficiency, and quickly schedules algorithms according to different scenarios.
Smart Images

Figure CN119741333B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an object tracking method based on a matching strategy of a task-oriented object tracking algorithm. Background Art
[0002] Single-modal object tracking usually relies on a single data source, such as visual images, which is simple to implement and has high computational efficiency, and is suitable for scenarios with relatively stable environmental conditions. However, it may perform poorly under conditions such as changing environmental illumination, occlusion, and complex backgrounds. In contrast, multi-modal object tracking can significantly improve the accuracy and robustness of tracking by integrating different types of data, especially performing better in complex or dynamic environments. However, the implementation complexity is high, the computational resource requirements are relatively large, and while improving the tracking accuracy, it will increase the computational burden during object tracking, resulting in lower tracking performance and efficiency.
[0003] In addition, there are numerous existing multi-modal object tracking algorithms, each with its own advantages and disadvantages. Different algorithms have different advantageous scenarios. In practical applications, there is a lack of an algorithm matching strategy for different tracking scenarios and task requirements. Summary of the Invention
[0004] In view of the above problems, the present invention provides an object tracking method based on a matching strategy of a task-oriented object tracking algorithm.
[0005] According to a first aspect of the present invention, there is provided an object tracking method based on a matching strategy of a task-oriented object tracking algorithm, including: obtaining an image sequence of a target object, the image sequence including I images, I being an integer greater than 1, the image sequence including visible light images and infrared images; respectively performing image detection on each image to generate attribute information of each image; in response to a difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image being greater than or equal to a predetermined threshold, inputting the attribute information of the i-th image into a target model to output information of a first target algorithm for processing the i-th image; wherein the target model is obtained by training an initial model based on the attribute information of sample images and algorithm information matching the attribute information of the sample images, 1 < i ≤ I; and respectively processing each image based on the respective target algorithms matching the attribute information of each image to obtain a tracking result of the target object.
[0006] A second aspect of the present invention provides an object tracking device based on a matching strategy of a task-oriented object tracking algorithm, including: an obtaining module, a detection module, an output module, and a processing module.
[0007] An acquisition module for acquiring an image sequence of a target object, the image sequence including I images, where I is an integer greater than 1, and the image sequence including visible light images and infrared images. A detection module for performing image detection on each image respectively to generate attribute information of each image. An output module for inputting the attribute information of the i-th image into a target model and outputting information of a first target algorithm for processing the i-th image in response to the difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image being greater than or equal to a predetermined threshold; wherein the target model is obtained by training an initial model based on the attribute information of sample images and algorithm information matching the attribute information of the sample images, 1 < i ≤ I. A processing module for processing each image respectively based on each target algorithm matching the attribute information of each image to obtain a tracking result of the target object.
[0008] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0009] A fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instruction stored thereon, and the above computer program or instruction implements the steps of the above method when executed by a processor.
[0010] A fifth aspect of the present invention further provides a computer program product including a computer program or instruction, and the above computer program or instruction implements the steps of the above method when executed by a processor.
[0011] According to an embodiment of the present disclosure, image detection is performed on the images in the acquired image sequence, and attribute information corresponding to each image is generated, converting a large amount of image information into attribute information of the images, rather than directly inputting the images, reducing the amount of data during information processing and improving the calculation efficiency. When the difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image is greater than or equal to a predetermined threshold, the attribute information of the i-th image is input into the target model, and information of a first target algorithm for processing the i-th image is output, and the images are processed according to the matching algorithm. During this process, only the attribute information of the images with a large difference is input into the model, and then the target tracking algorithm suitable for the image scene is matched, which can reduce the calculation amount, reduce the frequency of algorithm replacement, and improve the calculation efficiency. At the same time, through the trained target model matching algorithm, the algorithm scheduling can be quickly and accurately performed according to the scene task requirements of the images, improving the calculation efficiency of target tracking while ensuring the target tracking effect and optimizing the performance. Description of the Drawings
[0012] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0013] Figure 1 Shows an application scenario diagram of an object tracking method based on a task-oriented object tracking algorithm matching strategy according to an embodiment of the present invention;
[0014] Figure 2 Shows a flowchart of an object tracking method based on a task-oriented object tracking algorithm matching strategy according to an embodiment of the present invention;
[0015] Figure 3 Shows a flowchart of an object tracking method based on a task-oriented object tracking algorithm matching strategy according to another embodiment of the present invention;
[0016] Figure 4 Shows a flowchart of an object model training method according to an embodiment of the present invention;
[0017] Figure 5 Shows a block diagram of an evaluation system for the effects of various object tracking algorithms according to an embodiment of the present invention;
[0018] Figure 6 Shows a schematic diagram of an object algorithm matching process according to an embodiment of the present invention;
[0019] Figure 7 Shows a schematic diagram of generating image attribute information according to an embodiment of the present invention;
[0020] Figure 8 Shows a specific example flowchart of an object tracking method according to an embodiment of the present invention;
[0021] Figure 9 Shows a structural block diagram of an object tracking device based on a task-oriented object tracking algorithm matching strategy according to an embodiment of the present invention;
[0022] Figure 10 Shows a block diagram of an electronic device suitable for implementing an object tracking method based on a task-oriented object tracking algorithm matching strategy. Detailed Embodiments
[0023] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0024] Object tracking is a common computer vision task, which aims to continuously monitor the position of an object in video frames through different means and is widely applied in fields such as autonomous driving, intelligent monitoring, and robotics.
[0025] Single-modal object tracking usually relies on one data source, such as visual images, and has the advantages of simple implementation and high computational efficiency, being suitable for scenarios with relatively stable environmental conditions. However, it may perform poorly under conditions such as light changes, occlusions, and complex backgrounds.
[0026] In recent years, in order to address various challenges in complex environments, the idea of multi-modal fusion has achieved rapid development in the field of object tracking. Multi-modal object tracking technology is an important research direction in the field of computer vision, aiming to improve the accuracy and robustness of object tracking by combining data from different sensors or information sources. For example, multi-modal tracking integrates various information such as visual, infrared, lidar, and acoustic data. Through the information complementary strategy of multiple modal data, a more effective object tracking effect can be achieved in complex scenarios. For example, the RGBT (Red-Green-Blue Thermal, visible light and thermal infrared) object tracking algorithm that utilizes visible light images and infrared images takes into account both the texture information and thermal information of the tracking object, enabling high-precision object tracking in complex or dynamic environments and improving the accuracy and robustness of tracking. However, when applying the multi-modal information fusion strategy, the implementation complexity is relatively high, and the computational resource requirements are also relatively large. While improving the tracking accuracy, it will increase the computational burden during object tracking, resulting in relatively low tracking performance and efficiency.
[0027] In addition, there are numerous existing multi-modal object tracking algorithms, each with its own advantages and disadvantages, and different algorithms have different advantageous scenarios. In practical applications, for different tracking scenarios and task requirements, there is a lack of an algorithm matching strategy for different tracking scenarios to adaptively execute tracking tasks.
[0028] In view of this, an embodiment of the present invention provides an object tracking method based on a matching strategy of a task-oriented object tracking algorithm. The method performs image detection on images in an acquired image sequence and generates attribute information corresponding to each image, converting a large amount of image information into the attribute information of the images instead of directly inputting the images, reducing the amount of data during information processing and improving the calculation efficiency. When the difference between the attribute information of the i-th image and the attribute information of the (i-1)-th image is greater than or equal to a predetermined threshold, the attribute information of the i-th image is input into a target model, and information of a first target algorithm for processing the i-th image is output, and the images are processed according to the matched algorithm. During this process, only the image attribute information with a large difference is input into the model, and then a target tracking algorithm suitable for the image scene is matched, which can reduce the calculation amount, reduce the frequency of algorithm replacement, and improve the calculation efficiency. At the same time, through the trained target model matching algorithm, algorithm scheduling can be quickly and accurately performed according to the scene task requirements of the images, improving the calculation efficiency of object tracking and optimizing the performance while ensuring the object tracking effect.
[0029] Figure 1 FIG. shows an application scenario diagram of an object tracking method based on a matching strategy of a task-oriented object tracking algorithm according to an embodiment of the present invention.
[0030] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0031] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.
[0033] The server 105 may be a server that provides various services. For example, it may be a background management server (merely an example) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0034] It should be noted that the object tracking method based on the task-oriented object tracking algorithm matching strategy provided in the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the object tracking device based on the task-oriented object tracking algorithm matching strategy provided in the embodiments of the present invention can generally be set in the server 105. The object tracking method based on the task-oriented object tracking algorithm matching strategy provided in the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the object tracking device based on the task-oriented object tracking algorithm matching strategy provided in the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0035] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0036] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 8 scenario described below, the object tracking method based on the task-oriented object tracking algorithm matching strategy in the embodiments of the invention will be described in detail through
[0037] Figure 2 shows a flowchart of the object tracking method based on the task-oriented object tracking algorithm matching strategy according to an embodiment of the present invention.
[0038] As Figure 2 shown, the object tracking method 200 of this embodiment includes operations S210 to S240.
[0039] In operation S210, an image sequence for a target object is acquired.
[0040] In operation S220, image detection is performed on each image respectively to generate attribute information for each image.
[0041] In operation S230, in response to the difference between the attribute information of the i-th image and the attribute information of the (i-1)-th image being greater than or equal to a predetermined threshold, the attribute information of the i-th image is input into the target model, and information on the first target algorithm for processing the i-th image is output.
[0042] In operation S240, each image is processed respectively based on the respective target algorithms that match the attribute information of each image, and a tracking result for the target object is obtained.
[0043] According to an embodiment of the present invention, the target object is an object with tracking requirements. For example, the target object is an aerial tracking target in a complex environment. The image sequence includes I images, where I is an integer greater than 1. The image sequence includes visible light images and infrared images, and the visible light images and infrared images can be obtained through a binocular camera for simultaneously obtaining the texture information and thermal information of the same target object.
[0044] According to an embodiment of the present invention, for each image in the above image sequence, before image detection, the image can be preprocessed. Specifically, it includes operations such as image alignment, time synchronization, and image denoising.
[0045] Image detection includes detecting the brightness, contrast, sharpness, radiance, thermal noise, thermal contrast, etc. of the image, that is, evaluating the state or quality of the image, and then generating the attribute information of each image. The attribute information includes a set formed by the brightness standard deviation, contrast standard deviation, and sharpness variance obtained after the image is computationally processed. Brightness, contrast, and sharpness can be used as indicators for evaluating the state or quality of the image, and they form the corresponding index set. The above process can simplify the image information containing a large amount of data into attribute information (i.e., simplified into an index set) and then input it into the target model, rather than directly inputting the image, which can reduce the pressure on network calculation.
[0046] For example: The t-th image is an image of the target object entering the grass collected in the dark. The image is detected to generate the brightness standard deviation, contrast standard deviation, and sharpness variance corresponding to the image, which respectively reflect the brightness, contrast, and sharpness of the image, and a set is generated .
[0047] According to an embodiment of the present invention, when the difference between the attribute information of the i-th image and the attribute information of the (i-1)-th image is greater than or equal to a predetermined threshold, the attribute information of the i-th image is input into the target model, and information on the first target algorithm for processing the i-th image is output, where 1 < i ≤ I.
[0048] The predetermined threshold can be 10%, 20%, 30%, etc., without limitation here. The first target algorithm can be an RGB (Red-Green-Blue, visible light) image tracking algorithm, a TIR (Thermal Infrared) image tracking algorithm, or an RGBT image tracking algorithm, which is used to perform the target tracking task. The information of the first target algorithm includes the name of the algorithm.
[0049] The difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image can be obtained by calculating the differences in the standard deviation of brightness, the standard deviation of contrast, and the variance of sharpness in the attribute information of the two images, and then averaging all the differences, or by weighted summing all the differences, without limitation here.
[0050] Compare the above difference with the predetermined threshold. When the difference is greater than or equal to the predetermined threshold, it indicates that the change difference between the two images is generally large, the scene changes significantly, and it is necessary to input the model to match the target tracking algorithm corresponding to the image attribute information to select a more suitable algorithm for the current image task scenario to process the image and obtain the tracking result for this image.
[0051] For example: The predetermined threshold is 20%. The attribute information of the 5th image is , and the attribute information of the 6th image is . Calculate the difference between the attribute information as . The average value of the above difference value is 68%, which is greater than the predetermined threshold. Therefore, the attribute information of the 6th image is input into the target model, and the output information of the first target algorithm for processing the 6th image is the RGBT image tracking algorithm. The 6th image is processed using the RGBT image tracking algorithm to obtain the tracking result for this image.
[0052] The target model is obtained by training the initial model based on the attribute information of the sample image and the algorithm information matching the attribute information of the sample image. The target model can be a neural network model. In this model, the input is the attribute information of the sample image, and the output is the name of the optimal target tracking algorithm.
[0053] According to an embodiment of the present invention, image detection is performed on the images in the acquired image sequence, and attribute information corresponding to each image is generated, converting a large amount of image information into attribute information of the images instead of directly inputting the images, reducing the amount of data during information processing and improving the computational efficiency. When the difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image is greater than or equal to a predetermined threshold, the attribute information of the i-th image is input into the target model, and information on the first target algorithm for processing the i-th image is output, and the image is processed according to the matching algorithm. During this process, only the attribute information of the images with relatively large differences is input into the model, and then the target tracking algorithm suitable for the image scene is matched, which can reduce the computational amount, reduce the frequency of algorithm replacement, and improve the computational efficiency. At the same time, through the trained target model matching algorithm, algorithm scheduling can be quickly and accurately performed according to the scene task requirements of the image, improving the computational efficiency of target tracking while ensuring the target tracking effect and optimizing the performance.
[0054] Figure 3 FIG. shows a flowchart of a target tracking method based on a task-oriented target tracking algorithm matching strategy according to another embodiment of the present invention.
[0055] As Figure 3 shown, the target tracking method 300 of this embodiment includes operations S310 to S360.
[0056] In operation S310, an image sequence for a target object is acquired.
[0057] In operation S320, image detection is performed on each image respectively, and attribute information of each image is generated.
[0058] In operation S330, it is determined whether the difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image is less than a predetermined threshold. When the determination result is yes, operation S340 is executed, and when the determination result is no, operation S350 is executed.
[0059] In operation S340, the algorithm information matching the attribute information of the (i - 1)-th image is determined as the second target algorithm information of the i-th image.
[0060] In operation S350, the attribute information of the i-th image is input into the target model, and information on the first target algorithm for processing the i-th image is output.
[0061] In operation S360, each image is processed based on the respective target algorithms matching the attribute information of each image, and a tracking result for the target object is obtained.
[0062] According to an embodiment of the present invention, the second target algorithm may be an RGB image tracking algorithm, a TIR image tracking algorithm, or an RGB-T image tracking algorithm. The information of the second target algorithm includes the name of the algorithm.
[0063] In the target model application stage, after each image is subjected to image detection and generates attribute information, it is unreasonable to input the image attribute information of each image into the target model. On the one hand, the neural network needs to perform calculations on each image, which will incur a high computational cost and affect real-time performance; on the other hand, changing algorithms too frequently in the same image sequence also has a high computational burden.
[0064] For similar task scenarios, the adjacent images in the image sequence have relatively small changes. If each image is input into the target model for algorithm matching, it will cause excessive computational volume, and frequently changing algorithms will also affect the computational efficiency of the target model. Therefore, the present invention compares the differences in the attribute information of adjacent two images, and only when the attribute information of adjacent images changes greatly, the target model is re-input to match the target tracking algorithm.
[0065] For example: the predetermined threshold is 20%, the attribute information of the 5th image is , the attribute information of the 6th image is , calculate the difference between the attributes as , and the average value of the above difference value is 13%, which is less than the predetermined threshold. Therefore, the "TIR image tracking algorithm" that matches the attribute information of the 5th image is determined as the second target algorithm for the 6th image. Process the 6th image based on the TIR image tracking algorithm that matches the attribute information of the 6th image to obtain the tracking result for the 6th image.
[0066] According to an embodiment of the present invention, by calculating the differences in the attribute information of adjacent images, comparing the relationship between the differences and the predetermined threshold, it is determined whether the image needs to be input into the target model for algorithm matching. When the difference is less than the predetermined threshold, it indicates that the scene change is small at this time, the algorithm is not switched, and the algorithm of the previous image is used as the target algorithm, so as to quickly obtain the algorithm information that matches the attribute information of the image, without inputting the attribute information of each image into the target model, which can effectively reduce the computational pressure and consumption of computing resources, and improve the efficiency of target algorithm matching.
[0067] Figure 4 The flowchart of the target model training method according to an embodiment of the present invention is shown.
[0068] As Figure 4 shown, the method 400 of this embodiment includes operations S431 to S436.
[0069] In operation S431, multiple sample images and attribute information corresponding to the multiple sample images are obtained.
[0070] In operation S432, for each sample image, multiple algorithms are respectively called to process each sample image, and multiple tracking results for each sample image are obtained.
[0071] In operation S433, according to the multiple tracking results and the running information of the multiple algorithms, multiple effect information of each algorithm for processing each sample image is determined.
[0072] In operation S434, based on the multiple effect information, multiple algorithm evaluation information for each sample image is determined.
[0073] In operation S435, algorithm information that matches the attribute information of each sample image is determined from the multiple algorithm evaluation information.
[0074] In operation S436, the initial model is trained using the sample data set to obtain the target model.
[0075] According to an embodiment of the present invention, the multiple algorithms include multiple object tracking algorithms in the object tracking algorithm library. The tracking results include results such as the coincidence between the tracked object and the real object and the tracking success rate after executing the object tracking algorithm. The running information of the algorithm includes running information such as the tracking frame rate, the number of model parameters, and the video memory occupancy.
[0076] There are a large number of existing object tracking algorithm types and quantities. How to scientifically set the tracking effect evaluation method, effectively measure the effectiveness and reliability of the tracking system, and then select the most suitable object tracking algorithm according to the environment and target state is of great significance for the implementation of the tracking task. Usually, the algorithm effect evaluation is divided into multiple dimensions, including accuracy, robustness, and timeliness, etc. Accuracy is usually evaluated by calculating the error between the tracking result and the real object position, including the intersection over union and the average precision. Robustness refers to the stability of the algorithm under different environmental conditions (such as light changes, occlusion, and motion blur), and is usually evaluated on multiple scenarios or data sets. Timeliness refers to the speed at which the algorithm processes frames, usually measured in frames per second (FPS).
[0077] Figure 5 The block diagram of the effect evaluation system of each object tracking algorithm according to an embodiment of the present invention is shown.
[0078] In the design of the evaluation system for the effects of various object tracking algorithms, considering that in some complex scenarios, it is not only required that the accuracy of the object tracking algorithm meets the requirements, but also that it can adapt to problems such as rapid changes in the scenario, rapid movement of the object, and different degrees of occlusion. At the same time, there are also certain requirements for the execution speed of the algorithm. Therefore, there is an urgent need in the actual application scenario for how to reasonably combine these evaluation indicators to design a comprehensive evaluation system for the effects of object tracking algorithms that is applicable to different application scenarios and task requirements. Based on the above problems, the present invention proposes an evaluation system for the tracking effects of multiple object tracking algorithms.
[0079] As Figure 5 shown, in the embodiments of the present invention, the primary evaluation indicators include accuracy, timeliness, and robustness, and the secondary evaluation indicators include intersection over union (IOU) and average precision (AP), tracking frame rate, number of model parameters, and video memory occupancy, temporal robustness, spatial robustness, and long-term tracking ability. The above primary evaluation indicators and secondary evaluation indicators together constitute the evaluation system for the effects of each object tracking algorithm.
[0080] According to the embodiments of the present invention, the effect information includes at least one of the following: accurate information, timeliness information, and robustness information. Among them, the accurate information, timeliness information, and robustness information include the index values calculated from the three indicators of accuracy, timeliness, and robustness, which are used to evaluate the tracking effect of the algorithm. The following is a specific elaboration of the above effect information:
[0081] The accurate information includes intersection over union (IOU) and average precision (AP), which evaluate whether the algorithm can accurately track the target and are the key to evaluating whether the algorithm plays a role. Among them, the intersection over union measures the coincidence of the tracked target, which can be determined by the coincidence of the tracked target and the real target after executing the object tracking algorithm, and the average precision measures the Euclidean distance between the target center points. The index value corresponding to accuracy can be determined by the weighted sum of the intersection over union and the average precision, that is, the accurate information is determined, as shown in the following formula (1):
[0082] (1)
[0083] Among them, represents the index value of the accuracy of the algorithm, represents the intersection over union of the algorithm, represents the average precision of the algorithm, a represents the weight of the intersection over union, and b represents the weight of the average precision.
[0084] For example, for image a, algorithms x and y are called. Based on the tracking results of algorithms x and y and the running information of algorithms x and y, it is determined that the IOU of algorithm x is 0.8, the AP is 0.9, the weight of IOU is 1.2, and the weight of AP is 0.9. Thus, the index value for the accuracy of algorithm x can be obtained as 1.77. Similarly, it is determined that the IOU of algorithm y is 0.9, the AP is 0.95, the weight of IOU is 1.2, and the weight of AP is 0.9. Thus, the index value for the accuracy of algorithm y can be obtained as 1.94.
[0085] The timeliness information includes the tracking frame rate (Frame Per Second, FPS), the number of model parameters (Model Parameters, MP), and the memory usage (Memory Usage, MU) during the operation of the algorithm, which evaluates the task execution speed of the algorithm and the usage of computing resources, reflecting the real-time performance of the algorithm. The index value of timeliness can be determined by the weighted sum of the tracking frame rate, the number of model parameters, and the memory usage, that is, the timeliness information is determined as shown in the following formula (2):
[0086] (2)
[0087] Where, represents the index value of the timeliness of the algorithm, represents the tracking frame rate of the algorithm, represents the number of model parameters of the algorithm, represents the memory usage of the algorithm, c represents the weight of the tracking frame rate, d represents the weight of the number of model parameters, and e represents the weight of the memory usage.
[0088] For example, for image a, algorithms x and y are called. Based on the tracking results of algorithms x and y and the running information of algorithms x and y, it is determined that the FPS of algorithm x is 30, the MP is 100, the MU is 1024, the weight of FPS is 0.7, the weight of MP is 0.8, and the weight of MU is 0.75. Thus, the index value for the timeliness of algorithm x can be obtained as 869. Similarly, it is determined that the FPS of algorithm y is 25, the MP is 200, the MU is 2048, the weight of FPS is 0.7, the weight of MP is 0.8, and the weight of MU is 0.75. Thus, the index value for the timeliness of algorithm y can be obtained as 1713.
[0089] The robust information includes the Temporal Robustness Evaluation (TRE), Spatial Robustness Evaluation (SRE), and Long-term Tracking Capability (LTC) of the algorithm, which evaluate the stability of the algorithm in the face of anomalies and whether it can continuously and normally complete the target tracking task under perturbations. Among them, the temporal robustness can be determined by calculating the tracking success rate of the target in continuous time, the spatial robustness can be determined by calculating the tracking success rate of the target at different spatial positions (such as the center and edge of the image), and the long-term tracking capability can be determined by calculating the loss rate of the target during long-term tracking. The index value of the robustness can be determined by the weighted sum of the temporal robustness evaluation, spatial robustness evaluation, and long-term tracking capability, that is, the robust information is determined, as shown in the following formula (3):
[0090] (3)
[0091] Among them, represents the index value of the robustness of the algorithm, represents the temporal robustness of the algorithm, represents the spatial robustness of the algorithm, represents the long-term tracking capability of the algorithm, f represents the weight of the temporal robustness, g represents the weight of the spatial robustness, and h represents the weight of the long-term tracking capability.
[0092] For example: for image a, algorithms x and y are called. According to the tracking results of algorithms x and y and the running information of algorithms x and y, the TRE of algorithm x is determined to be 0.8, the SRE is 0.85, the LTC is 0.1, the weight of TRE is 0.6, the weight of SRE is 0.7, and the weight of LTC is 0.75. Thus, the index value of the robustness of algorithm x can be obtained as 1.15. Similarly, the TRE of algorithm y is determined to be 0.9, the SRE is 0.75, the LTC is 0.2, the weight of TRE is 0.6, the weight of SRE is 0.7, and the weight of LTC is 0.75. Thus, the index value of the robustness of algorithm y can be obtained as 1.21.
[0093] According to the embodiments of the present invention, the algorithm evaluation information includes the comprehensive score of the algorithm. Based on multiple effect information, determining the multiple algorithm evaluation information for each sample image includes obtaining the weighted sum of the calculated index values of the algorithm accuracy, timeliness, and robustness, as shown in the following formula (4):
[0094] (4)
[0095] Among them, Represents the comprehensive score of the algorithm, Represents the index value of the accuracy of the algorithm, Represents the index value of the timeliness of the algorithm, Represents the index value of the robustness of the algorithm. A represents the weight of accuracy, B represents the weight of timeliness, and C represents the weight of robustness.
[0096] For example: For image a, algorithms x and y are called, and the index values of accuracy, timeliness, and robustness for algorithm x are 1.77, 869, and 1.15 respectively. The index values of accuracy, timeliness, and robustness for algorithm y are 1.94, 1713, and 1.21 respectively. The weights of accuracy, timeliness, and robustness are 0.9, 0.8, and 0.7 respectively. Thus, the comprehensive scores of algorithms x and y are 697 and 1372 respectively.
[0097] According to an embodiment of the present invention, to determine the algorithm information that matches the attribute information of each sample image from multiple algorithm evaluation information, the comprehensive scores of multiple algorithms can be sorted, and the algorithm with the smallest sorting value can be determined as the sample target algorithm.
[0098] For example: For sample image a, the comprehensive score sorting of algorithms x and y is obtained. The sorting value of algorithm x is 1, and the sorting value of algorithm y is 2. Then algorithm x is determined as the sample target algorithm that matches the attribute information of sample image a, and the algorithm information of algorithm x is the algorithm information that matches the attribute information of sample image a.
[0099] According to an embodiment of the present invention, using a sample data set to train an initial model to obtain a target model, where the sample data set includes multiple groups of sample data, and each group of sample data includes the attribute information of the sample image and the algorithm information that matches the sample image.
[0100] Figure 6 Shows a schematic diagram of using a sample data set for target algorithm matching according to an embodiment of the present invention.
[0101] As Figure 6 shown, a sample image sequence is obtained, including visible light images and infrared images. Data preprocessing is performed on the images, such as image alignment, time synchronization, contrast adjustment, image enhancement, etc. Image detection is performed on the preprocessed images respectively. For example, brightness analysis, contrast analysis, and clarity analysis are performed on the visible light images to obtain the attribute information of the images, that is, the target state evaluation index set of the images. It is input into the decision-making system and algorithm switching module in the target model to obtain the comprehensive scores of multiple target tracking algorithms (target tracking algorithms 1 to q) in the target tracking algorithm library, and then the best target tracking algorithm is determined to perform the tracking task. Figure 6The decision-making system and algorithm switching module shown in the figure is a schematic structure of a neural network model.
[0102] Regarding the target model and training, it should be noted that the target model does not affect the existing various target tracking algorithms. The purpose is to make a comprehensive judgment on the existing various single-modal / multi-modal target tracking algorithms, and select the most suitable one of the algorithms to perform the target tracking task under different scenarios and different task requirements, ensuring that only one algorithm is selected to perform the target tracking task each time.
[0103] Specifically, the input and output of the target model are designed as follows: First, obtain each image state evaluation index set, that is, the attribute information of each image. At this time, each image will be transformed into a small index set, and then a large index set containing all the images in the image sequence will be formed; Second, according to the existing various target tracking algorithms, call each algorithm to calculate the tracking effect in the current scenario, and based on this effect information and the algorithm comprehensive evaluation method, calculate the comprehensive score of each algorithm as the module output. The state evaluation index of each image corresponds to the comprehensive score of each algorithm under this image, forming an input-output pair. Repeat the above operations for each image sequence, and the input and output of the entire data set will be obtained as the input-output training set of the target model. Each sample data set is as follows:
[0104]
[0105]
[0106]
[0107]
[0108]
[0109] Among them, represents the attribute information of the input image in a certain scenario, including the attribute information of each image. There are m images in this scenario; represents the attribute information of the t-th input image, including the attribute information of the visible light image and the infrared image of this image, represents the first attribute information, and so on. Each image has n attribute information. represents the selected target tracking algorithm library, represents the first algorithm, and so on. There are q target tracking algorithms to be selected. represents the scores of each algorithm in the tracking algorithm library, including the scores of each algorithm under each image; represents the scores of each algorithm for the t-th input image, Denote the tracking effect score of the first algorithm for the \(t\)-th input image. There are \(q\) target tracking algorithms to be selected, so each image has \(q\) scores.
[0110] Essentially, the above method calculates the tracking results of each tracking algorithm in advance. A large number of calculation processes are used to generate the training dataset, and the target model is obtained through the training of the neural network. When using this model, there is no need to calculate each tracking algorithm, and the best algorithm can be selected to perform the tracking task, saving calculation time and effectively improving the accuracy of target tracking.
[0111] According to an embodiment of the present invention, for each sample image, multiple algorithms are called to process each sample image, and multiple tracking results for each sample image are obtained. Then, the effect information of multiple algorithms is determined. Through the effect information, the evaluation information of the algorithms is obtained, and then the algorithm information matching each sample image is determined. The above information constitutes the sample dataset, and the matching relationship between the attribute information of the sample image and the sample target algorithm is established for the training of the target model. In the application of the target model, by inputting the attribute information of the image into the trained target model, the algorithm information matching the attribute information of the image can be output, realizing the rapid matching of the best algorithm for different image task scenarios and improving the efficiency of algorithm matching.
[0112] According to an embodiment of the present invention, based on multiple effect information, multiple algorithm evaluation information for each sample image is determined, including: using the analytic hierarchy process to process multiple effect information to obtain the first weight of each effect information for each sample image; using the objective weighting method to process multiple effect information to obtain the second weight of each effect information for each sample image; based on the first weight and the second weight of multiple effect information, determining the target weight of each effect information; and obtaining multiple algorithm evaluation information according to the target weight of each effect information.
[0113] According to an embodiment of the present invention, through the method of combining the subjective weighting method and the objective weighting method of the analytic hierarchy process, reasonable weights are assigned to the indicators in the above effect information, and the specific process is as follows:
[0114] Subjective weighting using the analytic hierarchy process:
[0115] First, compare the importance of the indicators of each algorithm to obtain the judgment matrix.
[0116] Table 1 is the relative scale table of the importance of algorithm effect information.
[0117] Table 1
[0118]
[0119] For example: According to expert experience, it is considered that accuracy is slightly more important than timeliness, accuracy is significantly more important than robustness, and timeliness is slightly more important than robustness. According to the relative scale in Table 1, the judgment matrix A is constructed as shown in the following formula (5):
[0120] (5)
[0121] Calculate the maximum eigenvalue of the judgment matrix , and the corresponding maximum eigenvector , for example, the calculated maximum eigenvalue is m, and the maximum eigenvector is .
[0122] Pass the consistency test to reduce the subjective random error. Specifically, first calculate the consistency index CI, and the calculation formula is as shown in the following formula (6):
[0123] (6)
[0124] Among them, represents the maximum eigenvalue of the judgment matrix, N represents the dimension of the judgment matrix, and CI represents the consistency index.
[0125] Judge the relationship between the random index RI and the consistency index CI. If , it means that the subjectively constructed eigenvector is reasonable. Otherwise, it is necessary to reconstruct the judgment matrix and perform the consistency test again until the calculated weight vector is reasonable. The value of the random index RI is related to the dimension of the matrix, as shown in Table 2.
[0126] Table 2 is the random index RI of the n-order matrix.
[0127] Table 2
[0128]
[0129] For example: For the three dimensions of accuracy, robustness, and timeliness, N = 3, CI = 0.03. Query Table 2 to get RI = 0.58, and get , which means that the subjectively constructed eigenvector is reasonable, and the weights of accuracy, robustness, and timeliness can be determined according to the eigenvector.
[0130] Objective weighting is carried out using the CRITIC (Criteria Importance Though Intercrieria Correlation) objective weighting method: The contrast intensity refers to the magnitude of the difference in the values of each evaluation scheme for the same indicator, which is expressed in the form of the standard deviation. The larger the standard deviation, the greater the fluctuation, that is, the greater the difference in the values between the schemes, and the higher the weight. The conflict between indicators is represented by the correlation coefficient. If there is a strong positive correlation between two indicators, it indicates that the conflict is smaller and the weight will be lower.
[0131] First, to eliminate the influence of different dimensions on the evaluation results, it is necessary to perform dimensionless processing on each indicator. The present invention adopts positive indicators, where the larger the value of the indicator, the better. The dimensionless processing formula is as shown in the following formula (7):
[0132] (7)
[0133] Among them, represents the test data value of the i-th indicator under the j-th algorithm, represents the data after dimensionless processing, and represent the maximum and minimum values in the test data of the i-th indicator under p algorithms.
[0134] Calculate the indicator variability: By using the standard deviation to measure the difference and fluctuation of the test data of the algorithm under different indicators, it can reflect the performance difference information of different algorithms. The calculation formula is as shown in the following formulas (8) and (9):
[0135] (8)
[0136] (9)
[0137] Among them, represents the mean value of the test data of p algorithms under the i-th indicator, represents the standard deviation.
[0138] Calculate the indicator conflict: Use the correlation coefficient to measure the statistical data correlation of the same algorithm model under different evaluation indicators, and reflect the overlap degree of the algorithm performance information. The calculation formula is as shown in the following formulas (10) and (11):
[0139] (10)
[0140] (11)
[0141] Among them, Represents the test data value of the j-th algorithm under the x index, Represents the test data value of the j-th algorithm under the y index, and respectively represent the average values of the test data of the algorithm under the two indexes, Represents the correlation coefficient between evaluation indexes, Represents the index conflict of evaluation index x, Represents the correlation coefficient between index x and index i, and n represents the number of other indexes.
[0142] Calculate the objective weight: Combining the conflict and variability of index evaluation, give the calculation of the objective weight value of the evaluation index. The calculation formula is shown in the following formula (12):
[0143] (12)
[0144] Among them, Represents the amount of information. The larger the value, the greater its role in the entire evaluation index system, and more weight should be assigned to it, W i Represents the objective weight.
[0145] Complement the subjective and objective weighting methods, so that the difference between the final weight vector and various weighting methods reaches the minimum, and give full play to the advantages of each method. Specifically, calculate the final target weight through the following formula (13) and formula (14):
[0146] (13)
[0147] (14)
[0148] Among them, and respectively represent the weight vectors of the subjective weighting method and the objective weighting method, and Represent their weight ratios, Represents the final weight vector. Solving formula (13) can obtain and The values of, and finally obtain the comprehensive weight vector.
[0149] It should be noted that in practical applications, if a certain index is not considered in the task scenario, the weight of this index can be set to 0; if the task scenario is special and it is necessary to increase the number of evaluation indexes or subjective and objective evaluation methods, it can also be supplemented.
[0150] According to an embodiment of the present invention, by combining subjective weighting and objective weighting, considering the internal statistical laws and authoritative values among index data, weights are assigned to multiple effect information to obtain the target weights of each effect information, which can not only reflect the objective laws of each effect information based on task requirements and environmental characteristics, reduce subjective influence, but also conform to certain value concepts of decision-makers, so as to more scientifically evaluate the algorithm and make up for the deficiencies brought by single weighting.
[0151] According to an embodiment of the present invention, image detection is performed on each image respectively, and the generated attribute information of each image includes: determining a detection strategy according to the type of each image; based on the detection strategy, detecting the image to obtain the attribute information of each image.
[0152] According to an embodiment of the present invention, determining a detection strategy according to the type of each image includes: in response to the type of the image being a visible light image, determining the detection strategy as performing visible light image detection on the image; in response to the type of the image being an infrared image, determining the detection strategy as performing infrared image detection on the image.
[0153] According to an embodiment of the present invention, determining a corresponding detection strategy according to different types of images and performing targeted detection on the image based on the detection strategy can reduce the detection error caused by different image types, can more accurately obtain the attribute information of the image, and improve the accuracy of image detection.
[0154] According to an embodiment of the present invention, the attribute information of each image includes at least one of the following: brightness information, contrast information, and clarity information; based on the detection strategy, detecting the image to obtain the attribute information of each image includes: performing gray-scale detection on the image to obtain the first gray-scale information of each image; obtaining the brightness information of each image by processing the first gray-scale information; performing contrast enhancement processing on each image; performing gray-scale detection on the enhanced image to obtain the second gray-scale information of each image; obtaining the contrast information of each image by processing the second gray-scale information; performing Laplace transform processing on each image; performing pixel detection on the transformed image to obtain the pixel information of each image; determining the clarity information of each image by processing the pixel information.
[0155] Figure 7 A schematic diagram showing the generation of image attribute information according to an embodiment of the present invention is shown.
[0156] The average brightness refers to the average value of the brightness values of all pixels in an image. For visible light images, the average brightness usually depends on the exposure settings and lighting conditions of the image. Under normal circumstances, the average brightness should be within a reasonable range to ensure that the image is not too dark or too bright. The standard deviation represents the distribution of the image brightness, that is, the range of brightness changes. A higher standard deviation means that the brightness of the image changes greatly, while a lower standard deviation means that the brightness changes little. For object detection tasks, the standard deviation usually needs to be large enough to ensure that the details and object features in the image can be effectively distinguished. Generally speaking, the standard deviation should be kept within a moderate range, usually between 40 and 100 is more common.
[0157] As Figure 7 shown, the gray-scale detection is performed on the image 720 to obtain the first gray-scale information 720-1 of the image 720. By processing the first gray-scale information 720-1, the brightness information 720-2 of the image 720 is obtained. The first gray-scale information includes the gray-scale values of the pixels in the image, and the brightness information includes the average value of the image brightness and the standard deviation of the image brightness.
[0158] The formula for calculating the average image brightness is shown in the following formula (15):
[0159] (15)
[0160] Where represents the average value of the image brightness, is the gray-scale value of the th pixel, and
[0161] is the total number of pixels in the image.
[0162] (16)
[0163] Where represents the standard deviation of the image brightness, represents the average value of the image brightness, is the gray-scale value of the th pixel, and
[0164] Contrast refers to the difference between the brightest and darkest regions in an image. High contrast helps to clearly distinguish the object from the background, while low contrast images may cause the object to be blurred.
[0165] As Figure 7As shown, contrast enhancement processing is performed on the image 720; gray level detection is performed on the enhanced image 721 to obtain the second gray level information 721-1 of the image 720; by processing the second gray level information 721-1, the contrast information 721-2 of the image 720 is obtained. The second gray level information includes the gray level value and the gray level mean of each pixel of the image after contrast enhancement processing. The contrast information includes the contrast standard deviation.
[0166] To better preserve image details and avoid the possible increase in noise contrast caused by global histogram equalization, CLAHE (Contrast Limited Adaptive Histogram Equalization) is selected as the contrast enhancement algorithm, and the enhanced contrast value is compared with a predefined threshold to evaluate whether the contrast is sufficient.
[0167] The input image is divided into small blocks, and for each block, its gray level histogram is statistically calculated. .
[0168] The cumulative distribution function is calculated using the following formula (17):
[0169] (17)
[0170] The equalization mapping normalizes the CDF (Cumulative Distribution Function value) to the target range, usually from 0 to L-1, and the calculation formula for the pixel value is as shown in the following formula (18):
[0171] (18)
[0172] Among them, represents the gray level histogram function of each block of the image, N represents the number of small blocks after the image is divided, represents the cumulative distribution function, L represents the number of gray levels, is the minimum CDF value in the local area, represents the pixel value function.
[0173] To avoid over-enhancing the contrast, CLAHE limits the histogram through histogram clipping and reallocating the clipped values. After applying histogram equalization to each block, the boundaries of the blocks are smoothed through bilinear interpolation to reduce the visual demarcation between blocks. This method enhances image details through local processing while limiting the excessive increase in contrast to avoid the amplification of image noise and solve the problems of low-contrast or images with non-uniform illumination.
[0174] After adaptive histogram equalization, for the pixel value matrix of the image , the global contrast can be determined by the standard deviation of contrast , and the calculation formula of the standard deviation of contrast is shown in the following formula (19):
[0175] (19)
[0176] Where is the gray value of the th pixel after adaptive histogram equalization processing, is the average gray value of the processed pixels, represents the standard deviation of contrast, and N represents the number of small blocks into which the image is divided.
[0177] As Figure 7 shown, perform Laplace transform processing on image 720; perform pixel detection on the processed image 722 to obtain the pixel information 722-1 of image 720; determine the sharpness information 722-2 of image 720 by processing the pixel information 722-1. The pixel information includes the pixel values and pixel means after Laplace transform processing, and the sharpness information includes the sharpness variance.
[0178] Calculate the sharpness of the image using the Laplace operator based on the edge and detail features of the image, and describe the high-frequency information of the image. First, convert the image to a grayscale image to simplify the calculation. Then, convolve the grayscale image with the selected Laplace convolution kernel to obtain the image after Laplace transform. The Laplace operator can be approximated by the following convolution kernel shown in formula (20):
[0179] (20)
[0180] These kernels will be applied to each pixel of the image, and the Laplace transform of the image is calculated through convolution operations. The image after Laplace transform represents the high-frequency components of the image, and its variance can reflect the edge and detail features of the image. Judge the sharpness of the image according to the calculated variance value. A higher variance value indicates that the image is clear, and a lower variance value indicates that the image is blurred. The calculation formula of the sharpness variance is as follows:
[0181] (21)
[0182] Where represents the pixel value of the th pixel after Laplace transform, represents the average value of these pixels, is the total number of pixels, represents the sharpness variance.
[0183] According to an embodiment of the present invention, by detecting an image, the brightness information, contrast information, and clarity information of the image are obtained. Among them, the brightness information can be obtained from the first grayscale information, the contrast information can be obtained from the second grayscale information after enhancing the contrast of the image, and the clarity information can be obtained from the pixel information after performing a Laplace transform on the image. This realizes the conversion of a large amount of image information into the attribute information of the image, rather than directly inputting the image, reduces the amount of data during information processing, and improves the calculation efficiency.
[0184] The technical solution of the present invention will be further elaborated below in conjunction with specific embodiments. It should be noted that these specific embodiments are only for facilitating those skilled in the art to better understand the technical solution of the present invention, and should not be construed as limiting the protection scope of the present invention.
[0185] Figure 8 A specific example flowchart of the target tracking method according to an embodiment of the present invention is shown.
[0186] As Figure 8 shown, for image sequence 1, data preprocessing is performed on the visible light and thermal infrared images of this sequence. Subsequently, quality detection of the visible light image and the infrared image is performed respectively, and the environmental and target state evaluation indicators of the visible light image and the infrared image are obtained respectively. These are input into the target model, and the target model will output the names of the corresponding target tracking algorithms, such as the RGB image tracking algorithm, the TIR image tracking algorithm, and the RGBT image tracking algorithm. Subsequently, the tracking task of the images in image sequence 1 is performed according to this algorithm. Secondly, in subsequent images, the above operations are repeated until the tracking task of this tracking scenario is completed.
[0187] According to an embodiment of the present invention, first, a target model for different modality target tracking algorithms is proposed, the input and output, training process, and application strategy of this model are designed, and it is realized to quickly select a target tracking algorithm that adapts to the characteristics of the scene and task requirements based on this model, which can play a role under different environmental and task conditions and has strong adaptability. In addition, this method only supports the expansion and replacement of algorithms in the algorithm library. After replacing the algorithm, the target model can be retrained after appropriately adjusting the input index set of the image, and it has strong expandability. Secondly, an expandable input image target state evaluation index set is proposed. This index set is for different modality target tracking fields and includes the basic characteristics of the input image data. As the input of the target model, it greatly reduces the calculation pressure of the model and supports the customized design requirements of special task scenarios. In addition, based on the three dimensions of the accuracy, timeliness, and robustness of the algorithm, the present invention proposes a comprehensive target tracking algorithm performance index evaluation system, which can be applicable under different application scenarios and task requirements, reflects the comprehensive scores of different algorithms, and combines with the target model to quickly select the optimal algorithm to perform tasks, and has strong balance.
[0188] Based on the above object tracking method with a task-oriented object tracking algorithm matching strategy, the present invention also provides an object tracking device based on a task-oriented object tracking algorithm matching strategy. The following will be combined with Figure 9 to describe this device in detail.
[0189] Figure 9 The structural block diagram of an object tracking device based on a task-oriented object tracking algorithm matching strategy according to an embodiment of the present invention is shown.
[0190] As Figure 9 shown, the object tracking device 900 based on the task-oriented object tracking algorithm matching strategy of this embodiment includes an acquisition module 910, a detection module 920, an output module 930, and a processing module 940.
[0191] The acquisition module 910 is configured to acquire an image sequence of a target object. The image sequence includes I images, where I is an integer greater than or equal to 1, and the image sequence includes visible light images and infrared images. In one embodiment, the acquisition module 910 can be used to perform the operation S210 described above, which will not be elaborated here.
[0192] The detection module 920 is configured to perform image detection on each image respectively to generate attribute information of each image. In one embodiment, the detection module 920 can be used to perform the operation S220 described above, which will not be elaborated here.
[0193] The output module 930 is configured to input the attribute information of the i-th image into the target model and output information of a first target algorithm for processing the i-th image in response to the difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image being greater than a predetermined threshold; wherein, the target model is obtained by training an initial model based on the attribute information of sample images and algorithm information matching the attribute information of the sample images, 1 < i ≤ I. In one embodiment, the output module 930 can be used to perform the operation S230 described above, which will not be elaborated here.
[0194] The processing module 940 is configured to process each image based on each target algorithm matching the attribute information of each image to obtain a tracking result of the target object. In one embodiment, the processing module 940 can be used to perform the operation S240 described above, which will not be elaborated here.
[0195] According to an embodiment of the present invention, the output module includes a second target algorithm determination sub-module, which is configured to determine the algorithm information matching the attribute information of the (i - 1)-th image as the information of the second target algorithm of the i-th image in response to the difference between the attribute information of the i-th image and the attribute information of the (i - 1)-th image being less than a predetermined threshold.
[0196] According to an embodiment of the present invention, the output module further includes an acquisition sub-module, a call sub-module, an effect information determination sub-module, an algorithm evaluation information determination sub-module, an algorithm information determination sub-module, and a training sub-module.
[0197] The acquisition sub-module is configured to acquire a plurality of sample images and attribute information corresponding to the plurality of sample images. The call sub-module is configured to, for each sample image, respectively call a plurality of algorithms to process each sample image, and obtain a plurality of tracking results for each sample image. The effect information determination sub-module is configured to determine, according to the plurality of tracking results and the running information of the plurality of algorithms, a plurality of effect information for each algorithm to process each sample image, where the plurality of effect information includes at least one of the following: accuracy information, timeliness information, and robustness information. The algorithm evaluation information determination sub-module is configured to determine, based on the plurality of effect information, a plurality of algorithm evaluation information for each sample image. The algorithm information determination sub-module is configured to determine, from the plurality of algorithm evaluation information, algorithm information that matches the attribute information of each sample image. The training sub-module is configured to use a sample data set to train an initial model to obtain a target model, where the sample data set includes multiple groups of sample data, and each group of sample data includes the attribute information of the sample image and the sample target algorithm that matches the sample image.
[0198] According to an embodiment of the present invention, the algorithm evaluation information determination sub-module includes a first weight determination unit, a second weight determination unit, a target weight determination unit, and an algorithm evaluation information determination unit. The first weight determination unit is configured to process the plurality of effect information by using the analytic hierarchy process to obtain a first weight of each effect information for each sample image. The second weight determination unit is configured to process the plurality of effect information by using an objective weighting method to obtain a second weight of each effect information for each sample image. The target weight determination unit is configured to determine a target weight of each effect information based on the first weight and the second weight of the plurality of effect information. The algorithm evaluation information determination unit is configured to obtain a plurality of algorithm evaluation information according to the target weight of each effect information.
[0199] According to an embodiment of the present invention, the detection module includes a detection strategy determination sub-module and a detection sub-module. The detection strategy determination sub-module is configured to determine a detection strategy according to the type of each image. The detection sub-module is configured to detect an image based on the detection strategy to obtain the attribute information of each image.
[0200] According to an embodiment of the present invention, the detection strategy determination sub-module includes a visible light image detection unit and an infrared image detection unit. The visible light image detection unit is configured to, in response to the type of the image being a visible light image, determine that the detection strategy is to perform visible light image detection on the image. The infrared image detection unit is configured to, in response to the type of the image being an infrared image, determine that the detection strategy is to perform infrared image detection on the image.
[0201] According to an embodiment of the present invention, the attribute information of each image includes at least one of the following: brightness information, contrast information, and sharpness information. The detection sub-module includes a first grayscale detection unit, a first grayscale information processing unit, a contrast enhancement processing unit, a second grayscale detection unit, a second grayscale information processing unit, a Laplace transform processing unit, a pixel detection unit, and a pixel information processing unit.
[0202] The first grayscale detection unit is configured to perform grayscale detection on the image to obtain the first grayscale information of each image. The first grayscale information processing unit is configured to obtain the brightness information of each image by processing the first grayscale information. The contrast enhancement processing unit is configured to perform contrast enhancement processing on each image. The second grayscale detection unit is configured to perform grayscale detection on the enhanced image to obtain the second grayscale information of each image. The second grayscale information processing unit is configured to obtain the contrast information of each image by processing the second grayscale information. The Laplace transform processing unit is configured to perform Laplace transform processing on each image. The pixel detection unit is configured to perform pixel detection on the transformed image to obtain the pixel information of each image. The pixel information processing unit is configured to determine the sharpness information of each image by processing the pixel information.
[0203] According to an embodiment of the present invention, any multiple of the acquisition module 910, the detection module 920, the output module 930, and the processing module 940 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 910, the detection module 920, the output module 930, and the processing module 940 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the acquisition module 910, the detection module 920, the output module 930, and the processing module 940 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0204] Figure 10 A block diagram of an electronic device suitable for implementing a target tracking method based on a task-oriented target tracking algorithm matching strategy according to an embodiment of the present invention is shown.
[0205] AsFigure 10 As shown, the electronic device 1000 according to an embodiment of the present invention includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. The processor 1001 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), and so on. The processor 1001 can also include on-board memory for caching purposes. The processor 1001 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0206] In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The processor 1001 performs various operations of the method flow according to an embodiment of the present invention by executing the program in the ROM 1002 and / or the RAM 1003. It should be noted that the program can also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 can also perform various operations of the method flow according to an embodiment of the present invention by executing the program stored in the one or more memories.
[0207] According to an embodiment of the present invention, the electronic device 1000 can further include an input / output (I / O) interface 1005, and the input / output (I / O) interface 1005 is also connected to the bus 1004. The electronic device 1000 can further include one or more of the following components connected to the input / output (I / O) interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output (I / O) interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.
[0208] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0209] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 1002 and / or RAM 1003 and / or one or more memories other than ROM 1002 and RAM 1003.
[0210] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs on a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present invention.
[0211] When the computer program is executed by the processor 1001, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0212] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 1009, and / or be installed from the removable medium 1011. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0213] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above functions defined in the system of the embodiment of the present invention are executed. According to the embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0214] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0215] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0216] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0217] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A target tracking method based on a task-oriented target tracking algorithm matching strategy, characterized in that: The method comprises: Acquire an image sequence for the target object, the image sequence comprising I images, where I is an integer greater than 1, and the image sequence comprises a visible light image and an infrared image; Performing image detection on each image respectively to generate attribute information of each image; In response to a difference between the attribute information of the ith image and the attribute information of the i-1th image being greater than or equal to a predetermined threshold, inputting the attribute information of the ith image into a target model, and outputting information of a first target algorithm for processing the ith image; wherein the target model is obtained by training an initial model based on the attribute information of a sample image and algorithm information matching the attribute information of the sample image, 1<i≤I; and Each image is processed separately based on each target algorithm matched with the attribute information of each image to obtain a tracking result for the target object.
2. The method according to claim 1, characterized in that The method further comprises: In response to the difference between the attribute information of the i-th image and the attribute information of the i-1-th image being less than a predetermined threshold, the algorithm information matching the attribute information of the i-1-th image is determined as the information of the second target algorithm of the i-th image.
3. The method according to claim 2, characterized in that The target model is obtained by training the initial model based on the attribute information of the sample image and the algorithm information matching the attribute information of the sample image, including: Acquire a plurality of sample images and attribute information corresponding to the plurality of sample images; For each sample image, calling multiple algorithms to process each sample image respectively to obtain multiple tracking results for each sample image; Determine, according to the multiple tracking results and the operation information of the multiple algorithms, multiple effect information of each algorithm processing each sample image, wherein the multiple effect information includes at least one of the following: accuracy information, timeliness information and robustness information; Based on the multiple effect information, determining multiple algorithm evaluation information of each sample image; Determining algorithm information matching the attribute information of each sample image from the plurality of algorithm evaluation information; The initial model is trained using a sample data set to obtain a target model, wherein the sample data set includes multiple groups of sample data, and each group of sample data includes attribute information of a sample image and algorithm information matching the sample image.
4. The method according to claim 3, characterized in that The step of determining a plurality of algorithm evaluation information of each sample image based on the plurality of effect information comprises: Processing the plurality of effect information by using a hierarchical analysis method to obtain a first weight of each piece of effect information for each sample image; Processing the plurality of effect information by using an objective weighting method to obtain a second weight of each piece of effect information for each sample image; Determining a target weight for each piece of effect information based on the first weight and the second weight of the plurality of pieces of effect information; According to the target weights of the various effect information, multiple algorithm evaluation information are obtained.
5. The method according to claim 1, characterized in that The performing image detection on each image respectively to generate attribute information of each image includes: Determine the detection strategy based on the type of each image; Based on the detection strategy, the images are detected to obtain attribute information of each image.
6. The method according to claim 5, characterized in that Determining a detection strategy according to the type of each image includes: In response to the type of the image being a visible light image, determining that the detection strategy is to perform visible light image detection on the image; In response to the type of the image being an infrared image, the detection strategy is determined to perform infrared image detection on the image.
7. The method according to claim 5, characterized in that The attribute information of each image includes at least one of the following: brightness information, contrast information and clarity information; the image is detected based on the detection strategy to obtain the attribute information of each image, including: Performing grayscale detection on the image to obtain first grayscale information of each image; Obtaining brightness information of each image by processing the first grayscale information; Perform contrast enhancement on each image; Performing grayscale detection on the enhanced image to obtain second grayscale information of each image; Obtaining contrast information of each image by processing the second grayscale information; Perform Laplace transform on each image; Performing pixel detection on the transformed image to obtain pixel information of each image; By processing the pixel information, the definition information of each image is determined.
8. A target tracking device based on a task-oriented target tracking algorithm matching strategy, characterized in that: The device comprises: An acquisition module, used to acquire an image sequence for a target object, wherein the image sequence includes I images, where I is an integer greater than 1, and the image sequence includes a visible light image and an infrared image; A detection module, used to perform image detection on each image and generate attribute information of each image; an output module, configured to input the attribute information of the i-th image into a target model in response to the difference between the attribute information of the i-th image and the attribute information of the i-1-th image being greater than or equal to a predetermined threshold, and output information of a first target algorithm for processing the i-th image; wherein the target model is obtained by training an initial model based on the attribute information of the sample image and the algorithm information matching the attribute information of the sample image, 1<i≤I; and The processing module is used to process each image based on each target algorithm matched with the attribute information of each image to obtain the tracking result for the target object.
9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Low-illumination scene multi-mode pedestrian detection tracking method based on decision-making level fusion
CN117636241A
Smart home detection and management method and device and computer equipment
CN118918535A