Image tracking method, device and system and computer readable storage medium

By using the graph heat source information and pre-trained models in the image tracking system, the contour information of the target image is determined, and the problem of insufficient accuracy of image tracking in the prior art is solved, and higher image tracking accuracy is achieved.

CN119991714APending Publication Date: 2025-05-13TCL TECHNOLOGY GROUP CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311494272.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing image tracking methods are susceptible to factors such as noise and light changes in complex environments, resulting in poor selection of image contour information and affecting the tracking effect.

Method used

By acquiring the trigger operation position information of the target image, the graph heat source information is generated, and the target image and graph heat source information are input to the pre-trained image tracking model to determine the image outline information corresponding to the position information.

Benefits of technology

Improve the accuracy of image tracking and avoid the problem of poor selection of image contour information caused by noise and lighting changes in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991714A_ABST
    Figure CN119991714A_ABST
Patent Text Reader

Abstract

The invention discloses an image tracking method, device and system and a computer readable storage medium, and the method comprises the steps: obtaining the position information of a triggering operation in response to the triggering operation of a target image; processing the target image according to the position information, and generating image heat source information; and inputting the target image and the image heat source information into an image tracking model, and determining image contour information corresponding to the position information. Through the pre-trained image tracking model, the image contour information of the corresponding image is determined and positioned according to the target image, the image heat source information corresponding to the target image and the position information in the target image, so that the influence of factors such as noise and illumination change in a complex environment is avoided; therefore, the image contour information selection of the image is not good, and the image tracking accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an image tracking method, device, system and computer-readable storage medium. Background Art

[0002] With the increasing demand for taking photos and selfies, as well as the continuous optimization and expansion of camera functions, manual tracking has become a very urgent need. However, existing image tracking methods mainly rely on traditional image processing technologies, such as edge detection and corner detection. These methods are easily affected by factors such as noise and lighting changes in complex environments, and often encounter situations where images are difficult to select and image positioning is not ideal, which in turn leads to poor selection of image contour information boundaries, which directly affects the tracking effect and ultimately affects the user experience. Therefore, how to improve the accuracy of image tracking is an urgent problem to be solved. Summary of the invention

[0003] The embodiments of the present disclosure provide an image tracking method, device, system and computer-readable storage medium, which can improve the accuracy of image tracking.

[0004] In a first aspect, an embodiment of the present disclosure provides an image tracking method, the method comprising:

[0005] In response to a trigger operation on a target image, acquiring position information of the trigger operation;

[0006] Process the target image according to the position information to generate image heat source information;

[0007] The target image and the image heat source information are input into an image tracking model to determine image contour information corresponding to the position information.

[0008] In a second aspect, an embodiment of the present disclosure provides an image tracking device comprising:

[0009] an acquisition unit, configured to acquire position information of a trigger operation in response to a trigger operation on a target image;

[0010] A generating unit, used for processing the target image according to the position information to generate image heat source information;

[0011] A determination unit is used to input the target image and the image heat source information into an image tracking model to determine the image contour information corresponding to the position information.

[0012] In a third aspect, an embodiment of the present disclosure further provides an image tracking system, including a memory storing a plurality of instructions; a processor loads the instructions from the memory to execute the steps of any one of the image tracking methods provided in the embodiment of the present disclosure.

[0013] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps of any one of the image tracking methods provided in the embodiments of the present disclosure.

[0014] In a fifth aspect, the embodiments of the present disclosure further provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of any one of the image tracking methods provided in the embodiments of the present disclosure.

[0015] The scheme of the application embodiment is adopted, in response to the trigger operation of the target image, the position information of the trigger operation is obtained; the target image is processed according to the position information to generate image heat source information; the target image and the image heat source information are input into the image tracking model to determine the image contour information corresponding to the position information. Through the pre-trained image tracking model, the image contour information of the corresponding image is determined and located according to the position information in the target image according to the target image and the image heat source information corresponding to the target image, so as to avoid the situation in which the image contour information of the image is poorly selected due to the influence of factors such as noise and illumination changes in complex environments, thereby improving the accuracy of image tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 is a schematic flow chart of a first embodiment of an image tracking method provided in an embodiment of the present disclosure;

[0018] Figure 2 is a schematic diagram of a sample image or a target image provided in an embodiment of the present disclosure;

[0019] Figure 3 is a schematic diagram of sample image heat source information or target image heat source information provided in an embodiment of the present disclosure;

[0020] Figure 4 is a flowchart of a second embodiment of the image tracking method provided in the embodiments of the present disclosure;

[0021] Figure 5 is a schematic diagram of the structure of an image tracking model provided in an embodiment of the present disclosure;

[0022] Figure 6is a schematic diagram of the structure of an image tracking device provided in an embodiment of the present disclosure;

[0023] Figure 7 It is a structural schematic diagram of the image tracking system provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure. At the same time, in the description of the embodiments of the present disclosure, the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present disclosure, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0025] Embodiments of the present disclosure provide an image tracking method, device, system, and computer-readable storage medium.

[0026] Specifically, the embodiments of the present disclosure will be described from the perspective of an image tracking system. The image tracking system can be integrated into an image tracking device, that is, the image tracking method of the embodiments of the present disclosure can be executed by the image tracking system.

[0027] The image tracking method provided by the embodiments of the present disclosure may be applied to an image tracking device, which may include, for example, a smart phone, a camera, and various devices with tracking requirements.

[0028] The following is a detailed description in conjunction with the accompanying drawings. In the embodiments of the present disclosure, the execution subject is an image tracking system as an example. It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in an order different from that shown in the drawings.

[0029] Please refer to Figure 1 The specific process of the first embodiment of the image tracking method includes the following steps:

[0030] Step 101, in response to a trigger operation on a target image, obtaining position information of the trigger operation;

[0031] Step 102, processing the target image according to the position information to generate image heat source information;

[0032] Step 103: input the target image and the image heat source information into an image tracking model to determine the image contour information corresponding to the position information.

[0033] In this embodiment, when the user needs to track an image, the image is displayed in the viewfinder of a device with a shooting function, such as a smart phone or a camera, to form a target image, and the screen of the device is clicked. The image tracking system responds to the user's trigger operation on the target image, obtains the location information of the trigger operation, processes the target image according to the location information, and generates image heat source information corresponding to the target image; the image tracking system inputs the target image and the image heat source information into the image tracking model, determines the contour information of the image corresponding to the location information on the target image, and thereby facilitates the subsequent tracking of the image in the contour information by devices with a shooting function, such as smart phones or cameras. It should be noted that the target image is an image that contains the image to be tracked; the target image will be displayed on the screen of the device, and the user clicks on the position of the image to be tracked on the screen of the device, and the image tracking system will obtain the position information, which is the coordinates of the click on the screen of the user's device; the shape of the image contour information of the image to be tracked can be the edge contour of the image, or it can be a rectangular frame, circular frame, etc. that can completely frame the image to be tracked, and there is no limitation here; the graph heat source information corresponding to the target image refers to the conversion of nodes or edges in the graph structure into a low-dimensional continuous vector representation. In this application, the graph heat source information is expressed in the form of a heat map.

[0034] The image tracking system of this embodiment responds to a trigger operation on a target image, obtains the location information of the trigger operation; processes the target image according to the location information to generate image heat source information; inputs the target image and the image heat source information into the image tracking model to determine the image contour information corresponding to the location information. Through the pre-trained image tracking model, the image contour information of the corresponding image is determined and located according to the location information in the target image according to the target image and the image heat source information corresponding to the target image, thereby avoiding the situation in which the image contour information of the image is poorly selected due to factors such as noise and lighting changes in complex environments, thereby improving the accuracy of image tracking.

[0035] Specifically, each step is described in detail below:

[0036] Step 101, in response to a trigger operation on a target image, obtaining position information of the trigger operation;

[0037] In this step, when the user needs to track an image, the image to be tracked is displayed in the viewfinder of a device with a shooting function such as a smart phone or a camera, and the user clicks the position of the image to be tracked on the screen of the device to complete the triggering operation of the target image; the image tracking system responds to the user's triggering operation on the target image and obtains the position information of the triggering operation; for example, Figure 2 As shown, Figure 2 is a schematic diagram of the target image, where the “STOP” sign is the image that the user needs to track. The user uses the camera function of the smartphone to display the “STOP” sign in the viewfinder of the smartphone, forming a Figure 2 The target image shown in the figure, at this time, the user clicks at any position on the "STOP" sign in the target image displayed on the smartphone screen, and the image tracking system determines the position where the user clicked on the smartphone screen as the position information.

[0038] Step 102, processing the target image according to the position information to generate image heat source information;

[0039] In this step, after determining the position information corresponding to the trigger operation on the target image, the image tracking system processes the target image according to the position information to generate image heat source information; specifically, the image tracking system generates a single-channel heat map of the image heat source information corresponding to the target image based on the position information on the target image and the length and width of the target image. It can be understood that the target image is an image with three RGB channels, and the heat map is a single-channel image. For example, Figure 3 As shown, Figure 3 is the heat map corresponding to the target image, where Figure 3 The white area in is the position information corresponding to the trigger operation on the target image.

[0040] Step 103: input the target image and the image heat source information into an image tracking model to determine the image contour information corresponding to the position information.

[0041] In this step, after obtaining the target image and the image heat source information corresponding to the target image, the image tracking system inputs the target image and the image heat source information into the image tracking model, determines the image corresponding to the position information according to the output of the image tracking model, and then determines the image contour information corresponding to the image.

[0042] It should be noted that the output of the image tracking model consists of two branches, of which branch 1 is composed of an average pooling layer and a fully connected layer, and branch 2 is composed of a 3x3 convolutional layer and a 1x1 convolutional layer; the image tracking model uses the lightweight network Touch Object Detection Network (TODNet) as the backbone network, which can speed up the reasoning of the model and reduce the number of network parameters, so that the model can run smoothly on a low-computing power platform; in order to further reduce the physical space occupied by the model and the reasoning time, the number of output channels of branch 1 of the backbone network Touch Object Detection Network (TODNet) is changed from the original 1024 to the current 256, so that the size of the model is only 756Kb (8-bit effect), and the accuracy of the model has basically not changed.

[0043] Specifically, step 103 includes:

[0044] Step 1031, inputting the target image and the image heat source information into an image tracking model, and judging whether there is an image at the position information according to a first prediction value output by the image tracking model;

[0045] In this step, the image tracking system inputs the target image and image heat source information into the image tracking model, and determines whether there is an image at the location information based on the first prediction value output by the image tracking model; specifically, the output of the image tracking model consists of two branches, where branch 1 consists of an average pooling layer and a fully connected layer, and branch 1 outputs the first prediction value. The image tracking system compares the first prediction value with the first preset threshold to further determine whether there is an image at the location information.

[0046] Further, step 1031 includes:

[0047] Step 10311, comparing the first predicted value with a first preset threshold;

[0048] In this step, the image tracking system compares the first prediction value output by the image tracking model with the first preset threshold; optionally, the first prediction value is a 1*2 tensor, that is, the first prediction value includes first data and second data, and the image tracking system compares the first data and the second data with the second preset threshold respectively; optionally, the first prediction value is a 1*1 tensor, that is, the first prediction value only includes one data, and the image tracking system compares the first prediction value with the first preset threshold.

[0049] Step 10312: if the first prediction value is greater than the first preset threshold, it is determined that an image exists at the location information;

[0050] In this step, the image tracking system compares the first prediction value output by the image tracking model with the first preset threshold value. If it is determined that the first prediction value is greater than the first preset threshold value, it is determined that an image exists at the location information.

[0051] Optionally, the first prediction value is a 1*2 tensor, that is, the first prediction value includes first data and second data, the first data is used to indicate whether there is an image at the position information and whether the image is located in the foreground of the target image, and the second data is used to indicate whether there is an image at the position information and whether the image is located in the background of the target image. If the first data is greater than the first preset threshold, it is determined that there is an image at the position information and the image is located in the foreground of the target image. If the second data is greater than the first preset threshold, it is determined that there is an image at the position information and the image is located in the background of the target image. Further, since the sum of the first data and the second data is equal to 1, the first preset threshold is usually 0.5. Therefore, when the first data is greater than 0.5, it is not necessary to compare the second data with the first preset threshold. When the first data is not greater than 0.5, it is necessary to compare the second data with the first preset threshold.

[0052] Optionally, the first prediction value is a 1*1 tensor, that is, the first prediction value only includes one data, and the first prediction value is used to indicate whether there is an image at the position information. If the first prediction value is greater than a first preset threshold, it is determined that there is an image at the position information.

[0053] Step 10313: If the first prediction value is not greater than the first preset threshold, it is determined that no image exists at the location information.

[0054] In this step, the image tracking system compares the first prediction value output by the image tracking model with the first preset threshold value. If it is determined that the first prediction value is not greater than the first preset threshold value, it is determined that there is no image at the location information.

[0055] Optionally, the first prediction value is a 1*2 tensor, that is, the first prediction value includes first data and second data. The image tracking system first compares the first data with a first preset threshold. If the first data is not greater than the first preset threshold, it is determined that no image exists at the position information. Then, the second data is compared with the second preset threshold. If the second data is not greater than the first preset threshold, it is determined that no image exists at the position information; that is, it can be determined that no image exists at the position information, whether in the foreground or the background.

[0056] Optionally, the first prediction value is a 1*1 tensor, that is, the first prediction value only includes one data. If the first prediction value is not greater than a first preset threshold, it is determined that there is no image at the position information.

[0057] Step 1032: if it exists, determine the position information of the image in the target image according to the second prediction value output by the image tracking model, and determine the image contour information corresponding to the image according to the position information of the image.

[0058] In this step, if the image tracking system determines that there is an image at the position information, the position information of the image is determined in the target image according to the second prediction value output by the image tracking model, and the image contour information corresponding to the image is determined according to the position information of the image. It can be understood that the image tracking model will output the second prediction value only when the image tracking system determines that there is an image at the position information, and the second prediction value is used to indicate the position information of the image to be tracked in the target image and the image contour information of the image to be tracked in the target image.

[0059] Specifically, step 1032 includes:

[0060] Step 10321, determining the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value, and the vertical coordinate offset prediction value in the second prediction value;

[0061] In this step, the image tracking system determines the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value and the vertical coordinate offset prediction value in the second prediction value; it should be noted that the preset downsampling factor is set in advance in the image tracking model. After receiving the input target image and image heat source information, the image tracking model performs feature fusion on the target image and image heat source information, and performs downsampling, convolution and other operations based on the preset downsampling factor, and finally outputs the second prediction value; the second prediction value includes the image area prediction value, the horizontal coordinate offset prediction value, the vertical coordinate offset prediction value, the image width prediction value and the image height prediction value.

[0062] Specifically, the second prediction value is a 5*52*52 tensor, representing keypoints heatmap, x_offset, y_offset, w and h respectively, where keypoints heatmap is the image area prediction value, keypointsheatmap is also called key point graph heat source information, which is a feature map obtained by the image tracking model by performing feature fusion on the target image and graph heat source information, and performing downsampling, convolution and other operations based on a preset downsampling factor. The feature map includes the feature value of the image corresponding to the position information in the target image; x_offset is the horizontal axis offset prediction value; y_offset is the vertical axis offset prediction value; w is the image width prediction value; h is the image height prediction value.

[0063] Further, step 10321 includes:

[0064] Step 103211, determining the predicted position information of the image in the target image according to the preset downsampling factor and the image area prediction value;

[0065] In this step, the image tracking system determines the predicted position information of the image in the target image according to the preset downsampling factor and the image area prediction value; specifically, the image tracking system first determines the area composed of pixels representing the image in the keypoints heatmap according to the feature value corresponding to each pixel in the keypoints heatmap in the image area prediction value, and then determines the coordinate value corresponding to each pixel in the area, and then multiplies the coordinate value corresponding to each pixel in the area by the preset downsampling factor, so that the coordinate value corresponding to each pixel in the area is restored to the target image, and then determines the predicted position information of the image in the target image.

[0066] Step 103212, correcting the coordinate value corresponding to each pixel point occupied by the predicted position information of the image according to the horizontal coordinate offset prediction value and the vertical coordinate offset prediction value, and determining the position information of the image in the target image.

[0067] In this step, the image tracking system corrects the coordinate value corresponding to each pixel point occupied by the predicted position information of the image according to the horizontal coordinate offset prediction value and the vertical coordinate offset prediction value, and determines the position information of the image in the target image; specifically, after restoring the coordinate value corresponding to each pixel point in the area composed of the pixel points representing the image in the keypoints heatmap to the target image to obtain the predicted position information of the image in the target image, the image tracking system adds the horizontal coordinate offset prediction value to the horizontal coordinate of the coordinate value corresponding to each pixel point in the predicted position information, and adds the vertical coordinate offset prediction value to the vertical coordinate of the coordinate value corresponding to each pixel point, and then corrects the coordinate value corresponding to each pixel point occupied by the predicted position information of the image, and determines the position information of the image in the target image.

[0068] Step 10322: Determine image contour information corresponding to the image based on the position information of the image and the image width prediction value and the image height prediction value in the second prediction value.

[0069] In this step, the image tracking system determines the image contour information corresponding to the image based on the position information of the image in the target image, and the image width prediction value and the image height prediction value in the second prediction value. Specifically, the image tracking system first determines the center point coordinates corresponding to the image based on the position information of the image in the target image, adds w / 2 (i.e., half of the image width prediction value) on both sides of the horizontal coordinates of the center point coordinates to determine the width of the contour information corresponding to the image, adds h / 2 (i.e., half of the image height prediction value) on both sides of the vertical coordinates of the center point coordinates to determine the height of the contour information corresponding to the image, and then determines the image contour information corresponding to the image on the target image.

[0070] The image tracking system of this embodiment responds to the trigger operation of the target image, obtains the location information of the trigger operation; processes the target image according to the location information to generate image heat source information; inputs the target image and the image heat source information into the image tracking model to determine the image contour information corresponding to the location information. Through the pre-trained image tracking model, the image contour information of the corresponding image is determined and located according to the target image and the image heat source information corresponding to the target image, avoiding the situation in which the image contour information of the image is poorly selected due to factors such as noise and illumination changes in complex environments, thereby improving the accuracy of image tracking; and the image tracking model uses a lightweight network as the backbone network, which can speed up the reasoning of the model and reduce the number of network parameters, so that the model can run smoothly on a low computing power platform, and the output channels of the image tracking model are reduced, further reducing the physical space occupied by the model and the reasoning time.

[0071] refer to Figure 4 In combination with the first embodiment, a second embodiment of the image tracking method is proposed. The specific process of the second embodiment of the image tracking method includes the following steps:

[0072] Step a, obtaining a sample image and sample graph heat source information of the sample image, wherein the sample image is provided with annotation data, wherein the annotation data includes position information and image contour information of the image corresponding to the position information, and the sample graph heat source information is generated by processing the sample image according to the position information;

[0073] In this step, the image tracking system obtains the sample image and the sample graph heat source information of the sample image. The sample image is provided with annotation data, which includes position information and image contour information of the image corresponding to the position information. The sample graph heat source information is generated by processing the sample image according to the position information. Specifically, the image tracking system can obtain a large number of sample images from a device with a shooting function or from an open source data set, and then randomly generate corresponding position information on each sample image based on the length and width of each sample image, generate the single-channel sample graph heat source information corresponding to each sample image according to the corresponding position information on each sample image and the length and width of each sample image, and then annotate the image contour information of the image corresponding to the position information on the sample image. If there are multiple images corresponding to the position information, the image with the smallest area is annotated; if there is no corresponding image for the position information, it is not annotated. The annotated sample image stores the coordinate information of the position information and the image contour information of the image. Next, the portrait segmentation deep learning model is used to filter out the contour information with an area less than 400pix, and the remaining sample images and sample graph heat source information are used as training data.

[0074] Step b, counting a first number of first sample images in which the image is located in the foreground, and a second number of second sample images in which the image is located in the background;

[0075] In this step, the image tracking system counts the first number of first sample images in the training data where the image is located in the foreground, and the second number of second sample images where the image is located in the background. Figure 2 As shown, the "STOP" sign belongs to the foreground of the image, while the road, mountain, and sky belong to the background of the image.

[0076] Step c, generating a target loss function based on the first number, the second number and an original loss function of a pre-created image tracking network;

[0077] In this step, the image tracking system generates a target loss function based on the first quantity, the second quantity, and the original loss function of the pre-created image tracking network; specifically, the original loss function of the pre-created image tracking network is the cross-entropy loss function. According to the generation process of the aforementioned training data set, the number of training data for images in the foreground and images in the background is usually not equal. The category with a larger data volume accounts for a larger proportion in the loss function, and the model tends to accurately predict the category, resulting in a larger prediction error rate for other categories. In order to reduce this effect, a weight factor of the inverse of the statistical quantity is introduced, as shown in the following formula:

[0078]

[0079] Among them, x i Indicates the number of training images of the i-th category. In this solution, i=0 or 1, 0 represents background (or foreground), and 1 represents foreground (or background).

[0080] The image tracking system determines the first weight factor of the sample image where the image is located in the foreground and the second weight factor of the sample image where the image is located in the background according to the above formula, and substitutes the first weight factor and the second weight factor into the cross-entropy loss function (cross-entropy loss) of the original loss function of the pre-created image tracking network to obtain the final model training loss function: L = -1 / N*Σ(Σ(y-true*log(y-pred)))*w0*w1, where N is the number of sample categories, in this scheme N = 2, y-true is the true label value, y-pred is the predicted value, w0 is the first weight factor of the sample image where the image is located in the foreground, and w1 is the second weight factor of the sample image where the image is located in the background.

[0081] Step d: training the image tracking network based on the sample image, the heat source information of the sample image and the loss function to obtain an image tracking model.

[0082] In this step, the image tracking system trains the image tracking network based on the sample image, the sample image heat source information and the loss function to obtain the image tracking model. Specifically, Figure 5 As shown in the figure, the image tracking network is a dual-input network, which is a lightweight network Touch Object Detection Network (TODNet). The input consists of a sample image and the sample image heat source information corresponding to the sample image. The sample image heat source information and the sample image are preprocessed and concat together, and convolution, downsampling, reverse residual and other operations are performed. The obtained data are output to two output branches respectively. One output branch is composed of a 3x3 convolution and a 1x1 convolution, and the other output branch is composed of an average pooling layer and a fully connected layer.

[0083] Specifically, step d comprises:

[0084] Step d1, inputting the sample image and the sample image heat source information into the image tracking network to obtain the prediction result of the image tracking network;

[0085] In this step, the image tracking system selects a sample image and the sample image heat source information corresponding to the sample image in the training data set in turn, and inputs the sample image and the sample image heat source information into the image tracking network to obtain the prediction result of the image tracking network. Specifically, the image tracking network first outputs the prediction result corresponding to the output branch composed of an average pooling layer and a fully connected layer, and then outputs the prediction result corresponding to the output branch composed of a 3x3 convolution and a 1x1 convolution.

[0086] Step d2, obtaining a loss value based on the prediction result, the labeled data of the sample image and the loss function;

[0087] In this step, the image tracking system determines the final prediction value based on the two prediction results output by the image tracking network, determines the true label value according to the annotation data of the input sample image, and substitutes the final prediction value and the true label value into the loss function to obtain the loss value.

[0088] Step d3, transferring the loss value along each layer of the image tracking network to update the parameters of the image tracking network;

[0089] In this step, the image tracking system transmits the obtained loss value along each layer of the image tracking network to update the parameters of the image tracking network.

[0090] Step d4, until the obtained loss value is less than the first preset threshold, an image tracking model is obtained.

[0091] In this step, after the image tracking system updates the parameters of the image tracking network, it continues to execute the above steps d1 to d3 until the loss value obtained is less than the first preset threshold, and the image tracking model is obtained. Alternatively, until the loss values ​​obtained by inputting the sample image and the sample image heat source information for a preset number of times are all less than the first preset threshold, and the image tracking model is obtained.

[0092] The image tracking system of this embodiment obtains a sample image and sample graph heat source information of the sample image. The sample image is provided with annotation data, which includes position information and image contour information of the image corresponding to the position information. The sample graph heat source information is generated by processing the sample image according to the position information; a first number of first sample images whose images are located in the foreground and a second number of second sample images whose images are located in the background are counted; a target loss function is generated based on the first number, the second number and the original loss function of the pre-created image tracking network; the image tracking network is trained based on the sample image, the sample graph heat source information and the loss function to obtain an image tracking model. The image tracking model is used to learn the relationship between the position information and the image contour information of the image on the position information, so as to avoid the situation in which the image contour information of the image is poorly selected due to factors such as noise and illumination changes in complex environments, thereby achieving accurate positioning of the image contour information of the image.

[0093] This embodiment also provides an image tracking device, which can be integrated into a smart phone, a camera, and various devices with tracking requirements, such as Figure 6 As shown, the image tracking device may include:

[0094] An acquisition unit 1001 is used for acquiring position information of a trigger operation in response to a trigger operation on a target image;

[0095] A generating unit 1002 is used to process the target image according to the position information to generate image heat source information;

[0096] The determination unit 1003 is used to input the target image and the image heat source information into the image tracking model to determine the image contour information corresponding to the position information.

[0097] In an optional example, the determining unit is further configured to:

[0098] Inputting the target image and the image heat source information into an image tracking model, and judging whether there is an image at the position information according to a first prediction value output by the image tracking model;

[0099] If it exists, the position information of the image is determined in the target image according to the second prediction value output by the image tracking model, and the image contour information corresponding to the image is determined according to the position information of the image.

[0100] In an optional example, the determining unit is further configured to:

[0101] Comparing the first predicted value with a first preset threshold;

[0102] If the first prediction value is greater than the first preset threshold, it is determined that an image exists at the location information;

[0103] If the first prediction value is not greater than the first preset threshold, it is determined that no image exists at the location information.

[0104] In an optional example, the determining unit is further configured to:

[0105] Determine the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value, and the vertical coordinate offset prediction value in the second prediction value;

[0106] Image contour information corresponding to the image is determined according to the position information of the image, and the image width prediction value and the image height prediction value in the second prediction value.

[0107] In an optional example, the determining unit is further configured to:

[0108] Determining predicted position information of the image in the target image according to the preset downsampling factor and the image area prediction value;

[0109] The coordinate value corresponding to each pixel point occupied by the predicted position information of the image is corrected according to the horizontal coordinate offset prediction value and the vertical coordinate offset prediction value, and the position information of the image is determined in the target image.

[0110] In an optional example, the acquisition unit further includes a training unit, and the training unit is used to:

[0111] Acquire a sample image and sample graph heat source information of the sample image, wherein the sample image is provided with annotation data, wherein the annotation data includes position information and image contour information of the image corresponding to the position information, and the sample graph heat source information is generated by processing the sample image according to the position information;

[0112] Counting a first number of first sample images in which the image is located in the foreground, and a second number of second sample images in which the image is located in the background;

[0113] generating a target loss function based on the first quantity, the second quantity, and an original loss function of a pre-created image tracking network;

[0114] The image tracking network is trained based on the sample image, the heat source information of the sample image and the loss function to obtain an image tracking model.

[0115] In an optional example, the training unit is further configured to:

[0116] Inputting the sample image and the sample image heat source information into the image tracking network to obtain a prediction result of the image tracking network;

[0117] Obtaining a loss value based on the prediction result, the labeled data of the sample image and the loss function;

[0118] Passing the loss value along each layer of the image tracking network to update the parameters of the image tracking network;

[0119] Until the obtained loss value is less than the second preset threshold, the image tracking model is obtained.

[0120] According to the scheme of this embodiment, in response to a trigger operation on a target image, the position information of the trigger operation is obtained; the target image is processed according to the position information to generate image heat source information; the target image and the image heat source information are input into the image tracking model to determine the image contour information corresponding to the position information. Through the pre-trained image tracking model, the image contour information of the corresponding image is determined and located according to the position information in the target image according to the target image and the image heat source information corresponding to the target image, so as to avoid the situation in which the image contour information of the image is poorly selected due to factors such as noise and illumination changes in complex environments, thereby improving the accuracy of image tracking.

[0121] Accordingly, the present disclosure also provides an image tracking system, such as Figure 7 As shown, Figure 7 A schematic diagram of the structure of an image tracking system provided in an embodiment of the present disclosure. The image tracking system 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. Those skilled in the art will appreciate that the image tracking system structure shown in the figure does not constitute a limitation on the image tracking system, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0122] The processor 1101 is the control center of the image tracking system 1100, and uses various interfaces and lines to connect various parts of the entire electronic device 1100. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, the processor 1101 executes various functions of the electronic device 1100 and processes data, thereby monitoring the image tracking system 1100 as a whole. The processor 1101 can be a processor CPU, a graphics processor GPU, a network processor (Network Processor, NP), etc., and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure.

[0123] In the embodiment of the present disclosure, the processor 1101 in the image tracking system 1100 will load instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions, such as:

[0124] In response to a trigger operation on a target image, acquiring position information of the trigger operation;

[0125] Process the target image according to the position information to generate image heat source information;

[0126] The target image and the image heat source information are input into an image tracking model to determine image contour information corresponding to the position information.

[0127] In an optional example, it also includes:

[0128] Inputting the target image and the image heat source information into an image tracking model, and judging whether there is an image at the position information according to a first prediction value output by the image tracking model;

[0129] If it exists, the position information of the image is determined in the target image according to the second prediction value output by the image tracking model, and the image contour information corresponding to the image is determined according to the position information of the image.

[0130] In an optional example, it also includes:

[0131] Comparing the first predicted value with a first preset threshold;

[0132] If the first prediction value is greater than the first preset threshold, it is determined that an image exists at the location information;

[0133] If the first prediction value is not greater than the first preset threshold, it is determined that no image exists at the location information.

[0134] In an optional example, it also includes:

[0135] Determine the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value, and the vertical coordinate offset prediction value in the second prediction value;

[0136] Image contour information corresponding to the image is determined according to the position information of the image, and the image width prediction value and the image height prediction value in the second prediction value.

[0137] In an optional example, it also includes:

[0138] Determining predicted position information of the image in the target image according to the preset downsampling factor and the image area prediction value;

[0139] The coordinate value corresponding to each pixel point occupied by the predicted position information of the image is corrected according to the predicted value of the horizontal coordinate offset and the predicted value of the vertical coordinate offset, and the position information of the image is determined in the target image.

[0140] In an optional example, it also includes:

[0141] Acquire a sample image and sample graph heat source information of the sample image, wherein the sample image is provided with annotation data, wherein the annotation data includes position information and image contour information of the image corresponding to the position information, and the sample graph heat source information is generated by processing the sample image according to the position information;

[0142] Counting a first number of first sample images in which the image is located in the foreground, and a second number of second sample images in which the image is located in the background;

[0143] generating a target loss function based on the first quantity, the second quantity, and an original loss function of a pre-created image tracking network;

[0144] The image tracking network is trained based on the sample image, the sample image heat source information and the loss function to obtain an image tracking model.

[0145] In an optional example, it also includes:

[0146] Inputting the sample image and the sample image heat source information into the image tracking network to obtain a prediction result of the image tracking network;

[0147] Obtaining a loss value based on the prediction result, the labeled data of the sample image and the loss function;

[0148] Passing the loss value along each layer of the image tracking network to update the parameters of the image tracking network;

[0149] Until the obtained loss value is less than the second preset threshold, the image tracking model is obtained.

[0150] Thus, the accuracy of image tracking can be improved.

[0151] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0152] Optional, such as Figure 7 As shown, the image tracking system 1100 further includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107, respectively. Those skilled in the art will appreciate that Figure 7 The image tracking system structure shown in the figure does not constitute a limitation on the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0153] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the electronic device, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light-emitting diode (OLED, Organic Light-EmittingDiode) and the like. The touch panel can be used to collect the user's touch operation on or near it (such as the user uses any suitable object or attachment such as a finger, a stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 1101, and can receive the command sent by the processor 1101 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event, and then the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiment of the present disclosure, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to realize the input function.

[0154] The radio frequency circuit 1104 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other electronic devices through wireless communication, and to send and receive signals between the network device or other electronic devices.

[0155] The audio circuit 1105 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and converted into audio data, and then the audio data is output to the processor 1101 for processing, and then sent to another electronic device through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earplug jack to provide communication between an external headset and an electronic device.

[0156] The input unit 1106 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0157] The power supply 1107 is used to supply power to various components of the electronic device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so that the power management system can manage charging, discharging, and power consumption. The power supply 1107 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0158] although Figure 7 Not shown, the electronic device 1100 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0159] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0160] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0161] To this end, the embodiment of the present disclosure provides a computer-readable storage medium, in which a plurality of computer programs are stored, and the computer program can be loaded by a processor to execute any one of the image tracking methods provided in the embodiment of the present disclosure. The computer program can execute the steps of the following image tracking method:

[0162] In response to a trigger operation on a target image, acquiring position information of the trigger operation;

[0163] Process the target image according to the position information to generate image heat source information;

[0164] The target image and the image heat source information are input into an image tracking model to determine image contour information corresponding to the position information.

[0165] In an optional example, it also includes:

[0166] Inputting the target image and the image heat source information into an image tracking model, and judging whether there is an image at the position information according to a first prediction value output by the image tracking model;

[0167] If it exists, the position information of the image is determined in the target image according to the second prediction value output by the image tracking model, and the image contour information corresponding to the image is determined according to the position information of the image.

[0168] In an optional example, it also includes:

[0169] Comparing the first predicted value with a first preset threshold;

[0170] If the first prediction value is greater than the first preset threshold, it is determined that an image exists at the location information;

[0171] If the first prediction value is not greater than the first preset threshold, it is determined that no image exists at the location information.

[0172] In an optional example, it also includes:

[0173] Determine the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value, and the vertical coordinate offset prediction value in the second prediction value;

[0174] Image contour information corresponding to the image is determined according to the position information of the image, and the image width prediction value and the image height prediction value in the second prediction value.

[0175] In an optional example, it also includes:

[0176] Determining predicted position information of the image in the target image according to the preset downsampling factor and the image area prediction value;

[0177] The coordinate value corresponding to each pixel point occupied by the predicted position information of the image is corrected according to the predicted value of the horizontal coordinate offset and the predicted value of the vertical coordinate offset, and the position information of the image is determined in the target image.

[0178] In an optional example, it also includes:

[0179] Acquire a sample image and sample graph heat source information of the sample image, wherein the sample image is provided with annotation data, wherein the annotation data includes position information and image contour information of the image corresponding to the position information, and the sample graph heat source information is generated by processing the sample image according to the position information;

[0180] Counting a first number of first sample images in which the image is located in the foreground, and a second number of second sample images in which the image is located in the background;

[0181] generating a target loss function based on the first quantity, the second quantity, and an original loss function of a pre-created image tracking network;

[0182] The image tracking network is trained based on the sample image, the heat source information of the sample image and the loss function to obtain an image tracking model.

[0183] In an optional example, it also includes:

[0184] Inputting the sample image and the sample image heat source information into the image tracking network to obtain a prediction result of the image tracking network;

[0185] Obtaining a loss value based on the prediction result, the labeled data of the sample image and the loss function;

[0186] Passing the loss value along each layer of the image tracking network to update the parameters of the image tracking network;

[0187] Until the obtained loss value is less than the second preset threshold, the image tracking model is obtained.

[0188] Thus, the accuracy of image tracking can be improved.

[0189] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0190] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0191] Since the computer program stored in the computer-readable storage medium can execute any one of the image tracking methods provided in the embodiments of the present disclosure, the beneficial effects that can be achieved by any one of the image tracking methods provided in the embodiments of the present disclosure can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0192] According to one aspect of the present disclosure, a computer program product or a computer program is also provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in various optional implementations in the above-mentioned embodiments.

[0193] In the above-mentioned image tracking device, computer-readable storage medium, image tracking system, and computer program product embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and beneficial effects of the above-mentioned image tracking device, computer-readable storage medium, computer program product, image tracking system, and corresponding units can refer to the description of the image tracking method in the above embodiment, and will not be repeated here.

[0194] The above is a detailed introduction to an image tracking method, device, system, computer-readable storage medium and computer program product provided by the embodiments of the present disclosure. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core idea. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

Claims

1. An image tracking method, characterized in that: The image tracking method comprises: In response to a trigger operation on a target image, acquiring position information of the trigger operation; Process the target image according to the position information to generate image heat source information; The target image and the image heat source information are input into an image tracking model to determine image contour information corresponding to the position information.

2. The image tracking method according to claim 1, characterized in that: The step of inputting the target image and the image heat source information into an image tracking model to determine the image contour information corresponding to the position information includes: Inputting the target image and the image heat source information into an image tracking model, and judging whether there is an image at the position information according to a first prediction value output by the image tracking model; If it exists, the position information of the image is determined in the target image according to the second prediction value output by the image tracking model, and the image contour information corresponding to the image is determined according to the position information of the image.

3. The image tracking method according to claim 2, characterized in that: The step of judging whether there is an image at the location information according to the first prediction value output by the image tracking model comprises: Comparing the first predicted value with a first preset threshold; If the first prediction value is greater than the first preset threshold, it is determined that an image exists at the location information; If the first prediction value is not greater than the first preset threshold, it is determined that no image exists at the location information.

4. The image tracking method according to claim 2, characterized in that: The method of determining the position information of the image in the target image according to the second prediction value output by the image tracking model, and determining the image contour information corresponding to the image according to the position information of the image, comprises: Determine the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value, and the vertical coordinate offset prediction value in the second prediction value; Image contour information corresponding to the image is determined according to the position information of the image, and the image width prediction value and the image height prediction value in the second prediction value.

5. The image tracking method according to claim 4, characterized in that: Determining the position information of the image in the target image according to the preset downsampling factor, and the image area prediction value, the horizontal coordinate offset prediction value, and the vertical coordinate offset prediction value in the second prediction value includes: Determining predicted position information of the image in the target image according to the preset downsampling factor and the image area prediction value; The coordinate value corresponding to each pixel point occupied by the predicted position information of the image is corrected according to the predicted value of the horizontal coordinate offset and the predicted value of the vertical coordinate offset, and the position information of the image is determined in the target image.

6. The image tracking method according to claim 1, characterized in that: Before responding to the user interaction operation and obtaining the location information of the triggering operation, the method includes: Acquire a sample image and sample graph heat source information of the sample image, wherein the sample image is provided with annotation data, wherein the annotation data includes position information and image contour information of the image corresponding to the position information, and the sample graph heat source information is generated by processing the sample image according to the position information; Counting a first number of first sample images in which the image is located in the foreground, and a second number of second sample images in which the image is located in the background; generating a target loss function based on the first quantity, the second quantity, and an original loss function of a pre-created image tracking network; The image tracking network is trained based on the sample image, the heat source information of the sample image and the loss function to obtain an image tracking model.

7. The image tracking method according to claim 6, characterized in that: The image tracking network is trained based on the sample image, the sample image heat source information and the loss function to obtain an image tracking model, including: Inputting the sample image and the sample image heat source information into the image tracking network to obtain a prediction result of the image tracking network; Obtaining a loss value based on the prediction result, the labeled data of the sample image and the loss function; Passing the loss value along each layer of the image tracking network to update the parameters of the image tracking network; Until the obtained loss value is less than the second preset threshold, the image tracking model is obtained.

8. An image tracking device, characterized in that: The device comprises: an acquisition unit, configured to acquire position information of a trigger operation in response to a trigger operation on a target image; A generating unit, used for processing the target image according to the position information to generate image heat source information; A determination unit is used to input the target image and the image heat source information into an image tracking model to determine the image contour information corresponding to the position information.

9. An image tracking system, characterized in that: It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the image tracking method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the image tracking method according to any one of claims 1 to 7.